建議 model:opus/effort:medium — 歸檔演算法是整案的價值交付點,冪等與部分命中的處理錯了會產生髒資料。

本卡屬 FR-107(母卡 CM-1843)第 .4 棒(子需求卡 CM-1847),子任務 T-4.2:做出「按一下把整批分好的檔搬進對應任務」與「清掉沒用到的剩餘檔」兩支 API。

這張卡做什麼(白話)

這張卡是整案的價值交付點——前作 FR-030 分類完只是把檔複製到 Google Drive 另一個資料夾,系統的任務證據一筆都沒寫,稽核員什麼都看不到。這張卡讓分類結果真的變成系統裡的證據。

歸檔:使用者在審閱頁確認完,按「歸檔進任務」,系統把每一份檔對到的每一個檢查點,各寫一筆任務證據。一份檔對到三個檢查點就是三筆證據紀錄,但指向同一個實體檔(不複製)。沒對到任何任務的檔留在批次裡標「無對應任務」。

清理:歸檔完系統問「剩下這 Z 份檔要清掉嗎」,選清掉就實刪實體檔,但批次檔表留一筆 removed 紀錄(檔名、雜湊、AI 判定結果都留著,之後還能回答「這批上傳過哪些檔、AI 怎麼判」)。已歸檔的檔永遠不刪。

動哪些檔

套件(走 path dependency 開發)
~/Projects/Jedicogy/module/jedi-python-package/jedi-evidence-classification/
  .../app/service/evidence_batch_service.py
      ← 加 archive(batch_uid, user)   §6.7 演算法,🔴 整段在單一 @transaction 內
      ← 加 purge(batch_uid, user)     D7
      ← ALLOWED_TRANSITIONS 補 review→archived、archived→archived
  .../api/routing.py
      ← POST /evidence-batches/<batch_uid>/archive
      ← POST /evidence-batches/<batch_uid>/purge
  統計回寫:evidence_batches 的 archived_count/no_task_count

歸檔演算法(design.md §6.7,逐字照做)

輸入:batch(status=review)、run=batch.current_run、actor

1. jobs = sink.resolve_jobs(batch.round_id)         # {part_id: job_execution_id}

2. for bf in batch_files where status in (pending, classified):
     parts = run.state[bf.file_uid].parts           # 審閱後最終判定(含人工加減;
                                                    #  標「不適用」的已排除)
     if not parts: continue                         # 沒判定的檔留 pending

     hit = [p for p in parts if p in jobs]
     if not hit:
         bf.status = no_task; bf.classification = parts; continue

     for p in hit:                                  # 一檔多 AO → 多筆
         r = sink.attach(job_execution_id=jobs[p], file_id=bf.file_id,
                         file_uid=bf.file_uid,
                         classification_run_id=run.id, actor_user_id=actor,
                         content_hash=bf.content_hash, description=f"AI 分類:{p}")
         created += (not r.existed); existed += r.existed
         bf.archived_evidence_uids.append(r.evidence_uid)

     bf.status = archived; bf.classification = parts
     miss = [p for p in parts if p not in jobs]     # 部分沒任務:檔仍算 archived,
                                                    # 缺的記在 classification 內 no_task=true

3. batch.status = archived;統計回寫(archived_count/no_task_count);run.archived_at

4. 回 {attached_files, evidence_created, evidence_existed, no_task_files}

🔴 冪等:整段在一個 @transaction 內;重按第二次時每個 attach 都回 existed=True、created=0、狀態不變。

🔴 不複製實體檔——job_evidences.file_id 指向同一筆 upload_files。

清理(purge,D7)

只刪 status IN (pending, classified, no_task) 的檔
🔴 archived 的列永不刪

每一筆:
  ① IEvidenceStorage.delete(tenant_id, file_uid)   實刪實體檔
  ② evidence_batch_files 那列:
       status = 'removed'
       classification = 判定快照(🔴 保留 AI 判定結果,D7 明文要求)
       removed_at / removed_by_user_id 填上
回:刪了幾檔

怎麼做(逐步)