▐▛███▜▌ Claude Code v2.1.224
▝▜█████▛▘ Opus 5 · Claude Pro
▘▘ ▝▝ ~/reTraining
Tackle your toughest work with Opus 5. Switch anytime with /model.
❯ /clear
❯ please create a deep understanding using workflow to understand this project
repo and bugs, give me recomendation of improving the code especially for
make this project funtion as i want. ask me a lot before you do that lot of
task, swarm cheaper agent model for helping you, you can use opus also but
its only for complex or task, i guess you can defined yourself. the goal is
to make a better workflow for reTraining my ai model, if i had a mistake just
say it, your goal is to help me make this app works well and ready to use.
Searched for 1 pattern, listed 1 directory (ctrl+o to expand)
● Repo is substantial (backend ~3.5k lines, 8 frontend pages, docs with REQ-xxx
already agreed). Before I spend a workflow on it, four things:
● User answered Claude's questions:
⎿ · What's actually broken or unsatisfying right now? Where should the audit
bite hardest? → Retraining loop itself, Auto-annotate / SAM3, Review editor
UX, Everything / I don't know yet
· Have you run a full round end-to-end yet (video → frames → annotate →
review → merge → train → compare)? → Finished once, but rough
· There's a lot of uncommitted work (24 modified files, 5 new untracked).
What's its status? → Half-broken experiments
· What do you want me to deliver at the end of the audit? → Report, then
fix the critical bugs
● Good — full sweep it is. Launching the workflow now.
● Running in the background. Structure:
- Map (4 parallel) — retraining core, SAM3/autolabel, review editor,
docs-drift + your dirty working tree
- Hunt (5 dimensions, each finding verified by an adversarial refuter as it
lands) — invariants/data corruption, concurrency & jobs, API contract,
frontend correctness, ops/portability
- Synthesize — mechanism-level system map, ranked bugs, per-file verdict on
the uncommitted work, workflow gaps where the design itself is wrong, and an
ordered fix plan with verification steps
Watch live with /workflows. When it lands I'll relay the report and then start
on the critical fixes.
✻ Brewed for 1m 26s
● Dynamic workflow "Deep audit of the reTraining dataset-enrichment app: map
architecture, find bugs, verify, prioritize fixes" completed · 7m 59s
⎿ You've hit your session limit · resets 12:10pm (Asia/Jakarta)
/upgrade to increase your usage limit.
✻ Crunched for 10m 22s
❯ continue
● The audit ran but hit the session usage limit near the end — the final
synthesis agent and 26 verifiers died. 22 bugs did survive verification. Let
me pull the full results and write the report myself.
Ran 3 shell commands (ctrl+o to expand)
● Two dimensions (concurrency, api-contract) lost all their verifiers to the
session limit — let me recover their raw findings from the journal so nothing
is silently dropped.
Ran 6 shell commands (ctrl+o to expand)
● Recovered everything. One key thing to confirm before I write it up — the two
CRITICALs point at jobs.py, which is in your dirty tree:
Ran 1 shell command (ctrl+o to expand)
● Confirmed — that uncommitted jobs.py edit is the root of both CRITICALs.
Writing the report.
● Write(docs/audit-2026-08-07.md)
⎿ Wrote 181 lines to docs/audit-2026-08-07.md
1 # Audit — reTraining, 2026-08-07
2
by dimension, adversarial
4 verification of each finding. 22 findings survived verification;
2 CRITICALs come from a
e marked *unverified* — both
6 were confirmed by hand afterwards.
7
9
10 ## 1. How the system actually works
… +343 lines (ctrl+o to expand)
● Report written to docs/audit-2026-08-07.md. 22 verified findings + 2
hand-confirmed criticals.
The headline: your biggest problem is uncommitted. You changed backend/jobs.py
to run every job in its own thread instead of through the single worker
queue, and narrowed GPU_JOB_TYPES to ("train",) with the comment "multiple
autolabel jobs can run in parallel." They can't — each SAM3 job holds a full
backbone in VRAM, and release_engine() only clears a module global, so it
cannot free memory held by a running job. That one edit also un-serialized
merge and train against the same dataset/ directory, which is the root of four
more HIGH findings.
Three design-level things you got wrong (you asked me to say so):
1. The val split is positional — _next_split counts dataset_items rows, so
every Nth insert goes to val. Deleting a batch shifts the split for everything
merged afterwards. CLAUDE.md §8 declares the invariant but nothing enforces
it. Make it hash(project + batch + stem) % 5 and it becomes true by
construction.
2. Merged labels are regenerated from live annotations on every training start
(sync_labels at dataset.py:117, no review_status filter). Re-run
auto-annotate on a merged batch, press Train, and raw unreviewed model output
overwrites your human-reviewed labels. A merged dataset should be a snapshot,
not a view.
3. The base-vs-new comparison degrades silently instead of refusing. Empty val
set → it validates on the training images. Deleted batch → orphaned JPEGs
still sit in images/val. For a tool whose only job is answering "did
retraining help?", every one of those returns a plausible-looking number.
Worth knowing: docker-compose.override.yml is committed with
nvidia.com/gpu=all (CPU hosts can't start), and start.sh deletes that tracked
file on every run — silently dropping your ./backend bind mount.
Before I start fixing:
● User answered Claude's questions:
⎿ · Fixing C1/C2 means reverting your jobs.py parallelism change. Was that
experiment solving a real problem I should address differently? → Revert,
but autolabel was too slow
· How far do I go in this pass? → Criticals + design + review editor
· The content-derived val split changes how existing merged frames are
assigned. Existing data? → Preserve existing assignments
● Understood: revert for correctness, then look at autolabel speed separately.
Steps 1–7, preserving existing splits.
Read 2 files (ctrl+o to expand)
● Now fixing. Step 1 — restore the job queue:
● Update(backend/jobs.py)
⎿ Added 3 lines, removed 5 lines
28 PROGRESS_FLUSH_SECONDS = 0.5
29
30 JOB_TYPES = ("extract", "autolabel", "merge", "train")
31 -GPU_JOB_TYPES = ("train",)
32 -"""`train` takes full GPU memory during model fitting. `autolab
-el` (SAM3/YOLO inference)
33 -runs fine-grained per-frame steps so multiple autolabel jobs ca
-n run in parallel."""
34 -
35 -
31 +GPU_JOB_TYPES = ("autolabel", "train")
32 +"""`extract` is ffmpeg and `merge` is file copying — neither to
+uches the card,
33 +so neither should be able to block an interactive assist."""
34 gpu_lock = threading.Lock()
35 """Held for the duration of any GPU work. The job worker takes
it around a
36 handler; the interactive assist route takes it around one SAM3
call. One card,
● Update(backend/jobs.py)
⎿ Added 2 lines, removed 12 lines
147 job_id = cur.lastrowid
148 job = get(job_id)
149 assert job is not None
150 - threading.Thread(target=_start_job, args=(job_id,), name=f
-"job-{job_type}-{job_id}", daemon=True).start()
150 + _queue.put(job_id)
151 + _ensure_worker()
152 return job
153
154
154 -def _start_job(job_id: int) -> None:
155 - job = get(job_id)
156 - if job is None:
157 - return
158 - if job.id in _cancelled:
159 - _finish(job, "cancelled")
160 - return
161 - _run(job)
162 -
163 -
164 -
155 def get(job_id: int) -> Optional[Job]:
156 with db.cursor() as cur:
157 cur.execute("SELECT * FROM jobs WHERE id = ?", (job_id
,))
● Update(backend/jobs.py)
⎿ Added 20 lines, removed 1 line
206 return cur.rowcount
207
208
209 -def _run(job: Job) -> None:
209 +def _ensure_worker() -> None:
210 + global _worker
211 + with _worker_lock:
212 + if _worker is None or not _worker.is_alive():
213 + _worker = threading.Thread(target=_worker_loop, na
+me="job-worker", daemon=True)
214 + _worker.start()
215
216 +
217 +def _worker_loop() -> None:
218 + while True:
219 + job_id = _queue.get()
220 + job = get(job_id)
221 + if job is None:
222 + continue
223 + if job.id in _cancelled:
224 + _finish(job, "cancelled")
225 + continue
226 + _run(job)
227 +
228 +
229 +def _run(job: Job) -> None:
230 job.status = "running"
231 job.started_at = time.time()
232 job.flush()
Read 1 file, ran 1 shell command (ctrl+o to expand)
● jobs.py is back to baseline. Step 2 — the synchronous GPU routes. Let me match
the existing review.assist pattern:
Ran 1 shell command (ctrl+o to expand)
● Update(backend/api/batches.py)
⎿ Added 8 lines, removed 1 line
162
163 @router.post("/api/batches/{batch_id}/preview")
164 def preview_autolabel(batch_id: int, request: PreviewRequest)
-> dict:
165 - from backend import autolabel
165 + from backend import autolabel, jobs
166 +
167 + if not jobs.gpu_lock.acquire(timeout=20):
168 + busy = jobs.running_types()
169 + kind = busy[0] if busy else "background"
170 + raise HTTPException(409, f"The GPU is busy with a {kin
+d} job — wait for it to finish")
171 try:
172 shapes = autolabel.preview_frame(
173 batch_id=batch_id,
...
182 return {"shapes": shapes}
183 except Exception as exc:
184 raise HTTPException(400, str(exc))
185 + finally:
186 + jobs.gpu_lock.release()
187
188 @router.post("/api/sam3/playground-test")
189 async def sam3_playground_test(
● Update(backend/api/batches.py)
⎿ Added 8 lines
207 if not prompt_list:
208 raise HTTPException(400, "At least one text prompt is
required")
209
210 + from backend import jobs
211 +
212 + if not jobs.gpu_lock.acquire(timeout=20):
213 + busy = jobs.running_types()
214 + kind = busy[0] if busy else "background"
215 + raise HTTPException(409, f"The GPU is busy with a {kin
+d} job — wait for it to finish")
216 try:
217 engine = get_engine()
218 raw_dets = engine.detect(image, prompt_list, threshold
)
219 kept_dets = labeling.deduplicate(raw_dets, iou_thresho
ld=iou_threshold)
220 except Exception as exc:
221 raise HTTPException(500, f"SAM3 inference failed: {exc
}")
222 + finally:
223 + jobs.gpu_lock.release()
224
225 results = []
226 for det in kept_dets:
● Step 3, 4 and 6 all live in dataset.py. Rewriting that section:
● Update(backend/dataset.py)
⎿ Added 72 lines, removed 23 lines
66 return f"{class_id} {coords}"
67
68
69 -def _next_split(cur, project_id: int, val_every: int) -> str:
70 - """Continue the every-Nth pattern from wherever the last m
-erge left off."""
69 +def split_for(project_id: int, batch_id: int, stem: str, val_e
+very: int) -> str:
70 + """Which split a frame belongs to, derived from its identi
+ty rather than from
71 + how many rows happen to precede it.
72 +
73 + A positional every-Nth rule makes membership depend on ins
+ertion history, so
74 + deleting or re-merging a batch silently reshuffles every l
+ater frame — and a
75 + frame that was in `val` for the last comparison could land
+ in `train` for the
76 + next one. Hashing the identity makes the stable-val-split
+invariant true by
77 + construction: the same frame always lands in the same spli
+t, whatever else
78 + happened to the dataset. Rows already in `dataset_items` k
+eep the split they
79 + were recorded with; nothing recomputes them.
80 + """
81 if val_every <= 0:
82 return "train"
73 - cur.execute("SELECT COUNT(*) FROM dataset_items WHERE proj
-ect_id = ?", (project_id,))
74 - position = cur.fetchone()[0]
75 - return "val" if position % val_every == val_every - 1 else
- "train"
83 + digest = hashlib.sha1(f"{project_id}/{batch_id}/{stem}".en
+code("utf-8")).hexdigest()
84 + return "val" if int(digest[:8], 16) % val_every == 0 else
+"train"
85
86
78 -def sync_labels(project_id: int, selected_class_ids: Optional[
-List[int]] = None) -> dict:
79 - """Re-sync label files on disk for all merged frames in th
-e project dataset."""
87 +def sync_labels(project_id: int) -> dict:
88 + """Re-sync label files on disk from the reviewed annotatio
+ns.
89 +
90 + Only frames a human actually signed off on are written. A
+merged frame whose
91 + batch was auto-annotated again drops back to `pending`, an
+d rewriting its
92 + label file from the fresh model output would push predicti
+ons nobody checked
93 + into the master dataset.
94 + """
95 project = projects.get(project_id)
96 root = dataset_dir(project["slug"])
97 with db.cursor() as cur:
98 cur.execute(
84 - "SELECT d.frame_id, d.label_rel FROM dataset_items
- d WHERE d.project_id = ?",
99 + """SELECT d.frame_id, d.label_rel FROM dataset_ite
+ms d
100 + JOIN frames f ON f.id = d.frame_id
101 + WHERE d.project_id = ? AND f.review_status = 'a
+pproved'""",
102 (project_id,),
103 )
104 items = cur.fetchall()
105
89 - class_map = None
90 - if selected_class_ids is not None and len(selected_class_i
-ds) > 0:
91 - class_map = {cid: idx for idx, cid in enumerate(sorted
-(selected_class_ids))}
92 -
106 synced_files = 0
107 total_lines = 0
108 for frame_id, label_rel in items:
96 - annotations = review.listing(frame_id)
97 - if class_map is not None:
98 - annotations = [a for a in annotations if a["class_
-id"] in class_map]
99 -
100 - lines = []
101 - for item in annotations:
102 - mapped_cid = class_map[item["class_id"]] if class_
-map is not None else item["class_id"]
103 - lines.append(_label_line(mapped_cid, item["geometr
-y"], project["label_type"]))
104 -
109 + lines = [_label_line(item["class_id"], item["geometry"
+], project["label_type"])
110 + for item in review.listing(frame_id)]
111 path = os.path.join(root, label_rel)
112 os.makedirs(os.path.dirname(path), exist_ok=True)
107 - with open(path, "w", encoding="utf-8") as f:
108 - f.write("\n".join(lines) + ("\n" if lines else "")
-)
113 + _write_atomic(path, "\n".join(lines) + ("\n" if lines
+else ""))
114 synced_files += 1
115 total_lines += len(lines)
116
117 return {"synced_files": synced_files, "total_lines": total
_lines}
118
119
120 +def _write_atomic(path: str, text: str) -> None:
121 + """Write via temp file + rename, so a training run never r
+eads a half-written
122 + label file or a truncated data.yaml."""
123 + tmp = f"{path}.tmp"
124 + with open(tmp, "w", encoding="utf-8") as handle:
125 + handle.write(text)
126 + os.replace(tmp, path)
127 +
128 +
129 +def _build_selected_tree(root: str, rows: list, class_map: Opt
+ional[dict],
130 + label_type: str) -> tuple:
131 + """Materialise a filtered view of the dataset under `selec
+ted/`.
132 +
133 + The master labels stay canonical. A per-run class filter r
+enumbers ids into a
134 + contiguous 0..k-1 space, which contradicts `project_classe
+s`; writing that
135 + into the shared label files would leave the master dataset
+ in the index space
136 + of whichever training run happened to go last.
137 + """
138 + selected_root = os.path.join(root, "selected")
139 + if os.path.isdir(selected_root):
140 + shutil.rmtree(selected_root)
141 + listed = {"train": [], "val": []}
142 + for image_rel, label_rel, frame_id, split in rows:
143 + stem = os.path.basename(image_rel)
144 + image_dst = os.path.join(selected_root, "images", spli
+t, stem)
145 + label_dst = os.path.join(selected_root, "labels", spli
+t,
146 + os.path.splitext(stem)[0] + "
+.txt")
147 + os.makedirs(os.path.dirname(image_dst), exist_ok=True)
148 + os.makedirs(os.path.dirname(label_dst), exist_ok=True)
149 + source = os.path.join(root, image_rel)
150 + if not os.path.exists(image_dst):
151 + os.symlink(source, image_dst)
152 +
153 + annotations = review.listing(frame_id)
154 + if class_map is not None:
155 + annotations = [a for a in annotations if a["class_
+id"] in class_map]
156 + lines = [_label_line(class_map[a["class_id"]] if class
+_map is not None else a["class_id"],
157 + a["geometry"], label_type)
158 + for a in annotations]
159 + _write_atomic(label_dst, "\n".join(lines) + ("\n" if l
+ines else ""))
160 + listed[split].append(image_dst)
161 + return selected_root, listed
162 +
163 +
164 def write_data_yaml(project: dict, batch_ids: list = None, sel
ected_class_ids: Optional[List[int]] = None) -> str:
165 """Rebuild data.yaml from the project's classes (REQ-051).
"""
166 sync_labels(project["id"], selected_class_ids=selected_cla
ss_ids)
● Now write_data_yaml — fail-loud on empty val, correct nc, use the selected
tree:
● Update(backend/dataset.py)
⎿ Added 46 lines, removed 24 lines
163
164 def write_data_yaml(project: dict, batch_ids: list = None, sel
ected_class_ids: Optional[List[int]] = None) -> str:
165 """Rebuild data.yaml from the project's classes (REQ-051).
"""
166 - sync_labels(project["id"], selected_class_ids=selected_cla
-ss_ids)
166 + sync_labels(project["id"])
167 root = dataset_dir(project["slug"])
168 os.makedirs(root, exist_ok=True)
169 counts = summary(project["id"])["splits"]
170
171 target_classes = project["classes"]
172 + class_map = None
173 if selected_class_ids is not None and len(selected_class_i
ds) > 0:
174 target_classes = [c for c in project["classes"] if c["
class_id"] in selected_class_ids]
175 + class_map = {cid: idx for idx, cid in enumerate(sorted
+(selected_class_ids))}
176
177 names = ", ".join(f"'{item['name']}'" for item in target_c
lasses)
178
177 - if batch_ids:
179 + if batch_ids or class_map is not None:
180 with db.cursor() as cur:
179 - placeholders = ",".join("?" for _ in batch_ids)
181 + where = "d.project_id = ?"
182 + args = [project["id"]]
183 + if batch_ids:
184 + where += f" AND f.batch_id IN ({','.join('?' f
+or _ in batch_ids)})"
185 + args += list(batch_ids)
186 cur.execute(
181 - f"""SELECT d.image_rel, d.split FROM dataset_i
-tems d
187 + f"""SELECT d.image_rel, d.label_rel, d.frame_i
+d, d.split FROM dataset_items d
188 JOIN frames f ON f.id = d.frame_id
183 - WHERE d.project_id = ? AND f.batch_id IN (
-{placeholders})""",
184 - [project["id"]] + list(batch_ids),
189 + WHERE {where}""",
190 + args,
191 )
192 rows = cur.fetchall()
193
188 - train_files = [row[0] for row in rows if row[1] == "tr
-ain"]
189 - val_files = [row[0] for row in rows if row[1] == "val"
-] or train_files
194 + selected_root, listed = _build_selected_tree(root, row
+s, class_map, project["label_type"])
195 + _require_val(len(listed["val"]), "the selected batches
+")
196
197 train_txt = os.path.join(root, "selected_train.txt")
198 val_txt = os.path.join(root, "selected_val.txt")
193 - with open(train_txt, "w", encoding="utf-8") as handle:
194 - handle.write("\n".join(os.path.join(root, rel) for
- rel in train_files) + "\n")
195 - with open(val_txt, "w", encoding="utf-8") as handle:
196 - handle.write("\n".join(os.path.join(root, rel) for
- rel in val_files) + "\n")
199 + _write_atomic(train_txt, "\n".join(listed["train"]) +
+"\n")
200 + _write_atomic(val_txt, "\n".join(listed["val"]) + "\n"
+)
201
202 path = os.path.join(root, "selected_data.yaml")
199 - with open(path, "w", encoding="utf-8") as handle:
200 - handle.write(f"path: {root}\n")
201 - handle.write(f"train: {train_txt}\n")
202 - handle.write(f"val: {val_txt}\n\n")
203 - handle.write(f"nc: {len(project['classes'])}\n")
204 - handle.write(f"names: [{names}]\n")
203 + _write_atomic(path,
204 + f"path: {selected_root}\n"
205 + f"train: {train_txt}\n"
206 + f"val: {val_txt}\n\n"
207 + f"nc: {len(target_classes)}\n"
208 + f"names: [{names}]\n")
209 return path
210
211 + _require_val(counts["val"], "this project")
212 +
213 path = os.path.join(root, "data.yaml")
208 - with open(path, "w", encoding="utf-8") as handle:
209 - handle.write(f"path: {root}\n")
210 - handle.write("train: images/train\n")
211 - handle.write(f"val: images/{'val' if counts['val'] > 0
- else 'train'}\n\n")
212 - handle.write(f"nc: {len(project['classes'])}\n")
213 - handle.write(f"names: [{names}]\n")
214 + _write_atomic(path,
215 + f"path: {root}\n"
216 + "train: images/train\n"
217 + "val: images/val\n\n"
218 + f"nc: {len(target_classes)}\n"
219 + f"names: [{names}]\n")
220 return path
221
222
223 +def _require_val(count: int, subject: str) -> None:
224 + """Refuse to build a dataset with an empty val split.
225 +
226 + Falling back to the training images produces a base-vs-new
+ mAP measured on
227 + data the model was fitted to — a number that looks fine an
+d means nothing.
228 + For a system whose whole purpose is answering "did retrain
+ing help?", this
229 + has to fail loudly.
230 + """
231 + if count == 0:
232 + raise DatasetError(
233 + f"There are no validation images in {subject}, so
+a base-vs-new comparison "
234 + "would be measured on the training images. Merge m
+ore frames, or lower the "
235 + "project's val_every."
236 + )
237 +
238 +
239 def summary(project_id: int) -> dict:
240 import json
241 with db.cursor() as cur:
● Now the merge handler — hash split, cancel bug, and the hashlib import:
● Update(backend/dataset.py)
⎿ Added 1 line
11 segment class_id x1 y1 x2 y2 … (normalized polygon
)
12 """
13
14 +import hashlib
15 import os
16 import shutil
17 import time
● Update(backend/dataset.py)
⎿ Added 3 lines, removed 1 line
436
437 added = {"train": 0, "val": 0}
438 skipped = 0
439 + cancelled = False
440 for index, frame in enumerate(frames):
441 if job.cancelled:
442 job.log(f"Cancelled after {index} frame(s)")
443 + cancelled = True
444 break
445
446 with db.cursor() as cur:
...
449 skipped += 1
450 job.progress(index + 1, len(frames))
451 continue
450 - split = _next_split(cur, project["id"], project["v
-al_every"])
452
453 stem = f"{batch['id']}__{os.path.splitext(frame['f
ilename'])[0]}"
454 + split = split_for(project["id"], batch["id"], stem
+, project["val_every"])
455 image_rel = f"images/{split}/{stem}.jpg"
456 label_rel = f"labels/{split}/{stem}.txt"
457 shutil.copyfile(os.path.join(source_dir, frame["fi
lename"]),
● Update(backend/dataset.py)
⎿ Added 8 lines
473 added[split] += 1
474 job.progress(index + 1, len(frames))
475
476 + if cancelled:
477 + # Leaving it 'merged' would be a lie: the frames after
+ the break point
478 + # have no dataset_items rows and no files, and approve
+() refuses to
479 + # re-merge a merged batch, so they could never be adde
+d. The per-frame
480 + # dataset_items guard already makes re-running the mer
+ge idempotent.
481 + job.log("Batch left approved — re-approve it to finish
+ the merge")
482 + return
483 +
484 with db.cursor() as cur:
485 cur.execute("UPDATE batches SET status = 'merged', mer
ged_at = ? WHERE id = ?",
486 (time.time(), batch["id"]))
● Step 5 — batch deletion orphans, plus the double-approve guard:
Read 1 file (ctrl+o to expand)
● Update(backend/batches.py)
⎿ Added 23 lines
238
239 def delete(batch_id: int) -> bool:
240 import shutil
241 + from backend import dataset
242 +
243 batch = get(batch_id)
244 if batch is None:
245 return False
246 +
247 + # The FK cascade takes the frames and dataset_items rows,
+but the copies this
248 + # batch put in the master dataset live in a sibling tree a
+nd would survive as
249 + # orphans — and data.yaml trains and validates on those di
+rectories, so the
250 + # model would keep learning from a batch the user believes
+ is gone.
251 + root = dataset.dataset_dir(batch["project_slug"])
252 with db.cursor() as cur:
253 + cur.execute(
254 + """SELECT d.image_rel, d.label_rel FROM dataset_it
+ems d
255 + JOIN frames f ON f.id = d.frame_id
256 + WHERE f.batch_id = ?""",
257 + (batch_id,),
258 + )
259 + merged_files = cur.fetchall()
260 + for image_rel, label_rel in merged_files:
261 + for rel in (image_rel, label_rel):
262 + try:
263 + os.unlink(os.path.join(root, rel))
264 + except FileNotFoundError:
265 + pass
266 +
267 + with db.cursor() as cur:
268 cur.execute("DELETE FROM batches WHERE id = ?", (batch
_id,))
269 shutil.rmtree(batch_dir(batch["project_slug"], batch_id),
ignore_errors=True)
270 return True
● Update(backend/dataset.py) ⎿ Added 8 lines
42 )
43 if batch["review"]["approved"] == 0:
44 raise DatasetError("Every frame was rejected — there is
nothing to merge")
45 + with db.cursor() as cur:
46 + cur.execute(
47 + """SELECT 1 FROM jobs WHERE batch_id = ? AND type =
+ 'merge'
48 + AND status IN ('queued', 'running')""",
49 + (batch_id,),
50 + )
51 + if cur.fetchone() is not None:
52 + raise DatasetError("A merge for this batch is alrea
+dy queued")
53
54 batches.set_status(batch_id, "approved")
55 job = jobs.create(
Searched for 2 patterns (ctrl+o to expand)
● Good catch — class edits and merges would now throw on a young dataset. Gating
the check to the training path only:
● Update(backend/dataset.py)
⎿ Added 3 lines, removed 1 line
170 return selected_root, listed
171
172
173 -def write_data_yaml(project: dict, batch_ids: list = None, sel
-ected_class_ids: Optional[List[int]] = None) -> str:
173 +def write_data_yaml(project: dict, batch_ids: list = None,
174 + selected_class_ids: Optional[List[int]] =
+None,
175 + require_val: bool = False) -> str:
176 """Rebuild data.yaml from the project's classes (REQ-051).
"""
177 sync_labels(project["id"])
178 root = dataset_dir(project["slug"])
● Update(backend/dataset.py)
⎿ Added 2 lines, removed 1 line
203 rows = cur.fetchall()
204
205 selected_root, listed = _build_selected_tree(root, row
s, class_map, project["label_type"])
206 - _require_val(len(listed["val"]), "the selected batches
-")
206 + if require_val:
207 + _require_val(len(listed["val"]), "the selected bat
+ches")
208
209 train_txt = os.path.join(root, "selected_train.txt")
210 val_txt = os.path.join(root, "selected_val.txt")
● Update(backend/dataset.py)
⎿ Added 2 lines, removed 1 line
220 f"names: [{names}]\n")
221 return path
222
223 - _require_val(counts["val"], "this project")
223 + if require_val:
224 + _require_val(counts["val"], "this project")
225
226 path = os.path.join(root, "data.yaml")
227 _write_atomic(path,
● Update(backend/training.py)
⎿ Added 2 lines, removed 1 line
103 settings = job.params["settings"]
104 batch_ids = job.params.get("batch_ids")
105 class_ids = job.params.get("class_ids")
106 - data_yaml = dataset.write_data_yaml(project, batch_ids=bat
-ch_ids, selected_class_ids=class_ids)
106 + data_yaml = dataset.write_data_yaml(project, batch_ids=bat
+ch_ids, selected_class_ids=class_ids,
107 + require_val=True)
108
109 # SAM3 and a training run must not hold VRAM at the same t
ime (REQ-065).
110 from backend.sam3_engine import release_engine
● Backend done. Step 7 — the review editor:
Read 1 file (ctrl+o to expand)
● Update(frontend/src/pages/ReviewPage.jsx)
⎿ Added 3 lines
234 useEffect(() => {
235 function onKeyDown(event) {
236 if (event.target?.matches?.('input, textarea, select, [c
ontenteditable="true"]')) return
237 + // Ctrl/Cmd/Alt combos belong to the browser and the OS:
+ without this,
238 + // Ctrl+S approves the frame and Ctrl+A/C/X/N/T all fire
+ review actions.
239 + if (event.ctrlKey || event.metaKey || event.altKey) retu
+rn
240 const { frames, project, setStatus, removeSelected, recl
ass, jumpToPending, jumpToNextAnnotated, setAssistMode, copyPr
evious, trackForward } = stateRef.current
241 const key = event.key
242 const isShortcutKey = /^[1-9]$/.test(key) || ['ArrowLeft
', 'ArrowRight', 'ArrowUp', 'ArrowDown', 'Delete', 'Backspace'
, 'a', 'A', 'x', 'X', 'u', 'U', 's', 'S', 'n', 'N', 'c', 'C',
't', 'T'].includes(key)
● Update(frontend/src/pages/ReviewPage.jsx)
⎿ Added 28 lines, removed 9 lines
124 return
125 }
126 if (!commit) return
127 - const current = annotations.find((row) => row.id === id)
128 - if (!current) return
129 - try { await api.patchAnnotation(id, { geometry: current.ge
-ometry }) } catch (exc) { setError(exc.message) }
127 + const previous = annotations.find((row) => row.id === id)
128 + if (!previous) return
129 + // A commit may carry its own geometry (delete-vertex send
+s the shortened
130 + // polygon); falling back to the row's geometry covers dra
+g/resize, which
131 + // already wrote through the {local:true} path.
132 + const next = geometry ?? previous.geometry
133 + setAnnotations((rows) => rows.map((row) => (row.id === id
+? { ...row, geometry: next } : row)))
134 + try {
135 + await api.patchAnnotation(id, { geometry: next })
136 + } catch (exc) {
137 + setAnnotations((rows) => rows.map((row) => (row.id === i
+d ? previous : row)))
138 + setError(exc.message)
139 + }
140 }
141
132 - const removeSelected = useCallback(async () => {
133 - if (selectedId == null || !frame) return
134 - const id = selectedId
142 + const removeAnnotation = useCallback(async (id) => {
143 + if (id == null || !frame) return
144 + const previous = annotations.find((row) => row.id === id)
145 + if (!previous) return
146 setAnnotations((rows) => rows.filter((row) => row.id !== i
d))
136 - setSelectedId(null)
147 + setSelectedId((current) => (current === id ? null : curren
+t))
148 setFrames((rows) => rows.map((row) => (row.id === frame.id
? { ...row, annotation_count: Math.max(0, (row.annotation_cou
nt ?? 1) - 1) } : row)))
138 - try { await api.deleteAnnotation(id) } catch (exc) { setEr
-ror(exc.message) }
139 - }, [selectedId, frame])
149 + try {
150 + await api.deleteAnnotation(id)
151 + } catch (exc) {
152 + setAnnotations((rows) => [...rows, previous])
153 + setFrames((rows) => rows.map((row) => (row.id === frame.
+id ? { ...row, annotation_count: (row.annotation_count ?? 0) +
+ 1 } : row)))
154 + setError(exc.message)
155 + }
156 + }, [annotations, frame])
157
158 + const removeSelected = useCallback(() => removeAnnotation(se
+lectedId), [removeAnnotation, selectedId])
159 +
160 const reclass = useCallback(async (classId) => {
161 setActiveClass(classId)
162 if (selectedId == null) return
● Now wire the sidebar to the id-explicit delete:
Ran 1 shell command (ctrl+o to expand)
● Update(frontend/src/components/ReviewSidebar.jsx)
⎿ Added 1 line, removed 4 lines
83
Ran 1 shell command (ctrl+o to expand)
● Update(frontend/src/components/ReviewSidebar.jsx)
⎿ Added 1 line, removed 1 line
11 annotations,
12 selectedId,
13 setSelectedId,
14 - removeSelected,
14 + removeAnnotation,
15 project,
16 jumpToNextAnnotated,
17 batchAnnotationCount,
● Update(frontend/src/pages/ReviewPage.jsx)
⎿ Added 1 line, removed 1 line
430 annotations={annotations}
431 selectedId={selectedId}
432 setSelectedId={setSelectedId}
433 - removeSelected={removeSelected}
433 + removeAnnotation={removeAnnotation}
434 project={project}
435 jumpToNextAnnotated={jumpToNextAnnotated}
436 batchAnnotationCount={batch?.annotation_count ?? 0}
● Now verify — imports, syntax, and the split logic:
Ran 3 shell commands (ctrl+o to expand)
● Done with steps 1–7. Both check out: backend imports clean, frontend builds.
What changed
Backend
- jobs.py — reverted to the single-worker queue; GPU_JOB_TYPES back to
("autolabel", "train"). Dead _start_job removed. The file is now
byte-identical to HEAD.
- api/batches.py — preview and /api/sam3/playground-test now take gpu_lock
with a 20 s timeout and return 409 with "GPU is busy with a X job", matching
review.assist.
- dataset.py:
- sync_labels filters to review_status='approved' and no longer takes
selected_class_ids — re-running auto-annotate on a merged batch can no longer
push unreviewed output into the master dataset.
- Class-filtered runs build a selected/ tree (symlinked images + remapped
labels). Master labels stay canonical.
- _next_split → split_for(project, batch, stem, val_every), hash-derived.
Verified deterministic, 18.8% val at val_every=5, val_every=0 → all train.
- All label/yaml writes go through _write_atomic (temp + os.replace).
- Cancelled merge leaves the batch approved and logs why, instead of marking
it merged and stranding the frames forever.
- nc now counts target_classes, not the full list.
- _require_val refuses an empty val split — gated behind require_val=True,
passed only from training.py.
- batches.py — delete() unlinks the batch's dataset_items files from the
master dataset before dropping the rows.
- dataset.approve() — rejects when a merge for that batch is already queued or
running.
Frontend
- removeSelected split into removeAnnotation(id) + a thin wrapper; sidebar
trash calls it with an explicit id. The stale-closure wrong-shape delete is
gone.
- Keydown handler early-returns on ctrlKey || metaKey || altKey.
- updateShape uses the geometry passed on commit (Alt-click delete-vertex now
actually works) and rolls back local state on a failed PATCH, as does
removeAnnotation.
On your split migration: no backfill was needed. dataset_items stores each
frame's split as a row and the merge loop skips any frame that already has
one, so existing assignments are preserved untouched — the hash only decides
new merges. Your existing base-vs-new comparison stays valid.
Two caveats, stated plainly. I verified imports, the split function's
behaviour, and a clean frontend build — I did not run a live end-to-end round,
so the merge/train paths are unexercised. And data.yaml now always says val:
images/val; if a project has merged data but zero val images, training raises
instead of silently validating on train. That's the intended behaviour, but it
will surface as a new error on any project in that state.
Not touched from the report: the custom_model_path critical (C3 — needs the
opaque-staging-id refactor across inspect-model, autolabel.start and
LibraryPage), the staged .pt leaks, library.py's unbounded ffmpeg threads,
datasetSummary returning every annotation, and the ops items (.dockerignore,
untracking docker-compose.override.yml, .env.example, pinning requirements).
Say the word on which to take next — and separately, tell me how slow
autolabel actually is and on what hardware, and I'll look at making a single
job faster.
✻ Sautéed for 6m 28s
※ recap: Goal is making your reTraining app reliable end-to-end; I audited it and fixed the seven critical issues in the job queue, dataset split, and review editor. Next up is your call on which remaining items to tackle. (disable recaps in /config)
❯ can you understand the data preparation pages and analyze what works and what
is not, what we should do to this pages, for this pages purposes is based on
the sam3 result, i can filter which is good for send to train and which to
ignore, but the ignored one is we can choose to just throw away or store them
as a different class, for example full-sack and half-sack. so the training
section can pick which they want to train based on that data preparation
filtering. do you understand my point?. ask me
● Let me read the page and its backend before I say I understand it.
Read 1 file (ctrl+o to expand)
● Yes, I understand the intent. Let me state it back, then flag one problem,
then ask.
What you want: Data Prep is the triage gate between SAM3's raw output and
training. You look at what SAM3 produced, filter it, and each shape gets one
of three fates — keep (goes to training), discard (thrown away), or relabel
(kept, but as a different class, e.g. the rejects become half-sack while the
keepers stay full-sack). Then Models/Training picks which of those classes it
actually trains on.
What the page does today: nothing of that. It is a read-only analytics
dashboard. minSizePct/maxSizePct are local React state
(DataPrepPage.jsx:10-11) that are never sent to the server — no apply button,
no persistence, no mutation. Reload the page and your filter is gone. The
class-distribution counts and the red/green box preview recompute live, but
they're a visualization of a filter that exists only in the browser.
What already exists that you can build on: training accepts selected_class_ids
(training.py:106 → write_data_yaml), which after this morning's fix
materializes a filtered selected/ tree. So "training picks which classes to
train" is already wired — the missing half is the mechanism that assigns
shapes to those classes in the first place.
One concern, and I think it matters. The only filter axis on this page is box
area as a percentage of the frame. Area is a proxy for distance from camera,
not for how full the sack is. A full sack at the far end of the conveyor and a
half sack up close can have identical area_pct. If full-vs-half is genuinely
the distinction you want, area alone will mislabel a lot of them. SAM3's
confidence score is the better "is this detection any good" signal — and note
it isn't even in the payload today (dataset.py sends id, frame_id, class_id,
area_pct, box — no score). Aspect ratio might separate full from half better
than area does. So I'd like to know what you're actually seeing in the data.
Questions:
● User answered Claude's questions:
⎿ · When you look at a SAM3 result and decide "this one is bad" or "this one
is half-sack", what are you actually judging it on? → SAM3 confidence
score, Box size / area, I have to look at the image, or anything you
recommend
· When you "relabel the rejects as half-sack", should that permanently
change the annotations, or be a reusable rule? → Saved filter rule / preset
just like you recommend. every section have their own rule, so if i
change in data prep section, the batches model doesnt have effect, but the
next step have
· Where should this triage happen in the loop? → Per batch, before merge
· What should "discard" actually do to a shape? → Mark ignored, decide
later
● Good — that clarifies a lot, but three things are still genuinely ambiguous
and they change the architecture, so let me pin them down.
Note one tension in your answers I want to resolve: you want triage per batch
before merge, but you also want a data-prep rule change to affect the next
step. If the rule is baked in at merge time, then editing it later can't
change anything downstream — the data is already written. So I need to know
when the rule actually gets evaluated.
Also, your "mark ignored" answer leaves the label-integrity trap open: if an
ignored box stays on an image that still goes to training, you're teaching the
model that a sack is background. That has to resolve somewhere.
● User answered Claude's questions:
⎿ · When is a data-prep rule actually evaluated? → Live, at training time
(Recommended)
· An 'ignored' box sits on an image. What happens to that image at training
time? → Exclude the whole image (Recommended)
· You said you often have to look at the image. How do manual decisions
coexist with rules? → Rules propose, you confirm in bulk
· For finding the full-vs-half boundary, what should I put on the page? →
Both
● That resolves it into a coherent design. One consequence you should see before
I write anything, because it touches your core invariant:
Live rule evaluation moves the val set. If a rule change excludes images, the
val split changes between runs — so a base-vs-new mAP from last week isn't
comparable to one from today. Both models in a single comparison are fine
(same data.yaml), but across rule edits the number silently shifts meaning.
This is fixable — stamp the rule-set version onto each training run and
refuse/warn on cross-version comparison — but it needs to be in the design,
not discovered later.
Per CLAUDE.md §5 I'm not writing feature code before a REQ covers it, so
here's the proposal:
● Write(docs/proposal-dataprep-triage.md)
⎿ Wrote 131 lines to docs/proposal-dataprep-triage.md
1 # Proposal — Data Prep as a triage stage
2
into `./requirements.md`
4 and the schema/API sections into `./design.md`.
5
7
8 - Signals: SAM3 confidence score, box area, aspect ratio, plus v
isual judgement.
orward only.
10 - Triage happens **per batch**, on SAM3 output.
… +175 lines (ctrl+o to expand)
● Written to docs/proposal-dataprep-triage.md. The short version:
The model is a three-way resolve, evaluated live: manual override > first
matching rule > default keep, producing keep / ignore / reclass → class.
annotations.class_id is never rewritten — SAM3's original output stays intact,
which is exactly what makes rules re-cuttable forever. Resolution happens in
one place only: dataset._build_selected_tree, which after this morning's fix
is already the thing that materializes what training sees.
Nine requirements (REQ-100…108), two new tables (triage_rules,
annotation_overrides), a rule_version column on training runs, six endpoints,
and a /crop endpoint to feed the thumbnail grid.
Three things worth your attention before you approve:
1. Rules are per-project, scoped to the dataprep stage — not per-batch. You
said triage happens per batch, and the UI is per batch (you look at one
batch's SAM3 output at a time). But a rule like "area < 5% → half-sack" is a
judgment about your domain, not about one batch, so it lives on the project
and applies everywhere. If you actually want per-batch rules, say so — it's a
different schema.
2. Score is already stored on annotations (ReviewSidebar renders item.score),
it's just missing from the Data Prep payload. Aspect ratio is free from the
box. So no new inference is needed to get all four signals — just plumbing.
3. REQ-104 vs REQ-107 is the real tension. Excluding whole images is the right
call for label integrity, but it makes your val set a function of your rules.
REQ-107 (stamp the rule version on each run, refuse cross-version comparison)
keeps that honest instead of silent. If it turns out to fire on every edit
and becomes noise, the fallback is to freeze val membership before rule
evaluation and let rules affect train only. I'd rather start strict and loosen
it.
Approve the REQs and I'll move them into requirements.md + design.md and start
building. If you'd rather see it working before committing to the
requirements, I can build the resolver and the scatter/crop UI against one
batch as a spike first — but say which, because CLAUDE.md says docs lead code
and I don't want to violate that silently.
✻ Baked for 2m 21s
❯ id rather see it working before commiting, i approve the plan
● Approved — building the spike. Docs stay as the proposal until you've seen it
run.
Read 1 file, ran 2 shell commands (ctrl+o to expand)
● Schema first:
● Update(backend/db.py)
⎿ Added 22 lines
125 finished_at REAL
126 )
127 """,
128 + """
129 + CREATE TABLE IF NOT EXISTS triage_rules (
130 + id INTEGER PRIMARY KEY AUTOINCREMENT,
131 + project_id INTEGER NOT NULL REFERENCES projects(id)
+ON DELETE CASCADE,
132 + stage TEXT NOT NULL DEFAULT 'dataprep',
133 + position INTEGER NOT NULL,
134 + name TEXT NOT NULL,
135 + predicate TEXT NOT NULL,
136 + action TEXT NOT NULL CHECK (action IN ('keep', '
+ignore', 'reclass')),
137 + target_class INTEGER,
138 + created_at REAL NOT NULL
139 + )
140 + """,
141 + """
142 + CREATE TABLE IF NOT EXISTS annotation_overrides (
143 + annotation_id INTEGER PRIMARY KEY REFERENCES annotatio
+ns(id) ON DELETE CASCADE,
144 + verdict TEXT NOT NULL CHECK (verdict IN ('keep',
+ 'ignore', 'reclass')),
145 + target_class INTEGER,
146 + decided_at REAL NOT NULL
147 + )
148 + """,
149 "CREATE INDEX IF NOT EXISTS idx_frames_batch ON frames(bat
ch_id, idx)",
150 + "CREATE INDEX IF NOT EXISTS idx_triage_rules_project ON tr
+iage_rules(project_id, stage, position)",
151 "CREATE INDEX IF NOT EXISTS idx_annotations_frame ON annot
ations(frame_id)",
152 "CREATE INDEX IF NOT EXISTS idx_batches_project ON batches
(project_id)",
153 "CREATE INDEX IF NOT EXISTS idx_jobs_project ON jobs(proje
ct_id, created_at)",
● Update(backend/db.py)
⎿ Added 5 lines
189 cur.execute("ALTER TABLE projects ADD COLUMN secon
dary_model_name TEXT")
190 if "secondary_model_classes" not in cols:
191 cur.execute("ALTER TABLE projects ADD COLUMN secon
dary_model_classes TEXT")
192 + # REQ-107: what rule set a run's numbers were measured
+ under.
193 + cur.execute("PRAGMA table_info(model_versions)")
194 + version_cols = [column[1] for column in cur.fetchall()
+]
195 + if "rule_version" not in version_cols:
196 + cur.execute("ALTER TABLE model_versions ADD COLUMN
+ rule_version TEXT")
197
198
199 def healthy() -> bool:
● Now the resolver — the core of the whole feature:
● Write(backend/triage.py)
⎿ Wrote 267 lines to backend/triage.py
1 """Triage: deciding what each SAM3 shape is actually worth (REQ-1
00…108).
3 A shape is never rewritten. Its verdict is *resolved* every time
it is needed:
4
6
7 so `annotations.class_id` keeps whatever SAM3 said, and any rule
can be re-cut
are evaluated
9 at training time rather than baked in at merge (REQ-102).
… +305 lines (ctrl+o to expand)
● Now wire the resolver into what training actually sees:
Read 1 file (ctrl+o to expand)
● Update(backend/dataset.py)
⎿ Added 36 lines, removed 10 lines
136
137
138 def _build_selected_tree(root: str, rows: list, class_map: Opt
ional[dict],
139 - label_type: str) -> tuple:
139 + label_type: str, resolver=None) -> tu
+ple:
140 """Materialise a filtered view of the dataset under `selec
ted/`.
141
142 The master labels stay canonical. A per-run class filter r
enumbers ids into a
143 contiguous 0..k-1 space, which contradicts `project_classe
s`; writing that
144 into the shared label files would leave the master dataset
in the index space
145 of whichever training run happened to go last.
146 +
147 + This is also the one place triage rules are applied (REQ-1
+02): resolve every
148 + shape, drop the images carrying an `ignore`, and emit the
+rest under their
149 + effective class. Because it happens here and not at merge
+time, editing a
150 + rule re-cuts the dataset on the next run without touching
+what was merged.
151 """
152 selected_root = os.path.join(root, "selected")
153 if os.path.isdir(selected_root):
154 shutil.rmtree(selected_root)
155 listed = {"train": [], "val": []}
156 + excluded = 0
157 for image_rel, label_rel, frame_id, split in rows:
158 + annotations = review.listing(frame_id)
159 +
160 + if resolver is not None:
161 + resolved = []
162 + skip_image = False
163 + for item in annotations:
164 + shape = {"id": item["id"], "class_id": item["c
+lass_id"],
165 + "score": float(item.get("score") or 1
+.0),
166 + **triage.metrics(item["geometry"])}
167 + effective = resolver.effective_class(shape)
168 + if effective is None:
169 + # REQ-104: a known object left unlabeled w
+ould teach the
170 + # model it is background, so the image goe
+s too.
171 + skip_image = True
172 + break
173 + resolved.append({**item, "class_id": effective
+})
174 + if skip_image:
175 + excluded += 1
176 + continue
177 + annotations = resolved
178 +
179 + if class_map is not None:
180 + annotations = [a for a in annotations if a["class_
+id"] in class_map]
181 + lines = [_label_line(class_map[a["class_id"]] if class
+_map is not None else a["class_id"],
182 + a["geometry"], label_type)
183 + for a in annotations]
184 +
185 stem = os.path.basename(image_rel)
186 image_dst = os.path.join(selected_root, "images", spli
t, stem)
187 label_dst = os.path.join(selected_root, "labels", spli
t,
188 os.path.splitext(stem)[0] + "
.txt")
189 os.makedirs(os.path.dirname(image_dst), exist_ok=True)
190 os.makedirs(os.path.dirname(label_dst), exist_ok=True)
158 - source = os.path.join(root, image_rel)
191 if not os.path.exists(image_dst):
160 - os.symlink(source, image_dst)
161 -
162 - annotations = review.listing(frame_id)
163 - if class_map is not None:
164 - annotations = [a for a in annotations if a["class_
-id"] in class_map]
165 - lines = [_label_line(class_map[a["class_id"]] if class
-_map is not None else a["class_id"],
166 - a["geometry"], label_type)
167 - for a in annotations]
192 + os.symlink(os.path.join(root, image_rel), image_ds
+t)
193 _write_atomic(label_dst, "\n".join(lines) + ("\n" if l
ines else ""))
194 listed[split].append(image_dst)
195 + listed["excluded"] = excluded
196 return selected_root, listed
197
198
● Update(backend/dataset.py)
⎿ Added 6 lines, removed 1 line
213
214 names = ", ".join(f"'{item['name']}'" for item in target_c
lasses)
215
216 - if batch_ids or class_map is not None:
216 + # Rules re-cut the dataset on every run (REQ-102), so thei
+r presence alone is
217 + # enough to need the `selected/` tree — the master labels
+must stay canonical.
218 + resolver = triage.Resolver(project["id"])
219 + has_rules = bool(resolver.rules or resolver.overrides)
220 +
221 + if batch_ids or class_map is not None or has_rules:
222 with db.cursor() as cur:
223 where = "d.project_id = ?"
224 args = [project["id"]]
● Update(backend/dataset.py)
⎿ Added 3 lines, removed 1 line
233 )
234 rows = cur.fetchall()
235
236 - selected_root, listed = _build_selected_tree(root, row
-s, class_map, project["label_type"])
236 + selected_root, listed = _build_selected_tree(
237 + root, rows, class_map, project["label_type"],
238 + resolver=resolver if has_rules else None)
239 if require_val:
240 _require_val(len(listed["val"]), "the selected bat
ches")
241
● Update(backend/dataset.py)
⎿ Added 1 line, removed 1 line
17 import time
18 from typing import List, Optional
19
20 -from backend import batches, config, db, jobs, projects, review
20 +from backend import batches, config, db, jobs, projects, review
+, triage
21
22
23 class DatasetError(Exception):
Ran 1 shell command (ctrl+o to expand)
● Now the API routes and crop endpoint:
● Write(backend/api/triage.py)
⎿ Wrote 118 lines to backend/api/triage.py
1 """Triage routes: rules, hand overrides, and the per-batch shape
view (REQ-100…108)."""
2
3 import io
4 import os
5
6 from fastapi import APIRouter, HTTPException
7 from fastapi.responses import StreamingResponse
8 from pydantic import BaseModel
9 from typing import List, Optional
… +115 lines (ctrl+o to expand)
Ran 1 shell command (ctrl+o to expand)
● Good. Registering the router:
● Update(backend/main.py)
⎿ Added 1 line, removed 1 line
15 from fastapi.middleware.cors import CORSMiddleware
16
17 from backend import config, db, jobs
18 -from backend.api import batches, jobs as job_routes, models, pr
-ojects, review
18 +from backend.api import batches, jobs as job_routes, models, pr
+ojects, review, triage
19
20
21 @asynccontextmanager
● Update(backend/main.py)
⎿ Added 1 line
43 app.include_router(review.router)
44 app.include_router(models.router)
45 app.include_router(job_routes.router)
46 +app.include_router(triage.router)
47
48
49 @app.get("/api/health")
● Let me verify the resolver logic end-to-end against a real scenario before
building UI on top of it:
Ran 1 shell command (ctrl+o to expand)
● Rules resolve correctly — the FK error is my test using fake annotation ids.
Let me redo with real rows:
Ran 3 shell commands (ctrl+o to expand)
● Resolver verified: precedence, first-match-wins, manual override, version
changes. Now the UI. Let me check the api client conventions:
Ran 1 shell command (ctrl+o to expand)
● Update(frontend/src/api.js)
⎿ Added 15 lines
116 approveAllBatchFrames: (batchId) => request(`/batches/${batc
hId}/approve-all`, { method: 'POST' }),
117 approveBatch: (batchId) => request(`/batches/${batchId}/appr
ove`, { method: 'POST' }),
118 datasetSummary: (projectId) => request(`/projects/${projectI
d}/dataset`),
119 +
120 + // Triage (REQ-100…108)
121 + triageRules: (projectId) => request(`/projects/${projectId}/
+triage/rules`),
122 + saveTriageRules: (projectId, rules) =>
123 + request(`/projects/${projectId}/triage/rules`, { method: '
+PUT', body: { rules } }),
124 + triageShapes: (batchId) => request(`/batches/${batchId}/tria
+ge/shapes`),
125 + triagePreview: (projectId) => request(`/projects/${projectId
+}/triage/preview`),
126 + setTriageOverrides: (annotationIds, verdict, targetClass = n
+ull) =>
127 + request('/triage/overrides', {
128 + method: 'POST',
129 + body: { annotation_ids: annotationIds, verdict, target_c
+lass: targetClass },
130 + }),
131 + clearTriageOverrides: (annotationIds) =>
132 + request('/triage/overrides', { method: 'DELETE', body: { a
+nnotation_ids: annotationIds } }),
133 + cropUrl: (annotationId) => `/api/annotations/${annotationId}
+/crop`,
134 datasetDownloadUrl: (projectId) => `/api/projects/${projectI
d}/dataset/download`,
135
136 hardware: () => request('/hardware'),
● Now the scatter plot with drag-selection:
● Write(frontend/src/components/TriageScatter.jsx)
⎿ Wrote 165 lines to frontend/src/components/TriageScatter.jsx
1 import React, { useMemo, useRef, useState } from 'react'
2
EQ-106).
4 *
5 * Area is on a log scale: box sizes span three orders of magnitu
6 * linear axis every shape piles into the left edge where no boun
dary is
7 * visible. */
9 const VERDICT_COLOR = { keep: '#4ade80', ignore: '#f87171', recla
ss: '#c084fc' }
… +179 lines (ctrl+o to expand)
● Now the crop grid and rule editor:
● Write(frontend/src/components/TriageCropGrid.jsx)
⎿ Wrote 128 lines to frontend/src/components/TriageCropGrid.jsx
1 import React, { useState } from 'react'
2 import { api } from '../api'
3
ck?" means
5 * looking at the object, not at the frame it came from. Sorted s
o the shapes a
6 * rule is about to act on sit together. */
8 const VERDICT_COLOR = { keep: '#4ade80', ignore: '#f87171', recla
ss: '#c084fc' }
9 const PAGE = 120
… +144 lines (ctrl+o to expand)
● Write(frontend/src/components/TriageRules.jsx)
⎿ Wrote 177 lines to frontend/src/components/TriageRules.jsx
1 import React from 'react'
2 import { TrashIcon } from './Icons'
3
4 /* The rule list (REQ-100). Ordered, first match wins — so the o
rder is part of
6
7 const ACTIONS = [
8 { value: 'keep', label: 'Keep' },
9 { value: 'ignore', label: 'Ignore' },
10 { value: 'reclass', label: 'Reclass to' },
… +198 lines (ctrl+o to expand)
● Now the page itself, wiring it together:
Searched for 2 patterns (ctrl+o to expand)
● Write(frontend/src/pages/DataPrepPage.jsx)
⎿ Added 237 lines, removed 569 lines
1 -import React, { useCallback, useEffect, useState } from 'react
-'
2 -import { api, classColor } from '../api'
1 +import React, { useCallback, useEffect, useMemo, useState } fr
+om 'react'
2 +import { api } from '../api'
3
4 -import { AlertIcon, BarChartIcon, DatabaseIcon, RocketIcon } f
-rom '../components/Icons'
4 +import { AlertIcon, DatabaseIcon, RocketIcon, SlidersIcon } fr
+om '../components/Icons'
5 +import TriageCropGrid from '../components/TriageCropGrid'
6 +import TriageRules from '../components/TriageRules'
7 +import TriageScatter from '../components/TriageScatter'
8
9 +/* Data Prep is the triage gate between SAM3's raw output and
+training
10 + * (REQ-100…108): look at what SAM3 produced for one batch, de
+cide what is worth
11 + * training on, and express the repeatable part of that decisi
+on as rules.
12 + *
13 + * Nothing here rewrites an annotation. Verdicts resolve as
14 + * `manual > first matching rule > keep`, and the dataset is r
+e-cut from them at
15 + * training time — so a rule stays editable forever. */
16 +
17 +const VERDICTS = [
18 + { value: 'keep', label: 'Keep', color: '#4ade80' },
19 + { value: 'ignore', label: 'Ignore', color: '#f87171' },
20 + { value: 'reclass', label: 'Reclass', color: '#c084fc' },
21 +]
22 +
23 export default function DataPrepPage({ projectId, onProject })
{
24 const [project, setProject] = useState(null)
8 - const [summary, setSummary] = useState(null)
9 - const [scaleMode, setScaleMode] = useState('normal') // 'nor
-mal' (Gaussian Bell Curve) | 'linear'
10 - const [minSizePct, setMinSizePct] = useState(0)
11 - const [maxSizePct, setMaxSizePct] = useState(100)
12 - const [previewFrameId, setPreviewFrameId] = useState(null)
13 - const [showFilteredOut, setShowFilteredOut] = useState(true)
25 + const [batches, setBatches] = useState([])
26 + const [batchId, setBatchId] = useState(null)
27 + const [view, setView] = useState(null)
28 + const [rules, setRules] = useState([])
29 + const [savedRules, setSavedRules] = useState([])
30 + const [preview, setPreview] = useState(null)
31 + const [selectedIds, setSelectedIds] = useState([])
32 + const [reclassTarget, setReclassTarget] = useState('')
33 + const [saving, setSaving] = useState(false)
34 + const [busy, setBusy] = useState(false)
35 const [error, setError] = useState('')
36
37 + useEffect(() => {
38 + let cancelled = false
39 + ;(async () => {
40 + try {
41 + const [loadedProject, loadedBatches, loadedRules] = aw
+ait Promise.all([
42 + api.getProject(projectId),
43 + api.listBatches(projectId),
44 + api.triageRules(projectId),
45 + ])
46 + if (cancelled) return
47 + setProject(loadedProject)
48 + onProject?.(loadedProject)
49 + const annotated = (loadedBatches.batches ?? loadedBatc
+hes).filter((b) => b.annotation_count > 0)
50 + setBatches(annotated)
51 + setBatchId((current) => current ?? annotated[0]?.id ??
+ null)
52 + setRules(loadedRules.rules)
53 + setSavedRules(loadedRules.rules)
54 + } catch (exc) {
55 + if (!cancelled) setError(exc.message)
56 + }
57 + })()
58 + return () => { cancelled = true }
59 + }, [projectId])
60
17 - const load = useCallback(async () => {
61 + const loadBatch = useCallback(async () => {
62 + if (!batchId) return
63 try {
19 - const [loadedProject, loadedSummary] = await Promise.all
-([
20 - api.getProject(projectId),
21 - api.datasetSummary(projectId),
64 + const [shapes, previewed] = await Promise.all([
65 + api.triageShapes(batchId),
66 + api.triagePreview(projectId),
67 ])
23 - setProject(loadedProject)
24 - onProject?.(loadedProject)
25 - setSummary(loadedSummary)
68 + setView(shapes)
69 + setPreview(previewed)
70 + setSelectedIds([])
71 } catch (exc) {
72 setError(exc.message)
73 }
29 - }, [projectId])
74 + }, [batchId, projectId])
75
31 - useEffect(() => {
32 - load()
33 - }, [load])
76 + useEffect(() => { loadBatch() }, [loadBatch])
77
35 - if (error && !project) {
36 - return