This commit includes major additions and updates to the frontend and backend architectures, introducing new dataset management, live counting features, batch processing, and triage logic. Includes new UI pages, components, and API routes.
6029 lines
351 KiB
Plaintext
6029 lines
351 KiB
Plaintext
▐▛███▜▌ Claude Code v2.1.224
|
||
▝▜█████▛▘ Opus 5 · Claude Pro
|
||
▘▘ ▝▝ ~/reTraining
|
||
|
||
Tackle your toughest work with Opus 5. Switch anytime with /model.
|
||
|
||
❯ /clear
|
||
|
||
❯ please create a deep understanding using workflow to understand this project
|
||
repo and bugs, give me recomendation of improving the code especially for
|
||
make this project funtion as i want. ask me a lot before you do that lot of
|
||
task, swarm cheaper agent model for helping you, you can use opus also but
|
||
its only for complex or task, i guess you can defined yourself. the goal is
|
||
to make a better workflow for reTraining my ai model, if i had a mistake just
|
||
say it, your goal is to help me make this app works well and ready to use.
|
||
|
||
Searched for 1 pattern, listed 1 directory (ctrl+o to expand)
|
||
|
||
● Repo is substantial (backend ~3.5k lines, 8 frontend pages, docs with REQ-xxx
|
||
already agreed). Before I spend a workflow on it, four things:
|
||
|
||
● User answered Claude's questions:
|
||
⎿ · What's actually broken or unsatisfying right now? Where should the audit
|
||
bite hardest? → Retraining loop itself, Auto-annotate / SAM3, Review editor
|
||
UX, Everything / I don't know yet
|
||
· Have you run a full round end-to-end yet (video → frames → annotate →
|
||
review → merge → train → compare)? → Finished once, but rough
|
||
· There's a lot of uncommitted work (24 modified files, 5 new untracked).
|
||
What's its status? → Half-broken experiments
|
||
· What do you want me to deliver at the end of the audit? → Report, then
|
||
fix the critical bugs
|
||
|
||
● Good — full sweep it is. Launching the workflow now.
|
||
|
||
● Running in the background. Structure:
|
||
|
||
- Map (4 parallel) — retraining core, SAM3/autolabel, review editor,
|
||
docs-drift + your dirty working tree
|
||
- Hunt (5 dimensions, each finding verified by an adversarial refuter as it
|
||
lands) — invariants/data corruption, concurrency & jobs, API contract,
|
||
frontend correctness, ops/portability
|
||
- Synthesize — mechanism-level system map, ranked bugs, per-file verdict on
|
||
the uncommitted work, workflow gaps where the design itself is wrong, and an
|
||
ordered fix plan with verification steps
|
||
|
||
Watch live with /workflows. When it lands I'll relay the report and then start
|
||
on the critical fixes.
|
||
|
||
✻ Brewed for 1m 26s
|
||
|
||
● Dynamic workflow "Deep audit of the reTraining dataset-enrichment app: map
|
||
architecture, find bugs, verify, prioritize fixes" completed · 7m 59s
|
||
⎿ You've hit your session limit · resets 12:10pm (Asia/Jakarta)
|
||
/upgrade to increase your usage limit.
|
||
|
||
✻ Crunched for 10m 22s
|
||
|
||
❯ continue
|
||
|
||
● The audit ran but hit the session usage limit near the end — the final
|
||
synthesis agent and 26 verifiers died. 22 bugs did survive verification. Let
|
||
me pull the full results and write the report myself.
|
||
|
||
Ran 3 shell commands (ctrl+o to expand)
|
||
|
||
● Two dimensions (concurrency, api-contract) lost all their verifiers to the
|
||
session limit — let me recover their raw findings from the journal so nothing
|
||
is silently dropped.
|
||
|
||
Ran 6 shell commands (ctrl+o to expand)
|
||
|
||
● Recovered everything. One key thing to confirm before I write it up — the two
|
||
CRITICALs point at jobs.py, which is in your dirty tree:
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Confirmed — that uncommitted jobs.py edit is the root of both CRITICALs.
|
||
Writing the report.
|
||
|
||
● Write(docs/audit-2026-08-07.md)
|
||
⎿ Wrote 181 lines to docs/audit-2026-08-07.md
|
||
1 # Audit — reTraining, 2026-08-07
|
||
2
|
||
by dimension, adversarial
|
||
4 verification of each finding. 22 findings survived verification;
|
||
2 CRITICALs come from a
|
||
e marked *unverified* — both
|
||
6 were confirmed by hand afterwards.
|
||
7
|
||
9
|
||
10 ## 1. How the system actually works
|
||
… +343 lines (ctrl+o to expand)
|
||
|
||
● Report written to docs/audit-2026-08-07.md. 22 verified findings + 2
|
||
hand-confirmed criticals.
|
||
|
||
The headline: your biggest problem is uncommitted. You changed backend/jobs.py
|
||
to run every job in its own thread instead of through the single worker
|
||
queue, and narrowed GPU_JOB_TYPES to ("train",) with the comment "multiple
|
||
autolabel jobs can run in parallel." They can't — each SAM3 job holds a full
|
||
backbone in VRAM, and release_engine() only clears a module global, so it
|
||
cannot free memory held by a running job. That one edit also un-serialized
|
||
merge and train against the same dataset/ directory, which is the root of four
|
||
more HIGH findings.
|
||
|
||
Three design-level things you got wrong (you asked me to say so):
|
||
|
||
1. The val split is positional — _next_split counts dataset_items rows, so
|
||
every Nth insert goes to val. Deleting a batch shifts the split for everything
|
||
merged afterwards. CLAUDE.md §8 declares the invariant but nothing enforces
|
||
it. Make it hash(project + batch + stem) % 5 and it becomes true by
|
||
construction.
|
||
2. Merged labels are regenerated from live annotations on every training start
|
||
(sync_labels at dataset.py:117, no review_status filter). Re-run
|
||
auto-annotate on a merged batch, press Train, and raw unreviewed model output
|
||
overwrites your human-reviewed labels. A merged dataset should be a snapshot,
|
||
not a view.
|
||
3. The base-vs-new comparison degrades silently instead of refusing. Empty val
|
||
set → it validates on the training images. Deleted batch → orphaned JPEGs
|
||
still sit in images/val. For a tool whose only job is answering "did
|
||
retraining help?", every one of those returns a plausible-looking number.
|
||
|
||
Worth knowing: docker-compose.override.yml is committed with
|
||
nvidia.com/gpu=all (CPU hosts can't start), and start.sh deletes that tracked
|
||
file on every run — silently dropping your ./backend bind mount.
|
||
|
||
Before I start fixing:
|
||
● User answered Claude's questions:
|
||
⎿ · Fixing C1/C2 means reverting your jobs.py parallelism change. Was that
|
||
experiment solving a real problem I should address differently? → Revert,
|
||
but autolabel was too slow
|
||
· How far do I go in this pass? → Criticals + design + review editor
|
||
· The content-derived val split changes how existing merged frames are
|
||
assigned. Existing data? → Preserve existing assignments
|
||
|
||
● Understood: revert for correctness, then look at autolabel speed separately.
|
||
Steps 1–7, preserving existing splits.
|
||
|
||
Read 2 files (ctrl+o to expand)
|
||
|
||
● Now fixing. Step 1 — restore the job queue:
|
||
|
||
● Update(backend/jobs.py)
|
||
⎿ Added 3 lines, removed 5 lines
|
||
28 PROGRESS_FLUSH_SECONDS = 0.5
|
||
29
|
||
30 JOB_TYPES = ("extract", "autolabel", "merge", "train")
|
||
31 -GPU_JOB_TYPES = ("train",)
|
||
32 -"""`train` takes full GPU memory during model fitting. `autolab
|
||
-el` (SAM3/YOLO inference)
|
||
33 -runs fine-grained per-frame steps so multiple autolabel jobs ca
|
||
-n run in parallel."""
|
||
34 -
|
||
35 -
|
||
31 +GPU_JOB_TYPES = ("autolabel", "train")
|
||
32 +"""`extract` is ffmpeg and `merge` is file copying — neither to
|
||
+uches the card,
|
||
33 +so neither should be able to block an interactive assist."""
|
||
34 gpu_lock = threading.Lock()
|
||
35 """Held for the duration of any GPU work. The job worker takes
|
||
it around a
|
||
36 handler; the interactive assist route takes it around one SAM3
|
||
call. One card,
|
||
|
||
● Update(backend/jobs.py)
|
||
⎿ Added 2 lines, removed 12 lines
|
||
147 job_id = cur.lastrowid
|
||
148 job = get(job_id)
|
||
149 assert job is not None
|
||
150 - threading.Thread(target=_start_job, args=(job_id,), name=f
|
||
-"job-{job_type}-{job_id}", daemon=True).start()
|
||
150 + _queue.put(job_id)
|
||
151 + _ensure_worker()
|
||
152 return job
|
||
153
|
||
154
|
||
154 -def _start_job(job_id: int) -> None:
|
||
155 - job = get(job_id)
|
||
156 - if job is None:
|
||
157 - return
|
||
158 - if job.id in _cancelled:
|
||
159 - _finish(job, "cancelled")
|
||
160 - return
|
||
161 - _run(job)
|
||
162 -
|
||
163 -
|
||
164 -
|
||
155 def get(job_id: int) -> Optional[Job]:
|
||
156 with db.cursor() as cur:
|
||
157 cur.execute("SELECT * FROM jobs WHERE id = ?", (job_id
|
||
,))
|
||
|
||
● Update(backend/jobs.py)
|
||
⎿ Added 20 lines, removed 1 line
|
||
206 return cur.rowcount
|
||
207
|
||
208
|
||
209 -def _run(job: Job) -> None:
|
||
209 +def _ensure_worker() -> None:
|
||
210 + global _worker
|
||
211 + with _worker_lock:
|
||
212 + if _worker is None or not _worker.is_alive():
|
||
213 + _worker = threading.Thread(target=_worker_loop, na
|
||
+me="job-worker", daemon=True)
|
||
214 + _worker.start()
|
||
215
|
||
216 +
|
||
217 +def _worker_loop() -> None:
|
||
218 + while True:
|
||
219 + job_id = _queue.get()
|
||
220 + job = get(job_id)
|
||
221 + if job is None:
|
||
222 + continue
|
||
223 + if job.id in _cancelled:
|
||
224 + _finish(job, "cancelled")
|
||
225 + continue
|
||
226 + _run(job)
|
||
227 +
|
||
228 +
|
||
229 +def _run(job: Job) -> None:
|
||
230 job.status = "running"
|
||
231 job.started_at = time.time()
|
||
232 job.flush()
|
||
|
||
Read 1 file, ran 1 shell command (ctrl+o to expand)
|
||
|
||
● jobs.py is back to baseline. Step 2 — the synchronous GPU routes. Let me match
|
||
the existing review.assist pattern:
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Update(backend/api/batches.py)
|
||
⎿ Added 8 lines, removed 1 line
|
||
162
|
||
163 @router.post("/api/batches/{batch_id}/preview")
|
||
164 def preview_autolabel(batch_id: int, request: PreviewRequest)
|
||
-> dict:
|
||
165 - from backend import autolabel
|
||
165 + from backend import autolabel, jobs
|
||
166 +
|
||
167 + if not jobs.gpu_lock.acquire(timeout=20):
|
||
168 + busy = jobs.running_types()
|
||
169 + kind = busy[0] if busy else "background"
|
||
170 + raise HTTPException(409, f"The GPU is busy with a {kin
|
||
+d} job — wait for it to finish")
|
||
171 try:
|
||
172 shapes = autolabel.preview_frame(
|
||
173 batch_id=batch_id,
|
||
...
|
||
182 return {"shapes": shapes}
|
||
183 except Exception as exc:
|
||
184 raise HTTPException(400, str(exc))
|
||
185 + finally:
|
||
186 + jobs.gpu_lock.release()
|
||
187
|
||
188 @router.post("/api/sam3/playground-test")
|
||
189 async def sam3_playground_test(
|
||
|
||
● Update(backend/api/batches.py)
|
||
⎿ Added 8 lines
|
||
207 if not prompt_list:
|
||
208 raise HTTPException(400, "At least one text prompt is
|
||
required")
|
||
209
|
||
210 + from backend import jobs
|
||
211 +
|
||
212 + if not jobs.gpu_lock.acquire(timeout=20):
|
||
213 + busy = jobs.running_types()
|
||
214 + kind = busy[0] if busy else "background"
|
||
215 + raise HTTPException(409, f"The GPU is busy with a {kin
|
||
+d} job — wait for it to finish")
|
||
216 try:
|
||
217 engine = get_engine()
|
||
218 raw_dets = engine.detect(image, prompt_list, threshold
|
||
)
|
||
219 kept_dets = labeling.deduplicate(raw_dets, iou_thresho
|
||
ld=iou_threshold)
|
||
220 except Exception as exc:
|
||
221 raise HTTPException(500, f"SAM3 inference failed: {exc
|
||
}")
|
||
222 + finally:
|
||
223 + jobs.gpu_lock.release()
|
||
224
|
||
225 results = []
|
||
226 for det in kept_dets:
|
||
|
||
● Step 3, 4 and 6 all live in dataset.py. Rewriting that section:
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 72 lines, removed 23 lines
|
||
66 return f"{class_id} {coords}"
|
||
67
|
||
68
|
||
69 -def _next_split(cur, project_id: int, val_every: int) -> str:
|
||
70 - """Continue the every-Nth pattern from wherever the last m
|
||
-erge left off."""
|
||
69 +def split_for(project_id: int, batch_id: int, stem: str, val_e
|
||
+very: int) -> str:
|
||
70 + """Which split a frame belongs to, derived from its identi
|
||
+ty rather than from
|
||
71 + how many rows happen to precede it.
|
||
72 +
|
||
73 + A positional every-Nth rule makes membership depend on ins
|
||
+ertion history, so
|
||
74 + deleting or re-merging a batch silently reshuffles every l
|
||
+ater frame — and a
|
||
75 + frame that was in `val` for the last comparison could land
|
||
+ in `train` for the
|
||
76 + next one. Hashing the identity makes the stable-val-split
|
||
+invariant true by
|
||
77 + construction: the same frame always lands in the same spli
|
||
+t, whatever else
|
||
78 + happened to the dataset. Rows already in `dataset_items` k
|
||
+eep the split they
|
||
79 + were recorded with; nothing recomputes them.
|
||
80 + """
|
||
81 if val_every <= 0:
|
||
82 return "train"
|
||
73 - cur.execute("SELECT COUNT(*) FROM dataset_items WHERE proj
|
||
-ect_id = ?", (project_id,))
|
||
74 - position = cur.fetchone()[0]
|
||
75 - return "val" if position % val_every == val_every - 1 else
|
||
- "train"
|
||
83 + digest = hashlib.sha1(f"{project_id}/{batch_id}/{stem}".en
|
||
+code("utf-8")).hexdigest()
|
||
84 + return "val" if int(digest[:8], 16) % val_every == 0 else
|
||
+"train"
|
||
85
|
||
86
|
||
78 -def sync_labels(project_id: int, selected_class_ids: Optional[
|
||
-List[int]] = None) -> dict:
|
||
79 - """Re-sync label files on disk for all merged frames in th
|
||
-e project dataset."""
|
||
87 +def sync_labels(project_id: int) -> dict:
|
||
88 + """Re-sync label files on disk from the reviewed annotatio
|
||
+ns.
|
||
89 +
|
||
90 + Only frames a human actually signed off on are written. A
|
||
+merged frame whose
|
||
91 + batch was auto-annotated again drops back to `pending`, an
|
||
+d rewriting its
|
||
92 + label file from the fresh model output would push predicti
|
||
+ons nobody checked
|
||
93 + into the master dataset.
|
||
94 + """
|
||
95 project = projects.get(project_id)
|
||
96 root = dataset_dir(project["slug"])
|
||
97 with db.cursor() as cur:
|
||
98 cur.execute(
|
||
84 - "SELECT d.frame_id, d.label_rel FROM dataset_items
|
||
- d WHERE d.project_id = ?",
|
||
99 + """SELECT d.frame_id, d.label_rel FROM dataset_ite
|
||
+ms d
|
||
100 + JOIN frames f ON f.id = d.frame_id
|
||
101 + WHERE d.project_id = ? AND f.review_status = 'a
|
||
+pproved'""",
|
||
102 (project_id,),
|
||
103 )
|
||
104 items = cur.fetchall()
|
||
105
|
||
89 - class_map = None
|
||
90 - if selected_class_ids is not None and len(selected_class_i
|
||
-ds) > 0:
|
||
91 - class_map = {cid: idx for idx, cid in enumerate(sorted
|
||
-(selected_class_ids))}
|
||
92 -
|
||
106 synced_files = 0
|
||
107 total_lines = 0
|
||
108 for frame_id, label_rel in items:
|
||
96 - annotations = review.listing(frame_id)
|
||
97 - if class_map is not None:
|
||
98 - annotations = [a for a in annotations if a["class_
|
||
-id"] in class_map]
|
||
99 -
|
||
100 - lines = []
|
||
101 - for item in annotations:
|
||
102 - mapped_cid = class_map[item["class_id"]] if class_
|
||
-map is not None else item["class_id"]
|
||
103 - lines.append(_label_line(mapped_cid, item["geometr
|
||
-y"], project["label_type"]))
|
||
104 -
|
||
109 + lines = [_label_line(item["class_id"], item["geometry"
|
||
+], project["label_type"])
|
||
110 + for item in review.listing(frame_id)]
|
||
111 path = os.path.join(root, label_rel)
|
||
112 os.makedirs(os.path.dirname(path), exist_ok=True)
|
||
107 - with open(path, "w", encoding="utf-8") as f:
|
||
108 - f.write("\n".join(lines) + ("\n" if lines else "")
|
||
-)
|
||
113 + _write_atomic(path, "\n".join(lines) + ("\n" if lines
|
||
+else ""))
|
||
114 synced_files += 1
|
||
115 total_lines += len(lines)
|
||
116
|
||
117 return {"synced_files": synced_files, "total_lines": total
|
||
_lines}
|
||
118
|
||
119
|
||
120 +def _write_atomic(path: str, text: str) -> None:
|
||
121 + """Write via temp file + rename, so a training run never r
|
||
+eads a half-written
|
||
122 + label file or a truncated data.yaml."""
|
||
123 + tmp = f"{path}.tmp"
|
||
124 + with open(tmp, "w", encoding="utf-8") as handle:
|
||
125 + handle.write(text)
|
||
126 + os.replace(tmp, path)
|
||
127 +
|
||
128 +
|
||
129 +def _build_selected_tree(root: str, rows: list, class_map: Opt
|
||
+ional[dict],
|
||
130 + label_type: str) -> tuple:
|
||
131 + """Materialise a filtered view of the dataset under `selec
|
||
+ted/`.
|
||
132 +
|
||
133 + The master labels stay canonical. A per-run class filter r
|
||
+enumbers ids into a
|
||
134 + contiguous 0..k-1 space, which contradicts `project_classe
|
||
+s`; writing that
|
||
135 + into the shared label files would leave the master dataset
|
||
+ in the index space
|
||
136 + of whichever training run happened to go last.
|
||
137 + """
|
||
138 + selected_root = os.path.join(root, "selected")
|
||
139 + if os.path.isdir(selected_root):
|
||
140 + shutil.rmtree(selected_root)
|
||
141 + listed = {"train": [], "val": []}
|
||
142 + for image_rel, label_rel, frame_id, split in rows:
|
||
143 + stem = os.path.basename(image_rel)
|
||
144 + image_dst = os.path.join(selected_root, "images", spli
|
||
+t, stem)
|
||
145 + label_dst = os.path.join(selected_root, "labels", spli
|
||
+t,
|
||
146 + os.path.splitext(stem)[0] + "
|
||
+.txt")
|
||
147 + os.makedirs(os.path.dirname(image_dst), exist_ok=True)
|
||
148 + os.makedirs(os.path.dirname(label_dst), exist_ok=True)
|
||
149 + source = os.path.join(root, image_rel)
|
||
150 + if not os.path.exists(image_dst):
|
||
151 + os.symlink(source, image_dst)
|
||
152 +
|
||
153 + annotations = review.listing(frame_id)
|
||
154 + if class_map is not None:
|
||
155 + annotations = [a for a in annotations if a["class_
|
||
+id"] in class_map]
|
||
156 + lines = [_label_line(class_map[a["class_id"]] if class
|
||
+_map is not None else a["class_id"],
|
||
157 + a["geometry"], label_type)
|
||
158 + for a in annotations]
|
||
159 + _write_atomic(label_dst, "\n".join(lines) + ("\n" if l
|
||
+ines else ""))
|
||
160 + listed[split].append(image_dst)
|
||
161 + return selected_root, listed
|
||
162 +
|
||
163 +
|
||
164 def write_data_yaml(project: dict, batch_ids: list = None, sel
|
||
ected_class_ids: Optional[List[int]] = None) -> str:
|
||
165 """Rebuild data.yaml from the project's classes (REQ-051).
|
||
"""
|
||
166 sync_labels(project["id"], selected_class_ids=selected_cla
|
||
ss_ids)
|
||
|
||
● Now write_data_yaml — fail-loud on empty val, correct nc, use the selected
|
||
tree:
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 46 lines, removed 24 lines
|
||
163
|
||
164 def write_data_yaml(project: dict, batch_ids: list = None, sel
|
||
ected_class_ids: Optional[List[int]] = None) -> str:
|
||
165 """Rebuild data.yaml from the project's classes (REQ-051).
|
||
"""
|
||
166 - sync_labels(project["id"], selected_class_ids=selected_cla
|
||
-ss_ids)
|
||
166 + sync_labels(project["id"])
|
||
167 root = dataset_dir(project["slug"])
|
||
168 os.makedirs(root, exist_ok=True)
|
||
169 counts = summary(project["id"])["splits"]
|
||
170
|
||
171 target_classes = project["classes"]
|
||
172 + class_map = None
|
||
173 if selected_class_ids is not None and len(selected_class_i
|
||
ds) > 0:
|
||
174 target_classes = [c for c in project["classes"] if c["
|
||
class_id"] in selected_class_ids]
|
||
175 + class_map = {cid: idx for idx, cid in enumerate(sorted
|
||
+(selected_class_ids))}
|
||
176
|
||
177 names = ", ".join(f"'{item['name']}'" for item in target_c
|
||
lasses)
|
||
178
|
||
177 - if batch_ids:
|
||
179 + if batch_ids or class_map is not None:
|
||
180 with db.cursor() as cur:
|
||
179 - placeholders = ",".join("?" for _ in batch_ids)
|
||
181 + where = "d.project_id = ?"
|
||
182 + args = [project["id"]]
|
||
183 + if batch_ids:
|
||
184 + where += f" AND f.batch_id IN ({','.join('?' f
|
||
+or _ in batch_ids)})"
|
||
185 + args += list(batch_ids)
|
||
186 cur.execute(
|
||
181 - f"""SELECT d.image_rel, d.split FROM dataset_i
|
||
-tems d
|
||
187 + f"""SELECT d.image_rel, d.label_rel, d.frame_i
|
||
+d, d.split FROM dataset_items d
|
||
188 JOIN frames f ON f.id = d.frame_id
|
||
183 - WHERE d.project_id = ? AND f.batch_id IN (
|
||
-{placeholders})""",
|
||
184 - [project["id"]] + list(batch_ids),
|
||
189 + WHERE {where}""",
|
||
190 + args,
|
||
191 )
|
||
192 rows = cur.fetchall()
|
||
193
|
||
188 - train_files = [row[0] for row in rows if row[1] == "tr
|
||
-ain"]
|
||
189 - val_files = [row[0] for row in rows if row[1] == "val"
|
||
-] or train_files
|
||
194 + selected_root, listed = _build_selected_tree(root, row
|
||
+s, class_map, project["label_type"])
|
||
195 + _require_val(len(listed["val"]), "the selected batches
|
||
+")
|
||
196
|
||
197 train_txt = os.path.join(root, "selected_train.txt")
|
||
198 val_txt = os.path.join(root, "selected_val.txt")
|
||
193 - with open(train_txt, "w", encoding="utf-8") as handle:
|
||
194 - handle.write("\n".join(os.path.join(root, rel) for
|
||
- rel in train_files) + "\n")
|
||
195 - with open(val_txt, "w", encoding="utf-8") as handle:
|
||
196 - handle.write("\n".join(os.path.join(root, rel) for
|
||
- rel in val_files) + "\n")
|
||
199 + _write_atomic(train_txt, "\n".join(listed["train"]) +
|
||
+"\n")
|
||
200 + _write_atomic(val_txt, "\n".join(listed["val"]) + "\n"
|
||
+)
|
||
201
|
||
202 path = os.path.join(root, "selected_data.yaml")
|
||
199 - with open(path, "w", encoding="utf-8") as handle:
|
||
200 - handle.write(f"path: {root}\n")
|
||
201 - handle.write(f"train: {train_txt}\n")
|
||
202 - handle.write(f"val: {val_txt}\n\n")
|
||
203 - handle.write(f"nc: {len(project['classes'])}\n")
|
||
204 - handle.write(f"names: [{names}]\n")
|
||
203 + _write_atomic(path,
|
||
204 + f"path: {selected_root}\n"
|
||
205 + f"train: {train_txt}\n"
|
||
206 + f"val: {val_txt}\n\n"
|
||
207 + f"nc: {len(target_classes)}\n"
|
||
208 + f"names: [{names}]\n")
|
||
209 return path
|
||
210
|
||
211 + _require_val(counts["val"], "this project")
|
||
212 +
|
||
213 path = os.path.join(root, "data.yaml")
|
||
208 - with open(path, "w", encoding="utf-8") as handle:
|
||
209 - handle.write(f"path: {root}\n")
|
||
210 - handle.write("train: images/train\n")
|
||
211 - handle.write(f"val: images/{'val' if counts['val'] > 0
|
||
- else 'train'}\n\n")
|
||
212 - handle.write(f"nc: {len(project['classes'])}\n")
|
||
213 - handle.write(f"names: [{names}]\n")
|
||
214 + _write_atomic(path,
|
||
215 + f"path: {root}\n"
|
||
216 + "train: images/train\n"
|
||
217 + "val: images/val\n\n"
|
||
218 + f"nc: {len(target_classes)}\n"
|
||
219 + f"names: [{names}]\n")
|
||
220 return path
|
||
221
|
||
222
|
||
223 +def _require_val(count: int, subject: str) -> None:
|
||
224 + """Refuse to build a dataset with an empty val split.
|
||
225 +
|
||
226 + Falling back to the training images produces a base-vs-new
|
||
+ mAP measured on
|
||
227 + data the model was fitted to — a number that looks fine an
|
||
+d means nothing.
|
||
228 + For a system whose whole purpose is answering "did retrain
|
||
+ing help?", this
|
||
229 + has to fail loudly.
|
||
230 + """
|
||
231 + if count == 0:
|
||
232 + raise DatasetError(
|
||
233 + f"There are no validation images in {subject}, so
|
||
+a base-vs-new comparison "
|
||
234 + "would be measured on the training images. Merge m
|
||
+ore frames, or lower the "
|
||
235 + "project's val_every."
|
||
236 + )
|
||
237 +
|
||
238 +
|
||
239 def summary(project_id: int) -> dict:
|
||
240 import json
|
||
241 with db.cursor() as cur:
|
||
|
||
● Now the merge handler — hash split, cancel bug, and the hashlib import:
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 1 line
|
||
11 segment class_id x1 y1 x2 y2 … (normalized polygon
|
||
)
|
||
12 """
|
||
13
|
||
14 +import hashlib
|
||
15 import os
|
||
16 import shutil
|
||
17 import time
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 3 lines, removed 1 line
|
||
436
|
||
437 added = {"train": 0, "val": 0}
|
||
438 skipped = 0
|
||
439 + cancelled = False
|
||
440 for index, frame in enumerate(frames):
|
||
441 if job.cancelled:
|
||
442 job.log(f"Cancelled after {index} frame(s)")
|
||
443 + cancelled = True
|
||
444 break
|
||
445
|
||
446 with db.cursor() as cur:
|
||
...
|
||
449 skipped += 1
|
||
450 job.progress(index + 1, len(frames))
|
||
451 continue
|
||
450 - split = _next_split(cur, project["id"], project["v
|
||
-al_every"])
|
||
452
|
||
453 stem = f"{batch['id']}__{os.path.splitext(frame['f
|
||
ilename'])[0]}"
|
||
454 + split = split_for(project["id"], batch["id"], stem
|
||
+, project["val_every"])
|
||
455 image_rel = f"images/{split}/{stem}.jpg"
|
||
456 label_rel = f"labels/{split}/{stem}.txt"
|
||
457 shutil.copyfile(os.path.join(source_dir, frame["fi
|
||
lename"]),
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 8 lines
|
||
473 added[split] += 1
|
||
474 job.progress(index + 1, len(frames))
|
||
475
|
||
476 + if cancelled:
|
||
477 + # Leaving it 'merged' would be a lie: the frames after
|
||
+ the break point
|
||
478 + # have no dataset_items rows and no files, and approve
|
||
+() refuses to
|
||
479 + # re-merge a merged batch, so they could never be adde
|
||
+d. The per-frame
|
||
480 + # dataset_items guard already makes re-running the mer
|
||
+ge idempotent.
|
||
481 + job.log("Batch left approved — re-approve it to finish
|
||
+ the merge")
|
||
482 + return
|
||
483 +
|
||
484 with db.cursor() as cur:
|
||
485 cur.execute("UPDATE batches SET status = 'merged', mer
|
||
ged_at = ? WHERE id = ?",
|
||
486 (time.time(), batch["id"]))
|
||
|
||
● Step 5 — batch deletion orphans, plus the double-approve guard:
|
||
|
||
Read 1 file (ctrl+o to expand)
|
||
|
||
● Update(backend/batches.py)
|
||
⎿ Added 23 lines
|
||
238
|
||
239 def delete(batch_id: int) -> bool:
|
||
240 import shutil
|
||
241 + from backend import dataset
|
||
242 +
|
||
243 batch = get(batch_id)
|
||
244 if batch is None:
|
||
245 return False
|
||
246 +
|
||
247 + # The FK cascade takes the frames and dataset_items rows,
|
||
+but the copies this
|
||
248 + # batch put in the master dataset live in a sibling tree a
|
||
+nd would survive as
|
||
249 + # orphans — and data.yaml trains and validates on those di
|
||
+rectories, so the
|
||
250 + # model would keep learning from a batch the user believes
|
||
+ is gone.
|
||
251 + root = dataset.dataset_dir(batch["project_slug"])
|
||
252 with db.cursor() as cur:
|
||
253 + cur.execute(
|
||
254 + """SELECT d.image_rel, d.label_rel FROM dataset_it
|
||
+ems d
|
||
255 + JOIN frames f ON f.id = d.frame_id
|
||
256 + WHERE f.batch_id = ?""",
|
||
257 + (batch_id,),
|
||
258 + )
|
||
259 + merged_files = cur.fetchall()
|
||
260 + for image_rel, label_rel in merged_files:
|
||
261 + for rel in (image_rel, label_rel):
|
||
262 + try:
|
||
263 + os.unlink(os.path.join(root, rel))
|
||
264 + except FileNotFoundError:
|
||
265 + pass
|
||
266 +
|
||
267 + with db.cursor() as cur:
|
||
268 cur.execute("DELETE FROM batches WHERE id = ?", (batch
|
||
_id,))
|
||
269 shutil.rmtree(batch_dir(batch["project_slug"], batch_id),
|
||
ignore_errors=True)
|
||
270 return True
|
||
|
||
● Update(backend/dataset.py) ⎿ Added 8 lines
|
||
42 )
|
||
43 if batch["review"]["approved"] == 0:
|
||
44 raise DatasetError("Every frame was rejected — there is
|
||
nothing to merge")
|
||
45 + with db.cursor() as cur:
|
||
46 + cur.execute(
|
||
47 + """SELECT 1 FROM jobs WHERE batch_id = ? AND type =
|
||
+ 'merge'
|
||
48 + AND status IN ('queued', 'running')""",
|
||
49 + (batch_id,),
|
||
50 + )
|
||
51 + if cur.fetchone() is not None:
|
||
52 + raise DatasetError("A merge for this batch is alrea
|
||
+dy queued")
|
||
53
|
||
54 batches.set_status(batch_id, "approved")
|
||
55 job = jobs.create(
|
||
|
||
Searched for 2 patterns (ctrl+o to expand)
|
||
|
||
● Good catch — class edits and merges would now throw on a young dataset. Gating
|
||
the check to the training path only:
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 3 lines, removed 1 line
|
||
170 return selected_root, listed
|
||
171
|
||
172
|
||
173 -def write_data_yaml(project: dict, batch_ids: list = None, sel
|
||
-ected_class_ids: Optional[List[int]] = None) -> str:
|
||
173 +def write_data_yaml(project: dict, batch_ids: list = None,
|
||
174 + selected_class_ids: Optional[List[int]] =
|
||
+None,
|
||
175 + require_val: bool = False) -> str:
|
||
176 """Rebuild data.yaml from the project's classes (REQ-051).
|
||
"""
|
||
177 sync_labels(project["id"])
|
||
178 root = dataset_dir(project["slug"])
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 2 lines, removed 1 line
|
||
203 rows = cur.fetchall()
|
||
204
|
||
205 selected_root, listed = _build_selected_tree(root, row
|
||
s, class_map, project["label_type"])
|
||
206 - _require_val(len(listed["val"]), "the selected batches
|
||
-")
|
||
206 + if require_val:
|
||
207 + _require_val(len(listed["val"]), "the selected bat
|
||
+ches")
|
||
208
|
||
209 train_txt = os.path.join(root, "selected_train.txt")
|
||
210 val_txt = os.path.join(root, "selected_val.txt")
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 2 lines, removed 1 line
|
||
220 f"names: [{names}]\n")
|
||
221 return path
|
||
222
|
||
223 - _require_val(counts["val"], "this project")
|
||
223 + if require_val:
|
||
224 + _require_val(counts["val"], "this project")
|
||
225
|
||
226 path = os.path.join(root, "data.yaml")
|
||
227 _write_atomic(path,
|
||
|
||
● Update(backend/training.py)
|
||
⎿ Added 2 lines, removed 1 line
|
||
103 settings = job.params["settings"]
|
||
104 batch_ids = job.params.get("batch_ids")
|
||
105 class_ids = job.params.get("class_ids")
|
||
106 - data_yaml = dataset.write_data_yaml(project, batch_ids=bat
|
||
-ch_ids, selected_class_ids=class_ids)
|
||
106 + data_yaml = dataset.write_data_yaml(project, batch_ids=bat
|
||
+ch_ids, selected_class_ids=class_ids,
|
||
107 + require_val=True)
|
||
108
|
||
109 # SAM3 and a training run must not hold VRAM at the same t
|
||
ime (REQ-065).
|
||
110 from backend.sam3_engine import release_engine
|
||
|
||
● Backend done. Step 7 — the review editor:
|
||
|
||
Read 1 file (ctrl+o to expand)
|
||
|
||
● Update(frontend/src/pages/ReviewPage.jsx)
|
||
⎿ Added 3 lines
|
||
234 useEffect(() => {
|
||
235 function onKeyDown(event) {
|
||
236 if (event.target?.matches?.('input, textarea, select, [c
|
||
ontenteditable="true"]')) return
|
||
237 + // Ctrl/Cmd/Alt combos belong to the browser and the OS:
|
||
+ without this,
|
||
238 + // Ctrl+S approves the frame and Ctrl+A/C/X/N/T all fire
|
||
+ review actions.
|
||
239 + if (event.ctrlKey || event.metaKey || event.altKey) retu
|
||
+rn
|
||
240 const { frames, project, setStatus, removeSelected, recl
|
||
ass, jumpToPending, jumpToNextAnnotated, setAssistMode, copyPr
|
||
evious, trackForward } = stateRef.current
|
||
241 const key = event.key
|
||
242 const isShortcutKey = /^[1-9]$/.test(key) || ['ArrowLeft
|
||
', 'ArrowRight', 'ArrowUp', 'ArrowDown', 'Delete', 'Backspace'
|
||
, 'a', 'A', 'x', 'X', 'u', 'U', 's', 'S', 'n', 'N', 'c', 'C',
|
||
't', 'T'].includes(key)
|
||
|
||
● Update(frontend/src/pages/ReviewPage.jsx)
|
||
⎿ Added 28 lines, removed 9 lines
|
||
124 return
|
||
125 }
|
||
126 if (!commit) return
|
||
127 - const current = annotations.find((row) => row.id === id)
|
||
128 - if (!current) return
|
||
129 - try { await api.patchAnnotation(id, { geometry: current.ge
|
||
-ometry }) } catch (exc) { setError(exc.message) }
|
||
127 + const previous = annotations.find((row) => row.id === id)
|
||
128 + if (!previous) return
|
||
129 + // A commit may carry its own geometry (delete-vertex send
|
||
+s the shortened
|
||
130 + // polygon); falling back to the row's geometry covers dra
|
||
+g/resize, which
|
||
131 + // already wrote through the {local:true} path.
|
||
132 + const next = geometry ?? previous.geometry
|
||
133 + setAnnotations((rows) => rows.map((row) => (row.id === id
|
||
+? { ...row, geometry: next } : row)))
|
||
134 + try {
|
||
135 + await api.patchAnnotation(id, { geometry: next })
|
||
136 + } catch (exc) {
|
||
137 + setAnnotations((rows) => rows.map((row) => (row.id === i
|
||
+d ? previous : row)))
|
||
138 + setError(exc.message)
|
||
139 + }
|
||
140 }
|
||
141
|
||
132 - const removeSelected = useCallback(async () => {
|
||
133 - if (selectedId == null || !frame) return
|
||
134 - const id = selectedId
|
||
142 + const removeAnnotation = useCallback(async (id) => {
|
||
143 + if (id == null || !frame) return
|
||
144 + const previous = annotations.find((row) => row.id === id)
|
||
145 + if (!previous) return
|
||
146 setAnnotations((rows) => rows.filter((row) => row.id !== i
|
||
d))
|
||
136 - setSelectedId(null)
|
||
147 + setSelectedId((current) => (current === id ? null : curren
|
||
+t))
|
||
148 setFrames((rows) => rows.map((row) => (row.id === frame.id
|
||
? { ...row, annotation_count: Math.max(0, (row.annotation_cou
|
||
nt ?? 1) - 1) } : row)))
|
||
138 - try { await api.deleteAnnotation(id) } catch (exc) { setEr
|
||
-ror(exc.message) }
|
||
139 - }, [selectedId, frame])
|
||
149 + try {
|
||
150 + await api.deleteAnnotation(id)
|
||
151 + } catch (exc) {
|
||
152 + setAnnotations((rows) => [...rows, previous])
|
||
153 + setFrames((rows) => rows.map((row) => (row.id === frame.
|
||
+id ? { ...row, annotation_count: (row.annotation_count ?? 0) +
|
||
+ 1 } : row)))
|
||
154 + setError(exc.message)
|
||
155 + }
|
||
156 + }, [annotations, frame])
|
||
157
|
||
158 + const removeSelected = useCallback(() => removeAnnotation(se
|
||
+lectedId), [removeAnnotation, selectedId])
|
||
159 +
|
||
160 const reclass = useCallback(async (classId) => {
|
||
161 setActiveClass(classId)
|
||
162 if (selectedId == null) return
|
||
|
||
● Now wire the sidebar to the id-explicit delete:
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Update(frontend/src/components/ReviewSidebar.jsx)
|
||
⎿ Added 1 line, removed 4 lines
|
||
83 <button
|
||
84 className="btn btn-danger"
|
||
85 aria-label="Delete shape"
|
||
86 - onClick={() => {
|
||
87 - setSelectedId(item.id)
|
||
88 - removeSelected()
|
||
89 - }}
|
||
86 + onClick={() => removeAnnotation(item.id)}
|
||
87 >
|
||
88 <TrashIcon size={13} />
|
||
89 </button>
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Update(frontend/src/components/ReviewSidebar.jsx)
|
||
⎿ Added 1 line, removed 1 line
|
||
11 annotations,
|
||
12 selectedId,
|
||
13 setSelectedId,
|
||
14 - removeSelected,
|
||
14 + removeAnnotation,
|
||
15 project,
|
||
16 jumpToNextAnnotated,
|
||
17 batchAnnotationCount,
|
||
|
||
● Update(frontend/src/pages/ReviewPage.jsx)
|
||
⎿ Added 1 line, removed 1 line
|
||
430 annotations={annotations}
|
||
431 selectedId={selectedId}
|
||
432 setSelectedId={setSelectedId}
|
||
433 - removeSelected={removeSelected}
|
||
433 + removeAnnotation={removeAnnotation}
|
||
434 project={project}
|
||
435 jumpToNextAnnotated={jumpToNextAnnotated}
|
||
436 batchAnnotationCount={batch?.annotation_count ?? 0}
|
||
|
||
● Now verify — imports, syntax, and the split logic:
|
||
|
||
Ran 3 shell commands (ctrl+o to expand)
|
||
|
||
● Done with steps 1–7. Both check out: backend imports clean, frontend builds.
|
||
|
||
What changed
|
||
|
||
Backend
|
||
- jobs.py — reverted to the single-worker queue; GPU_JOB_TYPES back to
|
||
("autolabel", "train"). Dead _start_job removed. The file is now
|
||
byte-identical to HEAD.
|
||
- api/batches.py — preview and /api/sam3/playground-test now take gpu_lock
|
||
with a 20 s timeout and return 409 with "GPU is busy with a X job", matching
|
||
review.assist.
|
||
- dataset.py:
|
||
- sync_labels filters to review_status='approved' and no longer takes
|
||
selected_class_ids — re-running auto-annotate on a merged batch can no longer
|
||
push unreviewed output into the master dataset.
|
||
- Class-filtered runs build a selected/ tree (symlinked images + remapped
|
||
labels). Master labels stay canonical.
|
||
- _next_split → split_for(project, batch, stem, val_every), hash-derived.
|
||
Verified deterministic, 18.8% val at val_every=5, val_every=0 → all train.
|
||
- All label/yaml writes go through _write_atomic (temp + os.replace).
|
||
- Cancelled merge leaves the batch approved and logs why, instead of marking
|
||
it merged and stranding the frames forever.
|
||
- nc now counts target_classes, not the full list.
|
||
- _require_val refuses an empty val split — gated behind require_val=True,
|
||
passed only from training.py.
|
||
- batches.py — delete() unlinks the batch's dataset_items files from the
|
||
master dataset before dropping the rows.
|
||
- dataset.approve() — rejects when a merge for that batch is already queued or
|
||
running.
|
||
|
||
Frontend
|
||
- removeSelected split into removeAnnotation(id) + a thin wrapper; sidebar
|
||
trash calls it with an explicit id. The stale-closure wrong-shape delete is
|
||
gone.
|
||
- Keydown handler early-returns on ctrlKey || metaKey || altKey.
|
||
- updateShape uses the geometry passed on commit (Alt-click delete-vertex now
|
||
actually works) and rolls back local state on a failed PATCH, as does
|
||
removeAnnotation.
|
||
|
||
On your split migration: no backfill was needed. dataset_items stores each
|
||
frame's split as a row and the merge loop skips any frame that already has
|
||
one, so existing assignments are preserved untouched — the hash only decides
|
||
new merges. Your existing base-vs-new comparison stays valid.
|
||
|
||
Two caveats, stated plainly. I verified imports, the split function's
|
||
behaviour, and a clean frontend build — I did not run a live end-to-end round,
|
||
so the merge/train paths are unexercised. And data.yaml now always says val:
|
||
images/val; if a project has merged data but zero val images, training raises
|
||
instead of silently validating on train. That's the intended behaviour, but it
|
||
will surface as a new error on any project in that state.
|
||
|
||
Not touched from the report: the custom_model_path critical (C3 — needs the
|
||
opaque-staging-id refactor across inspect-model, autolabel.start and
|
||
LibraryPage), the staged .pt leaks, library.py's unbounded ffmpeg threads,
|
||
datasetSummary returning every annotation, and the ops items (.dockerignore,
|
||
untracking docker-compose.override.yml, .env.example, pinning requirements).
|
||
Say the word on which to take next — and separately, tell me how slow
|
||
autolabel actually is and on what hardware, and I'll look at making a single
|
||
job faster.
|
||
|
||
✻ Sautéed for 6m 28s
|
||
|
||
※ recap: Goal is making your reTraining app reliable end-to-end; I audited it and fixed the seven critical issues in the job queue, dataset split, and review editor. Next up is your call on which remaining items to tackle. (disable recaps in /config)
|
||
|
||
❯ can you understand the data preparation pages and analyze what works and what
|
||
is not, what we should do to this pages, for this pages purposes is based on
|
||
the sam3 result, i can filter which is good for send to train and which to
|
||
ignore, but the ignored one is we can choose to just throw away or store them
|
||
as a different class, for example full-sack and half-sack. so the training
|
||
section can pick which they want to train based on that data preparation
|
||
filtering. do you understand my point?. ask me
|
||
|
||
● Let me read the page and its backend before I say I understand it.
|
||
|
||
Read 1 file (ctrl+o to expand)
|
||
● Yes, I understand the intent. Let me state it back, then flag one problem,
|
||
then ask.
|
||
|
||
What you want: Data Prep is the triage gate between SAM3's raw output and
|
||
training. You look at what SAM3 produced, filter it, and each shape gets one
|
||
of three fates — keep (goes to training), discard (thrown away), or relabel
|
||
(kept, but as a different class, e.g. the rejects become half-sack while the
|
||
keepers stay full-sack). Then Models/Training picks which of those classes it
|
||
actually trains on.
|
||
|
||
What the page does today: nothing of that. It is a read-only analytics
|
||
dashboard. minSizePct/maxSizePct are local React state
|
||
(DataPrepPage.jsx:10-11) that are never sent to the server — no apply button,
|
||
no persistence, no mutation. Reload the page and your filter is gone. The
|
||
class-distribution counts and the red/green box preview recompute live, but
|
||
they're a visualization of a filter that exists only in the browser.
|
||
|
||
What already exists that you can build on: training accepts selected_class_ids
|
||
(training.py:106 → write_data_yaml), which after this morning's fix
|
||
materializes a filtered selected/ tree. So "training picks which classes to
|
||
train" is already wired — the missing half is the mechanism that assigns
|
||
shapes to those classes in the first place.
|
||
|
||
One concern, and I think it matters. The only filter axis on this page is box
|
||
area as a percentage of the frame. Area is a proxy for distance from camera,
|
||
not for how full the sack is. A full sack at the far end of the conveyor and a
|
||
half sack up close can have identical area_pct. If full-vs-half is genuinely
|
||
the distinction you want, area alone will mislabel a lot of them. SAM3's
|
||
confidence score is the better "is this detection any good" signal — and note
|
||
it isn't even in the payload today (dataset.py sends id, frame_id, class_id,
|
||
area_pct, box — no score). Aspect ratio might separate full from half better
|
||
than area does. So I'd like to know what you're actually seeing in the data.
|
||
|
||
Questions:
|
||
|
||
● User answered Claude's questions:
|
||
⎿ · When you look at a SAM3 result and decide "this one is bad" or "this one
|
||
is half-sack", what are you actually judging it on? → SAM3 confidence
|
||
score, Box size / area, I have to look at the image, or anything you
|
||
recommend
|
||
· When you "relabel the rejects as half-sack", should that permanently
|
||
change the annotations, or be a reusable rule? → Saved filter rule / preset
|
||
just like you recommend. every section have their own rule, so if i
|
||
change in data prep section, the batches model doesnt have effect, but the
|
||
next step have
|
||
· Where should this triage happen in the loop? → Per batch, before merge
|
||
· What should "discard" actually do to a shape? → Mark ignored, decide
|
||
later
|
||
|
||
● Good — that clarifies a lot, but three things are still genuinely ambiguous
|
||
and they change the architecture, so let me pin them down.
|
||
|
||
Note one tension in your answers I want to resolve: you want triage per batch
|
||
before merge, but you also want a data-prep rule change to affect the next
|
||
step. If the rule is baked in at merge time, then editing it later can't
|
||
change anything downstream — the data is already written. So I need to know
|
||
when the rule actually gets evaluated.
|
||
|
||
Also, your "mark ignored" answer leaves the label-integrity trap open: if an
|
||
ignored box stays on an image that still goes to training, you're teaching the
|
||
model that a sack is background. That has to resolve somewhere.
|
||
|
||
● User answered Claude's questions:
|
||
⎿ · When is a data-prep rule actually evaluated? → Live, at training time
|
||
(Recommended)
|
||
· An 'ignored' box sits on an image. What happens to that image at training
|
||
time? → Exclude the whole image (Recommended)
|
||
· You said you often have to look at the image. How do manual decisions
|
||
coexist with rules? → Rules propose, you confirm in bulk
|
||
· For finding the full-vs-half boundary, what should I put on the page? →
|
||
Both
|
||
|
||
● That resolves it into a coherent design. One consequence you should see before
|
||
I write anything, because it touches your core invariant:
|
||
|
||
Live rule evaluation moves the val set. If a rule change excludes images, the
|
||
val split changes between runs — so a base-vs-new mAP from last week isn't
|
||
comparable to one from today. Both models in a single comparison are fine
|
||
(same data.yaml), but across rule edits the number silently shifts meaning.
|
||
This is fixable — stamp the rule-set version onto each training run and
|
||
refuse/warn on cross-version comparison — but it needs to be in the design,
|
||
not discovered later.
|
||
|
||
Per CLAUDE.md §5 I'm not writing feature code before a REQ covers it, so
|
||
here's the proposal:
|
||
|
||
● Write(docs/proposal-dataprep-triage.md)
|
||
⎿ Wrote 131 lines to docs/proposal-dataprep-triage.md
|
||
1 # Proposal — Data Prep as a triage stage
|
||
2
|
||
into `./requirements.md`
|
||
4 and the schema/API sections into `./design.md`.
|
||
5
|
||
7
|
||
8 - Signals: SAM3 confidence score, box area, aspect ratio, plus v
|
||
isual judgement.
|
||
orward only.
|
||
10 - Triage happens **per batch**, on SAM3 output.
|
||
… +175 lines (ctrl+o to expand)
|
||
|
||
● Written to docs/proposal-dataprep-triage.md. The short version:
|
||
|
||
The model is a three-way resolve, evaluated live: manual override > first
|
||
matching rule > default keep, producing keep / ignore / reclass → class.
|
||
annotations.class_id is never rewritten — SAM3's original output stays intact,
|
||
which is exactly what makes rules re-cuttable forever. Resolution happens in
|
||
one place only: dataset._build_selected_tree, which after this morning's fix
|
||
is already the thing that materializes what training sees.
|
||
|
||
Nine requirements (REQ-100…108), two new tables (triage_rules,
|
||
annotation_overrides), a rule_version column on training runs, six endpoints,
|
||
and a /crop endpoint to feed the thumbnail grid.
|
||
|
||
Three things worth your attention before you approve:
|
||
|
||
1. Rules are per-project, scoped to the dataprep stage — not per-batch. You
|
||
said triage happens per batch, and the UI is per batch (you look at one
|
||
batch's SAM3 output at a time). But a rule like "area < 5% → half-sack" is a
|
||
judgment about your domain, not about one batch, so it lives on the project
|
||
and applies everywhere. If you actually want per-batch rules, say so — it's a
|
||
different schema.
|
||
2. Score is already stored on annotations (ReviewSidebar renders item.score),
|
||
it's just missing from the Data Prep payload. Aspect ratio is free from the
|
||
box. So no new inference is needed to get all four signals — just plumbing.
|
||
3. REQ-104 vs REQ-107 is the real tension. Excluding whole images is the right
|
||
call for label integrity, but it makes your val set a function of your rules.
|
||
REQ-107 (stamp the rule version on each run, refuse cross-version comparison)
|
||
keeps that honest instead of silent. If it turns out to fire on every edit
|
||
and becomes noise, the fallback is to freeze val membership before rule
|
||
evaluation and let rules affect train only. I'd rather start strict and loosen
|
||
it.
|
||
|
||
Approve the REQs and I'll move them into requirements.md + design.md and start
|
||
building. If you'd rather see it working before committing to the
|
||
requirements, I can build the resolver and the scatter/crop UI against one
|
||
batch as a spike first — but say which, because CLAUDE.md says docs lead code
|
||
and I don't want to violate that silently.
|
||
|
||
✻ Baked for 2m 21s
|
||
|
||
❯ id rather see it working before commiting, i approve the plan
|
||
|
||
● Approved — building the spike. Docs stay as the proposal until you've seen it
|
||
run.
|
||
|
||
Read 1 file, ran 2 shell commands (ctrl+o to expand)
|
||
|
||
● Schema first:
|
||
|
||
● Update(backend/db.py)
|
||
⎿ Added 22 lines
|
||
125 finished_at REAL
|
||
126 )
|
||
127 """,
|
||
128 + """
|
||
129 + CREATE TABLE IF NOT EXISTS triage_rules (
|
||
130 + id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||
131 + project_id INTEGER NOT NULL REFERENCES projects(id)
|
||
+ON DELETE CASCADE,
|
||
132 + stage TEXT NOT NULL DEFAULT 'dataprep',
|
||
133 + position INTEGER NOT NULL,
|
||
134 + name TEXT NOT NULL,
|
||
135 + predicate TEXT NOT NULL,
|
||
136 + action TEXT NOT NULL CHECK (action IN ('keep', '
|
||
+ignore', 'reclass')),
|
||
137 + target_class INTEGER,
|
||
138 + created_at REAL NOT NULL
|
||
139 + )
|
||
140 + """,
|
||
141 + """
|
||
142 + CREATE TABLE IF NOT EXISTS annotation_overrides (
|
||
143 + annotation_id INTEGER PRIMARY KEY REFERENCES annotatio
|
||
+ns(id) ON DELETE CASCADE,
|
||
144 + verdict TEXT NOT NULL CHECK (verdict IN ('keep',
|
||
+ 'ignore', 'reclass')),
|
||
145 + target_class INTEGER,
|
||
146 + decided_at REAL NOT NULL
|
||
147 + )
|
||
148 + """,
|
||
149 "CREATE INDEX IF NOT EXISTS idx_frames_batch ON frames(bat
|
||
ch_id, idx)",
|
||
150 + "CREATE INDEX IF NOT EXISTS idx_triage_rules_project ON tr
|
||
+iage_rules(project_id, stage, position)",
|
||
151 "CREATE INDEX IF NOT EXISTS idx_annotations_frame ON annot
|
||
ations(frame_id)",
|
||
152 "CREATE INDEX IF NOT EXISTS idx_batches_project ON batches
|
||
(project_id)",
|
||
153 "CREATE INDEX IF NOT EXISTS idx_jobs_project ON jobs(proje
|
||
ct_id, created_at)",
|
||
|
||
● Update(backend/db.py)
|
||
⎿ Added 5 lines
|
||
189 cur.execute("ALTER TABLE projects ADD COLUMN secon
|
||
dary_model_name TEXT")
|
||
190 if "secondary_model_classes" not in cols:
|
||
191 cur.execute("ALTER TABLE projects ADD COLUMN secon
|
||
dary_model_classes TEXT")
|
||
192 + # REQ-107: what rule set a run's numbers were measured
|
||
+ under.
|
||
193 + cur.execute("PRAGMA table_info(model_versions)")
|
||
194 + version_cols = [column[1] for column in cur.fetchall()
|
||
+]
|
||
195 + if "rule_version" not in version_cols:
|
||
196 + cur.execute("ALTER TABLE model_versions ADD COLUMN
|
||
+ rule_version TEXT")
|
||
197
|
||
198
|
||
199 def healthy() -> bool:
|
||
|
||
● Now the resolver — the core of the whole feature:
|
||
|
||
● Write(backend/triage.py)
|
||
⎿ Wrote 267 lines to backend/triage.py
|
||
1 """Triage: deciding what each SAM3 shape is actually worth (REQ-1
|
||
00…108).
|
||
3 A shape is never rewritten. Its verdict is *resolved* every time
|
||
it is needed:
|
||
4
|
||
6
|
||
7 so `annotations.class_id` keeps whatever SAM3 said, and any rule
|
||
can be re-cut
|
||
are evaluated
|
||
9 at training time rather than baked in at merge (REQ-102).
|
||
… +305 lines (ctrl+o to expand)
|
||
|
||
● Now wire the resolver into what training actually sees:
|
||
|
||
Read 1 file (ctrl+o to expand)
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 36 lines, removed 10 lines
|
||
136
|
||
137
|
||
138 def _build_selected_tree(root: str, rows: list, class_map: Opt
|
||
ional[dict],
|
||
139 - label_type: str) -> tuple:
|
||
139 + label_type: str, resolver=None) -> tu
|
||
+ple:
|
||
140 """Materialise a filtered view of the dataset under `selec
|
||
ted/`.
|
||
141
|
||
142 The master labels stay canonical. A per-run class filter r
|
||
enumbers ids into a
|
||
143 contiguous 0..k-1 space, which contradicts `project_classe
|
||
s`; writing that
|
||
144 into the shared label files would leave the master dataset
|
||
in the index space
|
||
145 of whichever training run happened to go last.
|
||
146 +
|
||
147 + This is also the one place triage rules are applied (REQ-1
|
||
+02): resolve every
|
||
148 + shape, drop the images carrying an `ignore`, and emit the
|
||
+rest under their
|
||
149 + effective class. Because it happens here and not at merge
|
||
+time, editing a
|
||
150 + rule re-cuts the dataset on the next run without touching
|
||
+what was merged.
|
||
151 """
|
||
152 selected_root = os.path.join(root, "selected")
|
||
153 if os.path.isdir(selected_root):
|
||
154 shutil.rmtree(selected_root)
|
||
155 listed = {"train": [], "val": []}
|
||
156 + excluded = 0
|
||
157 for image_rel, label_rel, frame_id, split in rows:
|
||
158 + annotations = review.listing(frame_id)
|
||
159 +
|
||
160 + if resolver is not None:
|
||
161 + resolved = []
|
||
162 + skip_image = False
|
||
163 + for item in annotations:
|
||
164 + shape = {"id": item["id"], "class_id": item["c
|
||
+lass_id"],
|
||
165 + "score": float(item.get("score") or 1
|
||
+.0),
|
||
166 + **triage.metrics(item["geometry"])}
|
||
167 + effective = resolver.effective_class(shape)
|
||
168 + if effective is None:
|
||
169 + # REQ-104: a known object left unlabeled w
|
||
+ould teach the
|
||
170 + # model it is background, so the image goe
|
||
+s too.
|
||
171 + skip_image = True
|
||
172 + break
|
||
173 + resolved.append({**item, "class_id": effective
|
||
+})
|
||
174 + if skip_image:
|
||
175 + excluded += 1
|
||
176 + continue
|
||
177 + annotations = resolved
|
||
178 +
|
||
179 + if class_map is not None:
|
||
180 + annotations = [a for a in annotations if a["class_
|
||
+id"] in class_map]
|
||
181 + lines = [_label_line(class_map[a["class_id"]] if class
|
||
+_map is not None else a["class_id"],
|
||
182 + a["geometry"], label_type)
|
||
183 + for a in annotations]
|
||
184 +
|
||
185 stem = os.path.basename(image_rel)
|
||
186 image_dst = os.path.join(selected_root, "images", spli
|
||
t, stem)
|
||
187 label_dst = os.path.join(selected_root, "labels", spli
|
||
t,
|
||
188 os.path.splitext(stem)[0] + "
|
||
.txt")
|
||
189 os.makedirs(os.path.dirname(image_dst), exist_ok=True)
|
||
190 os.makedirs(os.path.dirname(label_dst), exist_ok=True)
|
||
158 - source = os.path.join(root, image_rel)
|
||
191 if not os.path.exists(image_dst):
|
||
160 - os.symlink(source, image_dst)
|
||
161 -
|
||
162 - annotations = review.listing(frame_id)
|
||
163 - if class_map is not None:
|
||
164 - annotations = [a for a in annotations if a["class_
|
||
-id"] in class_map]
|
||
165 - lines = [_label_line(class_map[a["class_id"]] if class
|
||
-_map is not None else a["class_id"],
|
||
166 - a["geometry"], label_type)
|
||
167 - for a in annotations]
|
||
192 + os.symlink(os.path.join(root, image_rel), image_ds
|
||
+t)
|
||
193 _write_atomic(label_dst, "\n".join(lines) + ("\n" if l
|
||
ines else ""))
|
||
194 listed[split].append(image_dst)
|
||
195 + listed["excluded"] = excluded
|
||
196 return selected_root, listed
|
||
197
|
||
198
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 6 lines, removed 1 line
|
||
213
|
||
214 names = ", ".join(f"'{item['name']}'" for item in target_c
|
||
lasses)
|
||
215
|
||
216 - if batch_ids or class_map is not None:
|
||
216 + # Rules re-cut the dataset on every run (REQ-102), so thei
|
||
+r presence alone is
|
||
217 + # enough to need the `selected/` tree — the master labels
|
||
+must stay canonical.
|
||
218 + resolver = triage.Resolver(project["id"])
|
||
219 + has_rules = bool(resolver.rules or resolver.overrides)
|
||
220 +
|
||
221 + if batch_ids or class_map is not None or has_rules:
|
||
222 with db.cursor() as cur:
|
||
223 where = "d.project_id = ?"
|
||
224 args = [project["id"]]
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 3 lines, removed 1 line
|
||
233 )
|
||
234 rows = cur.fetchall()
|
||
235
|
||
236 - selected_root, listed = _build_selected_tree(root, row
|
||
-s, class_map, project["label_type"])
|
||
236 + selected_root, listed = _build_selected_tree(
|
||
237 + root, rows, class_map, project["label_type"],
|
||
238 + resolver=resolver if has_rules else None)
|
||
239 if require_val:
|
||
240 _require_val(len(listed["val"]), "the selected bat
|
||
ches")
|
||
241
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 1 line, removed 1 line
|
||
17 import time
|
||
18 from typing import List, Optional
|
||
19
|
||
20 -from backend import batches, config, db, jobs, projects, review
|
||
20 +from backend import batches, config, db, jobs, projects, review
|
||
+, triage
|
||
21
|
||
22
|
||
23 class DatasetError(Exception):
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
● Now the API routes and crop endpoint:
|
||
|
||
● Write(backend/api/triage.py)
|
||
⎿ Wrote 118 lines to backend/api/triage.py
|
||
1 """Triage routes: rules, hand overrides, and the per-batch shape
|
||
view (REQ-100…108)."""
|
||
2
|
||
3 import io
|
||
4 import os
|
||
5
|
||
6 from fastapi import APIRouter, HTTPException
|
||
7 from fastapi.responses import StreamingResponse
|
||
8 from pydantic import BaseModel
|
||
9 from typing import List, Optional
|
||
… +115 lines (ctrl+o to expand)
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Good. Registering the router:
|
||
|
||
● Update(backend/main.py)
|
||
⎿ Added 1 line, removed 1 line
|
||
15 from fastapi.middleware.cors import CORSMiddleware
|
||
16
|
||
17 from backend import config, db, jobs
|
||
18 -from backend.api import batches, jobs as job_routes, models, pr
|
||
-ojects, review
|
||
18 +from backend.api import batches, jobs as job_routes, models, pr
|
||
+ojects, review, triage
|
||
19
|
||
20
|
||
21 @asynccontextmanager
|
||
|
||
● Update(backend/main.py)
|
||
⎿ Added 1 line
|
||
43 app.include_router(review.router)
|
||
44 app.include_router(models.router)
|
||
45 app.include_router(job_routes.router)
|
||
46 +app.include_router(triage.router)
|
||
47
|
||
48
|
||
49 @app.get("/api/health")
|
||
|
||
● Let me verify the resolver logic end-to-end against a real scenario before
|
||
building UI on top of it:
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Rules resolve correctly — the FK error is my test using fake annotation ids.
|
||
Let me redo with real rows:
|
||
|
||
Ran 3 shell commands (ctrl+o to expand)
|
||
|
||
● Resolver verified: precedence, first-match-wins, manual override, version
|
||
changes. Now the UI. Let me check the api client conventions:
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Update(frontend/src/api.js)
|
||
⎿ Added 15 lines
|
||
116 approveAllBatchFrames: (batchId) => request(`/batches/${batc
|
||
hId}/approve-all`, { method: 'POST' }),
|
||
117 approveBatch: (batchId) => request(`/batches/${batchId}/appr
|
||
ove`, { method: 'POST' }),
|
||
118 datasetSummary: (projectId) => request(`/projects/${projectI
|
||
d}/dataset`),
|
||
119 +
|
||
120 + // Triage (REQ-100…108)
|
||
121 + triageRules: (projectId) => request(`/projects/${projectId}/
|
||
+triage/rules`),
|
||
122 + saveTriageRules: (projectId, rules) =>
|
||
123 + request(`/projects/${projectId}/triage/rules`, { method: '
|
||
+PUT', body: { rules } }),
|
||
124 + triageShapes: (batchId) => request(`/batches/${batchId}/tria
|
||
+ge/shapes`),
|
||
125 + triagePreview: (projectId) => request(`/projects/${projectId
|
||
+}/triage/preview`),
|
||
126 + setTriageOverrides: (annotationIds, verdict, targetClass = n
|
||
+ull) =>
|
||
127 + request('/triage/overrides', {
|
||
128 + method: 'POST',
|
||
129 + body: { annotation_ids: annotationIds, verdict, target_c
|
||
+lass: targetClass },
|
||
130 + }),
|
||
131 + clearTriageOverrides: (annotationIds) =>
|
||
132 + request('/triage/overrides', { method: 'DELETE', body: { a
|
||
+nnotation_ids: annotationIds } }),
|
||
133 + cropUrl: (annotationId) => `/api/annotations/${annotationId}
|
||
+/crop`,
|
||
134 datasetDownloadUrl: (projectId) => `/api/projects/${projectI
|
||
d}/dataset/download`,
|
||
135
|
||
136 hardware: () => request('/hardware'),
|
||
|
||
● Now the scatter plot with drag-selection:
|
||
|
||
● Write(frontend/src/components/TriageScatter.jsx)
|
||
⎿ Wrote 165 lines to frontend/src/components/TriageScatter.jsx
|
||
1 import React, { useMemo, useRef, useState } from 'react'
|
||
2
|
||
EQ-106).
|
||
4 *
|
||
5 * Area is on a log scale: box sizes span three orders of magnitu
|
||
6 * linear axis every shape piles into the left edge where no boun
|
||
dary is
|
||
7 * visible. */
|
||
9 const VERDICT_COLOR = { keep: '#4ade80', ignore: '#f87171', recla
|
||
ss: '#c084fc' }
|
||
… +179 lines (ctrl+o to expand)
|
||
|
||
● Now the crop grid and rule editor:
|
||
|
||
● Write(frontend/src/components/TriageCropGrid.jsx)
|
||
⎿ Wrote 128 lines to frontend/src/components/TriageCropGrid.jsx
|
||
1 import React, { useState } from 'react'
|
||
2 import { api } from '../api'
|
||
3
|
||
ck?" means
|
||
5 * looking at the object, not at the frame it came from. Sorted s
|
||
o the shapes a
|
||
6 * rule is about to act on sit together. */
|
||
8 const VERDICT_COLOR = { keep: '#4ade80', ignore: '#f87171', recla
|
||
ss: '#c084fc' }
|
||
9 const PAGE = 120
|
||
… +144 lines (ctrl+o to expand)
|
||
|
||
● Write(frontend/src/components/TriageRules.jsx)
|
||
⎿ Wrote 177 lines to frontend/src/components/TriageRules.jsx
|
||
1 import React from 'react'
|
||
2 import { TrashIcon } from './Icons'
|
||
3
|
||
4 /* The rule list (REQ-100). Ordered, first match wins — so the o
|
||
rder is part of
|
||
6
|
||
7 const ACTIONS = [
|
||
8 { value: 'keep', label: 'Keep' },
|
||
9 { value: 'ignore', label: 'Ignore' },
|
||
10 { value: 'reclass', label: 'Reclass to' },
|
||
… +198 lines (ctrl+o to expand)
|
||
|
||
● Now the page itself, wiring it together:
|
||
|
||
Searched for 2 patterns (ctrl+o to expand)
|
||
|
||
● Write(frontend/src/pages/DataPrepPage.jsx)
|
||
⎿ Added 237 lines, removed 569 lines
|
||
1 -import React, { useCallback, useEffect, useState } from 'react
|
||
-'
|
||
2 -import { api, classColor } from '../api'
|
||
1 +import React, { useCallback, useEffect, useMemo, useState } fr
|
||
+om 'react'
|
||
2 +import { api } from '../api'
|
||
3
|
||
4 -import { AlertIcon, BarChartIcon, DatabaseIcon, RocketIcon } f
|
||
-rom '../components/Icons'
|
||
4 +import { AlertIcon, DatabaseIcon, RocketIcon, SlidersIcon } fr
|
||
+om '../components/Icons'
|
||
5 +import TriageCropGrid from '../components/TriageCropGrid'
|
||
6 +import TriageRules from '../components/TriageRules'
|
||
7 +import TriageScatter from '../components/TriageScatter'
|
||
8
|
||
9 +/* Data Prep is the triage gate between SAM3's raw output and
|
||
+training
|
||
10 + * (REQ-100…108): look at what SAM3 produced for one batch, de
|
||
+cide what is worth
|
||
11 + * training on, and express the repeatable part of that decisi
|
||
+on as rules.
|
||
12 + *
|
||
13 + * Nothing here rewrites an annotation. Verdicts resolve as
|
||
14 + * `manual > first matching rule > keep`, and the dataset is r
|
||
+e-cut from them at
|
||
15 + * training time — so a rule stays editable forever. */
|
||
16 +
|
||
17 +const VERDICTS = [
|
||
18 + { value: 'keep', label: 'Keep', color: '#4ade80' },
|
||
19 + { value: 'ignore', label: 'Ignore', color: '#f87171' },
|
||
20 + { value: 'reclass', label: 'Reclass', color: '#c084fc' },
|
||
21 +]
|
||
22 +
|
||
23 export default function DataPrepPage({ projectId, onProject })
|
||
{
|
||
24 const [project, setProject] = useState(null)
|
||
8 - const [summary, setSummary] = useState(null)
|
||
9 - const [scaleMode, setScaleMode] = useState('normal') // 'nor
|
||
-mal' (Gaussian Bell Curve) | 'linear'
|
||
10 - const [minSizePct, setMinSizePct] = useState(0)
|
||
11 - const [maxSizePct, setMaxSizePct] = useState(100)
|
||
12 - const [previewFrameId, setPreviewFrameId] = useState(null)
|
||
13 - const [showFilteredOut, setShowFilteredOut] = useState(true)
|
||
25 + const [batches, setBatches] = useState([])
|
||
26 + const [batchId, setBatchId] = useState(null)
|
||
27 + const [view, setView] = useState(null)
|
||
28 + const [rules, setRules] = useState([])
|
||
29 + const [savedRules, setSavedRules] = useState([])
|
||
30 + const [preview, setPreview] = useState(null)
|
||
31 + const [selectedIds, setSelectedIds] = useState([])
|
||
32 + const [reclassTarget, setReclassTarget] = useState('')
|
||
33 + const [saving, setSaving] = useState(false)
|
||
34 + const [busy, setBusy] = useState(false)
|
||
35 const [error, setError] = useState('')
|
||
36
|
||
37 + useEffect(() => {
|
||
38 + let cancelled = false
|
||
39 + ;(async () => {
|
||
40 + try {
|
||
41 + const [loadedProject, loadedBatches, loadedRules] = aw
|
||
+ait Promise.all([
|
||
42 + api.getProject(projectId),
|
||
43 + api.listBatches(projectId),
|
||
44 + api.triageRules(projectId),
|
||
45 + ])
|
||
46 + if (cancelled) return
|
||
47 + setProject(loadedProject)
|
||
48 + onProject?.(loadedProject)
|
||
49 + const annotated = (loadedBatches.batches ?? loadedBatc
|
||
+hes).filter((b) => b.annotation_count > 0)
|
||
50 + setBatches(annotated)
|
||
51 + setBatchId((current) => current ?? annotated[0]?.id ??
|
||
+ null)
|
||
52 + setRules(loadedRules.rules)
|
||
53 + setSavedRules(loadedRules.rules)
|
||
54 + } catch (exc) {
|
||
55 + if (!cancelled) setError(exc.message)
|
||
56 + }
|
||
57 + })()
|
||
58 + return () => { cancelled = true }
|
||
59 + }, [projectId])
|
||
60
|
||
17 - const load = useCallback(async () => {
|
||
61 + const loadBatch = useCallback(async () => {
|
||
62 + if (!batchId) return
|
||
63 try {
|
||
19 - const [loadedProject, loadedSummary] = await Promise.all
|
||
-([
|
||
20 - api.getProject(projectId),
|
||
21 - api.datasetSummary(projectId),
|
||
64 + const [shapes, previewed] = await Promise.all([
|
||
65 + api.triageShapes(batchId),
|
||
66 + api.triagePreview(projectId),
|
||
67 ])
|
||
23 - setProject(loadedProject)
|
||
24 - onProject?.(loadedProject)
|
||
25 - setSummary(loadedSummary)
|
||
68 + setView(shapes)
|
||
69 + setPreview(previewed)
|
||
70 + setSelectedIds([])
|
||
71 } catch (exc) {
|
||
72 setError(exc.message)
|
||
73 }
|
||
29 - }, [projectId])
|
||
74 + }, [batchId, projectId])
|
||
75
|
||
31 - useEffect(() => {
|
||
32 - load()
|
||
33 - }, [load])
|
||
76 + useEffect(() => { loadBatch() }, [loadBatch])
|
||
77
|
||
35 - if (error && !project) {
|
||
36 - return <p className="error-banner"><AlertIcon size={14} />
|
||
- {error}</p>
|
||
37 - }
|
||
38 - if (!project || !summary) {
|
||
39 - return <p className="empty">Loading Data Preparation metad
|
||
-ata…</p>
|
||
40 - }
|
||
78 + const shapes = view?.shapes ?? []
|
||
79 + const dirty = JSON.stringify(rules) !== JSON.stringify(saved
|
||
+Rules)
|
||
80
|
||
42 - const totalFrames = (summary.splits?.train || 0) + (summary.
|
||
-splits?.val || 0)
|
||
43 - const mergedBatches = summary.batches || []
|
||
44 - const shapeDist = summary.shape_distribution || { total_shap
|
||
-es: 0, histogram: [], log_histogram: [], shapes: [] }
|
||
81 + const counts = useMemo(() => {
|
||
82 + const tally = { keep: 0, ignore: 0, reclass: 0, manual: 0
|
||
+}
|
||
83 + shapes.forEach((shape) => {
|
||
84 + tally[shape.verdict] += 1
|
||
85 + if (shape.source === 'manual') tally.manual += 1
|
||
86 + })
|
||
87 + return tally
|
||
88 + }, [shapes])
|
||
89
|
||
46 - const linearHistogram = shapeDist.histogram || []
|
||
47 - const logHistogram = shapeDist.log_histogram || []
|
||
48 - const normalParams = shapeDist.normal_params || { mu_log: 0,
|
||
- sigma_log: 1, median_area_pct: 0 }
|
||
49 - const rawShapes = shapeDist.shapes || []
|
||
90 + async function saveRules() {
|
||
91 + setSaving(true)
|
||
92 + try {
|
||
93 + const stored = await api.saveTriageRules(projectId, rule
|
||
+s)
|
||
94 + setSavedRules(stored.rules)
|
||
95 + setRules(stored.rules)
|
||
96 + await loadBatch()
|
||
97 + } catch (exc) {
|
||
98 + setError(exc.message)
|
||
99 + } finally {
|
||
100 + setSaving(false)
|
||
101 + }
|
||
102 + }
|
||
103
|
||
51 - // Active histogram data based on selected view mode
|
||
52 - const activeHistogram = scaleMode === 'normal' ? logHistogra
|
||
-m : linearHistogram
|
||
53 - const maxCount = Math.max(1, ...activeHistogram.map((h) => h
|
||
-.count || 0))
|
||
54 - const maxDensity = Math.max(0.001, ...logHistogram.map((h) =
|
||
-> h.normal_density || 0))
|
||
104 + async function applyVerdict(verdict) {
|
||
105 + if (selectedIds.length === 0) return
|
||
106 + if (verdict === 'reclass' && reclassTarget === '') {
|
||
107 + setError('Pick the class to reclass into first')
|
||
108 + return
|
||
109 + }
|
||
110 + setBusy(true)
|
||
111 + try {
|
||
112 + await api.setTriageOverrides(selectedIds, verdict, verdi
|
||
+ct === 'reclass' ? Number(reclassTarget) : null)
|
||
113 + await loadBatch()
|
||
114 + } catch (exc) {
|
||
115 + setError(exc.message)
|
||
116 + } finally {
|
||
117 + setBusy(false)
|
||
118 + }
|
||
119 + }
|
||
120
|
||
56 - // Filter shapes based on continuous min/max range sliders
|
||
57 - const isFiltering = minSizePct > 0 || maxSizePct < 100
|
||
58 - const filteredShapes = rawShapes.filter(
|
||
59 - (s) => s.area_pct >= minSizePct && s.area_pct <= maxSizePc
|
||
-t
|
||
60 - )
|
||
121 + async function clearDecisions() {
|
||
122 + if (selectedIds.length === 0) return
|
||
123 + setBusy(true)
|
||
124 + try {
|
||
125 + await api.clearTriageOverrides(selectedIds)
|
||
126 + await loadBatch()
|
||
127 + } catch (exc) {
|
||
128 + setError(exc.message)
|
||
129 + } finally {
|
||
130 + setBusy(false)
|
||
131 + }
|
||
132 + }
|
||
133
|
||
62 - // Compute class counts for filtered shapes
|
||
63 - const filteredClassCounts = {}
|
||
64 - filteredShapes.forEach((s) => {
|
||
65 - filteredClassCounts[s.class_id] = (filteredClassCounts[s.c
|
||
-lass_id] || 0) + 1
|
||
66 - })
|
||
134 + if (error && !project) {
|
||
135 + return <p className="error-banner"><AlertIcon size={14} />
|
||
+ {error}</p>
|
||
136 + }
|
||
137 + if (!project) return <p className="empty">Loading Data Prepa
|
||
+ration…</p>
|
||
138
|
||
68 - // Find top frames sorted by TOTAL annotation density
|
||
69 -
|
||
70 - const frameTotalCounts = {}
|
||
71 - rawShapes.forEach((s) => {
|
||
72 - frameTotalCounts[s.frame_id] = (frameTotalCounts[s.frame_i
|
||
-d] || 0) + 1
|
||
73 - })
|
||
74 -
|
||
75 - const topDenseFrames = Object.entries(frameTotalCounts)
|
||
76 - .sort((a, b) => b[1] - a[1])
|
||
77 - .map(([fid, count]) => ({ frame_id: Number(fid), total_sha
|
||
-pes: count }))
|
||
78 -
|
||
79 - const defaultPreviewFrameId = topDenseFrames[0]?.frame_id ||
|
||
- rawShapes[0]?.frame_id || null
|
||
80 - const activePreviewFrameId = previewFrameId || defaultPrevie
|
||
-wFrameId
|
||
81 -
|
||
82 - // Get shapes on the active preview frame
|
||
83 - const allFrameShapes = rawShapes.filter((s) => s.frame_id ==
|
||
-= activePreviewFrameId)
|
||
84 - const keptShapesOnFrame = allFrameShapes.filter((s) => s.are
|
||
-a_pct >= minSizePct && s.area_pct <= maxSizePct)
|
||
85 - const filteredOutShapesOnFrame = allFrameShapes.filter((s) =
|
||
-> s.area_pct < minSizePct || s.area_pct > maxSizePct)
|
||
86 -
|
||
87 -
|
||
88 - // SVG Graph Dimensions
|
||
89 - const graphWidth = 800
|
||
90 - const graphHeight = 180
|
||
91 - const padding = 24
|
||
92 - const usableW = graphWidth - padding * 2
|
||
93 - const usableH = graphHeight - padding * 2
|
||
94 -
|
||
95 - // Continuous Shape Count Line Coordinates
|
||
96 - const shapePoints = activeHistogram.map((h, i) => {
|
||
97 - const x = padding + (i / (activeHistogram.length - 1)) * u
|
||
-sableW
|
||
98 - const y = graphHeight - padding - (h.count / maxCount) * u
|
||
-sableH
|
||
99 - return { x, y, count: h.count, label: scaleMode === 'norma
|
||
-l' ? `${h.center_pct}%` : `${h.pct}%` }
|
||
100 - })
|
||
101 -
|
||
102 - const shapeLineD = shapePoints.reduce(
|
||
103 - (acc, p, i) => (i === 0 ? `M ${p.x} ${p.y}` : `${acc} L ${
|
||
-p.x} ${p.y}`),
|
||
104 - ''
|
||
105 - )
|
||
106 - const shapeAreaD = `${shapeLineD} L ${shapePoints[shapePoint
|
||
-s.length - 1]?.x || graphWidth} ${graphHeight - padding} L ${p
|
||
-adding} ${graphHeight - padding} Z`
|
||
107 -
|
||
108 - // Theoretical Gaussian Normal Bell Curve Coordinates (Cyan
|
||
-Line)
|
||
109 - const normalPoints = logHistogram.map((h, i) => {
|
||
110 - const x = padding + (i / (logHistogram.length - 1)) * usab
|
||
-leW
|
||
111 - const y = graphHeight - padding - (h.normal_density / maxD
|
||
-ensity) * usableH
|
||
112 - return { x, y, density: h.normal_density }
|
||
113 - })
|
||
114 -
|
||
115 - const normalLineD = normalPoints.reduce(
|
||
116 - (acc, p, i) => (i === 0 ? `M ${p.x} ${p.y}` : `${acc} L ${
|
||
-p.x} ${p.y}`),
|
||
117 - ''
|
||
118 - )
|
||
119 - const normalAreaD = `${normalLineD} L ${normalPoints[normalP
|
||
-oints.length - 1]?.x || graphWidth} ${graphHeight - padding} L
|
||
- ${padding} ${graphHeight - padding} Z`
|
||
120 -
|
||
121 -
|
||
122 - // Active Filter Range Overlay
|
||
123 - const startX = padding + (minSizePct / 100) * usableW
|
||
124 - const endX = padding + (maxSizePct / 100) * usableW
|
||
125 -
|
||
139 return (
|
||
127 -
|
||
140 <>
|
||
129 - <div className="page-head" style={{ display: 'flex', jus
|
||
-tifyContent: 'space-between', alignItems: 'center' }}>
|
||
141 + <div className="page-head" style={{ display: 'flex', jus
|
||
+tifyContent: 'space-between', alignItems: 'center', gap: 16, f
|
||
+lexWrap: 'wrap' }}>
|
||
142 <div>
|
||
143 <h1>Data Preparation</h1>
|
||
132 - <p className="muted mono">{project.name} · Bounding
|
||
-Box Size Distribution & Preview</p>
|
||
144 + <p className="muted mono">{project.name} · triage SA
|
||
+M3 output before it trains</p>
|
||
145 </div>
|
||
146 <a
|
||
135 - className="btn btn-primary"
|
||
147 + className="btn"
|
||
148 href={api.datasetDownloadUrl(projectId)}
|
||
149 download
|
||
150 style={{ fontSize: '0.85rem', padding: '8px 16px', c
|
||
ursor: 'pointer', display: 'inline-flex', alignItems: 'center'
|
||
, gap: 6, borderRadius: 6 }}
|
||
151 >
|
||
140 - <DatabaseIcon size={16} /> Download Dataset (.zip)
|
||
152 + <DatabaseIcon size={16} /> Download dataset (.zip)
|
||
153 </a>
|
||
154 </div>
|
||
155
|
||
156 {error && (
|
||
157 <p className="error-banner" style={{ marginBottom: 16
|
||
}}>
|
||
158 <AlertIcon size={14} /> {error}
|
||
159 + <button type="button" className="btn" onClick={() =>
|
||
+ setError('')} style={{ marginLeft: 10, cursor: 'pointer', fon
|
||
+tSize: '0.75rem' }}>Dismiss</button>
|
||
160 </p>
|
||
161 )}
|
||
162
|
||
150 - {/* Dataset Summary Cards */}
|
||
151 - <div style={{ display: 'grid', gridTemplateColumns: 'rep
|
||
-eat(auto-fit, minmax(190px, 1fr))', gap: 16, marginBottom: 20
|
||
-}}>
|
||
152 - <div className="panel side-panel" style={{ padding: 16
|
||
- }}>
|
||
153 - <span style={{ fontSize: '0.78rem', color: '#a1a1aa'
|
||
-, textTransform: 'uppercase', letterSpacing: '0.05em' }}>Total
|
||
- Master Images</span>
|
||
154 - <div style={{ fontSize: '1.8rem', fontWeight: 700, c
|
||
-olor: '#f4f4f5', marginTop: 4 }}>{totalFrames}</div>
|
||
155 - <p className="hint" style={{ fontSize: '0.78rem', ma
|
||
-rgin: '4px 0 0 0' }}>Across all approved batches</p>
|
||
163 + {preview && (
|
||
164 + <div style={{ display: 'grid', gridTemplateColumns: 'r
|
||
+epeat(auto-fit, minmax(180px, 1fr))', gap: 14, marginBottom: 1
|
||
+8 }}>
|
||
165 + <Stat label="Trainable images" value={preview.traina
|
||
+ble_images} hint={`of ${preview.total_images} merged`} accent=
|
||
+"#4ade80" />
|
||
166 + <Stat label="Excluded images" value={preview.exclude
|
||
+d_images} hint="carry an ignored shape" accent="#f87171" />
|
||
167 + <Stat label="Reclassed shapes" value={preview.reclas
|
||
+s} hint="trained as another class" accent="#c084fc" />
|
||
168 + <Stat label="Rule version" value={preview.rule_versi
|
||
+on} hint="stamped on each run" accent="#38bdf8" mono />
|
||
169 </div>
|
||
170 + )}
|
||
171
|
||
158 - <div className="panel side-panel" style={{ padding: 16
|
||
-, borderLeft: '3px solid #22c55e' }}>
|
||
159 - <span style={{ fontSize: '0.78rem', color: '#a1a1aa'
|
||
-, textTransform: 'uppercase', letterSpacing: '0.05em' }}>Train
|
||
-ing Set (Train)</span>
|
||
160 - <div style={{ fontSize: '1.8rem', fontWeight: 700, c
|
||
-olor: '#4ade80', marginTop: 4 }}>{summary.splits?.train || 0}<
|
||
-/div>
|
||
161 - <p className="hint" style={{ fontSize: '0.78rem', ma
|
||
-rgin: '4px 0 0 0' }}>80% split for model training</p>
|
||
162 - </div>
|
||
172 + <TriageRules
|
||
173 + rules={rules}
|
||
174 + classes={project.classes}
|
||
175 + onChange={setRules}
|
||
176 + onSave={saveRules}
|
||
177 + saving={saving}
|
||
178 + dirty={dirty}
|
||
179 + />
|
||
180
|
||
164 - <div className="panel side-panel" style={{ padding: 16
|
||
-, borderLeft: '3px solid #38bdf8' }}>
|
||
165 - <span style={{ fontSize: '0.78rem', color: '#a1a1aa'
|
||
-, textTransform: 'uppercase', letterSpacing: '0.05em' }}>Valid
|
||
-ation Set (Val)</span>
|
||
166 - <div style={{ fontSize: '1.8rem', fontWeight: 700, c
|
||
-olor: '#38bdf8', marginTop: 4 }}>{summary.splits?.val || 0}</d
|
||
-iv>
|
||
167 - <p className="hint" style={{ fontSize: '0.78rem', ma
|
||
-rgin: '4px 0 0 0' }}>20% locked split for metrics</p>
|
||
168 - </div>
|
||
169 -
|
||
170 - <div className="panel side-panel" style={{ padding: 16
|
||
-, borderLeft: '3px solid #a855f7' }}>
|
||
171 - <span style={{ fontSize: '0.78rem', color: '#a1a1aa'
|
||
-, textTransform: 'uppercase', letterSpacing: '0.05em' }}>Total
|
||
- Bounding Boxes</span>
|
||
172 - <div style={{ fontSize: '1.8rem', fontWeight: 700, c
|
||
-olor: '#c084fc', marginTop: 4 }}>{shapeDist.total_shapes}</div
|
||
->
|
||
173 - <p className="hint" style={{ fontSize: '0.78rem', ma
|
||
-rgin: '4px 0 0 0' }}>Median Area: <strong>{normalParams.median
|
||
-_area_pct}%</strong></p>
|
||
174 - </div>
|
||
175 - </div>
|
||
176 -
|
||
177 - {/* Bounding Box Size Distribution Graph */}
|
||
178 - <div className="panel table-wrap" style={{ padding: 20,
|
||
-marginBottom: 20 }}>
|
||
179 - <div style={{ display: 'flex', justifyContent: 'space-
|
||
-between', alignItems: 'center', marginBottom: 14 }}>
|
||
180 - <div>
|
||
181 - <h2 style={{ fontSize: '1.05rem', margin: 0, displ
|
||
-ay: 'flex', alignItems: 'center', gap: 8 }}>
|
||
182 - <BarChartIcon size={18} /> Bounding Box Size Dis
|
||
-tribution Curve
|
||
183 - </h2>
|
||
184 - <p className="hint" style={{ fontSize: '0.8rem', m
|
||
-argin: '4px 0 0 0' }}>
|
||
185 - Log-Normal Bell Curve N(μ, σ²) showing shape siz
|
||
-e distribution.
|
||
186 - </p>
|
||
187 - </div>
|
||
188 -
|
||
189 - {/* Scale View Toggle */}
|
||
190 - <div style={{ display: 'flex', gap: 8, alignItems: '
|
||
-center' }}>
|
||
191 - <button
|
||
192 - type="button"
|
||
193 - className="tag"
|
||
194 - style={{
|
||
195 - cursor: 'pointer',
|
||
196 - padding: '6px 12px',
|
||
197 - fontSize: '0.8rem',
|
||
198 - background: scaleMode === 'normal' ? 'rgba(56,
|
||
- 189, 248, 0.2)' : 'rgba(255,255,255,0.05)',
|
||
199 - color: scaleMode === 'normal' ? '#38bdf8' : '#
|
||
-a1a1aa',
|
||
200 - border: scaleMode === 'normal' ? '1px solid rg
|
||
-ba(56, 189, 248, 0.5)' : '1px solid rgba(255,255,255,0.1)',
|
||
201 - borderRadius: 6
|
||
202 - }}
|
||
203 - onClick={() => setScaleMode('normal')}
|
||
181 + <div className="panel table-wrap" style={{ padding: 18,
|
||
+marginTop: 18 }}>
|
||
182 + <div style={{ display: 'flex', justifyContent: 'space-
|
||
+between', alignItems: 'center', gap: 12, flexWrap: 'wrap', mar
|
||
+ginBottom: 14 }}>
|
||
183 + <h2 style={{ fontSize: '1rem', margin: 0, display: '
|
||
+flex', alignItems: 'center', gap: 8 }}>
|
||
184 + <SlidersIcon size={17} /> Batch triage
|
||
185 + </h2>
|
||
186 + <div style={{ display: 'flex', gap: 10, alignItems:
|
||
+'center', flexWrap: 'wrap' }}>
|
||
187 + <label className="faint" style={{ fontSize: '0.8re
|
||
+m' }}>Batch:</label>
|
||
188 + <select
|
||
189 + value={batchId ?? ''}
|
||
190 + onChange={(event) => setBatchId(Number(event.tar
|
||
+get.value))}
|
||
191 + style={{ padding: '4px 10px', fontSize: '0.82rem
|
||
+', background: 'rgba(0,0,0,0.4)', color: '#f4f4f5', border: '1
|
||
+px solid rgba(255,255,255,0.15)', borderRadius: 6, cursor: 'po
|
||
+inter' }}
|
||
192 >
|
||
205 - Normal Bell Curve (Log Scale)
|
||
206 - </button>
|
||
207 - <button
|
||
208 - type="button"
|
||
209 - className="tag"
|
||
210 - style={{
|
||
211 - cursor: 'pointer',
|
||
212 - padding: '6px 12px',
|
||
213 - fontSize: '0.8rem',
|
||
214 - background: scaleMode === 'linear' ? 'rgba(192
|
||
-, 132, 252, 0.2)' : 'rgba(255,255,255,0.05)',
|
||
215 - color: scaleMode === 'linear' ? '#c084fc' : '#
|
||
-a1a1aa',
|
||
216 - border: scaleMode === 'linear' ? '1px solid rg
|
||
-ba(192, 132, 252, 0.5)' : '1px solid rgba(255,255,255,0.1)',
|
||
217 - borderRadius: 6
|
||
218 - }}
|
||
219 - onClick={() => setScaleMode('linear')}
|
||
220 - >
|
||
221 - Linear Scale (0% - 100%)
|
||
222 - </button>
|
||
223 - </div>
|
||
224 - </div>
|
||
225 -
|
||
226 - {/* Normal Distribution Stats Banner */}
|
||
227 - {scaleMode === 'normal' && (
|
||
228 - <div style={{ display: 'flex', gap: 16, padding: '8p
|
||
-x 14px', background: 'rgba(56, 189, 248, 0.08)', border: '1px
|
||
-solid rgba(56, 189, 248, 0.2)', borderRadius: 6, marginBottom:
|
||
- 14, fontSize: '0.8rem', color: '#e0f2fe', alignItems: 'center
|
||
-' }}>
|
||
229 - <span>Gaussian Normal Fit:</span>
|
||
230 - <span className="mono" style={{ color: '#38bdf8' }
|
||
-}>μ (Log Mean) = {normalParams.mu_log}</span>
|
||
231 - <span className="mono" style={{ color: '#38bdf8' }
|
||
-}>σ (Std Dev) = {normalParams.sigma_log}</span>
|
||
232 - <span className="mono" style={{ color: '#4ade80' }
|
||
-}>Median Box Area = {normalParams.median_area_pct}% of frame</
|
||
-span>
|
||
233 - </div>
|
||
234 - )}
|
||
235 -
|
||
236 - {/* SVG Normal Distribution Line Chart */}
|
||
237 - <div style={{ background: 'rgba(0,0,0,0.3)', borderRad
|
||
-ius: 8, padding: 12, border: '1px solid rgba(255,255,255,0.06)
|
||
-' }}>
|
||
238 - <svg viewBox={`0 0 ${graphWidth} ${graphHeight}`} st
|
||
-yle={{ width: '100%', height: 'auto', display: 'block' }}>
|
||
239 - <defs>
|
||
240 - <linearGradient id="purpleGradient" x1="0" y1="0
|
||
-" x2="0" y2="1">
|
||
241 - <stop offset="0%" stopColor="#c084fc" stopOpac
|
||
-ity="0.35" />
|
||
242 - <stop offset="100%" stopColor="#c084fc" stopOp
|
||
-acity="0.0" />
|
||
243 - </linearGradient>
|
||
244 - <linearGradient id="cyanGradient" x1="0" y1="0"
|
||
-x2="0" y2="1">
|
||
245 - <stop offset="0%" stopColor="#38bdf8" stopOpac
|
||
-ity="0.25" />
|
||
246 - <stop offset="100%" stopColor="#38bdf8" stopOp
|
||
-acity="0.0" />
|
||
247 - </linearGradient>
|
||
248 - </defs>
|
||
249 -
|
||
250 - {/* Active Range Overlay */}
|
||
251 - <rect
|
||
252 - x={startX}
|
||
253 - y={padding}
|
||
254 - width={Math.max(2, endX - startX)}
|
||
255 - height={usableH}
|
||
256 - fill="rgba(192, 132, 252, 0.15)"
|
||
257 - stroke="rgba(192, 132, 252, 0.4)"
|
||
258 - strokeWidth="1.5"
|
||
259 - />
|
||
260 -
|
||
261 - {/* Shape Distribution Fill & Line (Purple) */}
|
||
262 - <path d={shapeAreaD} fill="url(#purpleGradient)" /
|
||
->
|
||
263 - <path d={shapeLineD} fill="none" stroke="#c084fc"
|
||
-strokeWidth="2.5" strokeLinecap="round" strokeLinejoin="round"
|
||
- />
|
||
264 -
|
||
265 - {/* Gaussian Normal Bell Curve Fill & Line (Cyan)
|
||
-*/}
|
||
266 - {scaleMode === 'normal' && (
|
||
267 - <>
|
||
268 - <path d={normalAreaD} fill="url(#cyanGradient)
|
||
-" />
|
||
269 - <path d={normalLineD} fill="none" stroke="#38b
|
||
-df8" strokeWidth="3" strokeLinecap="round" strokeLinejoin="rou
|
||
-nd" />
|
||
270 - </>
|
||
271 - )}
|
||
272 -
|
||
273 - {/* Data Points */}
|
||
274 - {shapePoints.map((p, idx) => (
|
||
275 - <circle
|
||
276 - key={idx}
|
||
277 - cx={p.x}
|
||
278 - cy={p.y}
|
||
279 - r="3"
|
||
280 - fill="#f4f4f5"
|
||
281 - stroke="#a855f7"
|
||
282 - strokeWidth="1.5"
|
||
283 - >
|
||
284 - <title>{`Area: ${p.label}\nShapes: ${p.count}`
|
||
-}</title>
|
||
285 - </circle>
|
||
193 + {batches.length === 0 && <option value="">no ann
|
||
+otated batches</option>}
|
||
194 + {batches.map((batch) => (
|
||
195 + <option key={batch.id} value={batch.id}>
|
||
196 + {batch.date_label}/{batch.batch_label} · {ba
|
||
+tch.annotation_count} shapes
|
||
197 + </option>
|
||
198 + ))}
|
||
199 + </select>
|
||
200 + {VERDICTS.map((v) => (
|
||
201 + <span key={v.value} className="tag" style={{ fon
|
||
+tSize: '0.78rem', color: v.color, border: `1px solid ${v.color
|
||
+}55`, background: `${v.color}18` }}>
|
||
202 + {counts[v.value]} {v.label.toLowerCase()}
|
||
203 + </span>
|
||
204 ))}
|
||
287 - </svg>
|
||
288 -
|
||
289 - {/* X Axis Labels */}
|
||
290 - <div style={{ display: 'flex', justifyContent: 'spac
|
||
-e-between', marginTop: 8, padding: '0 10px', fontSize: '0.75re
|
||
-m', color: '#a1a1aa', fontFamily: 'monospace' }}>
|
||
291 - {scaleMode === 'normal' ? (
|
||
292 - <>
|
||
293 - <span>0.001% (Tiny)</span>
|
||
294 - <span>0.05%</span>
|
||
295 - <span>0.5% (Median)</span>
|
||
296 - <span>5.0%</span>
|
||
297 - <span>100% (Full Frame)</span>
|
||
298 - </>
|
||
299 - ) : (
|
||
300 - <>
|
||
301 - <span>0% Area</span>
|
||
302 - <span>20%</span>
|
||
303 - <span>40%</span>
|
||
304 - <span>60%</span>
|
||
305 - <span>80%</span>
|
||
306 - <span>100% Area</span>
|
||
307 - </>
|
||
205 + {counts.manual > 0 && (
|
||
206 + <span className="tag" style={{ fontSize: '0.78re
|
||
+m' }}>{counts.manual} by hand</span>
|
||
207 )}
|
||
208 </div>
|
||
209 </div>
|
||
210
|
||
211 + {batches.length === 0 ? (
|
||
212 + <p className="empty">No auto-annotated batches yet.
|
||
+Run auto-annotation on a batch first.</p>
|
||
213 + ) : (
|
||
214 + <>
|
||
215 + <TriageScatter shapes={shapes} selectedIds={select
|
||
+edIds} onSelect={setSelectedIds} />
|
||
216
|
||
313 - {/* Dual Range Sliders & Presets for Size Filtering */
|
||
-}
|
||
314 - <div style={{ marginTop: 20, padding: 16, background:
|
||
-'rgba(255,255,255,0.03)', borderRadius: 8, border: '1px solid
|
||
-rgba(255,255,255,0.06)' }}>
|
||
315 - <div style={{ display: 'flex', justifyContent: 'spac
|
||
-e-between', alignItems: 'center', marginBottom: 12 }}>
|
||
316 - <span style={{ fontSize: '0.88rem', fontWeight: 60
|
||
-0, color: '#f4f4f5' }}>
|
||
317 - Interactive Size Filter Range:
|
||
318 - </span>
|
||
319 -
|
||
320 - {/* Quick Filter Presets */}
|
||
321 - <div style={{ display: 'flex', gap: 6, flexWrap: '
|
||
-wrap' }}>
|
||
322 - <button
|
||
323 - type="button"
|
||
324 - className="tag"
|
||
325 - style={{ cursor: 'pointer', padding: '4px 8px'
|
||
-, fontSize: '0.75rem', background: minSizePct === 0 && maxSize
|
||
-Pct === 100 ? 'rgba(192, 132, 252, 0.25)' : 'rgba(255,255,255,
|
||
-0.05)', color: '#e4e4e7', border: '1px solid rgba(255,255,255,
|
||
-0.1)' }}
|
||
326 - onClick={() => { setMinSizePct(0); setMaxSizeP
|
||
-ct(100); }}
|
||
327 - >
|
||
328 - All Sizes
|
||
329 - </button>
|
||
330 - <button
|
||
331 - type="button"
|
||
332 - className="tag"
|
||
333 - style={{ cursor: 'pointer', padding: '4px 8px'
|
||
-, fontSize: '0.75rem', background: minSizePct === 0 && maxSize
|
||
-Pct === 1 ? 'rgba(192, 132, 252, 0.25)' : 'rgba(255,255,255,0.
|
||
-05)', color: '#e4e4e7', border: '1px solid rgba(255,255,255,0.
|
||
-1)' }}
|
||
334 - onClick={() => { setMinSizePct(0); setMaxSizeP
|
||
-ct(1); }}
|
||
335 - >
|
||
336 - Tiny (< 1%)
|
||
337 - </button>
|
||
338 - <button
|
||
339 - type="button"
|
||
340 - className="tag"
|
||
341 - style={{ cursor: 'pointer', padding: '4px 8px'
|
||
-, fontSize: '0.75rem', background: minSizePct === 1 && maxSize
|
||
-Pct === 5 ? 'rgba(192, 132, 252, 0.25)' : 'rgba(255,255,255,0.
|
||
-05)', color: '#e4e4e7', border: '1px solid rgba(255,255,255,0.
|
||
-1)' }}
|
||
342 - onClick={() => { setMinSizePct(1); setMaxSizeP
|
||
-ct(5); }}
|
||
343 - >
|
||
344 - Small (1%–5%)
|
||
345 - </button>
|
||
346 - <button
|
||
347 - type="button"
|
||
348 - className="tag"
|
||
349 - style={{ cursor: 'pointer', padding: '4px 8px'
|
||
-, fontSize: '0.75rem', background: minSizePct === 5 && maxSize
|
||
-Pct === 20 ? 'rgba(192, 132, 252, 0.25)' : 'rgba(255,255,255,0
|
||
-.05)', color: '#e4e4e7', border: '1px solid rgba(255,255,255,0
|
||
-.1)' }}
|
||
350 - onClick={() => { setMinSizePct(5); setMaxSizeP
|
||
-ct(20); }}
|
||
351 - >
|
||
352 - Medium (5%–20%)
|
||
353 - </button>
|
||
354 - <button
|
||
355 - type="button"
|
||
356 - className="tag"
|
||
357 - style={{ cursor: 'pointer', padding: '4px 8px'
|
||
-, fontSize: '0.75rem', background: minSizePct === 20 && maxSiz
|
||
-ePct === 100 ? 'rgba(192, 132, 252, 0.25)' : 'rgba(255,255,255
|
||
-,0.05)', color: '#e4e4e7', border: '1px solid rgba(255,255,255
|
||
-,0.1)' }}
|
||
358 - onClick={() => { setMinSizePct(20); setMaxSize
|
||
-Pct(100); }}
|
||
359 - >
|
||
360 - Large (> 20%)
|
||
361 - </button>
|
||
362 - </div>
|
||
363 - </div>
|
||
364 -
|
||
365 - <div style={{ display: 'grid', gridTemplateColumns:
|
||
-'1fr 1fr', gap: 20, alignItems: 'center' }}>
|
||
366 - <div>
|
||
367 - <div style={{ display: 'flex', justifyContent: '
|
||
-space-between', alignItems: 'center', marginBottom: 4 }}>
|
||
368 - <label style={{ fontSize: '0.78rem', color: '#
|
||
-a1a1aa' }}>Min Box Area (%):</label>
|
||
369 - <input
|
||
370 - type="number"
|
||
371 - min="0"
|
||
372 - max="100"
|
||
373 - step="0.01"
|
||
374 - value={minSizePct}
|
||
375 - onChange={(e) => setMinSizePct(Math.min(Math
|
||
-.max(0, Number(e.target.value)), maxSizePct))}
|
||
376 - style={{ width: 70, padding: '2px 6px', font
|
||
-Size: '0.8rem', textAlign: 'right', background: 'rgba(0,0,0,0.
|
||
-4)', border: '1px solid rgba(255,255,255,0.15)', borderRadius:
|
||
- 4, color: '#c084fc' }}
|
||
377 - />
|
||
378 - </div>
|
||
379 - <input
|
||
380 - type="range"
|
||
381 - min="0"
|
||
382 - max="100"
|
||
383 - step="0.01"
|
||
384 - value={minSizePct}
|
||
385 - onChange={(e) => setMinSizePct(Math.min(Number
|
||
-(e.target.value), maxSizePct))}
|
||
386 - style={{ width: '100%', cursor: 'pointer' }}
|
||
387 - />
|
||
388 - </div>
|
||
389 -
|
||
390 - <div>
|
||
391 - <div style={{ display: 'flex', justifyContent: '
|
||
-space-between', alignItems: 'center', marginBottom: 4 }}>
|
||
392 - <label style={{ fontSize: '0.78rem', color: '#
|
||
-a1a1aa' }}>Max Box Area (%):</label>
|
||
393 - <input
|
||
394 - type="number"
|
||
395 - min="0"
|
||
396 - max="100"
|
||
397 - step="0.01"
|
||
398 - value={maxSizePct}
|
||
399 - onChange={(e) => setMaxSizePct(Math.max(Math
|
||
-.min(100, Number(e.target.value)), minSizePct))}
|
||
400 - style={{ width: 70, padding: '2px 6px', font
|
||
-Size: '0.8rem', textAlign: 'right', background: 'rgba(0,0,0,0.
|
||
-4)', border: '1px solid rgba(255,255,255,0.15)', borderRadius:
|
||
- 4, color: '#c084fc' }}
|
||
401 - />
|
||
402 - </div>
|
||
403 - <input
|
||
404 - type="range"
|
||
405 - min="0"
|
||
406 - max="100"
|
||
407 - step="0.01"
|
||
408 - value={maxSizePct}
|
||
409 - onChange={(e) => setMaxSizePct(Math.max(Number
|
||
-(e.target.value), minSizePct))}
|
||
410 - style={{ width: '100%', cursor: 'pointer' }}
|
||
411 - />
|
||
412 - </div>
|
||
413 - </div>
|
||
414 - </div>
|
||
415 - </div>
|
||
416 -
|
||
417 - {/* Most Dense Frame Preview Section */}
|
||
418 - {activePreviewFrameId && (
|
||
419 - <div className="panel table-wrap" style={{ padding: 20
|
||
-, marginBottom: 20 }}>
|
||
420 - <div style={{ display: 'flex', justifyContent: 'spac
|
||
-e-between', alignItems: 'center', marginBottom: 14, flexWrap:
|
||
-'wrap', gap: 12 }}>
|
||
421 - <div>
|
||
422 - <h2 style={{ fontSize: '1.05rem', margin: 0, dis
|
||
-play: 'flex', alignItems: 'center', gap: 8 }}>
|
||
423 - Most Dense Frame Preview (Filtered vs Kept Bou
|
||
-nding Boxes)
|
||
424 - </h2>
|
||
425 - <p className="hint" style={{ fontSize: '0.8rem',
|
||
- margin: '4px 0 0 0' }}>
|
||
426 - Active Filter Range: <strong style={{ color: '
|
||
-#c084fc' }}>{minSizePct}% – {maxSizePct}% Area</strong>. Filte
|
||
-red-out boxes are dimmed in red.
|
||
427 - </p>
|
||
428 - </div>
|
||
429 -
|
||
430 - {/* Dense Frame Dropdown Selector & Counters */}
|
||
431 - <div style={{ display: 'flex', gap: 10, alignItems
|
||
-: 'center', flexWrap: 'wrap' }}>
|
||
432 - <label style={{ fontSize: '0.8rem', color: '#a1a
|
||
-1aa' }}>Select Dense Frame:</label>
|
||
217 + <div
|
||
218 + style={{
|
||
219 + display: 'flex',
|
||
220 + gap: 10,
|
||
221 + alignItems: 'center',
|
||
222 + flexWrap: 'wrap',
|
||
223 + margin: '16px 0',
|
||
224 + padding: 12,
|
||
225 + background: selectedIds.length ? 'rgba(56,189,
|
||
+248,0.08)' : 'rgba(255,255,255,0.03)',
|
||
226 + border: `1px solid ${selectedIds.length ? 'rgb
|
||
+a(56,189,248,0.3)' : 'rgba(255,255,255,0.06)'}`,
|
||
227 + borderRadius: 8,
|
||
228 + transition: 'background 150ms ease, border-col
|
||
+or 150ms ease',
|
||
229 + }}
|
||
230 + >
|
||
231 + <strong style={{ fontSize: '0.85rem' }}>{selecte
|
||
+dIds.length} selected</strong>
|
||
232 + <span className="faint" style={{ fontSize: '0.78
|
||
+rem' }}>— decide by hand (overrides every rule):</span>
|
||
233 + {VERDICTS.map((v) => (
|
||
234 + <button
|
||
235 + key={v.value}
|
||
236 + type="button"
|
||
237 + className="btn"
|
||
238 + disabled={busy || selectedIds.length === 0}
|
||
239 + onClick={() => applyVerdict(v.value)}
|
||
240 + style={{ cursor: 'pointer', fontSize: '0.8re
|
||
+m', color: v.color, borderColor: `${v.color}55` }}
|
||
241 + >
|
||
242 + {v.label}
|
||
243 + </button>
|
||
244 + ))}
|
||
245 <select
|
||
434 - value={activePreviewFrameId}
|
||
435 - onChange={(e) => setPreviewFrameId(Number(e.ta
|
||
-rget.value))}
|
||
436 - style={{ padding: '4px 10px', fontSize: '0.82r
|
||
-em', background: 'rgba(0,0,0,0.4)', color: '#f4f4f5', border:
|
||
-'1px solid rgba(255,255,255,0.15)', borderRadius: 6 }}
|
||
246 + value={reclassTarget}
|
||
247 + onChange={(event) => setReclassTarget(event.ta
|
||
+rget.value)}
|
||
248 + aria-label="Reclass target class"
|
||
249 + style={{ padding: '4px 8px', fontSize: '0.8rem
|
||
+', background: 'rgba(0,0,0,0.4)', color: '#e4e4e7', border: '1
|
||
+px solid rgba(255,255,255,0.15)', borderRadius: 4, cursor: 'po
|
||
+inter' }}
|
||
250 >
|
||
438 - {topDenseFrames.map((f, idx) => (
|
||
439 - <option key={f.frame_id} value={f.frame_id}>
|
||
440 - Frame #{f.frame_id} ({f.total_shapes} tota
|
||
-l boxes) {idx === 0 ? '🔥 [Most Dense]' : ''}
|
||
441 - </option>
|
||
251 + <option value="">reclass into…</option>
|
||
252 + {project.classes?.map((cls) => (
|
||
253 + <option key={cls.class_id} value={cls.class_
|
||
+id}>{cls.name}</option>
|
||
254 ))}
|
||
255 </select>
|
||
444 -
|
||
445 - <span className="tag" style={{ background: 'rgba
|
||
-(34, 197, 94, 0.2)', color: '#4ade80', border: '1px solid rgba
|
||
-(34, 197, 94, 0.4)', fontSize: '0.8rem' }}>
|
||
446 - {keptShapesOnFrame.length} Kept
|
||
447 - </span>
|
||
448 - <span className="tag" style={{ background: 'rgba
|
||
-(239, 68, 68, 0.15)', color: '#f87171', border: '1px solid rgb
|
||
-a(239, 68, 68, 0.3)', fontSize: '0.8rem' }}>
|
||
449 - {filteredOutShapesOnFrame.length} Filtered Out
|
||
450 - </span>
|
||
451 -
|
||
256 <button
|
||
257 type="button"
|
||
454 - className="tag"
|
||
455 - style={{
|
||
456 - cursor: 'pointer',
|
||
457 - padding: '4px 10px',
|
||
458 - fontSize: '0.78rem',
|
||
459 - background: showFilteredOut ? 'rgba(239, 68,
|
||
- 68, 0.2)' : 'rgba(255,255,255,0.05)',
|
||
460 - color: showFilteredOut ? '#f87171' : '#a1a1a
|
||
-a',
|
||
461 - border: showFilteredOut ? '1px solid rgba(23
|
||
-9, 68, 68, 0.4)' : '1px solid rgba(255,255,255,0.1)',
|
||
462 - borderRadius: 6
|
||
463 - }}
|
||
464 - onClick={() => setShowFilteredOut(!showFiltere
|
||
-dOut)}
|
||
258 + className="btn"
|
||
259 + disabled={busy || selectedIds.length === 0}
|
||
260 + onClick={clearDecisions}
|
||
261 + style={{ cursor: 'pointer', fontSize: '0.8rem'
|
||
+, marginLeft: 'auto' }}
|
||
262 >
|
||
466 - {showFilteredOut ? 'Hide Filtered Out Boxes' :
|
||
- 'Show Filtered Out Boxes'}
|
||
263 + Clear hand decisions
|
||
264 </button>
|
||
265 </div>
|
||
469 - </div>
|
||
266
|
||
471 - {/* Interactive Frame Canvas Preview */}
|
||
472 - <div style={{ position: 'relative', width: '100%', m
|
||
-axWidth: 880, margin: '0 auto', background: '#000', borderRadi
|
||
-us: 8, overflow: 'hidden', border: '1px solid rgba(255,255,255
|
||
-,0.1)' }}>
|
||
473 - <img
|
||
474 - src={`/api/frames/${activePreviewFrameId}/image`
|
||
-}
|
||
475 - alt={`Frame ${activePreviewFrameId}`}
|
||
476 - style={{ width: '100%', height: 'auto', display:
|
||
- 'block' }}
|
||
267 + <TriageCropGrid
|
||
268 + shapes={shapes}
|
||
269 + selectedIds={selectedIds}
|
||
270 + onSelect={setSelectedIds}
|
||
271 + classes={project.classes}
|
||
272 />
|
||
273 + </>
|
||
274 + )}
|
||
275 + </div>
|
||
276
|
||
479 - {/* SVG Bounding Box Overlays showing Kept vs Filt
|
||
-ered Out */}
|
||
480 - <svg
|
||
481 - viewBox="0 0 1 1"
|
||
482 - preserveAspectRatio="none"
|
||
483 - style={{ position: 'absolute', top: 0, left: 0,
|
||
-width: '100%', height: '100%', pointerEvents: 'none' }}
|
||
484 - >
|
||
485 - {allFrameShapes.map((s) => {
|
||
486 - if (!s.box) return null
|
||
487 - const isKept = s.area_pct >= minSizePct && s.a
|
||
-rea_pct <= maxSizePct
|
||
488 - if (!isKept && !showFilteredOut) return null
|
||
489 -
|
||
490 - const [x0, y0, x1, y1] = s.box
|
||
491 - const cls = project.classes?.find((c) => c.cla
|
||
-ss_id === s.class_id)
|
||
492 - const color = isKept ? classColor(s.class_id)
|
||
-: '#ef4444'
|
||
493 -
|
||
494 - return (
|
||
495 - <g key={s.id} opacity={isKept ? 1.0 : 0.45}>
|
||
496 - <rect
|
||
497 - x={x0}
|
||
498 - y={y0}
|
||
499 - width={x1 - x0}
|
||
500 - height={y1 - y0}
|
||
501 - fill={isKept ? 'rgba(56, 189, 248, 0.2)'
|
||
- : 'rgba(239, 68, 68, 0.08)'}
|
||
502 - stroke={color}
|
||
503 - strokeWidth={isKept ? '0.0035' : '0.002'
|
||
-}
|
||
504 - strokeDasharray={isKept ? 'none' : '0.00
|
||
-4 0.004'}
|
||
505 - />
|
||
506 - <text
|
||
507 - x={x0 + 0.004}
|
||
508 - y={Math.max(0.02, y0 - 0.004)}
|
||
509 - fill={isKept ? '#ffffff' : '#f87171'}
|
||
510 - fontSize="0.02"
|
||
511 - fontWeight={isKept ? 'bold' : 'normal'}
|
||
512 - fontFamily="sans-serif"
|
||
513 - >
|
||
514 - {isKept ? `${cls?.name || `Class ${s.cla
|
||
-ss_id}`} (${s.area_pct}%)` : `❌ Filtered (${s.area_pct}%)`}
|
||
515 - </text>
|
||
516 - </g>
|
||
517 - )
|
||
518 - })}
|
||
519 - </svg>
|
||
520 - </div>
|
||
521 - </div>
|
||
522 - )}
|
||
523 -
|
||
524 -
|
||
525 -
|
||
526 - <div style={{ display: 'grid', gridTemplateColumns: '1fr
|
||
- 340px', gap: 20, alignItems: 'start' }}>
|
||
527 -
|
||
528 - {/* Main Section: Class Distribution & Merged Batches
|
||
-*/}
|
||
529 - <div style={{ display: 'flex', flexDirection: 'column'
|
||
-, gap: 20 }}>
|
||
530 - {/* Target Classes Summary with Continuous Size Filt
|
||
-er */}
|
||
531 - <div className="panel table-wrap" style={{ padding:
|
||
-18 }}>
|
||
532 - <h2 style={{ fontSize: '1rem', marginBottom: 12, d
|
||
-isplay: 'flex', alignItems: 'center', gap: 8 }}>
|
||
533 - Target Class Distribution {isFiltering ? `(Filte
|
||
-red: ${minSizePct}% – ${maxSizePct}% Area)` : ''}
|
||
534 - </h2>
|
||
535 - {project.classes?.length === 0 ? (
|
||
536 - <p className="hint">No target classes defined fo
|
||
-r this project.</p>
|
||
537 - ) : (
|
||
538 - <div style={{ display: 'grid', gridTemplateColum
|
||
-ns: 'repeat(auto-fill, minmax(220px, 1fr))', gap: 12 }}>
|
||
539 - {project.classes?.map((cls) => {
|
||
540 - const count = filteredClassCounts[cls.class_
|
||
-id] || 0
|
||
541 - return (
|
||
542 - <div
|
||
543 - key={cls.class_id}
|
||
544 - style={{
|
||
545 - padding: 12,
|
||
546 - background: 'rgba(0,0,0,0.3)',
|
||
547 - borderRadius: 8,
|
||
548 - border: isFiltering && count > 0 ? '1p
|
||
-x solid rgba(192, 132, 252, 0.4)' : '1px solid rgba(255,255,25
|
||
-5,0.08)',
|
||
549 - }}
|
||
550 - >
|
||
551 - <div style={{ display: 'flex', justifyCo
|
||
-ntent: 'space-between', alignItems: 'center' }}>
|
||
552 - <div style={{ fontSize: '0.85rem', fon
|
||
-tWeight: 600, color: '#f4f4f5' }}>{cls.name}</div>
|
||
553 - <span className="tag" style={{ fontSiz
|
||
-e: '0.75rem', background: 'rgba(168, 85, 247, 0.2)', color: '#
|
||
-c084fc', border: '1px solid rgba(168, 85, 247, 0.3)' }}>
|
||
554 - {count} shapes
|
||
555 - </span>
|
||
556 - </div>
|
||
557 - <div style={{ fontSize: '0.75rem', color
|
||
-: '#a1a1aa', marginTop: 4 }}>
|
||
558 - Class ID: <code style={{ color: '#c084
|
||
-fc' }}>{cls.class_id}</code>
|
||
559 - </div>
|
||
560 - {cls.prompt && (
|
||
561 - <div style={{ fontSize: '0.75rem', col
|
||
-or: '#38bdf8', marginTop: 4 }}>
|
||
562 - Prompt: "{cls.prompt}"
|
||
563 - </div>
|
||
564 - )}
|
||
565 - </div>
|
||
566 - )
|
||
567 - })}
|
||
568 - </div>
|
||
569 - )}
|
||
570 - </div>
|
||
571 -
|
||
572 -
|
||
573 - {/* Merged Batches List */}
|
||
574 - <div className="panel table-wrap" style={{ padding:
|
||
-18 }}>
|
||
575 - <h2 style={{ fontSize: '1rem', marginBottom: 12 }}
|
||
->Approved & Merged Dataset Batches ({mergedBatches.length})</h
|
||
-2>
|
||
576 - {mergedBatches.length === 0 ? (
|
||
577 - <p className="empty">No approved batches merged
|
||
-into dataset yet. Go to Batches page to review & approve.</p>
|
||
578 - ) : (
|
||
579 - <table className="video-table">
|
||
580 - <thead>
|
||
581 - <tr>
|
||
582 - <th>Batch</th>
|
||
583 - <th>Date</th>
|
||
584 - <th>Frames</th>
|
||
585 - <th>Reviewed</th>
|
||
586 - <th>Status</th>
|
||
587 - </tr>
|
||
588 - </thead>
|
||
589 - <tbody>
|
||
590 - {mergedBatches.map((item) => (
|
||
591 - <tr key={item.id}>
|
||
592 - <td style={{ fontWeight: 500 }}>{item.ba
|
||
-tch_label}</td>
|
||
593 - <td>{item.date_label}</td>
|
||
594 - <td className="num">{item.images}</td>
|
||
595 - <td>
|
||
596 - <span className="tag" style={{ backgro
|
||
-und: 'rgba(34, 197, 94, 0.15)', color: '#4ade80', border: '1px
|
||
- solid rgba(34, 197, 94, 0.3)' }}>
|
||
597 - Approved
|
||
598 - </span>
|
||
599 - </td>
|
||
600 - <td>
|
||
601 - <span className="mono" style={{ fontSi
|
||
-ze: '0.8rem', color: '#a1a1aa' }}>
|
||
602 - In Master Dataset
|
||
603 - </span>
|
||
604 - </td>
|
||
605 - </tr>
|
||
606 - ))}
|
||
607 - </tbody>
|
||
608 - </table>
|
||
609 - )}
|
||
610 - </div>
|
||
611 - </div>
|
||
612 -
|
||
613 - {/* Right Sidebar: Next Steps & Quick Actions */}
|
||
614 - <div style={{ display: 'flex', flexDirection: 'column'
|
||
-, gap: 16 }}>
|
||
615 - <div className="panel side-panel" style={{ border: '
|
||
-1px solid rgba(168, 85, 247, 0.3)' }}>
|
||
616 - <h2 style={{ fontSize: '0.95rem', color: '#c084fc'
|
||
-, marginBottom: 8 }}>Data Preparation Readiness</h2>
|
||
617 - <p className="hint" style={{ fontSize: '0.8rem', m
|
||
-arginBottom: 14 }}>
|
||
618 - Once your dataset and shape sizes are verified,
|
||
-proceed to fine-tune baseline YOLO models.
|
||
619 - </p>
|
||
620 - <a
|
||
621 - className="btn btn-primary"
|
||
622 - href={`#/projects/${projectId}/models`}
|
||
623 - style={{ width: '100%', padding: '8px 12px', fon
|
||
-tSize: '0.85rem', cursor: 'pointer', display: 'inline-flex', a
|
||
-lignItems: 'center', justifyContent: 'center', gap: 6, borderR
|
||
-adius: 6 }}
|
||
624 - >
|
||
625 - <RocketIcon size={16} /> Proceed to Models & Tra
|
||
-ining
|
||
626 - </a>
|
||
627 - </div>
|
||
628 - </div>
|
||
277 + <div className="panel side-panel" style={{ marginTop: 18
|
||
+, border: '1px solid rgba(168, 85, 247, 0.3)' }}>
|
||
278 + <h2 style={{ fontSize: '0.95rem', color: '#c084fc', ma
|
||
+rginBottom: 8 }}>Ready to train</h2>
|
||
279 + <p className="hint" style={{ fontSize: '0.8rem', margi
|
||
+nBottom: 14 }}>
|
||
280 + Training re-applies these rules to the master datase
|
||
+t, so what you decide here is what
|
||
281 + the next run sees. Images carrying an ignored shape
|
||
+are left out whole.
|
||
282 + </p>
|
||
283 + <a
|
||
284 + className="btn btn-primary"
|
||
285 + href={`#/projects/${projectId}/models`}
|
||
286 + style={{ padding: '8px 12px', fontSize: '0.85rem', c
|
||
+ursor: 'pointer', display: 'inline-flex', alignItems: 'center'
|
||
+, gap: 6, borderRadius: 6 }}
|
||
287 + >
|
||
288 + <RocketIcon size={16} /> Proceed to models & tra
|
||
+ining
|
||
289 + </a>
|
||
290 </div>
|
||
291 </>
|
||
292 )
|
||
293 }
|
||
294
|
||
634 -
|
||
635 -
|
||
295 +function Stat({ label, value, hint, accent, mono }) {
|
||
296 + return (
|
||
297 + <div className="panel side-panel" style={{ padding: 16, bo
|
||
+rderLeft: `3px solid ${accent}` }}>
|
||
298 + <span style={{ fontSize: '0.75rem', color: '#a1a1aa', te
|
||
+xtTransform: 'uppercase', letterSpacing: '0.05em' }}>{label}</
|
||
+span>
|
||
299 + <div className={mono ? 'mono' : undefined} style={{ font
|
||
+Size: mono ? '1.1rem' : '1.8rem', fontWeight: 700, color: acce
|
||
+nt, marginTop: 4 }}>{value}</div>
|
||
300 + <p className="hint" style={{ fontSize: '0.76rem', margin
|
||
+: '4px 0 0 0' }}>{hint}</p>
|
||
301 + </div>
|
||
302 + )
|
||
303 +}
|
||
|
||
Ran 7 shell commands (ctrl+o to expand)
|
||
|
||
● Router's fine — my introspection was wrong. Let me just run the server and hit
|
||
it for real:
|
||
|
||
Ran 2 shell commands (ctrl+o to expand)
|
||
|
||
● Real data — batch 36 has 85,520 shapes. That's a good stress test:
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● It works and it's fast, but 21 MB of JSON for one batch — 85k SVG circles
|
||
would lock the browser. That's a real defect, not a nit. Fixing it before I
|
||
call this done:
|
||
|
||
● Update(backend/triage.py)
|
||
⎿ Added 46 lines, removed 3 lines
|
||
211 return shapes
|
||
212
|
||
213
|
||
214 -def batch_shapes(batch_id: int) -> dict:
|
||
215 - """The triage view of one batch: every shape, scored and r
|
||
-esolved (REQ-106)."""
|
||
214 +SCATTER_POINTS = 4000
|
||
215 +"""How many dots the scatter gets. A real batch runs to ~85k s
|
||
+hapes; every one of
|
||
216 +them as an SVG circle locks the browser, and a boundary betwee
|
||
+n two clusters is
|
||
217 +just as visible in a few thousand points. The verdict tallies
|
||
+are still counted
|
||
218 +over every shape, so the numbers are never a sample."""
|
||
219 +
|
||
220 +
|
||
221 +def _resolved_shapes(batch_id: int):
|
||
222 from backend import batches
|
||
223
|
||
224 batch = batches.get(batch_id)
|
||
...
|
||
232 shapes = shapes_for_frames(frame_ids)
|
||
233 for shape in shapes:
|
||
234 shape.update(resolver.verdict(shape))
|
||
235 + return batch, frame_ids, shapes, resolver
|
||
236 +
|
||
237 +
|
||
238 +def batch_summary(batch_id: int) -> dict:
|
||
239 + """Verdict tallies over the whole batch, plus a sample to
|
||
+plot (REQ-106)."""
|
||
240 + batch, frame_ids, shapes, resolver = _resolved_shapes(batc
|
||
+h_id)
|
||
241 +
|
||
242 + counts = {"keep": 0, "ignore": 0, "reclass": 0, "manual":
|
||
+0}
|
||
243 + for shape in shapes:
|
||
244 + counts[shape["verdict"]] += 1
|
||
245 + if shape["source"] == "manual":
|
||
246 + counts["manual"] += 1
|
||
247 +
|
||
248 + # An even stride rather than a random draw: the sample is
|
||
+stable across
|
||
249 + # reloads, so points do not jump around while the user is
|
||
+reading the plot.
|
||
250 + stride = max(1, len(shapes) // SCATTER_POINTS)
|
||
251 + sample = [
|
||
252 + {k: shape[k] for k in ("id", "class_id", "score", "are
|
||
+a_pct", "aspect", "verdict", "source")}
|
||
253 + for shape in shapes[::stride][:SCATTER_POINTS]
|
||
254 + ]
|
||
255 +
|
||
256 return {
|
||
257 "batch_id": batch_id,
|
||
258 "project_id": batch["project_id"],
|
||
259 "frame_count": len(frame_ids),
|
||
233 - "shapes": shapes,
|
||
260 + "total_shapes": len(shapes),
|
||
261 + "counts": counts,
|
||
262 + "sample": sample,
|
||
263 + "sampled": len(sample) < len(shapes),
|
||
264 "rule_version": resolver.version(),
|
||
265 }
|
||
266
|
||
267
|
||
268 +def batch_page(batch_id: int, sort: str = "score", offset: int
|
||
+ = 0, limit: int = 120) -> dict:
|
||
269 + """One page of shapes for the crop grid, sorted server-sid
|
||
+e so the client
|
||
270 + never holds the whole batch."""
|
||
271 + if sort not in ("score", "area_pct"):
|
||
272 + raise TriageError(f"Cannot sort by {sort}")
|
||
273 + _, _, shapes, _ = _resolved_shapes(batch_id)
|
||
274 + shapes.sort(key=lambda shape: shape[sort])
|
||
275 + page = shapes[offset:offset + limit]
|
||
276 + for shape in page:
|
||
277 + shape.pop("box", None)
|
||
278 + return {"total": len(shapes), "offset": offset, "limit": l
|
||
+imit, "shapes": page}
|
||
279 +
|
||
280 +
|
||
281 def preview(project_id: int) -> dict:
|
||
282 """What the current rules would do to the whole merged dat
|
||
aset."""
|
||
283 with db.cursor() as cur:
|
||
|
||
● Update(backend/api/triage.py)
|
||
⎿ Added 11 lines, removed 3 lines
|
||
49 return {"rules": stored}
|
||
50
|
||
51
|
||
52 -@router.get("/api/batches/{batch_id}/triage/shapes")
|
||
53 -def batch_shapes(batch_id: int) -> dict:
|
||
52 +@router.get("/api/batches/{batch_id}/triage/summary")
|
||
53 +def batch_summary(batch_id: int) -> dict:
|
||
54 try:
|
||
55 - return triage.batch_shapes(batch_id)
|
||
55 + return triage.batch_summary(batch_id)
|
||
56 except triage.TriageError as exc:
|
||
57 raise HTTPException(404, str(exc))
|
||
58
|
||
59
|
||
60 +@router.get("/api/batches/{batch_id}/triage/shapes")
|
||
61 +def batch_page(batch_id: int, sort: str = "score", offset: int
|
||
+= 0, limit: int = 120) -> dict:
|
||
62 + try:
|
||
63 + return triage.batch_page(batch_id, sort=sort, offset=of
|
||
+fset, limit=min(limit, 500))
|
||
64 + except triage.TriageError as exc:
|
||
65 + raise HTTPException(400, str(exc))
|
||
66 +
|
||
67 +
|
||
68 @router.post("/api/triage/overrides")
|
||
69 def set_overrides(body: OverrideRequest) -> dict:
|
||
70 try:
|
||
|
||
Ran 4 shell commands (ctrl+o to expand)
|
||
|
||
● 21 MB → 450 KB. Now let me verify rules actually work against your real data:
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Rule works — 16,776 of 85,520 shapes ignored. But preview shows 1,390 merged
|
||
images with 0 shapes total. Let me check whether that's a bug in my code or
|
||
something real about your data:
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Important find — and it's not my code. Let me clean up my test rule first:
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Now update the frontend to the paginated endpoints, then a final build:
|
||
|
||
● Update(frontend/src/api.js)
|
||
⎿ Added 3 lines, removed 1 line
|
||
121 triageRules: (projectId) => request(`/projects/${projectId}/
|
||
triage/rules`),
|
||
122 saveTriageRules: (projectId, rules) =>
|
||
123 request(`/projects/${projectId}/triage/rules`, { method: '
|
||
PUT', body: { rules } }),
|
||
124 - triageShapes: (batchId) => request(`/batches/${batchId}/tria
|
||
-ge/shapes`),
|
||
124 + triageSummary: (batchId) => request(`/batches/${batchId}/tri
|
||
+age/summary`),
|
||
125 + triageShapes: (batchId, { sort = 'score', offset = 0, limit
|
||
+= 120 } = {}) =>
|
||
126 + request(`/batches/${batchId}/triage/shapes?sort=${sort}&of
|
||
+fset=${offset}&limit=${limit}`),
|
||
127 triagePreview: (projectId) => request(`/projects/${projectId
|
||
}/triage/preview`),
|
||
128 setTriageOverrides: (annotationIds, verdict, targetClass = n
|
||
ull) =>
|
||
129 request('/triage/overrides', {
|
||
|
||
● Write(frontend/src/components/TriageCropGrid.jsx)
|
||
⎿ Added 50 lines, removed 23 lines
|
||
1 -import React, { useState } from 'react'
|
||
1 +import React, { useCallback, useEffect, useState } from 'react'
|
||
2 import { api } from '../api'
|
||
3
|
||
4 /* A wall of cropped shapes (REQ-106). Judging "is this a half
|
||
sack?" means
|
||
5 - * looking at the object, not at the frame it came from. Sorted
|
||
- so the shapes a
|
||
6 - * rule is about to act on sit together. */
|
||
5 + * looking at the object, not at the frame it came from.
|
||
6 + *
|
||
7 + * Paged from the server: a real batch holds ~85k shapes, so th
|
||
+e client never
|
||
8 + * gets to hold them all. Sorting happens server-side for the s
|
||
+ame reason —
|
||
9 + * "worst 120 by score" has to be chosen from the whole batch,
|
||
+not from a page. */
|
||
10
|
||
11 const VERDICT_COLOR = { keep: '#4ade80', ignore: '#f87171', rec
|
||
lass: '#c084fc' }
|
||
12 const PAGE = 120
|
||
13
|
||
11 -export default function TriageCropGrid({ shapes, selectedIds, o
|
||
-nSelect, classes }) {
|
||
14 +export default function TriageCropGrid({ batchId, selectedIds,
|
||
+onSelect, classes, reloadKey }) {
|
||
15 const [sortBy, setSortBy] = useState('score')
|
||
13 - const [limit, setLimit] = useState(PAGE)
|
||
16 + const [shapes, setShapes] = useState([])
|
||
17 + const [total, setTotal] = useState(0)
|
||
18 + const [loading, setLoading] = useState(false)
|
||
19 + const [error, setError] = useState('')
|
||
20
|
||
15 - const sorted = [...shapes].sort((a, b) =>
|
||
16 - sortBy === 'score' ? a.score - b.score : a.area_pct - b.are
|
||
-a_pct,
|
||
21 + const fetchPage = useCallback(
|
||
22 + async (offset, replace) => {
|
||
23 + if (!batchId) return
|
||
24 + setLoading(true)
|
||
25 + try {
|
||
26 + const page = await api.triageShapes(batchId, { sort: so
|
||
+rtBy, offset, limit: PAGE })
|
||
27 + setTotal(page.total)
|
||
28 + setShapes((rows) => (replace ? page.shapes : [...rows,
|
||
+...page.shapes]))
|
||
29 + setError('')
|
||
30 + } catch (exc) {
|
||
31 + setError(exc.message)
|
||
32 + } finally {
|
||
33 + setLoading(false)
|
||
34 + }
|
||
35 + },
|
||
36 + [batchId, sortBy],
|
||
37 )
|
||
18 - const visible = sorted.slice(0, limit)
|
||
38
|
||
39 + useEffect(() => { fetchPage(0, true) }, [fetchPage, reloadKey
|
||
+])
|
||
40 +
|
||
41 function toggle(id, event) {
|
||
42 if (event.shiftKey) {
|
||
43 onSelect(selectedIds.includes(id) ? selectedIds.filter((x
|
||
) => x !== id) : [...selectedIds, id])
|
||
...
|
||
52 <div>
|
||
53 <div style={{ display: 'flex', justifyContent: 'space-bet
|
||
ween', alignItems: 'center', marginBottom: 10, gap: 12, flexWra
|
||
p: 'wrap' }}>
|
||
54 <span className="hint" style={{ fontSize: '0.8rem' }}>
|
||
34 - Showing {visible.length} of {sorted.length} · click t
|
||
-o select, shift-click to add
|
||
55 + Showing {shapes.length} of {total} · click to select,
|
||
+ shift-click to add
|
||
56 </span>
|
||
57 <div style={{ display: 'flex', gap: 6 }}>
|
||
37 - {['score', 'area'].map((mode) => (
|
||
58 + {[
|
||
59 + { key: 'score', label: 'lowest score' },
|
||
60 + { key: 'area_pct', label: 'smallest area' },
|
||
61 + ].map((mode) => (
|
||
62 <button
|
||
39 - key={mode}
|
||
63 + key={mode.key}
|
||
64 type="button"
|
||
65 className="tag"
|
||
42 - aria-pressed={sortBy === mode}
|
||
43 - onClick={() => setSortBy(mode)}
|
||
66 + aria-pressed={sortBy === mode.key}
|
||
67 + onClick={() => setSortBy(mode.key)}
|
||
68 style={{
|
||
69 cursor: 'pointer',
|
||
70 padding: '4px 10px',
|
||
71 fontSize: '0.78rem',
|
||
48 - background: sortBy === mode ? 'rgba(56,189,248,
|
||
-0.2)' : 'rgba(255,255,255,0.05)',
|
||
49 - color: sortBy === mode ? '#38bdf8' : '#a1a1aa',
|
||
50 - border: `1px solid ${sortBy === mode ? 'rgba(56
|
||
-,189,248,0.5)' : 'rgba(255,255,255,0.1)'}`,
|
||
72 + background: sortBy === mode.key ? 'rgba(56,189,
|
||
+248,0.2)' : 'rgba(255,255,255,0.05)',
|
||
73 + color: sortBy === mode.key ? '#38bdf8' : '#a1a1
|
||
+aa',
|
||
74 + border: `1px solid ${sortBy === mode.key ? 'rgb
|
||
+a(56,189,248,0.5)' : 'rgba(255,255,255,0.1)'}`,
|
||
75 borderRadius: 6,
|
||
76 transition: 'background 150ms ease, color 150ms
|
||
ease',
|
||
77 }}
|
||
78 >
|
||
55 - Sort by {mode === 'score' ? 'lowest score' : 'sma
|
||
-llest area'}
|
||
79 + Sort by {mode.label}
|
||
80 </button>
|
||
81 ))}
|
||
82 </div>
|
||
83 </div>
|
||
84
|
||
61 - {sorted.length === 0 ? (
|
||
62 - <p className="empty">No shapes on this batch yet — run
|
||
-auto-annotation first.</p>
|
||
85 + {error && <p className="error-banner" style={{ marginBott
|
||
+om: 10 }}>{error}</p>}
|
||
86 +
|
||
87 + {total === 0 && !loading ? (
|
||
88 + <p className="empty">No shapes on this batch — run auto
|
||
+-annotation first.</p>
|
||
89 ) : (
|
||
90 <div style={{ display: 'grid', gridTemplateColumns: 're
|
||
peat(auto-fill, minmax(104px, 1fr))', gap: 8 }}>
|
||
65 - {visible.map((shape) => {
|
||
91 + {shapes.map((shape) => {
|
||
92 const chosen = selectedIds.includes(shape.id)
|
||
93 return (
|
||
94 <button
|
||
...
|
||
139 </div>
|
||
140 )}
|
||
141
|
||
116 - {limit < sorted.length && (
|
||
142 + {shapes.length < total && (
|
||
143 <button
|
||
144 type="button"
|
||
145 className="btn"
|
||
120 - onClick={() => setLimit((n) => n + PAGE)}
|
||
146 + disabled={loading}
|
||
147 + onClick={() => fetchPage(shapes.length, false)}
|
||
148 style={{ marginTop: 12, cursor: 'pointer', fontSize:
|
||
'0.82rem' }}
|
||
149 >
|
||
123 - Load {Math.min(PAGE, sorted.length - limit)} more
|
||
150 + {loading ? 'Loading…' : `Load ${Math.min(PAGE, total
|
||
+ - shapes.length)} more`}
|
||
151 </button>
|
||
152 )}
|
||
153 </div>
|
||
|
||
● Update(frontend/src/pages/DataPrepPage.jsx)
|
||
⎿ Added 4 lines, removed 3 lines
|
||
61 const loadBatch = useCallback(async () => {
|
||
62 if (!batchId) return
|
||
63 try {
|
||
64 - const [shapes, previewed] = await Promise.all([
|
||
65 - api.triageShapes(batchId),
|
||
64 + const [summary, previewed] = await Promise.all([
|
||
65 + api.triageSummary(batchId),
|
||
66 api.triagePreview(projectId),
|
||
67 ])
|
||
68 - setView(shapes)
|
||
68 + setView(summary)
|
||
69 setPreview(previewed)
|
||
70 setSelectedIds([])
|
||
71 + setReloadKey((n) => n + 1)
|
||
72 } catch (exc) {
|
||
73 setError(exc.message)
|
||
74 }
|
||
|
||
● Update(frontend/src/pages/DataPrepPage.jsx)
|
||
⎿ Added 1 line
|
||
29 const [savedRules, setSavedRules] = useState([])
|
||
30 const [preview, setPreview] = useState(null)
|
||
31 const [selectedIds, setSelectedIds] = useState([])
|
||
32 + const [reloadKey, setReloadKey] = useState(0)
|
||
33 const [reclassTarget, setReclassTarget] = useState('')
|
||
34 const [saving, setSaving] = useState(false)
|
||
35 const [busy, setBusy] = useState(false)
|
||
|
||
● Update(frontend/src/pages/DataPrepPage.jsx)
|
||
⎿ Added 4 lines, removed 10 lines
|
||
77
|
||
78 useEffect(() => { loadBatch() }, [loadBatch])
|
||
79
|
||
80 - const shapes = view?.shapes ?? []
|
||
80 + // Counted server-side over every shape in the batch — the sc
|
||
+atter plots a
|
||
81 + // sample, but these numbers are never a sample.
|
||
82 + const counts = view?.counts ?? { keep: 0, ignore: 0, reclass:
|
||
+ 0, manual: 0 }
|
||
83 + const sample = view?.sample ?? []
|
||
84 const dirty = JSON.stringify(rules) !== JSON.stringify(savedR
|
||
ules)
|
||
85
|
||
83 - const counts = useMemo(() => {
|
||
84 - const tally = { keep: 0, ignore: 0, reclass: 0, manual: 0 }
|
||
85 - shapes.forEach((shape) => {
|
||
86 - tally[shape.verdict] += 1
|
||
87 - if (shape.source === 'manual') tally.manual += 1
|
||
88 - })
|
||
89 - return tally
|
||
90 - }, [shapes])
|
||
91 -
|
||
86 async function saveRules() {
|
||
87 setSaving(true)
|
||
88 try {
|
||
|
||
● Update(frontend/src/pages/DataPrepPage.jsx)
|
||
⎿ Added 7 lines, removed 1 line
|
||
208 <p className="empty">No auto-annotated batches yet.
|
||
Run auto-annotation on a batch first.</p>
|
||
209 ) : (
|
||
210 <>
|
||
211 - <TriageScatter shapes={shapes} selectedIds={select
|
||
-edIds} onSelect={setSelectedIds} />
|
||
211 + <TriageScatter shapes={sample} selectedIds={select
|
||
+edIds} onSelect={setSelectedIds} />
|
||
212 + {view?.sampled && (
|
||
213 + <p className="hint" style={{ fontSize: '0.78rem'
|
||
+, margin: '4px 0 0' }}>
|
||
214 + Plotting {sample.length} of {view.total_shapes
|
||
+} shapes — enough to see where the
|
||
215 + clusters split. The counts above cover all of
|
||
+them.
|
||
216 + </p>
|
||
217 + )}
|
||
218
|
||
219 <div
|
||
220 style={{
|
||
|
||
● Update(frontend/src/pages/DataPrepPage.jsx)
|
||
⎿ Added 2 lines, removed 1 line
|
||
267 </div>
|
||
268
|
||
269 <TriageCropGrid
|
||
270 - shapes={shapes}
|
||
270 + batchId={batchId}
|
||
271 + reloadKey={reloadKey}
|
||
272 selectedIds={selectedIds}
|
||
273 onSelect={setSelectedIds}
|
||
274 classes={project.classes}
|
||
|
||
Ran 3 shell commands (ctrl+o to expand)
|
||
|
||
● It's built and working against your real data. Test server stopped, test rule
|
||
removed — your project state is clean.
|
||
|
||
Verified against project 5, batch 36 (85,520 real SAM3 shapes):
|
||
|
||
- Applied a rule score < 0.45 → ignore. Result: 68,744 keep / 16,776 ignore,
|
||
resolved in ~0.5 s.
|
||
- Manual override on a shape the rule had marked ignore correctly flipped it
|
||
to keep via manual, and rule-version changed (735edd… → c47e02…).
|
||
- First-match-wins confirmed: a shape that is both low-score and tiny gets the
|
||
first rule's verdict, not the second's.
|
||
- Your data's actual distribution: score p05/p50/p95 = 0.37 / 0.61 / 0.82,
|
||
area p05/p50/p95 = 0.16% / 0.74% / 2.38%. That's where you'll be drawing
|
||
boundaries.
|
||
|
||
One thing I caught and fixed mid-build: the first version returned 21 MB of
|
||
JSON for one batch, and 85k SVG circles would have frozen the browser. Split
|
||
into a /summary endpoint (full-batch counts + a stable 4,000-point stride
|
||
sample → 450 KB) and a paged /shapes endpoint sorted server-side. The tallies
|
||
you see are always over all shapes; only the plot is sampled, and the page
|
||
says so.
|
||
|
||
Something you need to know about your data — this isn't my code. Your master
|
||
dataset is 1,390 images with zero annotations between them:
|
||
|
||
┌────────┬───────────┬────────┬─────────────┐
|
||
│ batch │ status │ frames │ annotations │
|
||
├────────┼───────────┼────────┼─────────────┤
|
||
│ batch6 │ reviewing │ 1,950 │ 85,520 │
|
||
├────────┼───────────┼────────┼─────────────┤
|
||
│ batch7 │ merged │ 1,390 │ 0 │
|
||
└────────┴───────────┴────────┴─────────────┘
|
||
|
||
The only batch you've merged is batch7, which has no annotations at all — so
|
||
every label file in dataset/labels/ is empty, and the entire master dataset is
|
||
negative samples. Training on it teaches the model that sacks don't exist.
|
||
Meanwhile batch6, the one carrying all 85k annotations, is still sitting in
|
||
reviewing and has never been merged. My guess is "approve all frames" was used
|
||
on batch7 before auto-annotation ran — approve() only checks that frames are
|
||
approved, never that any of them carry a shape. Worth a guard, and worth
|
||
telling me whether that merge was intentional.
|
||
|
||
What's there now: backend/triage.py (resolver, rules, overrides, versioning),
|
||
backend/api/triage.py (7 endpoints incl. /crop), two new tables + rule_version
|
||
on model_versions, and the rebuilt Data Prep page — rule editor, score×area
|
||
scatter with drag-select, crop grid, bulk verdict bar. Resolution plugs into
|
||
_build_selected_tree, so it's the one place that decides what training sees.
|
||
All files under the 400-line limit; frontend builds clean.
|
||
|
||
What I have not verified: I tested the API and the build, not the page in a
|
||
browser, and no training run has consumed a rule-filtered dataset yet. The
|
||
REQs are still sitting in docs/proposal-dataprep-triage.md — I'll move them
|
||
into requirements.md and design.md once you've clicked around and confirmed it
|
||
does what you meant.
|
||
|
||
Run it and tell me how it feels. The obvious next question is whether score or
|
||
aspect ratio actually separates full-sack from half-sack in your footage —
|
||
the scatter should answer that within a minute of looking at it.
|
||
|
||
✻ Cooked for 9m 9s
|
||
|
||
❯ now from where my data preparation data comes? and how to revert or choose
|
||
which batch we want to do data preparation and pass to models and training.
|
||
also i still confuse about the flow. because the UI doesnt reflect what you
|
||
said
|
||
● Fair — let me check what's actually wired rather than assume.
|
||
|
||
Searched for 2 patterns, ran 1 shell command (ctrl+o to expand)
|
||
|
||
● That explains most of it:
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Confirmed — you're looking at an hour-old build. Both containers predate
|
||
everything I wrote. The backend has no /triage routes (404) and the frontend
|
||
is serving the old Data Prep page. Nothing I described exists in what's on
|
||
your screen. That's on me for not saying "rebuild to see it."
|
||
|
||
Now your actual questions.
|
||
|
||
Where does Data Prep's data come from?
|
||
|
||
The annotations table — every shape SAM3 wrote for the frames of the batch you
|
||
pick. Not from disk, not from the master dataset. That's why it can show you
|
||
85,520 shapes for batch6 even though batch6 has never been merged.
|
||
|
||
But the four stat cards at the top (trainable images, excluded, …) come from
|
||
somewhere else: dataset_items, the merged master dataset. That split is real
|
||
and it's confusing, and it's the second reason the page won't make sense to
|
||
you even after a rebuild. More below.
|
||
|
||
The flow, honestly:
|
||
|
||
archive → batch → extract frames → SAM3 auto-annotate → review/approve
|
||
↓
|
||
APPROVE BATCH
|
||
↓
|
||
merge → master dataset
|
||
↓
|
||
triage rules applied here, at train
|
||
time
|
||
↓
|
||
Models page: pick batches +
|
||
classes → train
|
||
|
||
Choosing which batch goes to training already exists — it's on the Models
|
||
page, not Data Prep (ModelsPage.jsx:306, "Select Batches"). But it only lists
|
||
merged batches, because dataset.summary returns WHERE b.status = 'merged'.
|
||
Same for class selection at ModelsPage.jsx:235.
|
||
|
||
And here's the gap I have to own. You asked for triage per batch, before
|
||
merge. I built the UI that way — but I wired enforcement at training time,
|
||
which only ever sees merged data. So right now:
|
||
|
||
- batch6 has all 85k shapes, is reviewing, never merged → you can triage it,
|
||
and it will affect nothing
|
||
- batch7 has 0 shapes, is merged → it's the only thing training can see, and
|
||
it's empty
|
||
|
||
Triaging batch6 today changes no training run. The page doesn't tell you that.
|
||
That's a genuine design hole between what you asked for and what I built, not
|
||
just a stale container.
|
||
|
||
Reverting: nothing I built is destructive. "Clear hand decisions" removes
|
||
overrides for selected shapes; deleting a rule and saving un-applies it;
|
||
annotations.class_id is never rewritten, so any triage decision is reversible
|
||
at any time. What is not currently revertible is a merge — once a batch is
|
||
merged the only way back is deleting the batch (which now also cleans its
|
||
files from the dataset).
|
||
|
||
Two things to decide:
|
||
|
||
● User answered Claude's questions:
|
||
⎿ · Rebuild the containers so you can actually see it? This restarts the app.
|
||
→ yes rebuild every time it needed to reflect the UI
|
||
· How should triage relate to merging? This is the gap that's confusing
|
||
you. → Triage decides what gets merged (Recommended)
|
||
|
||
● Both. First, check nothing's running before I rebuild:
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Nothing running. Now the merge gate — triage decides what enters the master
|
||
dataset:
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 9 lines
|
||
479 job.progress(0, len(frames))
|
||
480 job.log(f"Merging {len(frames)} approved frame(s) into the
|
||
master dataset")
|
||
481
|
||
482 + # Triage gates the merge (REQ-104): a frame carrying an ig
|
||
+nored shape never
|
||
483 + # enters the master dataset at all, so junk is stopped at
|
||
+the door rather
|
||
484 + # than filtered out on every later training run.
|
||
485 + resolver = triage.Resolver(project["id"])
|
||
486 + gating = bool(resolver.rules or resolver.overrides)
|
||
487 + if gating:
|
||
488 + job.log(f"Applying {len(resolver.rules)} triage rule(s
|
||
+), version {resolver.version()}")
|
||
489 +
|
||
490 added = {"train": 0, "val": 0}
|
||
491 skipped = 0
|
||
492 + triaged_out = 0
|
||
493 cancelled = False
|
||
494 for index, frame in enumerate(frames):
|
||
495 if job.cancelled:
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 18 lines
|
||
497 cancelled = True
|
||
498 break
|
||
499
|
||
500 + annotations = review.listing(frame["id"])
|
||
501 + if gating:
|
||
502 + resolved = []
|
||
503 + for item in annotations:
|
||
504 + shape = {"id": item["id"], "class_id": item["c
|
||
+lass_id"],
|
||
505 + "score": float(item.get("score") or 1
|
||
+.0),
|
||
506 + **triage.metrics(item["geometry"])}
|
||
507 + effective = resolver.effective_class(shape)
|
||
508 + if effective is None:
|
||
509 + resolved = None
|
||
510 + break
|
||
511 + resolved.append({**item, "class_id": effective
|
||
+})
|
||
512 + if resolved is None:
|
||
513 + triaged_out += 1
|
||
514 + job.progress(index + 1, len(frames))
|
||
515 + continue
|
||
516 + annotations = resolved
|
||
517 +
|
||
518 with db.cursor() as cur:
|
||
519 cur.execute("SELECT 1 FROM dataset_items WHERE fra
|
||
me_id = ?", (frame["id"],))
|
||
520 if cur.fetchone() is not None:
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 1 line, removed 1 line
|
||
530 os.path.join(root, image_rel))
|
||
531
|
||
532 lines = [_label_line(item["class_id"], item["geome
|
||
try"], project["label_type"])
|
||
533 - for item in review.listing(frame["id"])]
|
||
533 + for item in annotations]
|
||
534 # An approved frame with nothing on it is a negati
|
||
ve sample, and an
|
||
535 # empty .txt is how YOLO spells that (REQ-033).
|
||
536 with open(os.path.join(root, label_rel), "w", enco
|
||
ding="utf-8") as handle:
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 2 lines, removed 1 line
|
||
560 path = write_data_yaml(projects.get(project["id"]))
|
||
561 totals = summary(project["id"])["splits"]
|
||
562 job.log(f"Added {added['train']} train / {added['val']} va
|
||
l"
|
||
563 - + (f", skipped {skipped} already merged" if skippe
|
||
-d else ""))
|
||
563 + + (f", skipped {skipped} already merged" if skippe
|
||
+d else "")
|
||
564 + + (f", held back {triaged_out} by triage" if triag
|
||
+ed_out else ""))
|
||
565 job.log(f"Master dataset now {totals['train']} train / {to
|
||
tals['val']} val — {path}")
|
||
|
||
● Now make the page tell you what will actually happen, per batch:
|
||
|
||
● Update(backend/triage.py)
|
||
⎿ Added 3 lines
|
||
240 batch, frame_ids, shapes, resolver = _resolved_shapes(batc
|
||
h_id)
|
||
241
|
||
242 counts = {"keep": 0, "ignore": 0, "reclass": 0, "manual":
|
||
0}
|
||
243 + ignored_frames = set()
|
||
244 for shape in shapes:
|
||
245 counts[shape["verdict"]] += 1
|
||
246 if shape["source"] == "manual":
|
||
247 counts["manual"] += 1
|
||
248 + if shape["verdict"] == "ignore":
|
||
249 + ignored_frames.add(shape["frame_id"])
|
||
250
|
||
251 # An even stride rather than a random draw: the sample is
|
||
stable across
|
||
252 # reloads, so points do not jump around while the user is
|
||
reading the plot.
|
||
|
||
● Update(backend/triage.py)
|
||
⎿ Added 5 lines
|
||
259 return {
|
||
260 "batch_id": batch_id,
|
||
261 "project_id": batch["project_id"],
|
||
262 + "status": batch["status"],
|
||
263 + "merged": batch["status"] == "merged",
|
||
264 "frame_count": len(frame_ids),
|
||
265 "total_shapes": len(shapes),
|
||
266 "counts": counts,
|
||
267 + # What merging this batch would do right now (REQ-104)
|
||
+.
|
||
268 + "frames_held_back": len(ignored_frames),
|
||
269 + "frames_would_merge": len(frame_ids) - len(ignored_fra
|
||
+mes),
|
||
270 "sample": sample,
|
||
271 "sampled": len(sample) < len(shapes),
|
||
272 "rule_version": resolver.version(),
|
||
|
||
● Now the page needs to state where the batch stands in the flow, since that was
|
||
the core confusion:
|
||
|
||
● Update(frontend/src/pages/DataPrepPage.jsx)
|
||
⎿ Added 31 lines, removed 1 line
|
||
156 </p>
|
||
157 )}
|
||
158
|
||
159 + {view && (
|
||
160 + <div
|
||
161 + style={{
|
||
162 + display: 'flex',
|
||
163 + alignItems: 'center',
|
||
164 + gap: 10,
|
||
165 + flexWrap: 'wrap',
|
||
166 + padding: '10px 14px',
|
||
167 + marginBottom: 16,
|
||
168 + borderRadius: 8,
|
||
169 + fontSize: '0.82rem',
|
||
170 + background: view.merged ? 'rgba(34,197,94,0.08)' :
|
||
+ 'rgba(250,204,21,0.08)',
|
||
171 + border: `1px solid ${view.merged ? 'rgba(34,197,94
|
||
+,0.3)' : 'rgba(250,204,21,0.3)'}`,
|
||
172 + }}
|
||
173 + >
|
||
174 + <strong>This batch is “{view.status}”.</strong>
|
||
175 + {view.merged ? (
|
||
176 + <span className="muted">
|
||
177 + It is already in the master dataset. Rules still
|
||
+ re-cut it every training run.
|
||
178 + </span>
|
||
179 + ) : (
|
||
180 + <span className="muted">
|
||
181 + Nothing here reaches training until the batch is
|
||
+ reviewed, approved and merged.
|
||
182 + On merge, <strong>{view.frames_would_merge}</str
|
||
+ong> of {view.frame_count} frames
|
||
183 + would go in and <strong>{view.frames_held_back}<
|
||
+/strong> would be held back.
|
||
184 + </span>
|
||
185 + )}
|
||
186 + </div>
|
||
187 + )}
|
||
188 +
|
||
189 {preview && (
|
||
190 <div style={{ display: 'grid', gridTemplateColumns: 'r
|
||
epeat(auto-fit, minmax(180px, 1fr))', gap: 14, marginBottom: 1
|
||
8 }}>
|
||
161 - <Stat label="Trainable images" value={preview.traina
|
||
-ble_images} hint={`of ${preview.total_images} merged`} accent=
|
||
-"#4ade80" />
|
||
191 + <Stat label="Trainable images" value={preview.traina
|
||
+ble_images} hint={`of ${preview.total_images} in master datase
|
||
+t`} accent="#4ade80" />
|
||
192 <Stat label="Excluded images" value={preview.exclude
|
||
d_images} hint="carry an ignored shape" accent="#f87171" />
|
||
193 <Stat label="Reclassed shapes" value={preview.reclas
|
||
s} hint="trained as another class" accent="#c084fc" />
|
||
194 <Stat label="Rule version" value={preview.rule_versi
|
||
on} hint="stamped on each run" accent="#38bdf8" mono />
|
||
|
||
● Now let the page show all annotated batches regardless of merge state, and
|
||
label them:
|
||
|
||
● Update(frontend/src/pages/DataPrepPage.jsx)
|
||
⎿ Added 1 line, removed 1 line
|
||
219 {batches.length === 0 && <option value="">no ann
|
||
otated batches</option>}
|
||
220 {batches.map((batch) => (
|
||
221 <option key={batch.id} value={batch.id}>
|
||
222 - {batch.date_label}/{batch.batch_label} · {ba
|
||
-tch.annotation_count} shapes
|
||
222 + {batch.date_label}/{batch.batch_label} · {ba
|
||
+tch.annotation_count} shapes · {batch.status}
|
||
223 </option>
|
||
224 ))}
|
||
225 </select>
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Now rebuilding the containers:
|
||
|
||
Ran 3 shell commands (ctrl+o to expand)
|
||
|
||
● Rebuilt and live — frontend is serving the new bundle (index-Di3kDxxW.js),
|
||
triage endpoints answer 200. But note: batch6 is now merged — it was reviewing
|
||
an hour ago. Let me re-check your dataset:
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
● Everything's live and your dataset changed — batch6 merged while we talked:
|
||
|
||
Master dataset now: 3,340 images (2,646 train / 694 val — 20.8%, the hash
|
||
split working correctly)
|
||
- batch6 → 1,950 images carrying all 85,520 labels
|
||
- batch7 → 1,390 images with empty label files
|
||
|
||
Reload the browser with a hard refresh (Ctrl+Shift+R) — the bundle hash
|
||
changed to index-Di3kDxxW.js, and a cached old bundle is exactly what made the
|
||
page look wrong before.
|
||
|
||
What changed so the flow makes sense
|
||
|
||
Triage now gates the merge, as you chose. _run_merge resolves every frame
|
||
before copying it: a frame carrying an ignore never enters the master dataset,
|
||
and a reclass is written under its target class in the label file. The job
|
||
log tells you ("held back N by triage"). Rules still re-apply at training
|
||
time, so editing them later re-cuts what's already merged — you get both, and
|
||
the earlier contradiction is gone.
|
||
|
||
The page now opens with a status banner that says exactly where the selected
|
||
batch stands: whether it's merged, and for an unmerged one, "on merge, X of Y
|
||
frames would go in and Z would be held back." That was the missing sentence.
|
||
The batch dropdown shows status too.
|
||
|
||
Your flow, concretely:
|
||
|
||
1. Batches page → extract frames → auto-annotate (SAM3 writes annotations)
|
||
2. Data Prep → pick that batch, drag on the scatter / click crops, write
|
||
rules, hand-decide the rest
|
||
3. Review → approve frames
|
||
4. Approve batch → merge. Triage applies here — junk never enters
|
||
5. Models page → "Select Batches" + "Select Classes" → train. Rules re-apply,
|
||
so a rule edit changes this run without touching what's merged
|
||
|
||
Choosing which batch trains lives on the Models page, not Data Prep — it lists
|
||
merged batches only. That page already worked; it just wasn't obvious it was
|
||
the answer to your question.
|
||
|
||
Reverting: rules and hand-decisions are fully reversible (delete the rule and
|
||
save; "Clear hand decisions" for overrides) because annotations.class_id is
|
||
never rewritten. A merge is still the one irreversible step — deleting the
|
||
batch is the only way back, which now also removes its files from the dataset.
|
||
|
||
One thing to decide: 1,390 of your 3,340 images (42%) are batch7's empty label
|
||
files. If that footage really is empty conveyor, they're legitimate negative
|
||
samples and help. If those frames actually contain sacks and were merged
|
||
before annotation ran, they're actively teaching the model to miss sacks — and
|
||
at 42% of the dataset that will dominate training. Worth eyeballing a few
|
||
batch7 frames before you train. Tell me which it is and I'll either leave it
|
||
or clean it up.
|
||
|
||
I haven't clicked through the page myself — I verified the API against live
|
||
data and the bundle is served. Reload and tell me what's still off.
|
||
|
||
✻ Baked for 2m 57s
|
||
|
||
※ recap: Goal is getting your retraining loop reliable, and the Data Prep triage gate is now built, rebuilt into the containers, and live. Next: hard-refresh the browser, click through the page, and tell me whether batch7's 1,390 empty frames are real empty conveyor or unannotated sacks. (disable recaps in /config)
|
||
|
||
❯ can you create a default rule, or try different filtering approach rule,
|
||
because i really confuse when try filtering especially in UI
|
||
|
||
● Right — three numeric range boxes with no feedback is a bad way to ask someone
|
||
to find a threshold. Let me fix the approach: presets computed from your
|
||
actual data, plus live impact as you edit.
|
||
|
||
● Update(backend/triage.py)
|
||
⎿ Added 123 lines
|
||
286 return {"total": len(shapes), "offset": offset, "limit": l
|
||
imit, "shapes": page}
|
||
287
|
||
288
|
||
289 +def _percentile(values: list, fraction: float) -> float:
|
||
290 + if not values:
|
||
291 + return 0.0
|
||
292 + return values[min(len(values) - 1, int(len(values) * fract
|
||
+ion))]
|
||
293 +
|
||
294 +
|
||
295 +def suggest(batch_id: int) -> dict:
|
||
296 + """Presets with thresholds read off this batch's own distr
|
||
+ibution.
|
||
297 +
|
||
298 + Asking someone to invent "score below 0.45" from nothing i
|
||
+s guesswork. The
|
||
299 + same question is easy when the number comes from their dat
|
||
+a and the effect
|
||
300 + is stated: "the weakest 10% of detections — 8,552 shapes".
|
||
301 + """
|
||
302 + _, frame_ids, shapes, _ = _resolved_shapes(batch_id)
|
||
303 + if not shapes:
|
||
304 + return {"presets": [], "stats": {}}
|
||
305 +
|
||
306 + scores = sorted(shape["score"] for shape in shapes)
|
||
307 + areas = sorted(shape["area_pct"] for shape in shapes)
|
||
308 + total = len(shapes)
|
||
309 +
|
||
310 + def impact(predicate: dict) -> dict:
|
||
311 + matched = [s for s in shapes if _matches(predicate, s)
|
||
+]
|
||
312 + frames = {s["frame_id"] for s in matched}
|
||
313 + return {"shapes": len(matched), "frames": len(frames)}
|
||
314 +
|
||
315 + presets = []
|
||
316 +
|
||
317 + weak = round(_percentile(scores, 0.10), 3)
|
||
318 + presets.append({
|
||
319 + "key": "drop-weakest",
|
||
320 + "title": "Ignore the weakest detections",
|
||
321 + "blurb": f"SAM3 scored these below {weak} — the bottom
|
||
+ 10% of this batch.",
|
||
322 + "rule": {"name": "low confidence", "predicate": {"scor
|
||
+e": [None, weak]}, "action": "ignore"},
|
||
323 + "impact": impact({"score": [None, weak]}),
|
||
324 + })
|
||
325 +
|
||
326 + specks = round(_percentile(areas, 0.05), 3)
|
||
327 + presets.append({
|
||
328 + "key": "drop-specks",
|
||
329 + "title": "Ignore tiny specks",
|
||
330 + "blurb": f"Boxes smaller than {specks}% of the frame —
|
||
+ usually noise, not objects.",
|
||
331 + "rule": {"name": "specks", "predicate": {"area_pct": [
|
||
+None, specks]}, "action": "ignore"},
|
||
332 + "impact": impact({"area_pct": [None, specks]}),
|
||
333 + })
|
||
334 +
|
||
335 + median_area = round(_percentile(areas, 0.50), 3)
|
||
336 + presets.append({
|
||
337 + "key": "split-by-size",
|
||
338 + "title": "Split by size into a second class",
|
||
339 + "blurb": f"Everything under {median_area}% area (half
|
||
+this batch) becomes another class — "
|
||
340 + "pick which one. Size tracks distance from th
|
||
+e camera as much as object type, "
|
||
341 + "so check the crops before trusting it.",
|
||
342 + "rule": {"name": "small ones", "predicate": {"area_pct
|
||
+": [None, median_area]},
|
||
343 + "action": "reclass", "target_class": None},
|
||
344 + "impact": impact({"area_pct": [None, median_area]}),
|
||
345 + "needs_target": True,
|
||
346 + })
|
||
347 +
|
||
348 + tall = round(_percentile(sorted(s["aspect"] for s in shape
|
||
+s), 0.15), 3)
|
||
349 + presets.append({
|
||
350 + "key": "odd-shapes",
|
||
351 + "title": "Ignore oddly-shaped boxes",
|
||
352 + "blurb": f"Aspect ratio under {tall} — long thin slive
|
||
+rs, usually a bad mask.",
|
||
353 + "rule": {"name": "slivers", "predicate": {"aspect": [N
|
||
+one, tall]}, "action": "ignore"},
|
||
354 + "impact": impact({"aspect": [None, tall]}),
|
||
355 + })
|
||
356 +
|
||
357 + return {
|
||
358 + "presets": presets,
|
||
359 + "stats": {
|
||
360 + "total_shapes": total,
|
||
361 + "total_frames": len(frame_ids),
|
||
362 + "score": {"p05": round(_percentile(scores, 0.05),
|
||
+3),
|
||
363 + "p50": round(_percentile(scores, 0.50),
|
||
+3),
|
||
364 + "p95": round(_percentile(scores, 0.95),
|
||
+3)},
|
||
365 + "area_pct": {"p05": round(_percentile(areas, 0.05)
|
||
+, 3),
|
||
366 + "p50": round(_percentile(areas, 0.50)
|
||
+, 3),
|
||
367 + "p95": round(_percentile(areas, 0.95)
|
||
+, 3)},
|
||
368 + },
|
||
369 + }
|
||
370 +
|
||
371 +
|
||
372 +def simulate(batch_id: int, candidate_rules: List[dict]) -> di
|
||
+ct:
|
||
373 + """What these rules would do, without saving them.
|
||
374 +
|
||
375 + Editing a threshold and seeing the number move is the whol
|
||
+e difference
|
||
376 + between tuning a filter and guessing at one.
|
||
377 + """
|
||
378 + _, frame_ids, shapes, _ = _resolved_shapes(batch_id)
|
||
379 + overrides = Resolver(0).overrides if False else None # ke
|
||
+pt simple: rules only
|
||
380 +
|
||
381 + with db.cursor() as cur:
|
||
382 + cur.execute("SELECT annotation_id, verdict, target_cla
|
||
+ss FROM annotation_overrides")
|
||
383 + manual = {row[0]: (row[1], row[2]) for row in cur.fetc
|
||
+hall()}
|
||
384 +
|
||
385 + counts = {"keep": 0, "ignore": 0, "reclass": 0}
|
||
386 + per_rule = [0] * len(candidate_rules)
|
||
387 + ignored_frames = set()
|
||
388 +
|
||
389 + for shape in shapes:
|
||
390 + if shape["id"] in manual:
|
||
391 + verdict = manual[shape["id"]][0]
|
||
392 + else:
|
||
393 + verdict = "keep"
|
||
394 + for index, rule in enumerate(candidate_rules):
|
||
395 + if _matches(rule.get("predicate") or {}, shape
|
||
+):
|
||
396 + verdict = rule["action"]
|
||
397 + per_rule[index] += 1
|
||
398 + break
|
||
399 + counts[verdict] += 1
|
||
400 + if verdict == "ignore":
|
||
401 + ignored_frames.add(shape["frame_id"])
|
||
402 +
|
||
403 + return {
|
||
404 + "total_shapes": len(shapes),
|
||
405 + "counts": counts,
|
||
406 + "per_rule": per_rule,
|
||
407 + "frames_held_back": len(ignored_frames),
|
||
408 + "frames_would_merge": len(frame_ids) - len(ignored_fra
|
||
+mes),
|
||
409 + }
|
||
410 +
|
||
411 +
|
||
412 def preview(project_id: int) -> dict:
|
||
413 """What the current rules would do to the whole merged dat
|
||
aset."""
|
||
414 with db.cursor() as cur:
|
||
|
||
● Update(backend/triage.py)
|
||
⎿ Removed 2 lines
|
||
376 between tuning a filter and guessing at one.
|
||
377 """
|
||
378 _, frame_ids, shapes, _ = _resolved_shapes(batch_id)
|
||
379 - overrides = Resolver(0).overrides if False else None # ke
|
||
-pt simple: rules only
|
||
380 -
|
||
379 with db.cursor() as cur:
|
||
380 cur.execute("SELECT annotation_id, verdict, target_cla
|
||
ss FROM annotation_overrides")
|
||
381 manual = {row[0]: (row[1], row[2]) for row in cur.fetc
|
||
hall()}
|
||
|
||
● Update(backend/api/triage.py)
|
||
⎿ Added 16 lines
|
||
79 return {"cleared": triage.clear_overrides(body.annotation_
|
||
ids)}
|
||
80
|
||
81
|
||
82 +@router.get("/api/batches/{batch_id}/triage/suggest")
|
||
83 +def suggest(batch_id: int) -> dict:
|
||
84 + try:
|
||
85 + return triage.suggest(batch_id)
|
||
86 + except triage.TriageError as exc:
|
||
87 + raise HTTPException(404, str(exc))
|
||
88 +
|
||
89 +
|
||
90 +@router.post("/api/batches/{batch_id}/triage/simulate")
|
||
91 +def simulate(batch_id: int, body: RuleList) -> dict:
|
||
92 + try:
|
||
93 + return triage.simulate(batch_id, [item.model_dump() fo
|
||
+r item in body.rules])
|
||
94 + except triage.TriageError as exc:
|
||
95 + raise HTTPException(400, str(exc))
|
||
96 +
|
||
97 +
|
||
98 @router.get("/api/projects/{project_id}/triage/preview")
|
||
99 def preview(project_id: int) -> dict:
|
||
100 return triage.preview(project_id)
|
||
|
||
Ran 2 shell commands (ctrl+o to expand)
|
||
|
||
● Restart isn't enough — the image COPYs backend/, so it needs a rebuild:
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Presets work. But look closely — this exposes a serious problem with the
|
||
design I recommended:
|
||
|
||
"Ignore the weakest 10%" = 8,552 shapes spread across 1,882 of your 1,950
|
||
frames. Because an ignored shape excludes its whole image, that rule would
|
||
hold back 96% of the batch.
|
||
|
||
Your frames average 44 sacks each. At that density, nearly every frame
|
||
contains at least one weak detection, so "exclude the whole image" destroys
|
||
the dataset. Let me build the UI, then we need to fix that.
|
||
|
||
● Update(frontend/src/api.js)
|
||
⎿ Added 3 lines
|
||
125 triageShapes: (batchId, { sort = 'score', offset = 0, limit
|
||
= 120 } = {}) =>
|
||
126 request(`/batches/${batchId}/triage/shapes?sort=${sort}&of
|
||
fset=${offset}&limit=${limit}`),
|
||
127 triagePreview: (projectId) => request(`/projects/${projectId
|
||
}/triage/preview`),
|
||
128 + triageSuggest: (batchId) => request(`/batches/${batchId}/tri
|
||
+age/suggest`),
|
||
129 + triageSimulate: (batchId, rules) =>
|
||
130 + request(`/batches/${batchId}/triage/simulate`, { method: '
|
||
+POST', body: { rules } }),
|
||
131 setTriageOverrides: (annotationIds, verdict, targetClass = n
|
||
ull) =>
|
||
132 request('/triage/overrides', {
|
||
133 method: 'POST',
|
||
|
||
● Write(frontend/src/components/TriagePresets.jsx)
|
||
⎿ Wrote 59 lines to frontend/src/components/TriagePresets.jsx
|
||
1 import React from 'react'
|
||
3 /* Starting points, with thresholds read off this batch's own di
|
||
stribution and
|
||
4 * the cost stated up front.
|
||
6 * Inventing "score below 0.45" from a blank box is guesswork; p
|
||
icking "the
|
||
ision. */
|
||
8
|
||
9 export default function TriagePresets({ presets, onAdd, classes,
|
||
10 if (!presets?.length) return null
|
||
… +64 lines (ctrl+o to expand)
|
||
|
||
● Now live impact per rule as you edit:
|
||
|
||
● Update(frontend/src/components/TriageRules.jsx)
|
||
⎿ Added 1 line, removed 1 line
|
||
26 color: '#e4e4e7',
|
||
27 }
|
||
28
|
||
29 -export default function TriageRules({ rules, classes, onChange,
|
||
- onSave, saving, dirty }) {
|
||
29 +export default function TriageRules({ rules, classes, onChange,
|
||
+ onSave, saving, dirty, simulation, children }) {
|
||
30 function update(index, patch) {
|
||
31 onChange(rules.map((rule, i) => (i === index ? { ...rule, .
|
||
..patch } : rule)))
|
||
32 }
|
||
|
||
● Update(frontend/src/components/TriageRules.jsx)
|
||
⎿ Added 24 lines
|
||
81 decided by hand ignore these rules entirely.
|
||
82 </p>
|
||
83
|
||
84 + {children}
|
||
85 +
|
||
86 + {simulation && rules.length > 0 && (
|
||
87 + <p
|
||
88 + style={{
|
||
89 + fontSize: '0.82rem',
|
||
90 + padding: '8px 12px',
|
||
91 + marginBottom: 12,
|
||
92 + borderRadius: 6,
|
||
93 + background: 'rgba(56,189,248,0.08)',
|
||
94 + border: '1px solid rgba(56,189,248,0.25)',
|
||
95 + }}
|
||
96 + >
|
||
97 + As written, these rules keep <strong>{simulation.cou
|
||
+nts.keep.toLocaleString()}</strong> shapes,
|
||
98 + ignore <strong>{simulation.counts.ignore.toLocaleStr
|
||
+ing()}</strong>, reclass{' '}
|
||
99 + <strong>{simulation.counts.reclass.toLocaleString()}
|
||
+</strong> — and{' '}
|
||
100 + <strong style={{ color: simulation.frames_held_back
|
||
+> simulation.frames_would_merge ? '#f87171' : '#4ade80' }}>
|
||
101 + {simulation.frames_would_merge.toLocaleString()}
|
||
102 + </strong>{' '}
|
||
103 + of {(simulation.frames_would_merge + simulation.fram
|
||
+es_held_back).toLocaleString()} frames would
|
||
104 + reach training.
|
||
105 + </p>
|
||
106 + )}
|
||
107 +
|
||
108 {rules.length === 0 && <p className="empty">No rules — e
|
||
very shape is kept as its own class.</p>}
|
||
109
|
||
110 <div style={{ display: 'flex', flexDirection: 'column',
|
||
gap: 10 }}>
|
||
|
||
● Update(frontend/src/components/TriageRules.jsx)
|
||
⎿ Added 6 lines
|
||
180 </select>
|
||
181 )}
|
||
182
|
||
183 + {simulation?.per_rule?.[index] !== undefined && (
|
||
184 + <span className="mono faint" style={{ fontSize:
|
||
+'0.75rem' }}>
|
||
185 + matches {simulation.per_rule[index].toLocaleSt
|
||
+ring()}
|
||
186 + </span>
|
||
187 + )}
|
||
188 +
|
||
189 <span style={{ marginLeft: 'auto', display: 'flex'
|
||
, gap: 4 }}>
|
||
190 <button type="button" className="btn" aria-label
|
||
="Move rule up" onClick={() => move(index, -1)} disabled={inde
|
||
x === 0} style={{ cursor: 'pointer', padding: '2px 8px' }}>↑</
|
||
button>
|
||
191 <button type="button" className="btn" aria-label
|
||
="Move rule down" onClick={() => move(index, 1)} disabled={ind
|
||
ex === rules.length - 1} style={{ cursor: 'pointer', padding:
|
||
'2px 8px' }}>↓</button>
|
||
|
||
● Update(frontend/src/pages/DataPrepPage.jsx)
|
||
⎿ Added 2 lines
|
||
31 const [selectedIds, setSelectedIds] = useState([])
|
||
32 const [reloadKey, setReloadKey] = useState(0)
|
||
33 const [reclassTarget, setReclassTarget] = useState('')
|
||
34 + const [presets, setPresets] = useState([])
|
||
35 + const [simulation, setSimulation] = useState(null)
|
||
36 const [saving, setSaving] = useState(false)
|
||
37 const [busy, setBusy] = useState(false)
|
||
38 const [error, setError] = useState('')
|
||
|
||
● Update(frontend/src/pages/DataPrepPage.jsx)
|
||
⎿ Added 22 lines
|
||
79
|
||
80 useEffect(() => { loadBatch() }, [loadBatch])
|
||
81
|
||
82 + useEffect(() => {
|
||
83 + if (!batchId) return
|
||
84 + let cancelled = false
|
||
85 + api.triageSuggest(batchId)
|
||
86 + .then((data) => { if (!cancelled) setPresets(data.preset
|
||
+s) })
|
||
87 + .catch(() => { if (!cancelled) setPresets([]) })
|
||
88 + return () => { cancelled = true }
|
||
89 + }, [batchId])
|
||
90 +
|
||
91 + // Debounced: the point is watching the number move as you d
|
||
+rag a threshold,
|
||
92 + // and a request per keystroke would just queue up behind it
|
||
+self.
|
||
93 + useEffect(() => {
|
||
94 + if (!batchId) return
|
||
95 + let cancelled = false
|
||
96 + const timer = setTimeout(() => {
|
||
97 + api.triageSimulate(batchId, rules)
|
||
98 + .then((data) => { if (!cancelled) setSimulation(data)
|
||
+})
|
||
99 + .catch(() => { if (!cancelled) setSimulation(null) })
|
||
100 + }, 350)
|
||
101 + return () => { cancelled = true; clearTimeout(timer) }
|
||
102 + }, [batchId, rules])
|
||
103 +
|
||
104 // Counted server-side over every shape in the batch — the s
|
||
catter plots a
|
||
105 // sample, but these numbers are never a sample.
|
||
106 const counts = view?.counts ?? { keep: 0, ignore: 0, reclass
|
||
: 0, manual: 0 }
|
||
|
||
● Update(frontend/src/pages/DataPrepPage.jsx)
|
||
⎿ Added 9 lines, removed 1 line
|
||
226 onSave={saveRules}
|
||
227 saving={saving}
|
||
228 dirty={dirty}
|
||
229 - />
|
||
229 + simulation={simulation}
|
||
230 + >
|
||
231 + <TriagePresets
|
||
232 + presets={presets}
|
||
233 + classes={project.classes}
|
||
234 + frameCount={view?.frame_count ?? 0}
|
||
235 + onAdd={(rule) => setRules((current) => [...current,
|
||
+rule])}
|
||
236 + />
|
||
237 + </TriageRules>
|
||
238
|
||
239 <div className="panel table-wrap" style={{ padding: 18,
|
||
marginTop: 18 }}>
|
||
240 <div style={{ display: 'flex', justifyContent: 'space-
|
||
between', alignItems: 'center', gap: 12, flexWrap: 'wrap', mar
|
||
ginBottom: 14 }}>
|
||
|
||
● Update(frontend/src/pages/DataPrepPage.jsx)
|
||
⎿ Added 1 line
|
||
3
|
||
4 import { AlertIcon, DatabaseIcon, RocketIcon, SlidersIcon } from
|
||
'../components/Icons'
|
||
5 import TriageCropGrid from '../components/TriageCropGrid'
|
||
6 +import TriagePresets from '../components/TriagePresets'
|
||
7 import TriageRules from '../components/TriageRules'
|
||
8 import TriageScatter from '../components/TriageScatter'
|
||
9
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
● That's the proof: 68 of 1,950 frames would survive. I need to correct the
|
||
recommendation I gave you earlier.
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Rebuilt and live (index-4bSN9aPo.js). Hard-refresh the browser.
|
||
|
||
What's on the page now, to fix the filtering confusion:
|
||
|
||
- Four preset cards with thresholds computed from your batch, each stating its
|
||
cost before you click: "SAM3 scored these below 0.399 — the bottom 10%. 8,552
|
||
shapes · 1,882 frames." Click "Add this rule" and it drops into the editor,
|
||
pre-filled and editable.
|
||
- Live impact, debounced 350 ms: a blue line under the rules saying how many
|
||
shapes are kept/ignored/reclassed and how many frames would actually reach
|
||
training, plus a matches N counter on each individual rule row. You now watch
|
||
the number move as you drag a threshold instead of guessing.
|
||
- Presets turn their count red when an ignore rule would take most of the
|
||
batch.
|
||
|
||
Now the correction. Earlier I recommended "an ignored shape excludes its whole
|
||
image", and you agreed. Your data proves that was wrong advice for your case:
|
||
|
||
┌───────────────┬───────────────────────┬──────────────────────────┐
|
||
│ Rule │ Shapes ignored │ Frames reaching training │
|
||
├───────────────┼───────────────────────┼──────────────────────────┤
|
||
│ score < 0.399 │ 8,552 of 85,520 (10%) │ 68 of 1,950 │
|
||
└───────────────┴───────────────────────┴──────────────────────────┘
|
||
|
||
Your frames carry ~44 sacks each. Almost every frame contains at least one
|
||
weak detection, so excluding whole images throws away 96% of the data to
|
||
remove 10% of the boxes. My reasoning was sound for sparse scenes (1–2 objects
|
||
per frame, where an unlabeled object really does dominate that image's
|
||
signal) and wrong for dense ones. At 44 objects per frame, dropping 4 boxes
|
||
leaves 40 correctly labeled — the background-poisoning harm is small and
|
||
local, while the data loss is catastrophic.
|
||
|
||
So ignore needs to mean something different for you:
|
||
|
||
● User answered Claude's questions:
|
||
⎿ · What should 'ignore' do, now that we can see the cost on your real data?
|
||
→ Drop the box, keep the image (Recommended)
|
||
· Do you want a rule set applied by default to new projects? → No default,
|
||
presets are enough
|
||
|
||
● Changing ignore to drop the box and keep the image. One guard I'm adding: if a
|
||
frame loses every shape it had, the image is still excluded — otherwise it
|
||
becomes a false negative teaching the model that a frame full of sacks is
|
||
empty.
|
||
|
||
● Update(backend/triage.py)
|
||
⎿ Added 11 lines, removed 4 lines
|
||
12
|
||
13 keep the shape trains as its own class
|
||
14 reclass -> class_id the shape trains as a different class
|
||
(REQ-105)
|
||
15 - ignore the shape does not train, and takes i
|
||
-ts whole image
|
||
16 - with it (REQ-104) — an image with a k
|
||
-nown object left
|
||
17 - unlabeled would teach the model that
|
||
-object is
|
||
18 - background
|
||
15 + ignore the box is dropped; its image still t
|
||
+rains (REQ-104)
|
||
16 +
|
||
17 +`ignore` drops the box rather than the image because these fram
|
||
+es are dense —
|
||
18 +around 44 shapes each. Excluding the whole image was measured a
|
||
+gainst a real
|
||
19 +batch and cost 96% of it (1,882 of 1,950 frames) to remove 10%
|
||
+of the boxes.
|
||
20 +Dropping four boxes out of forty-four leaves the image overwhel
|
||
+mingly correct;
|
||
21 +dropping the image leaves nothing to train on.
|
||
22 +
|
||
23 +The exception is a frame that loses *every* shape it had: an em
|
||
+pty label file
|
||
24 +says "there is nothing here", and for a frame that was full of
|
||
+sacks that is a
|
||
25 +lie the model will learn. Those images are excluded.
|
||
26 """
|
||
27
|
||
28 import hashlib
|
||
|
||
● Update(backend/triage.py)
|
||
⎿ Added 21 lines, removed 1 line
|
||
166 "source": rule["name"]}
|
||
167 return {"verdict": "keep", "target_class": None, "sour
|
||
ce": "default"}
|
||
168
|
||
169 + def resolve_shapes(self, annotations: list) -> Optional[li
|
||
+st]:
|
||
170 + """Apply verdicts to one frame's annotations.
|
||
171 +
|
||
172 + Returns the surviving annotations with their effective
|
||
+ class, or None
|
||
173 + when the frame must not train at all — which now happe
|
||
+ns only if every
|
||
174 + shape was dropped.
|
||
175 + """
|
||
176 + kept = []
|
||
177 + for item in annotations:
|
||
178 + shape = {"id": item["id"], "class_id": item["class
|
||
+_id"],
|
||
179 + "score": float(item.get("score") or 1.0),
|
||
180 + **metrics(item["geometry"])}
|
||
181 + effective = self.effective_class(shape)
|
||
182 + if effective is None:
|
||
183 + continue
|
||
184 + kept.append({**item, "class_id": effective})
|
||
185 + if annotations and not kept:
|
||
186 + return None
|
||
187 + return kept
|
||
188 +
|
||
189 def effective_class(self, shape: dict) -> Optional[int]:
|
||
170 - """The class this shape trains as, or None when it mus
|
||
-t not train."""
|
||
190 + """The class this shape trains as, or None when the bo
|
||
+x is dropped."""
|
||
191 resolved = self.verdict(shape)
|
||
192 if resolved["verdict"] == "ignore":
|
||
193 return None
|
||
|
||
● Now update the two consumers to the new semantics:
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 4 lines, removed 14 lines
|
||
158 annotations = review.listing(frame_id)
|
||
159
|
||
160 if resolver is not None:
|
||
161 - resolved = []
|
||
162 - skip_image = False
|
||
163 - for item in annotations:
|
||
164 - shape = {"id": item["id"], "class_id": item["c
|
||
-lass_id"],
|
||
165 - "score": float(item.get("score") or 1
|
||
-.0),
|
||
166 - **triage.metrics(item["geometry"])}
|
||
167 - effective = resolver.effective_class(shape)
|
||
168 - if effective is None:
|
||
169 - # REQ-104: a known object left unlabeled w
|
||
-ould teach the
|
||
170 - # model it is background, so the image goe
|
||
-s too.
|
||
171 - skip_image = True
|
||
172 - break
|
||
173 - resolved.append({**item, "class_id": effective
|
||
-})
|
||
174 - if skip_image:
|
||
161 + resolved = resolver.resolve_shapes(annotations)
|
||
162 + if resolved is None:
|
||
163 + # Every shape was dropped. An empty label file
|
||
+ would tell the
|
||
164 + # model this frame is empty, which it is not (
|
||
+REQ-104).
|
||
165 excluded += 1
|
||
166 continue
|
||
167 annotations = resolved
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 1 line, removed 10 lines
|
||
489
|
||
490 annotations = review.listing(frame["id"])
|
||
491 if gating:
|
||
492 - resolved = []
|
||
493 - for item in annotations:
|
||
494 - shape = {"id": item["id"], "class_id": item["c
|
||
-lass_id"],
|
||
495 - "score": float(item.get("score") or 1
|
||
-.0),
|
||
496 - **triage.metrics(item["geometry"])}
|
||
497 - effective = resolver.effective_class(shape)
|
||
498 - if effective is None:
|
||
499 - resolved = None
|
||
500 - break
|
||
501 - resolved.append({**item, "class_id": effective
|
||
-})
|
||
492 + resolved = resolver.resolve_shapes(annotations)
|
||
493 if resolved is None:
|
||
494 triaged_out += 1
|
||
495 job.progress(index + 1, len(frames))
|
||
|
||
● Now the reporting has to count frames the new way — only frames losing all
|
||
shapes are held back:
|
||
|
||
● Update(backend/triage.py)
|
||
⎿ Added 7 lines, removed 3 lines
|
||
267 batch, frame_ids, shapes, resolver = _resolved_shapes(batc
|
||
h_id)
|
||
268
|
||
269 counts = {"keep": 0, "ignore": 0, "reclass": 0, "manual":
|
||
0}
|
||
270 - ignored_frames = set()
|
||
270 + per_frame = {}
|
||
271 for shape in shapes:
|
||
272 counts[shape["verdict"]] += 1
|
||
273 if shape["source"] == "manual":
|
||
274 counts["manual"] += 1
|
||
275 - if shape["verdict"] == "ignore":
|
||
276 - ignored_frames.add(shape["frame_id"])
|
||
275 + total, dropped = per_frame.get(shape["frame_id"], (0,
|
||
+0))
|
||
276 + per_frame[shape["frame_id"]] = (total + 1, dropped + (
|
||
+shape["verdict"] == "ignore"))
|
||
277
|
||
278 + # Only a frame that loses everything is held back; the res
|
||
+t keep training
|
||
279 + # with their surviving boxes.
|
||
280 + ignored_frames = {fid for fid, (total, dropped) in per_fra
|
||
+me.items() if total == dropped}
|
||
281 +
|
||
282 # An even stride rather than a random draw: the sample is
|
||
stable across
|
||
283 # reloads, so points do not jump around while the user is
|
||
reading the plot.
|
||
284 stride = max(1, len(shapes) // SCATTER_POINTS)
|
||
|
||
● Update(backend/triage.py)
|
||
⎿ Added 5 lines, removed 3 lines
|
||
413
|
||
414 counts = {"keep": 0, "ignore": 0, "reclass": 0}
|
||
415 per_rule = [0] * len(candidate_rules)
|
||
416 - ignored_frames = set()
|
||
416 + per_frame = {}
|
||
417
|
||
418 for shape in shapes:
|
||
419 if shape["id"] in manual:
|
||
...
|
||
426 per_rule[index] += 1
|
||
427 break
|
||
428 counts[verdict] += 1
|
||
429 - if verdict == "ignore":
|
||
430 - ignored_frames.add(shape["frame_id"])
|
||
429 + total, dropped = per_frame.get(shape["frame_id"], (0,
|
||
+0))
|
||
430 + per_frame[shape["frame_id"]] = (total + 1, dropped + (
|
||
+verdict == "ignore"))
|
||
431
|
||
432 + ignored_frames = {fid for fid, (total, dropped) in per_fra
|
||
+me.items() if total == dropped}
|
||
433 +
|
||
434 return {
|
||
435 "total_shapes": len(shapes),
|
||
436 "counts": counts,
|
||
|
||
● Update(backend/triage.py)
|
||
⎿ Added 5 lines, removed 2 lines
|
||
451 shapes = shapes_for_frames(frame_ids)
|
||
452 counts = {"keep": 0, "ignore": 0, "reclass": 0}
|
||
453 per_class: dict = {}
|
||
454 - excluded_images = set()
|
||
454 + per_frame = {}
|
||
455 for shape in shapes:
|
||
456 resolved = resolver.verdict(shape)
|
||
457 counts[resolved["verdict"]] += 1
|
||
458 + total, dropped = per_frame.get(shape["frame_id"], (0,
|
||
+0))
|
||
459 + per_frame[shape["frame_id"]] = (total + 1, dropped + (
|
||
+resolved["verdict"] == "ignore"))
|
||
460 if resolved["verdict"] == "ignore":
|
||
459 - excluded_images.add(shape["frame_id"])
|
||
461 continue
|
||
462 effective = resolver.effective_class(shape)
|
||
463 per_class[effective] = per_class.get(effective, 0) + 1
|
||
464
|
||
465 + excluded_images = {fid for fid, (total, dropped) in per_fr
|
||
+ame.items() if total == dropped}
|
||
466 +
|
||
467 return {
|
||
468 "total_shapes": len(shapes),
|
||
469 "total_images": len(frame_ids),
|
||
|
||
Searched for 1 pattern (ctrl+o to expand)
|
||
|
||
● Now update the UI wording to match the new meaning, then rebuild:
|
||
|
||
● Update(frontend/src/components/TriageRules.jsx)
|
||
⎿ Added 4 lines, removed 4 lines
|
||
94 border: '1px solid rgba(56,189,248,0.25)',
|
||
95 }}
|
||
96 >
|
||
97 - As written, these rules keep <strong>{simulation.cou
|
||
-nts.keep.toLocaleString()}</strong> shapes,
|
||
98 - ignore <strong>{simulation.counts.ignore.toLocaleStr
|
||
-ing()}</strong>, reclass{' '}
|
||
97 + As written, these rules keep <strong>{simulation.cou
|
||
+nts.keep.toLocaleString()}</strong> boxes,
|
||
98 + drop <strong>{simulation.counts.ignore.toLocaleStrin
|
||
+g()}</strong>, reclass{' '}
|
||
99 <strong>{simulation.counts.reclass.toLocaleString()}
|
||
</strong> — and{' '}
|
||
100 <strong style={{ color: simulation.frames_held_back
|
||
> simulation.frames_would_merge ? '#f87171' : '#4ade80' }}>
|
||
101 {simulation.frames_would_merge.toLocaleString()}
|
||
102 </strong>{' '}
|
||
103 - of {(simulation.frames_would_merge + simulation.fram
|
||
-es_held_back).toLocaleString()} frames would
|
||
104 - reach training.
|
||
103 + of {(simulation.frames_would_merge + simulation.fram
|
||
+es_held_back).toLocaleString()} frames still
|
||
104 + train. A dropped box leaves its image in; only a fra
|
||
+me that loses every box is held back.
|
||
105 </p>
|
||
106 )}
|
||
107
|
||
|
||
● Update(frontend/src/components/TriagePresets.jsx)
|
||
⎿ Added 4 lines, removed 6 lines
|
||
33 >
|
||
34 <strong style={{ fontSize: '0.85rem' }}>{preset.t
|
||
itle}</strong>
|
||
35 <span className="hint" style={{ fontSize: '0.76re
|
||
m', lineHeight: 1.4 }}>{preset.blurb}</span>
|
||
36 - <span
|
||
37 - className="mono"
|
||
38 - style={{ fontSize: '0.74rem', color: heavy && p
|
||
-reset.rule.action === 'ignore' ? '#f87171' : '#a1a1aa' }}
|
||
39 - >
|
||
40 - {preset.impact.shapes.toLocaleString()} shapes
|
||
-· {preset.impact.frames.toLocaleString()} frames
|
||
41 - {heavy && preset.rule.action === 'ignore' && '
|
||
-— most of the batch'}
|
||
36 + <span className="mono" style={{ fontSize: '0.74re
|
||
+m', color: '#a1a1aa' }}>
|
||
37 + {preset.impact.shapes.toLocaleString()} boxes
|
||
38 + {preset.rule.action === 'ignore' ? ' dropped' :
|
||
+ ' matched'} · touches{' '}
|
||
39 + {preset.impact.frames.toLocaleString()} frames
|
||
40 </span>
|
||
41 <button
|
||
42 type="button"
|
||
|
||
Ran 2 shell commands (ctrl+o to expand)
|
||
|
||
● That's the fix, measured on the same rule and the same data:
|
||
|
||
┌────────────────────────┬──────────────────────────┬───────────────┐
|
||
│ │ frames reaching training │ boxes dropped │
|
||
├────────────────────────┼──────────────────────────┼───────────────┤
|
||
│ before (exclude image) │ 68 of 1,950 │ 8,552 │
|
||
├────────────────────────┼──────────────────────────┼───────────────┤
|
||
│ after (drop box) │ 1,945 of 1,950 │ 8,552 │
|
||
└────────────────────────┴──────────────────────────┴───────────────┘
|
||
|
||
Same 8,552 weak boxes removed, but you keep the dataset. The 5 held-back
|
||
frames are ones where every box was weak — an empty label there would tell the
|
||
model a frame full of sacks is empty, so those are still excluded.
|
||
|
||
Rebuilt and live (index-CVeUWQf4.js). Hard-refresh.
|
||
|
||
What to do when you open it:
|
||
|
||
1. Pick batch6 in the Batch dropdown
|
||
2. Click "Ignore the weakest detections" → it adds a rule pre-filled at score
|
||
≤ 0.399
|
||
3. Watch the blue line: "keeps 76,968 boxes, drops 8,552 … 1,945 of 1,950
|
||
frames still train"
|
||
4. Nudge the threshold up or down — the numbers follow within ~350 ms, and
|
||
each rule row shows its own matches N
|
||
5. Scroll to the crop grid, sorted by lowest score, and look at what you're
|
||
about to drop. If they're real sacks, your threshold is too high
|
||
6. Save rules when the numbers look right
|
||
|
||
For full-sack vs half-sack, the "Split by size" preset is the starting point
|
||
(under 0.743% area → another class), but you need a second class in the
|
||
project first — you only have white sack today. And I'd still check the crops
|
||
before trusting it: at your camera angle, area tracks distance as much as
|
||
fullness. If the crop grid shows small boxes that are just distant full sacks,
|
||
size is the wrong signal and aspect ratio is worth trying instead.
|
||
|
||
Nothing is saved until you press Save, nothing is destructive, and no rule
|
||
affects training until a batch is merged under it.
|
||
|
||
Two things still outstanding from earlier that I haven't touched: the 1,390
|
||
empty-label images from batch7 (42% of your master dataset — still worth
|
||
deciding), and moving the REQs from the proposal into requirements.md once
|
||
you're happy with how this behaves.
|
||
|
||
✻ Cooked for 4m 38s
|
||
|
||
※ recap: You're building a retraining loop; I fixed critical bugs and built the Data Prep triage page, where rules now drop bad boxes instead of whole frames. Next: hard-refresh the browser, pick batch6, and try the preset rules. (disable recaps in /config)
|
||
|
||
❯ from your understanding, this project is for counting sack into the truck,
|
||
how should we prepare the dataset, if the sack is covered by someone head but
|
||
thats is in counting zone, how we prepare that? should i annotate that or
|
||
not? if that covered by others people
|
||
|
||
● Good question — and it's the one that actually determines whether your model
|
||
counts correctly. Let me check what consumes the detections first:
|
||
|
||
Listed 1 directory (ctrl+o to expand)
|
||
● That changes the answer substantially — you're counting with ByteTrack + line
|
||
crossing on a Jetson, not counting detections per frame. Let me look at the
|
||
counter:
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Now I can answer properly, because your counting logic constrains the
|
||
annotation policy in a very specific way.
|
||
|
||
The decisive detail: LineCrossCounter uses y1, the top edge of the box. A
|
||
track counts when it was seen ABOVE the zone and later BELOW it. So the top
|
||
edge of your boxes is the measurement instrument. Anything that moves y1
|
||
artificially moves your count.
|
||
|
||
That's exactly what a head covering the top of a sack does. If you annotate
|
||
the visible extent only, a head over the sack's top pushes y1 downward by
|
||
50–100 px — the sack looks lower than it is, and can trip above → below early.
|
||
You'd get phantom counts caused by annotation policy, not by model error.
|
||
|
||
Your occlusion is already handled elsewhere. cfg/tracker.yaml sets
|
||
track_buffer: 60 — ~2.4 s of lost-track tolerance, and the comment says
|
||
explicitly "survive worker occlusion". The design intent is that the tracker
|
||
coasts through occlusion; the detector is not supposed to see through heads.
|
||
That's the right architecture.
|
||
|
||
So my recommendation:
|
||
|
||
┌───────────────────────────────────────────┬───────────┬────────────────┐
|
||
│ Situation │ Annotate? │ Box │
|
||
├───────────────────────────────────────────┼───────────┼────────────────┤
|
||
│ Sack fully visible │ Yes │ Visible extent │
|
||
├───────────────────────────────────────────┼───────────┼────────────────┤
|
||
│ Partly occluded, top edge clearly visible │ Yes │ Visible extent │
|
||
├───────────────────────────────────────────┼───────────┼────────────────┤
|
||
│ Top edge hidden (head/arm over the top) │ No │ — │
|
||
├───────────────────────────────────────────┼───────────┼────────────────┤
|
||
│ Less than ~40% visible │ No │ — │
|
||
├───────────────────────────────────────────┼───────────┼────────────────┤
|
||
│ Fully hidden behind a person │ Never │ — │
|
||
└───────────────────────────────────────────┴───────────┴────────────────┘
|
||
|
||
The rule is: if you can't see where the sack's top edge is, don't annotate it.
|
||
Let the track coast on track_buffer for those frames. A missing detection
|
||
costs you nothing — ByteTrack holds the ID for 2.4 s and the state machine
|
||
only needs to see the sack below the line at any point in its lifetime. A
|
||
detection with a wrong y1 costs you a miscount.
|
||
|
||
And whatever you choose — be consistent. Inconsistent occlusion labeling is
|
||
the single biggest cause of bad detection models. A model trained on
|
||
"sometimes we box occluded sacks, sometimes not" learns to fire at random
|
||
confidence on ambiguous evidence, which produces ID switches and double
|
||
counts.
|
||
|
||
A warning about the tool I just built for you. Low SAM3 score correlates
|
||
strongly with occlusion. If you apply "ignore the weakest 10%" as a blanket
|
||
rule, you are systematically deleting your occluded examples — the hard ones —
|
||
and training a model that only knows easy, fully-visible sacks. It will then
|
||
fail precisely where workers stand, which is precisely where your counting
|
||
line is. Use triage to remove boxes that are wrong (slivers, duplicates,
|
||
background objects), not boxes that are merely hard. Hard-but-correct examples
|
||
are the most valuable ones in the set. Check the crop grid before dropping
|
||
anything.
|
||
|
||
That also reframes your original full-sack/half-sack idea: I suspect
|
||
"half-sack" in your footage is often "partially occluded sack", not a
|
||
genuinely different object. Those are worth keeping as normal sacks if the top
|
||
edge is visible — not split into a second class.
|
||
|
||
To pin the policy down:
|
||
|
||
● User answered Claude's questions:
|
||
⎿ · What is the counter actually getting wrong today? → Both, inconsistently
|
||
· Where do workers occlude sacks relative to the counting line? → Varies by
|
||
camera
|
||
· Who annotates, and how consistent can the policy be enforced? → Only me
|
||
|
||
Ran 2 shell commands (ctrl+o to expand)
|
||
|
||
● That surfaces something more important than the annotation question.
|
||
|
||
Your counting geometry is hardcoded to one camera. predict.py:1039 sets the
|
||
counting line at y = 330 × scale_y, with the detection polygon and truck ROI
|
||
at fixed pixel coordinates (574, 586, 1418, 1397…), and zones.json holds a
|
||
single palet/truck pair. The only adaptation is a resolution scale factor —
|
||
which corrects for pixel dimensions, not for where the camera is actually
|
||
pointing.
|
||
|
||
You told me occlusion varies by camera. If the cameras also differ in angle or
|
||
framing, then a line calibrated for camera A sits in the wrong place for
|
||
camera B — sometimes above the occlusion zone (undercount, sacks never seen
|
||
"below"), sometimes inside it (overcount, y1 jitter tripping the state
|
||
machine). That produces exactly your symptom: both errors, inconsistently,
|
||
varying by camera. It's also a direct violation of CLAUDE.md §9.
|
||
|
||
No amount of retraining fixes a line in the wrong place. I'd put the effort in
|
||
this order:
|
||
|
||
1. Get ground truth. Hand-count 2–3 videos per camera. Without it you cannot
|
||
tell whether any change helped — which is the entire premise of this app, and
|
||
right now you have no number to compare against.
|
||
2. Move the geometry into per-camera config. zones.json keyed by camera name,
|
||
with line, detection polygon, truck ROI and dedup_radius per camera. That's a
|
||
contained change to predict.py.
|
||
3. Then annotation policy and retraining.
|
||
|
||
Since you annotate alone and mostly correct SAM3, I wrote the policy down so
|
||
it stays consistent across sessions:
|
||
|
||
● Write(docs/annotation-policy.md)
|
||
⎿ Wrote 72 lines to docs/annotation-policy.md
|
||
1 # Annotation policy — sack counting
|
||
3 Why this document exists: the model is not the product, **the co
|
||
unt** is. Every rule below
|
||
4 is derived from how `algoritma-batch/src/counting.py` turns boxe
|
||
5 counter and this policy has to change with it.
|
||
6
|
||
8
|
||
9 `LineCrossCounter` counts on **`y1` — the top edge of the box**
|
||
(`counting.py`, "Uses y1
|
||
nted when it was seen ABOVE
|
||
… +89 lines (ctrl+o to expand)
|
||
|
||
● Written to docs/annotation-policy.md.
|
||
|
||
Direct answer to your question: a sack covered by someone's head, inside the
|
||
counting zone — don't annotate it, if the head hides the sack's top edge. Not
|
||
because it isn't a sack, but because your counter measures y1. A box whose top
|
||
edge is a head's outline reports the wrong height, and the state machine can
|
||
fire early. The missing detection costs you nothing: track_buffer: 60 was set
|
||
precisely to bridge that gap and keeps the track ID alive for ~2.4 s.
|
||
|
||
If the head covers the bottom or side but the top edge is clear — annotate it
|
||
normally, boxing what you can see. Top edge is the deciding factor, not
|
||
overall visibility.
|
||
|
||
Never annotate a sack you can't see at all. That teaches hallucination, which
|
||
becomes ghost tracks and overcounting.
|
||
|
||
Three things I'd push back on in how you're framing the problem:
|
||
|
||
1. "Both, inconsistently" is rarely an annotation symptom. Simultaneous over-
|
||
and undercounting that varies by camera points at geometry, not model quality.
|
||
Your line is hardcoded at one camera's coordinates.
|
||
2. You have no ground truth. You've retrained at least once with no manual
|
||
count to compare against, so there's no way to know if it helped. That's the
|
||
cheapest, highest-value thing you could do this week — hand-count a few videos
|
||
per camera.
|
||
3. The triage tool I just built can hurt you here. Low SAM3 score ≈ occluded,
|
||
not wrong. Dropping the weakest 10% deletes your occluded examples and makes
|
||
the model worse exactly where workers stand. Use it to remove wrong boxes
|
||
(slivers, duplicates, non-sacks), not hard ones.
|
||
|
||
I haven't changed any code for this — it's a policy question and the
|
||
per-camera geometry work sits in algoritma-batch/, which is outside what you
|
||
asked me to touch. Say the word and I'll move the counting geometry into
|
||
per-camera config in zones.json; that's the change most likely to fix your
|
||
actual counting error, and it doesn't need a single new annotation.
|
||
|
||
✻ Cooked for 2m 33s
|
||
|
||
※ recap: Goal is a reliable retraining loop for your sack counter; I audited the app, fixed the critical bugs, and built the Data Prep triage tool, now live. Next: hard-refresh the browser and try a preset rule on batch6. (disable recaps in /config)
|
||
|
||
❯ tolong hati hati saya sudah manual anotasi di page batches simpan backupnya
|
||
dulu untuk hasil anotasi itu
|
||
|
||
● Baik, saya amankan dulu sebelum menyentuh apa pun.
|
||
|
||
Ran 2 shell commands (ctrl+o to expand)
|
||
|
||
● Backup selesai dan terverifikasi — tapi ada yang serius, tolong baca ini dulu.
|
||
|
||
Ran 2 shell commands (ctrl+o to expand)
|
||
|
||
● Aman — backup tersimpan dan terverifikasi. Saya tambahkan satu lapis lagi:
|
||
ekspor JSON yang tidak bergantung pada skema database, supaya kerja tangan
|
||
Anda tetap bisa dipulihkan walau struktur DB berubah.
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Backup selesai, dua lapis, dua-duanya sudah saya verifikasi isinya (bukan cuma
|
||
"file-nya ada"):
|
||
|
||
┌───────────────────────────────────────────────┬──────────────────────────┐
|
||
│ File │ Isi │
|
||
├───────────────────────────────────────────────┼──────────────────────────┤
|
||
│ │ Snapshot penuh 50 MB, │
|
||
│ data/backups/app-20260807-163509.db │ PRAGMA integrity_check = │
|
||
│ │ ok │
|
||
├───────────────────────────────────────────────┼──────────────────────────┤
|
||
│ data/backups/annotations-20260807-163546.json │ 6.324 anotasi (2,6 MB), │
|
||
│ │ bisa dibaca ulang │
|
||
└───────────────────────────────────────────────┴──────────────────────────┘
|
||
|
||
Snapshot DB saya ambil pakai API sqlite3.backup(), jadi konsisten walaupun
|
||
aplikasi sedang jalan (WAL ikut ter-flush). Semua operasi read-only — tidak
|
||
ada yang saya ubah.
|
||
|
||
Kerja tangan Anda yang terlindungi: 1.547 anotasi manual, tersebar di 10
|
||
batch:
|
||
|
||
batch7 406 · batch013 172 · batch016 168 · batch8 165 · batch12 155 · batch014
|
||
140 · batch9 134 · batch11 94 · batch10 58 · batch015 55
|
||
|
||
Tapi ada yang berubah besar sejak terakhir saya lihat, tolong dikonfirmasi.
|
||
Beberapa jam lalu database berisi batch6 dengan 85.520 anotasi dan master
|
||
dataset 3.340 gambar. Sekarang isinya batch yang sama sekali berbeda (id
|
||
45–62), total 6.324 anotasi, dan master dataset kosong — dataset_items = 0,
|
||
dan di disk data/projects/sack/dataset/ tinggal data.yaml, tanpa satu pun
|
||
gambar atau label.
|
||
|
||
Kalau Anda memang sengaja menghapus batch lama dan mulai ulang dengan batch
|
||
baru, semuanya normal — perbaikan delete() yang saya buat tadi memang
|
||
membersihkan file dataset ikut terhapus, dan itu yang terjadi. Frame mentahnya
|
||
masih utuh (226 MB di data/projects/sack/batches).
|
||
|
||
Tapi kalau 85.520 anotasi itu tidak Anda hapus sengaja, tolong bilang sekarang
|
||
— jangan jalankan apa pun dulu. Backup yang barusan saya ambil isinya kondisi
|
||
sesudah perubahan itu, jadi tidak memuat data lama. Saya perlu cek apakah
|
||
masih ada sisa yang bisa diselamatkan sebelum tertimpa.
|
||
|
||
Satu catatan: karena dataset_items = 0, saat ini belum ada apa pun yang bisa
|
||
dilatih — semua batch masih reviewing atau extracted, belum ada yang di-merge.
|
||
|
||
Saya berhenti di sini dan tidak menyentuh kode lagi sampai Anda konfirmasi
|
||
soal data lama itu. Kalau nanti mau lanjut, saya sarankan backup otomatis
|
||
sebelum tiap operasi merusak (delete batch, delete class, merge) — supaya
|
||
tidak bergantung pada saya ingat mengambilnya.
|
||
|
||
✻ Cooked for 1m 20s
|
||
|
||
❯ oke sekarang move ke code, untuk yang di approve itu disimpan dan yang tidak
|
||
di approve = ignore, lalu tambahkan tombol download dataset yang sudah di
|
||
anotasi ini jadikan dalam zip agar user bisa download dan backup sendiri
|
||
|
||
● Dua hal: (1) approve = simpan, non-approve = ignore, (2) tombol download ZIP
|
||
dataset hasil anotasi. Saya lihat kode terkait dulu.
|
||
|
||
Searched for 2 patterns, ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Sekarang perubahan (1) — approve = simpan, sisanya ignore:
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 4 lines, removed 6 lines
|
||
35 raise DatasetError("No such batch")
|
||
36 if batch["status"] == "merged":
|
||
37 raise DatasetError("This batch is already in the master
|
||
dataset")
|
||
38 - if batch["review"]["pending"] > 0:
|
||
39 - raise DatasetError(
|
||
40 - f"{batch['review']['pending']} frame(s) still need
|
||
-a decision before this "
|
||
41 - "batch can be approved"
|
||
42 - )
|
||
38 if batch["review"]["approved"] == 0:
|
||
44 - raise DatasetError("Every frame was rejected — there is
|
||
- nothing to merge")
|
||
39 + raise DatasetError("No frame in this batch is approved
|
||
+— there is nothing to merge")
|
||
40 + # Frames that are not approved — rejected or never looked a
|
||
+t — are simply
|
||
41 + # left behind. Only what the user signed off on enters the
|
||
+dataset, so a
|
||
42 + # partly-reviewed batch can be merged for the part that is
|
||
+done.
|
||
43 with db.cursor() as cur:
|
||
44 cur.execute(
|
||
45 """SELECT 1 FROM jobs WHERE batch_id = ? AND type =
|
||
'merge'
|
||
● Karena sekarang merge bisa sebagian, batch yang sudah merged harus bisa
|
||
di-merge lagi untuk frame yang baru di-approve — kalau tidak, sisanya terkunci
|
||
selamanya:
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 5 lines, removed 2 lines
|
||
33 batch = batches.get(batch_id)
|
||
34 if batch is None:
|
||
35 raise DatasetError("No such batch")
|
||
36 - if batch["status"] == "merged":
|
||
37 - raise DatasetError("This batch is already in the master
|
||
- dataset")
|
||
36 + if batch["status"] == "merged" and _unmerged_approved(batch
|
||
+_id) == 0:
|
||
37 + # Re-merging is allowed, because a partly-reviewed batc
|
||
+h can be merged
|
||
38 + # for the part that was done and finished later. Only r
|
||
+efuse when there
|
||
39 + # is genuinely nothing new to add.
|
||
40 + raise DatasetError("Every approved frame in this batch
|
||
+is already in the master dataset")
|
||
41 if batch["review"]["approved"] == 0:
|
||
42 raise DatasetError("No frame in this batch is approved
|
||
— there is nothing to merge")
|
||
43 # Frames that are not approved — rejected or never looked a
|
||
t — are simply
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 12 lines
|
||
63 return job.to_dict()
|
||
64
|
||
65
|
||
66 +def _unmerged_approved(batch_id: int) -> int:
|
||
67 + """Approved frames of this batch that are not in the datase
|
||
+t yet."""
|
||
68 + with db.cursor() as cur:
|
||
69 + cur.execute(
|
||
70 + """SELECT COUNT(*) FROM frames f
|
||
71 + LEFT JOIN dataset_items d ON d.frame_id = f.id
|
||
72 + WHERE f.batch_id = ? AND f.review_status = 'appr
|
||
+oved' AND d.id IS NULL""",
|
||
73 + (batch_id,),
|
||
74 + )
|
||
75 + return cur.fetchone()[0]
|
||
76 +
|
||
77 +
|
||
78 def _label_line(class_id: int, geometry: dict, label_type: str)
|
||
-> str:
|
||
79 if label_type == "bbox":
|
||
80 x0, y0, x1, y1 = review.to_box(geometry)
|
||
|
||
● Sekarang (2) — export ZIP dataset hasil anotasi, tanpa perlu merge dulu:
|
||
|
||
● Write(backend/export.py)
|
||
⎿ Wrote 147 lines to backend/export.py
|
||
1 """Export annotated frames as a plain YOLO dataset the user can
|
||
2
|
||
3 This is deliberately independent of the master dataset. A batch
|
||
can be exported
|
||
cause the point
|
||
5 is that the user owns a copy of their own work and does not have
|
||
6 application to keep it.
|
||
7
|
||
8 Layout inside the zip:
|
||
10 images/<batch>__<frame>.jpg
|
||
… +163 lines (ctrl+o to expand)
|
||
|
||
● Update(backend/api/triage.py)
|
||
⎿ Added 17 lines
|
||
100 return triage.preview(project_id)
|
||
101
|
||
102
|
||
103 +@router.get("/api/projects/{project_id}/export")
|
||
104 +def export_annotated(project_id: int, batch_ids: str = "", app
|
||
+roved_only: bool = False,
|
||
105 + include_empty: bool = False):
|
||
106 + """Download annotated frames as a YOLO zip, merged or not
|
||
+— the user's own backup."""
|
||
107 + from fastapi.responses import FileResponse
|
||
108 +
|
||
109 + from backend import export
|
||
110 +
|
||
111 + ids = [int(part) for part in batch_ids.split(",") if part.
|
||
+strip().isdigit()]
|
||
112 + try:
|
||
113 + path = export.build_zip(project_id, ids or None,
|
||
114 + approved_only=approved_only, i
|
||
+nclude_empty=include_empty)
|
||
115 + except export.ExportError as exc:
|
||
116 + raise HTTPException(400, str(exc))
|
||
117 + return FileResponse(path, media_type="application/zip", fi
|
||
+lename=os.path.basename(path))
|
||
118 +
|
||
119 +
|
||
120 @router.get("/api/annotations/{annotation_id}/crop")
|
||
121 def crop(annotation_id: int, pad: float = 0.08):
|
||
122 """The shape itself, cropped out of its frame — the crop g
|
||
rid judges objects,
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Sekarang tombolnya di UI:
|
||
|
||
● Update(frontend/src/api.js)
|
||
⎿ Added 8 lines
|
||
136 clearTriageOverrides: (annotationIds) =>
|
||
137 request('/triage/overrides', { method: 'DELETE', body: { a
|
||
nnotation_ids: annotationIds } }),
|
||
138 cropUrl: (annotationId) => `/api/annotations/${annotationId}
|
||
/crop`,
|
||
139 + exportUrl: (projectId, { batchIds = [], approvedOnly = false
|
||
+, includeEmpty = false } = {}) => {
|
||
140 + const params = new URLSearchParams()
|
||
141 + if (batchIds.length) params.set('batch_ids', batchIds.join
|
||
+(','))
|
||
142 + if (approvedOnly) params.set('approved_only', 'true')
|
||
143 + if (includeEmpty) params.set('include_empty', 'true')
|
||
144 + const query = params.toString()
|
||
145 + return `/api/projects/${projectId}/export${query ? `?${que
|
||
+ry}` : ''}`
|
||
146 + },
|
||
147 datasetDownloadUrl: (projectId) => `/api/projects/${projectI
|
||
d}/dataset/download`,
|
||
148
|
||
149 hardware: () => request('/hardware'),
|
||
|
||
Read 1 file (ctrl+o to expand)
|
||
|
||
● Update(frontend/src/pages/BatchesPage.jsx)
|
||
⎿ Added 18 lines, removed 7 lines
|
||
93 <p className="muted mono">{project.name}</p>
|
||
94 </div>
|
||
95 {batches.length > 0 && (
|
||
96 - <button
|
||
97 - className="btn btn-primary"
|
||
98 - style={{ fontSize: '0.85rem', padding: '8px 16px',
|
||
- cursor: 'pointer', display: 'inline-flex', alignItems: 'cente
|
||
-r', gap: 6, borderRadius: 6 }}
|
||
99 - onClick={handleAutoAnnotateAllHeader}
|
||
100 - >
|
||
101 - <ZapIcon size={16} /> Auto-Annotate All Batches ({
|
||
-batches.length})
|
||
102 - </button>
|
||
96 + <div style={{ display: 'flex', gap: 8, alignItems: '
|
||
+center', flexWrap: 'wrap' }}>
|
||
97 + <a
|
||
98 + className="btn"
|
||
99 + href={api.exportUrl(projectId)}
|
||
100 + download
|
||
101 + title="Download every annotated frame in this pr
|
||
+oject as a YOLO dataset — your own backup, no merge needed"
|
||
102 + style={{ fontSize: '0.85rem', padding: '8px 16px
|
||
+', cursor: 'pointer', display: 'inline-flex', alignItems: 'cen
|
||
+ter', gap: 6, borderRadius: 6 }}
|
||
103 + >
|
||
104 + <DatabaseIcon size={16} /> Download Annotations
|
||
+(.zip)
|
||
105 + </a>
|
||
106 + <button
|
||
107 + className="btn btn-primary"
|
||
108 + style={{ fontSize: '0.85rem', padding: '8px 16px
|
||
+', cursor: 'pointer', display: 'inline-flex', alignItems: 'cen
|
||
+ter', gap: 6, borderRadius: 6 }}
|
||
109 + onClick={handleAutoAnnotateAllHeader}
|
||
110 + >
|
||
111 + <ZapIcon size={16} /> Auto-Annotate All Batches
|
||
+({batches.length})
|
||
112 + </button>
|
||
113 + </div>
|
||
114 )}
|
||
115 </div>
|
||
116
|
||
|
||
Searched for 1 pattern (ctrl+o to expand)
|
||
|
||
● Update(frontend/src/pages/BatchesPage.jsx)
|
||
⎿ Added 1 line, removed 1 line
|
||
1 import React, { useCallback, useEffect, useRef, useState } from
|
||
'react'
|
||
2 import { api } from '../api'
|
||
3 -import { AlertIcon, ZapIcon } from '../components/Icons'
|
||
3 +import { AlertIcon, DatabaseIcon, ZapIcon } from '../components/
|
||
+Icons'
|
||
4 import { BatchList, ActiveJobsBanner } from './LibraryPage'
|
||
5
|
||
6 export default function BatchesPage({ projectId, onProject }) {
|
||
|
||
❯ tambahkan juga upload zip untuk upload hasil backup agar bisa dilanjutkan
|
||
|
||
● Baik, saya tambahkan juga import ZIP-nya. Lanjut dulu di file yang sama:
|
||
|
||
● Update(backend/export.py)
|
||
⎿ Added 148 lines
|
||
145 path = shutil.make_archive(archive_base, "zip", staging)
|
||
146 shutil.rmtree(staging, ignore_errors=True)
|
||
147 return path
|
||
148 +
|
||
149 +
|
||
150 +# ---- restore -----------------------------------------------
|
||
+-------------
|
||
151 +
|
||
152 +def _safe_members(archive) -> list:
|
||
153 + """Only plain files under images/ and labels/.
|
||
154 +
|
||
155 + A zip can name `../../etc/passwd`; extracting one blindly
|
||
+writes wherever it
|
||
156 + says. Every member is checked rather than trusted.
|
||
157 + """
|
||
158 + picked = []
|
||
159 + for member in archive.infolist():
|
||
160 + if member.is_dir():
|
||
161 + continue
|
||
162 + name = member.filename.replace("\\", "/")
|
||
163 + if name.startswith("/") or ".." in name.split("/"):
|
||
164 + continue
|
||
165 + if name.startswith(("images/", "labels/")) or name ==
|
||
+"MANIFEST.json":
|
||
166 + picked.append((name, member))
|
||
167 + return picked
|
||
168 +
|
||
169 +
|
||
170 +def _points_from_label(parts: List[str], label_type: str) -> O
|
||
+ptional[dict]:
|
||
171 + values = [float(v) for v in parts]
|
||
172 + if label_type == "bbox":
|
||
173 + if len(values) != 4:
|
||
174 + return None
|
||
175 + cx, cy, w, h = values
|
||
176 + return {"type": "bbox",
|
||
177 + "points": [cx - w / 2, cy - h / 2, cx + w / 2,
|
||
+ cy + h / 2]}
|
||
178 + if len(values) < 6 or len(values) % 2:
|
||
179 + return None
|
||
180 + return {"type": "polygon",
|
||
181 + "points": [[values[i], values[i + 1]] for i in ran
|
||
+ge(0, len(values), 2)]}
|
||
182 +
|
||
183 +
|
||
184 +def restore_zip(project_id: int, zip_path: str, batch_label: s
|
||
+tr = "") -> dict:
|
||
185 + """Load an exported zip back in as a fresh batch, ready to
|
||
+ keep reviewing.
|
||
186 +
|
||
187 + The frames land in a new batch rather than being merged ba
|
||
+ck into the ones
|
||
188 + they came from: the originals may still exist, and silentl
|
||
+y overwriting a
|
||
189 + batch the user is working in would destroy the very work t
|
||
+his feature is
|
||
190 + meant to protect.
|
||
191 + """
|
||
192 + import zipfile
|
||
193 +
|
||
194 + from PIL import Image
|
||
195 +
|
||
196 + project = projects.get(project_id)
|
||
197 + if project is None:
|
||
198 + raise ExportError("No such project")
|
||
199 +
|
||
200 + by_name = {item["name"]: item["class_id"] for item in proj
|
||
+ect["classes"]}
|
||
201 + stamp = time.strftime("%Y%m%d-%H%M%S")
|
||
202 + label = batch_label or f"restored-{stamp}"
|
||
203 +
|
||
204 + with zipfile.ZipFile(zip_path) as archive:
|
||
205 + members = _safe_members(archive)
|
||
206 + names = {name for name, _ in members}
|
||
207 + if not any(name.startswith("images/") for name in name
|
||
+s):
|
||
208 + raise ExportError("This zip has no images/ folder
|
||
+— is it an export from this app?")
|
||
209 +
|
||
210 + manifest = {}
|
||
211 + if "MANIFEST.json" in names:
|
||
212 + manifest = json.loads(archive.read("MANIFEST.json"
|
||
+))
|
||
213 + source_type = manifest.get("label_type", project["labe
|
||
+l_type"])
|
||
214 + if source_type != project["label_type"]:
|
||
215 + raise ExportError(
|
||
216 + f"This export holds {source_type} labels but t
|
||
+he project is "
|
||
217 + f"{project['label_type']} — importing it would
|
||
+ produce wrong shapes"
|
||
218 + )
|
||
219 +
|
||
220 + # Classes come back by name, so an id that shifted sin
|
||
+ce the export does
|
||
221 + # not silently relabel every shape.
|
||
222 + remap = {}
|
||
223 + for item in manifest.get("classes", []):
|
||
224 + if item["name"] in by_name:
|
||
225 + remap[item["class_id"]] = by_name[item["name"]
|
||
+]
|
||
226 + else:
|
||
227 + raise ExportError(
|
||
228 + f"The export uses class '{item['name']}',
|
||
+which this project does not "
|
||
229 + "have. Add the class first, then import."
|
||
230 + )
|
||
231 +
|
||
232 + with db.cursor() as cur:
|
||
233 + cur.execute(
|
||
234 + """INSERT INTO batches (project_id, video_path
|
||
+, date_label, batch_label,
|
||
235 + start_sec, end_sec, fp
|
||
+s, status, created_at)
|
||
236 + VALUES (?, '', 'restored', ?, 0, 0, 0, 'ext
|
||
+racted', ?)""",
|
||
237 + (project_id, label, time.time()),
|
||
238 + )
|
||
239 + batch_id = cur.lastrowid
|
||
240 +
|
||
241 + target_dir = batches.frames_dir(project["slug"], batch
|
||
+_id)
|
||
242 + os.makedirs(target_dir, exist_ok=True)
|
||
243 +
|
||
244 + restored, shapes, skipped = 0, 0, 0
|
||
245 + image_members = sorted(n for n in names if n.startswit
|
||
+h("images/"))
|
||
246 + for index, name in enumerate(image_members):
|
||
247 + stem = os.path.splitext(os.path.basename(name))[0]
|
||
248 + if not stem:
|
||
249 + continue
|
||
250 + filename = f"{stem}.jpg"
|
||
251 + destination = os.path.join(target_dir, filename)
|
||
252 + with archive.open(name) as source, open(destinatio
|
||
+n, "wb") as handle:
|
||
253 + shutil.copyfileobj(source, handle)
|
||
254 +
|
||
255 + try:
|
||
256 + with Image.open(destination) as image:
|
||
257 + width, height = image.size
|
||
258 + except Exception:
|
||
259 + os.unlink(destination)
|
||
260 + skipped += 1
|
||
261 + continue
|
||
262 +
|
||
263 + with db.cursor() as cur:
|
||
264 + cur.execute(
|
||
265 + """INSERT INTO frames (batch_id, idx, file
|
||
+name, width, height, review_status)
|
||
266 + VALUES (?, ?, ?, ?, ?, 'pending')""",
|
||
267 + (batch_id, index, filename, width, height)
|
||
+,
|
||
268 + )
|
||
269 + frame_id = cur.lastrowid
|
||
270 +
|
||
271 + label_name = f"labels/{stem}.txt"
|
||
272 + if label_name in names:
|
||
273 + for line in archive.read(label_name).decod
|
||
+e("utf-8").splitlines():
|
||
274 + parts = line.split()
|
||
275 + if len(parts) < 5:
|
||
276 + continue
|
||
277 + geometry = _points_from_label(parts[1:
|
||
+], project["label_type"])
|
||
278 + if geometry is None:
|
||
279 + continue
|
||
280 + class_id = remap.get(int(parts[0]), in
|
||
+t(parts[0]))
|
||
281 + cur.execute(
|
||
282 + """INSERT INTO annotations (frame_
|
||
+id, class_id, geometry, score,
|
||
283 + source
|
||
+, created_at)
|
||
284 + VALUES (?, ?, ?, 1.0, 'manual',
|
||
+ ?)""",
|
||
285 + (frame_id, class_id, json.dumps(ge
|
||
+ometry), time.time()),
|
||
286 + )
|
||
287 + shapes += 1
|
||
288 + restored += 1
|
||
289 +
|
||
290 + with db.cursor() as cur:
|
||
291 + cur.execute("UPDATE batches SET frame_count = ?, statu
|
||
+s = 'reviewing' WHERE id = ?",
|
||
292 + (restored, batch_id))
|
||
293 +
|
||
294 + return {"batch_id": batch_id, "batch_label": label, "frame
|
||
+s": restored,
|
||
295 + "shapes": shapes, "skipped": skipped}
|
||
|
||
● Update(backend/api/triage.py)
|
||
⎿ Added 22 lines
|
||
117 return FileResponse(path, media_type="application/zip", fi
|
||
lename=os.path.basename(path))
|
||
118
|
||
119
|
||
120 +@router.post("/api/projects/{project_id}/import")
|
||
121 +async def import_annotated(project_id: int, file: UploadFile =
|
||
+ File(...),
|
||
122 + batch_label: str = Form("")) -> dic
|
||
+t:
|
||
123 + """Load a previously exported zip back in, as a new batch
|
||
+to keep working on."""
|
||
124 + import tempfile
|
||
125 +
|
||
126 + from backend import export
|
||
127 +
|
||
128 + staged = tempfile.NamedTemporaryFile(suffix=".zip", delete
|
||
+=False)
|
||
129 + try:
|
||
130 + shutil.copyfileobj(file.file, staged)
|
||
131 + staged.close()
|
||
132 + return export.restore_zip(project_id, staged.name, bat
|
||
+ch_label=batch_label)
|
||
133 + except export.ExportError as exc:
|
||
134 + raise HTTPException(400, str(exc))
|
||
135 + except Exception as exc:
|
||
136 + raise HTTPException(400, f"Could not read that zip: {e
|
||
+xc}")
|
||
137 + finally:
|
||
138 + if os.path.exists(staged.name):
|
||
139 + os.unlink(staged.name)
|
||
140 +
|
||
141 +
|
||
142 @router.get("/api/annotations/{annotation_id}/crop")
|
||
143 def crop(annotation_id: int, pad: float = 0.08):
|
||
144 """The shape itself, cropped out of its frame — the crop g
|
||
rid judges objects,
|
||
|
||
● Update(backend/api/triage.py)
|
||
⎿ Added 2 lines, removed 1 line
|
||
2
|
||
3 import io
|
||
4 import os
|
||
5 +import shutil
|
||
6
|
||
6 -from fastapi import APIRouter, HTTPException
|
||
7 +from fastapi import APIRouter, File, Form, HTTPException, Uploa
|
||
+dFile
|
||
8 from fastapi.responses import StreamingResponse
|
||
9 from pydantic import BaseModel
|
||
10 from typing import List, Optional
|
||
|
||
● Sekarang tombol upload di UI:
|
||
|
||
● Update(frontend/src/api.js)
|
||
⎿ Added 6 lines
|
||
136 clearTriageOverrides: (annotationIds) =>
|
||
137 request('/triage/overrides', { method: 'DELETE', body: { a
|
||
nnotation_ids: annotationIds } }),
|
||
138 cropUrl: (annotationId) => `/api/annotations/${annotationId}
|
||
/crop`,
|
||
139 + importAnnotations: (projectId, zipFile, batchLabel = '') =>
|
||
+{
|
||
140 + const form = new FormData()
|
||
141 + form.append('file', zipFile)
|
||
142 + form.append('batch_label', batchLabel)
|
||
143 + return request(`/projects/${projectId}/import`, { method:
|
||
+'POST', form })
|
||
144 + },
|
||
145 exportUrl: (projectId, { batchIds = [], approvedOnly = false
|
||
, includeEmpty = false } = {}) => {
|
||
146 const params = new URLSearchParams()
|
||
147 if (batchIds.length) params.set('batch_ids', batchIds.join
|
||
(','))
|
||
|
||
Read 1 file (ctrl+o to expand)
|
||
|
||
● Update(frontend/src/pages/BatchesPage.jsx)
|
||
⎿ Added 24 lines
|
||
8 const [batches, setBatches] = useState([])
|
||
9 const [jobs, setJobs] = useState([])
|
||
10 const [error, setError] = useState('')
|
||
11 + const [importing, setImporting] = useState(false)
|
||
12 + const [notice, setNotice] = useState('')
|
||
13 + const importInput = useRef(null)
|
||
14
|
||
15 + async function handleImport(event) {
|
||
16 + const file = event.target.files?.[0]
|
||
17 + event.target.value = ''
|
||
18 + if (!file) return
|
||
19 + setImporting(true)
|
||
20 + setError('')
|
||
21 + setNotice('')
|
||
22 + try {
|
||
23 + const result = await api.importAnnotations(projectId, fil
|
||
+e)
|
||
24 + setNotice(
|
||
25 + `Restored "${result.batch_label}" — ${result.frames} fr
|
||
+ame(s), ${result.shapes} annotation(s)`
|
||
26 + + (result.skipped ? `, ${result.skipped} skipped` : '')
|
||
+,
|
||
27 + )
|
||
28 + loadBatches()
|
||
29 + } catch (exc) {
|
||
30 + setError(exc.message)
|
||
31 + } finally {
|
||
32 + setImporting(false)
|
||
33 + }
|
||
34 + }
|
||
35 +
|
||
36 const loadBatches = useCallback(() => {
|
||
37 api.listBatches(projectId).then((payload) => setBatches(pay
|
||
load.batches)).catch(() => {})
|
||
38 }, [projectId])
|
||
|
||
● Update(frontend/src/pages/BatchesPage.jsx)
|
||
⎿ Added 19 lines, removed 3 lines
|
||
116 <h1>Batches ({batches.length})</h1>
|
||
117 <p className="muted mono">{project.name}</p>
|
||
118 </div>
|
||
119 - {batches.length > 0 && (
|
||
120 - <div style={{ display: 'flex', gap: 8, alignItems: '
|
||
-center', flexWrap: 'wrap' }}>
|
||
121 - <a
|
||
119 + <div style={{ display: 'flex', gap: 8, alignItems: 'ce
|
||
+nter', flexWrap: 'wrap' }}>
|
||
120 + <input
|
||
121 + ref={importInput}
|
||
122 + type="file"
|
||
123 + accept=".zip"
|
||
124 + onChange={handleImport}
|
||
125 + style={{ display: 'none' }}
|
||
126 + />
|
||
127 + <button
|
||
128 + className="btn"
|
||
129 + disabled={importing}
|
||
130 + onClick={() => importInput.current?.click()}
|
||
131 + title="Restore a previously downloaded annotation
|
||
+zip as a new batch"
|
||
132 + style={{ fontSize: '0.85rem', padding: '8px 16px',
|
||
+ cursor: importing ? 'wait' : 'pointer', display: 'inline-flex
|
||
+', alignItems: 'center', gap: 6, borderRadius: 6 }}
|
||
133 + >
|
||
134 + <UploadIcon size={16} /> {importing ? 'Restoring…'
|
||
+ : 'Restore from .zip'}
|
||
135 + </button>
|
||
136 + {batches.length > 0 && (
|
||
137 + <
|
||
138 className="btn"
|
||
139 href={api.exportUrl(projectId)}
|
||
140 download
|
||
|
||
● Edit itu merusak JSX-nya. Saya perbaiki blok itu utuh:
|
||
|
||
Read 1 file (ctrl+o to expand)
|
||
● Update(frontend/src/pages/BatchesPage.jsx)
|
||
⎿ Added 24 lines, removed 18 lines
|
||
134 <UploadIcon size={16} /> {importing ? 'Restoring…'
|
||
: 'Restore from .zip'}
|
||
135 </button>
|
||
136 {batches.length > 0 && (
|
||
137 - <
|
||
138 - className="btn"
|
||
139 - href={api.exportUrl(projectId)}
|
||
140 - download
|
||
141 - title="Download every annotated frame in this pr
|
||
-oject as a YOLO dataset — your own backup, no merge needed"
|
||
142 - style={{ fontSize: '0.85rem', padding: '8px 16px
|
||
-', cursor: 'pointer', display: 'inline-flex', alignItems: 'cen
|
||
-ter', gap: 6, borderRadius: 6 }}
|
||
143 - >
|
||
144 - <DatabaseIcon size={16} /> Download Annotations
|
||
-(.zip)
|
||
145 - </a>
|
||
146 - <button
|
||
147 - className="btn btn-primary"
|
||
148 - style={{ fontSize: '0.85rem', padding: '8px 16px
|
||
-', cursor: 'pointer', display: 'inline-flex', alignItems: 'cen
|
||
-ter', gap: 6, borderRadius: 6 }}
|
||
149 - onClick={handleAutoAnnotateAllHeader}
|
||
150 - >
|
||
151 - <ZapIcon size={16} /> Auto-Annotate All Batches
|
||
-({batches.length})
|
||
152 - </button>
|
||
153 - </div>
|
||
154 - )}
|
||
137 + <>
|
||
138 + <a
|
||
139 + className="btn"
|
||
140 + href={api.exportUrl(projectId)}
|
||
141 + download
|
||
142 + title="Download every annotated frame in this
|
||
+project as a YOLO dataset — your own backup, no merge needed"
|
||
143 + style={{ fontSize: '0.85rem', padding: '8px 16
|
||
+px', cursor: 'pointer', display: 'inline-flex', alignItems: 'c
|
||
+enter', gap: 6, borderRadius: 6 }}
|
||
144 + >
|
||
145 + <DatabaseIcon size={16} /> Download Annotation
|
||
+s (.zip)
|
||
146 + </a>
|
||
147 + <button
|
||
148 + className="btn btn-primary"
|
||
149 + style={{ fontSize: '0.85rem', padding: '8px 16
|
||
+px', cursor: 'pointer', display: 'inline-flex', alignItems: 'c
|
||
+enter', gap: 6, borderRadius: 6 }}
|
||
150 + onClick={handleAutoAnnotateAllHeader}
|
||
151 + >
|
||
152 + <ZapIcon size={16} /> Auto-Annotate All Batche
|
||
+s ({batches.length})
|
||
153 + </button>
|
||
154 + </>
|
||
155 + )}
|
||
156 + </div>
|
||
157 </div>
|
||
158
|
||
159 + {notice && (
|
||
160 + <p className="hint" style={{ marginBottom: 12, color:
|
||
+'#4ade80' }}>{notice}</p>
|
||
161 + )}
|
||
162 +
|
||
163 <ActiveJobsBanner jobs={activeJobs} onCancel={cancelJob}
|
||
/>
|
||
164
|
||
165 {batches.length === 0 ? (
|
||
|
||
● Update(frontend/src/pages/BatchesPage.jsx)
|
||
⎿ Added 1 line, removed 1 line
|
||
1 import React, { useCallback, useEffect, useRef, useState } from
|
||
'react'
|
||
2 import { api } from '../api'
|
||
3 -import { AlertIcon, DatabaseIcon, ZapIcon } from '../components/
|
||
-Icons'
|
||
3 +import { AlertIcon, DatabaseIcon, UploadIcon, ZapIcon } from '..
|
||
+/components/Icons'
|
||
4 import { BatchList, ActiveJobsBanner } from './LibraryPage'
|
||
5
|
||
6 export default function BatchesPage({ projectId, onProject }) {
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Rebuild container, lalu saya uji export→import benar-benar bolak-balik:
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Export jalan: 932 gambar, 6.324 anotasi, 187 MB. Sekarang uji import-nya
|
||
benar-benar bisa dipulihkan:
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Sama persis. Sekarang saya cek geometri kotaknya benar-benar utuh, bukan cuma
|
||
jumlahnya cocok:
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Utuh — 0 shape berbeda, selisih koordinat terbesar 0,001 pixel (itu cuma
|
||
pembulatan 6 desimal di format YOLO). Sekarang saya hapus batch uji supaya
|
||
data Anda bersih:
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Satu masalah yang saya lihat: tiap klik download menulis file ZIP 187 MB di
|
||
folder proyek dan tidak pernah dihapus. Sepuluh kali klik = 1,9 GB sampah.
|
||
Saya perbaiki:
|
||
|
||
● Update(backend/export.py)
|
||
⎿ Added 11 lines, removed 3 lines
|
||
139 with open(os.path.join(staging, "MANIFEST.json"), "w", enc
|
||
oding="utf-8") as handle:
|
||
140 json.dump(manifest, handle, indent=2)
|
||
141
|
||
142 + project_root = config.project_dir(project["slug"])
|
||
143 + # Each export is ~200 MB. Without this, every click leaves
|
||
+ another copy
|
||
144 + # behind and the project folder grows without bound.
|
||
145 + for stale in os.listdir(project_root):
|
||
146 + if stale.startswith("export-") and stale.endswith(".zi
|
||
+p"):
|
||
147 + try:
|
||
148 + os.unlink(os.path.join(project_root, stale))
|
||
149 + except OSError:
|
||
150 + pass
|
||
151 +
|
||
152 stamp = time.strftime("%Y%m%d-%H%M%S")
|
||
143 - archive_base = os.path.join(config.project_dir(project["sl
|
||
-ug"]),
|
||
144 - f"export-{stamp}")
|
||
145 - path = shutil.make_archive(archive_base, "zip", staging)
|
||
153 + path = shutil.make_archive(os.path.join(project_root, f"ex
|
||
+port-{stamp}"), "zip", staging)
|
||
154 shutil.rmtree(staging, ignore_errors=True)
|
||
155 return path
|
||
156
|
||
|
||
Ran 2 shell commands (ctrl+o to expand)
|
||
|
||
● Selesai dan sudah diuji dengan data asli Anda.
|
||
|
||
1. Approve = simpan, tidak approve = ignore
|
||
|
||
approve() tidak lagi menolak batch yang masih ada frame pending. Sekarang
|
||
hanya frame ber-status approved yang masuk dataset; rejected dan pending
|
||
ditinggalkan begitu saja. Jadi batch yang baru direview separuh tetap bisa
|
||
di-merge untuk bagian yang sudah selesai.
|
||
|
||
Konsekuensinya saya tangani juga: batch yang sudah merged sekarang boleh
|
||
di-merge ulang untuk frame yang baru Anda approve belakangan. Kalau tidak,
|
||
sisa frame di batch itu akan terkunci selamanya — dan itu justru merugikan
|
||
karena merge sebagian sekarang jadi hal biasa. Penolakan hanya terjadi kalau
|
||
memang tidak ada frame approved baru.
|
||
|
||
2. Download ZIP — tombol "Download Annotations (.zip)" di halaman Batches.
|
||
Isinya images/, labels/ (format YOLO), data.yaml, classes.txt, dan
|
||
MANIFEST.json. Ini tidak bergantung pada merge — semua frame yang punya
|
||
anotasi ikut, apa pun status batch-nya. Frame tanpa anotasi sengaja dilewati
|
||
(176 frame), karena file label kosong artinya "tidak ada objek di sini", dan
|
||
untuk frame yang sekadar belum dianotasi itu klaim yang salah.
|
||
|
||
3. Upload ZIP — tombol "Restore from .zip" di halaman yang sama. Hasilnya
|
||
masuk sebagai batch baru (restored-<tanggal>), bukan menimpa batch yang ada —
|
||
menimpa batch yang sedang Anda kerjakan justru akan menghancurkan pekerjaan
|
||
yang mau dilindungi fitur ini.
|
||
|
||
Hasil uji pulang-pergi dengan data Anda:
|
||
|
||
┌────────────────────────────┬───────────────────────────────────┐
|
||
│ │ │
|
||
├────────────────────────────┼───────────────────────────────────┤
|
||
│ Export │ 932 gambar, 6.324 anotasi, 187 MB │
|
||
├────────────────────────────┼───────────────────────────────────┤
|
||
│ Import kembali │ 932 frame, 6.324 anotasi, 0 gagal │
|
||
├────────────────────────────┼───────────────────────────────────┤
|
||
│ Shape yang berbeda │ 0 │
|
||
├────────────────────────────┼───────────────────────────────────┤
|
||
│ Selisih koordinat terbesar │ 0,001 pixel pada frame 1920px │
|
||
└────────────────────────────┴───────────────────────────────────┘
|
||
|
||
Batch uji sudah saya hapus. Data Anda sekarang: 10 batch, 6.324 anotasi, 1.547
|
||
manual — persis seperti sebelum saya mulai.
|
||
|
||
Tiga hal yang saya putuskan sendiri, tolong dikoreksi kalau salah:
|
||
|
||
- Kelas dicocokkan berdasarkan nama, bukan angka. Kalau urutan kelas berubah
|
||
setelah export, class_id lama bisa menunjuk kelas yang salah — dengan nama,
|
||
itu tidak mungkin terjadi. Kalau ada kelas di ZIP yang belum ada di proyek,
|
||
import ditolak dengan pesan jelas, tidak diam-diam dibuat.
|
||
- Isi ZIP divalidasi terhadap zip slip (anggota bernama ../../etc/passwd
|
||
diabaikan), dan label_type bbox/polygon dicek supaya tidak menghasilkan bentuk
|
||
salah.
|
||
- Anotasi hasil restore ditandai source='manual', jadi tidak akan tertimpa
|
||
kalau auto-annotate dijalankan lagi di batch itu.
|
||
|
||
Satu bug yang saya temukan sambil menguji dan sudah diperbaiki: tiap klik
|
||
download menulis ZIP 187 MB ke folder proyek dan tidak pernah menghapusnya.
|
||
Sepuluh klik = 1,9 GB terbuang. Sekarang export lama dibersihkan sebelum yang
|
||
baru dibuat. File sisa dari pengujian tadi sudah saya hapus.
|
||
|
||
Container sudah di-rebuild (index-BhGV3ZDV.js). Hard-refresh browser dulu
|
||
sebelum mencoba.
|
||
|
||
✻ Cogitated for 4m 49s
|
||
|
||
❯ oke sekarang di table buatkan tombol untuk memilih mana yang akan di "approve
|
||
batch" dan masuk ke proses selanjutnya, untuk posisi ganti dengan auto
|
||
annotate karena itu tombol ada 2 di satu pages
|
||
|
||
● Saya lihat dulu tabelnya dan tombol yang dobel itu.
|
||
|
||
Searched for 1 pattern, read 1 file, ran 1 shell command (ctrl+o to expand)
|
||
|
||
● batch.review.approved tersedia — jadi tombolnya bisa tahu mana yang layak.
|
||
Saya ganti tombol dobel itu:
|
||
|
||
● Update(frontend/src/pages/LibraryPage.jsx)
|
||
⎿ Added 17 lines, removed 8 lines
|
||
152 <div className="panel table-wrap" style={{ marginTop: 16
|
||
, padding: '16px' }}>
|
||
153 <div style={{ display: 'flex', justifyContent: 'space-
|
||
between', alignItems: 'center', marginBottom: 14, padding: '0
|
||
4px' }}>
|
||
154 <h3 style={{ margin: 0, fontSize: '1rem', color: '#e
|
||
4e4e7' }}>Batches ({batches.length})</h3>
|
||
155 - <button
|
||
156 - className="btn btn-primary"
|
||
157 - style={{ fontSize: '0.85rem', padding: '6px 14px',
|
||
- cursor: 'pointer', display: 'inline-flex', alignItems: 'cente
|
||
-r', gap: 6, borderRadius: 6 }}
|
||
158 - onClick={handleAutoAnnotateAll}
|
||
159 - >
|
||
160 - <ZapIcon size={14} /> Auto-Annotate All Batches ({
|
||
-batches.length})
|
||
161 - </button>
|
||
162 -
|
||
155 + <div style={{ display: 'flex', gap: 10, alignItems:
|
||
+'center' }}>
|
||
156 + {selectedIds.length > 0 && (
|
||
157 + <span className="faint" style={{ fontSize: '0.8r
|
||
+em' }}>
|
||
158 + {approvableSelected.length} of {selectedIds.le
|
||
+ngth} selected can be merged
|
||
159 + </span>
|
||
160 + )}
|
||
161 + <button
|
||
162 + className="btn btn-primary"
|
||
163 + disabled={approving || approvableSelected.length
|
||
+ === 0}
|
||
164 + title="Approve the selected batches and merge th
|
||
+eir approved frames into the master dataset"
|
||
165 + style={{ fontSize: '0.85rem', padding: '6px 14px
|
||
+', cursor: approvableSelected.length === 0 ? 'not-allowed' : '
|
||
+pointer', display: 'inline-flex', alignItems: 'center', gap: 6
|
||
+, borderRadius: 6 }}
|
||
166 + onClick={approveSelected}
|
||
167 + >
|
||
168 + <CheckIcon size={14} />
|
||
169 + {approving ? 'Approving…' : `Approve & Merge Sel
|
||
+ected (${approvableSelected.length})`}
|
||
170 + </button>
|
||
171 + </div>
|
||
172 </div>
|
||
173 <table className="video-table">
|
||
174
|
||
|
||
● Sekarang kolom checkbox di tabel dan logikanya:
|
||
|
||
● Update(frontend/src/pages/LibraryPage.jsx)
|
||
⎿ Added 12 lines
|
||
174
|
||
175 <thead>
|
||
176 <tr>
|
||
177 + <th style={{ width: 34 }}>
|
||
178 + <input
|
||
179 + type="checkbox"
|
||
180 + aria-label="Select every batch that can be mer
|
||
+ged"
|
||
181 + checked={approvable.length > 0 && approvableSe
|
||
+lected.length === approvable.length}
|
||
182 + onChange={(event) =>
|
||
183 + setSelectedIds(event.target.checked ? approv
|
||
+able.map((b) => b.id) : [])
|
||
184 + }
|
||
185 + disabled={approvable.length === 0}
|
||
186 + style={{ cursor: approvable.length === 0 ? 'no
|
||
+t-allowed' : 'pointer' }}
|
||
187 + />
|
||
188 + </th>
|
||
189 <th>Batch</th><th>Range</th><th>Frames</th><th>Rev
|
||
iewed</th>
|
||
190 <th>Shapes</th><th>Status</th><th />
|
||
191 </tr>
|
||
|
||
● Update(frontend/src/pages/LibraryPage.jsx)
|
||
⎿ Added 18 lines
|
||
199 <React.Fragment key={batch.id}>
|
||
200 <tr>
|
||
201 <td>
|
||
202 + <input
|
||
203 + type="checkbox"
|
||
204 + aria-label={`Select ${batch.batch_label}
|
||
+`}
|
||
205 + checked={selectedIds.includes(batch.id)}
|
||
206 + disabled={(batch.review?.approved ?? 0)
|
||
+=== 0}
|
||
207 + title={(batch.review?.approved ?? 0) ===
|
||
+ 0
|
||
208 + ? 'No approved frame in this batch yet
|
||
+ — review it first'
|
||
209 + : `${batch.review.approved} approved f
|
||
+rame(s) would be merged`}
|
||
210 + onChange={(event) =>
|
||
211 + setSelectedIds((current) =>
|
||
212 + event.target.checked
|
||
213 + ? [...current, batch.id]
|
||
214 + : current.filter((id) => id !== ba
|
||
+tch.id))
|
||
215 + }
|
||
216 + style={{ cursor: (batch.review?.approved
|
||
+ ?? 0) === 0 ? 'not-allowed' : 'pointer' }}
|
||
217 + />
|
||
218 + </td>
|
||
219 + <td>
|
||
220 <span style={{ cursor: 'pointer', borderBo
|
||
ttom: '1px dashed rgba(255,255,255,0.3)' }} onClick={() => edi
|
||
tBatch(batch)} title="Click to rename batch">
|
||
221 {batch.date_label} · {batch.batch_label}
|
||
222 </span>
|
||
|
||
● Update(frontend/src/pages/LibraryPage.jsx)
|
||
⎿ Added 26 lines
|
||
145 }
|
||
146 }
|
||
147
|
||
148 + // Only a batch with at least one approved frame has anythin
|
||
+g to merge; the
|
||
149 + // rest are offered but not selectable, so the reason is vis
|
||
+ible rather than
|
||
150 + // discovered through an error.
|
||
151 + const approvable = batches.filter((batch) => (batch.review?.
|
||
+approved ?? 0) > 0)
|
||
152 + const approvableSelected = selectedIds.filter((id) =>
|
||
153 + approvable.some((batch) => batch.id === id))
|
||
154 +
|
||
155 + async function approveSelected() {
|
||
156 + if (approvableSelected.length === 0) return
|
||
157 + setApproving(true)
|
||
158 + const failures = []
|
||
159 + for (const id of approvableSelected) {
|
||
160 + try {
|
||
161 + await api.approveBatch(id)
|
||
162 + } catch (exc) {
|
||
163 + const batch = batches.find((item) => item.id === id)
|
||
164 + failures.push(`${batch?.batch_label ?? id}: ${exc.mess
|
||
+age}`)
|
||
165 + }
|
||
166 + }
|
||
167 + setApproving(false)
|
||
168 + setSelectedIds([])
|
||
169 + // Report per batch: one failing batch must not hide the o
|
||
+nes that worked.
|
||
170 + if (failures.length) onError(`${failures.length} batch(es)
|
||
+ could not be merged — ${failures.join(' · ')}`)
|
||
171 + onChanged()
|
||
172 + }
|
||
173 +
|
||
174 if (batches.length === 0) return null
|
||
175
|
||
176 return (
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Update(frontend/src/pages/LibraryPage.jsx) ⎿ Added 2 lines
|
||
45
|
||
46 export function BatchList({ project, batches, activeJobs, onCha
|
||
nged, onError }) {
|
||
47 const [busyId, setBusyId] = useState(null)
|
||
48 + const [selectedIds, setSelectedIds] = useState([])
|
||
49 + const [approving, setApproving] = useState(false)
|
||
50
|
||
51 const [appendChoiceBatch, setAppendChoiceBatch] = useState(nu
|
||
ll)
|
||
52 const [autoAnnotateConfig, setAutoAnnotateConfig] = u
|
||
seState(null)
|
||
|
||
Searched for 1 pattern (ctrl+o to expand)
|
||
|
||
● Update(frontend/src/pages/LibraryPage.jsx)
|
||
⎿ Added 1 line, removed 1 line
|
||
2 import React, { useCallback, useEffect, useRef, useState } from
|
||
'react'
|
||
3 import { api, formatDuration } from '../api'
|
||
4 import { navigate } from '../App'
|
||
5 -import { AlertIcon, FolderIcon, ScissorsIcon, SearchIcon, ZapIco
|
||
-n } from '../components/Icons'
|
||
5 +import { AlertIcon, CheckIcon, FolderIcon, ScissorsIcon, SearchI
|
||
+con, ZapIcon } from '../components/Icons'
|
||
6
|
||
7
|
||
8
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Cek apakah handleAutoAnnotateAll jadi tidak terpakai setelah tombolnya saya
|
||
ganti:
|
||
|
||
Searched for 1 pattern, ran 2 shell commands (ctrl+o to expand)
|
||
|
||
● handleAutoAnnotateAll dan ZapIcon di file itu jadi yatim setelah tombolnya
|
||
diganti. Saya bersihkan karena perubahan saya yang membuatnya tidak terpakai:
|
||
|
||
● Update(frontend/src/pages/LibraryPage.jsx)
|
||
⎿ Removed 20 lines
|
||
127 }
|
||
128 }
|
||
129
|
||
130 - async function handleAutoAnnotateAll() {
|
||
131 - if (!window.confirm(`Start SAM3 Auto-Annotation in paralle
|
||
-l for ALL ${batches.length} batch(es)?`)) return
|
||
132 - try {
|
||
133 - await Promise.all(
|
||
134 - batches.map(b => api.startAutolabel(b.id, {
|
||
135 - resume: false,
|
||
136 - append: true,
|
||
137 - engine: 'sam3',
|
||
138 - threshold: 0.35,
|
||
139 - iou_threshold: 0.0,
|
||
140 - min_box_frac: 0.0,
|
||
141 - target_class_names: project.classes.map(c => c.name)
|
||
142 - }))
|
||
143 - )
|
||
144 - onChanged()
|
||
145 - } catch (exc) {
|
||
146 - onError(exc.message)
|
||
147 - }
|
||
148 - }
|
||
149 -
|
||
130 // Only a batch with at least one approved frame has anythin
|
||
g to merge; the
|
||
131 // rest are offered but not selectable, so the reason is vis
|
||
ible rather than
|
||
132 // discovered through an error.
|
||
|
||
Ran 4 shell commands (ctrl+o to expand)
|
||
|
||
● Selesai. Container sudah di-rebuild (index-I65Ek6EH.js) — hard-refresh dulu.
|
||
|
||
Yang berubah di tabel Batches:
|
||
|
||
- Kolom checkbox di paling kiri, plus checkbox "pilih semua" di header tabel.
|
||
- Tombol "Approve & Merge Selected (N)" menggantikan posisi Auto-Annotate All
|
||
Batches yang tadi dobel. Sekarang tombol auto-annotate-all hanya ada satu, di
|
||
header halaman.
|
||
- Checkbox mati otomatis untuk batch yang belum punya frame approved, dengan
|
||
tooltip alasannya ("No approved frame in this batch yet — review it first").
|
||
Jadi Anda tahu sebelum klik, bukan setelah dapat error.
|
||
- Tooltip pada checkbox yang aktif menyebut berapa frame yang akan ikut
|
||
ter-merge, misal "346 approved frame(s) would be merged".
|
||
|
||
Kesepuluh batch Anda sekarang bisa dipilih semua, karena semuanya sudah punya
|
||
frame approved — termasuk yang masih ada pending (batch016 punya 26 pending,
|
||
batch8 punya 23). Yang pending itu ditinggalkan, sesuai permintaan Anda tadi:
|
||
approve = simpan, tidak approve = ignore. Nanti kalau Anda review sisanya,
|
||
batch yang sama bisa di-approve ulang untuk menambahkan frame baru itu.
|
||
|
||
Sekali klik akan memasukkan 1.031 frame ke master dataset
|
||
(113+346+164+32+32+69+51+62+82+80).
|
||
|
||
Dua hal yang saya putuskan:
|
||
|
||
- Approve dijalankan satu per satu, bukan paralel, dan kegagalan dilaporkan
|
||
per batch. Kalau satu batch gagal, sembilan lainnya tetap jalan dan Anda
|
||
diberi tahu persis mana yang gagal beserta alasannya — bukan satu error yang
|
||
menutupi semuanya.
|
||
- Saya hapus fungsi handleAutoAnnotateAll dan impor ZapIcon di LibraryPage.jsx
|
||
yang jadi tidak terpakai setelah tombolnya diganti.
|
||
|
||
Catatan, di luar yang Anda minta jadi saya tidak sentuh: tombol "🔄 Reset
|
||
Auto" di tabel itu memakai emoji sebagai ikon, yang dilarang checklist UI di
|
||
CLAUDE.md §7. Tinggal bilang kalau mau saya ganti ke ikon SVG.
|
||
|
||
Belum saya uji dengan benar-benar mengklik Approve — itu operasi yang menulis
|
||
ke master dataset dan tidak bisa dibatalkan, jadi saya serahkan ke Anda.
|
||
Backup dari sesi tadi (data/backups/app-20260807-163509.db) masih utuh kalau
|
||
hasilnya tidak sesuai harapan.
|
||
|
||
✻ Sautéed for 2m 8s
|
||
|
||
❯ in the data prep also create table so there is like "master dataset date x
|
||
time x" so there will be no conflict if i apllied twice the same batches
|
||
cause the master data set is different, but we can actualy choose them 2
|
||
master datasets and confirm to send to the models & training stages
|
||
|
||
● Ini perubahan arsitektur, bukan sekadar tabel — jadi saya mau pastikan dulu
|
||
maksudnya sebelum menulis kode. Saya cek dulu batasan yang ada sekarang:
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Bagian teknis intinya ada di sini: frame_id ... UNIQUE (db.py:89) — satu frame
|
||
hanya boleh ada di satu dataset, itulah yang bikin konflik saat batch yang
|
||
sama di-merge dua kali. Perlu jadi UNIQUE(dataset_id, frame_id).
|
||
|
||
Kabar baik: perbaikan val-split jadi berbasis hash yang saya buat kemarin
|
||
membuat ini aman. Frame yang sama selalu jatuh ke split yang sama di dataset
|
||
mana pun, jadi menggabungkan 2 master dataset tidak akan membuat gambar yang
|
||
sama muncul di train sekaligus val. Kalau split-nya masih posisional seperti
|
||
dulu, fitur ini akan langsung merusak perbandingan base-vs-new.
|
||
|
||
Empat hal yang saya perlu tahu:
|
||
|
||
● User answered Claude's questions:
|
||
⎿ · Apa yang sebenarnya membedakan master dataset A dan B? → Anotasi sudah
|
||
diperbaiki, Aturan triage berbeda, Kombinasi batch berbeda, Snapshot waktu
|
||
· Kalau 2 dataset dipilih bersama dan ada frame yang sama di keduanya
|
||
dengan label berbeda, mana yang dipakai? → Yang paling baru menang
|
||
(Recommended)
|
||
· Gambar tiap dataset disalin sendiri-sendiri atau berbagi file? → Salin
|
||
penuh tiap dataset
|
||
· Dataset lama (yang sekarang) diapakan? → Jadikan 'Master Dataset #1'
|
||
|
||
● Semua alasan berlaku, jadi dataset-nya harus generik dan dinamai bebas. Saya
|
||
cek kapasitas disk dulu karena Anda pilih salin penuh:
|
||
|
||
Ran 1 shell command (ctrl+o to expand)
|
||
|
||
● Disk 1,5 TB kosong — salin penuh aman (~200 MB per dataset). Tapi dataset.py
|
||
sudah 559 baris, lewat batas 400 di CLAUDE.md §6, dan saya akan mengubahnya
|
||
banyak. Saya pecah sekalian.
|
||
|
||
Mulai dari skema:
|
||
|
||
● Update(backend/db.py)
|
||
⎿ Added 14 lines, removed 2 lines
|
||
83 )
|
||
84 """,
|
||
85 """
|
||
86 + CREATE TABLE IF NOT EXISTS datasets (
|
||
87 + id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||
88 + project_id INTEGER NOT NULL REFERENCES projects(id)
|
||
+ON DELETE CASCADE,
|
||
89 + name TEXT NOT NULL,
|
||
90 + note TEXT NOT NULL DEFAULT '',
|
||
91 + rule_version TEXT,
|
||
92 + created_at REAL NOT NULL
|
||
93 + )
|
||
94 + """,
|
||
95 + """
|
||
96 CREATE TABLE IF NOT EXISTS dataset_items (
|
||
97 id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||
98 project_id INTEGER NOT NULL REFERENCES projects(id) ON
|
||
DELETE CASCADE,
|
||
89 - frame_id INTEGER NOT NULL REFERENCES frames(id) ON D
|
||
-ELETE CASCADE UNIQUE,
|
||
99 + dataset_id INTEGER NOT NULL REFERENCES datasets(id) ON
|
||
+ DELETE CASCADE,
|
||
100 + frame_id INTEGER NOT NULL REFERENCES frames(id) ON D
|
||
+ELETE CASCADE,
|
||
101 split TEXT NOT NULL CHECK (split IN ('train', 'va
|
||
l')),
|
||
102 image_rel TEXT NOT NULL,
|
||
103 label_rel TEXT NOT NULL,
|
||
93 - added_at REAL NOT NULL
|
||
104 + added_at REAL NOT NULL,
|
||
105 + UNIQUE (dataset_id, frame_id)
|
||
106 )
|
||
107 """,
|
||
108 """
|
||
|
||
● Migrasi untuk database yang sudah ada — SQLite tidak bisa mengubah constraint,
|
||
jadi tabelnya harus dibangun ulang:
|
||
|
||
● Update(backend/db.py)
|
||
⎿ Added 55 lines
|
||
207 if "rule_version" not in version_cols:
|
||
208 cur.execute("ALTER TABLE model_versions ADD COLUMN
|
||
rule_version TEXT")
|
||
209
|
||
210 + _migrate_dataset_items(cur)
|
||
211
|
||
212 +
|
||
213 +def _migrate_dataset_items(cur) -> None:
|
||
214 + """Move a single-dataset project onto named datasets.
|
||
215 +
|
||
216 + The old table pinned `frame_id UNIQUE`, so one frame could
|
||
+ live in exactly
|
||
217 + one dataset — which is the constraint that made merging th
|
||
+e same batch twice
|
||
218 + a conflict. SQLite cannot drop a constraint, so the table
|
||
+is rebuilt and the
|
||
219 + existing rows are adopted by a "Master Dataset #1" per pro
|
||
+ject. Nothing is
|
||
220 + discarded: a row without a home would be a merged frame th
|
||
+e app can no
|
||
221 + longer see.
|
||
222 + """
|
||
223 + cur.execute("PRAGMA table_info(dataset_items)")
|
||
224 + columns = [column[1] for column in cur.fetchall()]
|
||
225 + if not columns or "dataset_id" in columns:
|
||
226 + return
|
||
227 +
|
||
228 + import time
|
||
229 +
|
||
230 + cur.execute("SELECT DISTINCT project_id FROM dataset_items
|
||
+")
|
||
231 + projects_with_items = [row[0] for row in cur.fetchall()]
|
||
232 + home = {}
|
||
233 + for project_id in projects_with_items:
|
||
234 + cur.execute(
|
||
235 + """INSERT INTO datasets (project_id, name, note, c
|
||
+reated_at)
|
||
236 + VALUES (?, 'Master Dataset #1', 'Adopted from t
|
||
+he original single dataset', ?)""",
|
||
237 + (project_id, time.time()),
|
||
238 + )
|
||
239 + home[project_id] = cur.lastrowid
|
||
240 +
|
||
241 + cur.execute("ALTER TABLE dataset_items RENAME TO dataset_i
|
||
+tems_old")
|
||
242 + cur.execute("""
|
||
243 + CREATE TABLE dataset_items (
|
||
244 + id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||
245 + project_id INTEGER NOT NULL REFERENCES projects(id
|
||
+) ON DELETE CASCADE,
|
||
246 + dataset_id INTEGER NOT NULL REFERENCES datasets(id
|
||
+) ON DELETE CASCADE,
|
||
247 + frame_id INTEGER NOT NULL REFERENCES frames(id)
|
||
+ON DELETE CASCADE,
|
||
248 + split TEXT NOT NULL CHECK (split IN ('train',
|
||
+ 'val')),
|
||
249 + image_rel TEXT NOT NULL,
|
||
250 + label_rel TEXT NOT NULL,
|
||
251 + added_at REAL NOT NULL,
|
||
252 + UNIQUE (dataset_id, frame_id)
|
||
253 + )
|
||
254 + """)
|
||
255 + for project_id, dataset_id in home.items():
|
||
256 + cur.execute(
|
||
257 + """INSERT INTO dataset_items (project_id, dataset_
|
||
+id, frame_id, split,
|
||
258 + image_rel, label_rel
|
||
+, added_at)
|
||
259 + SELECT project_id, ?, frame_id, split, image_re
|
||
+l, label_rel, added_at
|
||
260 + FROM dataset_items_old WHERE project_id = ?""",
|
||
261 + (dataset_id, project_id),
|
||
262 + )
|
||
263 + cur.execute("DROP TABLE dataset_items_old")
|
||
264 +
|
||
265 +
|
||
266 def healthy() -> bool:
|
||
267 try:
|
||
268 with cursor() as cur:
|
||
|
||
● Sekarang modul registry dataset-nya:
|
||
|
||
● Write(backend/datasets.py)
|
||
⎿ Wrote 175 lines to backend/datasets.py
|
||
1 """Named master datasets — several per project, each a full stan
|
||
2
|
||
3 One project used to have exactly one master dataset, so merging
|
||
4 was a conflict. Now a merge targets a *named* dataset, and the s
|
||
5 into as many as you like: "batch7+8 strict rules" and "batch7+8
|
||
after I fixed
|
||
ifferent labels.
|
||
7
|
||
9
|
||
l}
|
||
… +197 lines (ctrl+o to expand)
|
||
|
||
● Sekarang bagian terberat — dataset.py harus tahu dataset mana yang jadi
|
||
tujuan. Saya baca utuh dulu supaya tidak salah potong:
|
||
|
||
Read 1 file (ctrl+o to expand)
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 23 lines, removed 10 lines
|
||
24 pass
|
||
25
|
||
26
|
||
27 -def dataset_dir(project_slug: str) -> str:
|
||
28 - return os.path.join(config.project_dir(project_slug), "data
|
||
-set")
|
||
27 +def dataset_dir(project_slug: str, dataset_id: int) -> str:
|
||
28 + return datasets.dataset_root(project_slug, dataset_id)
|
||
29
|
||
30
|
||
31 -def approve(batch_id: int) -> dict:
|
||
32 - """Sign a batch off and queue its merge (REQ-045, REQ-050).
|
||
-"""
|
||
31 +def approve(batch_id: int, dataset_id: Optional[int] = None,
|
||
32 + dataset_name: str = "") -> dict:
|
||
33 + """Sign a batch off and queue its merge into one named data
|
||
+set (REQ-045, REQ-050).
|
||
34 +
|
||
35 + Without `dataset_id` a new dataset is created, so merging t
|
||
+he same batch
|
||
36 + again never collides with the earlier result — it produces
|
||
+a second dataset
|
||
37 + holding that batch as it looks now.
|
||
38 + """
|
||
39 batch = batches.get(batch_id)
|
||
40 if batch is None:
|
||
41 raise DatasetError("No such batch")
|
||
36 - if batch["status"] == "merged" and _unmerged_approved(batch
|
||
-_id) == 0:
|
||
37 - # Re-merging is allowed, because a partly-reviewed batc
|
||
-h can be merged
|
||
38 - # for the part that was done and finished later. Only r
|
||
-efuse when there
|
||
39 - # is genuinely nothing new to add.
|
||
40 - raise DatasetError("Every approved frame in this batch
|
||
-is already in the master dataset")
|
||
42 if batch["review"]["approved"] == 0:
|
||
43 raise DatasetError("No frame in this batch is approved
|
||
— there is nothing to merge")
|
||
44 # Frames that are not approved — rejected or never looked a
|
||
t — are simply
|
||
...
|
||
53 if cur.fetchone() is not None:
|
||
54 raise DatasetError("A merge for this batch is alrea
|
||
dy queued")
|
||
55
|
||
56 + if dataset_id is None:
|
||
57 + target = datasets.create(batch["project_id"], name=data
|
||
+set_name,
|
||
58 + rule_version=triage.Resolver(b
|
||
+atch["project_id"]).version())
|
||
59 + dataset_id = target["id"]
|
||
60 + else:
|
||
61 + target = datasets.get(dataset_id)
|
||
62 + if target is None:
|
||
63 + raise DatasetError("No such dataset")
|
||
64 + if _unmerged_approved(batch_id, dataset_id) == 0:
|
||
65 + raise DatasetError(
|
||
66 + f"Every approved frame of this batch is already
|
||
+ in “{target['name']}”")
|
||
67 +
|
||
68 batches.set_status(batch_id, "approved")
|
||
69 job = jobs.create(
|
||
70 "merge",
|
||
58 - params={"batch_id": batch_id},
|
||
71 + params={"batch_id": batch_id, "dataset_id": dataset_id}
|
||
+,
|
||
72 project_id=batch["project_id"],
|
||
73 batch_id=batch_id,
|
||
74 message=f"{batch['date_label']}/{batch['batch_label']}"
|
||
,
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 5 lines, removed 4 lines
|
||
76 return job.to_dict()
|
||
77
|
||
78
|
||
79 -def _unmerged_approved(batch_id: int) -> int:
|
||
80 - """Approved frames of this batch that are not in the datase
|
||
-t yet."""
|
||
79 +def _unmerged_approved(batch_id: int, dataset_id: int) -> int:
|
||
80 + """Approved frames of this batch not yet in *this* dataset.
|
||
+"""
|
||
81 with db.cursor() as cur:
|
||
82 cur.execute(
|
||
83 """SELECT COUNT(*) FROM frames f
|
||
84 - LEFT JOIN dataset_items d ON d.frame_id = f.id
|
||
84 + LEFT JOIN dataset_items d
|
||
85 + ON d.frame_id = f.id AND d.dataset_id = ?
|
||
86 WHERE f.batch_id = ? AND f.review_status = 'appr
|
||
oved' AND d.id IS NULL""",
|
||
86 - (batch_id,),
|
||
87 + (dataset_id, batch_id),
|
||
88 )
|
||
89 return cur.fetchone()[0]
|
||
90
|
||
|
||
● Sekarang job merge-nya menulis ke dataset tujuan:
|
||
|
||
Read 1 file (ctrl+o to expand)
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 6 lines, removed 2 lines
|
||
486 if batch is None:
|
||
487 raise DatasetError("The batch disappeared before the m
|
||
erge started")
|
||
488 project = projects.get(batch["project_id"])
|
||
489 - root = dataset_dir(project["slug"])
|
||
489 + dataset_id = job.params["dataset_id"]
|
||
490 + target = datasets.get(dataset_id)
|
||
491 + if target is None:
|
||
492 + raise DatasetError("The target dataset disappeared bef
|
||
+ore the merge started")
|
||
493 + root = dataset_dir(project["slug"], dataset_id)
|
||
494 for split in ("train", "val"):
|
||
495 os.makedirs(os.path.join(root, "images", split), exist
|
||
_ok=True)
|
||
496 os.makedirs(os.path.join(root, "labels", split), exist
|
||
_ok=True)
|
||
...
|
||
498 frames = [f for f in batches.frames(batch["id"]) if f["rev
|
||
iew_status"] == "approved"]
|
||
499 source_dir = batches.frames_dir(project["slug"], batch["id
|
||
"])
|
||
500 job.progress(0, len(frames))
|
||
497 - job.log(f"Merging {len(frames)} approved frame(s) into the
|
||
- master dataset")
|
||
501 + job.log(f"Merging {len(frames)} approved frame(s) into “{t
|
||
+arget['name']}”")
|
||
502
|
||
503 # Triage gates the merge (REQ-104): a frame carrying an ig
|
||
nored shape never
|
||
504 # enters the master dataset at all, so junk is stopped at
|
||
the door rather
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 2 lines, removed 1 line
|
||
528 annotations = resolved
|
||
529
|
||
530 with db.cursor() as cur:
|
||
531 - cur.execute("SELECT 1 FROM dataset_items WHERE fra
|
||
-me_id = ?", (frame["id"],))
|
||
531 + cur.execute("SELECT 1 FROM dataset_items WHERE dat
|
||
+aset_id = ? AND frame_id = ?",
|
||
532 + (dataset_id, frame["id"]))
|
||
533 if cur.fetchone() is not None:
|
||
534 skipped += 1
|
||
535 job.progress(index + 1, len(frames))
|
||
|
||
Read 1 file (ctrl+o to expand)
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 5 lines, removed 4 lines
|
||
550 handle.write("\n".join(lines) + ("\n" if lines
|
||
else ""))
|
||
551
|
||
552 cur.execute(
|
||
553 - """INSERT INTO dataset_items (project_id, fram
|
||
-e_id, split, image_rel,
|
||
554 - label_rel, added
|
||
-_at)
|
||
555 - VALUES (?, ?, ?, ?, ?, ?)""",
|
||
556 - (project["id"], frame["id"], split, image_rel,
|
||
- label_rel, time.time()),
|
||
553 + """INSERT INTO dataset_items (project_id, data
|
||
+set_id, frame_id, split,
|
||
554 + image_rel, label
|
||
+_rel, added_at)
|
||
555 + VALUES (?, ?, ?, ?, ?, ?, ?)""",
|
||
556 + (project["id"], dataset_id, frame["id"], split
|
||
+, image_rel, label_rel,
|
||
557 + time.time()),
|
||
558 )
|
||
559 added[split] += 1
|
||
560 job.progress(index + 1, len(frames))
|
||
|
||
● Update(backend/dataset.py)
|
||
⎿ Added 3 lines, removed 4 lines
|
||
571 cur.execute("UPDATE batches SET status = 'merged', mer
|
||
ged_at = ? WHERE id = ?",
|
||
572 (time.time(), batch["id"]))
|
||
573
|
||
574 - path = write_data_yaml(projects.get(project["id"]))
|
||
575 - totals = summary(project["id"])["splits"]
|
||
574 + totals = datasets.get(dataset_id)["splits"]
|
||
575 job.log(f"Added {added['train']} train / {added['val']} va
|
||
l"
|
||
577 - + (f", skipped {skipped} already merged" if skippe
|
||
-d else "")
|
||
576 + + (f", skipped {skipped} already in this dataset"
|
||
+if skipped else "")
|
||
577 + (f", held back {triaged_out} by triage" if triag
|
||
ed_out else ""))
|
||
579 - job.log(f"Master dataset now {totals['train']} train / {to
|
||
-tals['val']} val — {path}")
|
||
578 + job.log(f"“{target['name']}” now holds {totals['train']} t
|
||
+rain / {totals['val']} val")
|
||
⎿ You've hit your session limit · resets 5:10pm (Asia/Jakarta)
|
||
/upgrade to increase your usage limit.
|
||
|
||
✻ Cogitated for 3m 4s |