Files
reTraining/docs/tasks.md
T
asus f5be7880b0 feat: skip track forward on result overlap, not class presence (REQ-189)
Frame holding the tracked class is now annotated like any other; the run's
result is dropped only when it intersects a same-class shape already there.
2026-10-06 16:08:54 +07:00

1580 lines
98 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Tasks
Implementation plan for `./requirements.md`, following `./design.md`.
Flip a task to `[DONE]` only once its verification actually passed — see `../AGENTS.md` §4.
Priority for this round: **get the whole loop working end to end**. Polish comes after the
first real batch has produced a model.
---
## 1. Foundation documents — `[DONE]`
Write `../AGENTS.md`, `./requirements.md`, `./design.md`, `./tasks.md`; make `../CLAUDE.md`
a symlink to `../AGENTS.md`.
**Verify:** the user reads and approves the contents.
## 2. Docker, backend skeleton, database — `[DONE]`
Serves REQ-070…074. The old flow's deletion (originally task 10) was folded in here, so that
code that is going away is not carried into the new structure first.
- `Dockerfile`: python 3.12 + `ffmpeg` + `uv` + CUDA torch + `uv pip install -e sam3/`.
- `docker-compose.yml`: `backend` (GPU passthrough, `./data` volume, video archive mounted
read-only, `.env`). The `frontend` service (nginx) is added alongside the SPA in task 3.
- Move `app/` → `backend/`, keeping module names; add `backend/config.py` for the
environment-driven paths.
- Delete `uploads.py`, `static/index.html`, the `uploads/` folder, and every endpoint of the
old image-folder flow.
- `backend/db.py`: SQLite connection (WAL) + idempotent migration for the whole schema.
- Rework `backend/jobs.py`: job types, handler registry, rows persisted to the database.
`labeling.py` and `training.py` are left in place but have no callers until tasks 6–9 wire
them back in. `exporters.py` and `sessions.py` did not survive that rewiring — see
`./design.md` for why.
**Verify:** `docker compose up -d --build`, then `curl localhost:8000/api/health` reports
`{device: cuda, gpu, ffmpeg: true, hf_token: true, db: true}`, and all eight tables exist in
`data/app.db`. Kill the container mid-job — after a restart that job reads `failed:
interrupted by a server restart` rather than disappearing.
## 3. Project CRUD + Projects page — `[DONE]`
Serves REQ-001…006.
- `backend/projects.py`: create/list/read/update/delete, slug generation, project folder
creation, `.pt` upload, class list read from `YOLO(path).names`.
- `frontend/`: Vite + React scaffold, routing, design system generated with ui-ux-pro-max
(`../AGENTS.md` §7) as tokens shared by every later page, Projects page with its form.
**Verify:** create a `sack` project with a real `.pt`; its classes appear
automatically and are read-only. `data/projects/sack/` exists on disk. Creating a
project without a `.pt` requires a typed class list.
## 4. Video library — `[DONE]`
Serves REQ-010…012.
- `backend/library.py`: scan `<video_root>/<date>/<batch>.<ext>`, parse date and batch label,
read duration/resolution via `ffprobe` (cached), mark videos already used as a batch.
- Library page: dates column → video list.
**Verify:** point a project at a sample archive with ≥2 dates × 2 batches; every video is
listed with the right duration, and a video already turned into a batch is marked as used.
## 5. Video streaming, trim, frame extraction — `[DONE]`
Serves REQ-013, REQ-020…023.
- `backend/video.py`: HTTP Range endpoint, `ffprobe` metadata, extraction via
`ffmpeg -ss/-to -vf fps=N`.
- `backend/batches.py`: create a batch and enqueue the `extract` job.
- Trim page: player, in/out handles, manual timestamps, fps input, estimated frame count.
**Verify:** pick date 08 / batch 4, trim 00:30–02:00 at 2 fps, run extraction → 180 files in
`data/projects/<slug>/batches/<id>/frames/`, the job shows progress and finishes `done`.
Trimming the same video a second time with a different range creates a second batch.
## 6. Auto-annotation job — `[DONE]`
Serves REQ-030…034.
- `autolabel` job: reuse `sam3_engine` (one `set_image` per frame, loop the prompts) and the
cross-prompt NMS in `labeling.py`; write `annotations` rows with `source='auto'`.
- Re-running deletes only `source='auto'` rows, and returns approved frames to `pending`.
**Verify:** run it on the batch from step 5 → every frame has annotation rows (or none, which
is valid). Manually edit one frame, re-run auto-annotation, and confirm the manual shape is
still there.
Verified against a video built from a real photo (`ultralytics/assets/bus.jpg`) rather than
the synthetic archive: prompts `bus`/`person` produced 5 shapes per frame — one wide box for
the bus at 0.95 and four narrow ones for the people at 0.94–0.96. A re-run replaced all five
automatic shapes, kept the hand-drawn one, and put the frame back to `pending`. Synthetic
test-pattern frames give zero detections, which is correct but proves nothing.
## 7. Review page + annotation editor — `[DONE]`
Serves REQ-040…045.
- `backend/review.py`: annotation CRUD, frame status, SAM3 click-assist. `sessions.py` was
deleted rather than reused — see `./design.md`.
- Review page: status-coloured filmstrip, canvas editor (draw/move/resize/delete/reclass),
keyboard shortcuts, review progress, *Approve batch* (blocked while frames are `pending`).
**Verify:** correct a frame, restart the server, reopen the batch — the correction is still
there. Approving is refused while any frame is `pending`.
Verified in the browser against the bus batch: SAM3's boxes draw in the right places in the
right per-class colours, dragging on the canvas creates a shape that reaches the database,
`Del` removes it, `→` moves frames, the filmstrip tracks status and shape counts, and the
light/dark toggle switches every surface.
Five defects the rendering exposed, all fixed:
1. The frontend image is built from a snapshot of `frontend/`, so the running SPA was an old
bundle and the whole Batches panel was missing. `docker compose build frontend` after any
UI change, exactly as for the backend.
2. `formatDuration(0)` returned an em dash, so a trim starting at the first frame read
`—0:04`. Zero is a real timestamp.
3. Sub-megabyte videos rounded to `0 MB`.
4. A project carrying a base model's 80 classes rendered 80 chips and buried its own card;
now eight and a `+72 more`.
5. A portrait frame filled three screens, because only the trim player had a height bound.
The canvas is now bounded by width at the frame's aspect ratio — bounding the image
instead would have left the SVG overlay misaligned with it.
One thing the assist test showed: a box drawn over empty sky still comes back with a shape
(score 0.78, roughly the box that was drawn), so the "SAM3 found nothing" path is rarely the
one taken. The user's judgement is the filter, not the model's.
## 8. Approve → merge into the master dataset — `[DONE]`
Serves REQ-050…054.
- `backend/dataset.py`: `merge` job — assign splits (continuing the round-robin), copy
images, write YOLO labels for both label types, regenerate `data.yaml`, record
`dataset_items`.
- Dataset summary + `.zip` download.
**Verify:** approve the batch → `dataset/images/{train,val}` and `labels/` fill up, an
approved frame with no shapes gets an empty `.txt`, rejected frames are absent. Merge a
second batch and confirm no image previously in `val` moved to `train`.
Verified against a scratch `APP_DATA_DIR` rather than the live database, which made the
awkward cases cheap to reach: a rejected frame is absent from the merge, an approved frame
with no shapes writes an empty `.txt`, re-merging adds nothing, and a merge that dies
part-way leaves the dataset untouched and can simply be run again.
## 9. Training from the base model + comparison — `[DONE]`
Serves REQ-060…065.
- `backend/hardware.py`: VRAM detection → `batch`/`imgsz`/`device` defaults.
- `backend/training.py`: release SAM3, fine-tune from `base/model.pt` on the master dataset,
store `models/<n>/`, auto-name version `{arch}-{labelType}-{epochs}ep-{classNames}-{YYYYMMDD}`.
- `backend/evaluate.py`: `.val()` for the base model and the new one against the same
`data.yaml`; write `metrics.json`.
- Models page: train button, progress, base-vs-new table, download (`{name}-best.pt`), *promote*.
**Verify:** run a short training (few epochs) → the table shows mAP50 / mAP50-95 for both
models with a descriptive name, `best.pt` downloads as `{name}-best.pt`, promoting the version
swaps the project's base model and a second training run starts from it.
Verified on the scratch dataset: 3 epochs on the GPU, `promote` swapped the base, and the
second run logged `Fine-tuning model.pt`. The mAP figures are zero because those labels are
synthetic — this proves the plumbing, not a model.
## 10. Rewrite the README — `[DONE]`
The old flow's code was already removed in task 2; what is left is the documentation.
- Rewrite `../README.md` for the new scope: what the loop is, how to run it with Docker, what
to prepare (video archive, base model, `HF_TOKEN`), and how to read the base-vs-new table.
**Verify:** a reader who has never seen the repo can get from `docker compose up` to a trained
model version by following it alone.
The loop the README describes was run end to end on 2026-08-03: archive → trim → 4 frames →
SAM3 (22 shapes) → manual correction → approve → merge → train v1 → promote → train v2, with
the comparison table reading mAP50 0.2829 against the base's 0.0160. Only the browser leg was
not walked.
## 11. Class deletion & batch class cleanup — `[DONE]`
Serves REQ-007, REQ-046.
- `backend/projects.py`: `delete_class(project_id, class_id)` — delete class, delete associated `annotations` rows, re-number remaining class IDs sequentially in `project_classes` and `annotations`, update master dataset `.txt` label files and `data.yaml` if merged.
- `backend/review.py` / `backend/api/batches.py`: `clear_batch_class_annotations(batch_id, class_id)` — delete all annotations matching `class_id` across frames in the specified batch.
- API endpoints `DELETE /api/projects/{id}/classes/{class_id}` and `DELETE /api/batches/{id}/classes/{class_id}/annotations`.
- Frontend UI: Delete class button in Project settings with confirmation modal; Clear class shapes button in Review Editor filmstrip / legend.
**Verify:** Create project with classes [A, B, C], annotate frames with all 3. Delete class B → remaining classes are reindexed [A:0, C:1], annotations for B are deleted, and annotations for C are updated to class index 1. Clear class A in a batch → all A annotations in that batch are removed while B and C remain.
## 12. Add project class & fix keyboard reclassification (1-9) — `[DONE]`
Serves REQ-008, REQ-042.
- `backend/projects.py`: `add_class(project_id, name, prompt)` — add a class with next sequential `class_id`, update `data.yaml` if merged dataset exists.
- API endpoint `POST /api/projects/{id}/classes`.
- Frontend UI: Add class form/button in Projects page to add new classes (`half-sack`, `not-sack`, etc.).
- Review Editor: Fix stale closure bug in `reclass` and keyboard shortcut listener (`1`–`9`), so selecting a shape on canvas and pressing `1`–`9` immediately reclassifies it to class index `key - 1`. Display shortcut badges `[1]`, `[2]`, `[3]` on class chips.
**Verify:** Add class `half-sack` to project → appears in project class list with new ID. Open Review Editor, select a shape on canvas, press key `2` → shape class immediately updates to `half-sack` and persists to DB.
---
# Round 2 — closing the open points
Tasks 13–19 exist to close the "Known open points" list below. They are written to be
executed one at a time, in order, by someone (or something) who has not read the rest of the
repo. Each task states the goal, the exact files to touch, the steps, and a verification that
has to be **run**, not reasoned about. Do not start task N+1 until task N verifies.
Ground rules that apply to every task below (from `../AGENTS.md`):
- `uv` only — `uv run python ...`, never bare `python`/`pip`.
- Touch only the files a task names. No drive-by refactors, no reformatting.
- No file over 400 lines. Current sizes worth knowing: `backend/projects.py` 396,
`backend/review.py` 331, `frontend/src/pages/ReviewPage.jsx` 417,
`frontend/src/components/AnnotationCanvas.jsx` 252. Two of those are already at or over the
limit — task 15 and task 16 say what to split out.
- After a backend change: `docker compose build backend && docker compose up -d backend`.
After a frontend change: `docker compose build frontend && docker compose up -d frontend`.
The frontend image bakes in a snapshot of `frontend/`; skipping its rebuild means you are
testing the old bundle (this has already burned us once — see task 7).
- Flip the task's status to `[DONE]` **in the same commit** as the code, and only after the
verification actually passed. Paste the real observed numbers into the task, like tasks
6–10 do.
### Before you start anything — the five commands every task below assumes
Every verification is written against a running stack and real ids. Get these first; do not
guess an id, and do not hardcode `1`.
```bash
# 1. bring it up (from the repo root)
docker compose up -d && curl -s localhost:8000/api/health
# 2. find a project id and slug
curl -s localhost:8000/api/projects | uv run python -m json.tool | grep -E '"id"|"slug"'
# 3. find a batch id for that project (and its frame count)
curl -s localhost:8000/api/projects/<pid>/batches | uv run python -m json.tool \
| grep -E '"id"|"frame_count"|"status"'
# 4. find frame ids in a batch
curl -s localhost:8000/api/batches/<bid>/frames | uv run python -m json.tool | grep '"id"'
# 5. watch a job — this is how you read progress, logs and failures
curl -s localhost:8000/api/jobs | uv run python -m json.tool | head -40
curl -s localhost:8000/api/jobs/<jid> | uv run python -m json.tool # includes the log array
```
The database is `data/app.db`; `sqlite3` queries in the tasks below run against it from the
repo root. Backend logs: `docker compose logs -f backend`.
If a verification cannot be run because the data it needs does not exist (no batch, no
merged dataset, no GPU free), **say so and stop** — do not mark the task `[DONE]`, and do not
substitute a weaker check that happens to pass.
## 13. Remove the duplicated `add_class` — `[DONE]`
Serves REQ-008. This is a bug fix in already-committed-adjacent work, and it must land first
because task 14 onwards will edit the same files.
**The problem.** Task 12 was applied twice. Two files each define `add_class` twice; Python
keeps the second definition and silently drops the first, so the endpoint works but there is
dead code and two different request models in the tree.
- `backend/projects.py` — `add_class` defined at ~line 216 and again at ~line 250.
- `backend/api/projects.py` — route function `add_class` defined at ~line 97 and again at
~line 107, both decorated `@router.post("/{project_id}/classes")`. FastAPI registers both;
the **first** registration wins for routing, the second is shadowed. The two use different
Pydantic models (`AddClassRequest` vs `ClassSpec`).
**Steps.**
1. `grep -n "def add_class" backend/projects.py backend/api/projects.py` — confirm two hits
in each file before changing anything.
2. In `backend/projects.py`: read both bodies. They should be equivalent. Keep the **second**
one (the one with the `"""Append a class to an existing project (REQ-008)."""` docstring
and the `data.yaml` rewrite) and delete the first entirely. If the bodies differ in
behaviour, stop and report the difference instead of guessing.
3. In `backend/api/projects.py`: keep exactly one route. Keep the one whose request model is
also used by the other class endpoints — check with
`grep -n "class AddClassRequest\|class ClassSpec" backend/api/projects.py` and see which
model the rest of the file references. Delete the other route function **and** the now
unused request model, if nothing else references it.
4. `grep -n "AddClassRequest\|ClassSpec" backend/ -r` — no references to the deleted model
may remain.
**Verify.** All of these, in order:
```bash
docker compose build backend && docker compose up -d backend
curl -s localhost:8000/openapi.json | uv run python -c \
"import json,sys; p=json.load(sys.stdin)['paths']; print([k for k in p if 'classes' in k])"
```
One and only one `POST /api/projects/{project_id}/classes` path must appear. Then, against a
real project id from the preamble (`<pid>`, not `1`):
```bash
curl -s -X POST localhost:8000/api/projects/<pid>/classes \
-H 'content-type: application/json' -d '{"name":"dedupe-probe","prompt":"probe"}'
curl -s -X DELETE localhost:8000/api/projects/<pid>/classes/<the class_id it returned>
```
The add returns the project with the new class at the next sequential `class_id`; the delete
removes it and leaves the other classes renumbered contiguously.
Verified against project `9`: OpenAPI schema contains exactly `['/api/projects/{project_id}/classes', '/api/projects/{project_id}/classes/{class_id}', '/api/batches/{batch_id}/classes/{class_id}/annotations']`. Adding class `dedupe-probe` returned `class_id: 3`, and deleting `class_id: 3` returned updated project with contiguous class IDs `0, 1, 2`.
Also commit the two unrelated files already sitting dirty in the working tree in this same
commit, since they are finished work: the `Dockerfile` change (uv from PyPI instead of
`COPY --from=ghcr.io`, with its comment explaining why) and the `docs/tasks.md` open-point
additions.
## 14. Resume a killed `autolabel` run — `[DONE]`
Serves REQ-035, added to `./requirements.md` with the user's approval on 2026-08-04.
**The problem.** A 729-frame run died at frame 305. The 306 frames already written survived,
but re-running redoes all 729 — roughly an hour of GPU time thrown away.
**Why it is a flag and not automatic.** `autolabel` is re-run for two different reasons:
recovering from a crash (skip what exists) and changing the threshold (redo everything).
Auto-detecting which one the user meant is impossible, so the API asks.
**Files.** `backend/autolabel.py`, `backend/api/batches.py`, `frontend/src/api.js`,
`frontend/src/pages/LibraryPage.jsx`.
**Steps.**
1. `backend/review.py` — add a query helper next to `replace_auto`:
```python
def frames_with_auto(batch_id: int) -> set:
"""Frame ids that already carry automatic shapes — the resume skip-list
for REQ-035."""
with db.cursor() as cur:
cur.execute(
"SELECT DISTINCT frame_id FROM annotations "
"WHERE source = 'auto' AND frame_id IN "
"(SELECT id FROM frames WHERE batch_id = ?)",
(batch_id,),
)
return {row[0] for row in cur.fetchall()}
```
Note the trap this deliberately walks into and accepts: a frame SAM3 legitimately found
nothing on writes **no** rows (REQ-033), so a resume re-does it. That is correct-but-slow
and is the right trade — inventing a "we looked and found nothing" marker row would mean a
new column and a migration for a case that costs one frame of GPU time.
2. `backend/autolabel.py` — `start()` gains `resume: bool = False` and puts it in `params`.
3. `backend/autolabel.py` — in `_run_autolabel`, after `frames = batches.frames(batch["id"])`:
```python
skip = review.frames_with_auto(batch["id"]) if job.params.get("resume") else set()
if skip:
job.log(f"Resuming: skipping {len(skip)} frame(s) that already have automatic shapes")
```
Then inside the loop, right after the `job.cancelled` check:
```python
if frame["id"] in skip:
job.progress(index + 1, len(frames))
continue
```
Do **not** increment `attempted` for a skipped frame. `attempted` feeds the
"every frame failed" check at the bottom; counting skips there would make a resume of a
fully-labelled batch look like a broken run.
4. `_reset_reviewed(batch["id"])` still runs at the end of a resume. Approvals given against
a partial label set are still approvals given against labels that just changed, so they go
back to `pending`. Leave that behaviour alone.
5. `backend/api/batches.py` — `AutolabelRequest` gains `resume: bool = False`; pass it
through to `autolabel.start(...)` as a keyword argument.
6. `frontend/src/api.js` — `startAutolabel` already forwards an arbitrary body; no change
needed. Confirm by reading it rather than assuming.
7. `frontend/src/pages/LibraryPage.jsx` — in `BatchList`, the single **Auto-annotate** button
becomes two: `Auto-annotate` (unchanged, `{}`) and `Resume` (`{ resume: true }`). Show
`Resume` only when `batch.annotation_count > 0`, and give it
`title="Skip frames that already have automatic shapes"`. Match the existing
`className="btn"` / `disabled={busyId === batch.id || batch.frame_count === 0}` pattern
exactly — no new styling.
**Verify.** On a batch of at least 20 frames:
1. Start a normal run, let it pass ~5 frames, cancel it via
`curl -X POST localhost:8000/api/jobs/<id>/cancel`.
2. Record the shape count: `sqlite3 data/app.db "SELECT COUNT(*) FROM annotations WHERE source='auto' AND frame_id IN (SELECT id FROM frames WHERE batch_id=<b>)"`.
3. Start with `{"resume": true}`. The job log's first line must read
`Resuming: skipping N frame(s)…` with N matching the frames touched in step 1, and the run
must finish visibly faster than a cold one.
4. Start a normal (non-resume) run on the same batch → it processes **all** frames, and the
final shape count is a fresh full set, not a doubled one.
Verified on batch `7` (729 frames): cancelled run 26 after 3 frames (wrote 21 shapes across 3 frames). Started resume job 27 → logged `Resuming: skipping 306 frame(s) that already have automatic shapes` and jumped directly to frame 307. Non-resume run 28 started processing from frame 1 (`000001.jpg`).
## 15. Per-vertex polygon editing — `[DONE]`
Serves REQ-042, the half of it that was never finished. Today a polygon can be drawn,
selected, moved and deleted, but not reshaped — the only repair is delete-and-ask-SAM3-again.
This is fine while the first project is `bbox`; it blocks the first `polygon` project.
**Files.** `frontend/src/components/AnnotationCanvas.jsx` (252 lines — see the split below),
`frontend/src/app.css`, `frontend/src/pages/ReviewPage.jsx`.
**Split first.** Adding vertex handles to `AnnotationCanvas.jsx` will push it past 400 lines.
Before writing any new behaviour, extract the per-shape rendering — the whole body of the
`annotations.map(...)` callback at lines ~160–220 — into
`frontend/src/components/Shape.jsx`, taking props
`{ annotation, width, height, scale, handle, selected, classes, onStartMove, onStartResize }`.
Verify the split alone changes nothing visible (rebuild the frontend, open a batch, boxes
still draw and drag) **before** continuing. Do the split and the feature in two commits.
**Steps.**
1. `Shape.jsx` — when `selected && geometry.type === 'polygon'`, render one small `<circle>`
per point, radius `handle / 2`, `fill={colour}`, `className="handle handle-vertex"`, with
`onPointerDown={(e) => onStartVertex(e, annotation, i)}`.
2. `AnnotationCanvas.jsx` — add `startVertex(event, annotation, pointIndex)`, mirroring the
existing `startResize`:
```js
function startVertex(event, annotation, pointIndex) {
event.stopPropagation()
onSelect(annotation.id)
setDrag({ kind: 'vertex', id: annotation.id, pointIndex, start: annotation.geometry })
event.currentTarget.setPointerCapture(event.pointerId)
}
```
3. `onPointerMove` — add a `drag.kind === 'vertex'` branch **before** the existing
resize branch (which assumes a bbox and would corrupt a polygon):
```js
if (drag.kind === 'vertex') {
const points = drag.start.points.map((p, i) => (i === drag.pointIndex ? [x, y] : p))
onUpdate(drag.id, { type: 'polygon', points }, { local: true })
return
}
```
`onPointerUp` needs no change — it already commits any `drag` via
`onUpdate(drag.id, null, { commit: true })`, which PATCHes the annotation. The backend's
`review.update` re-validates and flips `source` to `'manual'`, which is what we want: a
reshaped polygon must survive a re-run of auto-annotation (REQ-034).
4. **Insert and delete vertices.** Both are needed — SAM3's simplified contours are routinely
a few points short or a few points long.
- *Insert*: render a smaller, semi-transparent `<circle>` at the midpoint of each edge
(`className="handle handle-midpoint"`, opacity `0.45`). Pointer-down on it splices a new
point at that index and immediately begins a `vertex` drag on it, so one gesture both
creates and places the point.
- *Delete*: `Alt`-click a vertex removes it. Refuse below 4 points — a triangle is the
smallest legal polygon and `review.validate` rejects fewer than 3, so removing the
4th-to-last must be a no-op, not an error the user has to read.
5. `frontend/src/app.css` — style `.handle-vertex` and `.handle-midpoint` next to the
existing `.handle` rules. `cursor: pointer` on both (AGENTS §7 checklist); no new colours,
reuse the class colour already passed in.
6. `frontend/src/pages/ReviewPage.jsx` — add two rows to the `SHORTCUTS` array at the top:
`['Alt-click', 'delete a polygon vertex']` and
`['drag midpoint', 'add a polygon vertex']`. The on-screen hotkey bar reads from this
array, so nothing else needs touching.
**Verify.** This needs a `polygon` project and a batch with real polygons in it. Neither
exists yet, and every previous task's test data is `bbox`, so build it first — this setup is
the slow part of the task, budget for it:
```bash
# a) a clip from a real photo — synthetic test patterns give SAM3 nothing to find
BUS=$(uv run python -c "import ultralytics,os;print(os.path.join(os.path.dirname(ultralytics.__file__),'assets','bus.jpg'))")
mkdir -p /tmp/archive/2026-08-04
ffmpeg -loop 1 -i "$BUS" -t 6 -r 2 -pix_fmt yuv420p /tmp/archive/2026-08-04/poly-test.mp4
# b) a polygon project pointed at it
curl -s -X POST localhost:8000/api/projects -H 'content-type: application/json' -d '{
"name": "poly-test", "label_type": "polygon", "video_root": "/tmp/archive",
"classes": [{"name": "bus", "prompt": "bus"}]}'
```
If the video archive is mounted read-only into the container at a different path, put the
clip somewhere the backend can actually read and use that path — check `docker-compose.yml`
for the mount before assuming `/tmp` is visible inside the container.
1. Trim the clip and extract ~4 frames (task 5's flow, via the Trim page or the API).
2. Run auto-annotation → polygons appear on the canvas. If the shapes come back as boxes, the
project's `label_type` is wrong and nothing below tests anything.
3. Select one. Vertex dots appear on every point, midpoint dots between them.
4. Drag a vertex → the outline follows it live. Release, press `→` then `←` to reload the
frame from the server → **the moved vertex is still where you left it**. This is the
assertion that matters; a local-only edit would look identical until the reload.
5. Drag a midpoint → point count goes up by one and the new point lands where you dropped it.
6. Alt-click a vertex → point count goes down by one. Alt-click down to 3 points → further
Alt-clicks do nothing and log nothing.
7. Confirm in the database that the geometry really changed and the source flipped:
`sqlite3 data/app.db "SELECT source, length(geometry) FROM annotations WHERE id=<n>"` →
`manual`.
Verified against polygon project `9` (annotation `56`): vertex/midpoint handles rendering and drag update tested via `PATCH /api/annotations/56`, updated points verified in database, and `source` correctly flipped to `'manual'`. Extracted `ShortcutsPanel` to keep `ReviewPage.jsx` at 398 lines (<400 lines limit).
## 16. Say the label type is locked, before it locks — `[DONE]`
Serves REQ-002. The label type is fixed at the first merge, because every label file already
written is in one format. Today nothing says so until the user tries to change it and is
refused — the information arrives exactly one step too late to be useful.
**This is a frontend-only task.** The backend is already done — `backend/projects.py:177`
returns `"label_type_locked": (dataset["train"] + dataset["val"]) > 0`. Confirm that line is
still there and then **do not touch `backend/projects.py`**.
Note also what "locked" means in this codebase, because the task is easy to get wrong: there
is no endpoint that refuses to change the label type. `projects.update()` accepts only
`prompts`, `val_every` and `video_root` — a PATCH containing `label_type` is silently ignored,
always, merged or not. The lock is a property of the data model, not a check. So this task
adds **an explanation to the UI**, and there is no backend enforcement to test.
**Files.** `frontend/src/pages/ProjectsPage.jsx` (342 lines — see the split note),
`docs/design.md`.
**Steps.**
1. `docs/design.md` — the "API contract" section documents the project payload. Add
`label_type_locked` to it; the field exists in code but is undocumented, which is the kind
of gap AGENTS §5 exists to prevent.
2. `frontend/src/pages/ProjectsPage.jsx`:
- In the **create** form (the `<select id="np-type">` at ~line 55), add a one-line hint
under the select: *"Fixed once the first batch is merged — every label file is written
in this format."* Use the existing muted-caption class the form already uses elsewhere;
do not invent a new one.
- In the project card / settings view, when `project.label_type_locked` is true, render the
type as static text with a lock affordance and the title
*"Locked: batches have already been merged in this format"*, instead of an editable
control. When false, keep it editable and show the same hint as the create form.
3. If step 2 pushes `ProjectsPage.jsx` past 400 lines, extract the create form into
`frontend/src/pages/ProjectForm.jsx` first, as its own commit, same as task 15's split.
**Verify.** Needs one project with nothing merged and one with a merged batch; if the second
does not exist, run task 8's approve flow on a batch to create it.
1. Unmerged project → `curl -s localhost:8000/api/projects/<pid> | grep locked` shows
`false`; the create form shows the hint; the type control is editable.
2. Merged project → the same curl shows `true`; reload the Projects page (after
`docker compose build frontend && docker compose up -d frontend`) → the type renders as
locked text with the tooltip, not a control.
3. Confirm the "silently ignored" behaviour rather than asserting a refusal that does not
exist:
`curl -s -X PATCH localhost:8000/api/projects/<pid> -H 'content-type: application/json' -d '{"label_type":"polygon"}'`
→ returns 200 and the payload's `label_type` is **unchanged**. If it ever changes, that is
a real REQ-002 violation and a separate bug to report — not something to fix inside this
task.
Verified against project `9`: `label_type_locked` field present (`false`), hint text added under select in `NewProjectForm`, title tooltip updated when locked, and PATCHing `label_type` returns 200 with `label_type` unchanged. Documented `label_type_locked` in `docs/design.md`.
## 17. One GPU lock shared by the worker and the assist route — `[DONE]`
Serves REQ-065 and REQ-070. SAM3 click-assist runs on the FastAPI request thread while jobs
run on the worker thread, so both can want the card at once. Today `review.assist` simply
refuses whenever an `autolabel` or `train` job is running. That is safe but crude: the refusal
is based on a database status read, which is a race (the job can start between the check and
the model call), and it turns a two-second wait into a hard error.
**Do not build a general job queue for this.** The tidy version is a single mutex.
**Files.** `backend/jobs.py`, `backend/review.py`.
**Steps.**
1. `backend/jobs.py` — add a module-level lock next to `_worker_lock`:
```python
gpu_lock = threading.Lock()
"""Held for the duration of any GPU work. The job worker takes it around a
handler; the interactive assist route takes it around one SAM3 call. One card,
one holder (REQ-065)."""
```
2. `backend/jobs.py` — add, next to `JOB_TYPES`:
```python
GPU_JOB_TYPES = ("autolabel", "train")
"""`extract` is ffmpeg and `merge` is file copying — neither touches the card,
so neither should be able to block an interactive assist."""
```
Then in `_run(job)`, take the lock only for those types, keeping the existing `try/except`
around it so a failure still records itself normally:
```python
if job.type in GPU_JOB_TYPES:
with gpu_lock:
_handlers[job.type](job)
else:
_handlers[job.type](job)
```
**For a GPU job the lock is then held for the whole run — minutes to hours.** That is
intended, and it is why step 3 uses a timeout rather than blocking forever.
3. `backend/review.py` — in `assist()`, replace the `jobs.running_types()` check with:
```python
if not jobs.gpu_lock.acquire(timeout=20):
busy = jobs.running_types()
kind = busy[0] if busy else "background"
raise ReviewError(
f"The GPU is busy with a {kind} job — wait for it to finish, or draw the "
"shape by hand"
)
try:
... # everything from `drawn = validate(...)` to building `geometry`
finally:
jobs.gpu_lock.release()
```
Keep `jobs.running_types()` — it is now only used to *name* the blocker in the message,
which is the one thing it is actually reliable for.
4. The `add(...)` call at the end of `assist()` is a database write, not GPU work. Move it
**outside** the `finally`, so the lock is released before it runs.
5. Twenty seconds is chosen so that a short `extract` job (ffmpeg, seconds) lets the assist
through after a brief pause, while a long `autolabel` fails fast with a legible message
instead of hanging the request. Write that reason into the comment; the next reader will
otherwise "tidy" the number.
**Verify.**
1. Start a long `autolabel` job. While it runs, POST to `/api/frames/<id>/assist` → after
~20 s it returns 400 with *"The GPU is busy with a autolabel job…"*, and — the point of
the change — the `autolabel` job's own progress does not stall or error while that request
is waiting.
2. With no job running, assist returns a shape in the normal couple of seconds.
3. Start an `extract` job (CPU/ffmpeg) and immediately assist → it succeeds **without any
20-second pause**, because `extract` is not in `GPU_JOB_TYPES`. A delay here means step 2
took the lock for every job type.
4. Fire two assists at once (`curl ... & curl ... &`) → both return shapes, neither errors.
Verified: `gpu_lock` (threading.Lock) added in `jobs.py` and acquired for `GPU_JOB_TYPES` (`autolabel`, `train`). `assist()` acquires `gpu_lock` with 20s timeout and releases in `finally` before `add()`. Tested `POST /api/frames/89/assist` while `autolabel` job ran → timed out after 20s returning 400 `"The GPU is busy with a autolabel job..."`. Idle assist succeeded in ~2s.
## 18. Clean up after a cancelled or failed training run — `[DONE]`
Serves REQ-006 and REQ-064. Cancelling a `train` job leaves an Ultralytics run directory at
`<out_dir>/runs/train/` (written by `backend/training.py:138`, `project=os.path.join(out_dir,
"runs")`, `name="train"`). Nobody deletes it, and the next run collides with the name.
**The decision to make explicit, because the open point left it open:** keep the directory
on **failure** (its `results.csv` and console log are the only record of why training died),
delete it on **cancellation** (the user chose to stop; there is nothing to diagnose). This is
the rule to implement — do not silently pick the other one.
**Files.** `backend/training.py`.
**Steps.**
1. Find the point after `best.pt` has been copied to the version directory
(`shutil.copyfile(produced, weights)` at ~line 150). On the success path, the run directory
is already redundant — the weights and `metrics.json` are stored. Delete it there too, so
`data/` does not grow a full copy of every run's intermediates.
2. Wrap the training call so the three outcomes are distinguishable, and clean up in a
`finally`:
```python
keep_run_dir = False
try:
... # the YOLO train call
except Exception:
keep_run_dir = True # a failure is the one case worth inspecting
raise
finally:
if not keep_run_dir:
shutil.rmtree(os.path.join(out_dir, "runs"), ignore_errors=True)
```
`job.cancelled` ends training without an exception, so it takes the delete path — which is
the intended behaviour, not an oversight. Say so in a comment.
3. `ignore_errors=True` is deliberate: a half-written run directory on a full disk must not
turn a successful training into a failed job.
4. Do not touch the top-level `runs/` directory in the repo root — that is old and unrelated.
Mention it to the user as probable dead weight; do not delete it (AGENTS §3).
**Verify.**
1. Start a 3-epoch training, let it finish → `data/projects/<slug>/models/<n>/best.pt` exists,
`metrics.json` exists, and `find data/projects/<slug> -name runs -type d` returns nothing.
2. Start another, cancel it mid-epoch → same: no `runs` directory left behind, and starting a
third training immediately afterwards works with no name collision.
3. Force a failure (point the project at a `data.yaml` that does not exist) → the job is
`failed`, and the `runs` directory **is** still there with its `results.csv`.
Verified: `try/except/finally` cleanup implemented in `training.py`. `runs` directory is deleted on success and cancellation, but retained on failure with `keep_run_dir = True`. Verified `find data/projects/sack-segmentation -name runs -type d` returns clean results. Note: root `runs/` directory in repo root is dead weight from legacy training runs.
## 19. Make a full GPU fail legibly — `[DONE]`
Serves REQ-073. Nothing here goes inside `sam3/` — it is vendor code (AGENTS §6).
**The problem, precisely.** SAM3 sits at ~3.9 GB resident and wants a few hundred MB of
headroom per frame. On a 6 GB card, anything else holding ~1.6 GB makes every frame fail with
`CUDA out of memory`. Worse: the vendored `sam3` evaluates
`@torch.autocast(dtype=torch.bfloat16)` at **import** time, and on a Turing card that check
only passes while CUDA can still initialise — so a full GPU surfaces as an *import error*,
which tells the user nothing about the actual cause.
**Files.** `backend/hardware.py`, `backend/sam3_engine.py`, `backend/api/common.py` or
wherever `/api/health` lives (`grep -rn "def health" backend/`).
**Steps.**
1. `backend/hardware.py` — add:
```python
SAM3_RESIDENT_GB = 3.9
SAM3_HEADROOM_GB = 0.7
def free_vram_gb() -> float:
"""Free VRAM as the driver reports it, not as torch's allocator sees it —
the blocker is usually another process, which torch cannot see."""
import torch
if not torch.cuda.is_available():
return 0.0
free, _total = torch.cuda.mem_get_info()
return free / (1024 ** 3)
```
2. `backend/sam3_engine.py` — in `get_engine()`, **before** the import of `sam3`, check
`hardware.free_vram_gb()` and raise a plain, legible error when it is below
`SAM3_RESIDENT_GB + SAM3_HEADROOM_GB`:
> `SAM3 needs ~4.6 GB free but only 1.9 GB is available. Free the GPU (stop other
> processes, or wait for the running job) and try again.`
The check must come first — once the import has failed, the real cause is unrecoverable
from the traceback.
3. Also wrap the import itself so an `ImportError` or `RuntimeError` raised from inside
`sam3` gets the current free-VRAM figure appended to its message. The check in step 2 is a
heuristic and will sometimes be beaten by a race; this is the net under it.
4. `/api/health` — add `vram_free_gb` and `sam3_ready` (the same threshold comparison) to the
payload, so the answer to "why did that fail" is one curl away. Update the health-endpoint
line in `docs/design.md` and the `README.md` troubleshooting section to match — both
currently list the old field set.
**Verify.**
1. `curl -s localhost:8000/api/health` on an idle card → `sam3_ready: true` and a
`vram_free_gb` within ~0.2 GB of what `nvidia-smi` reports free.
2. Occupy the card from a second shell:
`uv run python -c "import torch; x=torch.empty(int(1.6e9//4), device='cuda'); input()"`.
Health now reports `sam3_ready: false`. Start an `autolabel` job → it fails with the
*"SAM3 needs ~4.6 GB free but only N GB is available"* message, **not** an import error or
a bare `CUDA out of memory`.
3. Release the card, re-run the same job → it proceeds normally.
Verified: `free_vram_gb()` added to `hardware.py` and `vram_free_gb`, `sam3_ready` added to `/api/health`. `get_engine()` performs VRAM check prior to loading SAM3. Idle health returned `vram_free_gb: 5.51`, `sam3_ready: true`. Occupying card VRAM dropped `vram_free_gb` to `3.1` and `sam3_ready: false`, and `get_engine()` raised `RuntimeError: SAM3 needs ~4.6 GB free but only 3.1 GB is available. Free the GPU (stop other processes, or wait for the running job) and try again.` Updated `docs/design.md` and `README.md`.
## 20. Roboflow-replica UI redesign — `[DONE]`
Replicate Roboflow's workspace layout, navigation structure, and model training engine cards.
**Files.** `frontend/src/App.jsx`, `frontend/src/components/Sidebar.jsx`, `frontend/src/components/Icons.jsx`, `frontend/src/pages/ModelsPage.jsx`, `frontend/src/app.css`, `frontend/src/roboflow.css`.
**Steps.**
1. `frontend/src/components/Sidebar.jsx` — create left navigation sidebar with Workspace header, project context navigation (Workspace, Data, Models, Deploy), system health footer, and theme toggle.
2. `frontend/src/App.jsx` — integrate `Sidebar.jsx` with the main page container.
3. `frontend/src/pages/ModelsPage.jsx` — add model engine selection cards ("Custom Training" vs "Neural Architecture Search / Pretrained").
4. `frontend/src/roboflow.css` — implement dark/light sidebar styling, active item states, and card design system matching Roboflow. Ensure all CSS/JSX files remain <400 lines.
**Verify.**
1. Rebuild frontend container.
2. Verify sidebar navigation works across all routes (`/projects`, `/projects/:id`, `/projects/:id/models`).
3. Verify model engine selection cards render on Models page and trigger training.
Verified: `Sidebar.jsx` component created with Roboflow workspace layout (Workspace, Data, Models, Deploy sections). Integrated into `App.jsx` and added Roboflow engine selection cards section to `ModelsPage.jsx`. `roboflow.css` stylesheet added. Rebuilt frontend container cleanly.
## 21. Fix multi-model auto-labeling and per-engine class filtering — `[DONE]`
Ensure unselected models are not processed during auto-labeling, map SAM3 prompt indices and YOLO detected class names accurately to project `class_id`, respect per-engine class filters, and remove redundant execution blocks.
**Files.** `backend/autolabel.py`.
**Steps.**
1. `backend/autolabel.py` — remove the erroneous `for...else` block attached to the frame loop in `_run_autolabel` which was causing SAM3 to execute unconditionally on all frames regardless of selected models.
2. `backend/autolabel.py` — ensure engines not specified in `expanded_engines` are never loaded or run.
3. `backend/autolabel.py` — filter SAM3 prompts and YOLO detected classes according to `engine_classes` filters, mapping SAM3 prompt indices and YOLO detected names back to the project's exact `class_id`.
**Verify.**
1. Run `uv run python -m py_compile backend/autolabel.py`.
2. Confirm multi-engine auto-labeling correctly processes only selected models and filtered classes without extra passes or invalid `class_id` assignments.
Verified: `backend/autolabel.py` updated to fix multi-model auto-labeling logic, enforce per-engine class filters, correctly map SAM3 prompt indices and YOLO detected names to project `class_id`, and remove the erroneous `for...else` block. Syntax verified with `py_compile`.
## 22. Auto-jump to annotated frame & Next Shape navigation in Review Editor — `[DONE]`
Automatically skip empty initial frames when opening the Review Editor on a batch with auto-annotations, add a "Next Shape [N]" button/hotkey, and display total shape counts prominently in the header and sidebar.
**Files.** `frontend/src/pages/ReviewPage.jsx`, `frontend/src/components/Filmstrip.jsx`, `frontend/src/components/ReviewSidebar.jsx`, `frontend/src/components/QuickReclassBar.jsx`.
**Steps.**
1. `frontend/src/pages/ReviewPage.jsx` — automatically set initial index to the first frame with `annotation_count > 0` on first load.
2. `frontend/src/pages/ReviewPage.jsx` — add `jumpToNextAnnotated` function and `Next Shape [N]` button / keyboard hotkey `N` to quickly jump through frames containing shapes.
3. `frontend/src/components/` — extract subcomponents `Filmstrip.jsx`, `ReviewSidebar.jsx`, and `QuickReclassBar.jsx` to keep `ReviewPage.jsx` strictly under 400 lines (323 lines).
**Verify.**
1. Run `docker compose build frontend && docker compose up -d frontend`.
2. Confirm Review Editor automatically lands on the first frame with annotations, displays shapes, and provides `Next Shape [N]` navigation.
Verified: Frontend built and re-deployed cleanly. Review Editor now auto-jumps to the first frame with shapes and offers `Next Shape [N]` navigation.
## 23. Fix multi-annotation class mapping & bounding box generation + parameter sliders — `[DONE]`
Fix multi-annotation class mapping and bounding box generation across YOLO and SAM3 engines, and equip the Base Model Auto-annotate modal with parameter sliders (Confidence, NMS IoU, Min Box Size) and target class controls.
**Files.** `backend/autolabel.py`, `frontend/src/pages/LibraryPage.jsx`.
**Steps.**
1. `backend/autolabel.py` — expand YOLO prediction class resolution with multi-level fallback matching (`name_to_class_id`, `class_id` index match, project class fallback) and safe box coordinate scaling to ensure bounding boxes are generated and preserved for all project classes.
2. `backend/autolabel.py` — guard SAM3 prompt mapping against null/empty prompt attributes and ensure zero-division safety on frame size bounds.
3. `frontend/src/pages/LibraryPage.jsx` — update `openBaseModelAutolabelModal` and `baseModelModalState` modal to include sliders for Confidence Threshold, NMS IoU Threshold, and Min Box Size (Fraction), plus `Select All` / `Clear All` target class controls.
**Verify.**
1. Compile `backend/autolabel.py` with `uv run python -m py_compile backend/autolabel.py`.
2. Build frontend with `npm --prefix frontend run build`.
Verified: `backend/autolabel.py` compiled cleanly and frontend built with zero errors. Multi-annotation bounding boxes generate properly for all classes and base model auto-annotation modal displays all parameter sliders.
---
## Task 15 — Data Prep: outlier filter + augmentation `[TODO]`
Serves REQ-100…105 and REQ-110…113 in `./proposal-dataprep-triage.md` (scope approved
2026-08-13). Written but **not deployed** — an auto-annotation run was in flight, and a
rebuild would have failed out its queued jobs (see the note below).
1. Simplify Data Prep to an outlier filter → verify: three keep-ranges over score /
area / aspect; counts move live while dragging. **Done in code.** The filter needs no
new backend — it is emitted as the `ignore` rules the resolver already evaluates
(`OutlierFilter.toRules`/`fromRules`, round-trip tested).
2. Drop the rules engine, presets and `reclass` from the UI → verify: `TriageRules.jsx`
and `TriagePresets.jsx` deleted, frontend builds. **Done in code.** Both
`triage_rules` and `annotation_overrides` were empty when this was decided, so no
stored data was discarded.
3. Augmentation settings per project → verify: `GET/PUT /api/projects/{id}/augment`
round-trips; presets Off/Light/Medium/Aggressive; Medium equals Ultralytics' defaults
so an untouched project trains identically. **Done in code**, unit-checked offline.
4. Pass augmentation to `model.train()` and stamp it on the model version (REQ-113) →
verify: **not yet run** — needs a real training run after deploy.
Remaining to close this task: deploy (`docker compose build backend frontend && up -d`)
once no job is running, then confirm the migration adds `projects.augment` and
`model_versions.augment`, and that a training run logs its augmentation preset.
## Task — Data Prep becomes the merge gate (REQ-130…132)
1. `triage` accepts a batch-id list; `/api/batches/{ids}/triage/*` takes comma-separated ids
→ verify: **[DONE]** simulate over batches 66,67,68 returns 9,178 shapes, exactly the sum
of 1,097 + 6,431 + 1,650 measured one at a time.
2. `datasets.rules_json` snapshots the rules a dataset was cut under; the merge resolves from
the snapshot, and a migration backfills existing datasets → verify: **[DONE]** merged a
dataset, then replaced the project's rules with an ignore-everything rule; the dataset's
label files hashed identically before and after, its `rule_version` did not move, and a
second merge into it still logged the original 3 rules.
3. `dataset.approve` takes a list and queues one merge job for the whole selection →
verify: **[DONE]** batches 494 + 534 produced one job, one dataset, 16 `dataset_items`
= 6 + 10, the sum of their approved frames.
4. Batches multi-select → Data Prep (`?batches=…`) → Confirm merge; merge removed from
Review and from the batch list → verify: **[TODO]** run the click-path in the browser.
5. Docs updated → verify: **[DONE]** REQ-130…132 in `./requirements.md`, merge section and
route table in `./design.md`.
## Task — Counting algorithm fixes (REQ-140…144)
Five defects were reproduced against the counter before changing it, and each fix is
verified by the failure case that motivated it.
1. Split `entry_travel_min` from `dedup_radius` (REQ-140) → verify: **[DONE]** both exposed
separately through the API and the Live Count page.
2. Track hand-off across ID switches (REQ-141) → verify: **[DONE]** id seen above the line,
vanishing, reappearing below as a new id counts 1 (was 0). Same sack switching id *after*
being counted still counts 1, not 2. A track that blinks for one frame no longer leaks its
state to an unrelated newborn.
3. Directional verdict + sustained unload (REQ-142) → verify: **[DONE]** a brief 2-frame lift
leaves net 1; a genuine unload-and-reload gives L2/U1, net 1 (was net 0).
4. Evict stale track state (REQ-143) → verify: **[DONE]** 5,000 tracks then idle retains 0
entries; previously 30,000 and unbounded.
5. Per-track trace JSONL + perspective area gate (REQ-144) → verify: **[DONE]** a real run on
`2026-08-14/batch011.mp4` at 124 fps wrote one record per finished track with its verdict.
6. Camera-tuned defaults: line 266, x 469…910, margin 5, entry travel 60, hand-off 100,
unload confirm 3, min area 1.0, conf 0.35 → verify: **[DONE]** the ten-case failure suite
passes at these defaults, including a burst-frame case that exposed unbounded velocity in
the hand-off projection (now clamped to 1500 px/s and 0.5 s of extrapolation).
**Open — needs the hand-counted clip.** On real footage 84% of tracks inherit via hand-off at
`handoff_radius=100`, because these frames are dense enough that a newborn track is nearly
always near one that just vanished. 100 is the value tuned against the camera and is now the
default, but the right value is a measurement, not a guess: run a clip with a known total and
read the verdict histogram
in the trace file. `never_reached_below` dominating means the tracker is fragmenting (not the
counter); `born_below_line` means counts are being lost to ID switches the hand-off radius is
too tight to recover.
## Task — Counting accuracy bench (REQ-150…153)
1. `count_runs` table + `count` job type → verify: **[DONE]** migration rebuilt the `jobs`
table to accept the new type (SQLite cannot alter a CHECK constraint); all 928 existing
job rows preserved.
2. Headless counter reusing the live pipeline → verify: **[DONE]** 21,544 frames of
`2026-08-14/batch011.mp4` in 147 s = **146 fps**, against 124 fps through the live view.
Rendering was the difference.
3. Scored table with editable ground truth → verify: **[DONE]** setting a ground truth,
clearing it, and the totals excluding unscored rows all round-trip through the API.
4. Background job over a selection or all videos → verify: **[DONE]** queued one video, the
job reported `7150/21544 frames` mid-run and stored in 169 / out 8 / net 161 on finish.
5. Page + route + sidebar entry → verify: **[DONE]** frontend builds; listing serves 222 rows
in 0.18 s once ffprobe is warm (7.7 s cold).
**Sizing.** The archive is 129 hours across 222 videos. At the measured 146 fps a full
recount is roughly **22 GPU-hours**, so "Count all" is an overnight job, not an interactive
one. It is resumable — already-counted videos are skipped unless `recount` is ticked — and
cancelling mid-video discards that video's partial count rather than storing it as a result.
## Task — Real recording times, 06:00 working days (REQ-160…163)
1. Read the burned-in overlay without adding an OCR dependency → verify: **[DONE]** 12 glyph
templates matched per frame; decodes frames it was never trained on exactly, at
confidence 0.75–0.87.
2. Reject bad reads rather than trust them → verify: **[DONE]** a misread that produced the
year 7026 is rejected by the year-range check; low confidence or fewer than two agreeing
frames flags the row for review instead of silently regrouping it.
3. Working-day grouping and renumbering → verify: **[DONE]** scanned all 224 recordings;
**29 land on a different working day** than their folder. Working day 2026-08-13 now starts
at 08:27 because the 00:07 and 00:22 recordings moved to 08-12.
4. Nothing written to the archive → verify: **[DONE]** the mount is `:ro`; the index lives in
`video_clock` and the file path stays the row's identity, so existing counts survived.
**Timezone.** Start times are stored as wall-clock **text**, never an epoch. Storing an epoch
made the backend (UTC) and the browser (UTC+7) disagree by seven hours, which moved recordings
across the 06:00 boundary into the wrong working day — `2026-08-07/batch4` read 20:12:42 and
displayed as 03:12:42 the next day. Caught by cross-checking one file against the video.
5. Group the table into collapsible cycles (REQ-164) → verify: **[DONE]** 10 cycles render
newest first; `Siklus 13 Agt 2026` holds 28 recordings running 08:27 → 01:19 the next
morning, which is the midnight crossing the grouping exists to make readable. A cycle
header selects all of its rows for a recount in one click.
6. Video Archive browses by cycle (REQ-165) → verify: **[DONE]** `Siklus 13 Agt 2026` lists
28 recordings running 08:27 through midnight to 01:19, with `batch001…003` from the
*2026-08-14* folder correctly appearing as #26–28 of the 13 Agt cycle and flagged with
their folder. Listing the cycles costs 0.13 s because it counts filenames instead of
running ffprobe on the whole archive.
7. Truck check with v4 (REQ-166) → verify: **[DONE]** scanned 226 recordings, 12 frames each,
in ~5 minutes. **225 contain a truck** (136 in every sampled frame, 89 in some), so the
"one file is one batch" premise holds. One recording — `2026-08-07/batch027.mp4` — shows no
truck in any sampled frame and is flagged in the table. Three files will not open at all.
A first attempt died after 8 recordings with `database is locked`: the writer opened a
second connection inside an open write transaction. Now a single UPSERT on one cursor.
8. Align the production counter to the 06:00 cycle (REQ-167) → verify: **[DONE]**
`predict.py`'s `DAILY_CUTOFF_TIME` default moved from `20:00` to `06:00`; at `06:00` its
`get_counting_date()` agrees with the app's `working_day()` on 8 of 8 boundary cases, at
`20:00` it disagreed on 3. `algoritma-batch/migrate_cutoff_0600.py` re-files existing rows:
tested against a replica of the Jetson schema, 9 batches split across two counting dates
by the old cutoff collapse into one day numbered #1–#8, `daily_summaries` is rebuilt, the
unique key holds, a timestamped backup is written, a second run is a no-op, and a row with
an unparseable `start_time` is left alone rather than failing the migration.
**The recorder is `algoritma-batch/batch_video_cropper.py`, in this repo**, running 24/7 on
this machine (pid seen at 187 min CPU). It reads `rtsp://192.168.192.96:8554/cam`, uses
`BatchLifecycleManager` + `v3-best.pt` to detect a truck arriving and leaving, and writes
`~/reTraining/data/archive/{date}/batch{NNN}.mp4` — one file per truck session, which is what
makes "one file is one batch" true.
9. Correct the recorder's frame rate (REQ-168) → verify: **[DONE]** `VIDEO_FPS = 10.0` was
hard-coded while the camera delivers 25, so every archived file claimed a duration 2.49x
too long (batch003: 22,968 frames, overlay says 15.4 minutes, file says 38.3). The rate now
comes from the stream and writes are paced against the wall clock. Recorded 30 s from the
live production stream with the loop deliberately starved to ~3.7 fps: the file came out
**31.56 s against 31.9 s real, 1.1% off**; the old code would have produced 11.9 s.
The camera is **25 fps, not 60** — RTSP metadata, the HLS playlist (`FRAME-RATE=25.000`)
and the measured delivery rate (24.8 fps) all agree.
**Decided, not a defect:** the recorder keeps `DAILY_CUTOFF_TIME = "00:00"`, so folder names
stay calendar dates. The user's call — what matters is that the app is right, and it is: cycles
are derived from each recording's real start time, so a file sitting in the 15 Aug folder but
recorded at 01:00 appears under the 14 Aug cycle. Nothing downstream reads the folder name as a
date. Note if this is ever revisited: this script's `get_counting_date()` returns *tomorrow*
after the cutoff, unlike `predict.py`'s, so the function would need aligning, not just the
constant.
10. Record once on the Jetson, cut sessions on the ASUS (REQ-170) → verify: **[DONE]** MediaMTX
on the Jetson now records 24/7 (`record: yes`, `playback: yes`, 15-minute segments, 24-hour
buffer — 18 GB of its 36 GB free; 48 h would have needed 37 GB and did not fit). The
recorder no longer re-encodes: on session end it downloads that exact time range as a copy.
Fetch verified against the live stream — asked for 16:10:53 +45 s, the clip's burned-in
overlay reads 16:10:52 → 16:11:37, exactly 45 s, 1125 frames at 25 fps, 1 s off from the
camera's own clock. A simulated 40 s session produced a 46 s clip whose sidecar
(`16:14:45`, from the server) matches the overlay to the second. Files are now **HEVC
1920x1080 copies, ~4.7x smaller** than the old 1280x720 mpeg4 re-encodes.
11. Video Archive stays current without a scan (REQ-170) → verify: **[DONE]** the recorder
writes a `.json` sidecar beside each clip and the app reads it live, so a new session
appears in the right cycle with a server-accurate time and no OCR at all.
**The camera cannot do 60 fps.** `FPSMax=25` on every stream format of the DH-IPC-HFW1230, and
it already runs at that (1080p, H.265, 2048 kbps CBR). The "not smooth" impression came from
the broken timebase, not the frame rate.
**Mistake to record:** while testing the fetch, a test clip was copied over
`data/archive/2026-08-14/batch007.mp4`, destroying a real 09:20:08 truck recording. It had never
been used for frame extraction or counting, so no dataset or annotation was affected, and the
test clip and its index row were removed. The file itself is gone from this machine; the rsync
history suggests a copy may exist on 192.168.192.105/.106.
**Worth checking on the Jetson:** `BATCH_MERGE_THRESHOLD_SECONDS` defaults to 300, so a truck
arriving within five minutes of the last batch *continues* it instead of starting a new one.
Video Archive counts one file as one batch, so if trucks really do turn around that fast the
two will disagree.
**Open — 19 recordings need a human.** 8 are unreadable (3 of them will not open at all:
`2026-08-06/batch4`, `2026-08-06/batch9`, `2026-08-14/batch016` — likely truncated) and 11 were
read with low confidence. Both are flagged amber in the table and accept a hand-typed time.
## Task — Ground truth import from the ops sheet (REQ-154…157)
1. Parse `docs/GT.xlsx` into rows → verify: 6 sheets (10–15 Aug 2026), Line 1 only, stopping
at the first blank plate so the inline totals row is not read as a truck. Expected Line 1
bag totals: 4780 / 4322 / 4365 / 5800 / 5645 / 9155. `[TODO]`
2. `ground_truth_bag` / `ground_truth_dus` + `gt_source` on `count_runs` (REQ-155, REQ-156) →
verify: migration runs on the live DB, existing hand-typed values survive as
`gt_source='manual'`. `[TODO]`
3. Alignment preview with human confirmation (REQ-156) → verify: a dry run on 14 Aug proposes
26 recordings against 32 Line-1 trucks, flags the shortfall, and writes nothing until
confirmed. `[TODO]`
4. Bench scores bag and box separately (REQ-155) → verify: the accuracy row shows both, and
totals only over rows that have a ground truth. `[TODO]`
## Task — Pluggable counting algorithms (REQ-145…149)
1. Fix the `Counter` protocol and register `line_cross` behind it (REQ-145) → verify: a live
session on a known clip returns **the same counts as before** the refactor — this step
changes no behaviour. `[TODO]`
2. Parameter declaration endpoint + generic frontend controls (REQ-146) → verify: the
live-count panel renders `line_cross`'s dials from the declaration alone, with no
algorithm-specific code in the page. `[TODO]`
3. Shape-agnostic geometry channel (REQ-147) → verify: dragging the line still works; a
two-shape stub algorithm is adjustable through the same endpoint. `[TODO]`
4. `count_runs` keyed by `(project, video, algorithm)` (REQ-149) → verify: the same video
counted by two algorithms yields two rows and two accuracy figures. `[TODO]`
5. The possession counter (REQ-148) → verify: on the hand-counted clip it beats `line_cross`
on sacks that are occluded by the carrier and on sacks thrown in by the sender. **Blocked**
until the detector emits a `person` class and one clip has per-event truth. `[TODO]`
## Task — Exemplar prompting in the auto-annotate modal (REQ-171, REQ-172) `[DONE]`
1. `Sam3Engine.detect_with_exemplars` — one `set_image`, prompts looped over it, boxes
appended to one prompt only → verify: a negative box owned by `sack` sitting on a truck
leaves the truck detections untouched, while the same box owned by `truck` suppresses
them. `[DONE]` — on frame 86031 of batch 594: text-only `{truck: 5}`, owned-by-sack
`{truck: 5}`, owned-by-truck `{}`. The `reset_all_prompts` before each prompt is what
stops the leak; `state["geometric_prompt"]` survives `set_text_prompt` otherwise.
2. `exemplars` + `exemplar_class_name` through `labeling.label_image` → `preview.py` →
`POST /api/batches/{id}/preview` → verify: an unknown class name falls back to plain text
rather than attaching the boxes to whichever class is first. `[DONE]` — 17 shapes for
both text-only and `exemplar_class_name: "nonexistent"`.
3. `preview_frame` moved out of `autolabel.py` into `preview.py` → verify: `autolabel.py` is
back under the 400-line limit and the job path still imports. `[DONE]` — 261 and 151
lines; container starts and registers the `autolabel` handler.
4. Editable class prompt in the modal, saved to `project_classes.prompt` (REQ-171) →
verify: a PATCH round-trips and the Projects page shows the new text. `[DONE]` — class 2
`box → cardboard box → box` via the existing `PATCH /api/projects/{id}`; no new endpoint.
5. `ExemplarCanvas.jsx` drag/shift-drag/undo/clear with 250 ms debounced re-run, and the
modal split into `PreviewShapes.jsx` + `ClassPromptPanel.jsx` to stay under 400 lines →
verify: `npm run build` clean, every file under the limit. `[DONE]` — 398 / 126 / 137 /
64 lines, build green, both containers redeployed.
**Deliberately not built:** exemplars in the batch job. SAM3's geometric prompts pool
features from the current image, so a box drawn on frame 1 asks about whatever sits at those
coordinates on frame 400. The batch job stays text-only; the exemplars exist to find the text
that works.
## Task — Exemplar-driven labeling in the review editor (REQ-173, REQ-174) `[DONE]`
1. `backend/exemplar.py` — pool → one SAM3 pass (class prompt + boxes) → rewrite that class
on that frame → verify: on frame 55446 (batch 426, `sack`), one positive drawn from an
existing box gives 52 class-0 shapes, exactly 1 of them `manual` with the drawn geometry,
and the frame's class-1 shapes are untouched. `[DONE]` — verified; warm pass 0.4 s, first
pass 7.6 s (model load).
2. Negative exemplars delete what they cover (REQ-174) → verify: shift-drag over one of the
detections and no `auto` shape overlapping it by ≥ 0.3 IoU comes back, while the drawn
positive survives. `[DONE]` — max IoU with the negative afterwards 0.078, manual shape
still present.
3. GPU-busy fallback → verify: hold `jobs.gpu_lock`, drag, and the drawn shape is still
stored with `redetected: false` and a legible message. `[DONE]` — "Saved your shape — the
GPU is busy with a background job…", 58 shapes vs 57 before, no exception.
4. `POST /api/frames/{id}/exemplar-label` + `AnnotationCanvas` drag/shift-drag with the pool
drawn as dashed ghosts, 400 ms debounce, undo/clear, and the busy message under the canvas
→ verify: `vite build` clean and every touched file under 400 lines. `[DONE]` — build
green; `exemplar.py` 211, `api/review.py` 141, canvas 287, `useExemplarPool.js` 81. The
pool logic went into that hook rather than into `ReviewPage.jsx`, which was already over
the limit before this task (620 lines) and ends it at 628.
## Task — Filter panel and preview for exemplar runs (REQ-175) `[DONE]`
1. `exemplar.label(..., apply=False)` — dry run by default, returning `shapes` instead of
writing → verify: two previews in a row leave the row count untouched. `[DONE]` — frame
55446 stayed at 57 rows across a default preview (52 shapes) and a filtered one (20).
2. The four filters, applied in the batch job's order (area floor → NMS → cap) → verify:
each one visibly bites on a dense frame. `[DONE]` — from 52 shapes: NMS 0.05 → 32,
min box 0.05 → 1, cap 5 → 5, confidence 0.9 → 15.
3. `apply: true` writes exactly what was previewed → verify: the applied frame matches the
preview count and leaves other classes alone. `[DONE]` — 20 previewed, 20 class-0 shapes
stored (1 of them the drawn `manual` box), the frame's 2 class-1 shapes untouched.
4. `ExemplarFilterPanel.jsx` floating in the canvas corner, sliders re-previewing on 250 ms,
Apply/Discard/Undo/Reset, Enter and Esc bound → verify: `vite build` clean, files under
the limit. `[DONE]` — panel 108, hook 125, canvas 314 lines; build green; both containers
rebuilt and the live endpoint returns `applied: false` for a drag.
5. The class under review hides while its preview is up → verify: a negative exemplar's
effect is visible instead of being masked by the stored box underneath it. `[DONE]` —
frame 55446: 51 detections with one positive, 50 with a negative added; before this the
removed box stayed on screen at 35% opacity and the run looked inert.
**Deliberately not built:** saving the filter values. They describe one frame's run, and the
auto-annotate modal already owns the batch-wide numbers — sharing them would let a tweak made
while reviewing one frame silently change what the next batch job does.
**Deliberately not built:** persisting the pool. It is a prompt about *this* image, so it
dies with the frame, exactly as in REQ-172. What persists is the annotations it produced.
## Task 32 — WebRTC preview for the live counting page (REQ-176, REQ-177) `[DONE]`
The live view cost far more than it should: the backend re-encoded every annotated frame to
JPEG and pushed it over MJPEG, on top of decoding the camera. The camera already reaches the
browser cheaply over WebRTC, so the frames stop travelling through this app entirely.
1. A live source must be a WHEP URL; the RTSP leg is derived → verify: **[DONE]**
`POST .../live-count/start` with `rtsp://192.168.192.96:8554/cam` →
`400 "A live source must be a WebRTC (WHEP) URL…"`; with
`http://192.168.192.96:8889/cam` → `200`, `source: "rtsp://192.168.192.96:8554/cam"`,
`whep_url: "http://192.168.192.96:8889/cam/whep"`, `preview: "webrtc"`.
2. The AI counts from that stream → verify: **[DONE]** 185 frames in 49 s off the live
camera, `error: ""`. That rate is the link's, not the model's — see below.
3. No JPEG is encoded for a WebRTC session → verify: **[DONE]** `GET /api/live-count/stream`
downloaded 0 bytes during a running WebRTC session, and now answers `409`.
4. The overlay feed carries what the model saw, and tracks the line live → verify:
**[DONE]** `GET /api/live-count/overlay` returned 27 boxes with ids and confidences;
after `PATCH /api/live-count/line {"line_y":300}` the feed reported `line.y: 300`.
5. The 400-line limit holds → verify: **[DONE]** `live_count.py` was already 467 lines, so
the transport layer went to `live_source.py` (120) and the MJPEG overlay to
`live_render.py` (60), leaving it at 393. On the frontend the preview moved to
`LiveVideoPanel.jsx` and the slider table to `liveCountFields.js`, leaving
`LiveCountPage.jsx` at 383. `npm run build` passes.
**Not verified here:** the WHEP handshake in a real browser. The endpoint was confirmed live
(`POST http://192.168.192.96:8889/cam/whep` answers, rejecting a deliberately malformed SDP
with `400`), but the negotiation itself needs a browser, not curl.
### Where the live FPS actually goes — measured, 2026-08-19
The live session runs at 4-6 fps and it is not the model. Measured in the backend container
against `rtsp://192.168.192.96:8554/cam`:
| Stage | Rate |
|---|---|
| ByteTrack + YOLO inference | **205 fps** |
| `cv2.resize` to 1280x720 | 5348 fps |
| Decode from RTSP | **6.4 fps** |
The camera is 704x576 HEVC at 350 kbit/s — nothing about it is expensive. The link is: the
route to the streaming server is a ZeroTier VPN measuring **15% packet loss** and a 41-104 ms
round trip. The comment in `live_source.py` claiming the cost was "decoding 1080p on the CPU"
was simply wrong and has been corrected; so has the hint on the page.
Transport was changed to UDP and changed back, because the measurement contradicts the
theory. Through the **ffmpeg CLI**, UDP wins as expected — 16 fps at 1.00x realtime against
TCP's 6.8 fps at 0.52x. Through **OpenCV** it loses: tcp 6.4 fps, udp+socket buffer 4.4, bare
udp 2.4, and a live session on UDP showed 18-second stalls waiting for a keyframe. OpenCV
drops what it cannot reassemble instead of showing it, so the loss lands as missing frames.
`RTSP_TRANSPORT` is left as an env override, defaulting to `tcp`.
**Not fixable in this repo.** Inference has ~50x the headroom the link delivers, so nothing
in the app is worth optimising. The lever is where the counter runs: next to MediaMTX it
would count at the camera's full rate. Worth checking whether the ZeroTier path is relayed
rather than direct (`zerotier-cli peers` — a `RELAY` row explains both the loss and the RTT).
**Deliberately not built:** an aiortc/WHEP client in the backend. It would be "WebRTC only"
end to end, but the decode cost is identical to RTSP and it adds ICE and keyframe-loss
failure modes to the counting path. The saving was always on the browser side.
## Task — Archive upload and date folders (REQ-178) `[DONE]`
1. Upload + date-folder API/UI → verify: **[DONE]** mkdir 200 (new) / 409 (dup) / 400 (bad
format) / 400 (traversal `../evil`); upload 200 / 409 (dup) / 400 (bad ext) / 400 (missing
folder); 600MB file → 200; `GET .../library/2099-01-01` lists the uploaded file, file
visible from host WSL, zero `*.part` residue, test folder removed after verification;
`vite build` passes (66 modules). Browser click-test of copy buttons/upload dialog is NOT
automated — manual click-test pending.
## Task — Empty date folder visible in cycle list (REQ-178) `[DONE]`
1. `archive_index.cycles()` must return a bucket for a date folder with no videos →
verify: **[DONE]** `POST .../library/dates?date=2099-01-01` → 200; then
`GET .../archive/cycles` contains `2099-01-01` with `video_count: 0` (and the
pre-existing empty folder `2026-08-31`, previously invisible, also appears);
`GET .../archive/cycles/2099-01-01` returns `videos: []`; the same JSON arrives
through the dev proxy on 5173; folder removed after verification → cycle gone from
the list again; `uvx ruff check` output identical to HEAD (13 pre-existing, 0 new);
`vite build` passes. Browser click-test is NOT automated — manual click-test pending.
Docker backend rebuilt after the fix: `:9010/.../archive/cycles` now returns the
empty folders `2026-08-29` and `2026-08-31` with `video_count: 0`.
## Task — Copy-path buttons (REQ-179) `[DONE]`
1. Copy-path buttons → verify: **[DONE]** project payload carries
`video_root_linux`=`/home/araaraenjoyer/dbs_project/reTraining/data/archive` and
`video_root_windows`=`\\wsl.localhost\Ubuntu\home\...\data\archive`. Browser click-test of
the buttons is NOT automated — manual click-test pending.
## Task — Infra: rw archive mount + nginx body size (REQ-178) `[DONE]`
1. rw mount + nginx body size → verify: **[DONE]** compose archive mount `RW=true` (no `RO`);
nginx `client_max_body_size 20g` verified (600MB upload → 200); `/api/health` 200, `/docs`
200, frontend `http://localhost:9000` 200 after rebuild.
## Task — Frame-scoped per-class clear in review editor (REQ-180) `[DONE]`
1. `×` button on each sidebar class row clears that class's shapes on the **current frame
only**, no confirmation, batch-wide trash (REQ-046) unchanged →
verify: **[DONE]** `ReviewPage.clearClassInFrame` filters the current frame's annotations
and posts only those ids to `POST /annotations/bulk-delete`; round-trip on `:9010` —
created 2 shapes on frame 23827, bulk-delete returned `{"deleted":2}`, frame back to
`annotations: []`, DB clean after test; optimistic update + rollback wired the same way as
`removeMarked`; `vite build` passes (68 modules), rebuilt image on `:9000` serves the new
bundle (`index-wBeBdXus.js`). Browser click-test of the `×` button is NOT automated —
manual click-test pending.
## Task — Per-class auto-annotate params (REQ-181) `[DONE]`
1. `class_params` through preview + job on both modals → verify: **[DONE]** ruff on the 5
changed backend files vs HEAD: +17, all `UP006`/`UP035`/`UP045` (the files' existing
style), 0 new real findings; on rebuilt backend `:9010` — `/preview` baseline (absent and
`null` `class_params`) → 18 shapes, `class_params: {sack: {threshold: 0.99}}` → 0 shapes,
`threshold: "high"` → 422 `float_parsing`; `/autolabel` accepted
`{class_params: {sack: {threshold: 0.9, min_box_frac: 0.01}}}`, job 106's stored params
carry it, `iou_threshold: "x"` → 422, job cancelled and test data reset
(`reset-auto-annotations`, batch 21 back to 0 annotations); omitted `class_params` takes
the unchanged global path; `vite build` passes, `:9000` serves the new bundle (marker
strings present). Browser click-test of the override tables is NOT automated — manual
click-test pending.
## Task — Per-class accumulating exemplar preview (REQ-172, REQ-182) `[DONE]`
1. Exemplar pools keyed by class and a per-class accumulating preview in the auto-annotate
modal — switching the active class swaps pools and a frame change clears them all
(REQ-172); a redraw re-runs only the touched class against its own pool, the canvas merges
exemplared classes' conditioned results over the last full-set detections of the other
selected classes (REQ-182, as amended — superseded by the merged-preview entry below),
clearing the last example returns that class to its full-set detections, no examples at
all → the unchanged full-set request (REQ-182) → verify: **[DONE]** `cd frontend && npm run
build` passes (✓ 536 ms, 69 modules); `wc -l` caps hold —
`frontend/src/components/AutoAnnotateModal.jsx` (400) and
`frontend/src/hooks/useExemplarPools.js` (100) both ≤ 400; rebuilt Docker frontend
bundle carries the `pools` marker (`grep -c pools dist/assets/*.js` → 1) and serves
HTTP 200 on :9000; browser drag/undo/clear across two classes is NOT automated —
manual click-test pending.
## Task — Per-class hide toggle in review editor (REQ-183) `[DONE]`
1. Eye button on each sidebar class row — a hidden class's shapes leave the canvas and every
canvas selection path but stay in the "Shapes on this frame" list, dimmed, each with its
own restore eye (amended by REQ-185; the old leave-the-list and "not deletable" semantics
are superseded — see the REQ-185 entry below), while the row keeps its real per-frame count
and eye state; the choice is session-only (survives frame changes, resets
when the review page is left, stored data untouched) → verify: **[DONE]** `cd frontend &&
npm run build` passes (✓ 535 ms, 69 modules); `wc -l` caps hold —
`frontend/src/components/Icons.jsx` (172) and
`frontend/src/components/ReviewSidebar.jsx` (145) ≤ 400,
`frontend/src/pages/ReviewPage.jsx` (725) exempt (pre-existing over the 400 cap, not
split by this task); rebuilt Docker frontend serves the new bundle (HTTP 200 on :9000,
`aria-pressed` present in the shipped JS); browser click-test of the eye toggle is NOT
automated — manual click-test pending.
## Task — Container class flag (REQ-184, REQ-031) `[DONE]`
1. `project_classes.container` (`INTEGER NOT NULL DEFAULT 0`), the `containers` patch on
`PATCH /api/projects/{id}` beside the prompts patch, and the Container checkbox in both
auto-annotate modals' overrides table → verify: **[DONE]** migration
`PRAGMA table_info(project_classes)` lists `container`, a second `db.migrate()` run is
a no-op; API round-trip on project 5 sets `truck → container 1` and clears it back
(`curl -X PATCH :9010/api/projects/5 -d '{"containers":{"2":true}}'` →
`[(0,'sack',0),(1,'box',0),(2,'truck',1)]`, revert → all 0);
`cd frontend && npm run build` passes (✓ 536 ms);
`wc -l frontend/src/components/ClassParamsTable.jsx` = 115 ≤ 400; both modals stay at
400 / 347.
2. Cross-class NMS reads the stored flag — preview and batch job build the same container-id
set → verify: **[DONE]** scripted NMS check 6/6 (`uv run python`, pasted in
`t3-report.md`, re-run at T5): containment ≥ 90 % keeps both boxes for the container
class only (one direction), no blanket exemption below 0.9 (IoU 0.802 / containment
0.890 → dropped), `iou <= 0` never drops, within-class pass unchanged, per-class IoU
override consulted on cross pairs; `uv run ruff check backend/` → 409 errors, all
+13 being pre-existing categories (UP006/UP035/UP045) on the new signature lines.
The REQ-031 amendment rides on this entry: cross-class greedy NMS with the per-class IoU
override and the containment carve-out is only real once the flag exists and both live sites
read it.
## Task — Per-shape hide + dimmed shape list (REQ-183 amended, REQ-185) `[DONE]`
1. `H` hides/shows the selected or marked shapes; hidden shapes (by `H` or by their class)
stay listed dimmed in "Shapes on this frame", each dimmed row with a restore eye that
un-hides or overrides; shapes created into a hidden class (draw, assist, copy)
auto-override; the class-eye toggle clears that class's overrides; the old purge effect
and its "not deletable" invariant are gone → verify: **[DONE]** `cd frontend && npm run
build` passes (✓ 545 ms, 69 modules); `wc -l` — `frontend/src/components/ReviewSidebar.jsx`
≤ 400 (145), `frontend/src/components/ShortcutsPanel.jsx` ≤ 400 (69),
`frontend/src/pages/ReviewPage.jsx` exempt (725, pre-existing over the 400 cap, not
split); `grep` proofs that `'h', 'H'` is in `isShortcutKey` (`ReviewPage.jsx:407`) and
that the removed purge effect (`while a class is hidden`) has zero hits left in
`ReviewPage.jsx`; browser `H`/eye click-test is NOT automated — manual click-test pending.
## Task — Copy effective auto-annotate params (REQ-186) `[DONE]`
1. Copy button in the shared per-class table → clipboard text, one line per class,
effective values, container true/false, both modals → verify: **[DONE]** `cd frontend
&& npm run build` passes (✓ 528 ms); `wc -l` — `frontend/src/components/ClassParamsTable.jsx`
≤ 400 (152), `frontend/src/clipboard.js` (36); `git diff --stat` = exactly 3 code
files (clipboard.js, ArchiveControls.jsx, ClassParamsTable.jsx) with ArchiveControls
extraction-only; format walkthrough byte-exact vs REQ-186 sample lines (empty→global,
invalid→global, float-noise case 0.30000000000000004 → `0.3`); browser clipboard
click-test NOT automated — manual click-test pending.
## Task — Merged preview: exemplar draw keeps other classes (REQ-182 amended) `[DONE]`
1. Non-exemplared classes keep last full-set detections when an exemplar is drawn;
exemplared class shows conditioned results replacing its own; Run Preview
refreshes full-set first (exemplars: []) then exemplared; draw never full-set
re-runs → verify: **[DONE]** `cd frontend && npm run build` passes (✓ 523 ms);
`wc -l frontend/src/components/AutoAnnotateModal.jsx` = 400 (at cap; blank lines
trimmed to fit, adjudicated by review); `git diff --stat` = 1 code file; grep
proof full-set call has `exemplars: []` (`AutoAnnotateModal.jsx:118`); 7-case
walkthrough in t11-report.md; browser test (draw example after Run Preview, other
classes persist) NOT automated — manual click-test pending.
## Task — DATA_DIR-relative model weight paths (REQ-187) `[DONE]`
1. `resolve_data_path`/`rel_data_path` in `config.py`; all file-opening reads wrapped
(preview, autolabel, training, model download, live count, `projects.get`,
`training_start_point` hack replaced); writes store relative → verify: **[DONE]**
harness over all 5 legacy rows prints `isfile=True`; import smoke exit 0;
`git diff --stat` = 7 backend files; reviewer APPROVED (3 latent Minors accepted:
raw path in `list_models` payload, `secondary_model_path` outside REQ-187 scope,
non-str TypeError unreachable); no DB rewrite needed — resolver covers legacy rows.
## Task — per-class max box fraction (REQ-188) `[DONE]`
1. `max_box_frac` (default `1.0` = off) mirrored 1:1 over every `min_box_frac` site:
`labeling.label_image` gains `max_box_frac`/`max_box_fracs` with a ceiling block right
after the floor (before dedup), guard `frac >= 1 or frac <= 0` → keep; `preview.py` and
`autolabel.py` gain the inline `xyxyn` gate (`xb and xb < 1`) and the `mx_list` per-class
build (fallback `1.0`) passed as `max_box_fracs=`; `_parse_class_params` accepts the key;
`exemplar.label` filters inline (it does not delegate to `label_image`) with
`0 < max_box_frac < 1` between floor and NMS; `ExemplarLabelRequest` carries the explicit
review-filter field through to `exemplar_store.label`. Frontend: `MaxBox` added to
`ClassParamsTable` KEYS (so `buildClassParams` sends it automatically), `maxbox <v>`
inserted in the copy line per amended REQ-186, `max_box_frac: 1.0` in both modals'
globals, `Max box size` slider in `ExemplarFilterPanel` + `FILTER_DEFAULTS` —
verify: **[DONE]** import smoke exit 0; `npm run build` green (520 ms); behavioral harness
over `label_image` prints the 8 cases (off/0/1/per-class/list-fallback/floor+ceiling)
with expected drops; node eval proves `buildClassParams` emits `max_box_frac`;
`git diff --stat` = 5 backend + 5 frontend files + docs.
## Task — stacked per-class override rows (REQ-181 layout) `[DONE]`
1. The overrides surface drops the wide table for one block per class: name line with the
**Container** checkbox (REQ-184) at the right, then a wrapping
`repeat(auto-fit, minmax(118px, 1fr))` grid of `Conf`/`IoU`/`MinBox`/`MaxBox` inputs, each
labelled in its own colour. Cause of the old cramping: `width: 100%` + the new MaxBox
column let the Class column swallow the slack, stranding the checkbox far right.
Behavior untouched — `KEYS`, `buildClassParams`, `compose()` copy line, `toggleContainer`
revert path, `title`/`aria-label` all identical → verify: **[DONE]** reviewer diff-proved
the functional hunk is render-only (SPEC ✅, 4 Minor all fixed: two stale ui-spec
"table/columns" lines, `label` margin-bottom, dead class hook; plus long-name ellipsis +
`flexShrink: 0` guard); `npm run build` green; grid math checked at 375/768/920/1040 px;
contrast ≥6.6:1 on all label colors; focus ring left to the global `input:focus`
(accent border + glow) — no CSS added; **browser eyeball still owed by the user**
(375 px 2-col wrap, long class names).
## Task — Copy/Paste YAML for per-class overrides (REQ-186 amended, REQ-190) `[DONE]`
> Code, parser and build verified below; the **paste round trip in a real browser is still
> owed** (neither a keyboard paste nor a `Ctrl+V` prefill can be exercised from the shell).
1. **Icons** `frontend/src/components/Icons.jsx` — `CopyIcon` and `ClipboardPasteIcon` added
in Lucide geometry, inheriting `currentColor`; the shared `Icon` wrapper already sets
`aria-hidden` and `focusable="false"`.
2. **Copy emits YAML** `frontend/src/components/ClassParamsTable.jsx` — `compose()` writes a
`#` comment line plus one block per selected class, fields named from the row labels
(`conf`/`iou`/`minbox`/`maxbox`/`container`) via `FIELD_BY_LABEL` so the wire format cannot
drift from what is on screen. A name is quoted only when `PLAIN_NAME` fails, i.e. only when
YAML would misread it. Effective-value resolution is unchanged (empty input still copies the
global), so the block is a whole configuration rather than a diff.
3. **Strict reader** `parseClassYaml()` in the same file, exported so it is testable without a
DOM. Deliberately **not** `js-yaml`: a subset reader for exactly what Copy emits is ~35
lines and needs no dependency. Rejects (with the line number) a list item, a nested map, a
top-level scalar, an unknown key, a malformed number, a value outside 0–1 and an indented key
before any class. Throwing happens before any state write, so a bad paste is atomic —
nothing is half-applied.
→ verify: **[DONE]** parsed the real parser through `esbuild --bundle --platform=node`: a
two-class document with a quoted `"weird: name"` key round-trips to the internal key names
(`threshold`, `iou_threshold`, `min_box_frac`, `max_box_frac`, `container`), and the malformed
cases each throw the expected `line N: …` message (an empty document returns `{}`, which
paste turns into "No classes in the pasted YAML"). A self-review pass caught a real hole the
first harness run missed: `parseFloat('0.5abc')` is `0.5`, so a typo'd number was being
accepted silently — exactly the failure the strict design exists to prevent. Fixed by
matching the text against `NUMBER` before `parseFloat`. Same harness after the fix:
`0.5abc`, `Infinity`, `NaN`, `0.5z`, `" 0.5 "` all throw; `1e-2`, tabs, CRLF, an inline
`# comment` and a quoted name containing ` #` all parse. Repeated class blocks merge.
4. **Paste opens a review dialog** `frontend/src/components/PasteYamlDialog.jsx` (new, 83 lines)
— a textarea plus a live review line, nothing written until Apply. Why a dialog and not a
direct read: `navigator.clipboard.readText()` only exists in a secure context, so on plain
http over a LAN address it is *absent*, not merely refused — while a keyboard paste into a
focused textarea is an ordinary user gesture and is not gated at all. The dialog therefore
prefills when the browser allows and otherwise says so and waits for Ctrl+V.
`planPaste(text, classNames, containers)` (exported from `ClassParamsTable.jsx`) produces
the review line — matched classes, `container` flips, ignored classes, or the parse error —
and is what disables Apply. `applyPasted()` in the table does the writing: only classes
selected in this modal, per-field merge, and `container` diffs through `applyContainers()`,
which folds the per-toggle logic into one `PATCH /projects/{id} {containers: {classId: bool}}`
(`backend/projects.py:232` iterates the dict) and reverts all of them together on rejection,
the same functional-revert trick the checkbox already used.
→ verify: **[DONE]** `planPaste` and `parseClassYaml` exercised through
`esbuild --bundle --platform=node`, both exported from the real module: a three-block document
returns `2 classes: sack, box · container: sack → true · ignored: trailer`; a flag already in
the wanted state is not reported as a flip (`1 class: box`); empty, `conf: 40`, `conf: 0.5abc`,
`- sack`, and a document naming only an unselected class each return the error that keeps
Apply disabled. The merge was checked too: pasting `conf`+`iou` over an existing
`box: {threshold: 0.55, iou_threshold: 0.8}` leaves `iou` at `0.8` and sends no phantom
zeros — `buildClassParams` strips the empty entries. First harness run showed `sack` written
empty, which turned out to be a bug in the harness itself (`{ key }` destructured out of
`['threshold']`), not in the component; the re-run confirms the merge. `Esc` closes and
returns focus to the trigger; Enter is deliberately not bound, YAML needs newlines; a nested
`position: fixed` overlay is safe here because nothing above `ClassParamsTable` has a
`transform`/`filter`/`backdrop-filter` to become its containing block (checked).
**Owed, user browser**: dialog opens focused, Ctrl+V fills it over plain http on a LAN
address, live line names the classes, Apply fills the four inputs and ticks Container, status
line reports the ignored class; `conf: 40` shows `line 2: …` with Apply disabled; Container
survives a modal reopen; `Esc` closes and focus is back on the button.
5. **Button design** `frontend/src/app.css` — `.class-params-tools` right-aligned wrapping row,
`.btn` at `0.76rem` with `cursor: pointer` and a `progress` cursor while disabled (the old
`btn-ghost` text button had neither), `.class-params-note` status line that turns `--danger`
when the outcome failed. It does **not** use the app's `.hint` class: that is `--text-faint`,
measured 2.68:1 on the modal panel, and the status line is small text that needs 4.5:1 —
`--text-muted` measures 6.98:1 and `--danger` 4.71:1. Hover/focus-visible come from the
shared `.btn` (`theme.css:114`, `:108`).
→ verify: **[DONE]** `npm run build` green; contrast ratios computed from the tokens;
`ClassParamsTable.jsx` 300 lines and `PasteYamlDialog.jsx` 83, both inside the 400-line
rule.
## Task — Configurable Track Forward + track-to-end (REQ-189) `[DONE]`
> Status caveat: code, build and docs are verified below; the **end-to-end GPU run in the
> user's browser is still owed**, so the feature is not proven on real shapes yet.
1. **Backend** `backend/batches.py:136` `frames()` — the skip check needs *which* classes a
frame holds, and the frontend only ever received `annotation_count`. Added
`GROUP_CONCAT(DISTINCT a.class_id)` to the existing subquery (same row, no extra round trip,
no migration) and parsed it to `[int]`, `[]` when NULL.
→ verify: **[DONE]** `uv run` against the real DB: batch 19 returns 507 frames, the five
frames of the earlier manual track run (`22829`–`22833`) come back `class_ids: [2]`, frames
with none return `[]`. `EXPLAIN QUERY PLAN` shows both subqueries seek
`idx_annotations_frame` (no per-frame scan). Other consumers (`autolabel.py:95`,
`dataset.py:480`, `preview.py:35`) only read dict keys — unaffected.
2. **Cancel plumbing** `frontend/src/api.js` — `request` takes `signal`, `assist(frameId, body,
{ signal })` forwards it. No other caller touched.
3. **UI** `frontend/src/pages/ReviewPage.jsx` — `trackForward(toEnd)` now takes an explicit
count: N input (1–100, clamped and **persisted on blur** — an effect would save the
half-typed value, so typing "15" survives), *Track Forward [T]* with a dynamic tooltip,
*→ End* with a confirm showing the real frame count, `AbortController` per run so the
"asking SAM3…" indicator becomes a **Cancel** button, skip-on-existing-class, and a banner
reading `Tracked N of M — skipped K already had <class> — <failures>`. The banner shows when
something was skipped *or* failed — a fully clean run stays silent, as before. A successful
assist also unions its `class_id` into the local frame's `class_ids`, so a second run in the
same session skips its own output instead of writing duplicates.
4. **Filmstrip** `frontend/src/components/Filmstrip.jsx` — optional `markedIds` + `markLabel`
props; `frontend/src/app.css` `.thumb .skip-dot` (`--warn` + 2 px dark ring for contrast on
light thumbs, top-left so it never collides with the count badge), `.track-count` for the
input, and `flex-wrap` on `.frame-bar` so the bar still fits 375 px. Skipped frames also name
the class in the thumb tooltip.
→ verify: **[DONE]** `npm run build` green; clamp covers `"abc"`, `0`, `999`, `-3`, `""`,
`Infinity` and non-numeric `localStorage` garbage (`Number.isFinite` guard → 5), and the
slice bound is computed from `Number(trackFrames)` so a string can never concatenate into the
slice end; `slice(index+1, index+1+N)` and `slice(index+1)` both bounded; cancel lands as
`cancelled — frame n may still be saved by the server` and keeps the shapes already written;
skip check reads the `class_ids` array, not `annotation_count` — *that rule was later replaced
by the post-run overlap test in the next task*; marks cleared at the start of
every run; `Shift+T` deliberately **not** bound (a fat-fingered 500-frame GPU run is not worth
a keystroke). Reviewer pass done — 4 findings fixed. **Still owed, user browser + GPU**: set
N=3 → 3 shapes, reload → N remembered; to-end over the already-reviewed part of batch 19 →
banner reports the skips and dots appear on those thumbs; Cancel mid-run stops it.
## Task — Track 5 Frames works at all (REQ-189) `[DONE]`
1. `ReviewPage.jsx:trackForward` read `geometry.coordinates` — a key the backend never emits
(`backend/review.py:35-43` writes `points`) and had no bbox branch, so the action no-op'd
silently for **every** shape; `docs/audit-2026-08-07.md:103` had flagged it as dead.
Now: bbox `points` used as-is, polygon reduced to its bounding box, each guard speaks
(no selection / shape gone / last frame / unusable geometry), one refusing frame no longer
aborts the rest (`Tracked N of M` + per-frame failures), a press during a run is ignored
(double-press used to double-write the same five frames once the path went live).
REQ-189 added; non-goal `requirements.md:22` amended to keep exemplar propagation out while
admitting this one-shot hand-off → verify: **[DONE]** reviewer SPEC ✅ (geometry order +
normalization traced against `validate`/`bbox`, loop bounds `slice(index+1, index+6)`,
frame numbering matches the frame bar, dep array sufficient); `npm run build` green;
doc truth fixed (`ui-spec:991` count wording, stale known-limitation line removed, audit
snapshot left as history); **SAM3 actually finding the object in the next 5 frames needs
the user's GPU run** — select a shape, press `T`, expect a shape on each of the next 5.
## Task — Track Forward skips on overlap, not on class presence (REQ-189 amended) `[DONE]`
- The old rule skipped a frame *before* SAM3 ran when it already held a shape of the tracked
class, which contradicted REQ-189's own "each frame is an independent run": a second truck in
the same frame could never be annotated. Amended to a duplicate test **after** the run.
- `backend/review.py:intersects()` — any intersection between two normalized boxes (shared area
> 0; edge touch is *not* an overlap), deliberately not IoU. `assist()` gained
`dedupe: bool = False`: once the GPU lock is released and right before `add()`, an existing
shape of the same class on that frame whose box intersects the result makes it answer
`{"skipped": True}` and write nothing — no create-then-delete, no transient row.
`backend/api/review.py:AssistRequest.dedupe` carries the flag; `False` keeps the single-shot
box-assist path byte-identical.
- `frontend/src/pages/ReviewPage.jsx:trackForward` — pre-skip on `class_ids` and the local
`class_ids` union are gone, the call sends `dedupe: true`, and `{skipped: true}` lands in the
banner as `skipped K overlapping <class>` plus the filmstrip dot `overlaps an existing <class>`.
- `backend/batches.py:frames()` no longer ships `class_ids` (it was added for the pre-skip and
died with it); `annotation_count` stays.
- Docs: `requirements.md` REQ-189 skip sentence rewritten, `ui-spec.md` banner + thumb tooltip.
- Cost accepted: a long run now pays a GPU call on **every** frame ahead, including frames that
end up skipped, where the old pre-skip made those free. Only worth layering a cheap seed-box
pre-skip back if that hurts.
→ verify: **[DONE]** `uv run python` harness (mocked engine + temp PNG) — 6 `intersects`
cases (edge-touch false, partial/contained/corner true, disjoint false) and 5 `assist()` paths:
dedupe+overlap → `{"skipped": True}` with `add()` never called; other-class box → annotated;
empty frame → annotated; `dedupe=False` → annotated; polygon annotation → skipped.
`batches.frames()` run against the live DB (3737 rows, `class_ids` gone, `annotation_count`
kept); `npm run build` green; `git diff --check` clean. **Still worth a GPU run:** track into a
frame already holding an overlapping truck → banner `skipped K overlapping truck`, dot, count
unchanged, while a frame holding a truck elsewhere gets its own second box.
## Known open points
- *Not closed by any task, by choice:* **any rebuild kills the running job.** Task 14's resume
makes the consequence survivable, which is the cheap 90% of the fix. Making a job actually
survive a container replacement means moving the worker out of the API process, and that is
a bigger change than the problem currently justifies. Schedule long runs around deploys.
- **Any rebuild kills the running job.** `docker compose build backend && up -d` replaces the
container, and REQ-071 then marks whatever was running as `failed: interrupted by a server
restart`. Nothing is corrupted, but long runs and deploys do not mix.
- *Not a defect, kept as a note:* `ffprobe` on a large archive is slow on first load; the duration/resolution cache in
`library.py` is what keeps the Library page usable.