Files
reTraining/docs/design.md
T

487 lines
30 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Design
Serves `./requirements.md`. Each section names the REQs it fulfils.
## Overview
```
Browser (React + Vite, served by nginx)
│ /api → proxy
▼
FastAPI ──► jobs (1 worker thread, 1 GPU)
│ ├─ extract : ffmpeg (CPU)
│ ├─ autolabel: SAM3 (GPU)
│ ├─ merge : copy + write labels (CPU)
│ └─ train : Ultralytics + eval (GPU)
├──► SQLite (metadata & status)
└──► data/ (frames, master dataset, weights)
```
Storage split rule: **SQLite holds metadata and status; the disk holds pixels, final labels,
and weights.** The master dataset must stay useful even if the database is lost
(REQ-006, REQ-054). Model weight paths stored in SQLite (`projects.base_model_path`,
`model_versions.weights_path`, `model_versions.parent_model_path`) are written **relative
to `data/`** and resolved against `DATA_DIR` on every read (REQ-187), so a moved data
folder keeps every trained model reachable; legacy absolute rows from older installs go
through the same resolver.
## Disk layout (REQ-006)
```
data/ # Docker volume
app.db # SQLite (WAL)
projects/<slug>/
base/model.pt # the project's active base model (REQ-003)
dataset/ # MASTER, accumulative (REQ-050…052)
images/{train,val}/…jpg
labels/{train,val}/…txt
data.yaml
batches/<batch-id>/
frames/000001.jpg … # extraction output (REQ-022)
models/<n>/
best.pt
metrics.json # base vs new metrics (REQ-063)
runs/ # Ultralytics run directory
```
Master dataset filenames: `<batch-id>__<frame number>.jpg` — unique across batches and
self-documenting about where each image came from. The user's video archive is written only
by the user-initiated upload and date-folder endpoints (REQ-178); no other code path writes
into it (REQ-074).
## SQLite schema
Created by an idempotent migration in `backend/db.py` at startup.
```sql
projects(
id, slug UNIQUE, name, label_type CHECK(bbox|polygon),
base_model_path, base_model_kind CHECK(uploaded|pretrained|trained),
video_root, val_every DEFAULT 5, created_at)
project_classes(
id, project_id → projects, class_id INT, name, prompt,
container INTEGER NOT NULL DEFAULT 0, -- container class flag (REQ-184)
UNIQUE(project_id, class_id)) -- class_id = the YOLO class index (REQ-003/005)
batches(
id, project_id → projects, video_path, date_label, batch_label,
start_sec REAL, end_sec REAL, fps REAL,
status CHECK(extracting|extracted|labeling|reviewing|approved|merged|failed),
frame_count INT, created_at, merged_at)
frames(
id, batch_id → batches, idx INT, filename, width INT, height INT,
review_status CHECK(pending|approved|rejected) DEFAULT 'pending',
UNIQUE(batch_id, idx))
annotations(
id, frame_id → frames, class_id INT,
geometry TEXT, -- JSON; see "Geometry format"
score REAL, source CHECK(auto|manual), created_at)
dataset_items( -- master dataset membership (REQ-052)
id, project_id → projects, frame_id → frames UNIQUE,
split CHECK(train|val), image_rel, label_rel, added_at)
model_versions(
id, project_id → projects, version INT, name TEXT, weights_path,
parent_model_path, metrics TEXT, base_metrics TEXT, created_at,
UNIQUE(project_id, version))
jobs( -- persistent (REQ-071)
id, project_id, batch_id, type CHECK(extract|autolabel|merge|train),
status CHECK(queued|running|done|failed|cancelled),
progress INT, total INT, message, error, log TEXT,
created_at, started_at, finished_at)
```
**The stable val split (REQ-052)** is enforced by `dataset_items`: an existing row never
changes its `split`. On merge, only frames without a row are assigned, using a per-project
round-robin counter (`val_every`) that continues from the previous count.
**Geometry format.** One JSON column covers both label types (REQ-002):
- `bbox` → `{"type":"bbox","points":[x0,y0,x1,y1]}`
- `polygon` → `{"type":"polygon","points":[[x,y], …]}`
Coordinates are stored **normalized 0–1** against the frame size, so neither the editor nor
the exporter needs to know the display size. SAM3 mask → polygon conversion is
`review.mask_to_polygons()`; for `bbox` projects the mask is only used to take its bounding
box.
## Backend modules
`app/` moves to `backend/`. Reuse existing code wherever possible:
| Module | Role | Status |
|---|---|---|
| `sam3_engine.py` | SAM3 singleton, `open_state`/`apply_prompts`/`segment_at` | reused, plus `release_engine()` for REQ-065 and the idle-unload watcher (REQ-192): a daemon thread, started on first `get_engine()`, that takes the non-blocking `gpu_lock` and releases after `SAM3_IDLE_UNLOAD_S` seconds without use |
| `labeling.py` | per-frame detection + cross-class greedy NMS with per-class IoU override and the container containment carve-out (REQ-031, REQ-184) | reused; the folder-walking half went with the old flow |
| `exporters.py` | ~~YOLO label writing~~ | **deleted** — `dataset.py` writes labels, `mask_to_polygons` moved to `review.py` |
| `sessions.py` | ~~exemplar/tap interaction~~ | **deleted** — see below |
| `jobs.py` | single-worker queue | extended: job types + persistence |
| `training.py` | Ultralytics fine-tune | changed: starts from the base model, args from `hardware.py` |
| `db.py` | connection + migration | **new** |
| `projects.py` | project CRUD, reads classes from a `.pt` | **new** |
| `library.py` | scans `<video_root>/<date>/<batch>` | **new** |
| `video.py` | `ffprobe`, Range streaming, `ffmpeg` extraction | **new** |
| `batches.py` | batch lifecycle | **new** |
| `review.py` | annotation CRUD, frame status, click-assist | **new** |
| `autolabel.py` | the SAM3 job over a whole batch | **new** |
| `preview.py` | one-frame preview for the auto-annotate modal (REQ-171,172) | **new** |
| `exemplar.py` | exemplar-driven labeling in the review editor (REQ-173,174) | **new** |
| `dataset.py` | merge into the master dataset, stable split | **new** |
| `evaluate.py` | validate base vs new model | **new** |
| `hardware.py` | VRAM detection → training defaults | **new** |
| `live_count.py` | the live counting session: capture → track → count | **new** |
| `live_source.py` | what a source is, how it opens, WHEP↔RTSP (REQ-176) | **new** |
| `live_render.py` | the MJPEG overlay, archive-file preview only (REQ-177) | **new** |
| `api/` | the FastAPI routes, one module per domain | **new** |
Removed: `uploads.py`, `static/index.html`, and the old flow's endpoints.
`sessions.py` was meant to be reused for click-assist, but it existed to hold GPU-resident
state for an interactive session — a whole eviction policy, an undo stack, and a per-session
annotation store, all of which the database and the stateless `review.assist` now cover.
Adapting 357 lines to do what 60 lines do was not worth it, so the module is gone. The one
thing it knew that mattered — SAM3 wants exemplar boxes as normalized centre-x, centre-y,
width, height — moved with it.
The 400-line file limit (see `../AGENTS.md`) applies to all of the above. It is why the
routes live in `backend/api/{projects,batches,review,models,jobs}.py` rather than in
`main.py`, which now only builds the app and owns startup. Route modules import their
domain module under an alias (`from backend import projects as project_store`) so the two
namespaces stay distinguishable.
## API contract
```
GET /api/health REQ-073
POST /api/sam3/release # free SAM3's VRAM now; 409 while the GPU is busy (REQ-192)
GET /api/projects REQ-001
POST /api/projects REQ-001,002,004,005
GET /api/projects/{id} # includes label_type_locked: bool (REQ-002), video_root_linux/video_root_windows (REQ-179)
DELETE /api/projects/{id}
PATCH /api/projects/{id} # { prompts: {classId: text}, containers: {classId: bool} }, val_every (REQ-005, REQ-184)
POST /api/projects/{id}/classes # add new class {name, prompt} (REQ-008)
DELETE /api/projects/{id}/classes/{class_id} # delete class, delete shapes, reindex classes (REQ-007)
POST /api/projects/{id}/base-model # upload .pt, read classes (REQ-003)
GET /api/projects/{id}/dataset # master dataset summary (REQ-053)
GET /api/projects/{id}/dataset/download # zip (REQ-054)
GET /api/projects/{id}/library # list of dates (REQ-011)
GET /api/projects/{id}/library/{date} # videos + duration/resolution (REQ-012)
POST /api/projects/{id}/library/dates?date= # create YYYY-MM-DD folder → {date}; 409 if it exists, 400 bad format (REQ-178)
POST /api/projects/{id}/library/upload?date= # multipart `file`, streamed to disk → {rel}; 409 dup, 400 bad ext/missing date (REQ-178)
GET /api/projects/{id}/video?rel=… # Range streaming (REQ-013)
POST /api/projects/{id}/batches # {rel, start_sec, end_sec, fps} → extract job
GET /api/batches/{id} # status + review progress (REQ-045)
GET /api/batches/{id}/frames # frames + statuses
POST /api/batches/{id}/autolabel # {threshold, class_params} → job (REQ-030,032,034,181)
POST /api/batches/{id}/preview # one frame, run now, nothing written;
# +class_params {name: {threshold?, iou_threshold?, min_box_frac?, max_box_frac?}}
# empty override = global value, keys are class names (REQ-181)
# +exemplars[] {box:[cx,cy,w,h], positive} — the touched class's own pool only
# +exemplar_class_name, +target_class_names; per-class, accumulating
# on the client (REQ-171,172,182)
DELETE /api/batches/{id}/classes/{class_id}/annotations # clear all shapes of class in batch (REQ-046);
# frame-scoped clear = client filters this frame → bulk-delete (REQ-180)
POST /api/batches/{ids}/approve # one or many, comma-separated → one merge job (REQ-131)
GET /api/batches/{ids}/triage/summary # one or many, comma-separated (REQ-130)
GET /api/batches/{ids}/triage/shapes
GET /api/batches/{ids}/triage/suggest
POST /api/batches/{ids}/triage/simulate
GET /api/frames/{id}/image?w=… # frame image / thumbnail
GET /api/frames/{id}/annotations
POST /api/frames/{id}/annotations # add a manual shape (REQ-042)
PATCH /api/annotations/{id} # move/resize/reclass
DELETE /api/annotations/{id}
POST /api/annotations/bulk-delete # {ids[]} → N deletes (REQ-180 frame-clear uses it)
POST /api/frames/{id}/assist # click/box → SAM3 shape (REQ-043)
POST /api/frames/{id}/exemplar-label # drawn pool → re-detect one class (REQ-173,174)
POST /api/frames/{id}/status # approved | rejected | pending (REQ-041)
POST /api/projects/{id}/train # → train job (REQ-060,061,062)
GET /api/projects/{id}/models # versions + metrics (REQ-063,064)
GET /api/models/{id}/weights # download best.pt
POST /api/models/{id}/promote # make it the project's base model (REQ-064)
PATCH /api/models/{id}/rename # update model display name { name }
GET /api/projects/{id}/live-count/models # weights this project can count with
POST /api/projects/{id}/live-count/start # {source|source_rel, model_path, dials}
# source must be a WHEP URL (REQ-176)
POST /api/live-count/stop
PATCH /api/live-count/line # move the line mid-session
GET /api/live-count/status # counts + preview: "webrtc" | "mjpeg"
GET /api/live-count/overlay # boxes/line/counts for the canvas (REQ-177)
GET /api/live-count/stream # MJPEG; 409 on a WebRTC session (REQ-177)
GET /api/jobs?project_id=… REQ-070,071
GET /api/jobs/{id}
POST /api/jobs/{id}/cancel
```
### Live counting (REQ-176, REQ-177)
One camera, one ingest on the streaming server, two consumers:
```
camera ──▶ MediaMTX ──┬── WHEP :8889/cam/whep ──▶ browser <video> + <canvas> overlay
└── RTSP :8554/cam ──▶ backend decode → YOLO → counter
└─▶ GET /api/live-count/overlay
```
The user types only the WHEP URL; `live_source.whep_to_rtsp` derives the RTSP one, with the
ports coming from `MEDIAMTX_RTSP_PORT` / `MEDIAMTX_WHEP_PATH`. Because the browser plays the
camera directly, a live session encodes no JPEG at all — `live_render.py` runs only for
archive files, which have no WebRTC leg. The canvas draws exactly what `live_render.py`
would have burned in, so both previews describe the same session.
## Job flows
**extract (REQ-020…023).** `ffmpeg -ss <start> -to <end> -i <video> -vf fps=<n> -q:v 2
frames/%06d.jpg`. `frames` rows are written once the files exist; the batch's frame count is
updated. Range and fps live on the batch, so one video can be used repeatedly.
**autolabel (REQ-030…034).** Per frame: one `set_image`, then loop each class's prompt (see
the domain invariants in `../AGENTS.md`), cross-class greedy NMS (REQ-031), write `annotations` rows with
`source='auto'`. `class_params` (REQ-181) is consulted per class: a numeric
`threshold` / `iou_threshold` / `min_box_frac` / `max_box_frac` entry replaces the job's global for that class
only; classes without an entry — and a run with `class_params` empty or absent — take the
exact global path. The YOLO confidence floor uses `min([global] + overrides)`, the SAM3
threshold/dedup/min-box/max-box lists are built only from classes that actually override that key,
and `/preview` applies the same overrides so the preview and the job cannot disagree.
The job and `/preview` also read `project_classes.container` per class when building the
container-id set for the NMS containment carve-out — same stored flag for both (REQ-184,
amended REQ-031).
A re-run deletes only `source='auto'` rows — manual corrections
(`source='manual'`) survive — and returns already-approved frames to `pending`, because
that approval was given against labels that no longer exist. No overlay images are written:
the review canvas draws the shapes from the annotation rows, so a second rendering of the
same data on the server would only be a second thing to keep in sync.
**Deleting a class (REQ-007).** `projects.delete_class` removes the class's annotations,
decrements every `class_id` above it in `annotations` and `project_classes`, then calls
`dataset.drop_class_from_labels` to do the same edit to every `.txt` already written to
disk, and rewrites `data.yaml`. The renumbering is the whole job: a YOLO label is an integer
index, so a class list and a set of label files that disagree do not fail loudly — they
train a model on the wrong names. Refused for a project's last class.
**merge (REQ-050…053, REQ-131…132).** One job covers the whole selected set of batches, and
`datasets.rules_json` holds the triage rules frozen at the moment the merge was confirmed —
the resolver is built from that snapshot, never from the project's live rules. For every
`approved` frame not yet in `dataset_items`: assign a
split (continuing the round-robin), copy the JPEG to `dataset/images/<split>/`, write the
YOLO `.txt` from the frame's annotations, record the row. Finally rewrite `data.yaml`.
Frames with no annotations produce an empty `.txt` (REQ-033).
**train (REQ-060…065).** Release the SAM3 engine → `YOLO(base/model.pt)` (or pretrained if
the project has no base yet) → `.train(data=dataset/data.yaml, **hardware.defaults())` →
`evaluate.py` runs `.val()` for both the base model and the new one against the same
`data.yaml` → store `models/<n>/best.pt` + `metrics.json`. Each version is auto-named
`{arch}-{labelType}-{epochs}ep-{classNames}-{YYYYMMDD}` (e.g. `yolo11n-bbox-100ep-sack+box-20260909`).
`hardware.py` picks defaults from the detected VRAM:
| VRAM | batch | imgsz |
|---|---|---|
| < 8 GB | 8 | 640 |
| 8–16 GB | 16 | 640 |
| > 16 GB | 32 | 768 |
CPU-only: `batch=4`, `imgsz=512`, with a warning that training will be very slow. These are
form defaults only (REQ-062).
## Frontend
React + Vite, no heavy UI library; plain `fetch`, job polling once a second as today.
Design system and validation checklist follow ui-ux-pro-max — see `../AGENTS.md` §7.
Design constraints specific to this app:
- **Dense tool, not a landing page.** Neutral surface, one accent colour, tight spacing.
The frame and the canvas own the screen; panels are chrome.
- **Dark by default** — annotation work happens on video stills, and a bright surround
distorts the judgement of what's in the frame. Light mode is a toggle, not an afterthought.
- **Class colours are data, not decoration.** One fixed, colour-blind-safe hue per class
index, identical in the filmstrip, the canvas, and the class panel.
- **Keyboard first in Review.** Every action there has a shortcut and a visible focus state;
the mouse is for drawing shapes, not for navigating.
- **Progress is always visible.** Long jobs (extraction, auto-annotation, training) show
progress, elapsed time, and a cancel affordance — never a spinner with no end.
Layout Architecture:
- **Roboflow-Replica Layout**: Dense left sidebar (`Sidebar.jsx`) with Workspace switcher, navigation sections:
- **WORKSPACE**: Projects (`/projects`)
- **DATA**: Video Archive (`/projects/{id}`), Annotate / Review (`/batches/{id}`), Master Dataset (`/projects/{id}/dataset`)
- **MODELS**: Train & Select Engine (`/projects/{id}/models`), Model Monitoring & Analytics
- **DEPLOY**: Model Deployments & Base Model Promotion (`/projects/{id}/deploy`)
- **System Health Footer**: Hardware GPU, Free VRAM, SAM3 readiness, and ffmpeg status.
Pages:
1. **Projects** — workspace project cards + new project form (name, label type, base model, video root, classes & prompts).
2. **Library** — dates column → videos/batches with duration, resolution, used marker.
3. **Trim** — `<video>` + in/out timeline, fps input, estimated frame count, extract button.
4. **Review** — status-coloured filmstrip, canvas editor, class panel, shortcuts
(`←`/`→` frame, `A` approve, `X` reject, `Del` delete shape), *Approve batch* button.
5. **Models** — model engine selection cards (Custom Training vs NAS/Pretrained), train button, job progress, base-vs-new mAP comparison table, download & *promote*.
The canvas editor is hand-written; the normalized-coordinate conventions already exist in
`sessions.py` (`annotations_payload`, `detections_payload`) as a reference.
**Auto-annotate modal (REQ-171, REQ-172, REQ-182, REQ-184, REQ-186).** `AutoAnnotateModal.jsx` splits into
`PreviewShapes.jsx` (the result overlay, shared with the mass modal),
`ClassPromptPanel.jsx` (class chips + the editable SAM3 prompt) and `ExemplarCanvas.jsx`
(the drag-to-draw layer). All three overlays and the `<img>` share one shrink-wrapped
`position: relative` wrapper — they are sized to it, so nothing else may sit inside it or
every box shifts off the pixels it describes.
With SAM3 one selected chip is *active*: it owns the prompt field and any exemplars drawn on
the frame, so its chip is a pair of buttons — the name activates, the `×` deselects.
Exemplars are normalized `[cx, cy, w, h]`, one pool per class keyed by class name
(`hooks/useExemplarPools.js`), sent only to `/preview`, and dropped when the frame changes —
switching the active class swaps pools instead of discarding them (REQ-172). The preview is
per class and accumulates (REQ-182): exemplar changes share **one 250 ms debounce** that
re-runs the most recently touched class against its own pool and replaces only that class's
shapes — rapid alternation between classes re-runs only the last one — while the other
exemplared classes' results stay on screen and the canvas **merges** (amended REQ-182):
exemplared classes contribute their pool-conditioned results (replacing their own full-set
boxes), every other selected class keeps the detections of the last full-set run on that
frame. Undo is the same re-run with the shortened pool, because SAM3 can
only append geometric prompts. Clearing a class's last example sends no request and returns
that class to its full-set detections (blank if no full-set run covered it); clearing all
examples is then a plain full-set view again. Run Preview always refreshes the full-set for
every selected class first (with empty exemplars), then re-runs the exemplared classes
sequentially; drawing an example never triggers a full-set re-run. The
per-class overrides block (REQ-181) stacks one row per class — name plus **Container**
checkbox on the first line, the four override inputs in a wrapping grid below — and that
checkbox toggles
`project_classes.container` through `PATCH /api/projects/{id} { containers: … }`. The same
block carries a **Copy** and a **Paste** button (REQ-186, REQ-190): one click puts every
selected class's effective settings (per-class override where set, else the global slider;
container = the checkbox state) on the clipboard as **YAML** — one block per class, fields
named after the row labels (`conf`, `iou`, `minbox`, `maxbox`, `container`), through
`clipboard.js`'s `copyText` (Clipboard API with an `execCommand` fallback for insecure
contexts) — read-only, no network, no setting changed. **Paste** opens `PasteYamlDialog` — a textarea plus a live review line, nothing written until
Apply, and Apply disabled while the text does not parse. It prefills from the clipboard when
the browser allows a read and otherwise waits for Ctrl+V, which is what makes paste work on
plain http on a LAN address (`readText` needs a secure origin; a keyboard paste does not).
`planPaste()` in `ClassParamsTable.jsx` is what produces that review line; `applyPasted()`
writes it. `container` diffs go out as one
`PATCH /api/projects/{id} { containers: { classId: bool } }` and revert together on
rejection. Parsing is `parseClassYaml` in the same file — a strict subset reader for exactly
what Copy emits, no YAML dependency, all-or-nothing with a `line N: …` error. Because
`ClassParamsTable` is shared, the mass modal carries both buttons too.
**Review sidebar class rows (REQ-180, REQ-183, REQ-185, REQ-191).** Each class row's frame-clear `×`
(REQ-180) gains an eye toggle in front of it: session-only `Set` state in `ReviewPage`, so it
survives frame changes, resets when the review page is left, and never touches stored data.
Hidden shapes are never handed to `AnnotationCanvas` — not drawn, not clickable, not
marquee-selectable — but they stay listed, dimmed, in "Shapes on this frame". Two session
Sets drive the hide half: `hiddenShapeIds` (the `H` key, REQ-185, which toggles the single
selection or all marked shapes) and `overriddenShapeIds` (the row-eye restore and the
auto-override below); a shape is visible when it is not `H`-hidden and (its class is visible
or it is overridden). Dimmed rows stay selectable, reclassable and deletable from the list —
the old "not deletable while hidden" purge invariant is superseded — while the class row
keeps its real per-frame count. Creating a shape while its class is hidden (draw, assist,
copy) auto-overrides it so the user sees what they just made; reclassing into a hidden class
dims it unless it already carries an override; toggling a class eye clears that class's overrides.
A third, independent piece of state is `revealAll` (REQ-191): a boolean **overlay** in
`ReviewPage`, toggled by `Shift+H`, that makes `isShapeVisible` answer "yes" for every shape.
It reads neither Set and writes neither Set, which is what makes the second `Shift+H` restore
the previous hidden state exactly; a plain `H` while it is on only clears the overlay, so the
keystroke edits the real hidden state from then on. The sidebar shows a `revealing hidden` chip
in the "Shapes on this frame" header while it is on (REQ-191).
**Exemplar-driven labeling in review (REQ-173, REQ-174, REQ-175).** In `draw` mode a drag on
`AnnotationCanvas` is an exemplar, not a rectangle: `onExemplar(box, positive)` where
`positive` is `!event.shiftKey`. `hooks/useExemplarPool.js` keeps the pool in a ref as well
as state — the ref is what gets sent, so a drag that lands mid-flight is never lost — and posts
the **whole pool** to `/exemplar-label` 400 ms after the last drag. One pass runs at a time;
a drag arriving during a pass sets `rerunWanted` so exactly one rerun follows instead of a
queue. The pool is cleared by a frame change or a class change, and `Undo example` re-runs
with the shortened pool.
The run is a **dry run by default** (REQ-175). `ExemplarFilterPanel.jsx` is passed to
`ReviewSidebar` and rendered above the class list — not over the frame, which is where its
proposals are drawn. It is mounted only while a run is undecided (`pool.active`): the first
drag opens it, Apply and Discard close it, and the pool outlives it. Its five sliders
re-preview on a 250 ms debounce. While a preview is up the canvas **hides the stored shapes
of the class under review** — the run replaces them wholesale, and leaving them on screen made
a rejected detection look like it had never gone; other classes stay, dimmed
(`svg.previewing`). Proposals draw dashed on top, green for the boxes the user drew and
class-colored for what SAM3 found.
`Apply` re-runs with `apply: true` rather than posting the previewed geometry back — SAM3 is
deterministic for a pool and a threshold, and the browser should not be the authority on what
gets stored. `Discard` truncates the pool to `appliedRef`, the length it had at the last
successful apply, so a rejected run leaves neither shapes nor prompts behind. A successful
apply drops the negatives it sent and keeps the positives (REQ-174); drags that landed while
that request was in flight are not part of it and stay at the end of the pool, with a rerun
queued for them.
`.canvas-wrap`'s overlay rules are scoped to its direct child (`> svg`): they set
`position: absolute; width: 100%`, which any nested SVG — an icon in a panel, say — would
otherwise inherit and stretch across the whole frame.
Request/response:
```
POST /api/frames/{id}/exemplar-label
{ "exemplars": [{"box": [x0,y0,x1,y1], "positive": true}, …], # normalized xyxy
"class_id": 0,
"threshold": 0.5, "iou_threshold": 0.8, # the panel (REQ-175)
"min_box_frac": 0.002, "max_box_frac": 1.0, "max_detections": 100,
"apply": false }
→ { "shapes": [{"geometry": …, "score": …, "source": "manual"|"auto"}, …],
"applied": false, "redetected": true, "message": null,
"annotations": null } # the frame, on apply only
```
`backend/exemplar.py` converts each box to SAM3's normalized `[cx, cy, w, h]`, runs one
`open_state` + `apply_prompts` with the class's stored `prompt` as text plus every exemplar,
then rewrites the frame **for that class only**:
- positives are stored first as `source='manual'` — the literal rectangle in a `bbox`
project, the mask polygon of whatever SAM3 found inside it (IoU ≥ `SNAP_IOU`) in a
`polygon` one;
- detections overlapping a negative by ≥ `NEGATIVE_IOU` (0.3) are dropped, and ones
overlapping a positive by ≥ `DUPLICATE_IOU` (0.6) are dropped as the user's own shape
already covers them;
- what remains is written as `source='auto'`, so a later batch re-run replaces it (REQ-034)
while the drawn shapes survive.
The panel's filters run before any of that, in the order the batch job uses them: area floor,
then ceiling (REQ-188), then NMS (`labeling.deduplicate`), then the cap on how many survive.
The delete-and-reinsert happens in one `db.cursor()` transaction, so the frame is never
briefly empty. Other classes on the frame are never touched. The GPU lock is taken with a
**5 s** timeout — shorter than `assist`'s 20 s because this fires from a mouse gesture; on
timeout the response carries `redetected: false`, a message the panel shows, and the drawn
boxes alone as the preview. Applying that run **appends** the drawn boxes and honours the
negatives instead of taking the replace path — with no detections to put back, replacing
would wipe the class and leave only the drawings.
## Docker (REQ-072)
- `Dockerfile` — python 3.12, `ffmpeg`, `uv`, CUDA torch, `uv pip install -e sam3/`. The
known traps still apply: `setuptools<81` (because `sam3` imports `pkg_resources`) and the
undeclared `einops` + `pycocotools` dependencies of the vendored SAM3.
- `docker-compose.yml` — a `backend` service (GPU passthrough, `./data` volume, video archive
mounted read-write (REQ-178), `.env` for `HF_TOKEN`) and a `frontend` service (nginx: static
files + `/api` proxy, `client_max_body_size 20g`).
- GPU access uses compose **`gpus: all`** (same path as the `docker run --gpus all` CLI flag).
The earlier CDI form (`devices: nvidia.com/gpu=all`) fails on Docker Desktop's WSL2 backend
with `unresolvable CDI devices` because no CDI spec exists there; `gpus: all` works on both
Docker Desktop and a native engine with the NVIDIA container toolkit.
The runtime entry in `/etc/docker/daemon.json` is not registered with the daemon here.
- The frontend pins **Vite 7**. Vite 8's default bundler (Rolldown) ships a native binding
that dies with a bus error on this machine; Vite 7's Rollup path works.