This commit includes major additions and updates to the frontend and backend architectures, introducing new dataset management, live counting features, batch processing, and triage logic. Includes new UI pages, components, and API routes.
293 lines
15 KiB
Markdown
293 lines
15 KiB
Markdown
# Design
|
||
|
||
Serves `./requirements.md`. Each section names the REQs it fulfils.
|
||
|
||
## Overview
|
||
|
||
```
|
||
Browser (React + Vite, served by nginx)
|
||
│ /api → proxy
|
||
▼
|
||
FastAPI ──► jobs (1 worker thread, 1 GPU)
|
||
│ ├─ extract : ffmpeg (CPU)
|
||
│ ├─ autolabel: SAM3 (GPU)
|
||
│ ├─ merge : copy + write labels (CPU)
|
||
│ └─ train : Ultralytics + eval (GPU)
|
||
├──► SQLite (metadata & status)
|
||
└──► data/ (frames, master dataset, weights)
|
||
```
|
||
|
||
Storage split rule: **SQLite holds metadata and status; the disk holds pixels, final labels,
|
||
and weights.** The master dataset must stay useful even if the database is lost
|
||
(REQ-006, REQ-054).
|
||
|
||
## Disk layout (REQ-006)
|
||
|
||
```
|
||
data/ # Docker volume
|
||
app.db # SQLite (WAL)
|
||
projects/<slug>/
|
||
base/model.pt # the project's active base model (REQ-003)
|
||
dataset/ # MASTER, accumulative (REQ-050…052)
|
||
images/{train,val}/…jpg
|
||
labels/{train,val}/…txt
|
||
data.yaml
|
||
batches/<batch-id>/
|
||
frames/000001.jpg … # extraction output (REQ-022)
|
||
models/<n>/
|
||
best.pt
|
||
metrics.json # base vs new metrics (REQ-063)
|
||
runs/ # Ultralytics run directory
|
||
```
|
||
|
||
Master dataset filenames: `<batch-id>__<frame number>.jpg` — unique across batches and
|
||
self-documenting about where each image came from. The user's video archive is read-only
|
||
(REQ-074).
|
||
|
||
## SQLite schema
|
||
|
||
Created by an idempotent migration in `backend/db.py` at startup.
|
||
|
||
```sql
|
||
projects(
|
||
id, slug UNIQUE, name, label_type CHECK(bbox|polygon),
|
||
base_model_path, base_model_kind CHECK(uploaded|pretrained|trained),
|
||
video_root, val_every DEFAULT 5, created_at)
|
||
|
||
project_classes(
|
||
id, project_id → projects, class_id INT, name, prompt,
|
||
UNIQUE(project_id, class_id)) -- class_id = the YOLO class index (REQ-003/005)
|
||
|
||
batches(
|
||
id, project_id → projects, video_path, date_label, batch_label,
|
||
start_sec REAL, end_sec REAL, fps REAL,
|
||
status CHECK(extracting|extracted|labeling|reviewing|approved|merged|failed),
|
||
frame_count INT, created_at, merged_at)
|
||
|
||
frames(
|
||
id, batch_id → batches, idx INT, filename, width INT, height INT,
|
||
review_status CHECK(pending|approved|rejected) DEFAULT 'pending',
|
||
UNIQUE(batch_id, idx))
|
||
|
||
annotations(
|
||
id, frame_id → frames, class_id INT,
|
||
geometry TEXT, -- JSON; see "Geometry format"
|
||
score REAL, source CHECK(auto|manual), created_at)
|
||
|
||
dataset_items( -- master dataset membership (REQ-052)
|
||
id, project_id → projects, frame_id → frames UNIQUE,
|
||
split CHECK(train|val), image_rel, label_rel, added_at)
|
||
|
||
model_versions(
|
||
id, project_id → projects, version INT, weights_path,
|
||
parent_model_path, metrics TEXT, base_metrics TEXT, created_at,
|
||
UNIQUE(project_id, version))
|
||
|
||
jobs( -- persistent (REQ-071)
|
||
id, project_id, batch_id, type CHECK(extract|autolabel|merge|train),
|
||
status CHECK(queued|running|done|failed|cancelled),
|
||
progress INT, total INT, message, error, log TEXT,
|
||
created_at, started_at, finished_at)
|
||
```
|
||
|
||
**The stable val split (REQ-052)** is enforced by `dataset_items`: an existing row never
|
||
changes its `split`. On merge, only frames without a row are assigned, using a per-project
|
||
round-robin counter (`val_every`) that continues from the previous count.
|
||
|
||
**Geometry format.** One JSON column covers both label types (REQ-002):
|
||
|
||
- `bbox` → `{"type":"bbox","points":[x0,y0,x1,y1]}`
|
||
- `polygon` → `{"type":"polygon","points":[[x,y], …]}`
|
||
|
||
Coordinates are stored **normalized 0–1** against the frame size, so neither the editor nor
|
||
the exporter needs to know the display size. SAM3 mask → polygon conversion is
|
||
`review.mask_to_polygons()`; for `bbox` projects the mask is only used to take its bounding
|
||
box.
|
||
|
||
## Backend modules
|
||
|
||
`app/` moves to `backend/`. Reuse existing code wherever possible:
|
||
|
||
| Module | Role | Status |
|
||
|---|---|---|
|
||
| `sam3_engine.py` | SAM3 singleton, `open_state`/`apply_prompts`/`segment_at` | reused, plus a `release()` for REQ-065 |
|
||
| `labeling.py` | per-frame detection + cross-prompt NMS (REQ-031) | reused; the folder-walking half went with the old flow |
|
||
| `exporters.py` | ~~YOLO label writing~~ | **deleted** — `dataset.py` writes labels, `mask_to_polygons` moved to `review.py` |
|
||
| `sessions.py` | ~~exemplar/tap interaction~~ | **deleted** — see below |
|
||
| `jobs.py` | single-worker queue | extended: job types + persistence |
|
||
| `training.py` | Ultralytics fine-tune | changed: starts from the base model, args from `hardware.py` |
|
||
| `db.py` | connection + migration | **new** |
|
||
| `projects.py` | project CRUD, reads classes from a `.pt` | **new** |
|
||
| `library.py` | scans `<video_root>/<date>/<batch>` | **new** |
|
||
| `video.py` | `ffprobe`, Range streaming, `ffmpeg` extraction | **new** |
|
||
| `batches.py` | batch lifecycle | **new** |
|
||
| `review.py` | annotation CRUD, frame status, click-assist | **new** |
|
||
| `autolabel.py` | the SAM3 job over a whole batch | **new** |
|
||
| `dataset.py` | merge into the master dataset, stable split | **new** |
|
||
| `evaluate.py` | validate base vs new model | **new** |
|
||
| `hardware.py` | VRAM detection → training defaults | **new** |
|
||
| `api/` | the FastAPI routes, one module per domain | **new** |
|
||
|
||
Removed: `uploads.py`, `static/index.html`, and the old flow's endpoints.
|
||
|
||
`sessions.py` was meant to be reused for click-assist, but it existed to hold GPU-resident
|
||
state for an interactive session — a whole eviction policy, an undo stack, and a per-session
|
||
annotation store, all of which the database and the stateless `review.assist` now cover.
|
||
Adapting 357 lines to do what 60 lines do was not worth it, so the module is gone. The one
|
||
thing it knew that mattered — SAM3 wants exemplar boxes as normalized centre-x, centre-y,
|
||
width, height — moved with it.
|
||
|
||
The 400-line file limit (see `../AGENTS.md`) applies to all of the above. It is why the
|
||
routes live in `backend/api/{projects,batches,review,models,jobs}.py` rather than in
|
||
`main.py`, which now only builds the app and owns startup. Route modules import their
|
||
domain module under an alias (`from backend import projects as project_store`) so the two
|
||
namespaces stay distinguishable.
|
||
|
||
## API contract
|
||
|
||
```
|
||
GET /api/health REQ-073
|
||
|
||
GET /api/projects REQ-001
|
||
POST /api/projects REQ-001,002,004,005
|
||
GET /api/projects/{id} # includes label_type_locked: bool (REQ-002)
|
||
DELETE /api/projects/{id}
|
||
PATCH /api/projects/{id} # class prompts, val_every (REQ-005)
|
||
POST /api/projects/{id}/classes # add new class {name, prompt} (REQ-008)
|
||
DELETE /api/projects/{id}/classes/{class_id} # delete class, delete shapes, reindex classes (REQ-007)
|
||
POST /api/projects/{id}/base-model # upload .pt, read classes (REQ-003)
|
||
GET /api/projects/{id}/dataset # master dataset summary (REQ-053)
|
||
GET /api/projects/{id}/dataset/download # zip (REQ-054)
|
||
|
||
GET /api/projects/{id}/library # list of dates (REQ-011)
|
||
GET /api/projects/{id}/library/{date} # videos + duration/resolution (REQ-012)
|
||
GET /api/projects/{id}/video?rel=… # Range streaming (REQ-013)
|
||
|
||
POST /api/projects/{id}/batches # {rel, start_sec, end_sec, fps} → extract job
|
||
GET /api/batches/{id} # status + review progress (REQ-045)
|
||
GET /api/batches/{id}/frames # frames + statuses
|
||
POST /api/batches/{id}/autolabel # {threshold} → job (REQ-030,032,034)
|
||
DELETE /api/batches/{id}/classes/{class_id}/annotations # clear all shapes of class in batch (REQ-046)
|
||
POST /api/batches/{ids}/approve # one or many, comma-separated → one merge job (REQ-131)
|
||
GET /api/batches/{ids}/triage/summary # one or many, comma-separated (REQ-130)
|
||
GET /api/batches/{ids}/triage/shapes
|
||
GET /api/batches/{ids}/triage/suggest
|
||
POST /api/batches/{ids}/triage/simulate
|
||
|
||
GET /api/frames/{id}/image?w=… # frame image / thumbnail
|
||
GET /api/frames/{id}/annotations
|
||
POST /api/frames/{id}/annotations # add a manual shape (REQ-042)
|
||
PATCH /api/annotations/{id} # move/resize/reclass
|
||
DELETE /api/annotations/{id}
|
||
POST /api/frames/{id}/assist # click/box → SAM3 shape (REQ-043)
|
||
POST /api/frames/{id}/status # approved | rejected | pending (REQ-041)
|
||
|
||
POST /api/projects/{id}/train # → train job (REQ-060,061,062)
|
||
GET /api/projects/{id}/models # versions + metrics (REQ-063,064)
|
||
GET /api/models/{id}/weights # download best.pt
|
||
POST /api/models/{id}/promote # make it the project's base model (REQ-064)
|
||
|
||
GET /api/jobs?project_id=… REQ-070,071
|
||
GET /api/jobs/{id}
|
||
POST /api/jobs/{id}/cancel
|
||
```
|
||
|
||
## Job flows
|
||
|
||
**extract (REQ-020…023).** `ffmpeg -ss <start> -to <end> -i <video> -vf fps=<n> -q:v 2
|
||
frames/%06d.jpg`. `frames` rows are written once the files exist; the batch's frame count is
|
||
updated. Range and fps live on the batch, so one video can be used repeatedly.
|
||
|
||
**autolabel (REQ-030…034).** Per frame: one `set_image`, then loop each class's prompt (see
|
||
the domain invariants in `../AGENTS.md`), cross-prompt NMS, write `annotations` rows with
|
||
`source='auto'`. A re-run deletes only `source='auto'` rows — manual corrections
|
||
(`source='manual'`) survive — and returns already-approved frames to `pending`, because
|
||
that approval was given against labels that no longer exist. No overlay images are written:
|
||
the review canvas draws the shapes from the annotation rows, so a second rendering of the
|
||
same data on the server would only be a second thing to keep in sync.
|
||
|
||
**Deleting a class (REQ-007).** `projects.delete_class` removes the class's annotations,
|
||
decrements every `class_id` above it in `annotations` and `project_classes`, then calls
|
||
`dataset.drop_class_from_labels` to do the same edit to every `.txt` already written to
|
||
disk, and rewrites `data.yaml`. The renumbering is the whole job: a YOLO label is an integer
|
||
index, so a class list and a set of label files that disagree do not fail loudly — they
|
||
train a model on the wrong names. Refused for a project's last class.
|
||
|
||
**merge (REQ-050…053, REQ-131…132).** One job covers the whole selected set of batches, and
|
||
`datasets.rules_json` holds the triage rules frozen at the moment the merge was confirmed —
|
||
the resolver is built from that snapshot, never from the project's live rules. For every
|
||
`approved` frame not yet in `dataset_items`: assign a
|
||
split (continuing the round-robin), copy the JPEG to `dataset/images/<split>/`, write the
|
||
YOLO `.txt` from the frame's annotations, record the row. Finally rewrite `data.yaml`.
|
||
Frames with no annotations produce an empty `.txt` (REQ-033).
|
||
|
||
**train (REQ-060…065).** Release the SAM3 engine → `YOLO(base/model.pt)` (or pretrained if
|
||
the project has no base yet) → `.train(data=dataset/data.yaml, **hardware.defaults())` →
|
||
`evaluate.py` runs `.val()` for both the base model and the new one against the same
|
||
`data.yaml` → store `models/<n>/best.pt` + `metrics.json`.
|
||
|
||
`hardware.py` picks defaults from the detected VRAM:
|
||
|
||
| VRAM | batch | imgsz |
|
||
|---|---|---|
|
||
| < 8 GB | 8 | 640 |
|
||
| 8–16 GB | 16 | 640 |
|
||
| > 16 GB | 32 | 768 |
|
||
|
||
CPU-only: `batch=4`, `imgsz=512`, with a warning that training will be very slow. These are
|
||
form defaults only (REQ-062).
|
||
|
||
## Frontend
|
||
|
||
React + Vite, no heavy UI library; plain `fetch`, job polling once a second as today.
|
||
Design system and validation checklist follow ui-ux-pro-max — see `../AGENTS.md` §7.
|
||
|
||
Design constraints specific to this app:
|
||
|
||
- **Dense tool, not a landing page.** Neutral surface, one accent colour, tight spacing.
|
||
The frame and the canvas own the screen; panels are chrome.
|
||
- **Dark by default** — annotation work happens on video stills, and a bright surround
|
||
distorts the judgement of what's in the frame. Light mode is a toggle, not an afterthought.
|
||
- **Class colours are data, not decoration.** One fixed, colour-blind-safe hue per class
|
||
index, identical in the filmstrip, the canvas, and the class panel.
|
||
- **Keyboard first in Review.** Every action there has a shortcut and a visible focus state;
|
||
the mouse is for drawing shapes, not for navigating.
|
||
- **Progress is always visible.** Long jobs (extraction, auto-annotation, training) show
|
||
progress, elapsed time, and a cancel affordance — never a spinner with no end.
|
||
|
||
Layout Architecture:
|
||
|
||
- **Roboflow-Replica Layout**: Dense left sidebar (`Sidebar.jsx`) with Workspace switcher, navigation sections:
|
||
- **WORKSPACE**: Projects (`/projects`)
|
||
- **DATA**: Video Archive (`/projects/{id}`), Annotate / Review (`/batches/{id}`), Master Dataset (`/projects/{id}/dataset`)
|
||
- **MODELS**: Train & Select Engine (`/projects/{id}/models`), Model Monitoring & Analytics
|
||
- **DEPLOY**: Model Deployments & Base Model Promotion (`/projects/{id}/deploy`)
|
||
- **System Health Footer**: Hardware GPU, Free VRAM, SAM3 readiness, and ffmpeg status.
|
||
|
||
Pages:
|
||
|
||
1. **Projects** — workspace project cards + new project form (name, label type, base model, video root, classes & prompts).
|
||
2. **Library** — dates column → videos/batches with duration, resolution, used marker.
|
||
3. **Trim** — `<video>` + in/out timeline, fps input, estimated frame count, extract button.
|
||
4. **Review** — status-coloured filmstrip, canvas editor, class panel, shortcuts
|
||
(`←`/`→` frame, `A` approve, `X` reject, `Del` delete shape), *Approve batch* button.
|
||
5. **Models** — model engine selection cards (Custom Training vs NAS/Pretrained), train button, job progress, base-vs-new mAP comparison table, download & *promote*.
|
||
|
||
|
||
The canvas editor is hand-written; the normalized-coordinate conventions already exist in
|
||
`sessions.py` (`annotations_payload`, `detections_payload`) as a reference.
|
||
|
||
## Docker (REQ-072)
|
||
|
||
- `Dockerfile` — python 3.12, `ffmpeg`, `uv`, CUDA torch, `uv pip install -e sam3/`. The
|
||
known traps still apply: `setuptools<81` (because `sam3` imports `pkg_resources`) and the
|
||
undeclared `einops` + `pycocotools` dependencies of the vendored SAM3.
|
||
- `docker-compose.yml` — a `backend` service (GPU passthrough, `./data` volume, video archive
|
||
mounted read-only, `.env` for `HF_TOKEN`) and a `frontend` service (nginx: static files +
|
||
`/api` proxy).
|
||
- GPU access uses **CDI** (`devices: nvidia.com/gpu=all`), not the legacy `runtime: nvidia`.
|
||
Docker Engine 27+ discovers `nvidia.com/gpu` through the container toolkit; the runtime
|
||
entry in `/etc/docker/daemon.json` is not registered with the daemon here.
|
||
- The frontend pins **Vite 7**. Vite 8's default bundler (Rolldown) ships a native binding
|
||
that dies with a bus error on this machine; Vite 7's Rollup path works.
|