Files

354 lines
23 KiB
Markdown

# Requirements
Status: **agreed** (planning session, 2026-07-31). Changes only with the user's approval.
## Goal
A system for **enriching a dataset and improving an existing detection model**, iteratively,
from an archive of recorded video. One full round:
> pick a project → browse the video archive → pick a batch → trim a time range →
> extract frames → auto-annotate with SAM3 → review and correct every frame → approve →
> merge into the master dataset → fine-tune from the base model → compare against the base.
The system is **generic**: the sack case is only the first project. Other cases are
created as new projects with their own base model, classes, and video archive — no code
changes.
## Non-goals (for this version)
- Login, multi-user, tenants, quotas. The architecture leaves room for them; the features
are not built.
- Tracking or annotation propagation between frames.
- Collaborative annotation by several people at once.
- Public internet deployment.
---
## A. Project
- **REQ-001** — The user can create, list, and delete projects. A project has: name, label
type, base model, video archive root, and a class list.
- **REQ-002** — Each project picks a **label type**: `bbox` (YOLO detect) or `polygon`
(YOLO segment). This determines the export format, the editor's behaviour, and which
model variant is trained. It cannot be changed once a batch has been merged.
- **REQ-003** — The user uploads a **base model** `.pt`. The class list is read from the
model (`model.names`). It cannot drift on its own: nothing adds or removes a class as a
side effect of another action. Deliberate deletion is REQ-007.
- **REQ-004** — A project may be created **without** a base model. In that case the user
types the class list, and the first training starts from pretrained weights
(`yolo11n.pt` / `yolo11n-seg.pt`).
- **REQ-005** — Each class carries its own **SAM3 text prompt**, which may differ from the
class name (e.g. class `sack` with prompt `"woven plastic sack"`). Prompts can be
edited at any time without affecting existing data.
- **REQ-006** — All of a project's data (base model, master dataset, batch frames, trained
weights) lives under one project folder, so it can be copied or backed up whole.
- **REQ-007** — The user can **delete a class** at any point in a project's life, including
after batches have been merged. Deleting one:
- removes every annotation of that class, in batches under review and in the master
dataset alike;
- **renumbers the classes above it**, in the database *and* in every label file already
written to disk, because a YOLO label is an integer index and leaving a gap would make
old labels silently name the wrong class;
- regenerates `data.yaml`;
- is refused for a project's last remaining class.
Before confirming, the user is told how many shapes will be destroyed. The action cannot
be undone. If the project's base model was trained on the old class list, it stops being
comparable — which REQ-063 already reports rather than hides.
- **REQ-008** — The user can **add a new class** (name & prompt) to an existing project at any time.
The new class receives the next sequential `class_id`, and `data.yaml` is regenerated if a
master dataset exists.
## B. Video archive
- **REQ-010** — The video archive lives at a local path (disk or mount); videos are **not
uploaded** through the browser.
- **REQ-011** — Archive structure: `<video_root>/<date>/<batch>.<ext>`. The system lists the
dates, and within each date the videos with their batch labels parsed from the filename.
- **REQ-012** — Each video shows its duration, resolution, and whether it has already been
used as a batch in this project.
- **REQ-013** — Videos play in the browser with seeking (HTTP Range), without copying the
file first.
## C. Trim & frame extraction
- **REQ-020** — The user sets the in/out range with a timeline slider on the player, and can
also type precise timestamps.
- **REQ-021** — The user sets the extraction **frames per second** (default 1 fps). The
resulting frame count is shown before extraction runs.
- **REQ-022** — Extraction runs as a background job with progress, producing sequentially
numbered JPEG files inside the batch folder.
- **REQ-023** — One video may be used more than once with different time ranges; each
extraction produces its own batch.
## D. Auto-annotation
- **REQ-030** — Once frames are extracted, the system runs SAM3 over all of them using each
class's prompt, as a background job with progress and cancellation.
- **REQ-031** — Detections that overlap across prompts are deduplicated (greedy IoU NMS), so
one object is not labelled as two classes at once.
- **REQ-032** — The confidence threshold is configurable per job.
- **REQ-033** — A frame with no detections is valid and still enters the dataset as a
negative sample — it is not a failure.
- **REQ-034** — Auto-annotation can be re-run on the same batch; previous automatic results
are replaced, but **the user's manual corrections must never be lost**.
- **REQ-035** — Auto-annotation can be started in **resume** mode, which skips frames that
already carry automatic annotations. Resume is always an explicit choice and never the
default, because a full re-run is also how the confidence threshold (REQ-032) is changed —
the system cannot tell the two intentions apart, so it asks. A frame SAM3 legitimately
found nothing on (REQ-033) writes no annotations, so a resume re-does it; that is accepted
rather than tracked.
- **REQ-171** — In the auto-annotate modal the SAM3 text prompt of each selected class is
editable in place, next to the live preview. Saving it writes
`project_classes.prompt` — the same field the Projects page edits — so the batch job and
every later run send that text. The preview is the tuning surface; the stored prompt is the
artifact it produces.
- **REQ-172** — On the previewed frame the user can drag **positive** and **negative** box
exemplars (shift-drag for negative). They are appended to the active class's text prompt,
re-run immediately, and can be undone or cleared. Exemplars are a **tuning aid only**: they
are never written as annotations and never carried into the batch job, because SAM3's
geometric prompts pool features from the current image — replaying them on another frame
would ask about whatever happens to sit at those coordinates there. They belong to exactly
one class, so a new frame or a new active class discards them.
## E. Review & correction
- **REQ-040** — The user reviews frames one at a time, with fast navigation (left/right
arrows, thumbnail filmstrip, jump to the next unreviewed frame).
- **REQ-041** — Each frame has a status: `pending`, `approved`, or `rejected`. Rejected
frames never enter the dataset.
- **REQ-042** — The user can draw a new shape, move it, resize it, delete it, and change its
class.
- **REQ-043** — The user can ask SAM3 for help inside the editor: click or drag a box around
one object and the model produces its shape.
- **REQ-044** — All annotations and review statuses are **persistent** — they survive a
server restart, unlike today's in-memory sessions.
- **REQ-045** — Review progress is visible (e.g. "120/300 reviewed"), and a batch can only
be approved once no frame is still `pending`.
- **REQ-046** — The user can delete/clear all annotations of a specific class across all frames in
the current batch from the Review editor.
- **REQ-173** — In the review editor a plain drag on the canvas is an **exemplar-driven
label**, not just a rectangle. It proposes, in one action — and REQ-175's Apply is what
makes any of it real — that (a) the drawn shape becomes a `manual` annotation of the
active class — snapped to a SAM3 polygon first when the batch's
`label_type` is `polygon`, since a rectangle is a bad polygon label — (b) appends the box
to the frame's positive exemplar pool for that class, and (c) re-runs SAM3 over the whole
frame with the class's text prompt plus the pooled exemplars, deleting every existing shape
of that class on the frame and writing the detections in its place, then re-inserting the
pooled exemplar shapes verbatim so the user's own drawings always survive. The pool is
**frame-local and ephemeral** for the same reason as REQ-172 — SAM3's geometric prompts pool
features from the current image — so leaving the frame or switching the active class clears
it; the annotations it produced persist like any other. If the GPU lock (REQ-070) is not
free, the run comes back with the drawn shapes alone and says so, so the user can still
file them (REQ-175) and labeling is never blocked by a background job.
- **REQ-174** — **Shift**-drag in the review editor adds a **negative** exemplar. It is never
stored as an annotation; it deletes any existing shape of the active class that overlaps it,
and it is sent as a negative box in the REQ-173 re-detect. It is the "not this, and not
things like this" gesture, so it doubles as a delete. A negative is **spent on Apply**: the
frame it was applied to no longer carries what it rejected, so the drawing is dropped from
the pool while the positives stay on as prompts.
- **REQ-175** — An exemplar drag **previews**; it never writes on its own. The run's result
is drawn over the frame as proposals and a small panel floats on the canvas with the four
filters that decide what survives — confidence, NMS overlap, minimum box size, maximum
shapes — each re-running the preview as it moves. **Apply** writes the previewed set,
**Discard** rewinds the pool to whatever is already on the frame and leaves it untouched.
The panel is scoped to this gesture: its values are not stored, not shared with the
auto-annotate modal, and reset with the frame. Defaults are confidence `0.5`, NMS `0.8`,
min box `0.002`, max `100` — deliberately permissive, because on a dense frame an
aggressive NMS or area floor deletes real, touching objects rather than duplicates.
## E4. Live counting preview
- **REQ-176** — A **live** source on the Live Count page is a **WebRTC (WHEP) URL** and
nothing else; an RTSP URL is rejected with a message saying so. The backend derives the
RTSP leg of the same streaming-server path from it (`http://host:8889/cam` →
`rtsp://host:8554/cam`) and counts from that: WebRTC is what makes the browser preview
cheap, but pulling it into Python would add ICE and a jitter buffer on top of the identical
H.264 decode. One ingest on the streaming server, two consumers. The ports are read from
the environment (`MEDIAMTX_RTSP_PORT`, `MEDIAMTX_WHEP_PATH`), never hardcoded. Archive
files are unaffected — they are still opened as files.
- **REQ-177** — A live session is **watched over WebRTC**, played straight from the streaming
server by the browser: the frames never pass through this app and it encodes no JPEG for
them. What the model saw — boxes, ids, confidences, the counting line and its band, the
ignored region, the running totals — is served as geometry from
`GET /api/live-count/overlay` and drawn on a canvas over the video. The MJPEG endpoint
remains the preview for **archive files** only, and refuses a WebRTC session.
## F. Master dataset
- **REQ-050** — Approving a batch **merges** its approved frames and their labels into the
project's master dataset (accumulating across batches).
- **REQ-051** — The master dataset is train-ready YOLO format: `images/{train,val}`,
`labels/{train,val}`, and a `data.yaml` regenerated from the project's class list.
- **REQ-052** — **Stable val split**: once a frame is placed in `val`, it stays in `val`
across every later merge. New frames are split with an every-Nth pattern.
- **REQ-053** — The system records which batches have entered the master dataset, when, and
how many images/labels each added.
- **REQ-054** — The master dataset can be downloaded as a `.zip` (e.g. to import into
Roboflow or train on another machine).
## F2. Data Prep as the merge gate
- **REQ-130** — The Batches page supports multi-select. "Prepare & Merge Selected" opens
Data Prep scoped to exactly those batches (`#/projects/{id}/data-prep?batches=1,2,3`).
No dataset exists at this point.
- **REQ-131** — Data Prep is the merge gate. Filters and augmentation are tuned against the
selected batches' shapes; "Confirm merge" names or picks the target dataset and queues
**one** merge job for the whole selection. There is no path from Batches or Review
straight to a dataset.
- **REQ-132** — The rules in force when a merge is confirmed are **snapshotted onto the
dataset** (`datasets.rules_json`). The merge runs under the snapshot, and later edits to
the project's rules never rewrite an existing dataset. Only an explicit Resync adopts
today's rules — and it re-stamps the snapshot with them.
## F3. Counting correctness
- **REQ-140** — Ghost rejection (`entry_travel_min`) and spatial dedup
(`dedup_radius`) are separate parameters. They pull in opposite directions, so one
number cannot serve both.
- **REQ-141** — A track that vanishes parks its history; a new track id born within
`handoff_radius` of its velocity-projected position inherits it. This is what keeps an
ID switch at the counting line from either losing a count (the sack's "was above"
evidence dies with the old id) or duplicating one (the new id has no "already counted"
verdict).
- **REQ-142** — A counted direction is the track's *last* verdict, not a permanent one. A
sack genuinely taken back out and reloaded counts again; unloading requires
`unload_confirm_frames` sustained frames above the band, so repositioning by hand cannot
cancel a real count.
- **REQ-143** — Per-track state is evicted once a track has been gone for `track_ttl`, so a
long shift does not grow state without bound.
- **REQ-144** — Every finished track is written to a per-session JSONL with its trajectory
and the reason it did or did not count, so a miss can be attributed to the model, the
tracker, or the counter.
- **REQ-145** — Counting algorithms are **pluggable**. Each registers under a stable id
(`line_cross`, `possession`) and the session constructs one by id. The `Counter` protocol
in `src/interfaces.py` is the contract, corrected to match reality: `update()` returns the
frame's count events, not `None`. Adding an algorithm must not require editing
`live_count.py` or `counting_bench.py`.
- **REQ-146** — Each algorithm **declares its own parameters** — name, type, default, range —
and an endpoint serves that declaration, mirroring `live-count/models`. The frontend renders
its controls from the declaration and hardcodes no per-algorithm parameter list. The start
request carries `algorithm` plus an opaque `params` object validated against the
declaration, replacing today's flat line-specific fields.
- **REQ-147** — Geometry is generalised from a line to a **named shape set**. `line_cross`
declares one horizontal segment; `possession` declares a bed polygon and an approach zone.
The editor's drag channel (`move_line`) becomes shape-agnostic, so any algorithm's geometry
is adjustable live without a new endpoint.
- **REQ-148** — The **possession counter**: every sack track carries an `owner_id`, the person
track it currently overlaps, or none when at rest. A count fires on an ownership change that
crosses the bed boundary — person-outside to bed, or person-outside to person-inside.
Ownership is sticky with hysteresis, so occlusion by the carrier's back and the unowned
mid-air phase of a thrown sack do not break it. This requires a `person` class alongside
`sack` from the detector.
- **REQ-149** — Every count run records **which algorithm and parameter set** produced it, and
accuracy is comparable per algorithm against the same ground truth. Switching algorithms
adds results, it never invalidates stored ones — so `count_runs` is keyed by
`(project, video, algorithm)`, not by video alone.
## F4. Counting accuracy bench
- **REQ-150** — A page lists every archive video as a row: date, batch, length, and the
counter's `counted in` / `counted out` / `net` for it. Videos never counted are still
rows — the table is the work list.
- **REQ-151** — Each row has an editable **ground truth** (what a human counted). The
scored figure is the signed delta `counted_in - ground_truth`, so over- and under-counting
stay distinguishable. Accuracy is `1 - |delta| / ground_truth`.
- **REQ-152** — Accuracy totals are computed **only** over rows where a ground truth is
filled in. An uncounted or unscored video never enters the denominator.
- **REQ-153** — Counting runs as a queued background job over a selection of videos, or all
of them, holding the GPU lock. It renders nothing — no annotated frame, no JPEG encode —
which is what makes counting a 30-minute video practical. A run records the parameters and
model it used.
- **REQ-154** — Ground truth can be **imported in bulk** from the operations sheet
(`./GT.xlsx`, `DATA MUAT PAKAN PER LINE`). The camera watches **Line 1**; Line 2 is
recorded for completeness but never scored. Each sheet is one working day; a row is one
truck with a `BAG` count, a `DUS` count and a plate.
- **REQ-155** — `BAG` (sacks) and `DUS` (boxes) are **separate commodities**, counted and
scored separately. A box already resting in the truck bed is a legitimate object of a
different class, not a detection fault.
- **REQ-156** — An import never silently guesses. Recordings are aligned to sheet rows by
start time against row order, the proposed pairing is **shown for human confirmation**
before anything is written, and each imported value records that it came from the sheet
rather than from a hand count. A recording that merged two trucks
(`BATCH_MERGE_THRESHOLD_SECONDS`) is flagged, not paired.
- **REQ-157** — Sheet values are **order quantities, not hand counts** — 67% of them are
exactly 160 or 180 — so they score aggregate accuracy across many trucks and never
adjudicate a single video. Per-event truth for algorithm comparison comes from a
hand-counted clip, held separately.
## F5. Real recording times and working days
- **REQ-160** — Each recording's start time is read from the timestamp the camera burns into
the top-right of every frame. Folder names and file mtimes are both unreliable: mtimes are
file *copy* times, not recording times.
- **REQ-161** — A **working day runs 06:00 to 06:00**. A recording that started before 06:00
belongs to the previous working day. A recording spanning the boundary is assigned by its
start.
- **REQ-162** — Recordings are numbered 1..N within a working day, ordered by real start
time. The original file path stays the file's identity and is shown beside the number, so
results already recorded against it survive.
- **REQ-164** — The counting table is grouped into collapsible **cycles**, newest first, with
the recordings inside each one in the order they were made. A cycle's header carries its
video count, its AI/ground-truth totals, signed delta and accuracy, and how many of its rows
have an unverified start time. Only the newest cycle is expanded by default.
- **REQ-165** — The Video Archive page browses the archive **by cycle**, not by folder. The
left-hand list holds cycles newest first; the table shows the recordings of the selected
cycle in the order they were made, with their cycle batch number and the time read from the
overlay. A recording pulled in from another folder is marked with the folder it sits in.
- **REQ-166** — Each recording is checked for a truck with the project's newest model,
sampling a handful of frames rather than the whole file. The recording trigger is truck
arrival and departure, so one file is one batch — this check is what proves that assumption
per file, and flags any recording where it does not hold.
- **REQ-167** — The production counter's counting day turns over at the same hour as the
archive's cycles, 06:00, so `batch_number` on the Jetson and the batch order in Video
Archive mean the same thing. It stays overridable per deployment via `DAILY_CUTOFF_TIME`.
- **REQ-168** — The recorder writes each archive file at the frame rate the stream actually
delivers, and paces writes against the wall clock, so a file's duration equals the real
duration of the recording however unevenly the capture loop runs.
- **REQ-170** — Recording happens once, on the streaming server, and is stored as a rolling
buffer. The truck detector does not encode video: when a session ends it downloads that time
range as a **copy**, so the archive keeps the camera's own codec, resolution and frame rate.
Each clip carries a sidecar with the server's start time, which the app trusts over reading
the burned-in overlay. Archive folder names remain calendar dates; the app derives cycles
from the real start time, so the folder name is never read as a date.
- **REQ-163** — Nothing in the archive is moved, renamed or written to; it is mounted
read-only. The grouping lives in an index beside it. Timestamps that could not be read, or
were read with low confidence, are flagged and can be hand-entered; a hand-entered time
outranks any reading and is never overwritten by a rescan.
## G. Training & evaluation
- **REQ-060** — The user starts training from the project page. Training **fine-tunes from
the project's base model** on the merged master dataset (old + new).
- **REQ-061** — A fallback option "train on the latest batch only" (lower LR, fewer epochs)
exists for cases where the old dataset is unavailable. It is not the default, and the UI
warns about catastrophic forgetting.
- **REQ-062** — Default `batch`, `imgsz`, and `device` are derived from the hardware detected
at runtime (VRAM), and all of them can be overridden — so moving to a bigger machine needs
no code change.
- **REQ-063** — After training, the system validates **the base model and the new model on
the exact same val set**, then shows mAP50 and mAP50-95 for both side by side with the
delta.
- **REQ-064** — Each training run produces a stored model version (weights + metrics). The
user can download the weights and **promote that version to be the project's new base
model** for the next round.
- **REQ-065** — SAM3 and training must never hold VRAM at the same time; the system releases
the SAM3 model before training starts.
## H. System
- **REQ-070** — Heavy work (extraction, auto-annotation, training) runs as queued jobs, one
at a time, because there is a single GPU. Jobs show progress and logs, and can be cancelled.
- **REQ-071** — Jobs and their progress are persistent; after a server restart the job list
is still there with its final statuses.
- **REQ-072** — The application runs via `docker compose up` with GPU access, and every path
(video archive, data folder) is configured through environment/volumes — never hardcoded.
- **REQ-073** — The health endpoint reports: detected device/GPU, ffmpeg availability,
whether the HuggingFace token was picked up, and database reachability.
- **REQ-074** — The system never writes anything into the user's video archive folder.