354 lines
23 KiB
Markdown
354 lines
23 KiB
Markdown
# Requirements
|
|
|
|
Status: **agreed** (planning session, 2026-07-31). Changes only with the user's approval.
|
|
|
|
## Goal
|
|
|
|
A system for **enriching a dataset and improving an existing detection model**, iteratively,
|
|
from an archive of recorded video. One full round:
|
|
|
|
> pick a project → browse the video archive → pick a batch → trim a time range →
|
|
> extract frames → auto-annotate with SAM3 → review and correct every frame → approve →
|
|
> merge into the master dataset → fine-tune from the base model → compare against the base.
|
|
|
|
The system is **generic**: the sack case is only the first project. Other cases are
|
|
created as new projects with their own base model, classes, and video archive — no code
|
|
changes.
|
|
|
|
## Non-goals (for this version)
|
|
|
|
- Login, multi-user, tenants, quotas. The architecture leaves room for them; the features
|
|
are not built.
|
|
- Tracking or annotation propagation between frames.
|
|
- Collaborative annotation by several people at once.
|
|
- Public internet deployment.
|
|
|
|
---
|
|
|
|
## A. Project
|
|
|
|
- **REQ-001** — The user can create, list, and delete projects. A project has: name, label
|
|
type, base model, video archive root, and a class list.
|
|
- **REQ-002** — Each project picks a **label type**: `bbox` (YOLO detect) or `polygon`
|
|
(YOLO segment). This determines the export format, the editor's behaviour, and which
|
|
model variant is trained. It cannot be changed once a batch has been merged.
|
|
- **REQ-003** — The user uploads a **base model** `.pt`. The class list is read from the
|
|
model (`model.names`). It cannot drift on its own: nothing adds or removes a class as a
|
|
side effect of another action. Deliberate deletion is REQ-007.
|
|
- **REQ-004** — A project may be created **without** a base model. In that case the user
|
|
types the class list, and the first training starts from pretrained weights
|
|
(`yolo11n.pt` / `yolo11n-seg.pt`).
|
|
- **REQ-005** — Each class carries its own **SAM3 text prompt**, which may differ from the
|
|
class name (e.g. class `sack` with prompt `"woven plastic sack"`). Prompts can be
|
|
edited at any time without affecting existing data.
|
|
- **REQ-006** — All of a project's data (base model, master dataset, batch frames, trained
|
|
weights) lives under one project folder, so it can be copied or backed up whole.
|
|
- **REQ-007** — The user can **delete a class** at any point in a project's life, including
|
|
after batches have been merged. Deleting one:
|
|
- removes every annotation of that class, in batches under review and in the master
|
|
dataset alike;
|
|
- **renumbers the classes above it**, in the database *and* in every label file already
|
|
written to disk, because a YOLO label is an integer index and leaving a gap would make
|
|
old labels silently name the wrong class;
|
|
- regenerates `data.yaml`;
|
|
- is refused for a project's last remaining class.
|
|
|
|
Before confirming, the user is told how many shapes will be destroyed. The action cannot
|
|
be undone. If the project's base model was trained on the old class list, it stops being
|
|
comparable — which REQ-063 already reports rather than hides.
|
|
- **REQ-008** — The user can **add a new class** (name & prompt) to an existing project at any time.
|
|
The new class receives the next sequential `class_id`, and `data.yaml` is regenerated if a
|
|
master dataset exists.
|
|
|
|
## B. Video archive
|
|
|
|
- **REQ-010** — The video archive lives at a local path (disk or mount); videos are **not
|
|
uploaded** through the browser.
|
|
- **REQ-011** — Archive structure: `<video_root>/<date>/<batch>.<ext>`. The system lists the
|
|
dates, and within each date the videos with their batch labels parsed from the filename.
|
|
- **REQ-012** — Each video shows its duration, resolution, and whether it has already been
|
|
used as a batch in this project.
|
|
- **REQ-013** — Videos play in the browser with seeking (HTTP Range), without copying the
|
|
file first.
|
|
|
|
## C. Trim & frame extraction
|
|
|
|
- **REQ-020** — The user sets the in/out range with a timeline slider on the player, and can
|
|
also type precise timestamps.
|
|
- **REQ-021** — The user sets the extraction **frames per second** (default 1 fps). The
|
|
resulting frame count is shown before extraction runs.
|
|
- **REQ-022** — Extraction runs as a background job with progress, producing sequentially
|
|
numbered JPEG files inside the batch folder.
|
|
- **REQ-023** — One video may be used more than once with different time ranges; each
|
|
extraction produces its own batch.
|
|
|
|
## D. Auto-annotation
|
|
|
|
- **REQ-030** — Once frames are extracted, the system runs SAM3 over all of them using each
|
|
class's prompt, as a background job with progress and cancellation.
|
|
- **REQ-031** — Detections that overlap across prompts are deduplicated (greedy IoU NMS), so
|
|
one object is not labelled as two classes at once.
|
|
- **REQ-032** — The confidence threshold is configurable per job.
|
|
- **REQ-033** — A frame with no detections is valid and still enters the dataset as a
|
|
negative sample — it is not a failure.
|
|
- **REQ-034** — Auto-annotation can be re-run on the same batch; previous automatic results
|
|
are replaced, but **the user's manual corrections must never be lost**.
|
|
- **REQ-035** — Auto-annotation can be started in **resume** mode, which skips frames that
|
|
already carry automatic annotations. Resume is always an explicit choice and never the
|
|
default, because a full re-run is also how the confidence threshold (REQ-032) is changed —
|
|
the system cannot tell the two intentions apart, so it asks. A frame SAM3 legitimately
|
|
found nothing on (REQ-033) writes no annotations, so a resume re-does it; that is accepted
|
|
rather than tracked.
|
|
|
|
- **REQ-171** — In the auto-annotate modal the SAM3 text prompt of each selected class is
|
|
editable in place, next to the live preview. Saving it writes
|
|
`project_classes.prompt` — the same field the Projects page edits — so the batch job and
|
|
every later run send that text. The preview is the tuning surface; the stored prompt is the
|
|
artifact it produces.
|
|
- **REQ-172** — On the previewed frame the user can drag **positive** and **negative** box
|
|
exemplars (shift-drag for negative). They are appended to the active class's text prompt,
|
|
re-run immediately, and can be undone or cleared. Exemplars are a **tuning aid only**: they
|
|
are never written as annotations and never carried into the batch job, because SAM3's
|
|
geometric prompts pool features from the current image — replaying them on another frame
|
|
would ask about whatever happens to sit at those coordinates there. They belong to exactly
|
|
one class, so a new frame or a new active class discards them.
|
|
|
|
## E. Review & correction
|
|
|
|
- **REQ-040** — The user reviews frames one at a time, with fast navigation (left/right
|
|
arrows, thumbnail filmstrip, jump to the next unreviewed frame).
|
|
- **REQ-041** — Each frame has a status: `pending`, `approved`, or `rejected`. Rejected
|
|
frames never enter the dataset.
|
|
- **REQ-042** — The user can draw a new shape, move it, resize it, delete it, and change its
|
|
class.
|
|
- **REQ-043** — The user can ask SAM3 for help inside the editor: click or drag a box around
|
|
one object and the model produces its shape.
|
|
- **REQ-044** — All annotations and review statuses are **persistent** — they survive a
|
|
server restart, unlike today's in-memory sessions.
|
|
- **REQ-045** — Review progress is visible (e.g. "120/300 reviewed"), and a batch can only
|
|
be approved once no frame is still `pending`.
|
|
- **REQ-046** — The user can delete/clear all annotations of a specific class across all frames in
|
|
the current batch from the Review editor.
|
|
|
|
- **REQ-173** — In the review editor a plain drag on the canvas is an **exemplar-driven
|
|
label**, not just a rectangle. It proposes, in one action — and REQ-175's Apply is what
|
|
makes any of it real — that (a) the drawn shape becomes a `manual` annotation of the
|
|
active class — snapped to a SAM3 polygon first when the batch's
|
|
`label_type` is `polygon`, since a rectangle is a bad polygon label — (b) appends the box
|
|
to the frame's positive exemplar pool for that class, and (c) re-runs SAM3 over the whole
|
|
frame with the class's text prompt plus the pooled exemplars, deleting every existing shape
|
|
of that class on the frame and writing the detections in its place, then re-inserting the
|
|
pooled exemplar shapes verbatim so the user's own drawings always survive. The pool is
|
|
**frame-local and ephemeral** for the same reason as REQ-172 — SAM3's geometric prompts pool
|
|
features from the current image — so leaving the frame or switching the active class clears
|
|
it; the annotations it produced persist like any other. If the GPU lock (REQ-070) is not
|
|
free, the run comes back with the drawn shapes alone and says so, so the user can still
|
|
file them (REQ-175) and labeling is never blocked by a background job.
|
|
- **REQ-174** — **Shift**-drag in the review editor adds a **negative** exemplar. It is never
|
|
stored as an annotation; it deletes any existing shape of the active class that overlaps it,
|
|
and it is sent as a negative box in the REQ-173 re-detect. It is the "not this, and not
|
|
things like this" gesture, so it doubles as a delete. A negative is **spent on Apply**: the
|
|
frame it was applied to no longer carries what it rejected, so the drawing is dropped from
|
|
the pool while the positives stay on as prompts.
|
|
- **REQ-175** — An exemplar drag **previews**; it never writes on its own. The run's result
|
|
is drawn over the frame as proposals and a small panel floats on the canvas with the four
|
|
filters that decide what survives — confidence, NMS overlap, minimum box size, maximum
|
|
shapes — each re-running the preview as it moves. **Apply** writes the previewed set,
|
|
**Discard** rewinds the pool to whatever is already on the frame and leaves it untouched.
|
|
The panel is scoped to this gesture: its values are not stored, not shared with the
|
|
auto-annotate modal, and reset with the frame. Defaults are confidence `0.5`, NMS `0.8`,
|
|
min box `0.002`, max `100` — deliberately permissive, because on a dense frame an
|
|
aggressive NMS or area floor deletes real, touching objects rather than duplicates.
|
|
|
|
## E4. Live counting preview
|
|
|
|
- **REQ-176** — A **live** source on the Live Count page is a **WebRTC (WHEP) URL** and
|
|
nothing else; an RTSP URL is rejected with a message saying so. The backend derives the
|
|
RTSP leg of the same streaming-server path from it (`http://host:8889/cam` →
|
|
`rtsp://host:8554/cam`) and counts from that: WebRTC is what makes the browser preview
|
|
cheap, but pulling it into Python would add ICE and a jitter buffer on top of the identical
|
|
H.264 decode. One ingest on the streaming server, two consumers. The ports are read from
|
|
the environment (`MEDIAMTX_RTSP_PORT`, `MEDIAMTX_WHEP_PATH`), never hardcoded. Archive
|
|
files are unaffected — they are still opened as files.
|
|
- **REQ-177** — A live session is **watched over WebRTC**, played straight from the streaming
|
|
server by the browser: the frames never pass through this app and it encodes no JPEG for
|
|
them. What the model saw — boxes, ids, confidences, the counting line and its band, the
|
|
ignored region, the running totals — is served as geometry from
|
|
`GET /api/live-count/overlay` and drawn on a canvas over the video. The MJPEG endpoint
|
|
remains the preview for **archive files** only, and refuses a WebRTC session.
|
|
|
|
## F. Master dataset
|
|
|
|
- **REQ-050** — Approving a batch **merges** its approved frames and their labels into the
|
|
project's master dataset (accumulating across batches).
|
|
- **REQ-051** — The master dataset is train-ready YOLO format: `images/{train,val}`,
|
|
`labels/{train,val}`, and a `data.yaml` regenerated from the project's class list.
|
|
- **REQ-052** — **Stable val split**: once a frame is placed in `val`, it stays in `val`
|
|
across every later merge. New frames are split with an every-Nth pattern.
|
|
- **REQ-053** — The system records which batches have entered the master dataset, when, and
|
|
how many images/labels each added.
|
|
- **REQ-054** — The master dataset can be downloaded as a `.zip` (e.g. to import into
|
|
Roboflow or train on another machine).
|
|
|
|
## F2. Data Prep as the merge gate
|
|
|
|
- **REQ-130** — The Batches page supports multi-select. "Prepare & Merge Selected" opens
|
|
Data Prep scoped to exactly those batches (`#/projects/{id}/data-prep?batches=1,2,3`).
|
|
No dataset exists at this point.
|
|
- **REQ-131** — Data Prep is the merge gate. Filters and augmentation are tuned against the
|
|
selected batches' shapes; "Confirm merge" names or picks the target dataset and queues
|
|
**one** merge job for the whole selection. There is no path from Batches or Review
|
|
straight to a dataset.
|
|
- **REQ-132** — The rules in force when a merge is confirmed are **snapshotted onto the
|
|
dataset** (`datasets.rules_json`). The merge runs under the snapshot, and later edits to
|
|
the project's rules never rewrite an existing dataset. Only an explicit Resync adopts
|
|
today's rules — and it re-stamps the snapshot with them.
|
|
|
|
## F3. Counting correctness
|
|
|
|
- **REQ-140** — Ghost rejection (`entry_travel_min`) and spatial dedup
|
|
(`dedup_radius`) are separate parameters. They pull in opposite directions, so one
|
|
number cannot serve both.
|
|
- **REQ-141** — A track that vanishes parks its history; a new track id born within
|
|
`handoff_radius` of its velocity-projected position inherits it. This is what keeps an
|
|
ID switch at the counting line from either losing a count (the sack's "was above"
|
|
evidence dies with the old id) or duplicating one (the new id has no "already counted"
|
|
verdict).
|
|
- **REQ-142** — A counted direction is the track's *last* verdict, not a permanent one. A
|
|
sack genuinely taken back out and reloaded counts again; unloading requires
|
|
`unload_confirm_frames` sustained frames above the band, so repositioning by hand cannot
|
|
cancel a real count.
|
|
- **REQ-143** — Per-track state is evicted once a track has been gone for `track_ttl`, so a
|
|
long shift does not grow state without bound.
|
|
- **REQ-144** — Every finished track is written to a per-session JSONL with its trajectory
|
|
and the reason it did or did not count, so a miss can be attributed to the model, the
|
|
tracker, or the counter.
|
|
|
|
- **REQ-145** — Counting algorithms are **pluggable**. Each registers under a stable id
|
|
(`line_cross`, `possession`) and the session constructs one by id. The `Counter` protocol
|
|
in `src/interfaces.py` is the contract, corrected to match reality: `update()` returns the
|
|
frame's count events, not `None`. Adding an algorithm must not require editing
|
|
`live_count.py` or `counting_bench.py`.
|
|
- **REQ-146** — Each algorithm **declares its own parameters** — name, type, default, range —
|
|
and an endpoint serves that declaration, mirroring `live-count/models`. The frontend renders
|
|
its controls from the declaration and hardcodes no per-algorithm parameter list. The start
|
|
request carries `algorithm` plus an opaque `params` object validated against the
|
|
declaration, replacing today's flat line-specific fields.
|
|
- **REQ-147** — Geometry is generalised from a line to a **named shape set**. `line_cross`
|
|
declares one horizontal segment; `possession` declares a bed polygon and an approach zone.
|
|
The editor's drag channel (`move_line`) becomes shape-agnostic, so any algorithm's geometry
|
|
is adjustable live without a new endpoint.
|
|
- **REQ-148** — The **possession counter**: every sack track carries an `owner_id`, the person
|
|
track it currently overlaps, or none when at rest. A count fires on an ownership change that
|
|
crosses the bed boundary — person-outside to bed, or person-outside to person-inside.
|
|
Ownership is sticky with hysteresis, so occlusion by the carrier's back and the unowned
|
|
mid-air phase of a thrown sack do not break it. This requires a `person` class alongside
|
|
`sack` from the detector.
|
|
- **REQ-149** — Every count run records **which algorithm and parameter set** produced it, and
|
|
accuracy is comparable per algorithm against the same ground truth. Switching algorithms
|
|
adds results, it never invalidates stored ones — so `count_runs` is keyed by
|
|
`(project, video, algorithm)`, not by video alone.
|
|
|
|
## F4. Counting accuracy bench
|
|
|
|
- **REQ-150** — A page lists every archive video as a row: date, batch, length, and the
|
|
counter's `counted in` / `counted out` / `net` for it. Videos never counted are still
|
|
rows — the table is the work list.
|
|
- **REQ-151** — Each row has an editable **ground truth** (what a human counted). The
|
|
scored figure is the signed delta `counted_in - ground_truth`, so over- and under-counting
|
|
stay distinguishable. Accuracy is `1 - |delta| / ground_truth`.
|
|
- **REQ-152** — Accuracy totals are computed **only** over rows where a ground truth is
|
|
filled in. An uncounted or unscored video never enters the denominator.
|
|
- **REQ-153** — Counting runs as a queued background job over a selection of videos, or all
|
|
of them, holding the GPU lock. It renders nothing — no annotated frame, no JPEG encode —
|
|
which is what makes counting a 30-minute video practical. A run records the parameters and
|
|
model it used.
|
|
|
|
- **REQ-154** — Ground truth can be **imported in bulk** from the operations sheet
|
|
(`./GT.xlsx`, `DATA MUAT PAKAN PER LINE`). The camera watches **Line 1**; Line 2 is
|
|
recorded for completeness but never scored. Each sheet is one working day; a row is one
|
|
truck with a `BAG` count, a `DUS` count and a plate.
|
|
- **REQ-155** — `BAG` (sacks) and `DUS` (boxes) are **separate commodities**, counted and
|
|
scored separately. A box already resting in the truck bed is a legitimate object of a
|
|
different class, not a detection fault.
|
|
- **REQ-156** — An import never silently guesses. Recordings are aligned to sheet rows by
|
|
start time against row order, the proposed pairing is **shown for human confirmation**
|
|
before anything is written, and each imported value records that it came from the sheet
|
|
rather than from a hand count. A recording that merged two trucks
|
|
(`BATCH_MERGE_THRESHOLD_SECONDS`) is flagged, not paired.
|
|
- **REQ-157** — Sheet values are **order quantities, not hand counts** — 67% of them are
|
|
exactly 160 or 180 — so they score aggregate accuracy across many trucks and never
|
|
adjudicate a single video. Per-event truth for algorithm comparison comes from a
|
|
hand-counted clip, held separately.
|
|
|
|
## F5. Real recording times and working days
|
|
|
|
- **REQ-160** — Each recording's start time is read from the timestamp the camera burns into
|
|
the top-right of every frame. Folder names and file mtimes are both unreliable: mtimes are
|
|
file *copy* times, not recording times.
|
|
- **REQ-161** — A **working day runs 06:00 to 06:00**. A recording that started before 06:00
|
|
belongs to the previous working day. A recording spanning the boundary is assigned by its
|
|
start.
|
|
- **REQ-162** — Recordings are numbered 1..N within a working day, ordered by real start
|
|
time. The original file path stays the file's identity and is shown beside the number, so
|
|
results already recorded against it survive.
|
|
- **REQ-164** — The counting table is grouped into collapsible **cycles**, newest first, with
|
|
the recordings inside each one in the order they were made. A cycle's header carries its
|
|
video count, its AI/ground-truth totals, signed delta and accuracy, and how many of its rows
|
|
have an unverified start time. Only the newest cycle is expanded by default.
|
|
- **REQ-165** — The Video Archive page browses the archive **by cycle**, not by folder. The
|
|
left-hand list holds cycles newest first; the table shows the recordings of the selected
|
|
cycle in the order they were made, with their cycle batch number and the time read from the
|
|
overlay. A recording pulled in from another folder is marked with the folder it sits in.
|
|
- **REQ-166** — Each recording is checked for a truck with the project's newest model,
|
|
sampling a handful of frames rather than the whole file. The recording trigger is truck
|
|
arrival and departure, so one file is one batch — this check is what proves that assumption
|
|
per file, and flags any recording where it does not hold.
|
|
- **REQ-167** — The production counter's counting day turns over at the same hour as the
|
|
archive's cycles, 06:00, so `batch_number` on the Jetson and the batch order in Video
|
|
Archive mean the same thing. It stays overridable per deployment via `DAILY_CUTOFF_TIME`.
|
|
- **REQ-168** — The recorder writes each archive file at the frame rate the stream actually
|
|
delivers, and paces writes against the wall clock, so a file's duration equals the real
|
|
duration of the recording however unevenly the capture loop runs.
|
|
- **REQ-170** — Recording happens once, on the streaming server, and is stored as a rolling
|
|
buffer. The truck detector does not encode video: when a session ends it downloads that time
|
|
range as a **copy**, so the archive keeps the camera's own codec, resolution and frame rate.
|
|
Each clip carries a sidecar with the server's start time, which the app trusts over reading
|
|
the burned-in overlay. Archive folder names remain calendar dates; the app derives cycles
|
|
from the real start time, so the folder name is never read as a date.
|
|
- **REQ-163** — Nothing in the archive is moved, renamed or written to; it is mounted
|
|
read-only. The grouping lives in an index beside it. Timestamps that could not be read, or
|
|
were read with low confidence, are flagged and can be hand-entered; a hand-entered time
|
|
outranks any reading and is never overwritten by a rescan.
|
|
|
|
## G. Training & evaluation
|
|
|
|
- **REQ-060** — The user starts training from the project page. Training **fine-tunes from
|
|
the project's base model** on the merged master dataset (old + new).
|
|
- **REQ-061** — A fallback option "train on the latest batch only" (lower LR, fewer epochs)
|
|
exists for cases where the old dataset is unavailable. It is not the default, and the UI
|
|
warns about catastrophic forgetting.
|
|
- **REQ-062** — Default `batch`, `imgsz`, and `device` are derived from the hardware detected
|
|
at runtime (VRAM), and all of them can be overridden — so moving to a bigger machine needs
|
|
no code change.
|
|
- **REQ-063** — After training, the system validates **the base model and the new model on
|
|
the exact same val set**, then shows mAP50 and mAP50-95 for both side by side with the
|
|
delta.
|
|
- **REQ-064** — Each training run produces a stored model version (weights + metrics). The
|
|
user can download the weights and **promote that version to be the project's new base
|
|
model** for the next round.
|
|
- **REQ-065** — SAM3 and training must never hold VRAM at the same time; the system releases
|
|
the SAM3 model before training starts.
|
|
|
|
## H. System
|
|
|
|
- **REQ-070** — Heavy work (extraction, auto-annotation, training) runs as queued jobs, one
|
|
at a time, because there is a single GPU. Jobs show progress and logs, and can be cancelled.
|
|
- **REQ-071** — Jobs and their progress are persistent; after a server restart the job list
|
|
is still there with its final statuses.
|
|
- **REQ-072** — The application runs via `docker compose up` with GPU access, and every path
|
|
(video archive, data folder) is configured through environment/volumes — never hardcoded.
|
|
- **REQ-073** — The health endpoint reports: detected device/GPU, ffmpeg availability,
|
|
whether the HuggingFace token was picked up, and database reachability.
|
|
- **REQ-074** — The system never writes anything into the user's video archive folder.
|