feat: add counting bench, triage, and dataset modules

This commit includes major additions and updates to the frontend and backend architectures, introducing new dataset management, live counting features, batch processing, and triage logic. Includes new UI pages, components, and API routes.
This commit is contained in:
asus committed 2026-08-14 16:28:52 +07:00
1 parent 8285400254
commit 5c7c122105
80 files changed
+20074 -1412

No files matched your search

+89
View File
@@ -130,6 +130,95 @@ changes.
- **REQ-054** — The master dataset can be downloaded as a `.zip` (e.g. to import into
Roboflow or train on another machine).
## F2. Data Prep as the merge gate
- **REQ-130** — The Batches page supports multi-select. "Prepare & Merge Selected" opens
Data Prep scoped to exactly those batches (`#/projects/{id}/data-prep?batches=1,2,3`).
No dataset exists at this point.
- **REQ-131** — Data Prep is the merge gate. Filters and augmentation are tuned against the
selected batches' shapes; "Confirm merge" names or picks the target dataset and queues
**one** merge job for the whole selection. There is no path from Batches or Review
straight to a dataset.
- **REQ-132** — The rules in force when a merge is confirmed are **snapshotted onto the
dataset** (`datasets.rules_json`). The merge runs under the snapshot, and later edits to
the project's rules never rewrite an existing dataset. Only an explicit Resync adopts
today's rules — and it re-stamps the snapshot with them.
## F3. Counting correctness
- **REQ-140** — Ghost rejection (`entry_travel_min`) and spatial dedup
(`dedup_radius`) are separate parameters. They pull in opposite directions, so one
number cannot serve both.
- **REQ-141** — A track that vanishes parks its history; a new track id born within
`handoff_radius` of its velocity-projected position inherits it. This is what keeps an
ID switch at the counting line from either losing a count (the sack's "was above"
evidence dies with the old id) or duplicating one (the new id has no "already counted"
verdict).
- **REQ-142** — A counted direction is the track's *last* verdict, not a permanent one. A
sack genuinely taken back out and reloaded counts again; unloading requires
`unload_confirm_frames` sustained frames above the band, so repositioning by hand cannot
cancel a real count.
- **REQ-143** — Per-track state is evicted once a track has been gone for `track_ttl`, so a
long shift does not grow state without bound.
- **REQ-144** — Every finished track is written to a per-session JSONL with its trajectory
and the reason it did or did not count, so a miss can be attributed to the model, the
tracker, or the counter.
## F4. Counting accuracy bench
- **REQ-150** — A page lists every archive video as a row: date, batch, length, and the
counter's `counted in` / `counted out` / `net` for it. Videos never counted are still
rows — the table is the work list.
- **REQ-151** — Each row has an editable **ground truth** (what a human counted). The
scored figure is the signed delta `counted_in - ground_truth`, so over- and under-counting
stay distinguishable. Accuracy is `1 - |delta| / ground_truth`.
- **REQ-152** — Accuracy totals are computed **only** over rows where a ground truth is
filled in. An uncounted or unscored video never enters the denominator.
- **REQ-153** — Counting runs as a queued background job over a selection of videos, or all
of them, holding the GPU lock. It renders nothing — no annotated frame, no JPEG encode —
which is what makes counting a 30-minute video practical. A run records the parameters and
model it used.
## F5. Real recording times and working days
- **REQ-160** — Each recording's start time is read from the timestamp the camera burns into
the top-right of every frame. Folder names and file mtimes are both unreliable: mtimes are
file *copy* times, not recording times.
- **REQ-161** — A **working day runs 06:00 to 06:00**. A recording that started before 06:00
belongs to the previous working day. A recording spanning the boundary is assigned by its
start.
- **REQ-162** — Recordings are numbered 1..N within a working day, ordered by real start
time. The original file path stays the file's identity and is shown beside the number, so
results already recorded against it survive.
- **REQ-164** — The counting table is grouped into collapsible **cycles**, newest first, with
the recordings inside each one in the order they were made. A cycle's header carries its
video count, its AI/ground-truth totals, signed delta and accuracy, and how many of its rows
have an unverified start time. Only the newest cycle is expanded by default.
- **REQ-165** — The Video Archive page browses the archive **by cycle**, not by folder. The
left-hand list holds cycles newest first; the table shows the recordings of the selected
cycle in the order they were made, with their cycle batch number and the time read from the
overlay. A recording pulled in from another folder is marked with the folder it sits in.
- **REQ-166** — Each recording is checked for a truck with the project's newest model,
sampling a handful of frames rather than the whole file. The recording trigger is truck
arrival and departure, so one file is one batch — this check is what proves that assumption
per file, and flags any recording where it does not hold.
- **REQ-167** — The production counter's counting day turns over at the same hour as the
archive's cycles, 06:00, so `batch_number` on the Jetson and the batch order in Video
Archive mean the same thing. It stays overridable per deployment via `DAILY_CUTOFF_TIME`.
- **REQ-168** — The recorder writes each archive file at the frame rate the stream actually
delivers, and paces writes against the wall clock, so a file's duration equals the real
duration of the recording however unevenly the capture loop runs.
- **REQ-170** — Recording happens once, on the streaming server, and is stored as a rolling
buffer. The truck detector does not encode video: when a session ends it downloads that time
range as a **copy**, so the archive keeps the camera's own codec, resolution and frame rate.
Each clip carries a sidecar with the server's start time, which the app trusts over reading
the burned-in overlay. Archive folder names remain calendar dates; the app derives cycles
from the real start time, so the folder name is never read as a date.
- **REQ-163** — Nothing in the archive is moved, renamed or written to; it is mounted
read-only. The grouping lives in an index beside it. Timestamps that could not be read, or
were read with low confidence, are flagged and can be hand-entered; a hand-entered time
outranks any reading and is never overwritten by a rescan.
## G. Training & evaluation
- **REQ-060** — The user starts training from the project page. Training **fine-tunes from