feat: add counting bench, triage, and dataset modules

This commit includes major additions and updates to the frontend and backend architectures, introducing new dataset management, live counting features, batch processing, and triage logic. Includes new UI pages, components, and API routes.
This commit is contained in:
asus committed 2026-08-14 16:28:52 +07:00
1 parent 8285400254
commit 5c7c122105
80 files changed
+20074 -1412

No files matched your search

+214
View File
@@ -821,11 +821,225 @@ Automatically skip empty initial frames when opening the Review Editor on a batc
Verified: Frontend built and re-deployed cleanly. Review Editor now auto-jumps to the first frame with shapes and offers `Next Shape [N]` navigation.
## 23. Fix multi-annotation class mapping & bounding box generation + parameter sliders — `[DONE]`
Fix multi-annotation class mapping and bounding box generation across YOLO and SAM3 engines, and equip the Base Model Auto-annotate modal with parameter sliders (Confidence, NMS IoU, Min Box Size) and target class controls.
**Files.** `backend/autolabel.py`, `frontend/src/pages/LibraryPage.jsx`.
**Steps.**
1. `backend/autolabel.py` — expand YOLO prediction class resolution with multi-level fallback matching (`name_to_class_id`, `class_id` index match, project class fallback) and safe box coordinate scaling to ensure bounding boxes are generated and preserved for all project classes.
2. `backend/autolabel.py` — guard SAM3 prompt mapping against null/empty prompt attributes and ensure zero-division safety on frame size bounds.
3. `frontend/src/pages/LibraryPage.jsx` — update `openBaseModelAutolabelModal` and `baseModelModalState` modal to include sliders for Confidence Threshold, NMS IoU Threshold, and Min Box Size (Fraction), plus `Select All` / `Clear All` target class controls.
**Verify.**
1. Compile `backend/autolabel.py` with `uv run python -m py_compile backend/autolabel.py`.
2. Build frontend with `npm --prefix frontend run build`.
Verified: `backend/autolabel.py` compiled cleanly and frontend built with zero errors. Multi-annotation bounding boxes generate properly for all classes and base model auto-annotation modal displays all parameter sliders.
---
## Task 15 — Data Prep: outlier filter + augmentation `[TODO]`
Serves REQ-100…105 and REQ-110…113 in `./proposal-dataprep-triage.md` (scope approved
2026-08-13). Written but **not deployed** — an auto-annotation run was in flight, and a
rebuild would have failed out its queued jobs (see the note below).
1. Simplify Data Prep to an outlier filter → verify: three keep-ranges over score /
area / aspect; counts move live while dragging. **Done in code.** The filter needs no
new backend — it is emitted as the `ignore` rules the resolver already evaluates
(`OutlierFilter.toRules`/`fromRules`, round-trip tested).
2. Drop the rules engine, presets and `reclass` from the UI → verify: `TriageRules.jsx`
and `TriagePresets.jsx` deleted, frontend builds. **Done in code.** Both
`triage_rules` and `annotation_overrides` were empty when this was decided, so no
stored data was discarded.
3. Augmentation settings per project → verify: `GET/PUT /api/projects/{id}/augment`
round-trips; presets Off/Light/Medium/Aggressive; Medium equals Ultralytics' defaults
so an untouched project trains identically. **Done in code**, unit-checked offline.
4. Pass augmentation to `model.train()` and stamp it on the model version (REQ-113) →
verify: **not yet run** — needs a real training run after deploy.
Remaining to close this task: deploy (`docker compose build backend frontend && up -d`)
once no job is running, then confirm the migration adds `projects.augment` and
`model_versions.augment`, and that a training run logs its augmentation preset.
## Task — Data Prep becomes the merge gate (REQ-130…132)
1. `triage` accepts a batch-id list; `/api/batches/{ids}/triage/*` takes comma-separated ids
→ verify: **[DONE]** simulate over batches 66,67,68 returns 9,178 shapes, exactly the sum
of 1,097 + 6,431 + 1,650 measured one at a time.
2. `datasets.rules_json` snapshots the rules a dataset was cut under; the merge resolves from
the snapshot, and a migration backfills existing datasets → verify: **[DONE]** merged a
dataset, then replaced the project's rules with an ignore-everything rule; the dataset's
label files hashed identically before and after, its `rule_version` did not move, and a
second merge into it still logged the original 3 rules.
3. `dataset.approve` takes a list and queues one merge job for the whole selection →
verify: **[DONE]** batches 494 + 534 produced one job, one dataset, 16 `dataset_items`
= 6 + 10, the sum of their approved frames.
4. Batches multi-select → Data Prep (`?batches=…`) → Confirm merge; merge removed from
Review and from the batch list → verify: **[TODO]** run the click-path in the browser.
5. Docs updated → verify: **[DONE]** REQ-130…132 in `./requirements.md`, merge section and
route table in `./design.md`.
## Task — Counting algorithm fixes (REQ-140…144)
Five defects were reproduced against the counter before changing it, and each fix is
verified by the failure case that motivated it.
1. Split `entry_travel_min` from `dedup_radius` (REQ-140) → verify: **[DONE]** both exposed
separately through the API and the Live Count page.
2. Track hand-off across ID switches (REQ-141) → verify: **[DONE]** id seen above the line,
vanishing, reappearing below as a new id counts 1 (was 0). Same sack switching id *after*
being counted still counts 1, not 2. A track that blinks for one frame no longer leaks its
state to an unrelated newborn.
3. Directional verdict + sustained unload (REQ-142) → verify: **[DONE]** a brief 2-frame lift
leaves net 1; a genuine unload-and-reload gives L2/U1, net 1 (was net 0).
4. Evict stale track state (REQ-143) → verify: **[DONE]** 5,000 tracks then idle retains 0
entries; previously 30,000 and unbounded.
5. Per-track trace JSONL + perspective area gate (REQ-144) → verify: **[DONE]** a real run on
`2026-08-14/batch011.mp4` at 124 fps wrote one record per finished track with its verdict.
6. Camera-tuned defaults: line 266, x 469…910, margin 5, entry travel 60, hand-off 100,
unload confirm 3, min area 1.0, conf 0.35 → verify: **[DONE]** the ten-case failure suite
passes at these defaults, including a burst-frame case that exposed unbounded velocity in
the hand-off projection (now clamped to 1500 px/s and 0.5 s of extrapolation).
**Open — needs the hand-counted clip.** On real footage 84% of tracks inherit via hand-off at
`handoff_radius=100`, because these frames are dense enough that a newborn track is nearly
always near one that just vanished. 100 is the value tuned against the camera and is now the
default, but the right value is a measurement, not a guess: run a clip with a known total and
read the verdict histogram
in the trace file. `never_reached_below` dominating means the tracker is fragmenting (not the
counter); `born_below_line` means counts are being lost to ID switches the hand-off radius is
too tight to recover.
## Task — Counting accuracy bench (REQ-150…153)
1. `count_runs` table + `count` job type → verify: **[DONE]** migration rebuilt the `jobs`
table to accept the new type (SQLite cannot alter a CHECK constraint); all 928 existing
job rows preserved.
2. Headless counter reusing the live pipeline → verify: **[DONE]** 21,544 frames of
`2026-08-14/batch011.mp4` in 147 s = **146 fps**, against 124 fps through the live view.
Rendering was the difference.
3. Scored table with editable ground truth → verify: **[DONE]** setting a ground truth,
clearing it, and the totals excluding unscored rows all round-trip through the API.
4. Background job over a selection or all videos → verify: **[DONE]** queued one video, the
job reported `7150/21544 frames` mid-run and stored in 169 / out 8 / net 161 on finish.
5. Page + route + sidebar entry → verify: **[DONE]** frontend builds; listing serves 222 rows
in 0.18 s once ffprobe is warm (7.7 s cold).
**Sizing.** The archive is 129 hours across 222 videos. At the measured 146 fps a full
recount is roughly **22 GPU-hours**, so "Count all" is an overnight job, not an interactive
one. It is resumable — already-counted videos are skipped unless `recount` is ticked — and
cancelling mid-video discards that video's partial count rather than storing it as a result.
## Task — Real recording times, 06:00 working days (REQ-160…163)
1. Read the burned-in overlay without adding an OCR dependency → verify: **[DONE]** 12 glyph
templates matched per frame; decodes frames it was never trained on exactly, at
confidence 0.75–0.87.
2. Reject bad reads rather than trust them → verify: **[DONE]** a misread that produced the
year 7026 is rejected by the year-range check; low confidence or fewer than two agreeing
frames flags the row for review instead of silently regrouping it.
3. Working-day grouping and renumbering → verify: **[DONE]** scanned all 224 recordings;
**29 land on a different working day** than their folder. Working day 2026-08-13 now starts
at 08:27 because the 00:07 and 00:22 recordings moved to 08-12.
4. Nothing written to the archive → verify: **[DONE]** the mount is `:ro`; the index lives in
`video_clock` and the file path stays the row's identity, so existing counts survived.
**Timezone.** Start times are stored as wall-clock **text**, never an epoch. Storing an epoch
made the backend (UTC) and the browser (UTC+7) disagree by seven hours, which moved recordings
across the 06:00 boundary into the wrong working day — `2026-08-07/batch4` read 20:12:42 and
displayed as 03:12:42 the next day. Caught by cross-checking one file against the video.
5. Group the table into collapsible cycles (REQ-164) → verify: **[DONE]** 10 cycles render
newest first; `Siklus 13 Agt 2026` holds 28 recordings running 08:27 → 01:19 the next
morning, which is the midnight crossing the grouping exists to make readable. A cycle
header selects all of its rows for a recount in one click.
6. Video Archive browses by cycle (REQ-165) → verify: **[DONE]** `Siklus 13 Agt 2026` lists
28 recordings running 08:27 through midnight to 01:19, with `batch001…003` from the
*2026-08-14* folder correctly appearing as #26–28 of the 13 Agt cycle and flagged with
their folder. Listing the cycles costs 0.13 s because it counts filenames instead of
running ffprobe on the whole archive.
7. Truck check with v4 (REQ-166) → verify: **[DONE]** scanned 226 recordings, 12 frames each,
in ~5 minutes. **225 contain a truck** (136 in every sampled frame, 89 in some), so the
"one file is one batch" premise holds. One recording — `2026-08-07/batch027.mp4` — shows no
truck in any sampled frame and is flagged in the table. Three files will not open at all.
A first attempt died after 8 recordings with `database is locked`: the writer opened a
second connection inside an open write transaction. Now a single UPSERT on one cursor.
8. Align the production counter to the 06:00 cycle (REQ-167) → verify: **[DONE]**
`predict.py`'s `DAILY_CUTOFF_TIME` default moved from `20:00` to `06:00`; at `06:00` its
`get_counting_date()` agrees with the app's `working_day()` on 8 of 8 boundary cases, at
`20:00` it disagreed on 3. `algoritma-batch/migrate_cutoff_0600.py` re-files existing rows:
tested against a replica of the Jetson schema, 9 batches split across two counting dates
by the old cutoff collapse into one day numbered #1–#8, `daily_summaries` is rebuilt, the
unique key holds, a timestamped backup is written, a second run is a no-op, and a row with
an unparseable `start_time` is left alone rather than failing the migration.
**The recorder is `algoritma-batch/batch_video_cropper.py`, in this repo**, running 24/7 on
this machine (pid seen at 187 min CPU). It reads `rtsp://192.168.192.96:8554/cam`, uses
`BatchLifecycleManager` + `v3-best.pt` to detect a truck arriving and leaving, and writes
`~/reTraining/data/archive/{date}/batch{NNN}.mp4` — one file per truck session, which is what
makes "one file is one batch" true.
9. Correct the recorder's frame rate (REQ-168) → verify: **[DONE]** `VIDEO_FPS = 10.0` was
hard-coded while the camera delivers 25, so every archived file claimed a duration 2.49x
too long (batch003: 22,968 frames, overlay says 15.4 minutes, file says 38.3). The rate now
comes from the stream and writes are paced against the wall clock. Recorded 30 s from the
live production stream with the loop deliberately starved to ~3.7 fps: the file came out
**31.56 s against 31.9 s real, 1.1% off**; the old code would have produced 11.9 s.
The camera is **25 fps, not 60** — RTSP metadata, the HLS playlist (`FRAME-RATE=25.000`)
and the measured delivery rate (24.8 fps) all agree.
**Decided, not a defect:** the recorder keeps `DAILY_CUTOFF_TIME = "00:00"`, so folder names
stay calendar dates. The user's call — what matters is that the app is right, and it is: cycles
are derived from each recording's real start time, so a file sitting in the 15 Aug folder but
recorded at 01:00 appears under the 14 Aug cycle. Nothing downstream reads the folder name as a
date. Note if this is ever revisited: this script's `get_counting_date()` returns *tomorrow*
after the cutoff, unlike `predict.py`'s, so the function would need aligning, not just the
constant.
10. Record once on the Jetson, cut sessions on the ASUS (REQ-170) → verify: **[DONE]** MediaMTX
on the Jetson now records 24/7 (`record: yes`, `playback: yes`, 15-minute segments, 24-hour
buffer — 18 GB of its 36 GB free; 48 h would have needed 37 GB and did not fit). The
recorder no longer re-encodes: on session end it downloads that exact time range as a copy.
Fetch verified against the live stream — asked for 16:10:53 +45 s, the clip's burned-in
overlay reads 16:10:52 → 16:11:37, exactly 45 s, 1125 frames at 25 fps, 1 s off from the
camera's own clock. A simulated 40 s session produced a 46 s clip whose sidecar
(`16:14:45`, from the server) matches the overlay to the second. Files are now **HEVC
1920x1080 copies, ~4.7x smaller** than the old 1280x720 mpeg4 re-encodes.
11. Video Archive stays current without a scan (REQ-170) → verify: **[DONE]** the recorder
writes a `.json` sidecar beside each clip and the app reads it live, so a new session
appears in the right cycle with a server-accurate time and no OCR at all.
**The camera cannot do 60 fps.** `FPSMax=25` on every stream format of the DH-IPC-HFW1230, and
it already runs at that (1080p, H.265, 2048 kbps CBR). The "not smooth" impression came from
the broken timebase, not the frame rate.
**Mistake to record:** while testing the fetch, a test clip was copied over
`data/archive/2026-08-14/batch007.mp4`, destroying a real 09:20:08 truck recording. It had never
been used for frame extraction or counting, so no dataset or annotation was affected, and the
test clip and its index row were removed. The file itself is gone from this machine; the rsync
history suggests a copy may exist on 192.168.192.105/.106.
**Worth checking on the Jetson:** `BATCH_MERGE_THRESHOLD_SECONDS` defaults to 300, so a truck
arriving within five minutes of the last batch *continues* it instead of starting a new one.
Video Archive counts one file as one batch, so if trucks really do turn around that fast the
two will disagree.
**Open — 19 recordings need a human.** 8 are unreadable (3 of them will not open at all:
`2026-08-06/batch4`, `2026-08-06/batch9`, `2026-08-14/batch016` — likely truncated) and 11 were
read with low confidence. Both are flagged amber in the table and accept a hand-typed time.
## Known open points
- *Not closed by any task, by choice:* **any rebuild kills the running job.** Task 14's resume