feat: sync live-count, exemplar annotation modules, and update .gitignore

This commit is contained in:
ervanfahriaw committed 2026-08-24 14:57:42 +07:00
1 parent b6624eeff9
commit ac95674c07
39 files changed
+3979 -444

No files matched your search

+102
View File
@@ -100,6 +100,19 @@ changes.
found nothing on (REQ-033) writes no annotations, so a resume re-does it; that is accepted
rather than tracked.
- **REQ-171** — In the auto-annotate modal the SAM3 text prompt of each selected class is
editable in place, next to the live preview. Saving it writes
`project_classes.prompt` — the same field the Projects page edits — so the batch job and
every later run send that text. The preview is the tuning surface; the stored prompt is the
artifact it produces.
- **REQ-172** — On the previewed frame the user can drag **positive** and **negative** box
exemplars (shift-drag for negative). They are appended to the active class's text prompt,
re-run immediately, and can be undone or cleared. Exemplars are a **tuning aid only**: they
are never written as annotations and never carried into the batch job, because SAM3's
geometric prompts pool features from the current image — replaying them on another frame
would ask about whatever happens to sit at those coordinates there. They belong to exactly
one class, so a new frame or a new active class discards them.
## E. Review & correction
- **REQ-040** — The user reviews frames one at a time, with fast navigation (left/right
@@ -117,6 +130,53 @@ changes.
- **REQ-046** — The user can delete/clear all annotations of a specific class across all frames in
the current batch from the Review editor.
- **REQ-173** — In the review editor a plain drag on the canvas is an **exemplar-driven
label**, not just a rectangle. It proposes, in one action — and REQ-175's Apply is what
makes any of it real — that (a) the drawn shape becomes a `manual` annotation of the
active class — snapped to a SAM3 polygon first when the batch's
`label_type` is `polygon`, since a rectangle is a bad polygon label — (b) appends the box
to the frame's positive exemplar pool for that class, and (c) re-runs SAM3 over the whole
frame with the class's text prompt plus the pooled exemplars, deleting every existing shape
of that class on the frame and writing the detections in its place, then re-inserting the
pooled exemplar shapes verbatim so the user's own drawings always survive. The pool is
**frame-local and ephemeral** for the same reason as REQ-172 — SAM3's geometric prompts pool
features from the current image — so leaving the frame or switching the active class clears
it; the annotations it produced persist like any other. If the GPU lock (REQ-070) is not
free, the run comes back with the drawn shapes alone and says so, so the user can still
file them (REQ-175) and labeling is never blocked by a background job.
- **REQ-174** — **Shift**-drag in the review editor adds a **negative** exemplar. It is never
stored as an annotation; it deletes any existing shape of the active class that overlaps it,
and it is sent as a negative box in the REQ-173 re-detect. It is the "not this, and not
things like this" gesture, so it doubles as a delete. A negative is **spent on Apply**: the
frame it was applied to no longer carries what it rejected, so the drawing is dropped from
the pool while the positives stay on as prompts.
- **REQ-175** — An exemplar drag **previews**; it never writes on its own. The run's result
is drawn over the frame as proposals and a small panel floats on the canvas with the four
filters that decide what survives — confidence, NMS overlap, minimum box size, maximum
shapes — each re-running the preview as it moves. **Apply** writes the previewed set,
**Discard** rewinds the pool to whatever is already on the frame and leaves it untouched.
The panel is scoped to this gesture: its values are not stored, not shared with the
auto-annotate modal, and reset with the frame. Defaults are confidence `0.5`, NMS `0.8`,
min box `0.002`, max `100` — deliberately permissive, because on a dense frame an
aggressive NMS or area floor deletes real, touching objects rather than duplicates.
## E4. Live counting preview
- **REQ-176** — A **live** source on the Live Count page is a **WebRTC (WHEP) URL** and
nothing else; an RTSP URL is rejected with a message saying so. The backend derives the
RTSP leg of the same streaming-server path from it (`http://host:8889/cam` →
`rtsp://host:8554/cam`) and counts from that: WebRTC is what makes the browser preview
cheap, but pulling it into Python would add ICE and a jitter buffer on top of the identical
H.264 decode. One ingest on the streaming server, two consumers. The ports are read from
the environment (`MEDIAMTX_RTSP_PORT`, `MEDIAMTX_WHEP_PATH`), never hardcoded. Archive
files are unaffected — they are still opened as files.
- **REQ-177** — A live session is **watched over WebRTC**, played straight from the streaming
server by the browser: the frames never pass through this app and it encodes no JPEG for
them. What the model saw — boxes, ids, confidences, the counting line and its band, the
ignored region, the running totals — is served as geometry from
`GET /api/live-count/overlay` and drawn on a canvas over the video. The MJPEG endpoint
remains the preview for **archive files** only, and refuses a WebRTC session.
## F. Master dataset
- **REQ-050** — Approving a batch **merges** its approved frames and their labels into the
@@ -164,6 +224,31 @@ changes.
and the reason it did or did not count, so a miss can be attributed to the model, the
tracker, or the counter.
- **REQ-145** — Counting algorithms are **pluggable**. Each registers under a stable id
(`line_cross`, `possession`) and the session constructs one by id. The `Counter` protocol
in `src/interfaces.py` is the contract, corrected to match reality: `update()` returns the
frame's count events, not `None`. Adding an algorithm must not require editing
`live_count.py` or `counting_bench.py`.
- **REQ-146** — Each algorithm **declares its own parameters** — name, type, default, range —
and an endpoint serves that declaration, mirroring `live-count/models`. The frontend renders
its controls from the declaration and hardcodes no per-algorithm parameter list. The start
request carries `algorithm` plus an opaque `params` object validated against the
declaration, replacing today's flat line-specific fields.
- **REQ-147** — Geometry is generalised from a line to a **named shape set**. `line_cross`
declares one horizontal segment; `possession` declares a bed polygon and an approach zone.
The editor's drag channel (`move_line`) becomes shape-agnostic, so any algorithm's geometry
is adjustable live without a new endpoint.
- **REQ-148** — The **possession counter**: every sack track carries an `owner_id`, the person
track it currently overlaps, or none when at rest. A count fires on an ownership change that
crosses the bed boundary — person-outside to bed, or person-outside to person-inside.
Ownership is sticky with hysteresis, so occlusion by the carrier's back and the unowned
mid-air phase of a thrown sack do not break it. This requires a `person` class alongside
`sack` from the detector.
- **REQ-149** — Every count run records **which algorithm and parameter set** produced it, and
accuracy is comparable per algorithm against the same ground truth. Switching algorithms
adds results, it never invalidates stored ones — so `count_runs` is keyed by
`(project, video, algorithm)`, not by video alone.
## F4. Counting accuracy bench
- **REQ-150** — A page lists every archive video as a row: date, batch, length, and the
@@ -179,6 +264,23 @@ changes.
which is what makes counting a 30-minute video practical. A run records the parameters and
model it used.
- **REQ-154** — Ground truth can be **imported in bulk** from the operations sheet
(`./GT.xlsx`, `DATA MUAT PAKAN PER LINE`). The camera watches **Line 1**; Line 2 is
recorded for completeness but never scored. Each sheet is one working day; a row is one
truck with a `BAG` count, a `DUS` count and a plate.
- **REQ-155** — `BAG` (sacks) and `DUS` (boxes) are **separate commodities**, counted and
scored separately. A box already resting in the truck bed is a legitimate object of a
different class, not a detection fault.
- **REQ-156** — An import never silently guesses. Recordings are aligned to sheet rows by
start time against row order, the proposed pairing is **shown for human confirmation**
before anything is written, and each imported value records that it came from the sheet
rather than from a hand count. A recording that merged two trucks
(`BATCH_MERGE_THRESHOLD_SECONDS`) is flagged, not paired.
- **REQ-157** — Sheet values are **order quantities, not hand counts** — 67% of them are
exactly 160 or 180 — so they score aggregate accuracy across many trucks and never
adjudicate a single video. Per-event truth for algorithm comparison comes from a
hand-counted clip, held separately.
## F5. Real recording times and working days
- **REQ-160** — Each recording's start time is read from the timestamp the camera burns into