feat: sync live-count, exemplar annotation modules, and update .gitignore
This commit is contained in:
1 parent
b6624eeff9
commit
ac95674c07
39 files changed
+3979
-444
No files matched your search
@@ -100,6 +100,19 @@ changes.
|
||||
found nothing on (REQ-033) writes no annotations, so a resume re-does it; that is accepted
|
||||
rather than tracked.
|
||||
|
||||
- **REQ-171** — In the auto-annotate modal the SAM3 text prompt of each selected class is
|
||||
editable in place, next to the live preview. Saving it writes
|
||||
`project_classes.prompt` — the same field the Projects page edits — so the batch job and
|
||||
every later run send that text. The preview is the tuning surface; the stored prompt is the
|
||||
artifact it produces.
|
||||
- **REQ-172** — On the previewed frame the user can drag **positive** and **negative** box
|
||||
exemplars (shift-drag for negative). They are appended to the active class's text prompt,
|
||||
re-run immediately, and can be undone or cleared. Exemplars are a **tuning aid only**: they
|
||||
are never written as annotations and never carried into the batch job, because SAM3's
|
||||
geometric prompts pool features from the current image — replaying them on another frame
|
||||
would ask about whatever happens to sit at those coordinates there. They belong to exactly
|
||||
one class, so a new frame or a new active class discards them.
|
||||
|
||||
## E. Review & correction
|
||||
|
||||
- **REQ-040** — The user reviews frames one at a time, with fast navigation (left/right
|
||||
@@ -117,6 +130,53 @@ changes.
|
||||
- **REQ-046** — The user can delete/clear all annotations of a specific class across all frames in
|
||||
the current batch from the Review editor.
|
||||
|
||||
- **REQ-173** — In the review editor a plain drag on the canvas is an **exemplar-driven
|
||||
label**, not just a rectangle. It proposes, in one action — and REQ-175's Apply is what
|
||||
makes any of it real — that (a) the drawn shape becomes a `manual` annotation of the
|
||||
active class — snapped to a SAM3 polygon first when the batch's
|
||||
`label_type` is `polygon`, since a rectangle is a bad polygon label — (b) appends the box
|
||||
to the frame's positive exemplar pool for that class, and (c) re-runs SAM3 over the whole
|
||||
frame with the class's text prompt plus the pooled exemplars, deleting every existing shape
|
||||
of that class on the frame and writing the detections in its place, then re-inserting the
|
||||
pooled exemplar shapes verbatim so the user's own drawings always survive. The pool is
|
||||
**frame-local and ephemeral** for the same reason as REQ-172 — SAM3's geometric prompts pool
|
||||
features from the current image — so leaving the frame or switching the active class clears
|
||||
it; the annotations it produced persist like any other. If the GPU lock (REQ-070) is not
|
||||
free, the run comes back with the drawn shapes alone and says so, so the user can still
|
||||
file them (REQ-175) and labeling is never blocked by a background job.
|
||||
- **REQ-174** — **Shift**-drag in the review editor adds a **negative** exemplar. It is never
|
||||
stored as an annotation; it deletes any existing shape of the active class that overlaps it,
|
||||
and it is sent as a negative box in the REQ-173 re-detect. It is the "not this, and not
|
||||
things like this" gesture, so it doubles as a delete. A negative is **spent on Apply**: the
|
||||
frame it was applied to no longer carries what it rejected, so the drawing is dropped from
|
||||
the pool while the positives stay on as prompts.
|
||||
- **REQ-175** — An exemplar drag **previews**; it never writes on its own. The run's result
|
||||
is drawn over the frame as proposals and a small panel floats on the canvas with the four
|
||||
filters that decide what survives — confidence, NMS overlap, minimum box size, maximum
|
||||
shapes — each re-running the preview as it moves. **Apply** writes the previewed set,
|
||||
**Discard** rewinds the pool to whatever is already on the frame and leaves it untouched.
|
||||
The panel is scoped to this gesture: its values are not stored, not shared with the
|
||||
auto-annotate modal, and reset with the frame. Defaults are confidence `0.5`, NMS `0.8`,
|
||||
min box `0.002`, max `100` — deliberately permissive, because on a dense frame an
|
||||
aggressive NMS or area floor deletes real, touching objects rather than duplicates.
|
||||
|
||||
## E4. Live counting preview
|
||||
|
||||
- **REQ-176** — A **live** source on the Live Count page is a **WebRTC (WHEP) URL** and
|
||||
nothing else; an RTSP URL is rejected with a message saying so. The backend derives the
|
||||
RTSP leg of the same streaming-server path from it (`http://host:8889/cam` →
|
||||
`rtsp://host:8554/cam`) and counts from that: WebRTC is what makes the browser preview
|
||||
cheap, but pulling it into Python would add ICE and a jitter buffer on top of the identical
|
||||
H.264 decode. One ingest on the streaming server, two consumers. The ports are read from
|
||||
the environment (`MEDIAMTX_RTSP_PORT`, `MEDIAMTX_WHEP_PATH`), never hardcoded. Archive
|
||||
files are unaffected — they are still opened as files.
|
||||
- **REQ-177** — A live session is **watched over WebRTC**, played straight from the streaming
|
||||
server by the browser: the frames never pass through this app and it encodes no JPEG for
|
||||
them. What the model saw — boxes, ids, confidences, the counting line and its band, the
|
||||
ignored region, the running totals — is served as geometry from
|
||||
`GET /api/live-count/overlay` and drawn on a canvas over the video. The MJPEG endpoint
|
||||
remains the preview for **archive files** only, and refuses a WebRTC session.
|
||||
|
||||
## F. Master dataset
|
||||
|
||||
- **REQ-050** — Approving a batch **merges** its approved frames and their labels into the
|
||||
@@ -164,6 +224,31 @@ changes.
|
||||
and the reason it did or did not count, so a miss can be attributed to the model, the
|
||||
tracker, or the counter.
|
||||
|
||||
- **REQ-145** — Counting algorithms are **pluggable**. Each registers under a stable id
|
||||
(`line_cross`, `possession`) and the session constructs one by id. The `Counter` protocol
|
||||
in `src/interfaces.py` is the contract, corrected to match reality: `update()` returns the
|
||||
frame's count events, not `None`. Adding an algorithm must not require editing
|
||||
`live_count.py` or `counting_bench.py`.
|
||||
- **REQ-146** — Each algorithm **declares its own parameters** — name, type, default, range —
|
||||
and an endpoint serves that declaration, mirroring `live-count/models`. The frontend renders
|
||||
its controls from the declaration and hardcodes no per-algorithm parameter list. The start
|
||||
request carries `algorithm` plus an opaque `params` object validated against the
|
||||
declaration, replacing today's flat line-specific fields.
|
||||
- **REQ-147** — Geometry is generalised from a line to a **named shape set**. `line_cross`
|
||||
declares one horizontal segment; `possession` declares a bed polygon and an approach zone.
|
||||
The editor's drag channel (`move_line`) becomes shape-agnostic, so any algorithm's geometry
|
||||
is adjustable live without a new endpoint.
|
||||
- **REQ-148** — The **possession counter**: every sack track carries an `owner_id`, the person
|
||||
track it currently overlaps, or none when at rest. A count fires on an ownership change that
|
||||
crosses the bed boundary — person-outside to bed, or person-outside to person-inside.
|
||||
Ownership is sticky with hysteresis, so occlusion by the carrier's back and the unowned
|
||||
mid-air phase of a thrown sack do not break it. This requires a `person` class alongside
|
||||
`sack` from the detector.
|
||||
- **REQ-149** — Every count run records **which algorithm and parameter set** produced it, and
|
||||
accuracy is comparable per algorithm against the same ground truth. Switching algorithms
|
||||
adds results, it never invalidates stored ones — so `count_runs` is keyed by
|
||||
`(project, video, algorithm)`, not by video alone.
|
||||
|
||||
## F4. Counting accuracy bench
|
||||
|
||||
- **REQ-150** — A page lists every archive video as a row: date, batch, length, and the
|
||||
@@ -179,6 +264,23 @@ changes.
|
||||
which is what makes counting a 30-minute video practical. A run records the parameters and
|
||||
model it used.
|
||||
|
||||
- **REQ-154** — Ground truth can be **imported in bulk** from the operations sheet
|
||||
(`./GT.xlsx`, `DATA MUAT PAKAN PER LINE`). The camera watches **Line 1**; Line 2 is
|
||||
recorded for completeness but never scored. Each sheet is one working day; a row is one
|
||||
truck with a `BAG` count, a `DUS` count and a plate.
|
||||
- **REQ-155** — `BAG` (sacks) and `DUS` (boxes) are **separate commodities**, counted and
|
||||
scored separately. A box already resting in the truck bed is a legitimate object of a
|
||||
different class, not a detection fault.
|
||||
- **REQ-156** — An import never silently guesses. Recordings are aligned to sheet rows by
|
||||
start time against row order, the proposed pairing is **shown for human confirmation**
|
||||
before anything is written, and each imported value records that it came from the sheet
|
||||
rather than from a hand count. A recording that merged two trucks
|
||||
(`BATCH_MERGE_THRESHOLD_SECONDS`) is flagged, not paired.
|
||||
- **REQ-157** — Sheet values are **order quantities, not hand counts** — 67% of them are
|
||||
exactly 160 or 180 — so they score aggregate accuracy across many trucks and never
|
||||
adjudicate a single video. Per-event truth for algorithm comparison comes from a
|
||||
hand-counted clip, held separately.
|
||||
|
||||
## F5. Real recording times and working days
|
||||
|
||||
- **REQ-160** — Each recording's start time is read from the timestamp the camera burns into
|
||||
|
||||
Reference in new issue
Block a user