feat: sync live-count, exemplar annotation modules, and update .gitignore

This commit is contained in:
ervanfahriaw committed 2026-08-24 14:57:42 +07:00
1 parent b6624eeff9
commit ac95674c07
39 files changed
+3979 -444

No files matched your search

BIN
View File
Binary file not shown.
+115
View File
@@ -123,9 +123,14 @@ box.
| `batches.py` | batch lifecycle | **new** |
| `review.py` | annotation CRUD, frame status, click-assist | **new** |
| `autolabel.py` | the SAM3 job over a whole batch | **new** |
| `preview.py` | one-frame preview for the auto-annotate modal (REQ-171,172) | **new** |
| `exemplar.py` | exemplar-driven labeling in the review editor (REQ-173,174) | **new** |
| `dataset.py` | merge into the master dataset, stable split | **new** |
| `evaluate.py` | validate base vs new model | **new** |
| `hardware.py` | VRAM detection → training defaults | **new** |
| `live_count.py` | the live counting session: capture → track → count | **new** |
| `live_source.py` | what a source is, how it opens, WHEP↔RTSP (REQ-176) | **new** |
| `live_render.py` | the MJPEG overlay, archive-file preview only (REQ-177) | **new** |
| `api/` | the FastAPI routes, one module per domain | **new** |
Removed: `uploads.py`, `static/index.html`, and the old flow's endpoints.
@@ -167,6 +172,9 @@ POST /api/projects/{id}/batches # {rel, start_sec, end_sec, fps}
GET /api/batches/{id} # status + review progress (REQ-045)
GET /api/batches/{id}/frames # frames + statuses
POST /api/batches/{id}/autolabel # {threshold} → job (REQ-030,032,034)
POST /api/batches/{id}/preview # one frame, run now, nothing written;
# +exemplars[] {box:[cx,cy,w,h], positive}
# +exemplar_class_name (REQ-171,172)
DELETE /api/batches/{id}/classes/{class_id}/annotations # clear all shapes of class in batch (REQ-046)
POST /api/batches/{ids}/approve # one or many, comma-separated → one merge job (REQ-131)
GET /api/batches/{ids}/triage/summary # one or many, comma-separated (REQ-130)
@@ -180,6 +188,7 @@ POST /api/frames/{id}/annotations # add a manual shape (REQ-042)
PATCH /api/annotations/{id} # move/resize/reclass
DELETE /api/annotations/{id}
POST /api/frames/{id}/assist # click/box → SAM3 shape (REQ-043)
POST /api/frames/{id}/exemplar-label # drawn pool → re-detect one class (REQ-173,174)
POST /api/frames/{id}/status # approved | rejected | pending (REQ-041)
POST /api/projects/{id}/train # → train job (REQ-060,061,062)
@@ -187,11 +196,36 @@ GET /api/projects/{id}/models # versions + metrics (REQ-063,06
GET /api/models/{id}/weights # download best.pt
POST /api/models/{id}/promote # make it the project's base model (REQ-064)
GET /api/projects/{id}/live-count/models # weights this project can count with
POST /api/projects/{id}/live-count/start # {source|source_rel, model_path, dials}
# source must be a WHEP URL (REQ-176)
POST /api/live-count/stop
PATCH /api/live-count/line # move the line mid-session
GET /api/live-count/status # counts + preview: "webrtc" | "mjpeg"
GET /api/live-count/overlay # boxes/line/counts for the canvas (REQ-177)
GET /api/live-count/stream # MJPEG; 409 on a WebRTC session (REQ-177)
GET /api/jobs?project_id=… REQ-070,071
GET /api/jobs/{id}
POST /api/jobs/{id}/cancel
```
### Live counting (REQ-176, REQ-177)
One camera, one ingest on the streaming server, two consumers:
```
camera ──▶ MediaMTX ──┬── WHEP :8889/cam/whep ──▶ browser <video> + <canvas> overlay
└── RTSP :8554/cam ──▶ backend decode → YOLO → counter
└─▶ GET /api/live-count/overlay
```
The user types only the WHEP URL; `live_source.whep_to_rtsp` derives the RTSP one, with the
ports coming from `MEDIAMTX_RTSP_PORT` / `MEDIAMTX_WHEP_PATH`. Because the browser plays the
camera directly, a live session encodes no JPEG at all — `live_render.py` runs only for
archive files, which have no WebRTC leg. The canvas draws exactly what `live_render.py`
would have burned in, so both previews describe the same session.
## Job flows
**extract (REQ-020…023).** `ffmpeg -ss <start> -to <end> -i <video> -vf fps=<n> -q:v 2
@@ -277,6 +311,87 @@ Pages:
The canvas editor is hand-written; the normalized-coordinate conventions already exist in
`sessions.py` (`annotations_payload`, `detections_payload`) as a reference.
**Auto-annotate modal (REQ-171, REQ-172).** `AutoAnnotateModal.jsx` splits into
`PreviewShapes.jsx` (the result overlay, shared with the mass modal),
`ClassPromptPanel.jsx` (class chips + the editable SAM3 prompt) and `ExemplarCanvas.jsx`
(the drag-to-draw layer). All three overlays and the `<img>` share one shrink-wrapped
`position: relative` wrapper — they are sized to it, so nothing else may sit inside it or
every box shifts off the pixels it describes.
With SAM3 one selected chip is *active*: it owns the prompt field and any exemplars drawn on
the frame, so its chip is a pair of buttons — the name activates, the `×` deselects.
Exemplars are normalized `[cx, cy, w, h]`, sent only to `/preview`, and dropped whenever the
frame or the active class changes. A redraw is debounced 250 ms and re-runs the whole prompt
set from empty, which is also how undo works — SAM3 can only append geometric prompts.
**Exemplar-driven labeling in review (REQ-173, REQ-174, REQ-175).** In `draw` mode a drag on
`AnnotationCanvas` is an exemplar, not a rectangle: `onExemplar(box, positive)` where
`positive` is `!event.shiftKey`. `hooks/useExemplarPool.js` keeps the pool in a ref as well
as state — the ref is what gets sent, so a drag that lands mid-flight is never lost — and posts
the **whole pool** to `/exemplar-label` 400 ms after the last drag. One pass runs at a time;
a drag arriving during a pass sets `rerunWanted` so exactly one rerun follows instead of a
queue. The pool is cleared by a frame change or a class change, and `Undo example` re-runs
with the shortened pool.
The run is a **dry run by default** (REQ-175). `ExemplarFilterPanel.jsx` is passed to
`ReviewSidebar` and rendered above the class list — not over the frame, which is where its
proposals are drawn. It is mounted only while a run is undecided (`pool.active`): the first
drag opens it, Apply and Discard close it, and the pool outlives it. Its four sliders
re-preview on a 250 ms debounce. While a preview is up the canvas **hides the stored shapes
of the class under review** — the run replaces them wholesale, and leaving them on screen made
a rejected detection look like it had never gone; other classes stay, dimmed
(`svg.previewing`). Proposals draw dashed on top, green for the boxes the user drew and
class-colored for what SAM3 found.
`Apply` re-runs with `apply: true` rather than posting the previewed geometry back — SAM3 is
deterministic for a pool and a threshold, and the browser should not be the authority on what
gets stored. `Discard` truncates the pool to `appliedRef`, the length it had at the last
successful apply, so a rejected run leaves neither shapes nor prompts behind. A successful
apply drops the negatives it sent and keeps the positives (REQ-174); drags that landed while
that request was in flight are not part of it and stay at the end of the pool, with a rerun
queued for them.
`.canvas-wrap`'s overlay rules are scoped to its direct child (`> svg`): they set
`position: absolute; width: 100%`, which any nested SVG — an icon in a panel, say — would
otherwise inherit and stretch across the whole frame.
Request/response:
```
POST /api/frames/{id}/exemplar-label
{ "exemplars": [{"box": [x0,y0,x1,y1], "positive": true}, …], # normalized xyxy
"class_id": 0,
"threshold": 0.5, "iou_threshold": 0.8, # the panel (REQ-175)
"min_box_frac": 0.002, "max_detections": 100,
"apply": false }
→ { "shapes": [{"geometry": …, "score": …, "source": "manual"|"auto"}, …],
"applied": false, "redetected": true, "message": null,
"annotations": null } # the frame, on apply only
```
`backend/exemplar.py` converts each box to SAM3's normalized `[cx, cy, w, h]`, runs one
`open_state` + `apply_prompts` with the class's stored `prompt` as text plus every exemplar,
then rewrites the frame **for that class only**:
- positives are stored first as `source='manual'` — the literal rectangle in a `bbox`
project, the mask polygon of whatever SAM3 found inside it (IoU ≥ `SNAP_IOU`) in a
`polygon` one;
- detections overlapping a negative by ≥ `NEGATIVE_IOU` (0.3) are dropped, and ones
overlapping a positive by ≥ `DUPLICATE_IOU` (0.6) are dropped as the user's own shape
already covers them;
- what remains is written as `source='auto'`, so a later batch re-run replaces it (REQ-034)
while the drawn shapes survive.
The panel's filters run before any of that, in the order the batch job uses them: area floor,
then NMS (`labeling.deduplicate`), then the cap on how many survive.
The delete-and-reinsert happens in one `db.cursor()` transaction, so the frame is never
briefly empty. Other classes on the frame are never touched. The GPU lock is taken with a
**5 s** timeout — shorter than `assist`'s 20 s because this fires from a mouse gesture; on
timeout the response carries `redetected: false`, a message the panel shows, and the drawn
boxes alone as the preview. Applying that run **appends** the drawn boxes and honours the
negatives instead of taking the replace path — with no detections to put back, replacing
would wipe the class and leave only the drawings.
## Docker (REQ-072)
- `Dockerfile` — python 3.12, `ffmpeg`, `uv`, CUDA torch, `uv pip install -e sam3/`. The
+102
View File
@@ -100,6 +100,19 @@ changes.
found nothing on (REQ-033) writes no annotations, so a resume re-does it; that is accepted
rather than tracked.
- **REQ-171** — In the auto-annotate modal the SAM3 text prompt of each selected class is
editable in place, next to the live preview. Saving it writes
`project_classes.prompt` — the same field the Projects page edits — so the batch job and
every later run send that text. The preview is the tuning surface; the stored prompt is the
artifact it produces.
- **REQ-172** — On the previewed frame the user can drag **positive** and **negative** box
exemplars (shift-drag for negative). They are appended to the active class's text prompt,
re-run immediately, and can be undone or cleared. Exemplars are a **tuning aid only**: they
are never written as annotations and never carried into the batch job, because SAM3's
geometric prompts pool features from the current image — replaying them on another frame
would ask about whatever happens to sit at those coordinates there. They belong to exactly
one class, so a new frame or a new active class discards them.
## E. Review & correction
- **REQ-040** — The user reviews frames one at a time, with fast navigation (left/right
@@ -117,6 +130,53 @@ changes.
- **REQ-046** — The user can delete/clear all annotations of a specific class across all frames in
the current batch from the Review editor.
- **REQ-173** — In the review editor a plain drag on the canvas is an **exemplar-driven
label**, not just a rectangle. It proposes, in one action — and REQ-175's Apply is what
makes any of it real — that (a) the drawn shape becomes a `manual` annotation of the
active class — snapped to a SAM3 polygon first when the batch's
`label_type` is `polygon`, since a rectangle is a bad polygon label — (b) appends the box
to the frame's positive exemplar pool for that class, and (c) re-runs SAM3 over the whole
frame with the class's text prompt plus the pooled exemplars, deleting every existing shape
of that class on the frame and writing the detections in its place, then re-inserting the
pooled exemplar shapes verbatim so the user's own drawings always survive. The pool is
**frame-local and ephemeral** for the same reason as REQ-172 — SAM3's geometric prompts pool
features from the current image — so leaving the frame or switching the active class clears
it; the annotations it produced persist like any other. If the GPU lock (REQ-070) is not
free, the run comes back with the drawn shapes alone and says so, so the user can still
file them (REQ-175) and labeling is never blocked by a background job.
- **REQ-174** — **Shift**-drag in the review editor adds a **negative** exemplar. It is never
stored as an annotation; it deletes any existing shape of the active class that overlaps it,
and it is sent as a negative box in the REQ-173 re-detect. It is the "not this, and not
things like this" gesture, so it doubles as a delete. A negative is **spent on Apply**: the
frame it was applied to no longer carries what it rejected, so the drawing is dropped from
the pool while the positives stay on as prompts.
- **REQ-175** — An exemplar drag **previews**; it never writes on its own. The run's result
is drawn over the frame as proposals and a small panel floats on the canvas with the four
filters that decide what survives — confidence, NMS overlap, minimum box size, maximum
shapes — each re-running the preview as it moves. **Apply** writes the previewed set,
**Discard** rewinds the pool to whatever is already on the frame and leaves it untouched.
The panel is scoped to this gesture: its values are not stored, not shared with the
auto-annotate modal, and reset with the frame. Defaults are confidence `0.5`, NMS `0.8`,
min box `0.002`, max `100` — deliberately permissive, because on a dense frame an
aggressive NMS or area floor deletes real, touching objects rather than duplicates.
## E4. Live counting preview
- **REQ-176** — A **live** source on the Live Count page is a **WebRTC (WHEP) URL** and
nothing else; an RTSP URL is rejected with a message saying so. The backend derives the
RTSP leg of the same streaming-server path from it (`http://host:8889/cam` →
`rtsp://host:8554/cam`) and counts from that: WebRTC is what makes the browser preview
cheap, but pulling it into Python would add ICE and a jitter buffer on top of the identical
H.264 decode. One ingest on the streaming server, two consumers. The ports are read from
the environment (`MEDIAMTX_RTSP_PORT`, `MEDIAMTX_WHEP_PATH`), never hardcoded. Archive
files are unaffected — they are still opened as files.
- **REQ-177** — A live session is **watched over WebRTC**, played straight from the streaming
server by the browser: the frames never pass through this app and it encodes no JPEG for
them. What the model saw — boxes, ids, confidences, the counting line and its band, the
ignored region, the running totals — is served as geometry from
`GET /api/live-count/overlay` and drawn on a canvas over the video. The MJPEG endpoint
remains the preview for **archive files** only, and refuses a WebRTC session.
## F. Master dataset
- **REQ-050** — Approving a batch **merges** its approved frames and their labels into the
@@ -164,6 +224,31 @@ changes.
and the reason it did or did not count, so a miss can be attributed to the model, the
tracker, or the counter.
- **REQ-145** — Counting algorithms are **pluggable**. Each registers under a stable id
(`line_cross`, `possession`) and the session constructs one by id. The `Counter` protocol
in `src/interfaces.py` is the contract, corrected to match reality: `update()` returns the
frame's count events, not `None`. Adding an algorithm must not require editing
`live_count.py` or `counting_bench.py`.
- **REQ-146** — Each algorithm **declares its own parameters** — name, type, default, range —
and an endpoint serves that declaration, mirroring `live-count/models`. The frontend renders
its controls from the declaration and hardcodes no per-algorithm parameter list. The start
request carries `algorithm` plus an opaque `params` object validated against the
declaration, replacing today's flat line-specific fields.
- **REQ-147** — Geometry is generalised from a line to a **named shape set**. `line_cross`
declares one horizontal segment; `possession` declares a bed polygon and an approach zone.
The editor's drag channel (`move_line`) becomes shape-agnostic, so any algorithm's geometry
is adjustable live without a new endpoint.
- **REQ-148** — The **possession counter**: every sack track carries an `owner_id`, the person
track it currently overlaps, or none when at rest. A count fires on an ownership change that
crosses the bed boundary — person-outside to bed, or person-outside to person-inside.
Ownership is sticky with hysteresis, so occlusion by the carrier's back and the unowned
mid-air phase of a thrown sack do not break it. This requires a `person` class alongside
`sack` from the detector.
- **REQ-149** — Every count run records **which algorithm and parameter set** produced it, and
accuracy is comparable per algorithm against the same ground truth. Switching algorithms
adds results, it never invalidates stored ones — so `count_runs` is keyed by
`(project, video, algorithm)`, not by video alone.
## F4. Counting accuracy bench
- **REQ-150** — A page lists every archive video as a row: date, batch, length, and the
@@ -179,6 +264,23 @@ changes.
which is what makes counting a 30-minute video practical. A run records the parameters and
model it used.
- **REQ-154** — Ground truth can be **imported in bulk** from the operations sheet
(`./GT.xlsx`, `DATA MUAT PAKAN PER LINE`). The camera watches **Line 1**; Line 2 is
recorded for completeness but never scored. Each sheet is one working day; a row is one
truck with a `BAG` count, a `DUS` count and a plate.
- **REQ-155** — `BAG` (sacks) and `DUS` (boxes) are **separate commodities**, counted and
scored separately. A box already resting in the truck bed is a legitimate object of a
different class, not a detection fault.
- **REQ-156** — An import never silently guesses. Recordings are aligned to sheet rows by
start time against row order, the proposed pairing is **shown for human confirmation**
before anything is written, and each imported value records that it came from the sheet
rather than from a hand count. A recording that merged two trucks
(`BATCH_MERGE_THRESHOLD_SECONDS`) is flagged, not paired.
- **REQ-157** — Sheet values are **order quantities, not hand counts** — 67% of them are
exactly 160 or 180 — so they score aggregate accuracy across many trucks and never
adjudicate a single video. Per-event truth for algorithm comparison comes from a
hand-counted clip, held separately.
## F5. Real recording times and working days
- **REQ-160** — Each recording's start time is read from the timestamp the camera burns into
+167
View File
@@ -1040,6 +1040,173 @@ two will disagree.
`2026-08-06/batch4`, `2026-08-06/batch9`, `2026-08-14/batch016` — likely truncated) and 11 were
read with low confidence. Both are flagged amber in the table and accept a hand-typed time.
## Task — Ground truth import from the ops sheet (REQ-154…157)
1. Parse `docs/GT.xlsx` into rows → verify: 6 sheets (10–15 Aug 2026), Line 1 only, stopping
at the first blank plate so the inline totals row is not read as a truck. Expected Line 1
bag totals: 4780 / 4322 / 4365 / 5800 / 5645 / 9155. `[TODO]`
2. `ground_truth_bag` / `ground_truth_dus` + `gt_source` on `count_runs` (REQ-155, REQ-156) →
verify: migration runs on the live DB, existing hand-typed values survive as
`gt_source='manual'`. `[TODO]`
3. Alignment preview with human confirmation (REQ-156) → verify: a dry run on 14 Aug proposes
26 recordings against 32 Line-1 trucks, flags the shortfall, and writes nothing until
confirmed. `[TODO]`
4. Bench scores bag and box separately (REQ-155) → verify: the accuracy row shows both, and
totals only over rows that have a ground truth. `[TODO]`
## Task — Pluggable counting algorithms (REQ-145…149)
1. Fix the `Counter` protocol and register `line_cross` behind it (REQ-145) → verify: a live
session on a known clip returns **the same counts as before** the refactor — this step
changes no behaviour. `[TODO]`
2. Parameter declaration endpoint + generic frontend controls (REQ-146) → verify: the
live-count panel renders `line_cross`'s dials from the declaration alone, with no
algorithm-specific code in the page. `[TODO]`
3. Shape-agnostic geometry channel (REQ-147) → verify: dragging the line still works; a
two-shape stub algorithm is adjustable through the same endpoint. `[TODO]`
4. `count_runs` keyed by `(project, video, algorithm)` (REQ-149) → verify: the same video
counted by two algorithms yields two rows and two accuracy figures. `[TODO]`
5. The possession counter (REQ-148) → verify: on the hand-counted clip it beats `line_cross`
on sacks that are occluded by the carrier and on sacks thrown in by the sender. **Blocked**
until the detector emits a `person` class and one clip has per-event truth. `[TODO]`
## Task — Exemplar prompting in the auto-annotate modal (REQ-171, REQ-172) `[DONE]`
1. `Sam3Engine.detect_with_exemplars` — one `set_image`, prompts looped over it, boxes
appended to one prompt only → verify: a negative box owned by `sack` sitting on a truck
leaves the truck detections untouched, while the same box owned by `truck` suppresses
them. `[DONE]` — on frame 86031 of batch 594: text-only `{truck: 5}`, owned-by-sack
`{truck: 5}`, owned-by-truck `{}`. The `reset_all_prompts` before each prompt is what
stops the leak; `state["geometric_prompt"]` survives `set_text_prompt` otherwise.
2. `exemplars` + `exemplar_class_name` through `labeling.label_image` → `preview.py` →
`POST /api/batches/{id}/preview` → verify: an unknown class name falls back to plain text
rather than attaching the boxes to whichever class is first. `[DONE]` — 17 shapes for
both text-only and `exemplar_class_name: "nonexistent"`.
3. `preview_frame` moved out of `autolabel.py` into `preview.py` → verify: `autolabel.py` is
back under the 400-line limit and the job path still imports. `[DONE]` — 261 and 151
lines; container starts and registers the `autolabel` handler.
4. Editable class prompt in the modal, saved to `project_classes.prompt` (REQ-171) →
verify: a PATCH round-trips and the Projects page shows the new text. `[DONE]` — class 2
`box → cardboard box → box` via the existing `PATCH /api/projects/{id}`; no new endpoint.
5. `ExemplarCanvas.jsx` drag/shift-drag/undo/clear with 250 ms debounced re-run, and the
modal split into `PreviewShapes.jsx` + `ClassPromptPanel.jsx` to stay under 400 lines →
verify: `npm run build` clean, every file under the limit. `[DONE]` — 398 / 126 / 137 /
64 lines, build green, both containers redeployed.
**Deliberately not built:** exemplars in the batch job. SAM3's geometric prompts pool
features from the current image, so a box drawn on frame 1 asks about whatever sits at those
coordinates on frame 400. The batch job stays text-only; the exemplars exist to find the text
that works.
## Task — Exemplar-driven labeling in the review editor (REQ-173, REQ-174) `[DONE]`
1. `backend/exemplar.py` — pool → one SAM3 pass (class prompt + boxes) → rewrite that class
on that frame → verify: on frame 55446 (batch 426, `sack`), one positive drawn from an
existing box gives 52 class-0 shapes, exactly 1 of them `manual` with the drawn geometry,
and the frame's class-1 shapes are untouched. `[DONE]` — verified; warm pass 0.4 s, first
pass 7.6 s (model load).
2. Negative exemplars delete what they cover (REQ-174) → verify: shift-drag over one of the
detections and no `auto` shape overlapping it by ≥ 0.3 IoU comes back, while the drawn
positive survives. `[DONE]` — max IoU with the negative afterwards 0.078, manual shape
still present.
3. GPU-busy fallback → verify: hold `jobs.gpu_lock`, drag, and the drawn shape is still
stored with `redetected: false` and a legible message. `[DONE]` — "Saved your shape — the
GPU is busy with a background job…", 58 shapes vs 57 before, no exception.
4. `POST /api/frames/{id}/exemplar-label` + `AnnotationCanvas` drag/shift-drag with the pool
drawn as dashed ghosts, 400 ms debounce, undo/clear, and the busy message under the canvas
→ verify: `vite build` clean and every touched file under 400 lines. `[DONE]` — build
green; `exemplar.py` 211, `api/review.py` 141, canvas 287, `useExemplarPool.js` 81. The
pool logic went into that hook rather than into `ReviewPage.jsx`, which was already over
the limit before this task (620 lines) and ends it at 628.
## Task — Filter panel and preview for exemplar runs (REQ-175) `[DONE]`
1. `exemplar.label(..., apply=False)` — dry run by default, returning `shapes` instead of
writing → verify: two previews in a row leave the row count untouched. `[DONE]` — frame
55446 stayed at 57 rows across a default preview (52 shapes) and a filtered one (20).
2. The four filters, applied in the batch job's order (area floor → NMS → cap) → verify:
each one visibly bites on a dense frame. `[DONE]` — from 52 shapes: NMS 0.05 → 32,
min box 0.05 → 1, cap 5 → 5, confidence 0.9 → 15.
3. `apply: true` writes exactly what was previewed → verify: the applied frame matches the
preview count and leaves other classes alone. `[DONE]` — 20 previewed, 20 class-0 shapes
stored (1 of them the drawn `manual` box), the frame's 2 class-1 shapes untouched.
4. `ExemplarFilterPanel.jsx` floating in the canvas corner, sliders re-previewing on 250 ms,
Apply/Discard/Undo/Reset, Enter and Esc bound → verify: `vite build` clean, files under
the limit. `[DONE]` — panel 108, hook 125, canvas 314 lines; build green; both containers
rebuilt and the live endpoint returns `applied: false` for a drag.
5. The class under review hides while its preview is up → verify: a negative exemplar's
effect is visible instead of being masked by the stored box underneath it. `[DONE]` —
frame 55446: 51 detections with one positive, 50 with a negative added; before this the
removed box stayed on screen at 35% opacity and the run looked inert.
**Deliberately not built:** saving the filter values. They describe one frame's run, and the
auto-annotate modal already owns the batch-wide numbers — sharing them would let a tweak made
while reviewing one frame silently change what the next batch job does.
**Deliberately not built:** persisting the pool. It is a prompt about *this* image, so it
dies with the frame, exactly as in REQ-172. What persists is the annotations it produced.
## Task 32 — WebRTC preview for the live counting page (REQ-176, REQ-177) `[DONE]`
The live view cost far more than it should: the backend re-encoded every annotated frame to
JPEG and pushed it over MJPEG, on top of decoding the camera. The camera already reaches the
browser cheaply over WebRTC, so the frames stop travelling through this app entirely.
1. A live source must be a WHEP URL; the RTSP leg is derived → verify: **[DONE]**
`POST .../live-count/start` with `rtsp://192.168.192.96:8554/cam` →
`400 "A live source must be a WebRTC (WHEP) URL…"`; with
`http://192.168.192.96:8889/cam` → `200`, `source: "rtsp://192.168.192.96:8554/cam"`,
`whep_url: "http://192.168.192.96:8889/cam/whep"`, `preview: "webrtc"`.
2. The AI counts from that stream → verify: **[DONE]** 185 frames in 49 s off the live
camera, `error: ""`. That rate is the link's, not the model's — see below.
3. No JPEG is encoded for a WebRTC session → verify: **[DONE]** `GET /api/live-count/stream`
downloaded 0 bytes during a running WebRTC session, and now answers `409`.
4. The overlay feed carries what the model saw, and tracks the line live → verify:
**[DONE]** `GET /api/live-count/overlay` returned 27 boxes with ids and confidences;
after `PATCH /api/live-count/line {"line_y":300}` the feed reported `line.y: 300`.
5. The 400-line limit holds → verify: **[DONE]** `live_count.py` was already 467 lines, so
the transport layer went to `live_source.py` (120) and the MJPEG overlay to
`live_render.py` (60), leaving it at 393. On the frontend the preview moved to
`LiveVideoPanel.jsx` and the slider table to `liveCountFields.js`, leaving
`LiveCountPage.jsx` at 383. `npm run build` passes.
**Not verified here:** the WHEP handshake in a real browser. The endpoint was confirmed live
(`POST http://192.168.192.96:8889/cam/whep` answers, rejecting a deliberately malformed SDP
with `400`), but the negotiation itself needs a browser, not curl.
### Where the live FPS actually goes — measured, 2026-08-19
The live session runs at 4-6 fps and it is not the model. Measured in the backend container
against `rtsp://192.168.192.96:8554/cam`:
| Stage | Rate |
|---|---|
| ByteTrack + YOLO inference | **205 fps** |
| `cv2.resize` to 1280x720 | 5348 fps |
| Decode from RTSP | **6.4 fps** |
The camera is 704x576 HEVC at 350 kbit/s — nothing about it is expensive. The link is: the
route to the streaming server is a ZeroTier VPN measuring **15% packet loss** and a 41-104 ms
round trip. The comment in `live_source.py` claiming the cost was "decoding 1080p on the CPU"
was simply wrong and has been corrected; so has the hint on the page.
Transport was changed to UDP and changed back, because the measurement contradicts the
theory. Through the **ffmpeg CLI**, UDP wins as expected — 16 fps at 1.00x realtime against
TCP's 6.8 fps at 0.52x. Through **OpenCV** it loses: tcp 6.4 fps, udp+socket buffer 4.4, bare
udp 2.4, and a live session on UDP showed 18-second stalls waiting for a keyframe. OpenCV
drops what it cannot reassemble instead of showing it, so the loss lands as missing frames.
`RTSP_TRANSPORT` is left as an env override, defaulting to `tcp`.
**Not fixable in this repo.** Inference has ~50x the headroom the link delivers, so nothing
in the app is worth optimising. The lever is where the counter runs: next to MediaMTX it
would count at the camera's full rate. Worth checking whether the ZeroTier path is relayed
rather than direct (`zerotier-cli peers` — a `RELAY` row explains both the loss and the RTT).
**Deliberately not built:** an aiortc/WHEP client in the backend. It would be "WebRTC only"
end to end, but the decode cost is identical to RTSP and it adds ICE and keyframe-loss
failure modes to the counting path. The saving was always on the browser side.
## Known open points
- *Not closed by any task, by choice:* **any rebuild kills the running job.** Task 14's resume
+1550
View File
File diff suppressed because it is too large. Load diff