feat: sync live-count, exemplar annotation modules, and update .gitignore
This commit is contained in:
1 parent
b6624eeff9
commit
ac95674c07
39 files changed
+3979
-444
No files matched your search
Binary file not shown.
+115
@@ -123,9 +123,14 @@ box.
|
||||
| `batches.py` | batch lifecycle | **new** |
|
||||
| `review.py` | annotation CRUD, frame status, click-assist | **new** |
|
||||
| `autolabel.py` | the SAM3 job over a whole batch | **new** |
|
||||
| `preview.py` | one-frame preview for the auto-annotate modal (REQ-171,172) | **new** |
|
||||
| `exemplar.py` | exemplar-driven labeling in the review editor (REQ-173,174) | **new** |
|
||||
| `dataset.py` | merge into the master dataset, stable split | **new** |
|
||||
| `evaluate.py` | validate base vs new model | **new** |
|
||||
| `hardware.py` | VRAM detection → training defaults | **new** |
|
||||
| `live_count.py` | the live counting session: capture → track → count | **new** |
|
||||
| `live_source.py` | what a source is, how it opens, WHEP↔RTSP (REQ-176) | **new** |
|
||||
| `live_render.py` | the MJPEG overlay, archive-file preview only (REQ-177) | **new** |
|
||||
| `api/` | the FastAPI routes, one module per domain | **new** |
|
||||
|
||||
Removed: `uploads.py`, `static/index.html`, and the old flow's endpoints.
|
||||
@@ -167,6 +172,9 @@ POST /api/projects/{id}/batches # {rel, start_sec, end_sec, fps}
|
||||
GET /api/batches/{id} # status + review progress (REQ-045)
|
||||
GET /api/batches/{id}/frames # frames + statuses
|
||||
POST /api/batches/{id}/autolabel # {threshold} → job (REQ-030,032,034)
|
||||
POST /api/batches/{id}/preview # one frame, run now, nothing written;
|
||||
# +exemplars[] {box:[cx,cy,w,h], positive}
|
||||
# +exemplar_class_name (REQ-171,172)
|
||||
DELETE /api/batches/{id}/classes/{class_id}/annotations # clear all shapes of class in batch (REQ-046)
|
||||
POST /api/batches/{ids}/approve # one or many, comma-separated → one merge job (REQ-131)
|
||||
GET /api/batches/{ids}/triage/summary # one or many, comma-separated (REQ-130)
|
||||
@@ -180,6 +188,7 @@ POST /api/frames/{id}/annotations # add a manual shape (REQ-042)
|
||||
PATCH /api/annotations/{id} # move/resize/reclass
|
||||
DELETE /api/annotations/{id}
|
||||
POST /api/frames/{id}/assist # click/box → SAM3 shape (REQ-043)
|
||||
POST /api/frames/{id}/exemplar-label # drawn pool → re-detect one class (REQ-173,174)
|
||||
POST /api/frames/{id}/status # approved | rejected | pending (REQ-041)
|
||||
|
||||
POST /api/projects/{id}/train # → train job (REQ-060,061,062)
|
||||
@@ -187,11 +196,36 @@ GET /api/projects/{id}/models # versions + metrics (REQ-063,06
|
||||
GET /api/models/{id}/weights # download best.pt
|
||||
POST /api/models/{id}/promote # make it the project's base model (REQ-064)
|
||||
|
||||
GET /api/projects/{id}/live-count/models # weights this project can count with
|
||||
POST /api/projects/{id}/live-count/start # {source|source_rel, model_path, dials}
|
||||
# source must be a WHEP URL (REQ-176)
|
||||
POST /api/live-count/stop
|
||||
PATCH /api/live-count/line # move the line mid-session
|
||||
GET /api/live-count/status # counts + preview: "webrtc" | "mjpeg"
|
||||
GET /api/live-count/overlay # boxes/line/counts for the canvas (REQ-177)
|
||||
GET /api/live-count/stream # MJPEG; 409 on a WebRTC session (REQ-177)
|
||||
|
||||
GET /api/jobs?project_id=… REQ-070,071
|
||||
GET /api/jobs/{id}
|
||||
POST /api/jobs/{id}/cancel
|
||||
```
|
||||
|
||||
### Live counting (REQ-176, REQ-177)
|
||||
|
||||
One camera, one ingest on the streaming server, two consumers:
|
||||
|
||||
```
|
||||
camera ──▶ MediaMTX ──┬── WHEP :8889/cam/whep ──▶ browser <video> + <canvas> overlay
|
||||
└── RTSP :8554/cam ──▶ backend decode → YOLO → counter
|
||||
└─▶ GET /api/live-count/overlay
|
||||
```
|
||||
|
||||
The user types only the WHEP URL; `live_source.whep_to_rtsp` derives the RTSP one, with the
|
||||
ports coming from `MEDIAMTX_RTSP_PORT` / `MEDIAMTX_WHEP_PATH`. Because the browser plays the
|
||||
camera directly, a live session encodes no JPEG at all — `live_render.py` runs only for
|
||||
archive files, which have no WebRTC leg. The canvas draws exactly what `live_render.py`
|
||||
would have burned in, so both previews describe the same session.
|
||||
|
||||
## Job flows
|
||||
|
||||
**extract (REQ-020…023).** `ffmpeg -ss <start> -to <end> -i <video> -vf fps=<n> -q:v 2
|
||||
@@ -277,6 +311,87 @@ Pages:
|
||||
The canvas editor is hand-written; the normalized-coordinate conventions already exist in
|
||||
`sessions.py` (`annotations_payload`, `detections_payload`) as a reference.
|
||||
|
||||
**Auto-annotate modal (REQ-171, REQ-172).** `AutoAnnotateModal.jsx` splits into
|
||||
`PreviewShapes.jsx` (the result overlay, shared with the mass modal),
|
||||
`ClassPromptPanel.jsx` (class chips + the editable SAM3 prompt) and `ExemplarCanvas.jsx`
|
||||
(the drag-to-draw layer). All three overlays and the `<img>` share one shrink-wrapped
|
||||
`position: relative` wrapper — they are sized to it, so nothing else may sit inside it or
|
||||
every box shifts off the pixels it describes.
|
||||
|
||||
With SAM3 one selected chip is *active*: it owns the prompt field and any exemplars drawn on
|
||||
the frame, so its chip is a pair of buttons — the name activates, the `×` deselects.
|
||||
Exemplars are normalized `[cx, cy, w, h]`, sent only to `/preview`, and dropped whenever the
|
||||
frame or the active class changes. A redraw is debounced 250 ms and re-runs the whole prompt
|
||||
set from empty, which is also how undo works — SAM3 can only append geometric prompts.
|
||||
|
||||
**Exemplar-driven labeling in review (REQ-173, REQ-174, REQ-175).** In `draw` mode a drag on
|
||||
`AnnotationCanvas` is an exemplar, not a rectangle: `onExemplar(box, positive)` where
|
||||
`positive` is `!event.shiftKey`. `hooks/useExemplarPool.js` keeps the pool in a ref as well
|
||||
as state — the ref is what gets sent, so a drag that lands mid-flight is never lost — and posts
|
||||
the **whole pool** to `/exemplar-label` 400 ms after the last drag. One pass runs at a time;
|
||||
a drag arriving during a pass sets `rerunWanted` so exactly one rerun follows instead of a
|
||||
queue. The pool is cleared by a frame change or a class change, and `Undo example` re-runs
|
||||
with the shortened pool.
|
||||
|
||||
The run is a **dry run by default** (REQ-175). `ExemplarFilterPanel.jsx` is passed to
|
||||
`ReviewSidebar` and rendered above the class list — not over the frame, which is where its
|
||||
proposals are drawn. It is mounted only while a run is undecided (`pool.active`): the first
|
||||
drag opens it, Apply and Discard close it, and the pool outlives it. Its four sliders
|
||||
re-preview on a 250 ms debounce. While a preview is up the canvas **hides the stored shapes
|
||||
of the class under review** — the run replaces them wholesale, and leaving them on screen made
|
||||
a rejected detection look like it had never gone; other classes stay, dimmed
|
||||
(`svg.previewing`). Proposals draw dashed on top, green for the boxes the user drew and
|
||||
class-colored for what SAM3 found.
|
||||
`Apply` re-runs with `apply: true` rather than posting the previewed geometry back — SAM3 is
|
||||
deterministic for a pool and a threshold, and the browser should not be the authority on what
|
||||
gets stored. `Discard` truncates the pool to `appliedRef`, the length it had at the last
|
||||
successful apply, so a rejected run leaves neither shapes nor prompts behind. A successful
|
||||
apply drops the negatives it sent and keeps the positives (REQ-174); drags that landed while
|
||||
that request was in flight are not part of it and stay at the end of the pool, with a rerun
|
||||
queued for them.
|
||||
|
||||
`.canvas-wrap`'s overlay rules are scoped to its direct child (`> svg`): they set
|
||||
`position: absolute; width: 100%`, which any nested SVG — an icon in a panel, say — would
|
||||
otherwise inherit and stretch across the whole frame.
|
||||
|
||||
Request/response:
|
||||
|
||||
```
|
||||
POST /api/frames/{id}/exemplar-label
|
||||
{ "exemplars": [{"box": [x0,y0,x1,y1], "positive": true}, …], # normalized xyxy
|
||||
"class_id": 0,
|
||||
"threshold": 0.5, "iou_threshold": 0.8, # the panel (REQ-175)
|
||||
"min_box_frac": 0.002, "max_detections": 100,
|
||||
"apply": false }
|
||||
→ { "shapes": [{"geometry": …, "score": …, "source": "manual"|"auto"}, …],
|
||||
"applied": false, "redetected": true, "message": null,
|
||||
"annotations": null } # the frame, on apply only
|
||||
```
|
||||
|
||||
`backend/exemplar.py` converts each box to SAM3's normalized `[cx, cy, w, h]`, runs one
|
||||
`open_state` + `apply_prompts` with the class's stored `prompt` as text plus every exemplar,
|
||||
then rewrites the frame **for that class only**:
|
||||
|
||||
- positives are stored first as `source='manual'` — the literal rectangle in a `bbox`
|
||||
project, the mask polygon of whatever SAM3 found inside it (IoU ≥ `SNAP_IOU`) in a
|
||||
`polygon` one;
|
||||
- detections overlapping a negative by ≥ `NEGATIVE_IOU` (0.3) are dropped, and ones
|
||||
overlapping a positive by ≥ `DUPLICATE_IOU` (0.6) are dropped as the user's own shape
|
||||
already covers them;
|
||||
- what remains is written as `source='auto'`, so a later batch re-run replaces it (REQ-034)
|
||||
while the drawn shapes survive.
|
||||
|
||||
The panel's filters run before any of that, in the order the batch job uses them: area floor,
|
||||
then NMS (`labeling.deduplicate`), then the cap on how many survive.
|
||||
|
||||
The delete-and-reinsert happens in one `db.cursor()` transaction, so the frame is never
|
||||
briefly empty. Other classes on the frame are never touched. The GPU lock is taken with a
|
||||
**5 s** timeout — shorter than `assist`'s 20 s because this fires from a mouse gesture; on
|
||||
timeout the response carries `redetected: false`, a message the panel shows, and the drawn
|
||||
boxes alone as the preview. Applying that run **appends** the drawn boxes and honours the
|
||||
negatives instead of taking the replace path — with no detections to put back, replacing
|
||||
would wipe the class and leave only the drawings.
|
||||
|
||||
## Docker (REQ-072)
|
||||
|
||||
- `Dockerfile` — python 3.12, `ffmpeg`, `uv`, CUDA torch, `uv pip install -e sam3/`. The
|
||||
|
||||
@@ -100,6 +100,19 @@ changes.
|
||||
found nothing on (REQ-033) writes no annotations, so a resume re-does it; that is accepted
|
||||
rather than tracked.
|
||||
|
||||
- **REQ-171** — In the auto-annotate modal the SAM3 text prompt of each selected class is
|
||||
editable in place, next to the live preview. Saving it writes
|
||||
`project_classes.prompt` — the same field the Projects page edits — so the batch job and
|
||||
every later run send that text. The preview is the tuning surface; the stored prompt is the
|
||||
artifact it produces.
|
||||
- **REQ-172** — On the previewed frame the user can drag **positive** and **negative** box
|
||||
exemplars (shift-drag for negative). They are appended to the active class's text prompt,
|
||||
re-run immediately, and can be undone or cleared. Exemplars are a **tuning aid only**: they
|
||||
are never written as annotations and never carried into the batch job, because SAM3's
|
||||
geometric prompts pool features from the current image — replaying them on another frame
|
||||
would ask about whatever happens to sit at those coordinates there. They belong to exactly
|
||||
one class, so a new frame or a new active class discards them.
|
||||
|
||||
## E. Review & correction
|
||||
|
||||
- **REQ-040** — The user reviews frames one at a time, with fast navigation (left/right
|
||||
@@ -117,6 +130,53 @@ changes.
|
||||
- **REQ-046** — The user can delete/clear all annotations of a specific class across all frames in
|
||||
the current batch from the Review editor.
|
||||
|
||||
- **REQ-173** — In the review editor a plain drag on the canvas is an **exemplar-driven
|
||||
label**, not just a rectangle. It proposes, in one action — and REQ-175's Apply is what
|
||||
makes any of it real — that (a) the drawn shape becomes a `manual` annotation of the
|
||||
active class — snapped to a SAM3 polygon first when the batch's
|
||||
`label_type` is `polygon`, since a rectangle is a bad polygon label — (b) appends the box
|
||||
to the frame's positive exemplar pool for that class, and (c) re-runs SAM3 over the whole
|
||||
frame with the class's text prompt plus the pooled exemplars, deleting every existing shape
|
||||
of that class on the frame and writing the detections in its place, then re-inserting the
|
||||
pooled exemplar shapes verbatim so the user's own drawings always survive. The pool is
|
||||
**frame-local and ephemeral** for the same reason as REQ-172 — SAM3's geometric prompts pool
|
||||
features from the current image — so leaving the frame or switching the active class clears
|
||||
it; the annotations it produced persist like any other. If the GPU lock (REQ-070) is not
|
||||
free, the run comes back with the drawn shapes alone and says so, so the user can still
|
||||
file them (REQ-175) and labeling is never blocked by a background job.
|
||||
- **REQ-174** — **Shift**-drag in the review editor adds a **negative** exemplar. It is never
|
||||
stored as an annotation; it deletes any existing shape of the active class that overlaps it,
|
||||
and it is sent as a negative box in the REQ-173 re-detect. It is the "not this, and not
|
||||
things like this" gesture, so it doubles as a delete. A negative is **spent on Apply**: the
|
||||
frame it was applied to no longer carries what it rejected, so the drawing is dropped from
|
||||
the pool while the positives stay on as prompts.
|
||||
- **REQ-175** — An exemplar drag **previews**; it never writes on its own. The run's result
|
||||
is drawn over the frame as proposals and a small panel floats on the canvas with the four
|
||||
filters that decide what survives — confidence, NMS overlap, minimum box size, maximum
|
||||
shapes — each re-running the preview as it moves. **Apply** writes the previewed set,
|
||||
**Discard** rewinds the pool to whatever is already on the frame and leaves it untouched.
|
||||
The panel is scoped to this gesture: its values are not stored, not shared with the
|
||||
auto-annotate modal, and reset with the frame. Defaults are confidence `0.5`, NMS `0.8`,
|
||||
min box `0.002`, max `100` — deliberately permissive, because on a dense frame an
|
||||
aggressive NMS or area floor deletes real, touching objects rather than duplicates.
|
||||
|
||||
## E4. Live counting preview
|
||||
|
||||
- **REQ-176** — A **live** source on the Live Count page is a **WebRTC (WHEP) URL** and
|
||||
nothing else; an RTSP URL is rejected with a message saying so. The backend derives the
|
||||
RTSP leg of the same streaming-server path from it (`http://host:8889/cam` →
|
||||
`rtsp://host:8554/cam`) and counts from that: WebRTC is what makes the browser preview
|
||||
cheap, but pulling it into Python would add ICE and a jitter buffer on top of the identical
|
||||
H.264 decode. One ingest on the streaming server, two consumers. The ports are read from
|
||||
the environment (`MEDIAMTX_RTSP_PORT`, `MEDIAMTX_WHEP_PATH`), never hardcoded. Archive
|
||||
files are unaffected — they are still opened as files.
|
||||
- **REQ-177** — A live session is **watched over WebRTC**, played straight from the streaming
|
||||
server by the browser: the frames never pass through this app and it encodes no JPEG for
|
||||
them. What the model saw — boxes, ids, confidences, the counting line and its band, the
|
||||
ignored region, the running totals — is served as geometry from
|
||||
`GET /api/live-count/overlay` and drawn on a canvas over the video. The MJPEG endpoint
|
||||
remains the preview for **archive files** only, and refuses a WebRTC session.
|
||||
|
||||
## F. Master dataset
|
||||
|
||||
- **REQ-050** — Approving a batch **merges** its approved frames and their labels into the
|
||||
@@ -164,6 +224,31 @@ changes.
|
||||
and the reason it did or did not count, so a miss can be attributed to the model, the
|
||||
tracker, or the counter.
|
||||
|
||||
- **REQ-145** — Counting algorithms are **pluggable**. Each registers under a stable id
|
||||
(`line_cross`, `possession`) and the session constructs one by id. The `Counter` protocol
|
||||
in `src/interfaces.py` is the contract, corrected to match reality: `update()` returns the
|
||||
frame's count events, not `None`. Adding an algorithm must not require editing
|
||||
`live_count.py` or `counting_bench.py`.
|
||||
- **REQ-146** — Each algorithm **declares its own parameters** — name, type, default, range —
|
||||
and an endpoint serves that declaration, mirroring `live-count/models`. The frontend renders
|
||||
its controls from the declaration and hardcodes no per-algorithm parameter list. The start
|
||||
request carries `algorithm` plus an opaque `params` object validated against the
|
||||
declaration, replacing today's flat line-specific fields.
|
||||
- **REQ-147** — Geometry is generalised from a line to a **named shape set**. `line_cross`
|
||||
declares one horizontal segment; `possession` declares a bed polygon and an approach zone.
|
||||
The editor's drag channel (`move_line`) becomes shape-agnostic, so any algorithm's geometry
|
||||
is adjustable live without a new endpoint.
|
||||
- **REQ-148** — The **possession counter**: every sack track carries an `owner_id`, the person
|
||||
track it currently overlaps, or none when at rest. A count fires on an ownership change that
|
||||
crosses the bed boundary — person-outside to bed, or person-outside to person-inside.
|
||||
Ownership is sticky with hysteresis, so occlusion by the carrier's back and the unowned
|
||||
mid-air phase of a thrown sack do not break it. This requires a `person` class alongside
|
||||
`sack` from the detector.
|
||||
- **REQ-149** — Every count run records **which algorithm and parameter set** produced it, and
|
||||
accuracy is comparable per algorithm against the same ground truth. Switching algorithms
|
||||
adds results, it never invalidates stored ones — so `count_runs` is keyed by
|
||||
`(project, video, algorithm)`, not by video alone.
|
||||
|
||||
## F4. Counting accuracy bench
|
||||
|
||||
- **REQ-150** — A page lists every archive video as a row: date, batch, length, and the
|
||||
@@ -179,6 +264,23 @@ changes.
|
||||
which is what makes counting a 30-minute video practical. A run records the parameters and
|
||||
model it used.
|
||||
|
||||
- **REQ-154** — Ground truth can be **imported in bulk** from the operations sheet
|
||||
(`./GT.xlsx`, `DATA MUAT PAKAN PER LINE`). The camera watches **Line 1**; Line 2 is
|
||||
recorded for completeness but never scored. Each sheet is one working day; a row is one
|
||||
truck with a `BAG` count, a `DUS` count and a plate.
|
||||
- **REQ-155** — `BAG` (sacks) and `DUS` (boxes) are **separate commodities**, counted and
|
||||
scored separately. A box already resting in the truck bed is a legitimate object of a
|
||||
different class, not a detection fault.
|
||||
- **REQ-156** — An import never silently guesses. Recordings are aligned to sheet rows by
|
||||
start time against row order, the proposed pairing is **shown for human confirmation**
|
||||
before anything is written, and each imported value records that it came from the sheet
|
||||
rather than from a hand count. A recording that merged two trucks
|
||||
(`BATCH_MERGE_THRESHOLD_SECONDS`) is flagged, not paired.
|
||||
- **REQ-157** — Sheet values are **order quantities, not hand counts** — 67% of them are
|
||||
exactly 160 or 180 — so they score aggregate accuracy across many trucks and never
|
||||
adjudicate a single video. Per-event truth for algorithm comparison comes from a
|
||||
hand-counted clip, held separately.
|
||||
|
||||
## F5. Real recording times and working days
|
||||
|
||||
- **REQ-160** — Each recording's start time is read from the timestamp the camera burns into
|
||||
|
||||
+167
@@ -1040,6 +1040,173 @@ two will disagree.
|
||||
`2026-08-06/batch4`, `2026-08-06/batch9`, `2026-08-14/batch016` — likely truncated) and 11 were
|
||||
read with low confidence. Both are flagged amber in the table and accept a hand-typed time.
|
||||
|
||||
## Task — Ground truth import from the ops sheet (REQ-154…157)
|
||||
|
||||
1. Parse `docs/GT.xlsx` into rows → verify: 6 sheets (10–15 Aug 2026), Line 1 only, stopping
|
||||
at the first blank plate so the inline totals row is not read as a truck. Expected Line 1
|
||||
bag totals: 4780 / 4322 / 4365 / 5800 / 5645 / 9155. `[TODO]`
|
||||
2. `ground_truth_bag` / `ground_truth_dus` + `gt_source` on `count_runs` (REQ-155, REQ-156) →
|
||||
verify: migration runs on the live DB, existing hand-typed values survive as
|
||||
`gt_source='manual'`. `[TODO]`
|
||||
3. Alignment preview with human confirmation (REQ-156) → verify: a dry run on 14 Aug proposes
|
||||
26 recordings against 32 Line-1 trucks, flags the shortfall, and writes nothing until
|
||||
confirmed. `[TODO]`
|
||||
4. Bench scores bag and box separately (REQ-155) → verify: the accuracy row shows both, and
|
||||
totals only over rows that have a ground truth. `[TODO]`
|
||||
|
||||
## Task — Pluggable counting algorithms (REQ-145…149)
|
||||
|
||||
1. Fix the `Counter` protocol and register `line_cross` behind it (REQ-145) → verify: a live
|
||||
session on a known clip returns **the same counts as before** the refactor — this step
|
||||
changes no behaviour. `[TODO]`
|
||||
2. Parameter declaration endpoint + generic frontend controls (REQ-146) → verify: the
|
||||
live-count panel renders `line_cross`'s dials from the declaration alone, with no
|
||||
algorithm-specific code in the page. `[TODO]`
|
||||
3. Shape-agnostic geometry channel (REQ-147) → verify: dragging the line still works; a
|
||||
two-shape stub algorithm is adjustable through the same endpoint. `[TODO]`
|
||||
4. `count_runs` keyed by `(project, video, algorithm)` (REQ-149) → verify: the same video
|
||||
counted by two algorithms yields two rows and two accuracy figures. `[TODO]`
|
||||
5. The possession counter (REQ-148) → verify: on the hand-counted clip it beats `line_cross`
|
||||
on sacks that are occluded by the carrier and on sacks thrown in by the sender. **Blocked**
|
||||
until the detector emits a `person` class and one clip has per-event truth. `[TODO]`
|
||||
|
||||
## Task — Exemplar prompting in the auto-annotate modal (REQ-171, REQ-172) `[DONE]`
|
||||
|
||||
1. `Sam3Engine.detect_with_exemplars` — one `set_image`, prompts looped over it, boxes
|
||||
appended to one prompt only → verify: a negative box owned by `sack` sitting on a truck
|
||||
leaves the truck detections untouched, while the same box owned by `truck` suppresses
|
||||
them. `[DONE]` — on frame 86031 of batch 594: text-only `{truck: 5}`, owned-by-sack
|
||||
`{truck: 5}`, owned-by-truck `{}`. The `reset_all_prompts` before each prompt is what
|
||||
stops the leak; `state["geometric_prompt"]` survives `set_text_prompt` otherwise.
|
||||
2. `exemplars` + `exemplar_class_name` through `labeling.label_image` → `preview.py` →
|
||||
`POST /api/batches/{id}/preview` → verify: an unknown class name falls back to plain text
|
||||
rather than attaching the boxes to whichever class is first. `[DONE]` — 17 shapes for
|
||||
both text-only and `exemplar_class_name: "nonexistent"`.
|
||||
3. `preview_frame` moved out of `autolabel.py` into `preview.py` → verify: `autolabel.py` is
|
||||
back under the 400-line limit and the job path still imports. `[DONE]` — 261 and 151
|
||||
lines; container starts and registers the `autolabel` handler.
|
||||
4. Editable class prompt in the modal, saved to `project_classes.prompt` (REQ-171) →
|
||||
verify: a PATCH round-trips and the Projects page shows the new text. `[DONE]` — class 2
|
||||
`box → cardboard box → box` via the existing `PATCH /api/projects/{id}`; no new endpoint.
|
||||
5. `ExemplarCanvas.jsx` drag/shift-drag/undo/clear with 250 ms debounced re-run, and the
|
||||
modal split into `PreviewShapes.jsx` + `ClassPromptPanel.jsx` to stay under 400 lines →
|
||||
verify: `npm run build` clean, every file under the limit. `[DONE]` — 398 / 126 / 137 /
|
||||
64 lines, build green, both containers redeployed.
|
||||
|
||||
**Deliberately not built:** exemplars in the batch job. SAM3's geometric prompts pool
|
||||
features from the current image, so a box drawn on frame 1 asks about whatever sits at those
|
||||
coordinates on frame 400. The batch job stays text-only; the exemplars exist to find the text
|
||||
that works.
|
||||
|
||||
## Task — Exemplar-driven labeling in the review editor (REQ-173, REQ-174) `[DONE]`
|
||||
|
||||
1. `backend/exemplar.py` — pool → one SAM3 pass (class prompt + boxes) → rewrite that class
|
||||
on that frame → verify: on frame 55446 (batch 426, `sack`), one positive drawn from an
|
||||
existing box gives 52 class-0 shapes, exactly 1 of them `manual` with the drawn geometry,
|
||||
and the frame's class-1 shapes are untouched. `[DONE]` — verified; warm pass 0.4 s, first
|
||||
pass 7.6 s (model load).
|
||||
2. Negative exemplars delete what they cover (REQ-174) → verify: shift-drag over one of the
|
||||
detections and no `auto` shape overlapping it by ≥ 0.3 IoU comes back, while the drawn
|
||||
positive survives. `[DONE]` — max IoU with the negative afterwards 0.078, manual shape
|
||||
still present.
|
||||
3. GPU-busy fallback → verify: hold `jobs.gpu_lock`, drag, and the drawn shape is still
|
||||
stored with `redetected: false` and a legible message. `[DONE]` — "Saved your shape — the
|
||||
GPU is busy with a background job…", 58 shapes vs 57 before, no exception.
|
||||
4. `POST /api/frames/{id}/exemplar-label` + `AnnotationCanvas` drag/shift-drag with the pool
|
||||
drawn as dashed ghosts, 400 ms debounce, undo/clear, and the busy message under the canvas
|
||||
→ verify: `vite build` clean and every touched file under 400 lines. `[DONE]` — build
|
||||
green; `exemplar.py` 211, `api/review.py` 141, canvas 287, `useExemplarPool.js` 81. The
|
||||
pool logic went into that hook rather than into `ReviewPage.jsx`, which was already over
|
||||
the limit before this task (620 lines) and ends it at 628.
|
||||
|
||||
## Task — Filter panel and preview for exemplar runs (REQ-175) `[DONE]`
|
||||
|
||||
1. `exemplar.label(..., apply=False)` — dry run by default, returning `shapes` instead of
|
||||
writing → verify: two previews in a row leave the row count untouched. `[DONE]` — frame
|
||||
55446 stayed at 57 rows across a default preview (52 shapes) and a filtered one (20).
|
||||
2. The four filters, applied in the batch job's order (area floor → NMS → cap) → verify:
|
||||
each one visibly bites on a dense frame. `[DONE]` — from 52 shapes: NMS 0.05 → 32,
|
||||
min box 0.05 → 1, cap 5 → 5, confidence 0.9 → 15.
|
||||
3. `apply: true` writes exactly what was previewed → verify: the applied frame matches the
|
||||
preview count and leaves other classes alone. `[DONE]` — 20 previewed, 20 class-0 shapes
|
||||
stored (1 of them the drawn `manual` box), the frame's 2 class-1 shapes untouched.
|
||||
4. `ExemplarFilterPanel.jsx` floating in the canvas corner, sliders re-previewing on 250 ms,
|
||||
Apply/Discard/Undo/Reset, Enter and Esc bound → verify: `vite build` clean, files under
|
||||
the limit. `[DONE]` — panel 108, hook 125, canvas 314 lines; build green; both containers
|
||||
rebuilt and the live endpoint returns `applied: false` for a drag.
|
||||
|
||||
5. The class under review hides while its preview is up → verify: a negative exemplar's
|
||||
effect is visible instead of being masked by the stored box underneath it. `[DONE]` —
|
||||
frame 55446: 51 detections with one positive, 50 with a negative added; before this the
|
||||
removed box stayed on screen at 35% opacity and the run looked inert.
|
||||
|
||||
**Deliberately not built:** saving the filter values. They describe one frame's run, and the
|
||||
auto-annotate modal already owns the batch-wide numbers — sharing them would let a tweak made
|
||||
while reviewing one frame silently change what the next batch job does.
|
||||
|
||||
**Deliberately not built:** persisting the pool. It is a prompt about *this* image, so it
|
||||
dies with the frame, exactly as in REQ-172. What persists is the annotations it produced.
|
||||
|
||||
## Task 32 — WebRTC preview for the live counting page (REQ-176, REQ-177) `[DONE]`
|
||||
|
||||
The live view cost far more than it should: the backend re-encoded every annotated frame to
|
||||
JPEG and pushed it over MJPEG, on top of decoding the camera. The camera already reaches the
|
||||
browser cheaply over WebRTC, so the frames stop travelling through this app entirely.
|
||||
|
||||
1. A live source must be a WHEP URL; the RTSP leg is derived → verify: **[DONE]**
|
||||
`POST .../live-count/start` with `rtsp://192.168.192.96:8554/cam` →
|
||||
`400 "A live source must be a WebRTC (WHEP) URL…"`; with
|
||||
`http://192.168.192.96:8889/cam` → `200`, `source: "rtsp://192.168.192.96:8554/cam"`,
|
||||
`whep_url: "http://192.168.192.96:8889/cam/whep"`, `preview: "webrtc"`.
|
||||
2. The AI counts from that stream → verify: **[DONE]** 185 frames in 49 s off the live
|
||||
camera, `error: ""`. That rate is the link's, not the model's — see below.
|
||||
3. No JPEG is encoded for a WebRTC session → verify: **[DONE]** `GET /api/live-count/stream`
|
||||
downloaded 0 bytes during a running WebRTC session, and now answers `409`.
|
||||
4. The overlay feed carries what the model saw, and tracks the line live → verify:
|
||||
**[DONE]** `GET /api/live-count/overlay` returned 27 boxes with ids and confidences;
|
||||
after `PATCH /api/live-count/line {"line_y":300}` the feed reported `line.y: 300`.
|
||||
5. The 400-line limit holds → verify: **[DONE]** `live_count.py` was already 467 lines, so
|
||||
the transport layer went to `live_source.py` (120) and the MJPEG overlay to
|
||||
`live_render.py` (60), leaving it at 393. On the frontend the preview moved to
|
||||
`LiveVideoPanel.jsx` and the slider table to `liveCountFields.js`, leaving
|
||||
`LiveCountPage.jsx` at 383. `npm run build` passes.
|
||||
|
||||
**Not verified here:** the WHEP handshake in a real browser. The endpoint was confirmed live
|
||||
(`POST http://192.168.192.96:8889/cam/whep` answers, rejecting a deliberately malformed SDP
|
||||
with `400`), but the negotiation itself needs a browser, not curl.
|
||||
|
||||
### Where the live FPS actually goes — measured, 2026-08-19
|
||||
|
||||
The live session runs at 4-6 fps and it is not the model. Measured in the backend container
|
||||
against `rtsp://192.168.192.96:8554/cam`:
|
||||
|
||||
| Stage | Rate |
|
||||
|---|---|
|
||||
| ByteTrack + YOLO inference | **205 fps** |
|
||||
| `cv2.resize` to 1280x720 | 5348 fps |
|
||||
| Decode from RTSP | **6.4 fps** |
|
||||
|
||||
The camera is 704x576 HEVC at 350 kbit/s — nothing about it is expensive. The link is: the
|
||||
route to the streaming server is a ZeroTier VPN measuring **15% packet loss** and a 41-104 ms
|
||||
round trip. The comment in `live_source.py` claiming the cost was "decoding 1080p on the CPU"
|
||||
was simply wrong and has been corrected; so has the hint on the page.
|
||||
|
||||
Transport was changed to UDP and changed back, because the measurement contradicts the
|
||||
theory. Through the **ffmpeg CLI**, UDP wins as expected — 16 fps at 1.00x realtime against
|
||||
TCP's 6.8 fps at 0.52x. Through **OpenCV** it loses: tcp 6.4 fps, udp+socket buffer 4.4, bare
|
||||
udp 2.4, and a live session on UDP showed 18-second stalls waiting for a keyframe. OpenCV
|
||||
drops what it cannot reassemble instead of showing it, so the loss lands as missing frames.
|
||||
`RTSP_TRANSPORT` is left as an env override, defaulting to `tcp`.
|
||||
|
||||
**Not fixable in this repo.** Inference has ~50x the headroom the link delivers, so nothing
|
||||
in the app is worth optimising. The lever is where the counter runs: next to MediaMTX it
|
||||
would count at the camera's full rate. Worth checking whether the ZeroTier path is relayed
|
||||
rather than direct (`zerotier-cli peers` — a `RELAY` row explains both the loss and the RTT).
|
||||
|
||||
**Deliberately not built:** an aiortc/WHEP client in the backend. It would be "WebRTC only"
|
||||
end to end, but the decode cost is identical to RTSP and it adds ICE and keyframe-loss
|
||||
failure modes to the counting path. The saving was always on the browser side.
|
||||
|
||||
## Known open points
|
||||
|
||||
- *Not closed by any task, by choice:* **any rebuild kills the running job.** Task 14's resume
|
||||
|
||||
+1550
File diff suppressed because it is too large.
Load diff
Reference in new issue
Block a user