Files
reTraining/docs/ui-spec.md
T
Andrew-AAAA d170cff0e4 feat: add descriptive model naming and inline rename
- Auto-generate model names: {arch}-{labelType}-{epochs}ep-{classNames}-{YYYYMMDD}
- Add PATCH /api/models/{id}/rename endpoint
- Inline rename UI on Models & Training page
- Download filename uses model name instead of v{N}
- DB migration: add name column to model_versions
- Update all docs to reflect new naming convention
2026-09-10 09:14:55 +07:00

78 KiB
Raw Blame History

UI Specification — Dataset Enrichment Tool

Purpose of this document. It describes every screen of the existing frontend so the UI can be rebuilt on another platform without losing behaviour. It is written framework-agnostically: it says what a screen holds, what it does, and what it calls, not how React does it. Where a detail is load-bearing — removing it breaks the pipeline or silently corrupts training data — it is marked.

Read this together with: ./requirements.md (the numbered REQ-xxx this serves) and ./design.md (backend, disk layout, job flows). This file never contradicts them; if it does, design.md wins and this file is the bug.

Sections

  1. What the app is
  2. Global conventions
  3. Application shell
  4. Page specifications
  5. Review editor — deep specification
  6. Simplification rules
  7. API surface index
  8. Redesign brief — prompt siap-pakai

1. What the app is

A dense internal tool for one operator, running on one machine with one GPU. It turns CCTV recordings into a YOLO training dataset and then trains and scores models against it.

The pipeline is linear and the UI exists to walk it:

Video Archive → Trim → Batch (frames) → Auto-annotate → Review → Data Prep → Dataset → Train → Evaluate
     ①            ②         ③              ④              ⑤          ⑥          ⑦        ⑧        ⑨

Each arrow is a page. A user who cannot follow that order in the UI cannot use the app, so the order of the nav must survive any redesign.

It is not a landing page, not a SaaS dashboard, not multi-tenant. There is no login, no onboarding, no marketing surface. Sessions are long (hours of frame review), so density and keyboard reach beat whitespace and animation.


2. Global conventions

2.1 Routing

Hash-based, no router library. #/ + path, one optional query string.

Route Screen Notes
#/projects Projects default when hash is empty
#/projects/{id} Video Archive (Library)
#/projects/{id}/trim/{encodedRel} Trim rel is URI-encoded, e.g. 2026-03-01%2Fbatch4.mp4
#/projects/{id}/batches Batches
#/projects/{id}/review?batch={batchId} Review batch optional → falls back to first batch
#/batches/{batchId} Review alias; batch id in the path
#/projects/{id}/data-prep?batches=1,2,3 Data Preparation batches = the merge selection
#/projects/{id}/datasets Datasets
#/projects/{id}/models Models & Training
#/projects/{id}/live-count Live Counting
#/projects/{id}/counting-bench Counting Accuracy
#/sam3-playground SAM3 Playground no project context

Unknown paths fall back to Projects. ?batches= is parsed as comma-separated positive integers; anything else is dropped.

2.2 Design tokens

Dark by default; light is a toggle stored in localStorage under theme and applied as data-theme on the root element.

--bg              #0b0f19      page ground
--panel           rgba(17,24,39,0.5)
--panel-raised    rgba(30,41,59,0.7)
--border          rgba(255,255,255,0.12)
--text            #f3f4f6
--text-muted      #9ca3af
--text-faint      #6b7280
--accent          #a855f7   (purple; hover #c084fc)
--ok              #10b981
--warn            #f59e0b
--danger          #ef4444
--radius          12px  (--radius-sm 8px)
--space           8px
--font            Inter, system-ui
--mono            ui-monospace, JetBrains Mono, Menlo
--transition      160ms cubic-bezier(.4,0,.2,1)

Semantic colours used inline throughout, and they carry meaning — keep them:

Colour Means
#4ade80 green kept / approved / positive exemplar / better metric
#f87171 red dropped / rejected / negative exemplar / worse metric
#fbbf24 amber needs checking (untrusted clock, held-back frame, unsaved)
#38bdf8 blue selection, current marquee, "this run sees" figures
#c084fc purple machine work in progress (jobs, SAM3, auto-annotation)

2.3 Class colours — data, not decoration

classColor(classId) = ["#f59e0b","#38bdf8","#10b981","#facc15",
                       "#6366f1","#f97316","#ec4899","#9ca3af"][classId % 8]

The same class index must render the same hue in every place it appears: filmstrip badge, canvas stroke, class chip swatch, project card tag, shape list, crop grid border. A palette change is fine; per-page palettes are not.

2.4 Geometry

All shape coordinates are normalized 0–1 against the frame, everywhere, in both directions over the wire.

bbox     { "type": "bbox",    "points": [x0, y0, x1, y1] }
polygon  { "type": "polygon", "points": [[x,y], [x,y], …] }

Two exceptions, both deliberate:

  • Exemplars sent to /batches/{id}/preview use SAM3's own format: [cx, cy, w, h], still normalized.
  • Exemplars sent to /frames/{id}/exemplar-label use [x0, y0, x1, y1] (the backend converts). Do not "unify" these; the backend contract differs per endpoint.

The overlay SVG in the auto-annotate modals uses a 0 0 10000 10000 viewBox with preserveAspectRatio="none" — coordinates are multiplied by 10000. The review canvas instead uses the frame's real pixel dimensions as its viewBox. Both work; both must keep preserveAspectRatio="none" and be sized to a wrapper that hugs the image exactly.

2.5 Jobs and progress

Every long operation is a server-side job. There is one worker thread and one GPU, so jobs queue; the UI never assumes parallelism.

job = { id, project_id, batch_id, type, status, progress, total, message, error, log[] }
type   = extract | autolabel | merge | train | (counting bench, clock scan)
status = queued | running | done | failed | cancelled

Polling rules as implemented:

Screen Interval Condition
Shell health badge 3 s always
Library / Batches 2 s only while ≥1 job is queued/running
Review 2 s only while this batch has an active job
Trim 1 s only while the extraction job is unfinished
Models 2 s only while the training job is unfinished
Live Count status 1 s always (session may start at any time)
Counting Bench 2 s only while a count or clock-scan job runs

When the active-job count drops from >0 to 0, the page reloads its data once — that is how new frames, new shapes and new model versions appear without a manual refresh.

Progress UI must always show progress/total, a bar, and a cancel affordance. A bare spinner is not acceptable for anything that can run for minutes.

2.6 Errors, empty states, confirmations

  • Backend errors arrive as {"detail": "..."}. The UI shows that message, never "Error 500". A page that failed its first load renders the error banner instead of the page; an error after load renders a dismissible banner above the page.
  • Empty states are sentences that name the next action, e.g. "No extracted batches in this project yet. Go to Video Archive to trim frames into batches."
  • Destructive actions confirm, and the confirmation enumerates the consequences. Deleting a class says how many shapes die and how many label files get rewritten. Keep that specificity; a generic "Are you sure?" is a downgrade.

3. Application shell

A fixed 48 px top bar, full-height content area below it, page never scrolls horizontally.

Left — wordmark "Dataset Enrichment", clicking goes to #/projects.

Centre — the nav, in pipeline order. Every project-scoped link uses the current project id, falling back to route.projectId and then 1:

  1. Projects — #/projects
  2. Video Archive — #/projects/{id}
  3. Batches — #/projects/{id}/batches (also active on Review)
  4. Data Preparation — #/projects/{id}/data-prep
  5. Datasets — #/projects/{id}/datasets
  6. Models & Training — #/projects/{id}/models
  7. Live Counting — #/projects/{id}/live-count
  8. Counting Accuracy — #/projects/{id}/counting-bench
  9. SAM3 Playground — #/sam3-playground

Right — health badges from GET /api/health, polled every 3 s, plus the theme toggle:

GPU: <name, stripped of "NVIDIA GeForce" / "Laptop GPU">   |  CPU if none
VRAM: <free>GB                                             |  N/A
SAM3: Ready | Off      ← green when ready

Icons are SVG (Heroicons/Lucide-style), never emoji. (Some emoji survive inside page bodies — 🔄 Reset Auto, 🏷️ Next Shape, 📋 Copy Prev, 🚀 Track 5, 🤖, 📦, ⚡, 🖼️. These are legacy and should become SVG icons in a rebuild; that is the one place the current UI is off-spec.)

An error boundary wraps the page area, keyed by route, and renders a "View Exception Caught" card with the error text and a reload button. Keep it: a crash in one page must not blank the shell.


4. Page specifications

Each page below lists: what it is for, what it loads, its layout, its interactions and states, and its endpoints.


4.1 Projects — #/projects

Purpose. One project = one model you are improving: its base weights, its classes, its dataset. This is the only place a project is created or deleted.

Loads on mount. GET /api/projects → { projects: [...], video_root_default: "/videos" }.

Project shape:

{ id, name, label_type: "bbox"|"polygon", label_type_locked: bool,
  base_model_path, secondary_model_path, base_model_fallback: "yolo11n.pt",
  video_root, batch_count,
  dataset: { train, val },
  classes: [ { class_id, name, prompt, annotation_count } ] }

Layout. Page head (title + one-line description + "New project" button) → optional create-form panel → responsive card grid of projects.

New project form (inline panel, replaces the button while open):

Field Control Rules
Name text, autofocus, required
Label type select: bbox (YOLO detect) / polygon (YOLO segment) hint: fixed once the first batch is merged
Val split — every Nth frame number 0–50, default 5
Video archive root text, required, default from video_root_default hint: layout <date>/<batch>.mp4, read-only
Classes repeatable rows: swatch + name + SAM3 prompt + remove at least one row; blank names are dropped on submit; prompt defaults to the name

Actions: Create project (submits), Cancel, Add class.

Project card shows: name, a tag reading bbox / polygon or locked: <type> (with a tooltip explaining the lock), then a definition list:

  • Base models — "N model(s) inserted" plus which slots are filled (Model 1 Primary / Model 2 Secondary), or "none — default (yolo11n.pt)".
  • Archive — the video root, monospace.
  • Batches — count.
  • Dataset — X train / Y val, or "empty".

Then the class chips: swatch + name + annotation count + an × delete button (hidden when only one class remains). Only the first 8 render; the rest sit behind a +N more toggle, because a base model can carry 80 classes. A + add class chip opens an inline two-field form (name + optional prompt).

Card actions: Open (→ Video Archive), Model 1 upload (.pt), Model 2 upload (.pt), and a danger delete.

Confirmations.

  • Delete class: lists shapes to be deleted, label files to be rewritten, and warns that classes above it are renumbered. This renumbering is the whole point — see §6.
  • Delete project: "Delete X and everything under it?"

Endpoints. GET /projects, POST /projects, DELETE /projects/{id}, POST /projects/{id}/classes, DELETE /projects/{id}/classes/{classId}, POST /projects/{id}/base-model, POST /projects/{id}/secondary-model.


4.2 Video Archive (Library) — #/projects/{id}

Purpose. Browse the read-only recording archive, grouped into cycles, and pick a video to trim. Also hosts the batch table for this project.

The cycle concept — load-bearing domain logic. A production cycle runs 06:00 to 05:59 the next morning, so it always straddles midnight and covers two calendar dates. It is named after the date it starts (Siklus 3 Mar 2026). Grouping comes from the timestamp burned into the video image, not from the folder name — a recording made at 00:07 belongs to the cycle that began the previous morning. Files on disk are never moved or renamed; when a file's folder disagrees with its cycle, the folder date is shown beside it in amber.

Loads on mount (parallel): GET /projects/{id}, GET /projects/{id}/archive/cycles, GET /projects/{id}/batches, GET /jobs?project_id={id}. Then, whenever the selected cycle changes: GET /projects/{id}/archive/cycles/{cycle}.

Layout.

┌ page head: "Video Archive" + archive path + cycle explainer ──── [Cek truk (v4)] ┐
├ ActiveJobsBanner (only when jobs are running) ───────────────────────────────────┤
├ cycle list (left rail) │ video table (right, with filter box) ───────────────────┤
├ Batch table (BatchList) ─────────────────────────────────────────────────────────┤

Cycle rail. One button per cycle: folder icon, Siklus D Mon YYYY, an amber dot if any recording's clock is unverified, and the video count on the right. aria-current marks the selection.

Video table columns: Batch (running number within the cycle), Direkam (recorded time hh:mm:ss, amber when the clock is untrusted), File, Duration, Resolution, FPS, Size, Truk, Status, and a Trim action.

  • Truk shows hits/samples in green if a v4 truck-scan found trucks, "tanpa truk" in red if it found none, "belum dicek" if never scanned.
  • Status shows N batches in blue if this video has already been trimmed, else "Unused".
  • Trim is disabled when ffprobe could not read the duration.
  • A text box filters rows by batch label, client-side.

Cek truk (v4) posts POST /projects/{id}/archive/truck-scan — samples 12 frames per recording with the latest model to check a truck is actually present.

ActiveJobsBanner (shared with Batches): one row per running job — status dot, capitalised type, status, progress/total, Cancel, a progress bar, and the last log line.

BatchList (shared with Batches; see §4.4).

Endpoints. GET /projects/{id}, /archive/cycles, /archive/cycles/{cycle}, /batches, /jobs, POST /archive/truck-scan, POST /jobs/{id}/cancel.


4.3 Trim — #/projects/{id}/trim/{rel}

Purpose. Choose an in/out range and a sampling rate, then extract frames into a new batch.

Loads on mount. GET /projects/{id}/video/info?rel=… → { date_label, batch_label, duration, width, height, fps }. End defaults to min(duration, 60) seconds, start to 0, fps to 1.

Layout. Two columns: a <video> player (source GET /projects/{id}/video?rel=…, HTTP Range streaming, native controls) and a control stack.

Controls.

  • Start — range slider 0…duration step 0.1, a m:ss.d timecode text input, and a Use playhead button. Start is clamped to end - 0.1.
  • End — same, clamped to start + 0.1.
  • Moving either slider seeks the video to that position, so the boundary is visible.
  • Frames per second — number 0.1–30, step 0.1.
  • Live estimate: <range duration> of video → <N> frame(s) where N = round((end-start)*fps).
  • Extract frames — POST /projects/{id}/batches { rel, start_sec, end_sec, fps }, then finds its job in GET /jobs and polls it every second. Disabled while a job is in flight or when the estimate is 0; becomes "✓ Extracted" when done. A Cancel button appears while running.

Endpoints. GET /video/info, GET /video (Range), POST /projects/{id}/batches, GET /jobs, GET /jobs/{id}, POST /jobs/{id}/cancel.


4.4 Batches — #/projects/{id}/batches

Purpose. The work queue: everything extracted for this project, what has been annotated, what has been reviewed, and the entry point to auto-annotation and to the merge.

Loads on mount. GET /projects/{id}, GET /projects/{id}/batches, GET /jobs.

Batch shape:

{ id, project_id, date_label, batch_label, video_path,
  start_sec, end_sec, fps, status, frame_count, reviewed, annotation_count,
  review: { pending, approved, rejected } }
status = extracting | extracted | labeling | reviewing | approved | merged | failed

Page head actions. Restore from .zip (POST /projects/{id}/import, re-creates a batch from a previously downloaded annotation zip, reports frames/shapes/skipped), Download Annotations (.zip) (GET /projects/{id}/export, a backup that needs no merge), and Auto-Annotate All Batches (N) → the mass modal.

Batch table (component BatchList, also rendered on Video Archive):

Col Content
☑ select for merge; disabled when frame_count == 0; header checkbox selects all mergeable
Batch date_label · batch_label, click to rename (PATCH /batches/{id})
Range start–end @ fps
Frames frame_count
Reviewed reviewed / frame_count
Shapes annotation_count
Status active job type+status if any, else batch status
Actions Auto-annotate · Reset Auto · Review (n/N) · Delete

Above the table: X of Y selected can be merged and Prepare & Merge Selected (N) which navigates to #/projects/{id}/data-prep?batches=….

A batch is selectable if it has frames — not if it has been reviewed. Auto-annotation never approves anything (that is Review's job), so gating on approved frames made the button permanently dead. The checkbox tooltip says which case you are in: "N approved frame(s) would be merged" vs "Not reviewed — its N frame(s) would be approved as-is and merged".

Auto-annotate → engine chooser modal. Three cards, all appending to existing shapes:

  1. Project Base Model (only when base_model_path is set) — detect with this project's primary model.
  2. SAM3 Zero-Shot (Text Prompts) — detect by typing text.
  3. Custom YOLO Model (.pt) — opens a file picker, uploads to POST /batches/inspect-model, which stages the file and returns { staged_path, filename, classes[] }.

Whichever is picked opens the Auto-Annotate modal (§4.4.1) with that engine.

Reset Auto — POST /batches/{id}/reset-auto-annotations. Confirmation states that it clears auto shapes and resets frame review status to pending.

Endpoints. /projects/{id}, /projects/{id}/batches, /batches/{id} (PATCH, DELETE), /batches/{id}/reset-auto-annotations, /batches/inspect-model, /batches/{id}/autolabel, /projects/{id}/import, /projects/{id}/export, /jobs.

4.4.1 Auto-Annotate modal (single batch)

Two columns inside one dialog (920 px wide, max 96 vw / 90 vh, scrollable).

Left — live preview.

  • Frame image from GET /frames/{id}/image, wrapped in a container that hugs the image exactly. Nothing else may sit inside that wrapper — both overlays are sized to it, so an extra element shifts every box off the pixels it describes.
  • Overlay 1: PreviewShapes — detection boxes (dashed), polygon fills at 35 % opacity, and a className score% label per shape.
  • Overlay 2 (SAM3 only): ExemplarCanvas — drag to draw an example, Shift-drag for a counter-example. Positives are green solid, negatives red dashed, each numbered +1, −2… Escape abandons the in-progress drag. Drags shorter than 0.005 of the frame on either side are discarded as clicks.
  • Run Preview, and for SAM3: Undo box, Clear boxes, plus the hint Drag = example of <class> · Shift-drag = not this.
  • A frame slider across the whole batch; opens at the middle frame. Changing frames clears the preview and the exemplars.
  • An "Inferring…" chip while a request is in flight; a red banner for preview errors — a 409 here means a batch job holds the GPU lock, and that is the one failure the user can act on.

Right — parameters.

  • Confidence threshold — 0.05…0.95 step 0.05, default 0.35.
  • NMS IoU threshold — 0…0.9 step 0.05, default 0.0.
  • Min box size (fraction of frame) — 0…0.5 step 0.005, default 0, shown as a percentage.
  • ClassPromptPanel:
    • Target-class chips. Without SAM3 a chip is a simple toggle.
    • With SAM3 a chip is two buttons: the name selects-and-activates, the × deselects. Exactly one selected chip is active; it owns the prompt field and any exemplars drawn.
    • Prompt editor for the active class, saved with PATCH /projects/{id} { prompts: { [classId]: text } }. It writes to project_classes.prompt — the same field the Projects page edits — so what is tuned here is exactly what a batch run sends. Enter saves. Status line reads "Saved on the class" / "Unsaved — the batch run still uses the stored prompt until you save."
  • Cancel / Start Auto-Annotation → POST /batches/{id}/autolabel.

Preview auto-rerun. Adding, undoing or clearing an exemplar bumps a revision counter, debounced 250 ms, which re-runs the preview. Drawing a box is the question and the redrawn preview is the answer, so no button press sits between them. Re-running is cheap because set_image is already cached for this frame — only the grounding head runs.

Preview request — POST /batches/{id}/preview:

{ "frame_id": 12, "engine": "sam3|base_model|custom",
  "threshold": 0.35, "iou_threshold": 0.0, "min_box_frac": 0.0,
  "target_class_names": ["sack"],
  "custom_model_path": null,
  "exemplars": [ { "box": [cx, cy, w, h], "positive": true } ],
  "exemplar_class_name": "sack" }
→ { "shapes": [ { "class_id", "geometry", "score" } ] }

Nothing is written by a preview. Exemplars are never carried into the batch job — they tune the prompt, they are not labels.

Autolabel request — POST /batches/{id}/autolabel:

{ "resume": false, "append": true, "engine": "sam3",
  "threshold": 0.35, "iou_threshold": 0.0, "min_box_frac": 0.0,
  "target_class_names": ["sack", "half-sack"],
  "custom_model_path": null }

4.4.2 Mass Auto-Annotate modal

Same anatomy, wider (1040 px), with three differences:

  • An engine switch at the top (SAM3 / Base Model / Custom YOLO) rather than a pre-chosen engine; picking Custom opens the file picker and shows the staged model's name and classes.
  • A batch checklist — all batches ticked by default — plus a preview-batch selector.
  • An Append toggle.
  • Start submits one job per batch, sequentially in a loop, not Promise.all. One rejection must not abort the rest; the backend GPU lock serialises the real work anyway. Progress is reported as done/total (failed) and the result message is Queued N auto-annotation job(s).

4.5 Review — #/projects/{id}/review?batch={id} or #/batches/{id}

The core screen. Fully specified in §5.


4.6 Data Preparation — #/projects/{id}/data-prep[?batches=1,2,3]

Purpose. Two jobs and nothing else: throw out obviously-junk boxes, and decide how hard to augment what remains. It is also the merge gate — the only place a dataset gets created.

Two modes, from the query string:

Mode Trigger Behaviour
Gating ?batches=… present scope = those batches; shows Confirm merge
Browsing no query scope = one batch, chosen from a dropdown; shows Pick batches to merge

Loads on mount. GET /projects/{id}, GET /projects/{id}/batches (filtered to annotation_count > 0), GET /projects/{id}/triage/rules, GET /projects/{id}/augment, GET /projects (for the project switcher). Then per scope: GET /batches/{ids}/triage/summary, and on every filter change a debounced (300 ms) POST /batches/{ids}/triage/simulate.

Layout, top to bottom.

① Outlier filter. Three keep-ranges, one per signal, each a card with an enable checkbox and two sliders (keep from / up to):

Signal Range Meaning
score Confidence 0–1 step 0.01 how sure the detector was; low is usually junk
area_pct Box area 0–100 % step 0.1 catches specks and full-frame boxes
aspect Aspect ratio 0–10 step 0.1 width ÷ height; catches slivers

A disabled card dims to 0.4 opacity and stops accepting pointer events.

A keep-range is stored as two ignore rules, one per tail — the backend resolver already understood that shape, so the sliders are a friendlier face on existing machinery. toRules / fromRules are the entire translation:

range.min > signal.min  →  { name, predicate: { [key]: [null, range.min] }, action: "ignore" }
range.max < signal.max  →  { name, predicate: { [key]: [range.max, null] }, action: "ignore" }

A summary strip below shows kept, dropped (n · %), N frames held back of M — lost every box, an "unsaved" tag, Reset and Save filter (PUT /projects/{id}/triage/rules).

② Augmentation. Four preset cards — Off / Light / Medium / Aggressive — where Medium reproduces Ultralytics' own defaults exactly, so opening the page and saving nothing changes nothing. A Fine-tune individual settings disclosure reveals nine sliders: fliplr, flipud, degrees, translate, scale, hsv_h, hsv_s, hsv_v, mosaic, each with a plain-language label and a one-line hint. Any deviation from a preset shows a custom tag. Saved with PUT /projects/{id}/augment. Header note: applied to training images only — validation is never augmented, so mAP stays comparable.

③ Check the boundary. The scope indicator (list of selected batches, or the batch dropdown), a N decided by hand tag, then:

  • Scatter plot — SAM3 score (y) against box area (x, log scale, because box sizes span three orders of magnitude and a linear axis piles everything into the left edge). One dot per sampled shape. Drag a rectangle to select, shift-drag to add, click to clear. Red dashed shading marks what the current slider positions would drop — not what was last saved, because a plot that sits still while the thresholds move is useless. A gold ring marks a shape decided by hand. Caption states how many of the total shapes are plotted; the counts above always cover all of them.
  • Verdict bar — N selected — decide by hand (outranks the filter): Keep / Ignore / Clear hand decisions. Posts POST /triage/overrides / DELETE /triage/overrides.
  • Crop grid — a wall of cropped shapes, GET /annotations/{id}/crop, 120 per page, sorted server-side by lowest score or smallest area (a real batch holds ~85 000 shapes, so "worst 120" must be chosen from the whole batch, not from a page). Border colour = verdict; dropped crops go 35 % opacity and grayscale; each carries score and area in its footer. Click selects, shift-click adds. Load N more appends.

Verdict precedence, everywhere: manual > filter > keep.

④ Merge gate. In gating mode: Confirm merge (N boxes, M frames), disabled while the filter is dirty (the merge uses the saved rules). It opens MergeTargetModal:

  • radio list of existing datasets with image counts, plus Create a new dataset with an optional name (blank = today's date);
  • if any selected batch has zero approved frames, an amber warning naming the batch count and frame count, and a required "Merge them unreviewed" checkbox before Merge enables;
  • on confirm: POST /batches/{id}/approve-all for each unreviewed batch, then POST /batches/{ids}/approve { dataset_id | dataset_name }, which queues one merge job for the whole selection.

Explanatory copy that must survive: the dataset is cut now, from these rules; a dropped box leaves its image in the dataset and only a frame that loses every box is held back; the rules are frozen onto the dataset so editing them later never rewrites it; augmentation is read fresh at the start of every training run.

Endpoints. /triage/rules (GET, PUT), /triage/summary, /triage/shapes, /triage/simulate, /triage/overrides (POST, DELETE), /annotations/{id}/crop, /augment (GET, PUT), /batches/{id}/approve-all, /batches/{ids}/approve, /datasets.


4.7 Datasets — #/projects/{id}/datasets

Purpose. Named datasets: what a merge writes into, and what a training run picks from.

Loads. GET /projects/{id}/datasets; on selection change, POST /projects/{id}/datasets/combine-preview { dataset_ids }.

Dataset shape: { id, name, note, total, splits: { train, val }, rule_version, batches: [{ id, batch_label, images }] }.

Layout. Page head with totals and a Create empty dataset form (optional name) → an explanatory paragraph → a combine-preview strip when ≥1 card is ticked → a responsive card grid.

Card. Checkbox + name + N images · X train / Y val; optional note; the batches inside it, by label, with per-batch image counts — "Dataset #3" tells you nothing six weeks later, "batch9 + batch12, 908 images" is what you actually choose between. Then triage rules <version> and the actions: Rename, Resync, Download (zip), Delete.

Combine preview strip. N selected · T unique images · X train / Y val, plus, when datasets overlap, "K frame(s) appear in more than one — counted once, taking the labels from the newest dataset."

Resync is the only operation that changes an already-merged dataset. Its confirmation must say so, and must warn that the rule version is re-stamped — a model trained on it before this point was measured on different labels. The result alert reports labels rewritten, the new rule version, and how many frames would have lost every box and were left alone.

Delete confirmation: the frames and annotations stay, only this dataset's copy of them goes.

Explanatory paragraph that must survive: a dataset is what an approve/merge writes into; triage rules are applied at that moment and then frozen; editing rules in Data Prep does not reach back; a frame's train/val split is decided once per project and every dataset inherits it — otherwise base-vs-new mAP would be measured on images the new model had already trained on.

Endpoints. /projects/{id}/datasets (GET, POST), /datasets/{id} (PATCH, DELETE), /datasets/{id}/resync, /datasets/{id}/download, /datasets/combine-preview.


4.8 Models & Training — #/projects/{id}/models

Purpose. Configure and launch a fine-tune, watch it, and compare each version against the base model on the same val set.

Loads on mount (parallel): GET /projects/{id}, /projects/{id}/dataset, /projects/{id}/models, /hardware, /jobs?project_id=, then /projects/{id}/datasets and /projects/{id}/base-datasets. All datasets start ticked; base datasets start unticked — a base dataset is somebody else's labels and must be opted into.

Layout. Two columns: configuration (340–420 px) and results.

Left column.

  1. Base Model Configuration — status (Custom model.pt loaded / Default yolo11n.pt), the class list, and an Upload Base Model (.pt) control.
  2. Train Model — target-class chips (toggleable, all on by default), an Epochs number input (default 50, 1–500), a hardware line (<GPU or CPU> — default batch B, imgsz S) from GET /hardware, and Start Training. Disabled while running, or with no dataset selected, or with no class selected; inline red hints explain which.
  3. Select Datasets (n/N) — checklist with a Select/Deselect All toggle and per-dataset image counts, plus a live "this run sees T unique images (X train / Y val)" line and an overlap note. Below it, when any exist, Base datasets — externally labelled · train only — with image and box counts and the note that they never join the val split, so the base-vs-new mAP stays measured on this project's own frames.

Training request — POST /projects/{id}/train:

{ "epochs": 50, "dataset_ids": [1,2], "base_dataset_ids": [], "class_ids": [0,1] }

No batch filter: the chosen datasets already carry their batches, and filtering again could only subtract from them.

Right column.

  • Active job card — Training Job #id (status), progress/total, a bar, the error if any, the last 8 log lines in a monospace scroll box, a Cancel button while running, and a green "Training Finished" line when done (which also reloads the page data).

  • Trained Model Versions (N) — one card per version: the model name (auto-generated {arch}-{labelType}-{epochs}ep-{classNames}-{YYYYMMDD}, clickable to rename inline), the created timestamp, and a metrics table:

    Metric Base This version Δ
    mAP50, mAP50-95, precision, recall 4 dp 4 dp signed, green if >0, red if <0

    When there is no base column, an explicit line says the previous model could not be scored on this val set. Actions: Download best.pt ({name}-best.pt), Use as base model (POST /models/{id}/promote).

Endpoints. /projects/{id} , /dataset, /datasets, /base-datasets, /models, /hardware, /train, /models/{id}/promote, /models/{id}/rename, /models/{id}/weights, /jobs.


4.9 Live Counting — #/projects/{id}/live-count

Purpose. Point a trained model at an archive video or an RTSP camera and watch it count, using the same tracker, stabiliser and line-cross counter the production script uses. It answers "does the model count correctly", not "how many sacks today".

Loads. GET /projects/{id}/live-count/models, GET /projects/{id}/library (dates), then GET /projects/{id}/library/{date} (videos). GET /live-count/status polls every second, always.

Layout. Left control panel (320–380 px), right stats + video + events.

Source. Two chips: Archive video (default) and RTSP stream. A file decodes at ~100 fps, an RTSP camera at ~6 (OpenCV decodes 1080p on the CPU), so testing counting accuracy against a file is far quicker — the copy says this, and the default must stay "file". Archive mode gives date and video selects; stream mode a URL field.

Model. Select from live-count/models ({ path, label }).

Settings — sliders, in this order, each with a one-line hint:

Key Range Live? Note
line_y 0–720 ✅ the counting line
line_x_start 0–1280 ✅ ignore anything left of this
line_x_end 0–1280 ✅ ignore anything right of this
margin 0–120 px — dead band so jitter alone never counts
entry_travel_min 0–200 px — kills ghost boxes that blink in next to the line
handoff_radius 0–300 px — most sensitive dial: inheritance on track death; too large and unrelated sacks adopt each other
unload_confirm_frames 1–15 — stops a repositioned sack cancelling a real count
min_area_scale 0–2 — perspective-aware size gate; 0 disables
conf 0.05–0.95 — detector threshold

Live vs locked is not cosmetic. The three line fields can be changed mid-session and are pushed with PATCH /live-count/line (fire-and-forget — the next status poll confirms, and a request dropped during a fast drag is corrected by the one after). Every other field is disabled while running and needs a restart. Live fields show a green live marker.

Stat tiles (from the status poll): Counted in, Counted out, Net, FPS, Tracked now, Ignored, Too small, Tracks traced.

Video area. While running, an MJPEG <img> from GET /live-count/stream?k={key} — the key busts the browser cache so a restarted session gets a fresh connection. Clicking the video places whichever edge is armed (three chips: Counting line / Left edge / Right edge); the click is converted against a fixed 1280×720 coordinate space. Placing by eye beats guessing a pixel value on a slider. Not running → "Not running. Set a source and press Start."

Trace note. While running, if status.trace_path is set, a line points at it: every finished track and the reason it did or did not count is written there — that file is what separates a model miss from a tracker miss from a counter miss.

Recent counts. Chips of #trackId · direction · at s, newest first.

Footer note that must survive: the session holds the GPU, so training and auto-annotation wait their turn.

Endpoints. /live-count/models, /live-count/start, /live-count/stop, PATCH /live-count/line, /live-count/status, /live-count/stream.


4.10 Counting Accuracy — #/projects/{id}/counting-bench

Purpose. One row per archive video: what the counter predicted, what you actually counted, and the signed difference. The recount runs headless in a background job — no annotated frame, no JPEG encode — because the point is the number, not watching it happen.

Delta is deliberately signed. Counting 103 where the truth is 100 is a different failure from counting 97, and an absolute accuracy percentage hides which one you have.

Loads. GET /projects/{id}/counting-bench → { rows, totals, active_job, scan_job, unindexed }, plus GET /projects/{id}/live-count/models. Polls every 2 s only while a count or clock-scan job runs; the table is otherwise static and 222 rows/second is wasted work.

Totals row. Counted n/N, Scored, Predicted in, Ground truth, Delta (coloured: 0 green,

  • amber, − red), Accuracy % over scored videos only.

Toolbar. Model select · cycle filter · Recount videos that already have a result checkbox · Count selected (n) · Read timestamps (n left) / Re-read timestamps · Count all (N) (tooltip: this takes hours).

Clock scanning. When recordings have no verified start time, an amber notice explains that they are still grouped by folder name, that folder names are not when a recording happened, and points at Read timestamps (POST /counting-bench/scan-clock), which reads the clock burned into each video and regroups by the 06:00-to-06:00 working day.

Table, grouped by cycle. Each cycle is a clickable header row with: a select-all checkbox for the cycle, ▼/▶, the cycle label, the video count, N perlu dicek when start times are unverified, and the cycle's AI / GT / Δ / accuracy%. Only the newest cycle is expanded by default — eleven expanded cycles is the wall of rows the grouping exists to avoid. Cycles sort newest first; rows inside a cycle sort by batch_no, i.e. forwards in time.

Row columns: ☑ · # (batch no) · Recorded (editable hh:mm:ss text; amber border and text when the clock is untrusted; blur or Enter saves via PATCH /counting-bench/clock, combining the working day with the typed time) · File (folder_date/batch_label, plus a red "failed" tag carrying the error) · Length · Counted in · Counted out · Net · Ground truth (number input, blur/Enter saves via PATCH /counting-bench/ground-truth, must be a whole non-negative number, empty clears) · Delta (coloured).

Timezone rule — do not "fix" it. The stored start time is wall-clock text with no timezone and is displayed as-is by slicing the string. Parsing it into a Date re-applies the browser's offset and shifts every recording by hours; that is a bug this code already had once.

Endpoints. /counting-bench (GET), /counting-bench/run, /counting-bench/ground-truth, /counting-bench/scan-clock, /counting-bench/clock, /live-count/models, /jobs/{id}/cancel.


4.11 SAM3 Playground — #/sam3-playground

Purpose. Try SAM3 zero-shot text prompts on any image, with no project and no database writes. A scratchpad for finding the phrasing that works before committing it to a class prompt.

Layout. Left control panel (360 px), right canvas viewer.

Controls. A click-or-drag-and-drop image dropzone (.jpg, .png, shows filename and KB) · a comma-separated prompts textarea (default white plastic bag, forklift, pallet) with the tip use specific descriptive phrases ("woven white bag" beats "sack") · Confidence 0.05–0.95 · NMS IoU 0.1–0.9 · Run SAM3 Test. Results list underneath: one row per detection, coloured left border, prompt text, score %.

Viewer. The uploaded image with an SVG overlay at the response's width/height viewBox: polygon masks at 35 % fill, a dashed bounding box, and a solid label badge reading prompt (NN%). Empty state: "Upload an image on the left to start testing SAM3 text prompts."

Endpoint. POST /sam3/playground-test (multipart: file, prompts, threshold, iou_threshold) → { width, height, detections: [{ prompt, class_id, score, box, polygons }] }.


5. Review editor — deep specification

This is where the operator spends hours. Everything here is tuned for that.

5.1 Purpose and entry

Sign frames off one by one: fix the shapes auto-annotation produced, approve or reject the frame, move on. Review does not merge — it hands off to Data Prep, which is the gate that turns batches into a dataset.

Entry: #/batches/{id} or #/projects/{id}/review?batch={id}. With no batch id it loads the project's first batch.

On first load it jumps to the first frame that has any annotation. In a 600-frame batch where SAM3 found things in 40, opening on frame 1 shows an empty frame and reads as "nothing worked". The jump happens once per load, and again after an auto-label job finishes.

5.2 Data model in play

frame       { id, idx, filename, width, height, review_status, annotation_count }
annotation  { id, class_id, geometry, score, source: "auto"|"manual" }
batch       { …, review: { pending, approved, rejected }, annotation_count, status }

source matters: a batch re-run deletes only source="auto" rows, so manual corrections survive.

5.3 Layout

┌ head: date · batch  |  N frames · R/N reviewed · S shapes · status ┐
│                                    [Approve All Frames (p)] [Prepare & merge] │
├ active-job banner (while auto-labeling) ──────────────────────────────────────┤
├──────────────────────────── review grid ──────────────────────────────────────┤
│ main                                             │ sidebar                    │
│  canvas (image + SVG overlay)                    │  [exemplar filter panel]   │
│  frame-bar: mode switch + context hint           │  Classes                   │
│  [selection bar — select mode only]              │  Shapes on this frame (N)  │
│  frame-bar: ← n/N status → · N · C · T · Reject · Approve │  Shortcuts        │
│  [quick reclass bar — when a shape is selected]  │                            │
│  filmstrip                                       │                            │
└───────────────────────────────────────────────────────────────────────────────┘

The canvas is bounded by max-width: calc(72vh * width / height) — bounding the height by bounding the width at the frame's aspect ratio. A portrait frame would otherwise be three screens tall, and constraining the <img> itself would leave the SVG overlay misaligned.

Head actions. Approve All Frames (p) appears while pending frames remain and the batch is not merged (POST /batches/{id}/approve-all, confirmed). Prepare & merge is disabled while any frame is pending; it navigates to Data Prep with this batch selected, or reads "Merged".

5.4 The canvas

An SVG overlay over the frame image, not a <canvas>. Shapes are DOM elements, so selection, hover and focus come from the DOM instead of hand-written hit testing, and the whole thing stays keyboard-reachable. Rebuilds should keep SVG for the same reason.

viewBox = "0 0 frameWidth frameHeight", preserveAspectRatio="none". Only the handle size is converted back to frame units (12 px × scale), so handles stay the same physical size at any display size.

A shape renders as: a rect (bbox) or polygon, stroked in the class colour, #38bdf8 when marked. When selected and editable it also gets a label badge reading [classIndex+1] className and its handles:

  • bbox → four corner handles (nw/ne/se/sw), drag to resize;
  • polygon → a vertex circle per point (drag to move, Alt-click to delete, refused below 3 points) and a smaller half-opacity midpoint circle per edge (drag to insert a vertex there).

Gestures on the canvas.

Gesture Draw mode Select mode
Drag on empty canvas creates an exemplar (§5.6) — or a SAM3-assisted shape while S is held marquee: selects every shape it touches
Shift + drag on empty canvas negative exemplar adds to the selection instead of replacing
Drag on a shape moves it toggles its marked state — never moves it
Click on empty canvas clears the selected shape clears the marked set
Escape cancels the gesture in flight (never commits it) clears the marked set

Drags shorter than 0.004 normalized on both axes are treated as clicks.

Select mode never edits geometry. A stray drag on top of a box would silently rewrite the shape the user was only trying to tick off, so handles are not rendered there at all.

Move is clamped to the frame on both axes, for boxes and polygons alike.

Live edit vs commit. Dragging writes geometry locally on every pointer move (no request); pointer-up sends one PATCH /annotations/{id}. A failed patch rolls the shape back to its previous geometry and shows the error. Same optimistic-with-rollback pattern for delete.

5.5 Modes and bars

Mode switch (Draw / Select, V toggles). Switching clears the marked set.

  • Draw hint: Drag an example of this class · Shift-drag = not this, or, with a pool active, N example(s), M negative — tune the filters, then Apply.
  • Select hint: Select all (N) button plus N selected · Shift-drag adds · Esc clears.

Selection bar (select mode, ≥1 marked): the count, a reclass button per class labelled [n] name, Clear [Esc], and Delete N [Del] (confirmed). Bulk operations post POST /annotations/bulk-reclass and POST /annotations/bulk-delete.

Frame bar. ← · n / N + a status pill (pending/approved/rejected) · → · Next Shape [N] · Copy Prev [C] · Track 5 Frames [T] · an "asking SAM3…" indicator · Reject [X] · Approve [A].

  • Next Shape jumps to the next frame with annotations, wrapping to the first.
  • Copy Prev copies every annotation from frame n-1 onto this frame (one POST per shape); disabled at index 0 or when the previous frame is empty.
  • Track 5 Frames takes the selected shape's bounding box and runs SAM3 assist on the next 5 frames with it. Only implemented for polygon geometry in the current code — a bbox selection yields no box and the action no-ops. A rebuild should either keep that limitation or extend it deliberately, not accidentally.

Quick reclass bar appears whenever a shape is selected: one button per class ([n] name, class-coloured) plus Delete [Del].

Filmstrip. Every frame as a thumbnail (GET /frames/{id}/image?w=120, lazy-loaded), border coloured by review status, an annotation-count badge, aria-current on the active one, which is scrolled into view (block: nearest, inline: center) on every index change.

5.6 The exemplar flow — the most intricate part of the app

In draw mode, a drag is not a rectangle. It is an example. The user draws one instance of the class; SAM3 finds the rest; the result is previewed, tuned and only then written.

Pool semantics.

  • The pool is a list of { box: [x0,y0,x1,y1], positive: bool }. positive = !shiftKey.
  • It is held both in state and in a ref; the ref is what gets sent, so a drag that lands while a request is in flight is never lost.
  • It is cleared by a frame change or a class change — SAM3's geometric prompts pool features from this image and mean nothing on another, or for another class.
  • Every run posts the whole pool, not the newest box. SAM3 can only append geometric prompts, so re-running from empty is also how undo works.

Run scheduling.

  • 400 ms debounce after the last drag; 250 ms after the last slider move.
  • One pass at a time. A change arriving mid-pass sets rerunWanted so exactly one rerun follows — never a queue.
  • Every run is a dry run by default. Only Apply commits.

The filter panel is mounted only while a run is undecided: the first drag opens it, Apply and Discard close it. The pool outlives the panel. It lives at the top of the sidebar, not over the frame — the shapes it governs are drawn on the canvas, and a card on top of them covers the thing being judged. It scrolls itself into view on mount, because the sidebar scrolls and a decision the user cannot see is not a decision.

Its four sliders, re-previewing on a 250 ms debounce:

Key Label Range Default Hint
threshold Confidence 0.05–0.95 0.5 lower finds more, and more junk
iou_threshold Overlap (NMS) 0.1–1 0.8 boxes overlapping this much are one object; lower deletes more
min_box_frac Min box size 0–0.05 0.002 shown as % of frame; higher deletes more
max_detections Max shapes 5–300 100 keep only the highest-scoring N

Panel header: Auto-label "<class>" + Undo (drops the last example drawn). Status line: asking SAM3… / F found + D drawn · R rejected / Drag an example on the frame, plus "Apply replaces the N '' shapes on this frame — nothing is saved yet." Buttons: Reset, Discard (Esc), Apply (Enter, disabled while busy or with no preview).

While a preview is up, the canvas hides the stored shapes of the class under review. The run replaces them wholesale; leaving them on screen made a rejected detection look like it had never gone. Other classes stay, dimmed. Proposals draw dashed: green for the boxes the user drew, class-coloured for what SAM3 found.

Apply re-runs the request with apply: true rather than posting the previewed geometry back. SAM3 is deterministic for a given pool and threshold, and the browser should not be the authority on what gets stored.

After a successful apply: the negatives that were sent are dropped (spent — the frame no longer carries what they rejected) and the positives stay as prompts. Drags that landed while the apply was in flight are not part of that run, keep their place at the end of the pool, and trigger exactly one rerun.

Discard truncates the pool back to appliedRef — the length it had at the last successful apply — so a rejected run leaves neither shapes nor prompts behind.

Request / response.

POST /api/frames/{id}/exemplar-label
{ "exemplars": [{ "box": [x0,y0,x1,y1], "positive": true }],
  "class_id": 0,
  "threshold": 0.5, "iou_threshold": 0.8,
  "min_box_frac": 0.002, "max_detections": 100,
  "apply": false }
→ { "shapes": [{ "geometry", "score", "source": "manual"|"auto" }],
    "applied": false, "redetected": true, "message": null,
    "annotations": null }        // the whole frame, on apply only

Server-side, positives become source="manual" and survive a batch re-run; what SAM3 adds becomes source="auto" and is replaceable. Detections overlapping a negative by ≥0.3 are dropped; ones overlapping a positive by ≥0.6 are dropped as already covered. The panel's filters run first, in the batch job's order: area floor → NMS → cap.

GPU-lock timeout is 5 s here (versus 20 s for assist) because this fires from a mouse gesture. On timeout the response carries redetected: false, a message the panel shows, and the drawn boxes alone as the preview. Applying that run appends the drawn boxes and honours the negatives instead of taking the replace path — with no detections to put back, replacing would wipe the class and leave only the drawings.

5.7 SAM3 click-assist

Holding S switches the drag gesture from "exemplar" to "assist": the drawn box is sent to POST /frames/{id}/assist { box, class_id } and SAM3 returns one shape, which is added and selected. The canvas gets an assist class while S is held so the cursor and draft stroke change. Key-up releases it.

5.8 Keyboard map

Handlers are attached in the capture phase on document, and:

  • ignore events whose target is an input, textarea, select or contenteditable;
  • ignore anything with Ctrl / Cmd / Alt — those belong to the browser and the OS. Without this, Ctrl+S approves the frame and Ctrl+A/C/X/N/T all fire review actions;
  • preventDefault + stopPropagation only for keys that are actually shortcuts.
Key Draw mode Select mode
V → Select mode → Draw mode (clears marks)
← → previous / next frame same
A approve frame, then advance same
X reject frame, then advance same
U jump to next pending frame (GET /batches/{id}/next-pending?after_idx=) same
N jump to next annotated frame same
C copy annotations from previous frame same
T track selected shape forward 5 frames same
1–9 set active class; reclass the selected shape if any reclass all marked shapes
Del / Backspace delete the selected shape delete all marked shapes (confirmed)
Esc cancel the gesture in flight clear the marked set
hold S SAM3 assist mode while held —
Enter apply the exemplar preview (panel open) —
Esc discard the exemplar preview (panel open) —

In select mode the marquee owns Delete and the digits. Otherwise a 40-box selection is thrown away by one keystroke meant for it.

Approve and reject advance to the next frame, and patch the frame's status locally before the request returns — a review session is hundreds of A presses and must not wait on the network.

5.9 Sidebar

  • Exemplar filter panel — only while a run is undecided (§5.6).
  • Classes — one row per class: a chip (swatch + name + its 1-based hotkey) that sets the active class and reclasses the selected shape, plus a trash button that clears every shape of that class across the whole batch (DELETE /batches/{id}/classes/{classId}/annotations, confirmed).
  • Shapes on this frame (N) — one row per annotation: swatch, class name, and either the score to 2 dp (auto) or the word manual; click selects, trash deletes. When empty: "None on this frame — drag to draw a shape." plus, if the batch has shapes elsewhere, a Jump to Frame with Shapes [N] button.
  • Shortcuts — a static <dl> of the keyboard map. Keep it visible; this is a keyboard-first screen and the panel is the discovery mechanism.

6. Simplification rules

The UI may be redesigned. The following table separates what is free, what needs care, and what must not move.

6.1 Free to change

  • Visual style: spacing, radii, typography, the exact palette (as long as class colours stay consistent across screens and contrast stays ≥ 4.5:1).
  • Component library choice, CSS approach, icon set (SVG only — and the leftover emoji in buttons should be replaced).
  • Inline styles → design tokens/classes. Most of this codebase styles inline; that is technical debt, not intent.
  • Table vs card layouts, where a section sits on the page, panel collapse/expand behaviour.
  • Adding light-mode polish, responsive breakpoints, focus rings, reduced-motion handling.
  • Merging the two near-identical auto-annotate modals into one component with a batches: Batch[] prop — they already share PreviewShapes.
  • Replacing hash routing with a real router, as long as the URL shapes in §2.1 keep working (they are pasted between screens and bookmarked).

6.2 Change only with a decision recorded

These are real product choices, not accidents. Changing them is allowed but must be deliberate and written down in ./requirements.md.

Thing Why it is the way it is
Review opens on the first annotated frame opening on an empty frame 1 reads as "nothing worked"
Only the newest cycle expands on Counting Accuracy eleven expanded cycles is the wall of rows the grouping exists to avoid
Base datasets start unticked on Models they are somebody else's labels; they must be opted into
All project datasets start ticked on Models the common case is "train on everything I have"
Live Count defaults to archive video, not RTSP ~100 fps vs ~6 fps; accuracy testing against a file is far quicker
Delta is signed, not an absolute accuracy % over-count and under-count are different failures
The scatter uses a log area axis box sizes span three orders of magnitude; a linear axis piles everything at the left edge
The crop grid pages from the server, sorted server-side a batch holds ~85 k shapes; "worst 120" must come from the whole batch
Only 8 class chips show on a project card a base model can carry 80 classes
The exemplar filter panel sits in the sidebar, not over the frame it governs shapes drawn on the canvas; a card on top covers the thing being judged

6.3 Must not be simplified away

Removing any of these breaks the pipeline, corrupts data, or produces numbers that quietly lie.

Data integrity

  1. Normalized 0–1 geometry, end to end. The moment a screen stores pixels, that shape is wrong at any other display size. Only the two documented exemplar formats deviate.
  2. source: "auto" | "manual" must be preserved and displayed. A batch re-run deletes only auto rows; if the UI stops distinguishing them, manual corrections become invisible and destructible.
  3. The stable val split. A frame that lands in val stays in val forever, across every dataset. The UI must never offer a "reshuffle split" control. Without this, base-vs-new mAP is meaningless — the new model would be scored on images it trained on.
  4. Triage rules are frozen onto a dataset at merge time. Data Prep's copy that says so, and Resync's warning that it re-stamps the rule version, are part of the feature, not filler.
  5. Class deletion renumbers. A YOLO label is an integer index; a class list and a set of label files that disagree do not fail loudly — they train a model on the wrong names. The confirmation must keep enumerating exactly what will be rewritten.
  6. The merge gate lives in Data Prep, not in Review or Batches. Review signs frames off; Batches selects; only Confirm merge creates a dataset. Collapsing these puts an unfiltered merge one click from an annotation screen.
  7. Merging unreviewed batches requires an explicit acknowledgement checkbox. Wrong boxes become training labels.
  8. The wall-clock start time is displayed by string slicing, never parsed into a Date. Parsing re-applies the browser's timezone and shifts every recording by hours.

Machine behaviour

  1. One set_image per image. The N-prompt job calls set_image once per frame and loops prompts over that cached backbone state. Any UI that implies per-prompt image loading (e.g. "run each prompt separately") is pushing the backend toward a restructure that must not happen.
  2. Exemplars are tuning aids, never labels, in the auto-annotate modal. They go to /preview only, and are dropped on frame or class change.
  3. Exemplar runs in Review are dry runs until Apply, and Apply re-runs server-side rather than posting previewed geometry. The browser is not the authority on what gets stored.
  4. The whole pool is re-sent on every run. SAM3 can only append geometric prompts; this is also how undo works.
  5. One exemplar pass at a time, with a single queued rerun. Not a request per drag, not a stack of pending passes.
  6. Debounces exist for GPU cost, not for looks: 400 ms after a drag, 250 ms after a slider, 300 ms on the triage simulate. Removing them fires a GPU round trip per pixel.
  7. The mass auto-annotate loop submits sequentially. Promise.all drops every remaining batch on the first rejection.
  8. Jobs are polled only while jobs are running, and data is reloaded once when the last one finishes. Constant polling on a static 222-row table is wasted work; no reload means new frames and model versions never appear.
  9. A 409 from a preview or an assist means the GPU lock is held. It must reach the screen as its own message — it is the one failure the user can act on (wait, or stop the job).

Interaction contracts

  1. Every Review action has a keyboard shortcut, and the Shortcuts panel stays visible. The mouse is for drawing shapes, not for navigating.
  2. Ctrl/Cmd/Alt combos are never intercepted. Ctrl+S must not approve a frame.
  3. Typing in an input never triggers a shortcut.
  4. Select mode does not move shapes; draw mode does not marquee. Mixing them silently rewrites geometry the user was only ticking off.
  5. Escape cancels an in-flight gesture without committing it, on every canvas.
  6. Approve/reject advance the frame and update locally before the request returns.
  7. Optimistic edits roll back on failure, with the backend's own message shown.
  8. The overlay wrapper hugs the image and contains nothing else. Any extra element inside it shifts every box off the pixels it describes. Related: overlay CSS must be scoped to the direct child SVG (> svg), or a nested icon SVG inherits position:absolute; width:100% and stretches across the whole frame.
  9. Progress is always progress/total + bar + cancel. Never a spinner with no end.

Navigation

  1. The nav order is the pipeline order, and every project-scoped page keeps a working project id in its links.
  2. ?batches= and ?batch= survive. They are how Batches hands a selection to Data Prep and how Review is deep-linked.

6.4 Known rough edges (fix, don't preserve)

  • Emoji used as icons in buttons (🔄, 🏷️, 📋, 🚀, 📦, 🤖, ⚡, 🖼️) — replace with SVG.
  • window.confirm / window.prompt / window.alert for renames, deletes and the Resync report — replace with real dialogs, keeping the same enumerated consequences.
  • Mixed English/Indonesian copy (Video Archive, Counting Accuracy). Pick one language per screen; the domain words siklus, batch, truk are the operator's vocabulary and can stay if the rest follows.
  • Heavy inline styling throughout — move to tokens.
  • AutoAnnotateModal and MassAutoAnnotateModal duplicate ~70 % of their logic.
  • BatchList and ActiveJobsBanner are exported from LibraryPage and imported by BatchesPage; they belong in components/.
  • Track-5-frames only works for polygon geometry.

7. API surface index

Everything the frontend calls, grouped as the pages use it. Errors are {"detail": "..."}.

Health & hardware

GET  /api/health                 → { gpu, vram_free_gb, sam3_ready }
GET  /api/hardware               → { gpu, batch, imgsz }

Projects & classes

GET    /api/projects                              → { projects[], video_root_default }
POST   /api/projects                              { name, label_type, video_root, val_every, classes[] }
GET    /api/projects/{id}
PATCH  /api/projects/{id}                         { prompts: { classId: text }, val_every, … }
DELETE /api/projects/{id}
POST   /api/projects/{id}/classes                 { name, prompt? }
DELETE /api/projects/{id}/classes/{classId}
POST   /api/projects/{id}/base-model              multipart file (.pt)
POST   /api/projects/{id}/secondary-model         multipart file (.pt)

Archive & video

GET  /api/projects/{id}/library                   → { dates[] }
GET  /api/projects/{id}/library/{date}            → { videos[] }
GET  /api/projects/{id}/archive/cycles            → { cycles[] }
GET  /api/projects/{id}/archive/cycles/{cycle}    → { videos[] }
POST /api/projects/{id}/archive/truck-scan
GET  /api/projects/{id}/video/info?rel=…
GET  /api/projects/{id}/video?rel=…               Range streaming

Batches, frames, annotations

POST   /api/projects/{id}/batches                 { rel, start_sec, end_sec, fps }
GET    /api/projects/{id}/batches
GET    /api/batches/{id}
PATCH  /api/batches/{id}                          { batch_label }
DELETE /api/batches/{id}
GET    /api/batches/{id}/frames
POST   /api/batches/{id}/preview                  see §4.4.1
POST   /api/batches/{id}/autolabel                see §4.4.1
POST   /api/batches/{id}/autolabel-with-model     multipart
POST   /api/batches/inspect-model                 multipart → { staged_path, filename, classes[] }
POST   /api/batches/{id}/reset-auto-annotations
DELETE /api/batches/{id}/classes/{classId}/annotations
GET    /api/batches/{id}/next-pending?after_idx=
POST   /api/batches/{id}/approve-all
POST   /api/batches/{ids}/approve                 { dataset_id | dataset_name }
GET    /api/frames/{id}/image?w=
GET    /api/frames/{id}/annotations
POST   /api/frames/{id}/annotations               { class_id, geometry }
POST   /api/frames/{id}/assist                    { box, class_id }
POST   /api/frames/{id}/exemplar-label            see §5.6
POST   /api/frames/{id}/status                    { status }
PATCH  /api/annotations/{id}                      { geometry | class_id }
DELETE /api/annotations/{id}
POST   /api/annotations/bulk-delete               { annotation_ids }
POST   /api/annotations/bulk-reclass              { annotation_ids, class_id }
GET    /api/annotations/{id}/crop

Triage & augmentation

GET  /api/projects/{id}/triage/rules
PUT  /api/projects/{id}/triage/rules              { rules }
GET  /api/batches/{ids}/triage/summary
GET  /api/batches/{ids}/triage/shapes?sort=&offset=&limit=
POST /api/batches/{ids}/triage/simulate           { rules }
POST /api/triage/overrides                        { annotation_ids, verdict, target_class }
DELETE /api/triage/overrides                      { annotation_ids }
GET  /api/projects/{id}/augment
PUT  /api/projects/{id}/augment                   { settings }

Datasets

GET    /api/projects/{id}/dataset                 master summary
GET    /api/projects/{id}/datasets
POST   /api/projects/{id}/datasets                { name }
POST   /api/projects/{id}/datasets/combine-preview{ dataset_ids }
PATCH  /api/datasets/{id}                         { name }
DELETE /api/datasets/{id}
POST   /api/datasets/{id}/resync
GET    /api/datasets/{id}/download
GET    /api/projects/{id}/base-datasets
DELETE /api/base-datasets/{id}
POST   /api/projects/{id}/import                  multipart zip
GET    /api/projects/{id}/export?batch_ids=&approved_only=&include_empty=

Training & models

POST /api/projects/{id}/train      { epochs, dataset_ids, base_dataset_ids, class_ids }
GET  /api/projects/{id}/models
POST /api/models/{id}/promote
PATCH /api/models/{id}/rename      { name }
GET  /api/models/{id}/weights

Counting

GET   /api/projects/{id}/live-count/models
POST  /api/projects/{id}/live-count/start   { source_rel | source, model_path, ...cfg }
POST  /api/live-count/stop
PATCH /api/live-count/line                  { line_y, line_x_start, line_x_end }
GET   /api/live-count/status
GET   /api/live-count/stream?k=             MJPEG
GET   /api/projects/{id}/counting-bench[?date=]
POST  /api/projects/{id}/counting-bench/run { model_path, video_rels, all_videos, recount }
PATCH /api/projects/{id}/counting-bench/ground-truth { video_rel, ground_truth }
POST  /api/projects/{id}/counting-bench/scan-clock   { rescan }
PATCH /api/projects/{id}/counting-bench/clock        { video_rel, started_at }

SAM3 playground

POST /api/sam3/playground-test   multipart: file, prompts, threshold, iou_threshold

Jobs

GET  /api/jobs?project_id=
GET  /api/jobs/{id}
POST /api/jobs/{id}/cancel

8. Redesign brief — prompt siap-pakai

Bagian ini adalah satu prompt yang berdiri sendiri, untuk diserahkan ke seorang desainer atau ke sebuah model, ketika UI reTraining mau dirancang ulang dari nol. Isinya adalah pemadatan §1–§7: sengaja mengulang, supaya bisa disalin utuh tanpa membawa dokumen ini. Kalau suatu saat berbeda dengan §1–§7, §1–§7 yang benar dan bagian ini yang bug.

Copy mulai dari sini.


Apa ini

reTraining adalah internal tool self-hosted, single-user, single-GPU, yang mengubah rekaman CCTV mentah menjadi dataset YOLO dan model deteksi yang terbukti lebih baik dari model sebelumnya. Bukan SaaS, bukan landing page, tidak ada login/multi-tenant/onboarding. Satu operator, satu layar, sesi kerja berjam-jam (review ribuan frame).

Stack sekarang: React 19 + Vite, hash-routing tanpa router library, FastAPI, SAM3 (auto-label zero-shot), Ultralytics YOLO11 (training), Docker + GPU.

Domain kasus pertama: menghitung karung (sack) yang dimuat/dibongkar dari truk di sebuah line pakan — itu sebabnya ada modul counting.

Alur pipeline (urutan nav = urutan kerja, tidak boleh diacak)

Projects → Video Archive → Trim → Batches → Auto-annotate → Review → Data Prep → Datasets → Models & Training → (Live Count / Counting Accuracy)
   ①            ②            ③        ④           ⑤             ⑥         ⑦           ⑧              ⑨
  1. Projects — 1 project = 1 model yang sedang diperbaiki: base model .pt, daftar class, label_type (bbox/polygon, terkunci setelah merge pertama), root arsip video, val-split.
  2. Video Archive — browse arsip video read-only yang dikelompokkan per siklus (06:00–05:59 hari berikutnya, jadi selalu melewati tengah malam; nama siklus = tanggal mulainya). Pengelompokan diambil dari timestamp yang terbakar di gambar video, bukan nama folder.
  3. Trim — pilih rentang in/out di player + fps sampling → job ekstraksi frame → jadi satu batch.
  4. Batches — antrean kerja. Per batch: jumlah frame, reviewed, shapes, status. Dari sini auto-annotate (SAM3 text-prompt / base model / custom YOLO) dan pilih batch untuk merge.
  5. Review — layar inti. Perbaiki shape hasil auto-label, approve/reject tiap frame. Review tidak merge.
  6. Data Prep — gerbang merge satu-satunya: buang box sampah (filter outlier), atur augmentasi, lalu Confirm merge → membuat/menambah dataset.
  7. Datasets — dataset bernama; aturan triage dibekukan ke dataset saat merge.
  8. Models & Training — fine-tune dari base model, lalu skor base vs model baru di val set yang sama (mAP50, mAP50-95, precision, recall + delta bertanda).
  9. Live Counting & Counting Accuracy — jalankan counter (tracker + line-cross / possession) pada video atau RTSP; bandingkan hasil hitung AI vs ground truth manusia per rekaman, per siklus.

Plus SAM3 Playground (#/sam3-playground) — scratchpad tanpa project untuk mencari kalimat prompt yang jitu.

Rute yang harus tetap hidup

#/projects
#/projects/{id}                          → Video Archive
#/projects/{id}/trim/{encodedRel}
#/projects/{id}/batches
#/projects/{id}/review?batch={batchId}   (alias: #/batches/{batchId})
#/projects/{id}/data-prep?batches=1,2,3
#/projects/{id}/datasets
#/projects/{id}/models
#/projects/{id}/live-count
#/projects/{id}/counting-bench
#/sam3-playground

?batches= dan ?batch= adalah cara Batches menyerahkan seleksi ke Data Prep dan cara Review di-deep-link — jangan dihapus.

Shell

Top bar tetap 48 px: wordmark kiri → nav urut pipeline di tengah → kanan badge health (GPU: …, VRAM: … GB, SAM3: Ready/Off hijau saat siap, polling 3 detik) + toggle tema. Error boundary membungkus area halaman, di-key per rute.

Design token sekarang (boleh diganti, maknanya harus bertahan)

--bg #0b0f19 · --panel rgba(17,24,39,.5) · --panel-raised rgba(30,41,59,.7)
--border rgba(255,255,255,.12) · --text #f3f4f6 · --text-muted #9ca3af
--accent #a855f7 · --ok #10b981 · --warn #f59e0b · --danger #ef4444
--radius 12px · --space 8px · Inter / JetBrains Mono · transition 160ms

Warna semantik (jangan dipakai untuk dekorasi): hijau #4ade80 = kept/approved/positive exemplar/metrik lebih baik · merah #f87171 = dropped/rejected/negative exemplar/metrik lebih buruk · amber #fbbf24 = perlu dicek (jam tidak terverifikasi, frame tertahan, belum tersimpan) · biru #38bdf8 = seleksi/marquee · ungu #c084fc = kerja mesin sedang jalan (job, SAM3).

Warna class = data, bukan dekorasi: ["#f59e0b","#38bdf8","#10b981","#facc15","#6366f1","#f97316","#ec4899","#9ca3af"][classId % 8] — index class yang sama wajib berwarna sama di semua tempat (badge filmstrip, stroke canvas, chip class, tag kartu project, crop grid).

Review editor — layar terpenting

Layout: head (batch + progress + Approve All + Prepare & merge) → banner job → grid: kiri canvas

  • frame-bar + filmstrip, kanan sidebar (panel filter exemplar, Classes, Shapes on this frame, Shortcuts).
  • Overlay pakai SVG (bukan <canvas>), viewBox="0 0 frameW frameH", preserveAspectRatio="none". Wrapper overlay memeluk gambar persis dan tidak berisi elemen lain.
  • Canvas dibatasi max-width: calc(72vh * width / height).
  • Dua mode: Draw / Select (V). Select mode tidak pernah mengubah geometri; Draw mode tidak marquee.
  • Di draw mode, drag bukan menggambar kotak — itu adalah contoh. User menggambar satu instance, SAM3 mencari sisanya, hasilnya di-preview, di-tune lewat 4 slider (Confidence, Overlap/NMS, Min box size, Max shapes), lalu baru Apply yang menulis. Shift-drag = contoh negatif ("bukan ini"). Panel filter duduk di sidebar, bukan menutupi frame.
  • Tahan S = SAM3 click-assist (satu box → satu shape).
  • Keyboard-first: V ← → A X U N C T 1–9 Del Esc Enter. Semua aksi punya shortcut; panel Shortcuts selalu terlihat. Ctrl/Cmd/Alt tidak pernah di-intercept; mengetik di input tidak memicu shortcut.
  • Approve/Reject langsung pindah frame dan meng-update status lokal sebelum request selesai.

Aturan mutlak — melanggar ini merusak pipeline atau membuat angkanya berbohong

Versi panjangnya ada di §6.3.

Integritas data

  1. Geometri normalized 0–1 ujung-ke-ujung. bbox {points:[x0,y0,x1,y1]}, polygon {points:[[x,y]…]}. Dua pengecualian sengaja: exemplar ke /preview pakai [cx,cy,w,h], ke /exemplar-label pakai [x0,y0,x1,y1] — jangan "diseragamkan".
  2. source: "auto" | "manual" wajib disimpan dan ditampilkan — re-run batch hanya menghapus baris auto, koreksi manusia harus selamat.
  3. Val split stabil. Frame yang sudah masuk val selamanya val. UI tidak boleh pernah menyediakan tombol "reshuffle split".
  4. Aturan triage dibekukan ke dataset saat merge; copy yang menjelaskan ini, dan peringatan Resync, adalah bagian dari fitur.
  5. Menghapus class merenumber class di atasnya, di DB dan di file label. Konfirmasi harus menyebut jumlah shape yang mati dan file yang ditulis ulang.
  6. Gerbang merge hanya di Data Prep — bukan di Review, bukan di Batches.
  7. Merge batch yang belum di-review butuh checkbox pengakuan eksplisit.
  8. Waktu rekaman ditampilkan dengan slicing string, tidak pernah di-new Date() (offset browser menggeser semuanya berjam-jam).

Perilaku mesin

  1. Satu set_image per gambar; jangan ada UI yang menyiratkan "jalankan tiap prompt terpisah".
  2. Exemplar di modal auto-annotate = alat tuning, tidak pernah jadi label.
  3. Exemplar run di Review = dry-run sampai Apply; Apply menjalankan ulang di server, bukan mengirim balik geometri hasil preview.
  4. Seluruh pool exemplar dikirim ulang tiap run (itu juga mekanisme undo-nya).
  5. Satu pass sekali waktu, dengan maksimal satu rerun antre.
  6. Debounce ada karena biaya GPU, bukan estetika: 400 ms setelah drag, 250 ms setelah slider, 300 ms untuk triage simulate.
  7. Mass auto-annotate submit berurutan, bukan Promise.all.
  8. Job dipolling hanya selagi ada job jalan; data di-reload sekali saat job terakhir selesai.
  9. HTTP 409 = GPU lock dipegang job lain — harus muncul sebagai pesan tersendiri.
  10. Progress selalu progress/total + bar + tombol cancel. Spinner tanpa ujung tidak diterima.

Yang bebas diubah / yang harus diperbaiki

Bebas (§6.1): spacing, radius, tipografi, palet persisnya, pilihan component library, pendekatan CSS, tabel vs kartu, penempatan section, collapse/expand, breakpoint, focus ring, reduced-motion, menggabungkan dua modal auto-annotate yang 70% duplikat, mengganti hash routing dengan router sungguhan (asal bentuk URL di atas tetap jalan).

Wajib diperbaiki (§6.4):

  • Emoji sebagai ikon (🔄 🏷️ 📋 🚀 📦 🤖 ⚡ 🖼️) → ganti SVG (Heroicons/Lucide).
  • window.confirm / prompt / alert → dialog sungguhan, dengan konsekuensi tetap dienumerasi.
  • Copy campur Inggris–Indonesia → satu bahasa per layar (istilah siklus, batch, truk adalah kosakata operator, boleh tetap).
  • Inline style masif → design token.
  • Empty state harus berupa kalimat yang menyebut aksi berikutnya, bukan "No data".
  • Error backend datang sebagai {"detail": "..."} — tampilkan pesannya, bukan "Error 500".

Checklist pra-serah

  • Tidak ada emoji sebagai ikon (SVG saja)
  • cursor: pointer di setiap elemen yang bisa diklik
  • Hover state dengan transisi 150–300 ms
  • Kontras teks minimal 4.5:1
  • Focus state terlihat untuk navigasi keyboard
  • prefers-reduced-motion dihormati
  • Responsif di 375 / 768 / 1024 / 1440 px
  • Halaman tidak pernah scroll horizontal
  • Urutan nav = urutan pipeline, dan tiap link project-scoped membawa project id yang valid

Yang diminta

Rancang ulang tampilan reTraining sebagai dense internal tool untuk sesi kerja panjang: frame video dan canvas anotasi adalah kontennya, sisanya adalah chrome yang diam. Tanpa gradient dekoratif, tanpa motion marketing. Kepadatan informasi dan jangkauan keyboard lebih penting daripada whitespace. Konsistensi antar halaman lebih penting daripada kepintaran per halaman.

Kirimkan: design system (palet, tipografi, spacing, komponen dasar) lebih dulu, lalu layout tiap layar di daftar pipeline, lalu detail Review editor.