reTraining

Take a model you already have, and make it better with footage you already have.

A self-hosted loop for turning raw CCTV into a measurably better detector — record, extract, auto-label, review, filter, train, and prove the new model actually beat the old one.

Python FastAPI React Vite SAM3 YOLO Docker GPU

Quick start · How it works · The screens · Counting · Trust the numbers · Troubleshooting


⚡ Quick start

cp .env.example .env                       # paste your HF_TOKEN
VIDEO_ARCHIVE_HOST=/path/to/videos docker compose up -d --build

Open http://localhost:8080. The API is on :8000.

curl localhost:8000/api/health
{ "device": "cuda", "gpu": "NVIDIA GeForce RTX 5080 Laptop GPU",
  "vram_free_gb": 14.91, "sam3_ready": true, "ffmpeg": true,
  "hf_token": true, "db": true }

Important

The first auto-annotation job downloads the ~3.4 GB SAM3 checkpoint into a Docker volume. It happens once; later jobs take about 12 seconds to load the model into VRAM.

What you need first
Thing Why
A GPU with the NVIDIA container toolkit SAM3 is CUDA-only
A video archive laid out as <date>/<batch>.mp4 that structure is what the archive browser reads
HF_TOKEN with access to facebook/sam3 the weights are gated, and approval is manual
A base model .pt (optional) without one, training starts from yolo11n.pt and you type the classes yourself
videos/
  2026-08-13/
    batch001.mp4
    batch002.mp4
  2026-08-14/
    batch001.mp4
Run it without Docker (development)
uv pip install -r requirements.txt        # backend
uv pip install -e sam3/
uv run uvicorn backend.main:app --reload  # :8000

cd frontend && npm install
npm run dev                               # :5173, proxies /api to :8000

The frontend pins Vite 7 on purpose — Vite 8's Rolldown binding crashes on this machine. The package manager is uv; there is no pip/poetry path.


🔄 How it works

One project owns its base model, its class list, its video archive and its own accumulating datasets. A second use case is a second project — not a second copy of the code.

flowchart LR
    A[📹 Archive<br/>one file per truck session] --> B[✂️ Trim<br/>pick a range + fps]
    B --> C[🖼️ Extract<br/>frames to disk]
    C --> D[🤖 Auto-label<br/>SAM3 text prompts]
    D --> E[👁️ Review<br/>fix every frame]
    E --> F[🧹 Data Prep<br/>filter + augment]
    F --> G[📦 Dataset<br/>named, immutable]
    G --> H[🎯 Train<br/>fine-tune from base]
    H --> I[📊 Compare<br/>base vs new, same val set]
    I -.->|promote| H

Note

Data Prep is the gate. Selecting batches does not create a dataset — it opens Data Prep scoped to that selection. Only Confirm merge cuts the dataset, and the filter rules in force at that moment are frozen onto it, so editing them later can never rewrite a dataset you already trained on.


🖥️ The screens

1. Projects — name, label type, archive root, classes

Upload a base model and its classes are read from the checkpoint and locked, so the dataset and the model can never drift apart.

2. Video Archive — browse by cycle, not by folder

A cycle is one shift: 06:00 → 05:59 the next morning. It always crosses midnight, so it always spans two calendar dates, and is named after the date it starts on.

Folder names are not when a recording was made, and neither are file mtimes — those are file copy times. The real start comes from the timestamp the camera burns into every frame, so a file sitting in the 2026-08-14 folder but recorded at 00:11 shows up as part of the 13 Aug cycle, numbered in the order it was actually made.

Siklus 13 Agt 2026                                        28 rekaman
  #1    08:27:27   batch003    2026-08-13
  ...
  #25   23:53:45   batch027    2026-08-13
  #26   00:11:02   batch001    2026-08-14  ← pulled in from the next folder
  #27   00:34:01   batch002    2026-08-14

Nothing in the archive is moved or renamed — it is mounted read-only. The grouping lives in an index beside it, and the original path stays the file's identity.

3. Trim — play the video, set in/out, pick a frame rate

It tells you how many frames that produces before you commit to it.

4. Review — the frame with its shapes on top
Key Action Key Action
A approve ← → previous / next frame
X reject U jump to next unreviewed
Del delete shape 1–9 pick class
S + drag SAM3-assisted shape drag add / move / resize a box

Approving no longer merges. It marks frames approved; merging happens in Data Prep.

5. Data Prep — throw out the junk, then set augmentation

Tune an outlier filter over score / area / aspect against exactly the batches you picked, watching the counts move as you drag. A dropped box leaves its image in the dataset; only a frame that loses every box is held back — these frames hold ~44 objects each, and excluding the whole image was measured to cost 96% of a batch to remove 10% of its boxes.

Then Confirm merge creates the dataset under a frozen copy of those rules.

6. Datasets — several per project, each a standalone copy

batch7+8 strict rules and batch7+8 after I fixed the annotations are two datasets holding the same frames with different labels. Combining them for a run is newest wins, so a frame appearing twice is emitted once rather than teaching the model two contradictory labels.

7. Models & Training — run, then read the comparison

Pick datasets, pick classes, train. Batch size, image size and device default from the hardware actually detected, so a bigger GPU changes the numbers in the form, not the code.

8. Live Counting — point a model at a camera and watch it count

Same tracker, stabiliser and line-cross counter the production script uses. Click the video to place the counting line; it moves live without losing the counts. Every finished track is written to a JSONL with the reason it did or did not count — which is what separates a model miss from a tracker miss from a counter miss.

9. Counting Accuracy — scored, per cycle

One row per recording, grouped into collapsible cycles. Type in the ground truth you counted by hand and the table shows the signed delta — +3 and -3 are different failures, and a single accuracy percentage hides which one you have. Totals only ever count rows where a ground truth is filled in.

Recount runs headless in a background job — no annotated frame, no JPEG encode, which is worth 146 fps vs 124 in the live view.


🎥 Counting

The counter is deliberately robust to low frame rates: it never needs to catch the exact frame of a crossing, only that a track was seen above the line at some point in its life.

stateDiagram-v2
    [*] --> UNKNOWN
    UNKNOWN --> ABOVE: y1 above the band
    UNKNOWN --> BELOW: born below (ghost — never counts)
    ABOVE --> COUNTED: seen below + travelled far enough
    COUNTED --> ABOVE: sustained frames above (real unload)
Four layers, and what each one is actually for
Layer Guard Stops
1 Must have been above the line at some point a box that appears inside the truck
2 Must have travelled entry_travel_min from where it first appeared ghost boxes that blink into existence next to the line
3 Track hand-off — a dying track parks its history for a newborn nearby to inherit an ID switch at the line losing the count or duplicating it
4 One count per direction per track, and the verdict is the track's last direction double counting, while still letting a genuine unload-and-reload count again

Every one of these was written against a reproduced failure. Hand-off replaced a spatial dedup that did not dedup: a blocked track simply retried each frame and counted anyway once it drifted out of the circle — late, at the wrong position, seeding the next circle in the wrong place.

handoff_radius is the dial that matters most. These frames hold ~44 objects, so a newborn track is nearly always near one that just vanished; calibrate it against a clip with a hand-counted total rather than by eye.

Recording pipeline — record once, cut sessions afterwards
flowchart LR
    CAM[📷 Dahua 1080p<br/>H.265 @ 25fps] --> MTX[MediaMTX on Jetson<br/>records 24/7 · 24h buffer]
    MTX -->|RTSP| DET[Truck detector<br/>on the GPU box]
    DET -->|session ends| FETCH[Download that exact<br/>time range as a copy]
    MTX --> FETCH
    FETCH --> ARC[📁 archive/date/batchNNN.mp4<br/>+ .json sidecar]

The detector does not encode video. When a truck session ends it downloads that range from the recording server, so the archive keeps the camera's own codec, resolution and frame rate.

Before After
Codec mpeg4 re-encode HEVC copy
Resolution 1280×720 1920×1080
Size 8.0 Mbps 1.72 Mbps (4.7× smaller)
Timebase declared 10 fps at 25 fps real → 2.49× slow true 25 fps
Start time read from the burned-in overlay by OCR from the server, exact

Each clip carries a .json sidecar with the server's start time, which the app trusts over reading the overlay — so a new session appears in the right cycle with no scan at all.


📊 Why the numbers are trustworthy

After training, the base model and the new one are validated on the same val set, and mAP50 / mAP50-95 / precision / recall appear side by side with the difference.

Tip

The val split is stable. Once a frame is in val it stays there for every later merge — derived from the frame's identity, not from how many rows precede it. A rising score cannot be an easier val set.

  • Training uses whole datasets, old batches included. Fine-tuning on the newest batch alone tends to raise the score on new footage while quietly losing the old.
  • Datasets are snapshots. Labels are copied from what the dataset holds on disk, not re-derived from today's rules, so two runs over the same dataset cannot disagree.
  • An empty Base column is honest. It means the previous model's classes did not match this dataset's, so scoring it would have compared two different things. The message says which.

Use as base model promotes a version, and the next round fine-tunes from it.


🗂️ Where things live

data/
  app.db                                  # 14 tables: projects, frames, annotations,
                                          # datasets, jobs, count_runs, video_clock …
  archive/<date>/batchNNN.mp4             # recordings (read-only to the app)
                 batchNNN.json            # sidecar: true start time from the server
  recorder.log                            # the 24/7 recorder's output
  live-count/session-*.jsonl              # per-track traces from the live counter
  projects/<slug>/
    base/model.pt                         # the base model
    datasets/<id>/                        # one folder per named dataset
      images/{train,val}/  labels/{train,val}/
    batches/<id>/frames/                  # extracted frames
    models/<n>/best.pt + metrics.json     # each training run

The database holds status; the disk holds pixels, labels and weights. A dataset trains as-is with Ultralytics, or imports into Roboflow, without this application.

Warning

Your video archive is mounted read-only and nothing is ever written back into it.


⚙️ Jobs

Heavy work runs as queued jobs, one at a time, so the GPU is never double-booked.

Type GPU What it does
extract — ffmpeg pulls frames out of a range
autolabel ✅ SAM3 over a batch, one set_image per image
merge — copies approved frames into a dataset under frozen rules
train ✅ fine-tunes, then validates base and new on the same val set
count ✅ headless recount of archive videos for the accuracy table
clock-scan — reads each recording's real start time
truck-scan ✅ checks every recording actually contains a truck

Jobs and their progress are persistent; after a restart the list is still there.


🔧 When something goes wrong

Deployment and GPU
Symptom Cause / fix
API returns 404 for routes you just added the image copies backend/ at build time — docker compose build backend again
A UI change doesn't show up same trap on the other side: docker compose build frontend, then hard-reload
could not select device driver the NVIDIA container toolkit is not installed, or Docker is older than the CDI support compose relies on (devices: nvidia.com/gpu=all)
CUDA out of memory while training lower epochs/batch on the Models page, or free the card — SAM3 is released before training, but another process may still hold it
A job reads interrupted by a server restart it was running when the process died — jobs are not resumable, start it again
SAM3 and auto-labelling
Symptom Cause / fix
Job fails at loading model with a 401 access to facebook/sam3 not granted yet, or HF_TOKEN missing
SAM3 download crawls at a few KB/s HuggingFace's Xet transfer throttling itself; HF_HUB_DISABLE_XET=1 is already set in compose for that reason
Archive and counting
Symptom Cause / fix
Videos listed as unreadable ffprobe could not parse them; they are still listed rather than hidden, so the archive never looks emptier than it is
A recording's time shows amber its overlay was read with low confidence, or not at all — type the time you can see in the video; a hand-entered time is never overwritten by a rescan
A cycle looks short check whether its recordings moved to the neighbouring cycle — anything before 06:00 belongs to the previous shift
Live counting says GPU busy a training or auto-label job holds the card; it waits 30 s before giving up

🤝 Contributing

The documents drive the repo, not the other way round:

Document Contents
docs/requirements.md numbered REQ-xxx, changed only with the owner's approval
docs/design.md schema, API contract, disk layout — each section names the REQ-xxx it serves
docs/tasks.md implementation steps and how each was verified, [TODO] / [DONE]
AGENTS.md working rules: simplicity, surgical changes, verify by running something

Two invariants are easy to break and make the whole system lie:

  1. A frame in val stays in val — otherwise the base-vs-new comparison is meaningless.
  2. One set_image per image — set_text_prompt() re-runs only the grounding head against the cached backbone output. An N-prompt job calls set_image once and loops prompts over that state.
S
Description
used for retraining and annotation of karung feedmill project
Readme MIT
30 MiB
0 Stars 1 Watchers 0 Forks
Languages
Python 64.4%
JavaScript 32.6%
CSS 2.4%
Shell 0.4%
Dockerfile 0.2%