Files
reTraining/README.md
T

445 lines
17 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<div align="center">
# reTraining
**Take a model you already have, and make it better with footage you already have.**
A self-hosted loop for turning raw CCTV into a measurably better detector — record, extract,
auto-label, review, filter, train, and prove the new model actually beat the old one.
<p>
<img alt="Python" src="https://img.shields.io/badge/python-3.12-3776AB?logo=python&logoColor=white">
<img alt="FastAPI" src="https://img.shields.io/badge/FastAPI-72%20endpoints-009688?logo=fastapi&logoColor=white">
<img alt="React" src="https://img.shields.io/badge/React-19-61DAFB?logo=react&logoColor=black">
<img alt="Vite" src="https://img.shields.io/badge/Vite-7-646CFF?logo=vite&logoColor=white">
<img alt="SAM3" src="https://img.shields.io/badge/SAM3-auto--label-FF6F00">
<img alt="YOLO" src="https://img.shields.io/badge/Ultralytics-YOLO11-00BFA5">
<img alt="Docker" src="https://img.shields.io/badge/docker-compose-2496ED?logo=docker&logoColor=white">
<img alt="GPU" src="https://img.shields.io/badge/GPU-required-76B900?logo=nvidia&logoColor=white">
</p>
[Quick start](#-quick-start) · [How it works](#-how-it-works) · [The screens](#-the-screens) ·
[Counting](#-counting) · [Trust the numbers](#-why-the-numbers-are-trustworthy) ·
[Troubleshooting](#-when-something-goes-wrong)
</div>
---
## ⚡ Quick start
```bash
cp .env.example .env # paste your HF_TOKEN
VIDEO_ARCHIVE_HOST=/path/to/videos docker compose up -d --build
```
Open **<http://localhost:8080>**. The API is on `:8000`.
```bash
curl localhost:8000/api/health
```
```json
{ "device": "cuda", "gpu": "NVIDIA GeForce RTX 5080 Laptop GPU",
"vram_free_gb": 14.91, "sam3_ready": true, "ffmpeg": true,
"hf_token": true, "db": true }
```
> [!IMPORTANT]
> The first auto-annotation job downloads the **~3.4 GB SAM3 checkpoint** into a Docker
> volume. It happens once; later jobs take about 12 seconds to load the model into VRAM.
<details>
<summary><b>What you need first</b></summary>
<br>
| Thing | Why |
|---|---|
| A GPU with the NVIDIA container toolkit | SAM3 is CUDA-only |
| A video archive laid out as `<date>/<batch>.mp4` | that structure is what the archive browser reads |
| `HF_TOKEN` with access to [facebook/sam3](https://huggingface.co/facebook/sam3) | the weights are gated, and approval is manual |
| A base model `.pt` (optional) | without one, training starts from `yolo11n.pt` and you type the classes yourself |
```
videos/
2026-08-13/
batch001.mp4
batch002.mp4
2026-08-14/
batch001.mp4
```
</details>
<details>
<summary><b>Run it without Docker (development)</b></summary>
<br>
```bash
uv pip install -r requirements.txt # backend
uv pip install -e sam3/
uv run uvicorn backend.main:app --reload # :8000
cd frontend && npm install
npm run dev # :5173, proxies /api to :8000
```
The frontend pins Vite 7 on purpose — Vite 8's Rolldown binding crashes on this machine.
The package manager is `uv`; there is no `pip`/`poetry` path.
</details>
---
## 🔄 How it works
One project owns its base model, its class list, its video archive and its own accumulating
datasets. A second use case is a second project — not a second copy of the code.
```mermaid
flowchart LR
A[📹 Archive<br/>one file per truck session] --> B[✂️ Trim<br/>pick a range + fps]
B --> C[🖼️ Extract<br/>frames to disk]
C --> D[🤖 Auto-label<br/>SAM3 text prompts]
D --> E[👁️ Review<br/>fix every frame]
E --> F[🧹 Data Prep<br/>filter + augment]
F --> G[📦 Dataset<br/>named, immutable]
G --> H[🎯 Train<br/>fine-tune from base]
H --> I[📊 Compare<br/>base vs new, same val set]
I -.->|promote| H
```
> [!NOTE]
> **Data Prep is the gate.** Selecting batches does not create a dataset — it opens Data Prep
> scoped to that selection. Only *Confirm merge* cuts the dataset, and the filter rules in
> force at that moment are **frozen onto it**, so editing them later can never rewrite a
> dataset you already trained on.
---
## 🖥️ The screens
<details open>
<summary><b>1. Projects</b> — name, label type, archive root, classes</summary>
<br>
Upload a base model and its classes are read from the checkpoint and locked, so the dataset
and the model can never drift apart.
</details>
<details>
<summary><b>2. Video Archive</b> — browse by <i>cycle</i>, not by folder</summary>
<br>
A **cycle** is one shift: `06:00 → 05:59` the next morning. It always crosses midnight, so it
always spans two calendar dates, and is named after the date it starts on.
Folder names are not when a recording was made, and neither are file mtimes — those are file
*copy* times. The real start comes from the timestamp the camera burns into every frame, so a
file sitting in the `2026-08-14` folder but recorded at `00:11` shows up as part of the
**13 Aug** cycle, numbered in the order it was actually made.
```
Siklus 13 Agt 2026 28 rekaman
#1 08:27:27 batch003 2026-08-13
...
#25 23:53:45 batch027 2026-08-13
#26 00:11:02 batch001 2026-08-14 ← pulled in from the next folder
#27 00:34:01 batch002 2026-08-14
```
Nothing in the archive is moved or renamed — it is mounted **read-only**. The grouping lives
in an index beside it, and the original path stays the file's identity.
</details>
<details>
<summary><b>3. Trim</b> — play the video, set in/out, pick a frame rate</summary>
<br>
It tells you how many frames that produces before you commit to it.
</details>
<details>
<summary><b>4. Review</b> — the frame with its shapes on top</summary>
<br>
| Key | Action | | Key | Action |
|---|---|---|---|---|
| `A` | approve | | `←` `→` | previous / next frame |
| `X` | reject | | `U` | jump to next unreviewed |
| `Del` | delete shape | | `1`–`9` | pick class |
| `S` + drag | SAM3-assisted shape | | drag | add / move / resize a box |
Approving no longer merges. It marks frames approved; merging happens in Data Prep.
</details>
<details>
<summary><b>5. Data Prep</b> — throw out the junk, then set augmentation</summary>
<br>
Tune an outlier filter over **score / area / aspect** against exactly the batches you picked,
watching the counts move as you drag. A dropped box leaves its image in the dataset; only a
frame that loses *every* box is held back — these frames hold ~44 objects each, and excluding
the whole image was measured to cost 96% of a batch to remove 10% of its boxes.
Then **Confirm merge** creates the dataset under a frozen copy of those rules.
</details>
<details>
<summary><b>6. Datasets</b> — several per project, each a standalone copy</summary>
<br>
`batch7+8 strict rules` and `batch7+8 after I fixed the annotations` are two datasets holding
the same frames with different labels. Combining them for a run is *newest wins*, so a frame
appearing twice is emitted once rather than teaching the model two contradictory labels.
</details>
<details>
<summary><b>7. Models & Training</b> — run, then read the comparison</summary>
<br>
Pick datasets, pick classes, train. Batch size, image size and device default from the
hardware actually detected, so a bigger GPU changes the numbers in the form, not the code.
</details>
<details>
<summary><b>8. Live Counting</b> — point a model at a camera and watch it count</summary>
<br>
Same tracker, stabiliser and line-cross counter the production script uses. Click the video to
place the counting line; it moves live without losing the counts. Every finished track is
written to a JSONL with the reason it did or did not count — which is what separates a model
miss from a tracker miss from a counter miss.
</details>
<details>
<summary><b>9. Counting Accuracy</b> — scored, per cycle</summary>
<br>
One row per recording, grouped into collapsible cycles. Type in the ground truth you counted
by hand and the table shows the **signed delta** — `+3` and `-3` are different failures, and a
single accuracy percentage hides which one you have. Totals only ever count rows where a
ground truth is filled in.
Recount runs headless in a background job — no annotated frame, no JPEG encode, which is worth
**146 fps vs 124** in the live view.
</details>
---
## 🎥 Counting
The counter is deliberately robust to low frame rates: it never needs to catch the exact frame
of a crossing, only that a track was seen *above* the line at some point in its life.
```mermaid
stateDiagram-v2
[*] --> UNKNOWN
UNKNOWN --> ABOVE: y1 above the band
UNKNOWN --> BELOW: born below (ghost — never counts)
ABOVE --> COUNTED: seen below + travelled far enough
COUNTED --> ABOVE: sustained frames above (real unload)
```
<details>
<summary><b>Four layers, and what each one is actually for</b></summary>
<br>
| Layer | Guard | Stops |
|---|---|---|
| 1 | Must have been **above** the line at some point | a box that appears inside the truck |
| 2 | Must have travelled `entry_travel_min` from where it first appeared | ghost boxes that blink into existence next to the line |
| 3 | **Track hand-off** — a dying track parks its history for a newborn nearby to inherit | an ID switch at the line losing the count *or* duplicating it |
| 4 | One count per direction per track, and the verdict is the track's *last* direction | double counting, while still letting a genuine unload-and-reload count again |
Every one of these was written against a reproduced failure. Hand-off replaced a spatial dedup
that did not dedup: a blocked track simply retried each frame and counted anyway once it
drifted out of the circle — late, at the wrong position, seeding the next circle in the wrong
place.
`handoff_radius` is the dial that matters most. These frames hold ~44 objects, so a newborn
track is nearly always near one that just vanished; calibrate it against a clip with a
hand-counted total rather than by eye.
</details>
<details>
<summary><b>Recording pipeline</b> — record once, cut sessions afterwards</summary>
<br>
```mermaid
flowchart LR
CAM[📷 Dahua 1080p<br/>H.265 @ 25fps] --> MTX[MediaMTX on Jetson<br/>records 24/7 · 24h buffer]
MTX -->|RTSP| DET[Truck detector<br/>on the GPU box]
DET -->|session ends| FETCH[Download that exact<br/>time range as a copy]
MTX --> FETCH
FETCH --> ARC[📁 archive/date/batchNNN.mp4<br/>+ .json sidecar]
```
The detector does **not** encode video. When a truck session ends it downloads that range from
the recording server, so the archive keeps the camera's own codec, resolution and frame rate.
| | Before | After |
|---|---|---|
| Codec | mpeg4 re-encode | HEVC copy |
| Resolution | 1280×720 | 1920×1080 |
| Size | 8.0 Mbps | **1.72 Mbps** (4.7× smaller) |
| Timebase | declared 10 fps at 25 fps real → **2.49× slow** | true 25 fps |
| Start time | read from the burned-in overlay by OCR | from the server, exact |
Each clip carries a `.json` sidecar with the server's start time, which the app trusts over
reading the overlay — so a new session appears in the right cycle with no scan at all.
</details>
---
## 📊 Why the numbers are trustworthy
After training, the base model and the new one are validated **on the same val set**, and
mAP50 / mAP50-95 / precision / recall appear side by side with the difference.
> [!TIP]
> **The val split is stable.** Once a frame is in `val` it stays there for every later merge —
> derived from the frame's identity, not from how many rows precede it. A rising score cannot
> be an easier val set.
- **Training uses whole datasets, old batches included.** Fine-tuning on the newest batch alone
tends to raise the score on new footage while quietly losing the old.
- **Datasets are snapshots.** Labels are copied from what the dataset holds on disk, not
re-derived from today's rules, so two runs over the same dataset cannot disagree.
- **An empty Base column is honest.** It means the previous model's classes did not match this
dataset's, so scoring it would have compared two different things. The message says which.
`Use as base model` promotes a version, and the next round fine-tunes from it.
---
## 🗂️ Where things live
```
data/
app.db # 14 tables: projects, frames, annotations,
# datasets, jobs, count_runs, video_clock …
archive/<date>/batchNNN.mp4 # recordings (read-only to the app)
batchNNN.json # sidecar: true start time from the server
recorder.log # the 24/7 recorder's output
live-count/session-*.jsonl # per-track traces from the live counter
projects/<slug>/
base/model.pt # the base model
datasets/<id>/ # one folder per named dataset
images/{train,val}/ labels/{train,val}/
batches/<id>/frames/ # extracted frames
models/<n>/best.pt + metrics.json # each training run
```
The database holds status; the disk holds pixels, labels and weights. A dataset trains as-is
with Ultralytics, or imports into Roboflow, without this application.
> [!WARNING]
> Your video archive is mounted **read-only** and nothing is ever written back into it.
---
## ⚙️ Jobs
Heavy work runs as queued jobs, one at a time, so the GPU is never double-booked.
| Type | GPU | What it does |
|---|---|---|
| `extract` | — | ffmpeg pulls frames out of a range |
| `autolabel` | ✅ | SAM3 over a batch, one `set_image` per image |
| `merge` | — | copies approved frames into a dataset under frozen rules |
| `train` | ✅ | fine-tunes, then validates base and new on the same val set |
| `count` | ✅ | headless recount of archive videos for the accuracy table |
| `clock-scan` | — | reads each recording's real start time |
| `truck-scan` | ✅ | checks every recording actually contains a truck |
Jobs and their progress are persistent; after a restart the list is still there.
---
## 🔧 When something goes wrong
<details>
<summary><b>Deployment and GPU</b></summary>
<br>
| Symptom | Cause / fix |
|---|---|
| API returns 404 for routes you just added | the image copies `backend/` at build time — `docker compose build backend` again |
| A UI change doesn't show up | same trap on the other side: `docker compose build frontend`, then hard-reload |
| `could not select device driver` | the NVIDIA container toolkit is not installed, or Docker is older than the CDI support compose relies on (`devices: nvidia.com/gpu=all`) |
| `CUDA out of memory` while training | lower epochs/batch on the Models page, or free the card — SAM3 is released before training, but another process may still hold it |
| A job reads *interrupted by a server restart* | it was running when the process died — jobs are not resumable, start it again |
</details>
<details>
<summary><b>SAM3 and auto-labelling</b></summary>
<br>
| Symptom | Cause / fix |
|---|---|
| Job fails at *loading model* with a 401 | access to `facebook/sam3` not granted yet, or `HF_TOKEN` missing |
| SAM3 download crawls at a few KB/s | HuggingFace's Xet transfer throttling itself; `HF_HUB_DISABLE_XET=1` is already set in compose for that reason |
</details>
<details>
<summary><b>Archive and counting</b></summary>
<br>
| Symptom | Cause / fix |
|---|---|
| Videos listed as *unreadable* | ffprobe could not parse them; they are still listed rather than hidden, so the archive never looks emptier than it is |
| A recording's time shows amber | its overlay was read with low confidence, or not at all — type the time you can see in the video; a hand-entered time is never overwritten by a rescan |
| A cycle looks short | check whether its recordings moved to the neighbouring cycle — anything before 06:00 belongs to the previous shift |
| Live counting says *GPU busy* | a training or auto-label job holds the card; it waits 30 s before giving up |
</details>
---
## 🤝 Contributing
The documents drive the repo, not the other way round:
| Document | Contents |
|---|---|
| [`docs/requirements.md`](docs/requirements.md) | numbered `REQ-xxx`, changed only with the owner's approval |
| [`docs/design.md`](docs/design.md) | schema, API contract, disk layout — each section names the `REQ-xxx` it serves |
| [`docs/tasks.md`](docs/tasks.md) | implementation steps and how each was *verified*, `[TODO]` / `[DONE]` |
| [`AGENTS.md`](AGENTS.md) | working rules: simplicity, surgical changes, verify by running something |
Two invariants are easy to break and make the whole system lie:
1. **A frame in `val` stays in `val`** — otherwise the base-vs-new comparison is meaningless.
2. **One `set_image` per image** — `set_text_prompt()` re-runs only the grounding head against
the cached backbone output. An N-prompt job calls `set_image` once and loops prompts over
that state.