docs: update README and add .env.example

This commit is contained in:
asus committed 2026-08-14 16:37:27 +07:00
1 parent 5c7c122105
commit b6624eeff9
2 files changed
+423 -70

No files matched your search

+34
View File
@@ -0,0 +1,34 @@
# Copy to .env and fill in. Only HF_TOKEN is required to get started.
#
# cp .env.example .env
# ---- required ---------------------------------------------------------------
# HuggingFace token with access to facebook/sam3. The weights are gated and
# approval is manual, so request access before the first auto-annotation job:
# https://huggingface.co/facebook/sam3
HF_TOKEN=
# ---- host paths -------------------------------------------------------------
# Where your recordings live, laid out as <date>/<batch>.mp4. Mounted read-only
# into the container; nothing is ever written back into it.
VIDEO_ARCHIVE_HOST=./data/archive
# ---- ports ------------------------------------------------------------------
# The UI. The API is always on :8000.
WEB_PORT=8080
# Origins allowed to call the API. Add your machine's LAN address to reach the
# dev server from another device.
CORS_ORIGINS=http://localhost:5173,http://localhost:8080
# ---- recorder (algoritma-batch/batch_video_cropper.py) -----------------------
# Only needed if you run the 24/7 truck-session recorder. It reads the RTSP
# stream to detect sessions, then downloads each session from the recording
# server as a copy rather than re-encoding it.
# MediaMTX playback endpoint and the path name to pull from.
PLAYBACK_URL=http://192.168.192.96:9996/get
PLAYBACK_PATH=cam
+389 -70
View File
@@ -1,125 +1,444 @@
# Dataset Enrichment - DEV
<div align="center">
Take a model you already have, and make it better with footage you already have.
# reTraining
One round of the loop:
**Take a model you already have, and make it better with footage you already have.**
> pick a project → browse the video archive → pick a batch → trim a range →
> extract frames → auto-annotate with SAM3 → review and correct every frame → approve →
> merge into the master dataset → fine-tune from the base model → compare against the base.
A self-hosted loop for turning raw CCTV into a measurably better detector — record, extract,
auto-label, review, filter, train, and prove the new model actually beat the old one.
Nothing here is specific to one dataset. A project owns its base model, its class list, its
video archive and its own accumulating dataset, so a second use case is a second project —
not a second copy of the code.
<p>
<img alt="Python" src="https://img.shields.io/badge/python-3.12-3776AB?logo=python&logoColor=white">
<img alt="FastAPI" src="https://img.shields.io/badge/FastAPI-72%20endpoints-009688?logo=fastapi&logoColor=white">
<img alt="React" src="https://img.shields.io/badge/React-19-61DAFB?logo=react&logoColor=black">
<img alt="Vite" src="https://img.shields.io/badge/Vite-7-646CFF?logo=vite&logoColor=white">
<img alt="SAM3" src="https://img.shields.io/badge/SAM3-auto--label-FF6F00">
<img alt="YOLO" src="https://img.shields.io/badge/Ultralytics-YOLO11-00BFA5">
<img alt="Docker" src="https://img.shields.io/badge/docker-compose-2496ED?logo=docker&logoColor=white">
<img alt="GPU" src="https://img.shields.io/badge/GPU-required-76B900?logo=nvidia&logoColor=white">
</p>
## Run it
[Quick start](#-quick-start) · [How it works](#-how-it-works) · [The screens](#-the-screens) ·
[Counting](#-counting) · [Trust the numbers](#-why-the-numbers-are-trustworthy) ·
[Troubleshooting](#-when-something-goes-wrong)
</div>
---
## ⚡ Quick start
```bash
cp .env.example .env # paste HF_TOKEN
cp .env.example .env # paste your HF_TOKEN
VIDEO_ARCHIVE_HOST=/path/to/videos docker compose up -d --build
```
Then open <http://localhost:8080>. The API is on `:8000` if you want to poke at it directly;
`curl localhost:8000/api/health` reports the GPU, free VRAM (`vram_free_gb`), SAM3 readiness (`sam3_ready`), whether ffmpeg is present, whether the token was picked up, and whether the database is reachable.
Open **<http://localhost:8080>**. The API is on `:8000`.
```bash
curl localhost:8000/api/health
```
```json
{ "device": "cuda", "gpu": "NVIDIA GeForce RTX 5080 Laptop GPU",
"vram_free_gb": 14.91, "sam3_ready": true, "ffmpeg": true,
"hf_token": true, "db": true }
```
The first auto-annotation job downloads the ~3.4 GB SAM3 checkpoint into a Docker volume.
It happens once; later jobs take about 12 seconds to load the model into VRAM.
> [!IMPORTANT]
> The first auto-annotation job downloads the **~3.4 GB SAM3 checkpoint** into a Docker
> volume. It happens once; later jobs take about 12 seconds to load the model into VRAM.
### What you need first
<details>
<summary><b>What you need first</b></summary>
<br>
| Thing | Why |
|---|---|
| A GPU with the NVIDIA container toolkit | SAM3 is CUDA-only |
| A video archive laid out as `<date>/<batch>.mp4` | that structure is what the Library reads |
| A video archive laid out as `<date>/<batch>.mp4` | that structure is what the archive browser reads |
| `HF_TOKEN` with access to [facebook/sam3](https://huggingface.co/facebook/sam3) | the weights are gated, and approval is manual |
| A base model `.pt` (optional) | without one, training starts from `yolo11n.pt` and you type the classes yourself |
Example archive:
```
videos/
2026-07-08/
batch-1.mp4
batch-4.mp4
2026-07-09/
batch-2.mp4
2026-08-13/
batch001.mp4
batch002.mp4
2026-08-14/
batch001.mp4
```
## The screens
</details>
1. **Projects** — name, label type (`bbox` or `polygon`), archive root, classes. Upload a
base model and its classes are read from the checkpoint and locked, so the dataset and
the model can never drift apart.
2. **Library** — dates on the left, that date's recordings on the right with duration,
resolution and how many batches already came out of each.
3. **Trim** — play the video, set in/out, pick a frame rate. It tells you how many frames
that produces before you commit to it.
4. **Review** — the frame with its shapes on top. Drag to add a box, drag a corner to
resize, drag the middle to move, `Del` to remove. Hold `S` and drag for a SAM3-assisted
shape. `A` approves, `X` rejects, `←`/`→` move, `U` jumps to the next unreviewed frame,
`1`–`9` pick the class. A batch can only be approved once no frame is still pending.
5. **Dataset & models** — what the master dataset holds, a training run, and the
base-versus-new table.
<details>
<summary><b>Run it without Docker (development)</b></summary>
## Reading the comparison
<br>
```bash
uv pip install -r requirements.txt # backend
uv pip install -e sam3/
uv run uvicorn backend.main:app --reload # :8000
cd frontend && npm install
npm run dev # :5173, proxies /api to :8000
```
The frontend pins Vite 7 on purpose — Vite 8's Rolldown binding crashes on this machine.
The package manager is `uv`; there is no `pip`/`poetry` path.
</details>
---
## 🔄 How it works
One project owns its base model, its class list, its video archive and its own accumulating
datasets. A second use case is a second project — not a second copy of the code.
```mermaid
flowchart LR
A[📹 Archive<br/>one file per truck session] --> B[✂️ Trim<br/>pick a range + fps]
B --> C[🖼️ Extract<br/>frames to disk]
C --> D[🤖 Auto-label<br/>SAM3 text prompts]
D --> E[👁️ Review<br/>fix every frame]
E --> F[🧹 Data Prep<br/>filter + augment]
F --> G[📦 Dataset<br/>named, immutable]
G --> H[🎯 Train<br/>fine-tune from base]
H --> I[📊 Compare<br/>base vs new, same val set]
I -.->|promote| H
```
> [!NOTE]
> **Data Prep is the gate.** Selecting batches does not create a dataset — it opens Data Prep
> scoped to that selection. Only *Confirm merge* cuts the dataset, and the filter rules in
> force at that moment are **frozen onto it**, so editing them later can never rewrite a
> dataset you already trained on.
---
## 🖥️ The screens
<details open>
<summary><b>1. Projects</b> — name, label type, archive root, classes</summary>
<br>
Upload a base model and its classes are read from the checkpoint and locked, so the dataset
and the model can never drift apart.
</details>
<details>
<summary><b>2. Video Archive</b> — browse by <i>cycle</i>, not by folder</summary>
<br>
A **cycle** is one shift: `06:00 → 05:59` the next morning. It always crosses midnight, so it
always spans two calendar dates, and is named after the date it starts on.
Folder names are not when a recording was made, and neither are file mtimes — those are file
*copy* times. The real start comes from the timestamp the camera burns into every frame, so a
file sitting in the `2026-08-14` folder but recorded at `00:11` shows up as part of the
**13 Aug** cycle, numbered in the order it was actually made.
```
Siklus 13 Agt 2026 28 rekaman
#1 08:27:27 batch003 2026-08-13
...
#25 23:53:45 batch027 2026-08-13
#26 00:11:02 batch001 2026-08-14 ← pulled in from the next folder
#27 00:34:01 batch002 2026-08-14
```
Nothing in the archive is moved or renamed — it is mounted **read-only**. The grouping lives
in an index beside it, and the original path stays the file's identity.
</details>
<details>
<summary><b>3. Trim</b> — play the video, set in/out, pick a frame rate</summary>
<br>
It tells you how many frames that produces before you commit to it.
</details>
<details>
<summary><b>4. Review</b> — the frame with its shapes on top</summary>
<br>
| Key | Action | | Key | Action |
|---|---|---|---|---|
| `A` | approve | | `←` `→` | previous / next frame |
| `X` | reject | | `U` | jump to next unreviewed |
| `Del` | delete shape | | `1`–`9` | pick class |
| `S` + drag | SAM3-assisted shape | | drag | add / move / resize a box |
Approving no longer merges. It marks frames approved; merging happens in Data Prep.
</details>
<details>
<summary><b>5. Data Prep</b> — throw out the junk, then set augmentation</summary>
<br>
Tune an outlier filter over **score / area / aspect** against exactly the batches you picked,
watching the counts move as you drag. A dropped box leaves its image in the dataset; only a
frame that loses *every* box is held back — these frames hold ~44 objects each, and excluding
the whole image was measured to cost 96% of a batch to remove 10% of its boxes.
Then **Confirm merge** creates the dataset under a frozen copy of those rules.
</details>
<details>
<summary><b>6. Datasets</b> — several per project, each a standalone copy</summary>
<br>
`batch7+8 strict rules` and `batch7+8 after I fixed the annotations` are two datasets holding
the same frames with different labels. Combining them for a run is *newest wins*, so a frame
appearing twice is emitted once rather than teaching the model two contradictory labels.
</details>
<details>
<summary><b>7. Models & Training</b> — run, then read the comparison</summary>
<br>
Pick datasets, pick classes, train. Batch size, image size and device default from the
hardware actually detected, so a bigger GPU changes the numbers in the form, not the code.
</details>
<details>
<summary><b>8. Live Counting</b> — point a model at a camera and watch it count</summary>
<br>
Same tracker, stabiliser and line-cross counter the production script uses. Click the video to
place the counting line; it moves live without losing the counts. Every finished track is
written to a JSONL with the reason it did or did not count — which is what separates a model
miss from a tracker miss from a counter miss.
</details>
<details>
<summary><b>9. Counting Accuracy</b> — scored, per cycle</summary>
<br>
One row per recording, grouped into collapsible cycles. Type in the ground truth you counted
by hand and the table shows the **signed delta** — `+3` and `-3` are different failures, and a
single accuracy percentage hides which one you have. Totals only ever count rows where a
ground truth is filled in.
Recount runs headless in a background job — no annotated frame, no JPEG encode, which is worth
**146 fps vs 124** in the live view.
</details>
---
## 🎥 Counting
The counter is deliberately robust to low frame rates: it never needs to catch the exact frame
of a crossing, only that a track was seen *above* the line at some point in its life.
```mermaid
stateDiagram-v2
[*] --> UNKNOWN
UNKNOWN --> ABOVE: y1 above the band
UNKNOWN --> BELOW: born below (ghost — never counts)
ABOVE --> COUNTED: seen below + travelled far enough
COUNTED --> ABOVE: sustained frames above (real unload)
```
<details>
<summary><b>Four layers, and what each one is actually for</b></summary>
<br>
| Layer | Guard | Stops |
|---|---|---|
| 1 | Must have been **above** the line at some point | a box that appears inside the truck |
| 2 | Must have travelled `entry_travel_min` from where it first appeared | ghost boxes that blink into existence next to the line |
| 3 | **Track hand-off** — a dying track parks its history for a newborn nearby to inherit | an ID switch at the line losing the count *or* duplicating it |
| 4 | One count per direction per track, and the verdict is the track's *last* direction | double counting, while still letting a genuine unload-and-reload count again |
Every one of these was written against a reproduced failure. Hand-off replaced a spatial dedup
that did not dedup: a blocked track simply retried each frame and counted anyway once it
drifted out of the circle — late, at the wrong position, seeding the next circle in the wrong
place.
`handoff_radius` is the dial that matters most. These frames hold ~44 objects, so a newborn
track is nearly always near one that just vanished; calibrate it against a clip with a
hand-counted total rather than by eye.
</details>
<details>
<summary><b>Recording pipeline</b> — record once, cut sessions afterwards</summary>
<br>
```mermaid
flowchart LR
CAM[📷 Dahua 1080p<br/>H.265 @ 25fps] --> MTX[MediaMTX on Jetson<br/>records 24/7 · 24h buffer]
MTX -->|RTSP| DET[Truck detector<br/>on the GPU box]
DET -->|session ends| FETCH[Download that exact<br/>time range as a copy]
MTX --> FETCH
FETCH --> ARC[📁 archive/date/batchNNN.mp4<br/>+ .json sidecar]
```
The detector does **not** encode video. When a truck session ends it downloads that range from
the recording server, so the archive keeps the camera's own codec, resolution and frame rate.
| | Before | After |
|---|---|---|
| Codec | mpeg4 re-encode | HEVC copy |
| Resolution | 1280×720 | 1920×1080 |
| Size | 8.0 Mbps | **1.72 Mbps** (4.7× smaller) |
| Timebase | declared 10 fps at 25 fps real → **2.49× slow** | true 25 fps |
| Start time | read from the burned-in overlay by OCR | from the server, exact |
Each clip carries a `.json` sidecar with the server's start time, which the app trusts over
reading the overlay — so a new session appears in the right cycle with no scan at all.
</details>
---
## 📊 Why the numbers are trustworthy
After training, the base model and the new one are validated **on the same val set**, and
mAP50 / mAP50-95 / precision / recall appear side by side with the difference.
Two things make that number trustworthy:
> [!TIP]
> **The val split is stable.** Once a frame is in `val` it stays there for every later merge —
> derived from the frame's identity, not from how many rows precede it. A rising score cannot
> be an easier val set.
- **The val split is stable.** Once a frame is in `val`, it stays there for every later
merge. A rising score cannot be an easier val set.
- **Training uses the whole master dataset**, old batches included. Fine-tuning on the newest
batch alone tends to raise the score on new footage while quietly losing the old.
If the Base column is empty, the previous model's classes did not match this dataset's, so
scoring it here would have compared two different things. The message says which case it was.
- **Training uses whole datasets, old batches included.** Fine-tuning on the newest batch alone
tends to raise the score on new footage while quietly losing the old.
- **Datasets are snapshots.** Labels are copied from what the dataset holds on disk, not
re-derived from today's rules, so two runs over the same dataset cannot disagree.
- **An empty Base column is honest.** It means the previous model's classes did not match this
dataset's, so scoring it would have compared two different things. The message says which.
`Use as base model` promotes a version, and the next round fine-tunes from it.
## Where things live
---
## 🗂️ Where things live
```
data/
app.db # projects, batches, frames, annotations, jobs
app.db # 14 tables: projects, frames, annotations,
# datasets, jobs, count_runs, video_clock …
archive/<date>/batchNNN.mp4 # recordings (read-only to the app)
batchNNN.json # sidecar: true start time from the server
recorder.log # the 24/7 recorder's output
live-count/session-*.jsonl # per-track traces from the live counter
projects/<slug>/
base/model.pt # the base model
dataset/ # master dataset, accumulating
images/{train,val}/ labels/{train,val}/ data.yaml
datasets/<id>/ # one folder per named dataset
images/{train,val}/ labels/{train,val}/
batches/<id>/frames/ # extracted frames
models/<n>/best.pt + metrics.json # each training run
```
The database holds status; the disk holds pixels, labels and weights. The master dataset
trains as-is with Ultralytics, or imports into Roboflow, without this application.
The database holds status; the disk holds pixels, labels and weights. A dataset trains as-is
with Ultralytics, or imports into Roboflow, without this application.
Your video archive is mounted read-only. Nothing is ever written back into it.
> [!WARNING]
> Your video archive is mounted **read-only** and nothing is ever written back into it.
## When something goes wrong
---
## ⚙️ Jobs
Heavy work runs as queued jobs, one at a time, so the GPU is never double-booked.
| Type | GPU | What it does |
|---|---|---|
| `extract` | — | ffmpeg pulls frames out of a range |
| `autolabel` | ✅ | SAM3 over a batch, one `set_image` per image |
| `merge` | — | copies approved frames into a dataset under frozen rules |
| `train` | ✅ | fine-tunes, then validates base and new on the same val set |
| `count` | ✅ | headless recount of archive videos for the accuracy table |
| `clock-scan` | — | reads each recording's real start time |
| `truck-scan` | ✅ | checks every recording actually contains a truck |
Jobs and their progress are persistent; after a restart the list is still there.
---
## 🔧 When something goes wrong
<details>
<summary><b>Deployment and GPU</b></summary>
<br>
| Symptom | Cause / fix |
|---|---|
| API returns 404 for routes you just added | the image copies `backend/` at build time — `docker compose build backend` again |
| A UI change doesn't show up | same trap on the other side: `docker compose build frontend`, then hard-reload |
| Job fails at *loading model* with a 401 | access to `facebook/sam3` not granted yet, or `HF_TOKEN` missing |
| SAM3 download crawls at a few KB/s | HuggingFace's Xet transfer throttling itself; `HF_HUB_DISABLE_XET=1` is already set in compose for that reason |
| `could not select device driver` | the NVIDIA container toolkit is not installed, or this Docker is older than the CDI support compose relies on (`devices: nvidia.com/gpu=all`) |
| `CUDA out of memory` while training | lower **Epochs**/batch on the Models page, or free the card — SAM3 is released before training, but another process may still hold it |
| Videos listed as *unreadable* | ffprobe could not parse them; they are still listed rather than hidden, so the archive never looks emptier than it is |
| `could not select device driver` | the NVIDIA container toolkit is not installed, or Docker is older than the CDI support compose relies on (`devices: nvidia.com/gpu=all`) |
| `CUDA out of memory` while training | lower epochs/batch on the Models page, or free the card — SAM3 is released before training, but another process may still hold it |
| A job reads *interrupted by a server restart* | it was running when the process died — jobs are not resumable, start it again |
## Development
</details>
```bash
uv pip install -r requirements.txt # backend
uv pip install -e sam3/
cd frontend && npm install # frontend
npm run dev # :5173, proxies /api to :8000
```
<details>
<summary><b>SAM3 and auto-labelling</b></summary>
The frontend pins Vite 7 on purpose — Vite 8's Rolldown binding crashes on this machine.
<br>
Working rules for agents and the documents that drive this repo are in `AGENTS.md` and
`docs/` (`requirements.md` → `design.md` → `tasks.md`).
| Symptom | Cause / fix |
|---|---|
| Job fails at *loading model* with a 401 | access to `facebook/sam3` not granted yet, or `HF_TOKEN` missing |
| SAM3 download crawls at a few KB/s | HuggingFace's Xet transfer throttling itself; `HF_HUB_DISABLE_XET=1` is already set in compose for that reason |
</details>
<details>
<summary><b>Archive and counting</b></summary>
<br>
| Symptom | Cause / fix |
|---|---|
| Videos listed as *unreadable* | ffprobe could not parse them; they are still listed rather than hidden, so the archive never looks emptier than it is |
| A recording's time shows amber | its overlay was read with low confidence, or not at all — type the time you can see in the video; a hand-entered time is never overwritten by a rescan |
| A cycle looks short | check whether its recordings moved to the neighbouring cycle — anything before 06:00 belongs to the previous shift |
| Live counting says *GPU busy* | a training or auto-label job holds the card; it waits 30 s before giving up |
</details>
---
## 🤝 Contributing
The documents drive the repo, not the other way round:
| Document | Contents |
|---|---|
| [`docs/requirements.md`](docs/requirements.md) | numbered `REQ-xxx`, changed only with the owner's approval |
| [`docs/design.md`](docs/design.md) | schema, API contract, disk layout — each section names the `REQ-xxx` it serves |
| [`docs/tasks.md`](docs/tasks.md) | implementation steps and how each was *verified*, `[TODO]` / `[DONE]` |
| [`AGENTS.md`](AGENTS.md) | working rules: simplicity, surgical changes, verify by running something |
Two invariants are easy to break and make the whole system lie:
1. **A frame in `val` stays in `val`** — otherwise the base-vs-new comparison is meaningless.
2. **One `set_image` per image** — `set_text_prompt()` re-runs only the grounding head against
the cached backbone output. An N-prompt job calls `set_image` once and loops prompts over
that state.