From b6624eeff99a7ff4496aa3717e985890b29145c2 Mon Sep 17 00:00:00 2001 From: asus Date: Fri, 14 Aug 2026 16:37:27 +0700 Subject: [PATCH] docs: update README and add .env.example --- .env.example | 34 ++++ README.md | 459 +++++++++++++++++++++++++++++++++++++++++++-------- 2 files changed, 423 insertions(+), 70 deletions(-) create mode 100644 .env.example diff --git a/.env.example b/.env.example new file mode 100644 index 0000000..76ba1e5 --- /dev/null +++ b/.env.example @@ -0,0 +1,34 @@ +# Copy to .env and fill in. Only HF_TOKEN is required to get started. +# +# cp .env.example .env + +# ---- required --------------------------------------------------------------- + +# HuggingFace token with access to facebook/sam3. The weights are gated and +# approval is manual, so request access before the first auto-annotation job: +# https://huggingface.co/facebook/sam3 +HF_TOKEN= + +# ---- host paths ------------------------------------------------------------- + +# Where your recordings live, laid out as /.mp4. Mounted read-only +# into the container; nothing is ever written back into it. +VIDEO_ARCHIVE_HOST=./data/archive + +# ---- ports ------------------------------------------------------------------ + +# The UI. The API is always on :8000. +WEB_PORT=8080 + +# Origins allowed to call the API. Add your machine's LAN address to reach the +# dev server from another device. +CORS_ORIGINS=http://localhost:5173,http://localhost:8080 + +# ---- recorder (algoritma-batch/batch_video_cropper.py) ----------------------- +# Only needed if you run the 24/7 truck-session recorder. It reads the RTSP +# stream to detect sessions, then downloads each session from the recording +# server as a copy rather than re-encoding it. + +# MediaMTX playback endpoint and the path name to pull from. +PLAYBACK_URL=http://192.168.192.96:9996/get +PLAYBACK_PATH=cam diff --git a/README.md b/README.md index e047c94..a063f14 100644 --- a/README.md +++ b/README.md @@ -1,125 +1,444 @@ -# Dataset Enrichment - DEV +
-Take a model you already have, and make it better with footage you already have. +# reTraining -One round of the loop: +**Take a model you already have, and make it better with footage you already have.** -> pick a project → browse the video archive → pick a batch → trim a range → -> extract frames → auto-annotate with SAM3 → review and correct every frame → approve → -> merge into the master dataset → fine-tune from the base model → compare against the base. +A self-hosted loop for turning raw CCTV into a measurably better detector — record, extract, +auto-label, review, filter, train, and prove the new model actually beat the old one. -Nothing here is specific to one dataset. A project owns its base model, its class list, its -video archive and its own accumulating dataset, so a second use case is a second project — -not a second copy of the code. +

+Python +FastAPI +React +Vite +SAM3 +YOLO +Docker +GPU +

-## Run it +[Quick start](#-quick-start) · [How it works](#-how-it-works) · [The screens](#-the-screens) · +[Counting](#-counting) · [Trust the numbers](#-why-the-numbers-are-trustworthy) · +[Troubleshooting](#-when-something-goes-wrong) + +
+ +--- + +## ⚡ Quick start ```bash -cp .env.example .env # paste HF_TOKEN +cp .env.example .env # paste your HF_TOKEN VIDEO_ARCHIVE_HOST=/path/to/videos docker compose up -d --build ``` -Then open . The API is on `:8000` if you want to poke at it directly; -`curl localhost:8000/api/health` reports the GPU, free VRAM (`vram_free_gb`), SAM3 readiness (`sam3_ready`), whether ffmpeg is present, whether the token was picked up, and whether the database is reachable. +Open ****. The API is on `:8000`. +```bash +curl localhost:8000/api/health +``` +```json +{ "device": "cuda", "gpu": "NVIDIA GeForce RTX 5080 Laptop GPU", + "vram_free_gb": 14.91, "sam3_ready": true, "ffmpeg": true, + "hf_token": true, "db": true } +``` -The first auto-annotation job downloads the ~3.4 GB SAM3 checkpoint into a Docker volume. -It happens once; later jobs take about 12 seconds to load the model into VRAM. +> [!IMPORTANT] +> The first auto-annotation job downloads the **~3.4 GB SAM3 checkpoint** into a Docker +> volume. It happens once; later jobs take about 12 seconds to load the model into VRAM. -### What you need first +
+What you need first + +
| Thing | Why | |---|---| | A GPU with the NVIDIA container toolkit | SAM3 is CUDA-only | -| A video archive laid out as `/.mp4` | that structure is what the Library reads | +| A video archive laid out as `/.mp4` | that structure is what the archive browser reads | | `HF_TOKEN` with access to [facebook/sam3](https://huggingface.co/facebook/sam3) | the weights are gated, and approval is manual | | A base model `.pt` (optional) | without one, training starts from `yolo11n.pt` and you type the classes yourself | -Example archive: - ``` videos/ - 2026-07-08/ - batch-1.mp4 - batch-4.mp4 - 2026-07-09/ - batch-2.mp4 + 2026-08-13/ + batch001.mp4 + batch002.mp4 + 2026-08-14/ + batch001.mp4 ``` -## The screens +
-1. **Projects** — name, label type (`bbox` or `polygon`), archive root, classes. Upload a - base model and its classes are read from the checkpoint and locked, so the dataset and - the model can never drift apart. -2. **Library** — dates on the left, that date's recordings on the right with duration, - resolution and how many batches already came out of each. -3. **Trim** — play the video, set in/out, pick a frame rate. It tells you how many frames - that produces before you commit to it. -4. **Review** — the frame with its shapes on top. Drag to add a box, drag a corner to - resize, drag the middle to move, `Del` to remove. Hold `S` and drag for a SAM3-assisted - shape. `A` approves, `X` rejects, `←`/`→` move, `U` jumps to the next unreviewed frame, - `1`–`9` pick the class. A batch can only be approved once no frame is still pending. -5. **Dataset & models** — what the master dataset holds, a training run, and the - base-versus-new table. +
+Run it without Docker (development) -## Reading the comparison +
+ +```bash +uv pip install -r requirements.txt # backend +uv pip install -e sam3/ +uv run uvicorn backend.main:app --reload # :8000 + +cd frontend && npm install +npm run dev # :5173, proxies /api to :8000 +``` + +The frontend pins Vite 7 on purpose — Vite 8's Rolldown binding crashes on this machine. +The package manager is `uv`; there is no `pip`/`poetry` path. + +
+ +--- + +## 🔄 How it works + +One project owns its base model, its class list, its video archive and its own accumulating +datasets. A second use case is a second project — not a second copy of the code. + +```mermaid +flowchart LR + A[📹 Archive
one file per truck session] --> B[✂️ Trim
pick a range + fps] + B --> C[🖼️ Extract
frames to disk] + C --> D[🤖 Auto-label
SAM3 text prompts] + D --> E[👁️ Review
fix every frame] + E --> F[🧹 Data Prep
filter + augment] + F --> G[📦 Dataset
named, immutable] + G --> H[🎯 Train
fine-tune from base] + H --> I[📊 Compare
base vs new, same val set] + I -.->|promote| H +``` + +> [!NOTE] +> **Data Prep is the gate.** Selecting batches does not create a dataset — it opens Data Prep +> scoped to that selection. Only *Confirm merge* cuts the dataset, and the filter rules in +> force at that moment are **frozen onto it**, so editing them later can never rewrite a +> dataset you already trained on. + +--- + +## 🖥️ The screens + +
+1. Projects — name, label type, archive root, classes + +
+ +Upload a base model and its classes are read from the checkpoint and locked, so the dataset +and the model can never drift apart. + +
+ +
+2. Video Archive — browse by cycle, not by folder + +
+ +A **cycle** is one shift: `06:00 → 05:59` the next morning. It always crosses midnight, so it +always spans two calendar dates, and is named after the date it starts on. + +Folder names are not when a recording was made, and neither are file mtimes — those are file +*copy* times. The real start comes from the timestamp the camera burns into every frame, so a +file sitting in the `2026-08-14` folder but recorded at `00:11` shows up as part of the +**13 Aug** cycle, numbered in the order it was actually made. + +``` +Siklus 13 Agt 2026 28 rekaman + #1 08:27:27 batch003 2026-08-13 + ... + #25 23:53:45 batch027 2026-08-13 + #26 00:11:02 batch001 2026-08-14 ← pulled in from the next folder + #27 00:34:01 batch002 2026-08-14 +``` + +Nothing in the archive is moved or renamed — it is mounted **read-only**. The grouping lives +in an index beside it, and the original path stays the file's identity. + +
+ +
+3. Trim — play the video, set in/out, pick a frame rate + +
+ +It tells you how many frames that produces before you commit to it. + +
+ +
+4. Review — the frame with its shapes on top + +
+ +| Key | Action | | Key | Action | +|---|---|---|---|---| +| `A` | approve | | `←` `→` | previous / next frame | +| `X` | reject | | `U` | jump to next unreviewed | +| `Del` | delete shape | | `1`–`9` | pick class | +| `S` + drag | SAM3-assisted shape | | drag | add / move / resize a box | + +Approving no longer merges. It marks frames approved; merging happens in Data Prep. + +
+ +
+5. Data Prep — throw out the junk, then set augmentation + +
+ +Tune an outlier filter over **score / area / aspect** against exactly the batches you picked, +watching the counts move as you drag. A dropped box leaves its image in the dataset; only a +frame that loses *every* box is held back — these frames hold ~44 objects each, and excluding +the whole image was measured to cost 96% of a batch to remove 10% of its boxes. + +Then **Confirm merge** creates the dataset under a frozen copy of those rules. + +
+ +
+6. Datasets — several per project, each a standalone copy + +
+ +`batch7+8 strict rules` and `batch7+8 after I fixed the annotations` are two datasets holding +the same frames with different labels. Combining them for a run is *newest wins*, so a frame +appearing twice is emitted once rather than teaching the model two contradictory labels. + +
+ +
+7. Models & Training — run, then read the comparison + +
+ +Pick datasets, pick classes, train. Batch size, image size and device default from the +hardware actually detected, so a bigger GPU changes the numbers in the form, not the code. + +
+ +
+8. Live Counting — point a model at a camera and watch it count + +
+ +Same tracker, stabiliser and line-cross counter the production script uses. Click the video to +place the counting line; it moves live without losing the counts. Every finished track is +written to a JSONL with the reason it did or did not count — which is what separates a model +miss from a tracker miss from a counter miss. + +
+ +
+9. Counting Accuracy — scored, per cycle + +
+ +One row per recording, grouped into collapsible cycles. Type in the ground truth you counted +by hand and the table shows the **signed delta** — `+3` and `-3` are different failures, and a +single accuracy percentage hides which one you have. Totals only ever count rows where a +ground truth is filled in. + +Recount runs headless in a background job — no annotated frame, no JPEG encode, which is worth +**146 fps vs 124** in the live view. + +
+ +--- + +## 🎥 Counting + +The counter is deliberately robust to low frame rates: it never needs to catch the exact frame +of a crossing, only that a track was seen *above* the line at some point in its life. + +```mermaid +stateDiagram-v2 + [*] --> UNKNOWN + UNKNOWN --> ABOVE: y1 above the band + UNKNOWN --> BELOW: born below (ghost — never counts) + ABOVE --> COUNTED: seen below + travelled far enough + COUNTED --> ABOVE: sustained frames above (real unload) +``` + +
+Four layers, and what each one is actually for + +
+ +| Layer | Guard | Stops | +|---|---|---| +| 1 | Must have been **above** the line at some point | a box that appears inside the truck | +| 2 | Must have travelled `entry_travel_min` from where it first appeared | ghost boxes that blink into existence next to the line | +| 3 | **Track hand-off** — a dying track parks its history for a newborn nearby to inherit | an ID switch at the line losing the count *or* duplicating it | +| 4 | One count per direction per track, and the verdict is the track's *last* direction | double counting, while still letting a genuine unload-and-reload count again | + +Every one of these was written against a reproduced failure. Hand-off replaced a spatial dedup +that did not dedup: a blocked track simply retried each frame and counted anyway once it +drifted out of the circle — late, at the wrong position, seeding the next circle in the wrong +place. + +`handoff_radius` is the dial that matters most. These frames hold ~44 objects, so a newborn +track is nearly always near one that just vanished; calibrate it against a clip with a +hand-counted total rather than by eye. + +
+ +
+Recording pipeline — record once, cut sessions afterwards + +
+ +```mermaid +flowchart LR + CAM[📷 Dahua 1080p
H.265 @ 25fps] --> MTX[MediaMTX on Jetson
records 24/7 · 24h buffer] + MTX -->|RTSP| DET[Truck detector
on the GPU box] + DET -->|session ends| FETCH[Download that exact
time range as a copy] + MTX --> FETCH + FETCH --> ARC[📁 archive/date/batchNNN.mp4
+ .json sidecar] +``` + +The detector does **not** encode video. When a truck session ends it downloads that range from +the recording server, so the archive keeps the camera's own codec, resolution and frame rate. + +| | Before | After | +|---|---|---| +| Codec | mpeg4 re-encode | HEVC copy | +| Resolution | 1280×720 | 1920×1080 | +| Size | 8.0 Mbps | **1.72 Mbps** (4.7× smaller) | +| Timebase | declared 10 fps at 25 fps real → **2.49× slow** | true 25 fps | +| Start time | read from the burned-in overlay by OCR | from the server, exact | + +Each clip carries a `.json` sidecar with the server's start time, which the app trusts over +reading the overlay — so a new session appears in the right cycle with no scan at all. + +
+ +--- + +## 📊 Why the numbers are trustworthy After training, the base model and the new one are validated **on the same val set**, and mAP50 / mAP50-95 / precision / recall appear side by side with the difference. -Two things make that number trustworthy: +> [!TIP] +> **The val split is stable.** Once a frame is in `val` it stays there for every later merge — +> derived from the frame's identity, not from how many rows precede it. A rising score cannot +> be an easier val set. -- **The val split is stable.** Once a frame is in `val`, it stays there for every later - merge. A rising score cannot be an easier val set. -- **Training uses the whole master dataset**, old batches included. Fine-tuning on the newest - batch alone tends to raise the score on new footage while quietly losing the old. - -If the Base column is empty, the previous model's classes did not match this dataset's, so -scoring it here would have compared two different things. The message says which case it was. +- **Training uses whole datasets, old batches included.** Fine-tuning on the newest batch alone + tends to raise the score on new footage while quietly losing the old. +- **Datasets are snapshots.** Labels are copied from what the dataset holds on disk, not + re-derived from today's rules, so two runs over the same dataset cannot disagree. +- **An empty Base column is honest.** It means the previous model's classes did not match this + dataset's, so scoring it would have compared two different things. The message says which. `Use as base model` promotes a version, and the next round fine-tunes from it. -## Where things live +--- + +## 🗂️ Where things live ``` data/ - app.db # projects, batches, frames, annotations, jobs + app.db # 14 tables: projects, frames, annotations, + # datasets, jobs, count_runs, video_clock … + archive//batchNNN.mp4 # recordings (read-only to the app) + batchNNN.json # sidecar: true start time from the server + recorder.log # the 24/7 recorder's output + live-count/session-*.jsonl # per-track traces from the live counter projects// base/model.pt # the base model - dataset/ # master dataset, accumulating - images/{train,val}/ labels/{train,val}/ data.yaml + datasets// # one folder per named dataset + images/{train,val}/ labels/{train,val}/ batches//frames/ # extracted frames models//best.pt + metrics.json # each training run ``` -The database holds status; the disk holds pixels, labels and weights. The master dataset -trains as-is with Ultralytics, or imports into Roboflow, without this application. +The database holds status; the disk holds pixels, labels and weights. A dataset trains as-is +with Ultralytics, or imports into Roboflow, without this application. -Your video archive is mounted read-only. Nothing is ever written back into it. +> [!WARNING] +> Your video archive is mounted **read-only** and nothing is ever written back into it. -## When something goes wrong +--- + +## ⚙️ Jobs + +Heavy work runs as queued jobs, one at a time, so the GPU is never double-booked. + +| Type | GPU | What it does | +|---|---|---| +| `extract` | — | ffmpeg pulls frames out of a range | +| `autolabel` | ✅ | SAM3 over a batch, one `set_image` per image | +| `merge` | — | copies approved frames into a dataset under frozen rules | +| `train` | ✅ | fine-tunes, then validates base and new on the same val set | +| `count` | ✅ | headless recount of archive videos for the accuracy table | +| `clock-scan` | — | reads each recording's real start time | +| `truck-scan` | ✅ | checks every recording actually contains a truck | + +Jobs and their progress are persistent; after a restart the list is still there. + +--- + +## 🔧 When something goes wrong + +
+Deployment and GPU + +
| Symptom | Cause / fix | |---|---| | API returns 404 for routes you just added | the image copies `backend/` at build time — `docker compose build backend` again | | A UI change doesn't show up | same trap on the other side: `docker compose build frontend`, then hard-reload | -| Job fails at *loading model* with a 401 | access to `facebook/sam3` not granted yet, or `HF_TOKEN` missing | -| SAM3 download crawls at a few KB/s | HuggingFace's Xet transfer throttling itself; `HF_HUB_DISABLE_XET=1` is already set in compose for that reason | -| `could not select device driver` | the NVIDIA container toolkit is not installed, or this Docker is older than the CDI support compose relies on (`devices: nvidia.com/gpu=all`) | -| `CUDA out of memory` while training | lower **Epochs**/batch on the Models page, or free the card — SAM3 is released before training, but another process may still hold it | -| Videos listed as *unreadable* | ffprobe could not parse them; they are still listed rather than hidden, so the archive never looks emptier than it is | +| `could not select device driver` | the NVIDIA container toolkit is not installed, or Docker is older than the CDI support compose relies on (`devices: nvidia.com/gpu=all`) | +| `CUDA out of memory` while training | lower epochs/batch on the Models page, or free the card — SAM3 is released before training, but another process may still hold it | | A job reads *interrupted by a server restart* | it was running when the process died — jobs are not resumable, start it again | -## Development +
-```bash -uv pip install -r requirements.txt # backend -uv pip install -e sam3/ -cd frontend && npm install # frontend -npm run dev # :5173, proxies /api to :8000 -``` +
+SAM3 and auto-labelling -The frontend pins Vite 7 on purpose — Vite 8's Rolldown binding crashes on this machine. +
-Working rules for agents and the documents that drive this repo are in `AGENTS.md` and -`docs/` (`requirements.md` → `design.md` → `tasks.md`). +| Symptom | Cause / fix | +|---|---| +| Job fails at *loading model* with a 401 | access to `facebook/sam3` not granted yet, or `HF_TOKEN` missing | +| SAM3 download crawls at a few KB/s | HuggingFace's Xet transfer throttling itself; `HF_HUB_DISABLE_XET=1` is already set in compose for that reason | + +
+ +
+Archive and counting + +
+ +| Symptom | Cause / fix | +|---|---| +| Videos listed as *unreadable* | ffprobe could not parse them; they are still listed rather than hidden, so the archive never looks emptier than it is | +| A recording's time shows amber | its overlay was read with low confidence, or not at all — type the time you can see in the video; a hand-entered time is never overwritten by a rescan | +| A cycle looks short | check whether its recordings moved to the neighbouring cycle — anything before 06:00 belongs to the previous shift | +| Live counting says *GPU busy* | a training or auto-label job holds the card; it waits 30 s before giving up | + +
+ +--- + +## 🤝 Contributing + +The documents drive the repo, not the other way round: + +| Document | Contents | +|---|---| +| [`docs/requirements.md`](docs/requirements.md) | numbered `REQ-xxx`, changed only with the owner's approval | +| [`docs/design.md`](docs/design.md) | schema, API contract, disk layout — each section names the `REQ-xxx` it serves | +| [`docs/tasks.md`](docs/tasks.md) | implementation steps and how each was *verified*, `[TODO]` / `[DONE]` | +| [`AGENTS.md`](AGENTS.md) | working rules: simplicity, surgical changes, verify by running something | + +Two invariants are easy to break and make the whole system lie: + +1. **A frame in `val` stays in `val`** — otherwise the base-vs-new comparison is meaningless. +2. **One `set_image` per image** — `set_text_prompt()` re-runs only the grounding head against + the cached backbone output. An N-prompt job calls `set_image` once and loops prompts over + that state.