first commit

This commit is contained in:
asus committed 2026-08-05 11:15:05 +07:00
commit b5c28cc98a
1 file changed
+125
+125
View File
@@ -0,0 +1,125 @@
# Dataset Enrichment
Take a model you already have, and make it better with footage you already have.
One round of the loop:
> pick a project → browse the video archive → pick a batch → trim a range →
> extract frames → auto-annotate with SAM3 → review and correct every frame → approve →
> merge into the master dataset → fine-tune from the base model → compare against the base.
Nothing here is specific to one dataset. A project owns its base model, its class list, its
video archive and its own accumulating dataset, so a second use case is a second project —
not a second copy of the code.
## Run it
```bash
cp .env.example .env # paste HF_TOKEN
VIDEO_ARCHIVE_HOST=/path/to/videos docker compose up -d --build
```
Then open <http://localhost:8080>. The API is on `:8000` if you want to poke at it directly;
`curl localhost:8000/api/health` reports the GPU, free VRAM (`vram_free_gb`), SAM3 readiness (`sam3_ready`), whether ffmpeg is present, whether the token was picked up, and whether the database is reachable.
The first auto-annotation job downloads the ~3.4 GB SAM3 checkpoint into a Docker volume.
It happens once; later jobs take about 12 seconds to load the model into VRAM.
### What you need first
| Thing | Why |
|---|---|
| A GPU with the NVIDIA container toolkit | SAM3 is CUDA-only |
| A video archive laid out as `<date>/<batch>.mp4` | that structure is what the Library reads |
| `HF_TOKEN` with access to [facebook/sam3](https://huggingface.co/facebook/sam3) | the weights are gated, and approval is manual |
| A base model `.pt` (optional) | without one, training starts from `yolo11n.pt` and you type the classes yourself |
Example archive:
```
videos/
2026-07-08/
batch-1.mp4
batch-4.mp4
2026-07-09/
batch-2.mp4
```
## The screens
1. **Projects** — name, label type (`bbox` or `polygon`), archive root, classes. Upload a
base model and its classes are read from the checkpoint and locked, so the dataset and
the model can never drift apart.
2. **Library** — dates on the left, that date's recordings on the right with duration,
resolution and how many batches already came out of each.
3. **Trim** — play the video, set in/out, pick a frame rate. It tells you how many frames
that produces before you commit to it.
4. **Review** — the frame with its shapes on top. Drag to add a box, drag a corner to
resize, drag the middle to move, `Del` to remove. Hold `S` and drag for a SAM3-assisted
shape. `A` approves, `X` rejects, `←`/`→` move, `U` jumps to the next unreviewed frame,
`1`–`9` pick the class. A batch can only be approved once no frame is still pending.
5. **Dataset & models** — what the master dataset holds, a training run, and the
base-versus-new table.
## Reading the comparison
After training, the base model and the new one are validated **on the same val set**, and
mAP50 / mAP50-95 / precision / recall appear side by side with the difference.
Two things make that number trustworthy:
- **The val split is stable.** Once a frame is in `val`, it stays there for every later
merge. A rising score cannot be an easier val set.
- **Training uses the whole master dataset**, old batches included. Fine-tuning on the newest
batch alone tends to raise the score on new footage while quietly losing the old.
If the Base column is empty, the previous model's classes did not match this dataset's, so
scoring it here would have compared two different things. The message says which case it was.
`Use as base model` promotes a version, and the next round fine-tunes from it.
## Where things live
```
data/
app.db # projects, batches, frames, annotations, jobs
projects/<slug>/
base/model.pt # the base model
dataset/ # master dataset, accumulating
images/{train,val}/ labels/{train,val}/ data.yaml
batches/<id>/frames/ # extracted frames
models/<n>/best.pt + metrics.json # each training run
```
The database holds status; the disk holds pixels, labels and weights. The master dataset
trains as-is with Ultralytics, or imports into Roboflow, without this application.
Your video archive is mounted read-only. Nothing is ever written back into it.
## When something goes wrong
| Symptom | Cause / fix |
|---|---|
| API returns 404 for routes you just added | the image copies `backend/` at build time — `docker compose build backend` again |
| A UI change doesn't show up | same trap on the other side: `docker compose build frontend`, then hard-reload |
| Job fails at *loading model* with a 401 | access to `facebook/sam3` not granted yet, or `HF_TOKEN` missing |
| SAM3 download crawls at a few KB/s | HuggingFace's Xet transfer throttling itself; `HF_HUB_DISABLE_XET=1` is already set in compose for that reason |
| `could not select device driver` | the NVIDIA container toolkit is not installed, or this Docker is older than the CDI support compose relies on (`devices: nvidia.com/gpu=all`) |
| `CUDA out of memory` while training | lower **Epochs**/batch on the Models page, or free the card — SAM3 is released before training, but another process may still hold it |
| Videos listed as *unreadable* | ffprobe could not parse them; they are still listed rather than hidden, so the archive never looks emptier than it is |
| A job reads *interrupted by a server restart* | it was running when the process died — jobs are not resumable, start it again |
## Development
```bash
uv pip install -r requirements.txt # backend
uv pip install -e sam3/
cd frontend && npm install # frontend
npm run dev # :5173, proxies /api to :8000
```
The frontend pins Vite 7 on purpose — Vite 8's Rolldown binding crashes on this machine.
Working rules for agents and the documents that drive this repo are in `AGENTS.md` and
`docs/` (`requirements.md` → `design.md` → `tasks.md`).