126 lines
5.9 KiB
Markdown
126 lines
5.9 KiB
Markdown
# Dataset Enrichment
|
||
|
||
Take a model you already have, and make it better with footage you already have.
|
||
|
||
One round of the loop:
|
||
|
||
> pick a project → browse the video archive → pick a batch → trim a range →
|
||
> extract frames → auto-annotate with SAM3 → review and correct every frame → approve →
|
||
> merge into the master dataset → fine-tune from the base model → compare against the base.
|
||
|
||
Nothing here is specific to one dataset. A project owns its base model, its class list, its
|
||
video archive and its own accumulating dataset, so a second use case is a second project —
|
||
not a second copy of the code.
|
||
|
||
## Run it
|
||
|
||
```bash
|
||
cp .env.example .env # paste HF_TOKEN
|
||
VIDEO_ARCHIVE_HOST=/path/to/videos docker compose up -d --build
|
||
```
|
||
|
||
Then open <http://localhost:8080>. The API is on `:8000` if you want to poke at it directly;
|
||
`curl localhost:8000/api/health` reports the GPU, free VRAM (`vram_free_gb`), SAM3 readiness (`sam3_ready`), whether ffmpeg is present, whether the token was picked up, and whether the database is reachable.
|
||
|
||
|
||
The first auto-annotation job downloads the ~3.4 GB SAM3 checkpoint into a Docker volume.
|
||
It happens once; later jobs take about 12 seconds to load the model into VRAM.
|
||
|
||
### What you need first
|
||
|
||
| Thing | Why |
|
||
|---|---|
|
||
| A GPU with the NVIDIA container toolkit | SAM3 is CUDA-only |
|
||
| A video archive laid out as `<date>/<batch>.mp4` | that structure is what the Library reads |
|
||
| `HF_TOKEN` with access to [facebook/sam3](https://huggingface.co/facebook/sam3) | the weights are gated, and approval is manual |
|
||
| A base model `.pt` (optional) | without one, training starts from `yolo11n.pt` and you type the classes yourself |
|
||
|
||
Example archive:
|
||
|
||
```
|
||
videos/
|
||
2026-07-08/
|
||
batch-1.mp4
|
||
batch-4.mp4
|
||
2026-07-09/
|
||
batch-2.mp4
|
||
```
|
||
|
||
## The screens
|
||
|
||
1. **Projects** — name, label type (`bbox` or `polygon`), archive root, classes. Upload a
|
||
base model and its classes are read from the checkpoint and locked, so the dataset and
|
||
the model can never drift apart.
|
||
2. **Library** — dates on the left, that date's recordings on the right with duration,
|
||
resolution and how many batches already came out of each.
|
||
3. **Trim** — play the video, set in/out, pick a frame rate. It tells you how many frames
|
||
that produces before you commit to it.
|
||
4. **Review** — the frame with its shapes on top. Drag to add a box, drag a corner to
|
||
resize, drag the middle to move, `Del` to remove. Hold `S` and drag for a SAM3-assisted
|
||
shape. `A` approves, `X` rejects, `←`/`→` move, `U` jumps to the next unreviewed frame,
|
||
`1`–`9` pick the class. A batch can only be approved once no frame is still pending.
|
||
5. **Dataset & models** — what the master dataset holds, a training run, and the
|
||
base-versus-new table.
|
||
|
||
## Reading the comparison
|
||
|
||
After training, the base model and the new one are validated **on the same val set**, and
|
||
mAP50 / mAP50-95 / precision / recall appear side by side with the difference.
|
||
|
||
Two things make that number trustworthy:
|
||
|
||
- **The val split is stable.** Once a frame is in `val`, it stays there for every later
|
||
merge. A rising score cannot be an easier val set.
|
||
- **Training uses the whole master dataset**, old batches included. Fine-tuning on the newest
|
||
batch alone tends to raise the score on new footage while quietly losing the old.
|
||
|
||
If the Base column is empty, the previous model's classes did not match this dataset's, so
|
||
scoring it here would have compared two different things. The message says which case it was.
|
||
|
||
`Use as base model` promotes a version, and the next round fine-tunes from it.
|
||
|
||
## Where things live
|
||
|
||
```
|
||
data/
|
||
app.db # projects, batches, frames, annotations, jobs
|
||
projects/<slug>/
|
||
base/model.pt # the base model
|
||
dataset/ # master dataset, accumulating
|
||
images/{train,val}/ labels/{train,val}/ data.yaml
|
||
batches/<id>/frames/ # extracted frames
|
||
models/<n>/best.pt + metrics.json # each training run
|
||
```
|
||
|
||
The database holds status; the disk holds pixels, labels and weights. The master dataset
|
||
trains as-is with Ultralytics, or imports into Roboflow, without this application.
|
||
|
||
Your video archive is mounted read-only. Nothing is ever written back into it.
|
||
|
||
## When something goes wrong
|
||
|
||
| Symptom | Cause / fix |
|
||
|---|---|
|
||
| API returns 404 for routes you just added | the image copies `backend/` at build time — `docker compose build backend` again |
|
||
| A UI change doesn't show up | same trap on the other side: `docker compose build frontend`, then hard-reload |
|
||
| Job fails at *loading model* with a 401 | access to `facebook/sam3` not granted yet, or `HF_TOKEN` missing |
|
||
| SAM3 download crawls at a few KB/s | HuggingFace's Xet transfer throttling itself; `HF_HUB_DISABLE_XET=1` is already set in compose for that reason |
|
||
| `could not select device driver` | the NVIDIA container toolkit is not installed, or this Docker is older than the CDI support compose relies on (`devices: nvidia.com/gpu=all`) |
|
||
| `CUDA out of memory` while training | lower **Epochs**/batch on the Models page, or free the card — SAM3 is released before training, but another process may still hold it |
|
||
| Videos listed as *unreadable* | ffprobe could not parse them; they are still listed rather than hidden, so the archive never looks emptier than it is |
|
||
| A job reads *interrupted by a server restart* | it was running when the process died — jobs are not resumable, start it again |
|
||
|
||
## Development
|
||
|
||
```bash
|
||
uv pip install -r requirements.txt # backend
|
||
uv pip install -e sam3/
|
||
cd frontend && npm install # frontend
|
||
npm run dev # :5173, proxies /api to :8000
|
||
```
|
||
|
||
The frontend pins Vite 7 on purpose — Vite 8's Rolldown binding crashes on this machine.
|
||
|
||
Working rules for agents and the documents that drive this repo are in `AGENTS.md` and
|
||
`docs/` (`requirements.md` → `design.md` → `tasks.md`).
|