Dataset Enrichment - DEV
Take a model you already have, and make it better with footage you already have.
One round of the loop:
pick a project → browse the video archive → pick a batch → trim a range → extract frames → auto-annotate with SAM3 → review and correct every frame → approve → merge into the master dataset → fine-tune from the base model → compare against the base.
Nothing here is specific to one dataset. A project owns its base model, its class list, its video archive and its own accumulating dataset, so a second use case is a second project — not a second copy of the code.
Run it
cp .env.example .env # paste HF_TOKEN
VIDEO_ARCHIVE_HOST=/path/to/videos docker compose up -d --build
Then open http://localhost:8080. The API is on :8000 if you want to poke at it directly;
curl localhost:8000/api/health reports the GPU, free VRAM (vram_free_gb), SAM3 readiness (sam3_ready), whether ffmpeg is present, whether the token was picked up, and whether the database is reachable.
The first auto-annotation job downloads the ~3.4 GB SAM3 checkpoint into a Docker volume. It happens once; later jobs take about 12 seconds to load the model into VRAM.
What you need first
| Thing | Why |
|---|---|
| A GPU with the NVIDIA container toolkit | SAM3 is CUDA-only |
A video archive laid out as <date>/<batch>.mp4 |
that structure is what the Library reads |
HF_TOKEN with access to facebook/sam3 |
the weights are gated, and approval is manual |
A base model .pt (optional) |
without one, training starts from yolo11n.pt and you type the classes yourself |
Example archive:
videos/
2026-07-08/
batch-1.mp4
batch-4.mp4
2026-07-09/
batch-2.mp4
The screens
- Projects — name, label type (
bboxorpolygon), archive root, classes. Upload a base model and its classes are read from the checkpoint and locked, so the dataset and the model can never drift apart. - Library — dates on the left, that date's recordings on the right with duration, resolution and how many batches already came out of each.
- Trim — play the video, set in/out, pick a frame rate. It tells you how many frames that produces before you commit to it.
- Review — the frame with its shapes on top. Drag to add a box, drag a corner to
resize, drag the middle to move,
Delto remove. HoldSand drag for a SAM3-assisted shape.Aapproves,Xrejects,←/→move,Ujumps to the next unreviewed frame,1–9pick the class. A batch can only be approved once no frame is still pending. - Dataset & models — what the master dataset holds, a training run, and the base-versus-new table.
Reading the comparison
After training, the base model and the new one are validated on the same val set, and mAP50 / mAP50-95 / precision / recall appear side by side with the difference.
Two things make that number trustworthy:
- The val split is stable. Once a frame is in
val, it stays there for every later merge. A rising score cannot be an easier val set. - Training uses the whole master dataset, old batches included. Fine-tuning on the newest batch alone tends to raise the score on new footage while quietly losing the old.
If the Base column is empty, the previous model's classes did not match this dataset's, so scoring it here would have compared two different things. The message says which case it was.
Use as base model promotes a version, and the next round fine-tunes from it.
Where things live
data/
app.db # projects, batches, frames, annotations, jobs
projects/<slug>/
base/model.pt # the base model
dataset/ # master dataset, accumulating
images/{train,val}/ labels/{train,val}/ data.yaml
batches/<id>/frames/ # extracted frames
models/<n>/best.pt + metrics.json # each training run
The database holds status; the disk holds pixels, labels and weights. The master dataset trains as-is with Ultralytics, or imports into Roboflow, without this application.
Your video archive is mounted read-only. Nothing is ever written back into it.
When something goes wrong
| Symptom | Cause / fix |
|---|---|
| API returns 404 for routes you just added | the image copies backend/ at build time — docker compose build backend again |
| A UI change doesn't show up | same trap on the other side: docker compose build frontend, then hard-reload |
| Job fails at loading model with a 401 | access to facebook/sam3 not granted yet, or HF_TOKEN missing |
| SAM3 download crawls at a few KB/s | HuggingFace's Xet transfer throttling itself; HF_HUB_DISABLE_XET=1 is already set in compose for that reason |
could not select device driver |
the NVIDIA container toolkit is not installed, or this Docker is older than the CDI support compose relies on (devices: nvidia.com/gpu=all) |
CUDA out of memory while training |
lower Epochs/batch on the Models page, or free the card — SAM3 is released before training, but another process may still hold it |
| Videos listed as unreadable | ffprobe could not parse them; they are still listed rather than hidden, so the archive never looks emptier than it is |
| A job reads interrupted by a server restart | it was running when the process died — jobs are not resumable, start it again |
Development
uv pip install -r requirements.txt # backend
uv pip install -e sam3/
cd frontend && npm install # frontend
npm run dev # :5173, proxies /api to :8000
The frontend pins Vite 7 on purpose — Vite 8's Rolldown binding crashes on this machine.
Working rules for agents and the documents that drive this repo are in AGENTS.md and
docs/ (requirements.md → design.md → tasks.md).