2026-08-05 11:15:05 +07:00
2026-08-05 11:15:05 +07:00

Dataset Enrichment

Take a model you already have, and make it better with footage you already have.

One round of the loop:

pick a project → browse the video archive → pick a batch → trim a range → extract frames → auto-annotate with SAM3 → review and correct every frame → approve → merge into the master dataset → fine-tune from the base model → compare against the base.

Nothing here is specific to one dataset. A project owns its base model, its class list, its video archive and its own accumulating dataset, so a second use case is a second project — not a second copy of the code.

Run it

cp .env.example .env          # paste HF_TOKEN
VIDEO_ARCHIVE_HOST=/path/to/videos docker compose up -d --build

Then open http://localhost:8080. The API is on :8000 if you want to poke at it directly; curl localhost:8000/api/health reports the GPU, free VRAM (vram_free_gb), SAM3 readiness (sam3_ready), whether ffmpeg is present, whether the token was picked up, and whether the database is reachable.

The first auto-annotation job downloads the ~3.4 GB SAM3 checkpoint into a Docker volume. It happens once; later jobs take about 12 seconds to load the model into VRAM.

What you need first

Thing Why
A GPU with the NVIDIA container toolkit SAM3 is CUDA-only
A video archive laid out as <date>/<batch>.mp4 that structure is what the Library reads
HF_TOKEN with access to facebook/sam3 the weights are gated, and approval is manual
A base model .pt (optional) without one, training starts from yolo11n.pt and you type the classes yourself

Example archive:

videos/
  2026-07-08/
    batch-1.mp4
    batch-4.mp4
  2026-07-09/
    batch-2.mp4

The screens

  1. Projects — name, label type (bbox or polygon), archive root, classes. Upload a base model and its classes are read from the checkpoint and locked, so the dataset and the model can never drift apart.
  2. Library — dates on the left, that date's recordings on the right with duration, resolution and how many batches already came out of each.
  3. Trim — play the video, set in/out, pick a frame rate. It tells you how many frames that produces before you commit to it.
  4. Review — the frame with its shapes on top. Drag to add a box, drag a corner to resize, drag the middle to move, Del to remove. Hold S and drag for a SAM3-assisted shape. A approves, X rejects, ←/→ move, U jumps to the next unreviewed frame, 1–9 pick the class. A batch can only be approved once no frame is still pending.
  5. Dataset & models — what the master dataset holds, a training run, and the base-versus-new table.

Reading the comparison

After training, the base model and the new one are validated on the same val set, and mAP50 / mAP50-95 / precision / recall appear side by side with the difference.

Two things make that number trustworthy:

  • The val split is stable. Once a frame is in val, it stays there for every later merge. A rising score cannot be an easier val set.
  • Training uses the whole master dataset, old batches included. Fine-tuning on the newest batch alone tends to raise the score on new footage while quietly losing the old.

If the Base column is empty, the previous model's classes did not match this dataset's, so scoring it here would have compared two different things. The message says which case it was.

Use as base model promotes a version, and the next round fine-tunes from it.

Where things live

data/
  app.db                                  # projects, batches, frames, annotations, jobs
  projects/<slug>/
    base/model.pt                         # the base model
    dataset/                              # master dataset, accumulating
      images/{train,val}/  labels/{train,val}/  data.yaml
    batches/<id>/frames/                  # extracted frames
    models/<n>/best.pt + metrics.json     # each training run

The database holds status; the disk holds pixels, labels and weights. The master dataset trains as-is with Ultralytics, or imports into Roboflow, without this application.

Your video archive is mounted read-only. Nothing is ever written back into it.

When something goes wrong

Symptom Cause / fix
API returns 404 for routes you just added the image copies backend/ at build time — docker compose build backend again
A UI change doesn't show up same trap on the other side: docker compose build frontend, then hard-reload
Job fails at loading model with a 401 access to facebook/sam3 not granted yet, or HF_TOKEN missing
SAM3 download crawls at a few KB/s HuggingFace's Xet transfer throttling itself; HF_HUB_DISABLE_XET=1 is already set in compose for that reason
could not select device driver the NVIDIA container toolkit is not installed, or this Docker is older than the CDI support compose relies on (devices: nvidia.com/gpu=all)
CUDA out of memory while training lower Epochs/batch on the Models page, or free the card — SAM3 is released before training, but another process may still hold it
Videos listed as unreadable ffprobe could not parse them; they are still listed rather than hidden, so the archive never looks emptier than it is
A job reads interrupted by a server restart it was running when the process died — jobs are not resumable, start it again

Development

uv pip install -r requirements.txt        # backend
uv pip install -e sam3/
cd frontend && npm install                # frontend
npm run dev                               # :5173, proxies /api to :8000

The frontend pins Vite 7 on purpose — Vite 8's Rolldown binding crashes on this machine.

Working rules for agents and the documents that drive this repo are in AGENTS.md and docs/ (requirements.md → design.md → tasks.md).

S
Description
used for retraining and annotation of karung feedmill project
Readme MIT
30 MiB
0 Stars 1 Watchers 0 Forks
Languages
Python 65.3%
JavaScript 31.7%
CSS 2.3%
Shell 0.4%
Dockerfile 0.2%