From b5c28cc98afb42a0ca256667d32af04137453015 Mon Sep 17 00:00:00 2001 From: asus Date: Wed, 5 Aug 2026 11:15:05 +0700 Subject: [PATCH] first commit --- README.md | 125 ++++++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 125 insertions(+) create mode 100644 README.md diff --git a/README.md b/README.md new file mode 100644 index 0000000..4e05eaa --- /dev/null +++ b/README.md @@ -0,0 +1,125 @@ +# Dataset Enrichment + +Take a model you already have, and make it better with footage you already have. + +One round of the loop: + +> pick a project → browse the video archive → pick a batch → trim a range → +> extract frames → auto-annotate with SAM3 → review and correct every frame → approve → +> merge into the master dataset → fine-tune from the base model → compare against the base. + +Nothing here is specific to one dataset. A project owns its base model, its class list, its +video archive and its own accumulating dataset, so a second use case is a second project — +not a second copy of the code. + +## Run it + +```bash +cp .env.example .env # paste HF_TOKEN +VIDEO_ARCHIVE_HOST=/path/to/videos docker compose up -d --build +``` + +Then open . The API is on `:8000` if you want to poke at it directly; +`curl localhost:8000/api/health` reports the GPU, free VRAM (`vram_free_gb`), SAM3 readiness (`sam3_ready`), whether ffmpeg is present, whether the token was picked up, and whether the database is reachable. + + +The first auto-annotation job downloads the ~3.4 GB SAM3 checkpoint into a Docker volume. +It happens once; later jobs take about 12 seconds to load the model into VRAM. + +### What you need first + +| Thing | Why | +|---|---| +| A GPU with the NVIDIA container toolkit | SAM3 is CUDA-only | +| A video archive laid out as `/.mp4` | that structure is what the Library reads | +| `HF_TOKEN` with access to [facebook/sam3](https://huggingface.co/facebook/sam3) | the weights are gated, and approval is manual | +| A base model `.pt` (optional) | without one, training starts from `yolo11n.pt` and you type the classes yourself | + +Example archive: + +``` +videos/ + 2026-07-08/ + batch-1.mp4 + batch-4.mp4 + 2026-07-09/ + batch-2.mp4 +``` + +## The screens + +1. **Projects** — name, label type (`bbox` or `polygon`), archive root, classes. Upload a + base model and its classes are read from the checkpoint and locked, so the dataset and + the model can never drift apart. +2. **Library** — dates on the left, that date's recordings on the right with duration, + resolution and how many batches already came out of each. +3. **Trim** — play the video, set in/out, pick a frame rate. It tells you how many frames + that produces before you commit to it. +4. **Review** — the frame with its shapes on top. Drag to add a box, drag a corner to + resize, drag the middle to move, `Del` to remove. Hold `S` and drag for a SAM3-assisted + shape. `A` approves, `X` rejects, `←`/`→` move, `U` jumps to the next unreviewed frame, + `1`–`9` pick the class. A batch can only be approved once no frame is still pending. +5. **Dataset & models** — what the master dataset holds, a training run, and the + base-versus-new table. + +## Reading the comparison + +After training, the base model and the new one are validated **on the same val set**, and +mAP50 / mAP50-95 / precision / recall appear side by side with the difference. + +Two things make that number trustworthy: + +- **The val split is stable.** Once a frame is in `val`, it stays there for every later + merge. A rising score cannot be an easier val set. +- **Training uses the whole master dataset**, old batches included. Fine-tuning on the newest + batch alone tends to raise the score on new footage while quietly losing the old. + +If the Base column is empty, the previous model's classes did not match this dataset's, so +scoring it here would have compared two different things. The message says which case it was. + +`Use as base model` promotes a version, and the next round fine-tunes from it. + +## Where things live + +``` +data/ + app.db # projects, batches, frames, annotations, jobs + projects// + base/model.pt # the base model + dataset/ # master dataset, accumulating + images/{train,val}/ labels/{train,val}/ data.yaml + batches//frames/ # extracted frames + models//best.pt + metrics.json # each training run +``` + +The database holds status; the disk holds pixels, labels and weights. The master dataset +trains as-is with Ultralytics, or imports into Roboflow, without this application. + +Your video archive is mounted read-only. Nothing is ever written back into it. + +## When something goes wrong + +| Symptom | Cause / fix | +|---|---| +| API returns 404 for routes you just added | the image copies `backend/` at build time — `docker compose build backend` again | +| A UI change doesn't show up | same trap on the other side: `docker compose build frontend`, then hard-reload | +| Job fails at *loading model* with a 401 | access to `facebook/sam3` not granted yet, or `HF_TOKEN` missing | +| SAM3 download crawls at a few KB/s | HuggingFace's Xet transfer throttling itself; `HF_HUB_DISABLE_XET=1` is already set in compose for that reason | +| `could not select device driver` | the NVIDIA container toolkit is not installed, or this Docker is older than the CDI support compose relies on (`devices: nvidia.com/gpu=all`) | +| `CUDA out of memory` while training | lower **Epochs**/batch on the Models page, or free the card — SAM3 is released before training, but another process may still hold it | +| Videos listed as *unreadable* | ffprobe could not parse them; they are still listed rather than hidden, so the archive never looks emptier than it is | +| A job reads *interrupted by a server restart* | it was running when the process died — jobs are not resumable, start it again | + +## Development + +```bash +uv pip install -r requirements.txt # backend +uv pip install -e sam3/ +cd frontend && npm install # frontend +npm run dev # :5173, proxies /api to :8000 +``` + +The frontend pins Vite 7 on purpose — Vite 8's Rolldown binding crashes on this machine. + +Working rules for agents and the documents that drive this repo are in `AGENTS.md` and +`docs/` (`requirements.md` → `design.md` → `tasks.md`).