first commit
This commit is contained in:
commit
b5c28cc98a
1 file changed
+125
@@ -0,0 +1,125 @@
|
||||
# Dataset Enrichment
|
||||
|
||||
Take a model you already have, and make it better with footage you already have.
|
||||
|
||||
One round of the loop:
|
||||
|
||||
> pick a project → browse the video archive → pick a batch → trim a range →
|
||||
> extract frames → auto-annotate with SAM3 → review and correct every frame → approve →
|
||||
> merge into the master dataset → fine-tune from the base model → compare against the base.
|
||||
|
||||
Nothing here is specific to one dataset. A project owns its base model, its class list, its
|
||||
video archive and its own accumulating dataset, so a second use case is a second project —
|
||||
not a second copy of the code.
|
||||
|
||||
## Run it
|
||||
|
||||
```bash
|
||||
cp .env.example .env # paste HF_TOKEN
|
||||
VIDEO_ARCHIVE_HOST=/path/to/videos docker compose up -d --build
|
||||
```
|
||||
|
||||
Then open <http://localhost:8080>. The API is on `:8000` if you want to poke at it directly;
|
||||
`curl localhost:8000/api/health` reports the GPU, free VRAM (`vram_free_gb`), SAM3 readiness (`sam3_ready`), whether ffmpeg is present, whether the token was picked up, and whether the database is reachable.
|
||||
|
||||
|
||||
The first auto-annotation job downloads the ~3.4 GB SAM3 checkpoint into a Docker volume.
|
||||
It happens once; later jobs take about 12 seconds to load the model into VRAM.
|
||||
|
||||
### What you need first
|
||||
|
||||
| Thing | Why |
|
||||
|---|---|
|
||||
| A GPU with the NVIDIA container toolkit | SAM3 is CUDA-only |
|
||||
| A video archive laid out as `<date>/<batch>.mp4` | that structure is what the Library reads |
|
||||
| `HF_TOKEN` with access to [facebook/sam3](https://huggingface.co/facebook/sam3) | the weights are gated, and approval is manual |
|
||||
| A base model `.pt` (optional) | without one, training starts from `yolo11n.pt` and you type the classes yourself |
|
||||
|
||||
Example archive:
|
||||
|
||||
```
|
||||
videos/
|
||||
2026-07-08/
|
||||
batch-1.mp4
|
||||
batch-4.mp4
|
||||
2026-07-09/
|
||||
batch-2.mp4
|
||||
```
|
||||
|
||||
## The screens
|
||||
|
||||
1. **Projects** — name, label type (`bbox` or `polygon`), archive root, classes. Upload a
|
||||
base model and its classes are read from the checkpoint and locked, so the dataset and
|
||||
the model can never drift apart.
|
||||
2. **Library** — dates on the left, that date's recordings on the right with duration,
|
||||
resolution and how many batches already came out of each.
|
||||
3. **Trim** — play the video, set in/out, pick a frame rate. It tells you how many frames
|
||||
that produces before you commit to it.
|
||||
4. **Review** — the frame with its shapes on top. Drag to add a box, drag a corner to
|
||||
resize, drag the middle to move, `Del` to remove. Hold `S` and drag for a SAM3-assisted
|
||||
shape. `A` approves, `X` rejects, `←`/`→` move, `U` jumps to the next unreviewed frame,
|
||||
`1`–`9` pick the class. A batch can only be approved once no frame is still pending.
|
||||
5. **Dataset & models** — what the master dataset holds, a training run, and the
|
||||
base-versus-new table.
|
||||
|
||||
## Reading the comparison
|
||||
|
||||
After training, the base model and the new one are validated **on the same val set**, and
|
||||
mAP50 / mAP50-95 / precision / recall appear side by side with the difference.
|
||||
|
||||
Two things make that number trustworthy:
|
||||
|
||||
- **The val split is stable.** Once a frame is in `val`, it stays there for every later
|
||||
merge. A rising score cannot be an easier val set.
|
||||
- **Training uses the whole master dataset**, old batches included. Fine-tuning on the newest
|
||||
batch alone tends to raise the score on new footage while quietly losing the old.
|
||||
|
||||
If the Base column is empty, the previous model's classes did not match this dataset's, so
|
||||
scoring it here would have compared two different things. The message says which case it was.
|
||||
|
||||
`Use as base model` promotes a version, and the next round fine-tunes from it.
|
||||
|
||||
## Where things live
|
||||
|
||||
```
|
||||
data/
|
||||
app.db # projects, batches, frames, annotations, jobs
|
||||
projects/<slug>/
|
||||
base/model.pt # the base model
|
||||
dataset/ # master dataset, accumulating
|
||||
images/{train,val}/ labels/{train,val}/ data.yaml
|
||||
batches/<id>/frames/ # extracted frames
|
||||
models/<n>/best.pt + metrics.json # each training run
|
||||
```
|
||||
|
||||
The database holds status; the disk holds pixels, labels and weights. The master dataset
|
||||
trains as-is with Ultralytics, or imports into Roboflow, without this application.
|
||||
|
||||
Your video archive is mounted read-only. Nothing is ever written back into it.
|
||||
|
||||
## When something goes wrong
|
||||
|
||||
| Symptom | Cause / fix |
|
||||
|---|---|
|
||||
| API returns 404 for routes you just added | the image copies `backend/` at build time — `docker compose build backend` again |
|
||||
| A UI change doesn't show up | same trap on the other side: `docker compose build frontend`, then hard-reload |
|
||||
| Job fails at *loading model* with a 401 | access to `facebook/sam3` not granted yet, or `HF_TOKEN` missing |
|
||||
| SAM3 download crawls at a few KB/s | HuggingFace's Xet transfer throttling itself; `HF_HUB_DISABLE_XET=1` is already set in compose for that reason |
|
||||
| `could not select device driver` | the NVIDIA container toolkit is not installed, or this Docker is older than the CDI support compose relies on (`devices: nvidia.com/gpu=all`) |
|
||||
| `CUDA out of memory` while training | lower **Epochs**/batch on the Models page, or free the card — SAM3 is released before training, but another process may still hold it |
|
||||
| Videos listed as *unreadable* | ffprobe could not parse them; they are still listed rather than hidden, so the archive never looks emptier than it is |
|
||||
| A job reads *interrupted by a server restart* | it was running when the process died — jobs are not resumable, start it again |
|
||||
|
||||
## Development
|
||||
|
||||
```bash
|
||||
uv pip install -r requirements.txt # backend
|
||||
uv pip install -e sam3/
|
||||
cd frontend && npm install # frontend
|
||||
npm run dev # :5173, proxies /api to :8000
|
||||
```
|
||||
|
||||
The frontend pins Vite 7 on purpose — Vite 8's Rolldown binding crashes on this machine.
|
||||
|
||||
Working rules for agents and the documents that drive this repo are in `AGENTS.md` and
|
||||
`docs/` (`requirements.md` → `design.md` → `tasks.md`).
|
||||
Reference in new issue
Block a user