feat: setup dataset enrichment app codebase and scripts

This commit is contained in:
asus committed 2026-08-05 11:52:27 +07:00
1 parent b5c28cc98a
commit d07578462e
72 files changed
+11370

No files matched your search

+162
View File
@@ -0,0 +1,162 @@
# Requirements
Status: **agreed** (planning session, 2026-07-31). Changes only with the user's approval.
## Goal
A system for **enriching a dataset and improving an existing detection model**, iteratively,
from an archive of recorded video. One full round:
> pick a project → browse the video archive → pick a batch → trim a time range →
> extract frames → auto-annotate with SAM3 → review and correct every frame → approve →
> merge into the master dataset → fine-tune from the base model → compare against the base.
The system is **generic**: the sack case is only the first project. Other cases are
created as new projects with their own base model, classes, and video archive — no code
changes.
## Non-goals (for this version)
- Login, multi-user, tenants, quotas. The architecture leaves room for them; the features
are not built.
- Tracking or annotation propagation between frames.
- Collaborative annotation by several people at once.
- Public internet deployment.
---
## A. Project
- **REQ-001** — The user can create, list, and delete projects. A project has: name, label
type, base model, video archive root, and a class list.
- **REQ-002** — Each project picks a **label type**: `bbox` (YOLO detect) or `polygon`
(YOLO segment). This determines the export format, the editor's behaviour, and which
model variant is trained. It cannot be changed once a batch has been merged.
- **REQ-003** — The user uploads a **base model** `.pt`. The class list is read from the
model (`model.names`). It cannot drift on its own: nothing adds or removes a class as a
side effect of another action. Deliberate deletion is REQ-007.
- **REQ-004** — A project may be created **without** a base model. In that case the user
types the class list, and the first training starts from pretrained weights
(`yolo11n.pt` / `yolo11n-seg.pt`).
- **REQ-005** — Each class carries its own **SAM3 text prompt**, which may differ from the
class name (e.g. class `sack` with prompt `"woven plastic sack"`). Prompts can be
edited at any time without affecting existing data.
- **REQ-006** — All of a project's data (base model, master dataset, batch frames, trained
weights) lives under one project folder, so it can be copied or backed up whole.
- **REQ-007** — The user can **delete a class** at any point in a project's life, including
after batches have been merged. Deleting one:
- removes every annotation of that class, in batches under review and in the master
dataset alike;
- **renumbers the classes above it**, in the database *and* in every label file already
written to disk, because a YOLO label is an integer index and leaving a gap would make
old labels silently name the wrong class;
- regenerates `data.yaml`;
- is refused for a project's last remaining class.
Before confirming, the user is told how many shapes will be destroyed. The action cannot
be undone. If the project's base model was trained on the old class list, it stops being
comparable — which REQ-063 already reports rather than hides.
- **REQ-008** — The user can **add a new class** (name & prompt) to an existing project at any time.
The new class receives the next sequential `class_id`, and `data.yaml` is regenerated if a
master dataset exists.
## B. Video archive
- **REQ-010** — The video archive lives at a local path (disk or mount); videos are **not
uploaded** through the browser.
- **REQ-011** — Archive structure: `<video_root>/<date>/<batch>.<ext>`. The system lists the
dates, and within each date the videos with their batch labels parsed from the filename.
- **REQ-012** — Each video shows its duration, resolution, and whether it has already been
used as a batch in this project.
- **REQ-013** — Videos play in the browser with seeking (HTTP Range), without copying the
file first.
## C. Trim & frame extraction
- **REQ-020** — The user sets the in/out range with a timeline slider on the player, and can
also type precise timestamps.
- **REQ-021** — The user sets the extraction **frames per second** (default 1 fps). The
resulting frame count is shown before extraction runs.
- **REQ-022** — Extraction runs as a background job with progress, producing sequentially
numbered JPEG files inside the batch folder.
- **REQ-023** — One video may be used more than once with different time ranges; each
extraction produces its own batch.
## D. Auto-annotation
- **REQ-030** — Once frames are extracted, the system runs SAM3 over all of them using each
class's prompt, as a background job with progress and cancellation.
- **REQ-031** — Detections that overlap across prompts are deduplicated (greedy IoU NMS), so
one object is not labelled as two classes at once.
- **REQ-032** — The confidence threshold is configurable per job.
- **REQ-033** — A frame with no detections is valid and still enters the dataset as a
negative sample — it is not a failure.
- **REQ-034** — Auto-annotation can be re-run on the same batch; previous automatic results
are replaced, but **the user's manual corrections must never be lost**.
- **REQ-035** — Auto-annotation can be started in **resume** mode, which skips frames that
already carry automatic annotations. Resume is always an explicit choice and never the
default, because a full re-run is also how the confidence threshold (REQ-032) is changed —
the system cannot tell the two intentions apart, so it asks. A frame SAM3 legitimately
found nothing on (REQ-033) writes no annotations, so a resume re-does it; that is accepted
rather than tracked.
## E. Review & correction
- **REQ-040** — The user reviews frames one at a time, with fast navigation (left/right
arrows, thumbnail filmstrip, jump to the next unreviewed frame).
- **REQ-041** — Each frame has a status: `pending`, `approved`, or `rejected`. Rejected
frames never enter the dataset.
- **REQ-042** — The user can draw a new shape, move it, resize it, delete it, and change its
class.
- **REQ-043** — The user can ask SAM3 for help inside the editor: click or drag a box around
one object and the model produces its shape.
- **REQ-044** — All annotations and review statuses are **persistent** — they survive a
server restart, unlike today's in-memory sessions.
- **REQ-045** — Review progress is visible (e.g. "120/300 reviewed"), and a batch can only
be approved once no frame is still `pending`.
- **REQ-046** — The user can delete/clear all annotations of a specific class across all frames in
the current batch from the Review editor.
## F. Master dataset
- **REQ-050** — Approving a batch **merges** its approved frames and their labels into the
project's master dataset (accumulating across batches).
- **REQ-051** — The master dataset is train-ready YOLO format: `images/{train,val}`,
`labels/{train,val}`, and a `data.yaml` regenerated from the project's class list.
- **REQ-052** — **Stable val split**: once a frame is placed in `val`, it stays in `val`
across every later merge. New frames are split with an every-Nth pattern.
- **REQ-053** — The system records which batches have entered the master dataset, when, and
how many images/labels each added.
- **REQ-054** — The master dataset can be downloaded as a `.zip` (e.g. to import into
Roboflow or train on another machine).
## G. Training & evaluation
- **REQ-060** — The user starts training from the project page. Training **fine-tunes from
the project's base model** on the merged master dataset (old + new).
- **REQ-061** — A fallback option "train on the latest batch only" (lower LR, fewer epochs)
exists for cases where the old dataset is unavailable. It is not the default, and the UI
warns about catastrophic forgetting.
- **REQ-062** — Default `batch`, `imgsz`, and `device` are derived from the hardware detected
at runtime (VRAM), and all of them can be overridden — so moving to a bigger machine needs
no code change.
- **REQ-063** — After training, the system validates **the base model and the new model on
the exact same val set**, then shows mAP50 and mAP50-95 for both side by side with the
delta.
- **REQ-064** — Each training run produces a stored model version (weights + metrics). The
user can download the weights and **promote that version to be the project's new base
model** for the next round.
- **REQ-065** — SAM3 and training must never hold VRAM at the same time; the system releases
the SAM3 model before training starts.
## H. System
- **REQ-070** — Heavy work (extraction, auto-annotation, training) runs as queued jobs, one
at a time, because there is a single GPU. Jobs show progress and logs, and can be cancelled.
- **REQ-071** — Jobs and their progress are persistent; after a server restart the job list
is still there with its final statuses.
- **REQ-072** — The application runs via `docker compose up` with GPU access, and every path
(video archive, data folder) is configured through environment/volumes — never hardcoded.
- **REQ-073** — The health endpoint reports: detected device/GPU, ffmpeg availability,
whether the HuggingFace token was picked up, and database reachability.
- **REQ-074** — The system never writes anything into the user's video archive folder.