This commit includes major additions and updates to the frontend and backend architectures, introducing new dataset management, live counting features, batch processing, and triage logic. Includes new UI pages, components, and API routes.
6.6 KiB
Proposal — Data Prep: outlier filter + augmentation
Status: scope approved 2026-08-13. Supersedes the rules-engine proposal agreed on
2026-08-07, which was never implemented into ./requirements.md. If accepted as written,
REQ-100…REQ-105 and REQ-110…REQ-113 move into ./requirements.md.
What changed from the 2026-08-07 draft, and why:
- The
reclassaction, the ordered rule list, and the preset suggestions are dropped. In practice the only decision being made on this page is "is this box junk?" — a reclass target and first-match-wins ordering were machinery for a decision nobody was making. Bothtriage_rulesandannotation_overrideswere empty when this was decided, so nothing was lost. - Data Prep gains augmentation settings, which the pipeline previously left entirely to Ultralytics' defaults.
Part 1 — Outlier filter
A shape's verdict is still resolved, never stored destructively:
manual override (if any) > outlier filter > default: keep
Verdict is one of keep · ignore. annotations.class_id is never rewritten, so a
filter can be re-cut at any time.
The filter is expressed over the same three signals the resolver already computes —
score, area_pct, aspect — as a keep-range per signal. Anything outside an enabled
range is ignore.
Requirements
- REQ-100 — A project has one outlier filter: an optional keep-range
(
[min, max], either edge blank) over each ofscore,area_pctandaspect. A shape falling outside any enabled range resolves toignore. - REQ-101 — The filter is stage-scoped to
dataprep. Editing it never alters what the batches or review stages display; it changes only what the stages after it consume. - REQ-102 — The filter is applied at merge time, against the live annotations. It is
never baked into the master dataset, and
annotations.class_idis never rewritten. - REQ-103 — The user can override any individual shape by hand (
keep/ignore). A manual override outranks the filter and survives any later filter edit. - REQ-104 — A shape resolving to
ignoredrops that box. A frame that loses every box is held back from the dataset — an image is never trained on with a known object left unlabeled. - REQ-105 — The Data Prep page shows, for one batch: a scatter of score × area with rectangular selection, and a grid of cropped shape thumbnails. Selecting in either assigns a manual verdict in bulk. Filter edits update the kept/ignored counts live, before anything is saved.
Storage reuses the existing triage_rules table: the UI emits the filter as
ignore rules and reads them back. No schema change, no second code path in the resolver.
Part 2 — Augmentation
Ultralytics augments during training whether or not we ask it to. backend/training.py
passes no augmentation arguments, so every run so far has used library defaults
(mosaic=1.0, fliplr=0.5, HSV jitter, scale=0.5) — invisibly, and unrecorded.
- REQ-110 — A project stores augmentation settings:
fliplr,flipud,degrees,translate,scale,hsv_h,hsv_s,hsv_v,mosaic. They are passed tomodel.train()on every run. - REQ-111 — The UI offers presets — Off, Light, Medium, Aggressive — and lets any single value be adjusted afterwards. A project with no stored settings uses Medium, which reproduces Ultralytics' defaults, so behaviour does not change until the user changes it.
- REQ-112 — Augmentation applies to training images only. Validation is never augmented, so a base-vs-new mAP comparison stays a like-for-like measurement (this is Ultralytics' own behaviour; the requirement is that we must not defeat it).
- REQ-113 — Each stored model version records the augmentation settings it trained under, so two runs can be told apart.
Schema
ALTER TABLE projects ADD COLUMN augment TEXT; -- JSON, null = Medium preset
ALTER TABLE model_versions ADD COLUMN augment TEXT; -- REQ-113: what this run used
API
GET /api/projects/{id}/augment -> {settings, preset}
PUT /api/projects/{id}/augment <- {settings}
Part 3 — Base datasets
Externally-labelled images the user already trusts, registered against a project and offered as a checkbox beside the project's own datasets. Not a batch, and never becomes one: no frames, no review, no triage.
- REQ-120 — A project may register any number of base datasets, each a folder of images with YOLO label files, imported from an unpacked export. Class ids are kept as they are; the importer only drops classes the user did not ask to keep.
- REQ-121 — A base dataset is opt-in per training run. Selecting no dataset at all never silently pulls one in.
- REQ-122 — A base dataset contributes training images only. It can never supply validation images, because REQ-063's base-vs-new comparison is only meaningful measured on this project's own stable val split (REQ-052). A run with base datasets but no project dataset is refused — there would be nothing to validate on.
- REQ-123 — Labels are normalised to the project's
label_typeon import. A segmentation polygon imported into abboxproject is collapsed to its bounding box, because a detect model reads the first four numbers of a polygon line as a box and would otherwise train on nonsense.
Schema
CREATE TABLE base_datasets (
id INTEGER PRIMARY KEY,
project_id INTEGER NOT NULL REFERENCES projects(id) ON DELETE CASCADE,
name TEXT NOT NULL,
source TEXT NOT NULL DEFAULT '',
image_count INTEGER NOT NULL DEFAULT 0,
box_count INTEGER NOT NULL DEFAULT 0,
classes TEXT NOT NULL DEFAULT '[]',
created_at REAL NOT NULL
);
They stay out of dataset_items deliberately: that table keys on frame_id, and inventing
frame rows for images this app never extracted would put fake batches in front of the user
forever. The rows go straight to dataset._build_selected_tree, which only ever needed an
image path and a label path.
API
GET /api/projects/{id}/base-datasets -> [base_dataset]
DELETE /api/base-datasets/{id}
Import is an offline call to base_dataset.import_tree — a 150 MB export is not an HTTP
request worth holding open.
Open risk
REQ-104 is unchanged from the earlier draft and carries the same interaction: excluding
whole frames means the val set is a function of the filter. The earlier REQ-107
(rule_version stamped on each run) already exists in the schema and keeps that honest —
it is retained even though the rules engine around it is gone.