# Proposal — Data Prep: outlier filter + augmentation Status: **scope approved 2026-08-13**. Supersedes the rules-engine proposal agreed on 2026-08-07, which was never implemented into `./requirements.md`. If accepted as written, REQ-100…REQ-105 and REQ-110…REQ-113 move into `./requirements.md`. What changed from the 2026-08-07 draft, and why: - The `reclass` action, the ordered rule list, and the preset suggestions are **dropped**. In practice the only decision being made on this page is "is this box junk?" — a reclass target and first-match-wins ordering were machinery for a decision nobody was making. Both `triage_rules` and `annotation_overrides` were empty when this was decided, so nothing was lost. - Data Prep gains **augmentation settings**, which the pipeline previously left entirely to Ultralytics' defaults. --- ## Part 1 — Outlier filter A shape's verdict is still resolved, never stored destructively: ``` manual override (if any) > outlier filter > default: keep ``` Verdict is one of `keep` · `ignore`. `annotations.class_id` is never rewritten, so a filter can be re-cut at any time. The filter is expressed over the same three signals the resolver already computes — `score`, `area_pct`, `aspect` — as a keep-range per signal. Anything outside an enabled range is `ignore`. ### Requirements - **REQ-100** — A project has one **outlier filter**: an optional keep-range (`[min, max]`, either edge blank) over each of `score`, `area_pct` and `aspect`. A shape falling outside any enabled range resolves to `ignore`. - **REQ-101** — The filter is **stage-scoped** to `dataprep`. Editing it never alters what the batches or review stages display; it changes only what the stages after it consume. - **REQ-102** — The filter is applied **at merge time**, against the live annotations. It is never baked into the master dataset, and `annotations.class_id` is never rewritten. - **REQ-103** — The user can **override any individual shape** by hand (`keep`/`ignore`). A manual override outranks the filter and survives any later filter edit. - **REQ-104** — A shape resolving to `ignore` drops that box. A frame that loses **every** box is held back from the dataset — an image is never trained on with a known object left unlabeled. - **REQ-105** — The Data Prep page shows, for one batch: a scatter of score × area with rectangular selection, and a grid of cropped shape thumbnails. Selecting in either assigns a manual verdict in bulk. Filter edits update the kept/ignored counts live, before anything is saved. Storage reuses the existing `triage_rules` table: the UI emits the filter as `ignore` rules and reads them back. No schema change, no second code path in the resolver. ## Part 2 — Augmentation Ultralytics augments during training whether or not we ask it to. `backend/training.py` passes no augmentation arguments, so every run so far has used library defaults (`mosaic=1.0`, `fliplr=0.5`, HSV jitter, `scale=0.5`) — invisibly, and unrecorded. - **REQ-110** — A project stores **augmentation settings**: `fliplr`, `flipud`, `degrees`, `translate`, `scale`, `hsv_h`, `hsv_s`, `hsv_v`, `mosaic`. They are passed to `model.train()` on every run. - **REQ-111** — The UI offers presets — **Off**, **Light**, **Medium**, **Aggressive** — and lets any single value be adjusted afterwards. A project with no stored settings uses **Medium**, which reproduces Ultralytics' defaults, so behaviour does not change until the user changes it. - **REQ-112** — Augmentation applies to **training images only**. Validation is never augmented, so a base-vs-new mAP comparison stays a like-for-like measurement (this is Ultralytics' own behaviour; the requirement is that we must not defeat it). - **REQ-113** — Each stored model version records the augmentation settings it trained under, so two runs can be told apart. ### Schema ```sql ALTER TABLE projects ADD COLUMN augment TEXT; -- JSON, null = Medium preset ALTER TABLE model_versions ADD COLUMN augment TEXT; -- REQ-113: what this run used ``` ### API ``` GET /api/projects/{id}/augment -> {settings, preset} PUT /api/projects/{id}/augment <- {settings} ``` ## Part 3 — Base datasets Externally-labelled images the user already trusts, registered against a project and offered as a checkbox beside the project's own datasets. Not a batch, and never becomes one: no frames, no review, no triage. - **REQ-120** — A project may register any number of **base datasets**, each a folder of images with YOLO label files, imported from an unpacked export. Class ids are kept as they are; the importer only drops classes the user did not ask to keep. - **REQ-121** — A base dataset is **opt-in per training run**. Selecting no dataset at all never silently pulls one in. - **REQ-122** — A base dataset contributes **training images only**. It can never supply validation images, because REQ-063's base-vs-new comparison is only meaningful measured on this project's own stable val split (REQ-052). A run with base datasets but no project dataset is refused — there would be nothing to validate on. - **REQ-123** — Labels are normalised to the project's `label_type` on import. A segmentation polygon imported into a `bbox` project is collapsed to its bounding box, because a detect model reads the first four numbers of a polygon line as a box and would otherwise train on nonsense. ### Schema ```sql CREATE TABLE base_datasets ( id INTEGER PRIMARY KEY, project_id INTEGER NOT NULL REFERENCES projects(id) ON DELETE CASCADE, name TEXT NOT NULL, source TEXT NOT NULL DEFAULT '', image_count INTEGER NOT NULL DEFAULT 0, box_count INTEGER NOT NULL DEFAULT 0, classes TEXT NOT NULL DEFAULT '[]', created_at REAL NOT NULL ); ``` They stay out of `dataset_items` deliberately: that table keys on `frame_id`, and inventing frame rows for images this app never extracted would put fake batches in front of the user forever. The rows go straight to `dataset._build_selected_tree`, which only ever needed an image path and a label path. ### API ``` GET /api/projects/{id}/base-datasets -> [base_dataset] DELETE /api/base-datasets/{id} ``` Import is an offline call to `base_dataset.import_tree` — a 150 MB export is not an HTTP request worth holding open. --- ## Open risk REQ-104 is unchanged from the earlier draft and carries the same interaction: excluding whole frames means the val set is a function of the filter. The earlier REQ-107 (`rule_version` stamped on each run) already exists in the schema and keeps that honest — it is retained even though the rules engine around it is gone.