This commit includes major additions and updates to the frontend and backend architectures, introducing new dataset management, live counting features, batch processing, and triage logic. Includes new UI pages, components, and API routes.
147 lines
6.6 KiB
Markdown
147 lines
6.6 KiB
Markdown
# Proposal — Data Prep: outlier filter + augmentation
|
||
|
||
Status: **scope approved 2026-08-13**. Supersedes the rules-engine proposal agreed on
|
||
2026-08-07, which was never implemented into `./requirements.md`. If accepted as written,
|
||
REQ-100…REQ-105 and REQ-110…REQ-113 move into `./requirements.md`.
|
||
|
||
What changed from the 2026-08-07 draft, and why:
|
||
|
||
- The `reclass` action, the ordered rule list, and the preset suggestions are **dropped**.
|
||
In practice the only decision being made on this page is "is this box junk?" — a
|
||
reclass target and first-match-wins ordering were machinery for a decision nobody was
|
||
making. Both `triage_rules` and `annotation_overrides` were empty when this was decided,
|
||
so nothing was lost.
|
||
- Data Prep gains **augmentation settings**, which the pipeline previously left entirely to
|
||
Ultralytics' defaults.
|
||
|
||
---
|
||
|
||
## Part 1 — Outlier filter
|
||
|
||
A shape's verdict is still resolved, never stored destructively:
|
||
|
||
```
|
||
manual override (if any) > outlier filter > default: keep
|
||
```
|
||
|
||
Verdict is one of `keep` · `ignore`. `annotations.class_id` is never rewritten, so a
|
||
filter can be re-cut at any time.
|
||
|
||
The filter is expressed over the same three signals the resolver already computes —
|
||
`score`, `area_pct`, `aspect` — as a keep-range per signal. Anything outside an enabled
|
||
range is `ignore`.
|
||
|
||
### Requirements
|
||
|
||
- **REQ-100** — A project has one **outlier filter**: an optional keep-range
|
||
(`[min, max]`, either edge blank) over each of `score`, `area_pct` and `aspect`. A shape
|
||
falling outside any enabled range resolves to `ignore`.
|
||
- **REQ-101** — The filter is **stage-scoped** to `dataprep`. Editing it never alters what
|
||
the batches or review stages display; it changes only what the stages after it consume.
|
||
- **REQ-102** — The filter is applied **at merge time**, against the live annotations. It is
|
||
never baked into the master dataset, and `annotations.class_id` is never rewritten.
|
||
- **REQ-103** — The user can **override any individual shape** by hand (`keep`/`ignore`).
|
||
A manual override outranks the filter and survives any later filter edit.
|
||
- **REQ-104** — A shape resolving to `ignore` drops that box. A frame that loses **every**
|
||
box is held back from the dataset — an image is never trained on with a known object
|
||
left unlabeled.
|
||
- **REQ-105** — The Data Prep page shows, for one batch: a scatter of score × area with
|
||
rectangular selection, and a grid of cropped shape thumbnails. Selecting in either
|
||
assigns a manual verdict in bulk. Filter edits update the kept/ignored counts live,
|
||
before anything is saved.
|
||
|
||
Storage reuses the existing `triage_rules` table: the UI emits the filter as
|
||
`ignore` rules and reads them back. No schema change, no second code path in the resolver.
|
||
|
||
## Part 2 — Augmentation
|
||
|
||
Ultralytics augments during training whether or not we ask it to. `backend/training.py`
|
||
passes no augmentation arguments, so every run so far has used library defaults
|
||
(`mosaic=1.0`, `fliplr=0.5`, HSV jitter, `scale=0.5`) — invisibly, and unrecorded.
|
||
|
||
- **REQ-110** — A project stores **augmentation settings**: `fliplr`, `flipud`, `degrees`,
|
||
`translate`, `scale`, `hsv_h`, `hsv_s`, `hsv_v`, `mosaic`. They are passed to
|
||
`model.train()` on every run.
|
||
- **REQ-111** — The UI offers presets — **Off**, **Light**, **Medium**, **Aggressive** —
|
||
and lets any single value be adjusted afterwards. A project with no stored settings uses
|
||
**Medium**, which reproduces Ultralytics' defaults, so behaviour does not change until
|
||
the user changes it.
|
||
- **REQ-112** — Augmentation applies to **training images only**. Validation is never
|
||
augmented, so a base-vs-new mAP comparison stays a like-for-like measurement
|
||
(this is Ultralytics' own behaviour; the requirement is that we must not defeat it).
|
||
- **REQ-113** — Each stored model version records the augmentation settings it trained
|
||
under, so two runs can be told apart.
|
||
|
||
### Schema
|
||
|
||
```sql
|
||
ALTER TABLE projects ADD COLUMN augment TEXT; -- JSON, null = Medium preset
|
||
ALTER TABLE model_versions ADD COLUMN augment TEXT; -- REQ-113: what this run used
|
||
```
|
||
|
||
### API
|
||
|
||
```
|
||
GET /api/projects/{id}/augment -> {settings, preset}
|
||
PUT /api/projects/{id}/augment <- {settings}
|
||
```
|
||
|
||
## Part 3 — Base datasets
|
||
|
||
Externally-labelled images the user already trusts, registered against a project and
|
||
offered as a checkbox beside the project's own datasets. Not a batch, and never becomes
|
||
one: no frames, no review, no triage.
|
||
|
||
- **REQ-120** — A project may register any number of **base datasets**, each a folder of
|
||
images with YOLO label files, imported from an unpacked export. Class ids are kept as
|
||
they are; the importer only drops classes the user did not ask to keep.
|
||
- **REQ-121** — A base dataset is **opt-in per training run**. Selecting no dataset at all
|
||
never silently pulls one in.
|
||
- **REQ-122** — A base dataset contributes **training images only**. It can never supply
|
||
validation images, because REQ-063's base-vs-new comparison is only meaningful measured
|
||
on this project's own stable val split (REQ-052). A run with base datasets but no project
|
||
dataset is refused — there would be nothing to validate on.
|
||
- **REQ-123** — Labels are normalised to the project's `label_type` on import. A
|
||
segmentation polygon imported into a `bbox` project is collapsed to its bounding box,
|
||
because a detect model reads the first four numbers of a polygon line as a box and would
|
||
otherwise train on nonsense.
|
||
|
||
### Schema
|
||
|
||
```sql
|
||
CREATE TABLE base_datasets (
|
||
id INTEGER PRIMARY KEY,
|
||
project_id INTEGER NOT NULL REFERENCES projects(id) ON DELETE CASCADE,
|
||
name TEXT NOT NULL,
|
||
source TEXT NOT NULL DEFAULT '',
|
||
image_count INTEGER NOT NULL DEFAULT 0,
|
||
box_count INTEGER NOT NULL DEFAULT 0,
|
||
classes TEXT NOT NULL DEFAULT '[]',
|
||
created_at REAL NOT NULL
|
||
);
|
||
```
|
||
|
||
They stay out of `dataset_items` deliberately: that table keys on `frame_id`, and inventing
|
||
frame rows for images this app never extracted would put fake batches in front of the user
|
||
forever. The rows go straight to `dataset._build_selected_tree`, which only ever needed an
|
||
image path and a label path.
|
||
|
||
### API
|
||
|
||
```
|
||
GET /api/projects/{id}/base-datasets -> [base_dataset]
|
||
DELETE /api/base-datasets/{id}
|
||
```
|
||
|
||
Import is an offline call to `base_dataset.import_tree` — a 150 MB export is not an HTTP
|
||
request worth holding open.
|
||
|
||
---
|
||
|
||
## Open risk
|
||
|
||
REQ-104 is unchanged from the earlier draft and carries the same interaction: excluding
|
||
whole frames means the val set is a function of the filter. The earlier REQ-107
|
||
(`rule_version` stamped on each run) already exists in the schema and keeps that honest —
|
||
it is retained even though the rules engine around it is gone.
|