Files
reTraining/docs/proposal-dataprep-triage.md
T
asus 5c7c122105 feat: add counting bench, triage, and dataset modules
This commit includes major additions and updates to the frontend and backend architectures, introducing new dataset management, live counting features, batch processing, and triage logic. Includes new UI pages, components, and API routes.
2026-08-14 16:28:52 +07:00

147 lines
6.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Proposal — Data Prep: outlier filter + augmentation
Status: **scope approved 2026-08-13**. Supersedes the rules-engine proposal agreed on
2026-08-07, which was never implemented into `./requirements.md`. If accepted as written,
REQ-100…REQ-105 and REQ-110…REQ-113 move into `./requirements.md`.
What changed from the 2026-08-07 draft, and why:
- The `reclass` action, the ordered rule list, and the preset suggestions are **dropped**.
In practice the only decision being made on this page is "is this box junk?" — a
reclass target and first-match-wins ordering were machinery for a decision nobody was
making. Both `triage_rules` and `annotation_overrides` were empty when this was decided,
so nothing was lost.
- Data Prep gains **augmentation settings**, which the pipeline previously left entirely to
Ultralytics' defaults.
---
## Part 1 — Outlier filter
A shape's verdict is still resolved, never stored destructively:
```
manual override (if any) > outlier filter > default: keep
```
Verdict is one of `keep` · `ignore`. `annotations.class_id` is never rewritten, so a
filter can be re-cut at any time.
The filter is expressed over the same three signals the resolver already computes —
`score`, `area_pct`, `aspect` — as a keep-range per signal. Anything outside an enabled
range is `ignore`.
### Requirements
- **REQ-100** — A project has one **outlier filter**: an optional keep-range
(`[min, max]`, either edge blank) over each of `score`, `area_pct` and `aspect`. A shape
falling outside any enabled range resolves to `ignore`.
- **REQ-101** — The filter is **stage-scoped** to `dataprep`. Editing it never alters what
the batches or review stages display; it changes only what the stages after it consume.
- **REQ-102** — The filter is applied **at merge time**, against the live annotations. It is
never baked into the master dataset, and `annotations.class_id` is never rewritten.
- **REQ-103** — The user can **override any individual shape** by hand (`keep`/`ignore`).
A manual override outranks the filter and survives any later filter edit.
- **REQ-104** — A shape resolving to `ignore` drops that box. A frame that loses **every**
box is held back from the dataset — an image is never trained on with a known object
left unlabeled.
- **REQ-105** — The Data Prep page shows, for one batch: a scatter of score × area with
rectangular selection, and a grid of cropped shape thumbnails. Selecting in either
assigns a manual verdict in bulk. Filter edits update the kept/ignored counts live,
before anything is saved.
Storage reuses the existing `triage_rules` table: the UI emits the filter as
`ignore` rules and reads them back. No schema change, no second code path in the resolver.
## Part 2 — Augmentation
Ultralytics augments during training whether or not we ask it to. `backend/training.py`
passes no augmentation arguments, so every run so far has used library defaults
(`mosaic=1.0`, `fliplr=0.5`, HSV jitter, `scale=0.5`) — invisibly, and unrecorded.
- **REQ-110** — A project stores **augmentation settings**: `fliplr`, `flipud`, `degrees`,
`translate`, `scale`, `hsv_h`, `hsv_s`, `hsv_v`, `mosaic`. They are passed to
`model.train()` on every run.
- **REQ-111** — The UI offers presets — **Off**, **Light**, **Medium**, **Aggressive** —
and lets any single value be adjusted afterwards. A project with no stored settings uses
**Medium**, which reproduces Ultralytics' defaults, so behaviour does not change until
the user changes it.
- **REQ-112** — Augmentation applies to **training images only**. Validation is never
augmented, so a base-vs-new mAP comparison stays a like-for-like measurement
(this is Ultralytics' own behaviour; the requirement is that we must not defeat it).
- **REQ-113** — Each stored model version records the augmentation settings it trained
under, so two runs can be told apart.
### Schema
```sql
ALTER TABLE projects ADD COLUMN augment TEXT; -- JSON, null = Medium preset
ALTER TABLE model_versions ADD COLUMN augment TEXT; -- REQ-113: what this run used
```
### API
```
GET /api/projects/{id}/augment -> {settings, preset}
PUT /api/projects/{id}/augment <- {settings}
```
## Part 3 — Base datasets
Externally-labelled images the user already trusts, registered against a project and
offered as a checkbox beside the project's own datasets. Not a batch, and never becomes
one: no frames, no review, no triage.
- **REQ-120** — A project may register any number of **base datasets**, each a folder of
images with YOLO label files, imported from an unpacked export. Class ids are kept as
they are; the importer only drops classes the user did not ask to keep.
- **REQ-121** — A base dataset is **opt-in per training run**. Selecting no dataset at all
never silently pulls one in.
- **REQ-122** — A base dataset contributes **training images only**. It can never supply
validation images, because REQ-063's base-vs-new comparison is only meaningful measured
on this project's own stable val split (REQ-052). A run with base datasets but no project
dataset is refused — there would be nothing to validate on.
- **REQ-123** — Labels are normalised to the project's `label_type` on import. A
segmentation polygon imported into a `bbox` project is collapsed to its bounding box,
because a detect model reads the first four numbers of a polygon line as a box and would
otherwise train on nonsense.
### Schema
```sql
CREATE TABLE base_datasets (
id INTEGER PRIMARY KEY,
project_id INTEGER NOT NULL REFERENCES projects(id) ON DELETE CASCADE,
name TEXT NOT NULL,
source TEXT NOT NULL DEFAULT '',
image_count INTEGER NOT NULL DEFAULT 0,
box_count INTEGER NOT NULL DEFAULT 0,
classes TEXT NOT NULL DEFAULT '[]',
created_at REAL NOT NULL
);
```
They stay out of `dataset_items` deliberately: that table keys on `frame_id`, and inventing
frame rows for images this app never extracted would put fake batches in front of the user
forever. The rows go straight to `dataset._build_selected_tree`, which only ever needed an
image path and a label path.
### API
```
GET /api/projects/{id}/base-datasets -> [base_dataset]
DELETE /api/base-datasets/{id}
```
Import is an offline call to `base_dataset.import_tree` — a 150 MB export is not an HTTP
request worth holding open.
---
## Open risk
REQ-104 is unchanged from the earlier draft and carries the same interaction: excluding
whole frames means the val set is a function of the filter. The earlier REQ-107
(`rule_version` stamped on each run) already exists in the schema and keeps that honest —
it is retained even though the rules engine around it is gone.