Files
reTraining/docs/proposal-dataprep-triage.md
T
asus 5c7c122105 feat: add counting bench, triage, and dataset modules
This commit includes major additions and updates to the frontend and backend architectures, introducing new dataset management, live counting features, batch processing, and triage logic. Includes new UI pages, components, and API routes.
2026-08-14 16:28:52 +07:00

6.6 KiB
Raw Blame History

Proposal — Data Prep: outlier filter + augmentation

Status: scope approved 2026-08-13. Supersedes the rules-engine proposal agreed on 2026-08-07, which was never implemented into ./requirements.md. If accepted as written, REQ-100…REQ-105 and REQ-110…REQ-113 move into ./requirements.md.

What changed from the 2026-08-07 draft, and why:

  • The reclass action, the ordered rule list, and the preset suggestions are dropped. In practice the only decision being made on this page is "is this box junk?" — a reclass target and first-match-wins ordering were machinery for a decision nobody was making. Both triage_rules and annotation_overrides were empty when this was decided, so nothing was lost.
  • Data Prep gains augmentation settings, which the pipeline previously left entirely to Ultralytics' defaults.

Part 1 — Outlier filter

A shape's verdict is still resolved, never stored destructively:

manual override (if any)  >  outlier filter  >  default: keep

Verdict is one of keep · ignore. annotations.class_id is never rewritten, so a filter can be re-cut at any time.

The filter is expressed over the same three signals the resolver already computes — score, area_pct, aspect — as a keep-range per signal. Anything outside an enabled range is ignore.

Requirements

  • REQ-100 — A project has one outlier filter: an optional keep-range ([min, max], either edge blank) over each of score, area_pct and aspect. A shape falling outside any enabled range resolves to ignore.
  • REQ-101 — The filter is stage-scoped to dataprep. Editing it never alters what the batches or review stages display; it changes only what the stages after it consume.
  • REQ-102 — The filter is applied at merge time, against the live annotations. It is never baked into the master dataset, and annotations.class_id is never rewritten.
  • REQ-103 — The user can override any individual shape by hand (keep/ignore). A manual override outranks the filter and survives any later filter edit.
  • REQ-104 — A shape resolving to ignore drops that box. A frame that loses every box is held back from the dataset — an image is never trained on with a known object left unlabeled.
  • REQ-105 — The Data Prep page shows, for one batch: a scatter of score × area with rectangular selection, and a grid of cropped shape thumbnails. Selecting in either assigns a manual verdict in bulk. Filter edits update the kept/ignored counts live, before anything is saved.

Storage reuses the existing triage_rules table: the UI emits the filter as ignore rules and reads them back. No schema change, no second code path in the resolver.

Part 2 — Augmentation

Ultralytics augments during training whether or not we ask it to. backend/training.py passes no augmentation arguments, so every run so far has used library defaults (mosaic=1.0, fliplr=0.5, HSV jitter, scale=0.5) — invisibly, and unrecorded.

  • REQ-110 — A project stores augmentation settings: fliplr, flipud, degrees, translate, scale, hsv_h, hsv_s, hsv_v, mosaic. They are passed to model.train() on every run.
  • REQ-111 — The UI offers presets — Off, Light, Medium, Aggressive — and lets any single value be adjusted afterwards. A project with no stored settings uses Medium, which reproduces Ultralytics' defaults, so behaviour does not change until the user changes it.
  • REQ-112 — Augmentation applies to training images only. Validation is never augmented, so a base-vs-new mAP comparison stays a like-for-like measurement (this is Ultralytics' own behaviour; the requirement is that we must not defeat it).
  • REQ-113 — Each stored model version records the augmentation settings it trained under, so two runs can be told apart.

Schema

ALTER TABLE projects        ADD COLUMN augment TEXT;   -- JSON, null = Medium preset
ALTER TABLE model_versions  ADD COLUMN augment TEXT;   -- REQ-113: what this run used

API

GET  /api/projects/{id}/augment   -> {settings, preset}
PUT  /api/projects/{id}/augment   <- {settings}

Part 3 — Base datasets

Externally-labelled images the user already trusts, registered against a project and offered as a checkbox beside the project's own datasets. Not a batch, and never becomes one: no frames, no review, no triage.

  • REQ-120 — A project may register any number of base datasets, each a folder of images with YOLO label files, imported from an unpacked export. Class ids are kept as they are; the importer only drops classes the user did not ask to keep.
  • REQ-121 — A base dataset is opt-in per training run. Selecting no dataset at all never silently pulls one in.
  • REQ-122 — A base dataset contributes training images only. It can never supply validation images, because REQ-063's base-vs-new comparison is only meaningful measured on this project's own stable val split (REQ-052). A run with base datasets but no project dataset is refused — there would be nothing to validate on.
  • REQ-123 — Labels are normalised to the project's label_type on import. A segmentation polygon imported into a bbox project is collapsed to its bounding box, because a detect model reads the first four numbers of a polygon line as a box and would otherwise train on nonsense.

Schema

CREATE TABLE base_datasets (
  id          INTEGER PRIMARY KEY,
  project_id  INTEGER NOT NULL REFERENCES projects(id) ON DELETE CASCADE,
  name        TEXT NOT NULL,
  source      TEXT NOT NULL DEFAULT '',
  image_count INTEGER NOT NULL DEFAULT 0,
  box_count   INTEGER NOT NULL DEFAULT 0,
  classes     TEXT NOT NULL DEFAULT '[]',
  created_at  REAL NOT NULL
);

They stay out of dataset_items deliberately: that table keys on frame_id, and inventing frame rows for images this app never extracted would put fake batches in front of the user forever. The rows go straight to dataset._build_selected_tree, which only ever needed an image path and a label path.

API

GET    /api/projects/{id}/base-datasets  -> [base_dataset]
DELETE /api/base-datasets/{id}

Import is an offline call to base_dataset.import_tree — a 150 MB export is not an HTTP request worth holding open.


Open risk

REQ-104 is unchanged from the earlier draft and carries the same interaction: excluding whole frames means the val set is a function of the filter. The earlier REQ-107 (rule_version stamped on each run) already exists in the schema and keeps that honest — it is retained even though the rules engine around it is gone.