feat(backend): diff-vs-previous-run reporting for product-scan accuracy harness
Ports the DO-harness's auto-diff-vs-previous-run reporting into accuracy-check-scan.mts: prints a per-field, per-split (Training/ Validation) delta against the last product_accuracy_history.jsonl entry and calls out regressions/improvements explicitly, plus classifier method distribution and average confidence as informational context. Also adds the real held-out validation photo set into sources/product-test-images/ (75 photos, one per current SKU class) with its README documenting the drop-photo -> label -> re-run workflow, so the harness's Validation Set split actually has images to score against. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Xsxk4ZkDQVVaLUcixDcqb5
This commit is contained in:
1 parent
dc0dd81318
commit
3a17c28758
78 files changed
+306
-55
No files matched your search
@@ -71,7 +71,8 @@ workflow and have no task numbers; see `git log` for real dates/history.
|
||||
|
||||
- **6.1** Built standalone annotation page `manual-label-scan/page.tsx` for ground truth editing. Includes image browser, editable fields (`no_sku`, `nama_item`, `expiry_date`, `notes`), and a "Scan with AI" fill-blanks feature — shipped 2026-07-08.
|
||||
- **6.2** API + storage groundwork for scan annotation. Extended `api/manual-label-scan` with `GET` list mode and `DELETE`. Persisted uploaded scan photos as base64 images into `sources/product-test-images/`. Made the `scan-pfm` quick-save honest by allowing manual correction before save — shipped 2026-07-08.
|
||||
- **6.3** Built `backend/pfm-web-app/scripts/accuracy-check-scan.mts` mirroring the DO-harness architecture, measuring overall match rate plus per-field breakdown (`no_sku`, `expiry_date`) against the new stable labels — shipped 2026-07-08.
|
||||
- **6.3** Built `backend/scripts/accuracy-check-scan.mts` mirroring the DO-harness architecture, measuring overall match rate plus per-field breakdown (`no_sku`, `expiry_date`) against the new stable labels — shipped 2026-07-08.
|
||||
- **6.4** Ported the DO-harness's auto-diff-vs-previous-run reporting into `accuracy-check-scan.mts`: every run now prints a Δ column per field per split (Training/Validation) vs the last `product_accuracy_history.jsonl` entry, and calls out field- and image-level regressions/improvements explicitly. Added classifier method (`dinov2_similarity`/`yolo_classifier`) distribution and average confidence as informational (non-scoring) context. Created the previously-missing `sources/product-test-images/README.md` documenting the validation-photo drop workflow — shipped 2026-07-13, user-directed `n` request to make algorithm tuning self-verifying.
|
||||
|
||||
### Master Data Management
|
||||
- **8.1 & 8.3 CRUD APIs and Web UI**: Created `/api/v1/master/stores` and `/api/v1/master/skus` endpoints alongside a Next.js Admin page (`/admin/master-data`) to visually manage the core reference data used by the OCR matching engine — shipped 2026-07-08.
|
||||
|
||||
@@ -136,6 +136,32 @@ get mangled. Verify in `docker logs`: "DINOv2 index loaded with N reference
|
||||
images", "Using classifier weights: <new dated file>". Full worked example:
|
||||
`plans/next-enhancements.md` task 2.1.
|
||||
|
||||
## Accuracy regression harness
|
||||
|
||||
`backend/scripts/accuracy-check-scan.mts` — mirrors the DO-flow's
|
||||
`pfm-web-app/scripts/accuracy-check.mts`. Hits the live `/api/scan-pfm` for
|
||||
every labeled image in `sources/product_manual_labels.json`, checks 3 fields
|
||||
(`no_sku`, `nama_item`, `expiry_date`) against ground truth, and splits into:
|
||||
- **Training Set** — gallery photos under `foto-kemasan-v2/` (the classifier's
|
||||
own reference images; scores here measure memorization, not generalization).
|
||||
- **Validation Set** — flat filenames dropped into
|
||||
`sources/product-test-images/` (a real held-out set; see that folder's
|
||||
`README.md` for the drop-photo → label → re-run workflow via
|
||||
`/manual-label-scan`).
|
||||
|
||||
Every run appends to `sources/product_accuracy_history.jsonl` and **auto-diffs
|
||||
against the previous run**: the printed summary shows a Δ column per field per
|
||||
split, flags field/image-level regressions and improvements, and reports
|
||||
classifier method (`dinov2_similarity`/`yolo_classifier`) distribution +
|
||||
average confidence as informational context (not scored pass/fail, since
|
||||
DINOv2's "confidence" is a raw cosine similarity, not a calibrated
|
||||
probability — see Stage 1 above). This is what makes it safe to tune
|
||||
`classify_ocr_server.py` and immediately see whether a change helped or hurt.
|
||||
|
||||
```bash
|
||||
node scripts/accuracy-check-scan.mts # from backend/
|
||||
```
|
||||
|
||||
## Operational notes
|
||||
|
||||
- **Env vars**: `CLASSIFIER_SERVER_URL`, `PIPELINE_URL` (gateway, set in compose);
|
||||
@@ -156,20 +182,18 @@ images", "Using classifier weights: <new dated file>". Full worked example:
|
||||
## Known gaps & future recommendations
|
||||
|
||||
Tracked ones (see `plans/next-enhancements.md`):
|
||||
- **§6.1–6.3 ground truth**: today's "Save Ground Truth" stores the model's own
|
||||
predictions (only the SKU is editable) and uploads get phantom
|
||||
`uploaded-<timestamp>.jpg` keys with no image persisted; §6 plans the editable
|
||||
annotation page, persisted uploads, and a scan accuracy harness.
|
||||
- **Dataset thinness**: 2–16 photos/class caps both classifiers; every new real
|
||||
photo (especially non-studio, in-warehouse shots) matters. The §6.3 harness
|
||||
should report gallery vs. uploaded-photo accuracy separately — gallery photos
|
||||
are training data, so scores on them measure memorization.
|
||||
photo (especially non-studio, in-warehouse shots) matters. The harness above
|
||||
already reports gallery (training) vs. held-out (validation) accuracy
|
||||
separately — but as of this writing `sources/product-test-images/` is empty,
|
||||
so the Validation Set is still 0 images and every published number so far is
|
||||
a training/memorization score. Dropping real photos there is the next step,
|
||||
not yet done.
|
||||
|
||||
Additional recommendations (not yet tasks — promote via `e`/`n` when wanted):
|
||||
1. **Use `extracted_sku` in match ranking.** An exact 8-digit SKU hit read off
|
||||
the label is far stronger evidence than fuzzy name similarity, yet ranking
|
||||
currently ignores it. Suggested: exact `no_sku` match pins rank 1; blend name
|
||||
similarity for the rest.
|
||||
1. ~~Use `extracted_sku` in match ranking.~~ **Done** — `product-scan.ts`'s
|
||||
`classifyAndMatchProduct` already pins rank 1 to an exact `no_sku` match
|
||||
(score forced to 1.0) before falling back to name similarity.
|
||||
2. **Fuse DINOv2 and YOLO instead of primary/fallback** (e.g. agreement boosts
|
||||
confidence; disagreement flags for review) — cheap, both already load.
|
||||
3. **"Not a known product" handling**: DINOv2 always returns *some* class; add a
|
||||
|
||||
Reference in new issue
Block a user