feat(backend): diff-vs-previous-run reporting for product-scan accuracy harness

Ports the DO-harness's auto-diff-vs-previous-run reporting into
accuracy-check-scan.mts: prints a per-field, per-split (Training/
Validation) delta against the last product_accuracy_history.jsonl entry
and calls out regressions/improvements explicitly, plus classifier
method distribution and average confidence as informational context.

Also adds the real held-out validation photo set into
sources/product-test-images/ (75 photos, one per current SKU class) with
its README documenting the drop-photo -> label -> re-run workflow, so the
harness's Validation Set split actually has images to score against.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xsxk4ZkDQVVaLUcixDcqb5
This commit is contained in:
Rafhan Mazaya FathurrahmanandClaude Sonnet 5 committed 2026-07-14 08:34:02 +07:00
1 parent dc0dd81318
commit 3a17c28758
78 files changed
+306 -55

No files matched your search

+35 -11
View File
@@ -136,6 +136,32 @@ get mangled. Verify in `docker logs`: "DINOv2 index loaded with N reference
images", "Using classifier weights: <new dated file>". Full worked example:
`plans/next-enhancements.md` task 2.1.
## Accuracy regression harness
`backend/scripts/accuracy-check-scan.mts` — mirrors the DO-flow's
`pfm-web-app/scripts/accuracy-check.mts`. Hits the live `/api/scan-pfm` for
every labeled image in `sources/product_manual_labels.json`, checks 3 fields
(`no_sku`, `nama_item`, `expiry_date`) against ground truth, and splits into:
- **Training Set** — gallery photos under `foto-kemasan-v2/` (the classifier's
own reference images; scores here measure memorization, not generalization).
- **Validation Set** — flat filenames dropped into
`sources/product-test-images/` (a real held-out set; see that folder's
`README.md` for the drop-photo → label → re-run workflow via
`/manual-label-scan`).
Every run appends to `sources/product_accuracy_history.jsonl` and **auto-diffs
against the previous run**: the printed summary shows a Δ column per field per
split, flags field/image-level regressions and improvements, and reports
classifier method (`dinov2_similarity`/`yolo_classifier`) distribution +
average confidence as informational context (not scored pass/fail, since
DINOv2's "confidence" is a raw cosine similarity, not a calibrated
probability — see Stage 1 above). This is what makes it safe to tune
`classify_ocr_server.py` and immediately see whether a change helped or hurt.
```bash
node scripts/accuracy-check-scan.mts # from backend/
```
## Operational notes
- **Env vars**: `CLASSIFIER_SERVER_URL`, `PIPELINE_URL` (gateway, set in compose);
@@ -156,20 +182,18 @@ images", "Using classifier weights: <new dated file>". Full worked example:
## Known gaps & future recommendations
Tracked ones (see `plans/next-enhancements.md`):
- **§6.1–6.3 ground truth**: today's "Save Ground Truth" stores the model's own
predictions (only the SKU is editable) and uploads get phantom
`uploaded-<timestamp>.jpg` keys with no image persisted; §6 plans the editable
annotation page, persisted uploads, and a scan accuracy harness.
- **Dataset thinness**: 2–16 photos/class caps both classifiers; every new real
photo (especially non-studio, in-warehouse shots) matters. The §6.3 harness
should report gallery vs. uploaded-photo accuracy separately — gallery photos
are training data, so scores on them measure memorization.
photo (especially non-studio, in-warehouse shots) matters. The harness above
already reports gallery (training) vs. held-out (validation) accuracy
separately — but as of this writing `sources/product-test-images/` is empty,
so the Validation Set is still 0 images and every published number so far is
a training/memorization score. Dropping real photos there is the next step,
not yet done.
Additional recommendations (not yet tasks — promote via `e`/`n` when wanted):
1. **Use `extracted_sku` in match ranking.** An exact 8-digit SKU hit read off
the label is far stronger evidence than fuzzy name similarity, yet ranking
currently ignores it. Suggested: exact `no_sku` match pins rank 1; blend name
similarity for the rest.
1. ~~Use `extracted_sku` in match ranking.~~ **Done** — `product-scan.ts`'s
`classifyAndMatchProduct` already pins rank 1 to an exact `no_sku` match
(score forced to 1.0) before falling back to name similarity.
2. **Fuse DINOv2 and YOLO instead of primary/fallback** (e.g. agreement boosts
confidence; disagreement flags for review) — cheap, both already load.
3. **"Not a known product" handling**: DINOv2 always returns *some* class; add a