feat(backend): scan-product accuracy 66.2% -> 79.7% + frozen validation benchmark
Accuracy work on the 79-image product-scan validation set (user goal: 90%): - classify_ocr_server.py: 0/90/180/270-degree expiry-date search (stops at first hit, 0-degree fallback); classification decoupled onto the upright image (rotated frames regressed DINOv2 -6pts until this); cross-line date stitching; tiled full-res OCR pass (defeats the 4000px downscale that killed small inkjet dates); VL-pipeline expiry fallback with keyword-anchored anti-hallucination guard; VL text lines merged into text_lines + VL SKU retry. Visualization endpoints removed entirely (Visual/Spotting grids - unused by frontend, 3x per-scan GPU cost). - product-scan.ts: coverage-normalized OCR-evidence re-ranking of DINOv2 top-K (tuned offline: +8/-0 on top-1 misses), re-ranked class mapped to sku_master by SKU prefix; classifier timeout 90s->240s for fallback paths. - Frozen benchmark: product-test-images-fixed/ (79 renamed images) + freeze/seed/build-undetected/capture/experiment scripts; labels trimmed to the 79 validation entries (training rows kept in .bak-with-training); 5 TRAINED-ON SKUs replaced with fresh held-out photos. - manual-label-scan page: shows last batch-test AI prediction under every field by default (new /api/product-scan-results); serves the fixed folder; fixed total hydration failure via allowedDevOrigins 127.0.0.1. - Measured (all-79, zero failures): sku/name 87.3%, expiry 64.6%, overall 79.7%. Tiles/VL-evidence/VL-SKU deployed but not yet batch-measured. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gr6HH7JrdsXX8AARejQboM
This commit is contained in:
1 parent
19f1facf9b
commit
e76ccb60a6
156 files changed
+17150
-1405
No files matched your search
@@ -119,9 +119,9 @@ pipeline call with `promptLabel: "spotting"`, no layout detection).
|
||||
|
||||
| File (`pfm-web-app/public/produk-pfm/`) | What |
|
||||
|---|---|
|
||||
| `foto-kemasan-v2/<SKU or class>/…` | Reference photo dataset — 16 classes, 118 photos (2–16 each) |
|
||||
| `models/dinov2_index.pkl` | DINOv2 embeddings + metadata (rebuild after adding photos) |
|
||||
| `models/produk-pfm-classifier-26n-100e-2026-07-08.pt` / `.onnx` | Fine-tuned YOLO classifier (83.3% top-1 / 90% top-5 val on the thin dataset) |
|
||||
| `foto-kemasan-v2/<SKU or class>/…` | Reference photo dataset — 81 classes, 2,493 photos (target ~230 SKU) |
|
||||
| `models/dinov2_index.pkl` | DINOv2 embeddings + metadata (rebuild after adding photos) — currently indexes all 2,493 photos across 81 classes |
|
||||
| `models/produk-pfm-classifier-26n-100e-2026-07-14.pt` / `.onnx` | Fine-tuned YOLO classifier (85.8% top-1 / 94.4% top-5 val across all 81 classes; retrained 2026-07-14, 54m21s on an RTX 2060, up from the prior 2026-07-08 model's 83.3%/90% on only 16 classes) |
|
||||
| `index_dinov2.py` | Rebuilds the pickle index from `foto-kemasan-v2/` |
|
||||
| `train_classifier.py` | Splits 80/20 into `yolo_dataset/`, fine-tunes `yolo26n-cls.pt` (default 100 epochs, `--imgsz 224`), writes a dated checkpoint |
|
||||
|
||||
@@ -144,10 +144,16 @@ every labeled image in `sources/product_manual_labels.json`, checks 3 fields
|
||||
(`no_sku`, `nama_item`, `expiry_date`) against ground truth, and splits into:
|
||||
- **Training Set** — gallery photos under `foto-kemasan-v2/` (the classifier's
|
||||
own reference images; scores here measure memorization, not generalization).
|
||||
- **Validation Set** — flat filenames dropped into
|
||||
`sources/product-test-images/` (a real held-out set; see that folder's
|
||||
`README.md` for the drop-photo → label → re-run workflow via
|
||||
`/manual-label-scan`).
|
||||
- **Validation Set** — flat filenames, scored from the frozen
|
||||
`sources/product-test-images-fixed/` snapshot (renamed `<index> <no_sku>.<ext>`,
|
||||
built by `scripts/freeze-validation-set.mjs`) so a rerun always grades the
|
||||
same 79 images regardless of what's since been dropped into the live-intake
|
||||
`sources/product-test-images/` folder. See each folder's `README.md` — the
|
||||
live folder documents the drop-photo → label → re-run-freeze-script workflow
|
||||
via `/manual-label-scan`; the fixed folder documents the freeze/promote step
|
||||
and flags 5 SKUs (12010801, 12012504, 12130504, 13050101, 15040102) whose
|
||||
only available photo was already used to train the classifier, so their
|
||||
scores aren't a clean held-out result.
|
||||
|
||||
Every run appends to `sources/product_accuracy_history.jsonl` and **auto-diffs
|
||||
against the previous run**: the printed summary shows a Δ column per field per
|
||||
@@ -185,10 +191,9 @@ Tracked ones (see `plans/next-enhancements.md`):
|
||||
- **Dataset thinness**: 2–16 photos/class caps both classifiers; every new real
|
||||
photo (especially non-studio, in-warehouse shots) matters. The harness above
|
||||
already reports gallery (training) vs. held-out (validation) accuracy
|
||||
separately — but as of this writing `sources/product-test-images/` is empty,
|
||||
so the Validation Set is still 0 images and every published number so far is
|
||||
a training/memorization score. Dropping real photos there is the next step,
|
||||
not yet done.
|
||||
separately, and as of 2026-07-14 the Validation Set has 79 labeled images
|
||||
(74 genuinely held out, 5 flagged trained-on — see above) — the first real
|
||||
(non-zero) Validation Set numbers.
|
||||
|
||||
Additional recommendations (not yet tasks — promote via `e`/`n` when wanted):
|
||||
1. ~~Use `extracted_sku` in match ranking.~~ **Done** — `product-scan.ts`'s
|
||||
|
||||
Reference in new issue
Block a user