feat(backend): scan-product accuracy 66.2% -> 79.7% + frozen validation benchmark
Accuracy work on the 79-image product-scan validation set (user goal: 90%): - classify_ocr_server.py: 0/90/180/270-degree expiry-date search (stops at first hit, 0-degree fallback); classification decoupled onto the upright image (rotated frames regressed DINOv2 -6pts until this); cross-line date stitching; tiled full-res OCR pass (defeats the 4000px downscale that killed small inkjet dates); VL-pipeline expiry fallback with keyword-anchored anti-hallucination guard; VL text lines merged into text_lines + VL SKU retry. Visualization endpoints removed entirely (Visual/Spotting grids - unused by frontend, 3x per-scan GPU cost). - product-scan.ts: coverage-normalized OCR-evidence re-ranking of DINOv2 top-K (tuned offline: +8/-0 on top-1 misses), re-ranked class mapped to sku_master by SKU prefix; classifier timeout 90s->240s for fallback paths. - Frozen benchmark: product-test-images-fixed/ (79 renamed images) + freeze/seed/build-undetected/capture/experiment scripts; labels trimmed to the 79 validation entries (training rows kept in .bak-with-training); 5 TRAINED-ON SKUs replaced with fresh held-out photos. - manual-label-scan page: shows last batch-test AI prediction under every field by default (new /api/product-scan-results); serves the fixed folder; fixed total hydration failure via allowedDevOrigins 127.0.0.1. - Measured (all-79, zero failures): sku/name 87.3%, expiry 64.6%, overall 79.7%. Tiles/VL-evidence/VL-SKU deployed but not yet batch-measured. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gr6HH7JrdsXX8AARejQboM
This commit is contained in:
1 parent
19f1facf9b
commit
e76ccb60a6
156 files changed
+17150
-1405
No files matched your search
@@ -1,18 +1,22 @@
|
||||
# Product-scan validation images
|
||||
# Product-scan validation images — live intake (staging)
|
||||
|
||||
This folder is the **validation/test set** for the product-scan accuracy
|
||||
harness (`backend/scripts/accuracy-check-scan.mts`) — real-world photos that
|
||||
are *not* part of the classifier's reference dataset, so scoring against them
|
||||
measures actual accuracy instead of memorization.
|
||||
This folder is the **live-intake / staging area** for the product-scan
|
||||
validation set. It's still where the `/manual-label-scan` page saves new
|
||||
photo drops, and still what real-world photos get dropped into by hand — but
|
||||
it is **no longer what the accuracy harness scores**. That's
|
||||
`../product-test-images-fixed/` (a frozen, sequentially-renamed snapshot) —
|
||||
see that folder's README for why the split exists.
|
||||
|
||||
**Workflow:**
|
||||
**Workflow (adding a new SKU or photo):**
|
||||
1. Drop a photo here directly (flat, no subfolders — a filename with no `/`
|
||||
is what marks an image as "validation" instead of "training").
|
||||
2. Label it via the `/manual-label-scan` page (correct `no_sku`, `nama_item`,
|
||||
`expiry_date` by hand — don't just accept the AI-scan prefill, that would
|
||||
make the ground truth equal to the model's own prediction).
|
||||
3. Run `node scripts/accuracy-check-scan.mts` from `backend/` — the photo now
|
||||
scores under "Validation Set", separate from "Training Set".
|
||||
3. Re-run `node scripts/freeze-validation-set.mjs` from `backend/` to promote
|
||||
the new photo into `product-test-images-fixed/` (renamed to
|
||||
`<index> <no_sku>.<ext>`) so it actually gets scored on the next
|
||||
`accuracy-check-scan.mts` run.
|
||||
|
||||
**This is not where new training photos go.** To improve the classifier
|
||||
itself (DINOv2 index / YOLO fine-tune), add photos to
|
||||
|
||||
Reference in new issue
Block a user