feat(backend): scan-product accuracy 66.2% -> 79.7% + frozen validation benchmark
Accuracy work on the 79-image product-scan validation set (user goal: 90%): - classify_ocr_server.py: 0/90/180/270-degree expiry-date search (stops at first hit, 0-degree fallback); classification decoupled onto the upright image (rotated frames regressed DINOv2 -6pts until this); cross-line date stitching; tiled full-res OCR pass (defeats the 4000px downscale that killed small inkjet dates); VL-pipeline expiry fallback with keyword-anchored anti-hallucination guard; VL text lines merged into text_lines + VL SKU retry. Visualization endpoints removed entirely (Visual/Spotting grids - unused by frontend, 3x per-scan GPU cost). - product-scan.ts: coverage-normalized OCR-evidence re-ranking of DINOv2 top-K (tuned offline: +8/-0 on top-1 misses), re-ranked class mapped to sku_master by SKU prefix; classifier timeout 90s->240s for fallback paths. - Frozen benchmark: product-test-images-fixed/ (79 renamed images) + freeze/seed/build-undetected/capture/experiment scripts; labels trimmed to the 79 validation entries (training rows kept in .bak-with-training); 5 TRAINED-ON SKUs replaced with fresh held-out photos. - manual-label-scan page: shows last batch-test AI prediction under every field by default (new /api/product-scan-results); serves the fixed folder; fixed total hydration failure via allowedDevOrigins 127.0.0.1. - Measured (all-79, zero failures): sku/name 87.3%, expiry 64.6%, overall 79.7%. Tiles/VL-evidence/VL-SKU deployed but not yet batch-measured. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gr6HH7JrdsXX8AARejQboM
No files matched your search
|
After Width: | Height: | Size: 3.4 MiB |
|
After Width: | Height: | Size: 140 KiB |
|
After Width: | Height: | Size: 1.3 MiB |
|
After Width: | Height: | Size: 257 KiB |
|
After Width: | Height: | Size: 3.7 MiB |
|
After Width: | Height: | Size: 229 KiB |
|
After Width: | Height: | Size: 254 KiB |
|
After Width: | Height: | Size: 268 KiB |
|
After Width: | Height: | Size: 4.2 MiB |
|
After Width: | Height: | Size: 184 KiB |
|
After Width: | Height: | Size: 220 KiB |
|
After Width: | Height: | Size: 2.3 MiB |
|
After Width: | Height: | Size: 227 KiB |
|
After Width: | Height: | Size: 170 KiB |
|
After Width: | Height: | Size: 217 KiB |
|
After Width: | Height: | Size: 217 KiB |
|
After Width: | Height: | Size: 201 KiB |
|
After Width: | Height: | Size: 3.9 MiB |
|
After Width: | Height: | Size: 239 KiB |
|
After Width: | Height: | Size: 194 KiB |
|
After Width: | Height: | Size: 200 KiB |
|
After Width: | Height: | Size: 207 KiB |
|
After Width: | Height: | Size: 3.1 MiB |
|
After Width: | Height: | Size: 275 KiB |
|
After Width: | Height: | Size: 231 KiB |
|
After Width: | Height: | Size: 304 KiB |
|
After Width: | Height: | Size: 3.3 MiB |
|
After Width: | Height: | Size: 278 KiB |
|
After Width: | Height: | Size: 2.0 MiB |
|
After Width: | Height: | Size: 178 KiB |
|
After Width: | Height: | Size: 293 KiB |
|
After Width: | Height: | Size: 284 KiB |
|
After Width: | Height: | Size: 223 KiB |
|
After Width: | Height: | Size: 164 KiB |
|
After Width: | Height: | Size: 3.4 MiB |
|
After Width: | Height: | Size: 3.4 MiB |
|
After Width: | Height: | Size: 1.8 MiB |
|
After Width: | Height: | Size: 225 KiB |
|
After Width: | Height: | Size: 4.5 MiB |
|
After Width: | Height: | Size: 232 KiB |
|
After Width: | Height: | Size: 3.6 MiB |
|
After Width: | Height: | Size: 235 KiB |
|
After Width: | Height: | Size: 2.4 MiB |
|
After Width: | Height: | Size: 149 KiB |
|
After Width: | Height: | Size: 285 KiB |
|
After Width: | Height: | Size: 3.6 MiB |
|
After Width: | Height: | Size: 3.8 MiB |
|
After Width: | Height: | Size: 3.7 MiB |
|
After Width: | Height: | Size: 106 KiB |
|
After Width: | Height: | Size: 2.5 MiB |
|
After Width: | Height: | Size: 175 KiB |
|
After Width: | Height: | Size: 2.9 MiB |
|
After Width: | Height: | Size: 208 KiB |
|
After Width: | Height: | Size: 143 KiB |
|
After Width: | Height: | Size: 251 KiB |
|
After Width: | Height: | Size: 2.4 MiB |
|
After Width: | Height: | Size: 176 KiB |
|
After Width: | Height: | Size: 230 KiB |
|
After Width: | Height: | Size: 2.3 MiB |
|
After Width: | Height: | Size: 302 KiB |
|
After Width: | Height: | Size: 114 KiB |
|
After Width: | Height: | Size: 217 KiB |
|
After Width: | Height: | Size: 208 KiB |
|
After Width: | Height: | Size: 129 KiB |
|
After Width: | Height: | Size: 288 KiB |
|
After Width: | Height: | Size: 4.1 MiB |
|
After Width: | Height: | Size: 2.5 MiB |
|
After Width: | Height: | Size: 232 KiB |
|
After Width: | Height: | Size: 286 KiB |
|
After Width: | Height: | Size: 3.8 MiB |
|
After Width: | Height: | Size: 171 KiB |
|
After Width: | Height: | Size: 265 KiB |
|
After Width: | Height: | Size: 263 KiB |
|
After Width: | Height: | Size: 261 KiB |
|
After Width: | Height: | Size: 227 KiB |
|
After Width: | Height: | Size: 1.8 MiB |
|
After Width: | Height: | Size: 222 KiB |
|
After Width: | Height: | Size: 215 KiB |
|
After Width: | Height: | Size: 3.1 MiB |
@@ -0,0 +1,29 @@
|
||||
# Product-scan validation images — frozen benchmark set
|
||||
|
||||
This is the **actual Validation Set the accuracy harness scores**
|
||||
(`accuracy-check-scan.mts`'s `getImagePath()` points here, not at
|
||||
`../product-test-images/`). It exists so re-running the harness always grades
|
||||
the exact same images — the live-intake folder can keep growing from new
|
||||
`/manual-label-scan` drops without silently shifting the benchmark underfoot.
|
||||
|
||||
**Naming**: each file is `<index> <no_sku>.<ext>` (e.g. `1 11110059.jpeg`),
|
||||
where `<index>` is just this file's stable position in the set — it carries
|
||||
no other meaning. `product_manual_labels.json`'s ground-truth entries for
|
||||
these images use this same filename.
|
||||
|
||||
**Do not hand-edit this folder.** It's fully generated by
|
||||
`node scripts/freeze-validation-set.mjs` (run from `backend/`), which copies
|
||||
every flat (non-training) entry out of `product_manual_labels.json` from
|
||||
`../product-test-images/`, renames it, and rewrites those entries'
|
||||
`filename` fields to match. To add a new SKU/photo to the benchmark:
|
||||
label it in the live-intake folder first (see that folder's README), then
|
||||
re-run the freeze script.
|
||||
|
||||
**79 images as of 2026-07-14.** 5 of them (SKUs 12010801, 12012504, 12130504,
|
||||
13050101, 15040102) are flagged `TRAINED-ON` in `product_manual_labels.json`'s
|
||||
`notes` field — their only available source photo (from an external
|
||||
`research-sam3` segmentation project) was already used to train the
|
||||
classifier (as a SAM3 crop + augmentations), so they are **not** a clean
|
||||
held-out test. Their per-image scores will read as memorization, not real
|
||||
generalization, until a fresh, never-trained-on photo is dropped for those
|
||||
SKUs. The other 74 are genuinely held out.
|
||||
|
After Width: | Height: | Size: 3.4 MiB |
|
After Width: | Height: | Size: 1.3 MiB |
|
After Width: | Height: | Size: 254 KiB |
|
After Width: | Height: | Size: 268 KiB |
|
After Width: | Height: | Size: 4.2 MiB |
|
After Width: | Height: | Size: 2.3 MiB |
|
After Width: | Height: | Size: 227 KiB |
|
After Width: | Height: | Size: 170 KiB |
|
After Width: | Height: | Size: 217 KiB |
|
After Width: | Height: | Size: 3.9 MiB |
|
After Width: | Height: | Size: 194 KiB |
|
After Width: | Height: | Size: 200 KiB |
|
After Width: | Height: | Size: 231 KiB |
|
After Width: | Height: | Size: 304 KiB |
|
After Width: | Height: | Size: 178 KiB |
|
After Width: | Height: | Size: 293 KiB |
|
After Width: | Height: | Size: 284 KiB |
|
After Width: | Height: | Size: 3.4 MiB |
|
After Width: | Height: | Size: 1.8 MiB |
|
After Width: | Height: | Size: 4.5 MiB |