Files
pfm-ocr/backend/scripts/freeze-validation-set.mjs
T
Rafhan Mazaya FathurrahmanandClaude Fable 5 e76ccb60a6 feat(backend): scan-product accuracy 66.2% -> 79.7% + frozen validation benchmark
Accuracy work on the 79-image product-scan validation set (user goal: 90%):
- classify_ocr_server.py: 0/90/180/270-degree expiry-date search (stops at
  first hit, 0-degree fallback); classification decoupled onto the upright
  image (rotated frames regressed DINOv2 -6pts until this); cross-line date
  stitching; tiled full-res OCR pass (defeats the 4000px downscale that
  killed small inkjet dates); VL-pipeline expiry fallback with
  keyword-anchored anti-hallucination guard; VL text lines merged into
  text_lines + VL SKU retry. Visualization endpoints removed entirely
  (Visual/Spotting grids - unused by frontend, 3x per-scan GPU cost).
- product-scan.ts: coverage-normalized OCR-evidence re-ranking of DINOv2
  top-K (tuned offline: +8/-0 on top-1 misses), re-ranked class mapped to
  sku_master by SKU prefix; classifier timeout 90s->240s for fallback paths.
- Frozen benchmark: product-test-images-fixed/ (79 renamed images) +
  freeze/seed/build-undetected/capture/experiment scripts; labels trimmed to
  the 79 validation entries (training rows kept in .bak-with-training);
  5 TRAINED-ON SKUs replaced with fresh held-out photos.
- manual-label-scan page: shows last batch-test AI prediction under every
  field by default (new /api/product-scan-results); serves the fixed folder;
  fixed total hydration failure via allowedDevOrigins 127.0.0.1.
- Measured (all-79, zero failures): sku/name 87.3%, expiry 64.6%, overall
  79.7%. Tiles/VL-evidence/VL-SKU deployed but not yet batch-measured.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gr6HH7JrdsXX8AARejQboM
2026-07-14 19:55:17 +07:00

48 lines
2.1 KiB
JavaScript

// One-off script: freeze the current 74-image product-scan Validation Set into
// a dedicated, stable folder (sources/product-test-images-fixed/) so re-running
// the accuracy harness always scores the exact same images, independent of
// whatever new photos get dropped into the live-intake folder
// (sources/product-test-images/, still fed by the /manual-label-scan page).
// Renames each image "<index> <no_sku>.<ext>" (index = its stable position,
// 1-based) and updates product_manual_labels.json's flat-filename entries to
// match. Run once from backend/: node scripts/freeze-validation-set.mjs
import fs from "fs";
import path from "path";
const LIVE_DIR = path.join("sources", "product-test-images");
const FIXED_DIR = path.join("sources", "product-test-images-fixed");
const LABELS_PATH = path.join("sources", "product_manual_labels.json");
const labels = JSON.parse(fs.readFileSync(LABELS_PATH, "utf8"));
const flatEntries = labels.filter((l) => !l.filename.includes("/"));
if (!fs.existsSync(FIXED_DIR)) fs.mkdirSync(FIXED_DIR, { recursive: true });
// Resolve by SKU prefix (not entry.filename directly) so this script is
// idempotent/rerunnable even after a previous run already renamed
// entry.filename to "<index> <sku>.<ext>" — the live-intake folder always
// keeps its original "<sku> <product>__<camera-filename>.<ext>" names.
const liveFiles = fs.readdirSync(LIVE_DIR);
function findSourceFile(no_sku) {
const match = liveFiles.find((f) => f.startsWith(`${no_sku} `) || f.startsWith(`${no_sku}__`));
if (!match) return null;
return path.join(LIVE_DIR, match);
}
let copied = 0;
flatEntries.forEach((entry, i) => {
const index = i + 1;
const srcPath = findSourceFile(entry.no_sku);
if (!srcPath) {
throw new Error(`Missing source image for ${entry.no_sku} in ${LIVE_DIR}`);
}
const ext = path.extname(srcPath);
const newFilename = `${index} ${entry.no_sku}${ext}`;
fs.copyFileSync(srcPath, path.join(FIXED_DIR, newFilename));
entry.filename = newFilename;
copied++;
});
fs.writeFileSync(LABELS_PATH, JSON.stringify(labels, null, 2), "utf8");
console.log(`Copied ${copied} images into ${FIXED_DIR} and updated ${LABELS_PATH}.`);