Files
pfm-ocr/backend/sources/product-test-images/README.md
T
Rafhan Mazaya FathurrahmanandClaude Sonnet 5 3a17c28758 feat(backend): diff-vs-previous-run reporting for product-scan accuracy harness
Ports the DO-harness's auto-diff-vs-previous-run reporting into
accuracy-check-scan.mts: prints a per-field, per-split (Training/
Validation) delta against the last product_accuracy_history.jsonl entry
and calls out regressions/improvements explicitly, plus classifier
method distribution and average confidence as informational context.

Also adds the real held-out validation photo set into
sources/product-test-images/ (75 photos, one per current SKU class) with
its README documenting the drop-photo -> label -> re-run workflow, so the
harness's Validation Set split actually has images to score against.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xsxk4ZkDQVVaLUcixDcqb5
2026-07-14 08:34:02 +07:00

1.1 KiB

Product-scan validation images

This folder is the validation/test set for the product-scan accuracy harness (backend/scripts/accuracy-check-scan.mts) — real-world photos that are not part of the classifier's reference dataset, so scoring against them measures actual accuracy instead of memorization.

Workflow:

  1. Drop a photo here directly (flat, no subfolders — a filename with no / is what marks an image as "validation" instead of "training").
  2. Label it via the /manual-label-scan page (correct no_sku, nama_item, expiry_date by hand — don't just accept the AI-scan prefill, that would make the ground truth equal to the model's own prediction).
  3. Run node scripts/accuracy-check-scan.mts from backend/ — the photo now scores under "Validation Set", separate from "Training Set".

This is not where new training photos go. To improve the classifier itself (DINOv2 index / YOLO fine-tune), add photos to pfm-web-app/public/produk-pfm/foto-kemasan-v2/<SKU folder>/ instead, then reindex/retrain per docs/scan-product.md's "Model artifacts & retraining" section.