Files
pfm-ocr/backend/docs/scan-product.md
T
Rafhan Mazaya FathurrahmanandClaude Fable 5 e76ccb60a6 feat(backend): scan-product accuracy 66.2% -> 79.7% + frozen validation benchmark
Accuracy work on the 79-image product-scan validation set (user goal: 90%):
- classify_ocr_server.py: 0/90/180/270-degree expiry-date search (stops at
  first hit, 0-degree fallback); classification decoupled onto the upright
  image (rotated frames regressed DINOv2 -6pts until this); cross-line date
  stitching; tiled full-res OCR pass (defeats the 4000px downscale that
  killed small inkjet dates); VL-pipeline expiry fallback with
  keyword-anchored anti-hallucination guard; VL text lines merged into
  text_lines + VL SKU retry. Visualization endpoints removed entirely
  (Visual/Spotting grids - unused by frontend, 3x per-scan GPU cost).
- product-scan.ts: coverage-normalized OCR-evidence re-ranking of DINOv2
  top-K (tuned offline: +8/-0 on top-1 misses), re-ranked class mapped to
  sku_master by SKU prefix; classifier timeout 90s->240s for fallback paths.
- Frozen benchmark: product-test-images-fixed/ (79 renamed images) +
  freeze/seed/build-undetected/capture/experiment scripts; labels trimmed to
  the 79 validation entries (training rows kept in .bak-with-training);
  5 TRAINED-ON SKUs replaced with fresh held-out photos.
- manual-label-scan page: shows last batch-test AI prediction under every
  field by default (new /api/product-scan-results); serves the fixed folder;
  fixed total hydration failure via allowedDevOrigins 127.0.0.1.
- Measured (all-79, zero failures): sku/name 87.3%, expiry 64.6%, overall
  79.7%. Tiles/VL-evidence/VL-SKU deployed but not yet batch-measured.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gr6HH7JrdsXX8AARejQboM
2026-07-14 19:55:17 +07:00

13 KiB
Raw Blame History

Product Scan (scan-pfm) — How It Works

End-to-end reference for the Product/SKU scanning feature: a photo of a Primafood product package goes in; the SKU class, product name, expiry date, and a ranked SKU-master match list come out. Written 2026-07-08 against the live code. Related: plans/next-enhancements.md §2 (build history) and §6 (ground-truth roadmap); docs/feature-list.md tasks 2.1/2.3.

High-level flow

flowchart LR
    A[Browser: /scan-pfm page] -->|"POST /api/scan-pfm {image_base64}"| B[Next.js gateway<br/>pfm-web-app :3000]
    B -->|"POST :8120/classify-ocr"| C[classify_ocr_server.py<br/>FastAPI, in pipeline-api]
    C --> C1[1. DINOv2 similarity<br/>fallback: YOLO classifier]
    C --> C2[2. PaddleOCR + regex<br/>SKU / expiry / name]
    C -->|"POST localhost:8090/layout-parsing<br/>promptLabel: spotting"| D[PaddleX pipeline<br/>same container]
    B -->|"POST :8090/layout-parsing"| D
    B -->|"SELECT sku_master"| E[(Postgres)]
    B -->|Levenshtein ranking| A

Two processes live in the paddleocr-pipeline-api container, both started by scripts/serve-pipeline.sh: the PaddleX layout-parsing pipeline on :8090 (shared with the DO flow; VL recognition goes out to the vLLM server on :8118) and config/classify_ocr_server.py on :8120 (product scan only). The gateway reaches them via Docker DNS (CLASSIFIER_SERVER_URL, PIPELINE_URL in root docker-compose.yml:87-88); nginx (:8000) proxies /scan-pfm to the Next.js app.

Request walkthrough

  1. Page (pfm-web-app/src/app/scan-pfm/page.tsx, desktop-only test UI): pick a sample from the gallery (GET /api/produk-pfm) or upload/rotate a photo (rotation is done client-side on a canvas), then send it as a base64 data-URL.
  2. Gateway (api/scan-pfm/route.ts):
    • forwards {image_base64} to the classifier server (/classify-ocr);
    • separately calls the layout-parsing pipeline with useLayoutDetection: true for the Visual Grid tab's output images (failure here is non-fatal — logged, layoutParsingResult returns null);
    • loads the full sku_master table and ranks every SKU by Levenshtein similarity between nama_item and the classifier's top1_name (lowercased, alphanumerics only). Top 5 with score > 0.1 are returned; rank 1 gets isBestMatch: true. Note: ocr.extracted_sku and ocr.extracted_product_name are read but not used in this ranking — see Future recommendations.
  3. Classifier server (config/classify_ocr_server.py) does classification, OCR extraction, and visualization — detailed below — and returns {classification, ocr}.
  4. Page renders four tabs: Summary (classification card + top-5 override "Use" buttons + OCR fields + SKU matches), Visual Grid, Spotting Grid, Raw Response (JSON). "Save Ground Truth" posts to /api/manual-label-scan.

Stage 1 — classification (which product is this?)

Primary: DINOv2 similarity search (method: "dinov2_similarity"). At startup the server loads dinov2_vits14 from torch.hub (network fetch on first run) plus models/dinov2_index.pkl — precomputed L2-normalized 384-dim embeddings of all 118 reference photos across 16 SKU class folders. Per request: embed the query image (resize 224², ImageNet normalization), dot-product against all reference embeddings (= cosine similarity), then aggregate per class = max similarity of any reference photo in that class. Classes sorted by similarity become all_probabilities. Caveat: these "confidences" are cosine similarities, not probabilities — they don't sum to 1 and are typically all high (0.4–0.9); compare relatively, not against an absolute threshold.

Fallback: YOLO classifier (method: "yolo_classifier") — only when DINOv2 is unavailable (no index/model) or throws. A fine-tuned yolo26n-cls checkpoint; its all_probabilities are real softmax probabilities. Weights are auto-discovered: CLASSIFIER_MODEL_PATH env wins; otherwise the newest produk-pfm-classifier-26n-*e-*.pt in models/ by (date-in-filename, mtime) — so retraining just drops a new dated file, no config change.

If both are unavailable, classification carries an error field instead.

Stage 2 — OCR extraction (SKU, expiry date, product name)

PaddleOCR (lang='en', textline orientation on) produces rec_texts lines + rec_polys boxes. Three extractors run over the lines:

  • SKU (extract_sku): first 8-digit number anywhere; else first 7–9 digit number. (Primafood SKUs are 8 digits, printed near the label top.)
  • Expiry date (extract_expired_date): each line is first noise-cleaned (clean_date_line: 1)→0, ()→0, B8/8B/88→BB before digits, o→0, I/l/|→1, S→5, Z→2, B→8 when digit-flanked, plus 012/112 month-misread repairs), then a 6-level priority cascade runs: (1) BB/EXP-keyword line with compact DDMMYYYY; (2) keyword line with spaced DD MM YYYY; (3) keyword + 6–8 digit run; (3.5) keyword line, lenient noisy match; (4) any line spaced date; (5) any line compact DDMMYYYY — skipping lines that look like a SKU-on-product-name; (6) legacy formats (slashes, 05 MAR 2027). Recognized keywords: EXP, EXPIRED, TGL, EXPIRY, BBD, BEST BEFORE, BB, BAIK DIGUNAKAN. Output normalized to DD/MM/YYYY.
  • Product name (extract_product_name): longest line containing a brand/ product keyword (FIESTA, CHAMP, OKEY, AKUMO, ASIMO, NUGGET, SOSIS, …) after stripping SKU digits and date fragments; falls back to the classifier's top1_name, then the longest non-numeric line, then "Unknown Product".

Visualization artifacts built server-side: vis_image_base64 (all OCR boxes drawn teal TEXT, the expiry line amber EXP, on the orientation-corrected image so boxes align), expired_date_crop_base64 (padded crop of the expiry line for eyeball verification — find_expired_crop_index prefers the box whose digits actually contain the date), and spotting_image_base64 (a second pipeline call with promptLabel: "spotting", no layout detection).

Endpoint reference

Endpoint Where Purpose
POST /api/scan-pfm gateway Main scan. Body {image_base64} (data-URL ok). Returns {classification, ocr, possibleMatches[], layoutParsingResult}
POST http://paddleocr-pipeline-api:8120/classify-ocr classifier server Internal. Body {image_base64}. Returns {classification: {top1_name, top1_confidence, all_probabilities[], method}, ocr: {text_lines[], extracted_product_name, extracted_sku, extracted_expired_date, expired_line_index, expired_source_line, expired_date_crop_base64, vis_image_base64, spotting_image_base64}}
GET /api/produk-pfm gateway Gallery: SKU folders under public/produk-pfm/foto-kemasan-v2/ with image + thumb URLs
GET/POST /api/manual-label-scan gateway Ground-truth read/upsert to sources/product_manual_labels.json (host-visible via the ./backend/sources:/sources mount)
POST :8090/layout-parsing pipeline Shared PaddleX pipeline; used here for Visual Grid images and (with promptLabel: "spotting") the Spotting Grid
/scan-pfm nginx :8000 Proxies the page to Next.js :3000

possibleMatches[] items: {no_sku, nama_item, score, yoloSimilarity, isBestMatch} — score currently equals yoloSimilarity (name-vs-name Levenshtein, 0..1).

Model artifacts & retraining

File (pfm-web-app/public/produk-pfm/) What
foto-kemasan-v2/<SKU or class>/… Reference photo dataset — 81 classes, 2,493 photos (target ~230 SKU)
models/dinov2_index.pkl DINOv2 embeddings + metadata (rebuild after adding photos) — currently indexes all 2,493 photos across 81 classes
models/produk-pfm-classifier-26n-100e-2026-07-14.pt / .onnx Fine-tuned YOLO classifier (85.8% top-1 / 94.4% top-5 val across all 81 classes; retrained 2026-07-14, 54m21s on an RTX 2060, up from the prior 2026-07-08 model's 83.3%/90% on only 16 classes)
index_dinov2.py Rebuilds the pickle index from foto-kemasan-v2/
train_classifier.py Splits 80/20 into yolo_dataset/, fine-tunes yolo26n-cls.pt (default 100 epochs, --imgsz 224), writes a dated checkpoint

Retraining procedure (Windows host — bare-metal doesn't work here, paddlepaddle-gpu wheels are Linux-only): add photos to foto-kemasan-v2/, docker compose build pipeline-api from the repo root, run a one-off docker run --gpus all from that image with models/ mounted writable (the live service mounts it :ro), run index_dinov2.py then train_classifier.py train --imgsz 224, then docker compose restart pipeline-api. From Git Bash prefix MSYS_NO_PATHCONV=1 or /app/... arguments get mangled. Verify in docker logs: "DINOv2 index loaded with N reference images", "Using classifier weights: ". Full worked example: plans/next-enhancements.md task 2.1.

Accuracy regression harness

backend/scripts/accuracy-check-scan.mts — mirrors the DO-flow's pfm-web-app/scripts/accuracy-check.mts. Hits the live /api/scan-pfm for every labeled image in sources/product_manual_labels.json, checks 3 fields (no_sku, nama_item, expiry_date) against ground truth, and splits into:

  • Training Set — gallery photos under foto-kemasan-v2/ (the classifier's own reference images; scores here measure memorization, not generalization).
  • Validation Set — flat filenames, scored from the frozen sources/product-test-images-fixed/ snapshot (renamed <index> <no_sku>.<ext>, built by scripts/freeze-validation-set.mjs) so a rerun always grades the same 79 images regardless of what's since been dropped into the live-intake sources/product-test-images/ folder. See each folder's README.md — the live folder documents the drop-photo → label → re-run-freeze-script workflow via /manual-label-scan; the fixed folder documents the freeze/promote step and flags 5 SKUs (12010801, 12012504, 12130504, 13050101, 15040102) whose only available photo was already used to train the classifier, so their scores aren't a clean held-out result.

Every run appends to sources/product_accuracy_history.jsonl and auto-diffs against the previous run: the printed summary shows a Δ column per field per split, flags field/image-level regressions and improvements, and reports classifier method (dinov2_similarity/yolo_classifier) distribution + average confidence as informational context (not scored pass/fail, since DINOv2's "confidence" is a raw cosine similarity, not a calibrated probability — see Stage 1 above). This is what makes it safe to tune classify_ocr_server.py and immediately see whether a change helped or hurt.

node scripts/accuracy-check-scan.mts   # from backend/

Operational notes

  • Env vars: CLASSIFIER_SERVER_URL, PIPELINE_URL (gateway, set in compose); CLASSIFIER_MODELS_DIR, CLASSIFIER_MODEL_PATH (classifier server overrides). The gateway's in-code default PIPELINE_URL (localhost:7871) is stale — the compose env always overrides it in Docker.
  • Startup order/health: the classifier server loads DINOv2 (torch.hub → needs network/cache), YOLO, and PaddleOCR at import time; until done, :8120 refuses connections and /api/scan-pfm 500s. No healthcheck exists yet (plan task 4.2 / 1.6).
  • GPU: DINOv2 + YOLO + PaddleOCR share the container/GPU with the PaddleX pipeline; all are small (ViT-S/14, nano YOLO) next to the vLLM server's footprint, but they do add VRAM on the same PIPELINE_DEVICE.
  • Failure isolation: layout-vis and spotting calls are best-effort (null/absent on failure); classification and OCR errors surface as error fields inside their sections rather than failing the whole scan.

Known gaps & future recommendations

Tracked ones (see plans/next-enhancements.md):

  • Dataset thinness: 2–16 photos/class caps both classifiers; every new real photo (especially non-studio, in-warehouse shots) matters. The harness above already reports gallery (training) vs. held-out (validation) accuracy separately, and as of 2026-07-14 the Validation Set has 79 labeled images (74 genuinely held out, 5 flagged trained-on — see above) — the first real (non-zero) Validation Set numbers.

Additional recommendations (not yet tasks — promote via e/n when wanted):

  1. Use extracted_sku in match ranking. Done — product-scan.ts's classifyAndMatchProduct already pins rank 1 to an exact no_sku match (score forced to 1.0) before falling back to name similarity.
  2. Fuse DINOv2 and YOLO instead of primary/fallback (e.g. agreement boosts confidence; disagreement flags for review) — cheap, both already load.
  3. "Not a known product" handling: DINOv2 always returns some class; add a minimum-similarity threshold below which the response says unknown rather than confidently misclassifying a foreign package.
  4. Pin the DINOv2 backbone offline (vendor the weights or pre-bake the torch.hub cache into the image) — startup currently depends on an internet fetch on cold cache, bad for on-prem deploys.
  5. Batch/lot number extraction — explicitly out of scope so far (plan §2 note); if requested, follow the expiry-date regex-cascade pattern.
  6. Mobile: no web mobile page by design (task 2.2 cancelled) — real mobile scanning should go through the Flutter app calling POST /api/scan-pfm (would need an authenticated /api/v1 variant; the classic route has no auth).