# Product Scan (scan-pfm) — How It Works End-to-end reference for the Product/SKU scanning feature: a photo of a Primafood product package goes in; the SKU class, product name, expiry date, and a ranked SKU-master match list come out. Written 2026-07-08 against the live code. Related: `plans/next-enhancements.md` §2 (build history) and §6 (ground-truth roadmap); `docs/feature-list.md` tasks 2.1/2.3. ## High-level flow ```mermaid flowchart LR A[Browser: /scan-pfm page] -->|"POST /api/scan-pfm {image_base64}"| B[Next.js gateway
pfm-web-app :3000] B -->|"POST :8120/classify-ocr"| C[classify_ocr_server.py
FastAPI, in pipeline-api] C --> C1[1. DINOv2 similarity
fallback: YOLO classifier] C --> C2[2. PaddleOCR + regex
SKU / expiry / name] C -->|"POST localhost:8090/layout-parsing
promptLabel: spotting"| D[PaddleX pipeline
same container] B -->|"POST :8090/layout-parsing"| D B -->|"SELECT sku_master"| E[(Postgres)] B -->|Levenshtein ranking| A ``` Two processes live in the `paddleocr-pipeline-api` container, both started by `scripts/serve-pipeline.sh`: the PaddleX layout-parsing pipeline on **:8090** (shared with the DO flow; VL recognition goes out to the vLLM server on :8118) and `config/classify_ocr_server.py` on **:8120** (product scan only). The gateway reaches them via Docker DNS (`CLASSIFIER_SERVER_URL`, `PIPELINE_URL` in root `docker-compose.yml:87-88`); nginx (:8000) proxies `/scan-pfm` to the Next.js app. ## Request walkthrough 1. **Page** (`pfm-web-app/src/app/scan-pfm/page.tsx`, desktop-only test UI): pick a sample from the gallery (`GET /api/produk-pfm`) or upload/rotate a photo (rotation is done client-side on a canvas), then send it as a base64 data-URL. 2. **Gateway** (`api/scan-pfm/route.ts`): - forwards `{image_base64}` to the classifier server (`/classify-ocr`); - separately calls the layout-parsing pipeline with `useLayoutDetection: true` for the Visual Grid tab's output images (failure here is non-fatal — logged, `layoutParsingResult` returns `null`); - loads the full `sku_master` table and ranks every SKU by **Levenshtein similarity between `nama_item` and the classifier's `top1_name`** (lowercased, alphanumerics only). Top 5 with score > 0.1 are returned; rank 1 gets `isBestMatch: true`. Note: `ocr.extracted_sku` and `ocr.extracted_product_name` are read but **not used** in this ranking — see Future recommendations. 3. **Classifier server** (`config/classify_ocr_server.py`) does classification, OCR extraction, and visualization — detailed below — and returns `{classification, ocr}`. 4. **Page renders** four tabs: Summary (classification card + top-5 override "Use" buttons + OCR fields + SKU matches), Visual Grid, Spotting Grid, Raw Response (JSON). "Save Ground Truth" posts to `/api/manual-label-scan`. ## Stage 1 — classification (which product is this?) **Primary: DINOv2 similarity search** (`method: "dinov2_similarity"`). At startup the server loads `dinov2_vits14` **from `torch.hub` (network fetch on first run)** plus `models/dinov2_index.pkl` — precomputed L2-normalized 384-dim embeddings of all 118 reference photos across 16 SKU class folders. Per request: embed the query image (resize 224², ImageNet normalization), dot-product against all reference embeddings (= cosine similarity), then aggregate **per class = max similarity of any reference photo in that class**. Classes sorted by similarity become `all_probabilities`. Caveat: these "confidences" are cosine similarities, **not probabilities** — they don't sum to 1 and are typically all high (0.4–0.9); compare relatively, not against an absolute threshold. **Fallback: YOLO classifier** (`method: "yolo_classifier"`) — only when DINOv2 is unavailable (no index/model) or throws. A fine-tuned `yolo26n-cls` checkpoint; its `all_probabilities` are real softmax probabilities. Weights are **auto-discovered**: `CLASSIFIER_MODEL_PATH` env wins; otherwise the newest `produk-pfm-classifier-26n-*e-*.pt` in `models/` by (date-in-filename, mtime) — so retraining just drops a new dated file, no config change. If both are unavailable, `classification` carries an `error` field instead. ## Stage 2 — OCR extraction (SKU, expiry date, product name) PaddleOCR (`lang='en'`, textline orientation on) produces `rec_texts` lines + `rec_polys` boxes. Three extractors run over the lines: - **SKU** (`extract_sku`): first 8-digit number anywhere; else first 7–9 digit number. (Primafood SKUs are 8 digits, printed near the label top.) - **Expiry date** (`extract_expired_date`): each line is first noise-cleaned (`clean_date_line`: `1)`→`0`, `()`→`0`, `B8/8B/88`→`BB` before digits, o→0, I/l/|→1, S→5, Z→2, B→8 when digit-flanked, plus `012`/`112` month-misread repairs), then a **6-level priority cascade** runs: (1) BB/EXP-keyword line with compact `DDMMYYYY`; (2) keyword line with spaced `DD MM YYYY`; (3) keyword + 6–8 digit run; (3.5) keyword line, lenient noisy match; (4) any line spaced date; (5) any line compact `DDMMYYYY` — skipping lines that look like a SKU-on-product-name; (6) legacy formats (slashes, `05 MAR 2027`). Recognized keywords: `EXP`, `EXPIRED`, `TGL`, `EXPIRY`, `BBD`, `BEST BEFORE`, `BB`, `BAIK DIGUNAKAN`. Output normalized to `DD/MM/YYYY`. - **Product name** (`extract_product_name`): longest line containing a brand/ product keyword (FIESTA, CHAMP, OKEY, AKUMO, ASIMO, NUGGET, SOSIS, …) after stripping SKU digits and date fragments; falls back to the classifier's `top1_name`, then the longest non-numeric line, then `"Unknown Product"`. Visualization artifacts built server-side: `vis_image_base64` (all OCR boxes drawn teal `TEXT`, the expiry line amber `EXP`, on the orientation-corrected image so boxes align), `expired_date_crop_base64` (padded crop of the expiry line for eyeball verification — `find_expired_crop_index` prefers the box whose digits actually contain the date), and `spotting_image_base64` (a second pipeline call with `promptLabel: "spotting"`, no layout detection). ## Endpoint reference | Endpoint | Where | Purpose | |---|---|---| | `POST /api/scan-pfm` | gateway | Main scan. Body `{image_base64}` (data-URL ok). Returns `{classification, ocr, possibleMatches[], layoutParsingResult}` | | `POST http://paddleocr-pipeline-api:8120/classify-ocr` | classifier server | Internal. Body `{image_base64}`. Returns `{classification: {top1_name, top1_confidence, all_probabilities[], method}, ocr: {text_lines[], extracted_product_name, extracted_sku, extracted_expired_date, expired_line_index, expired_source_line, expired_date_crop_base64, vis_image_base64, spotting_image_base64}}` | | `GET /api/produk-pfm` | gateway | Gallery: SKU folders under `public/produk-pfm/foto-kemasan-v2/` with image + thumb URLs | | `GET/POST /api/manual-label-scan` | gateway | Ground-truth read/upsert to `sources/product_manual_labels.json` (host-visible via the `./backend/sources:/sources` mount) | | `POST :8090/layout-parsing` | pipeline | Shared PaddleX pipeline; used here for Visual Grid images and (with `promptLabel: "spotting"`) the Spotting Grid | | `/scan-pfm` | nginx :8000 | Proxies the page to Next.js :3000 | `possibleMatches[]` items: `{no_sku, nama_item, score, yoloSimilarity, isBestMatch}` — `score` currently equals `yoloSimilarity` (name-vs-name Levenshtein, 0..1). ## Model artifacts & retraining | File (`pfm-web-app/public/produk-pfm/`) | What | |---|---| | `foto-kemasan-v2//…` | Reference photo dataset — 16 classes, 118 photos (2–16 each) | | `models/dinov2_index.pkl` | DINOv2 embeddings + metadata (rebuild after adding photos) | | `models/produk-pfm-classifier-26n-100e-2026-07-08.pt` / `.onnx` | Fine-tuned YOLO classifier (83.3% top-1 / 90% top-5 val on the thin dataset) | | `index_dinov2.py` | Rebuilds the pickle index from `foto-kemasan-v2/` | | `train_classifier.py` | Splits 80/20 into `yolo_dataset/`, fine-tunes `yolo26n-cls.pt` (default 100 epochs, `--imgsz 224`), writes a dated checkpoint | **Retraining procedure (Windows host — bare-metal doesn't work here, `paddlepaddle-gpu` wheels are Linux-only):** add photos to `foto-kemasan-v2/`, `docker compose build pipeline-api` from the **repo root**, run a one-off `docker run --gpus all` from that image with `models/` mounted **writable** (the live service mounts it `:ro`), run `index_dinov2.py` then `train_classifier.py train --imgsz 224`, then `docker compose restart pipeline-api`. From Git Bash prefix `MSYS_NO_PATHCONV=1` or `/app/...` arguments get mangled. Verify in `docker logs`: "DINOv2 index loaded with N reference images", "Using classifier weights: ". Full worked example: `plans/next-enhancements.md` task 2.1. ## Accuracy regression harness `backend/scripts/accuracy-check-scan.mts` — mirrors the DO-flow's `pfm-web-app/scripts/accuracy-check.mts`. Hits the live `/api/scan-pfm` for every labeled image in `sources/product_manual_labels.json`, checks 3 fields (`no_sku`, `nama_item`, `expiry_date`) against ground truth, and splits into: - **Training Set** — gallery photos under `foto-kemasan-v2/` (the classifier's own reference images; scores here measure memorization, not generalization). - **Validation Set** — flat filenames dropped into `sources/product-test-images/` (a real held-out set; see that folder's `README.md` for the drop-photo → label → re-run workflow via `/manual-label-scan`). Every run appends to `sources/product_accuracy_history.jsonl` and **auto-diffs against the previous run**: the printed summary shows a Δ column per field per split, flags field/image-level regressions and improvements, and reports classifier method (`dinov2_similarity`/`yolo_classifier`) distribution + average confidence as informational context (not scored pass/fail, since DINOv2's "confidence" is a raw cosine similarity, not a calibrated probability — see Stage 1 above). This is what makes it safe to tune `classify_ocr_server.py` and immediately see whether a change helped or hurt. ```bash node scripts/accuracy-check-scan.mts # from backend/ ``` ## Operational notes - **Env vars**: `CLASSIFIER_SERVER_URL`, `PIPELINE_URL` (gateway, set in compose); `CLASSIFIER_MODELS_DIR`, `CLASSIFIER_MODEL_PATH` (classifier server overrides). The gateway's in-code default `PIPELINE_URL` (`localhost:7871`) is stale — the compose env always overrides it in Docker. - **Startup order/health**: the classifier server loads DINOv2 (torch.hub → needs network/cache), YOLO, and PaddleOCR at import time; until done, :8120 refuses connections and `/api/scan-pfm` 500s. No healthcheck exists yet (plan task 4.2 / 1.6). - **GPU**: DINOv2 + YOLO + PaddleOCR share the container/GPU with the PaddleX pipeline; all are small (ViT-S/14, nano YOLO) next to the vLLM server's footprint, but they do add VRAM on the same `PIPELINE_DEVICE`. - **Failure isolation**: layout-vis and spotting calls are best-effort (`null`/absent on failure); classification and OCR errors surface as `error` fields inside their sections rather than failing the whole scan. ## Known gaps & future recommendations Tracked ones (see `plans/next-enhancements.md`): - **Dataset thinness**: 2–16 photos/class caps both classifiers; every new real photo (especially non-studio, in-warehouse shots) matters. The harness above already reports gallery (training) vs. held-out (validation) accuracy separately — but as of this writing `sources/product-test-images/` is empty, so the Validation Set is still 0 images and every published number so far is a training/memorization score. Dropping real photos there is the next step, not yet done. Additional recommendations (not yet tasks — promote via `e`/`n` when wanted): 1. ~~Use `extracted_sku` in match ranking.~~ **Done** — `product-scan.ts`'s `classifyAndMatchProduct` already pins rank 1 to an exact `no_sku` match (score forced to 1.0) before falling back to name similarity. 2. **Fuse DINOv2 and YOLO instead of primary/fallback** (e.g. agreement boosts confidence; disagreement flags for review) — cheap, both already load. 3. **"Not a known product" handling**: DINOv2 always returns *some* class; add a minimum-similarity threshold below which the response says unknown rather than confidently misclassifying a foreign package. 4. **Pin the DINOv2 backbone offline** (vendor the weights or pre-bake the torch.hub cache into the image) — startup currently depends on an internet fetch on cold cache, bad for on-prem deploys. 5. **Batch/lot number extraction** — explicitly out of scope so far (plan §2 note); if requested, follow the expiry-date regex-cascade pattern. 6. **Mobile**: no web mobile page by design (task 2.2 cancelled) — real mobile scanning should go through the Flutter app calling `POST /api/scan-pfm` (would need an authenticated `/api/v1` variant; the classic route has no auth).