# Product Scan (scan-pfm) — How It Works
End-to-end reference for the Product/SKU scanning feature: a photo of a Primafood
product package goes in; the SKU class, product name, expiry date, and a ranked
SKU-master match list come out. Written 2026-07-08 against the live code. Related:
`plans/next-enhancements.md` §2 (build history) and §6 (ground-truth roadmap);
`docs/feature-list.md` tasks 2.1/2.3.
## High-level flow
```mermaid
flowchart LR
A[Browser: /scan-pfm page] -->|"POST /api/scan-pfm {image_base64}"| B[Next.js gateway
pfm-web-app :3000]
B -->|"POST :8120/classify-ocr"| C[classify_ocr_server.py
FastAPI, in pipeline-api]
C --> C1[1. DINOv2 similarity
fallback: YOLO classifier]
C --> C2[2. PaddleOCR + regex
SKU / expiry / name]
C -->|"POST localhost:8090/layout-parsing
promptLabel: spotting"| D[PaddleX pipeline
same container]
B -->|"POST :8090/layout-parsing"| D
B -->|"SELECT sku_master"| E[(Postgres)]
B -->|Levenshtein ranking| A
```
Two processes live in the `paddleocr-pipeline-api` container, both started by
`scripts/serve-pipeline.sh`: the PaddleX layout-parsing pipeline on **:8090**
(shared with the DO flow; VL recognition goes out to the vLLM server on :8118) and
`config/classify_ocr_server.py` on **:8120** (product scan only). The gateway
reaches them via Docker DNS (`CLASSIFIER_SERVER_URL`, `PIPELINE_URL` in root
`docker-compose.yml:87-88`); nginx (:8000) proxies `/scan-pfm` to the Next.js app.
## Request walkthrough
1. **Page** (`pfm-web-app/src/app/scan-pfm/page.tsx`, desktop-only test UI): pick a
sample from the gallery (`GET /api/produk-pfm`) or upload/rotate a photo (rotation
is done client-side on a canvas), then send it as a base64 data-URL.
2. **Gateway** (`api/scan-pfm/route.ts`):
- forwards `{image_base64}` to the classifier server (`/classify-ocr`);
- separately calls the layout-parsing pipeline with `useLayoutDetection: true`
for the Visual Grid tab's output images (failure here is non-fatal — logged,
`layoutParsingResult` returns `null`);
- loads the full `sku_master` table and ranks every SKU by **Levenshtein
similarity between `nama_item` and the classifier's `top1_name`**
(lowercased, alphanumerics only). Top 5 with score > 0.1 are returned;
rank 1 gets `isBestMatch: true`. Note: `ocr.extracted_sku` and
`ocr.extracted_product_name` are read but **not used** in this ranking —
see Future recommendations.
3. **Classifier server** (`config/classify_ocr_server.py`) does classification,
OCR extraction, and visualization — detailed below — and returns
`{classification, ocr}`.
4. **Page renders** four tabs: Summary (classification card + top-5 override
"Use" buttons + OCR fields + SKU matches), Visual Grid, Spotting Grid, Raw
Response (JSON). "Save Ground Truth" posts to `/api/manual-label-scan`.
## Stage 1 — classification (which product is this?)
**Primary: DINOv2 similarity search** (`method: "dinov2_similarity"`). At startup
the server loads `dinov2_vits14` **from `torch.hub` (network fetch on first run)**
plus `models/dinov2_index.pkl` — precomputed L2-normalized 384-dim embeddings of
all 118 reference photos across 16 SKU class folders. Per request: embed the query
image (resize 224², ImageNet normalization), dot-product against all reference
embeddings (= cosine similarity), then aggregate **per class = max similarity of
any reference photo in that class**. Classes sorted by similarity become
`all_probabilities`. Caveat: these "confidences" are cosine similarities, **not
probabilities** — they don't sum to 1 and are typically all high (0.4–0.9);
compare relatively, not against an absolute threshold.
**Fallback: YOLO classifier** (`method: "yolo_classifier"`) — only when DINOv2 is
unavailable (no index/model) or throws. A fine-tuned `yolo26n-cls` checkpoint;
its `all_probabilities` are real softmax probabilities. Weights are
**auto-discovered**: `CLASSIFIER_MODEL_PATH` env wins; otherwise the newest
`produk-pfm-classifier-26n-*e-*.pt` in `models/` by (date-in-filename, mtime) —
so retraining just drops a new dated file, no config change.
If both are unavailable, `classification` carries an `error` field instead.
## Stage 2 — OCR extraction (SKU, expiry date, product name)
PaddleOCR (`lang='en'`, textline orientation on) produces `rec_texts` lines +
`rec_polys` boxes. Three extractors run over the lines:
- **SKU** (`extract_sku`): first 8-digit number anywhere; else first 7–9 digit
number. (Primafood SKUs are 8 digits, printed near the label top.)
- **Expiry date** (`extract_expired_date`): each line is first noise-cleaned
(`clean_date_line`: `1)`→`0`, `()`→`0`, `B8/8B/88`→`BB` before digits, o→0,
I/l/|→1, S→5, Z→2, B→8 when digit-flanked, plus `012`/`112` month-misread
repairs), then a **6-level priority cascade** runs: (1) BB/EXP-keyword line
with compact `DDMMYYYY`; (2) keyword line with spaced `DD MM YYYY`; (3)
keyword + 6–8 digit run; (3.5) keyword line, lenient noisy match; (4) any line
spaced date; (5) any line compact `DDMMYYYY` — skipping lines that look like a
SKU-on-product-name; (6) legacy formats (slashes, `05 MAR 2027`). Recognized
keywords: `EXP`, `EXPIRED`, `TGL`, `EXPIRY`, `BBD`, `BEST BEFORE`, `BB`,
`BAIK DIGUNAKAN`. Output normalized to `DD/MM/YYYY`.
- **Product name** (`extract_product_name`): longest line containing a brand/
product keyword (FIESTA, CHAMP, OKEY, AKUMO, ASIMO, NUGGET, SOSIS, …) after
stripping SKU digits and date fragments; falls back to the classifier's
`top1_name`, then the longest non-numeric line, then `"Unknown Product"`.
Visualization artifacts built server-side: `vis_image_base64` (all OCR boxes
drawn teal `TEXT`, the expiry line amber `EXP`, on the orientation-corrected
image so boxes align), `expired_date_crop_base64` (padded crop of the expiry
line for eyeball verification — `find_expired_crop_index` prefers the box whose
digits actually contain the date), and `spotting_image_base64` (a second
pipeline call with `promptLabel: "spotting"`, no layout detection).
## Endpoint reference
| Endpoint | Where | Purpose |
|---|---|---|
| `POST /api/scan-pfm` | gateway | Main scan. Body `{image_base64}` (data-URL ok). Returns `{classification, ocr, possibleMatches[], layoutParsingResult}` |
| `POST http://paddleocr-pipeline-api:8120/classify-ocr` | classifier server | Internal. Body `{image_base64}`. Returns `{classification: {top1_name, top1_confidence, all_probabilities[], method}, ocr: {text_lines[], extracted_product_name, extracted_sku, extracted_expired_date, expired_line_index, expired_source_line, expired_date_crop_base64, vis_image_base64, spotting_image_base64}}` |
| `GET /api/produk-pfm` | gateway | Gallery: SKU folders under `public/produk-pfm/foto-kemasan-v2/` with image + thumb URLs |
| `GET/POST /api/manual-label-scan` | gateway | Ground-truth read/upsert to `sources/product_manual_labels.json` (host-visible via the `./backend/sources:/sources` mount) |
| `POST :8090/layout-parsing` | pipeline | Shared PaddleX pipeline; used here for Visual Grid images and (with `promptLabel: "spotting"`) the Spotting Grid |
| `/scan-pfm` | nginx :8000 | Proxies the page to Next.js :3000 |
`possibleMatches[]` items: `{no_sku, nama_item, score, yoloSimilarity, isBestMatch}` —
`score` currently equals `yoloSimilarity` (name-vs-name Levenshtein, 0..1).
## Model artifacts & retraining
| File (`pfm-web-app/public/produk-pfm/`) | What |
|---|---|
| `foto-kemasan-v2//…` | Reference photo dataset — 16 classes, 118 photos (2–16 each) |
| `models/dinov2_index.pkl` | DINOv2 embeddings + metadata (rebuild after adding photos) |
| `models/produk-pfm-classifier-26n-100e-2026-07-08.pt` / `.onnx` | Fine-tuned YOLO classifier (83.3% top-1 / 90% top-5 val on the thin dataset) |
| `index_dinov2.py` | Rebuilds the pickle index from `foto-kemasan-v2/` |
| `train_classifier.py` | Splits 80/20 into `yolo_dataset/`, fine-tunes `yolo26n-cls.pt` (default 100 epochs, `--imgsz 224`), writes a dated checkpoint |
**Retraining procedure (Windows host — bare-metal doesn't work here,
`paddlepaddle-gpu` wheels are Linux-only):** add photos to `foto-kemasan-v2/`,
`docker compose build pipeline-api` from the **repo root**, run a one-off
`docker run --gpus all` from that image with `models/` mounted **writable** (the
live service mounts it `:ro`), run `index_dinov2.py` then
`train_classifier.py train --imgsz 224`, then `docker compose restart
pipeline-api`. From Git Bash prefix `MSYS_NO_PATHCONV=1` or `/app/...` arguments
get mangled. Verify in `docker logs`: "DINOv2 index loaded with N reference
images", "Using classifier weights: ". Full worked example:
`plans/next-enhancements.md` task 2.1.
## Accuracy regression harness
`backend/scripts/accuracy-check-scan.mts` — mirrors the DO-flow's
`pfm-web-app/scripts/accuracy-check.mts`. Hits the live `/api/scan-pfm` for
every labeled image in `sources/product_manual_labels.json`, checks 3 fields
(`no_sku`, `nama_item`, `expiry_date`) against ground truth, and splits into:
- **Training Set** — gallery photos under `foto-kemasan-v2/` (the classifier's
own reference images; scores here measure memorization, not generalization).
- **Validation Set** — flat filenames dropped into
`sources/product-test-images/` (a real held-out set; see that folder's
`README.md` for the drop-photo → label → re-run workflow via
`/manual-label-scan`).
Every run appends to `sources/product_accuracy_history.jsonl` and **auto-diffs
against the previous run**: the printed summary shows a Δ column per field per
split, flags field/image-level regressions and improvements, and reports
classifier method (`dinov2_similarity`/`yolo_classifier`) distribution +
average confidence as informational context (not scored pass/fail, since
DINOv2's "confidence" is a raw cosine similarity, not a calibrated
probability — see Stage 1 above). This is what makes it safe to tune
`classify_ocr_server.py` and immediately see whether a change helped or hurt.
```bash
node scripts/accuracy-check-scan.mts # from backend/
```
## Operational notes
- **Env vars**: `CLASSIFIER_SERVER_URL`, `PIPELINE_URL` (gateway, set in compose);
`CLASSIFIER_MODELS_DIR`, `CLASSIFIER_MODEL_PATH` (classifier server overrides).
The gateway's in-code default `PIPELINE_URL` (`localhost:7871`) is stale — the
compose env always overrides it in Docker.
- **Startup order/health**: the classifier server loads DINOv2 (torch.hub →
needs network/cache), YOLO, and PaddleOCR at import time; until done, :8120
refuses connections and `/api/scan-pfm` 500s. No healthcheck exists yet (plan
task 4.2 / 1.6).
- **GPU**: DINOv2 + YOLO + PaddleOCR share the container/GPU with the PaddleX
pipeline; all are small (ViT-S/14, nano YOLO) next to the vLLM server's
footprint, but they do add VRAM on the same `PIPELINE_DEVICE`.
- **Failure isolation**: layout-vis and spotting calls are best-effort
(`null`/absent on failure); classification and OCR errors surface as `error`
fields inside their sections rather than failing the whole scan.
## Known gaps & future recommendations
Tracked ones (see `plans/next-enhancements.md`):
- **Dataset thinness**: 2–16 photos/class caps both classifiers; every new real
photo (especially non-studio, in-warehouse shots) matters. The harness above
already reports gallery (training) vs. held-out (validation) accuracy
separately — but as of this writing `sources/product-test-images/` is empty,
so the Validation Set is still 0 images and every published number so far is
a training/memorization score. Dropping real photos there is the next step,
not yet done.
Additional recommendations (not yet tasks — promote via `e`/`n` when wanted):
1. ~~Use `extracted_sku` in match ranking.~~ **Done** — `product-scan.ts`'s
`classifyAndMatchProduct` already pins rank 1 to an exact `no_sku` match
(score forced to 1.0) before falling back to name similarity.
2. **Fuse DINOv2 and YOLO instead of primary/fallback** (e.g. agreement boosts
confidence; disagreement flags for review) — cheap, both already load.
3. **"Not a known product" handling**: DINOv2 always returns *some* class; add a
minimum-similarity threshold below which the response says unknown rather
than confidently misclassifying a foreign package.
4. **Pin the DINOv2 backbone offline** (vendor the weights or pre-bake the
torch.hub cache into the image) — startup currently depends on an internet
fetch on cold cache, bad for on-prem deploys.
5. **Batch/lot number extraction** — explicitly out of scope so far (plan §2
note); if requested, follow the expiry-date regex-cascade pattern.
6. **Mobile**: no web mobile page by design (task 2.2 cancelled) — real mobile
scanning should go through the Flutter app calling `POST /api/scan-pfm`
(would need an authenticated `/api/v1` variant; the classic route has no auth).