Files
pfm-ocr/backend/docs/scan-product.md
T
Rafhan Mazaya FathurrahmanandClaude Sonnet 5 3a17c28758 feat(backend): diff-vs-previous-run reporting for product-scan accuracy harness
Ports the DO-harness's auto-diff-vs-previous-run reporting into
accuracy-check-scan.mts: prints a per-field, per-split (Training/
Validation) delta against the last product_accuracy_history.jsonl entry
and calls out regressions/improvements explicitly, plus classifier
method distribution and average confidence as informational context.

Also adds the real held-out validation photo set into
sources/product-test-images/ (75 photos, one per current SKU class) with
its README documenting the drop-photo -> label -> re-run workflow, so the
harness's Validation Set split actually has images to score against.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xsxk4ZkDQVVaLUcixDcqb5
2026-07-14 08:34:02 +07:00

210 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Product Scan (scan-pfm) — How It Works
End-to-end reference for the Product/SKU scanning feature: a photo of a Primafood
product package goes in; the SKU class, product name, expiry date, and a ranked
SKU-master match list come out. Written 2026-07-08 against the live code. Related:
`plans/next-enhancements.md` §2 (build history) and §6 (ground-truth roadmap);
`docs/feature-list.md` tasks 2.1/2.3.
## High-level flow
```mermaid
flowchart LR
A[Browser: /scan-pfm page] -->|"POST /api/scan-pfm {image_base64}"| B[Next.js gateway<br/>pfm-web-app :3000]
B -->|"POST :8120/classify-ocr"| C[classify_ocr_server.py<br/>FastAPI, in pipeline-api]
C --> C1[1. DINOv2 similarity<br/>fallback: YOLO classifier]
C --> C2[2. PaddleOCR + regex<br/>SKU / expiry / name]
C -->|"POST localhost:8090/layout-parsing<br/>promptLabel: spotting"| D[PaddleX pipeline<br/>same container]
B -->|"POST :8090/layout-parsing"| D
B -->|"SELECT sku_master"| E[(Postgres)]
B -->|Levenshtein ranking| A
```
Two processes live in the `paddleocr-pipeline-api` container, both started by
`scripts/serve-pipeline.sh`: the PaddleX layout-parsing pipeline on **:8090**
(shared with the DO flow; VL recognition goes out to the vLLM server on :8118) and
`config/classify_ocr_server.py` on **:8120** (product scan only). The gateway
reaches them via Docker DNS (`CLASSIFIER_SERVER_URL`, `PIPELINE_URL` in root
`docker-compose.yml:87-88`); nginx (:8000) proxies `/scan-pfm` to the Next.js app.
## Request walkthrough
1. **Page** (`pfm-web-app/src/app/scan-pfm/page.tsx`, desktop-only test UI): pick a
sample from the gallery (`GET /api/produk-pfm`) or upload/rotate a photo (rotation
is done client-side on a canvas), then send it as a base64 data-URL.
2. **Gateway** (`api/scan-pfm/route.ts`):
- forwards `{image_base64}` to the classifier server (`/classify-ocr`);
- separately calls the layout-parsing pipeline with `useLayoutDetection: true`
for the Visual Grid tab's output images (failure here is non-fatal — logged,
`layoutParsingResult` returns `null`);
- loads the full `sku_master` table and ranks every SKU by **Levenshtein
similarity between `nama_item` and the classifier's `top1_name`**
(lowercased, alphanumerics only). Top 5 with score > 0.1 are returned;
rank 1 gets `isBestMatch: true`. Note: `ocr.extracted_sku` and
`ocr.extracted_product_name` are read but **not used** in this ranking —
see Future recommendations.
3. **Classifier server** (`config/classify_ocr_server.py`) does classification,
OCR extraction, and visualization — detailed below — and returns
`{classification, ocr}`.
4. **Page renders** four tabs: Summary (classification card + top-5 override
"Use" buttons + OCR fields + SKU matches), Visual Grid, Spotting Grid, Raw
Response (JSON). "Save Ground Truth" posts to `/api/manual-label-scan`.
## Stage 1 — classification (which product is this?)
**Primary: DINOv2 similarity search** (`method: "dinov2_similarity"`). At startup
the server loads `dinov2_vits14` **from `torch.hub` (network fetch on first run)**
plus `models/dinov2_index.pkl` — precomputed L2-normalized 384-dim embeddings of
all 118 reference photos across 16 SKU class folders. Per request: embed the query
image (resize 224², ImageNet normalization), dot-product against all reference
embeddings (= cosine similarity), then aggregate **per class = max similarity of
any reference photo in that class**. Classes sorted by similarity become
`all_probabilities`. Caveat: these "confidences" are cosine similarities, **not
probabilities** — they don't sum to 1 and are typically all high (0.4–0.9);
compare relatively, not against an absolute threshold.
**Fallback: YOLO classifier** (`method: "yolo_classifier"`) — only when DINOv2 is
unavailable (no index/model) or throws. A fine-tuned `yolo26n-cls` checkpoint;
its `all_probabilities` are real softmax probabilities. Weights are
**auto-discovered**: `CLASSIFIER_MODEL_PATH` env wins; otherwise the newest
`produk-pfm-classifier-26n-*e-*.pt` in `models/` by (date-in-filename, mtime) —
so retraining just drops a new dated file, no config change.
If both are unavailable, `classification` carries an `error` field instead.
## Stage 2 — OCR extraction (SKU, expiry date, product name)
PaddleOCR (`lang='en'`, textline orientation on) produces `rec_texts` lines +
`rec_polys` boxes. Three extractors run over the lines:
- **SKU** (`extract_sku`): first 8-digit number anywhere; else first 7–9 digit
number. (Primafood SKUs are 8 digits, printed near the label top.)
- **Expiry date** (`extract_expired_date`): each line is first noise-cleaned
(`clean_date_line`: `1)`→`0`, `()`→`0`, `B8/8B/88`→`BB` before digits, o→0,
I/l/|→1, S→5, Z→2, B→8 when digit-flanked, plus `012`/`112` month-misread
repairs), then a **6-level priority cascade** runs: (1) BB/EXP-keyword line
with compact `DDMMYYYY`; (2) keyword line with spaced `DD MM YYYY`; (3)
keyword + 6–8 digit run; (3.5) keyword line, lenient noisy match; (4) any line
spaced date; (5) any line compact `DDMMYYYY` — skipping lines that look like a
SKU-on-product-name; (6) legacy formats (slashes, `05 MAR 2027`). Recognized
keywords: `EXP`, `EXPIRED`, `TGL`, `EXPIRY`, `BBD`, `BEST BEFORE`, `BB`,
`BAIK DIGUNAKAN`. Output normalized to `DD/MM/YYYY`.
- **Product name** (`extract_product_name`): longest line containing a brand/
product keyword (FIESTA, CHAMP, OKEY, AKUMO, ASIMO, NUGGET, SOSIS, …) after
stripping SKU digits and date fragments; falls back to the classifier's
`top1_name`, then the longest non-numeric line, then `"Unknown Product"`.
Visualization artifacts built server-side: `vis_image_base64` (all OCR boxes
drawn teal `TEXT`, the expiry line amber `EXP`, on the orientation-corrected
image so boxes align), `expired_date_crop_base64` (padded crop of the expiry
line for eyeball verification — `find_expired_crop_index` prefers the box whose
digits actually contain the date), and `spotting_image_base64` (a second
pipeline call with `promptLabel: "spotting"`, no layout detection).
## Endpoint reference
| Endpoint | Where | Purpose |
|---|---|---|
| `POST /api/scan-pfm` | gateway | Main scan. Body `{image_base64}` (data-URL ok). Returns `{classification, ocr, possibleMatches[], layoutParsingResult}` |
| `POST http://paddleocr-pipeline-api:8120/classify-ocr` | classifier server | Internal. Body `{image_base64}`. Returns `{classification: {top1_name, top1_confidence, all_probabilities[], method}, ocr: {text_lines[], extracted_product_name, extracted_sku, extracted_expired_date, expired_line_index, expired_source_line, expired_date_crop_base64, vis_image_base64, spotting_image_base64}}` |
| `GET /api/produk-pfm` | gateway | Gallery: SKU folders under `public/produk-pfm/foto-kemasan-v2/` with image + thumb URLs |
| `GET/POST /api/manual-label-scan` | gateway | Ground-truth read/upsert to `sources/product_manual_labels.json` (host-visible via the `./backend/sources:/sources` mount) |
| `POST :8090/layout-parsing` | pipeline | Shared PaddleX pipeline; used here for Visual Grid images and (with `promptLabel: "spotting"`) the Spotting Grid |
| `/scan-pfm` | nginx :8000 | Proxies the page to Next.js :3000 |
`possibleMatches[]` items: `{no_sku, nama_item, score, yoloSimilarity, isBestMatch}` —
`score` currently equals `yoloSimilarity` (name-vs-name Levenshtein, 0..1).
## Model artifacts & retraining
| File (`pfm-web-app/public/produk-pfm/`) | What |
|---|---|
| `foto-kemasan-v2/<SKU or class>/…` | Reference photo dataset — 16 classes, 118 photos (2–16 each) |
| `models/dinov2_index.pkl` | DINOv2 embeddings + metadata (rebuild after adding photos) |
| `models/produk-pfm-classifier-26n-100e-2026-07-08.pt` / `.onnx` | Fine-tuned YOLO classifier (83.3% top-1 / 90% top-5 val on the thin dataset) |
| `index_dinov2.py` | Rebuilds the pickle index from `foto-kemasan-v2/` |
| `train_classifier.py` | Splits 80/20 into `yolo_dataset/`, fine-tunes `yolo26n-cls.pt` (default 100 epochs, `--imgsz 224`), writes a dated checkpoint |
**Retraining procedure (Windows host — bare-metal doesn't work here,
`paddlepaddle-gpu` wheels are Linux-only):** add photos to `foto-kemasan-v2/`,
`docker compose build pipeline-api` from the **repo root**, run a one-off
`docker run --gpus all` from that image with `models/` mounted **writable** (the
live service mounts it `:ro`), run `index_dinov2.py` then
`train_classifier.py train --imgsz 224`, then `docker compose restart
pipeline-api`. From Git Bash prefix `MSYS_NO_PATHCONV=1` or `/app/...` arguments
get mangled. Verify in `docker logs`: "DINOv2 index loaded with N reference
images", "Using classifier weights: <new dated file>". Full worked example:
`plans/next-enhancements.md` task 2.1.
## Accuracy regression harness
`backend/scripts/accuracy-check-scan.mts` — mirrors the DO-flow's
`pfm-web-app/scripts/accuracy-check.mts`. Hits the live `/api/scan-pfm` for
every labeled image in `sources/product_manual_labels.json`, checks 3 fields
(`no_sku`, `nama_item`, `expiry_date`) against ground truth, and splits into:
- **Training Set** — gallery photos under `foto-kemasan-v2/` (the classifier's
own reference images; scores here measure memorization, not generalization).
- **Validation Set** — flat filenames dropped into
`sources/product-test-images/` (a real held-out set; see that folder's
`README.md` for the drop-photo → label → re-run workflow via
`/manual-label-scan`).
Every run appends to `sources/product_accuracy_history.jsonl` and **auto-diffs
against the previous run**: the printed summary shows a Δ column per field per
split, flags field/image-level regressions and improvements, and reports
classifier method (`dinov2_similarity`/`yolo_classifier`) distribution +
average confidence as informational context (not scored pass/fail, since
DINOv2's "confidence" is a raw cosine similarity, not a calibrated
probability — see Stage 1 above). This is what makes it safe to tune
`classify_ocr_server.py` and immediately see whether a change helped or hurt.
```bash
node scripts/accuracy-check-scan.mts # from backend/
```
## Operational notes
- **Env vars**: `CLASSIFIER_SERVER_URL`, `PIPELINE_URL` (gateway, set in compose);
`CLASSIFIER_MODELS_DIR`, `CLASSIFIER_MODEL_PATH` (classifier server overrides).
The gateway's in-code default `PIPELINE_URL` (`localhost:7871`) is stale — the
compose env always overrides it in Docker.
- **Startup order/health**: the classifier server loads DINOv2 (torch.hub →
needs network/cache), YOLO, and PaddleOCR at import time; until done, :8120
refuses connections and `/api/scan-pfm` 500s. No healthcheck exists yet (plan
task 4.2 / 1.6).
- **GPU**: DINOv2 + YOLO + PaddleOCR share the container/GPU with the PaddleX
pipeline; all are small (ViT-S/14, nano YOLO) next to the vLLM server's
footprint, but they do add VRAM on the same `PIPELINE_DEVICE`.
- **Failure isolation**: layout-vis and spotting calls are best-effort
(`null`/absent on failure); classification and OCR errors surface as `error`
fields inside their sections rather than failing the whole scan.
## Known gaps & future recommendations
Tracked ones (see `plans/next-enhancements.md`):
- **Dataset thinness**: 2–16 photos/class caps both classifiers; every new real
photo (especially non-studio, in-warehouse shots) matters. The harness above
already reports gallery (training) vs. held-out (validation) accuracy
separately — but as of this writing `sources/product-test-images/` is empty,
so the Validation Set is still 0 images and every published number so far is
a training/memorization score. Dropping real photos there is the next step,
not yet done.
Additional recommendations (not yet tasks — promote via `e`/`n` when wanted):
1. ~~Use `extracted_sku` in match ranking.~~ **Done** — `product-scan.ts`'s
`classifyAndMatchProduct` already pins rank 1 to an exact `no_sku` match
(score forced to 1.0) before falling back to name similarity.
2. **Fuse DINOv2 and YOLO instead of primary/fallback** (e.g. agreement boosts
confidence; disagreement flags for review) — cheap, both already load.
3. **"Not a known product" handling**: DINOv2 always returns *some* class; add a
minimum-similarity threshold below which the response says unknown rather
than confidently misclassifying a foreign package.
4. **Pin the DINOv2 backbone offline** (vendor the weights or pre-bake the
torch.hub cache into the image) — startup currently depends on an internet
fetch on cold cache, bad for on-prem deploys.
5. **Batch/lot number extraction** — explicitly out of scope so far (plan §2
note); if requested, follow the expiry-date regex-cascade pattern.
6. **Mobile**: no web mobile page by design (task 2.2 cancelled) — real mobile
scanning should go through the Flutter app calling `POST /api/scan-pfm`
(would need an authenticated `/api/v1` variant; the classic route has no auth).