Ports the DO-harness's auto-diff-vs-previous-run reporting into accuracy-check-scan.mts: prints a per-field, per-split (Training/ Validation) delta against the last product_accuracy_history.jsonl entry and calls out regressions/improvements explicitly, plus classifier method distribution and average confidence as informational context. Also adds the real held-out validation photo set into sources/product-test-images/ (75 photos, one per current SKU class) with its README documenting the drop-photo -> label -> re-run workflow, so the harness's Validation Set split actually has images to score against. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Xsxk4ZkDQVVaLUcixDcqb5
210 lines
13 KiB
Markdown
210 lines
13 KiB
Markdown
# Product Scan (scan-pfm) — How It Works
|
||
|
||
End-to-end reference for the Product/SKU scanning feature: a photo of a Primafood
|
||
product package goes in; the SKU class, product name, expiry date, and a ranked
|
||
SKU-master match list come out. Written 2026-07-08 against the live code. Related:
|
||
`plans/next-enhancements.md` §2 (build history) and §6 (ground-truth roadmap);
|
||
`docs/feature-list.md` tasks 2.1/2.3.
|
||
|
||
## High-level flow
|
||
|
||
```mermaid
|
||
flowchart LR
|
||
A[Browser: /scan-pfm page] -->|"POST /api/scan-pfm {image_base64}"| B[Next.js gateway<br/>pfm-web-app :3000]
|
||
B -->|"POST :8120/classify-ocr"| C[classify_ocr_server.py<br/>FastAPI, in pipeline-api]
|
||
C --> C1[1. DINOv2 similarity<br/>fallback: YOLO classifier]
|
||
C --> C2[2. PaddleOCR + regex<br/>SKU / expiry / name]
|
||
C -->|"POST localhost:8090/layout-parsing<br/>promptLabel: spotting"| D[PaddleX pipeline<br/>same container]
|
||
B -->|"POST :8090/layout-parsing"| D
|
||
B -->|"SELECT sku_master"| E[(Postgres)]
|
||
B -->|Levenshtein ranking| A
|
||
```
|
||
|
||
Two processes live in the `paddleocr-pipeline-api` container, both started by
|
||
`scripts/serve-pipeline.sh`: the PaddleX layout-parsing pipeline on **:8090**
|
||
(shared with the DO flow; VL recognition goes out to the vLLM server on :8118) and
|
||
`config/classify_ocr_server.py` on **:8120** (product scan only). The gateway
|
||
reaches them via Docker DNS (`CLASSIFIER_SERVER_URL`, `PIPELINE_URL` in root
|
||
`docker-compose.yml:87-88`); nginx (:8000) proxies `/scan-pfm` to the Next.js app.
|
||
|
||
## Request walkthrough
|
||
|
||
1. **Page** (`pfm-web-app/src/app/scan-pfm/page.tsx`, desktop-only test UI): pick a
|
||
sample from the gallery (`GET /api/produk-pfm`) or upload/rotate a photo (rotation
|
||
is done client-side on a canvas), then send it as a base64 data-URL.
|
||
2. **Gateway** (`api/scan-pfm/route.ts`):
|
||
- forwards `{image_base64}` to the classifier server (`/classify-ocr`);
|
||
- separately calls the layout-parsing pipeline with `useLayoutDetection: true`
|
||
for the Visual Grid tab's output images (failure here is non-fatal — logged,
|
||
`layoutParsingResult` returns `null`);
|
||
- loads the full `sku_master` table and ranks every SKU by **Levenshtein
|
||
similarity between `nama_item` and the classifier's `top1_name`**
|
||
(lowercased, alphanumerics only). Top 5 with score > 0.1 are returned;
|
||
rank 1 gets `isBestMatch: true`. Note: `ocr.extracted_sku` and
|
||
`ocr.extracted_product_name` are read but **not used** in this ranking —
|
||
see Future recommendations.
|
||
3. **Classifier server** (`config/classify_ocr_server.py`) does classification,
|
||
OCR extraction, and visualization — detailed below — and returns
|
||
`{classification, ocr}`.
|
||
4. **Page renders** four tabs: Summary (classification card + top-5 override
|
||
"Use" buttons + OCR fields + SKU matches), Visual Grid, Spotting Grid, Raw
|
||
Response (JSON). "Save Ground Truth" posts to `/api/manual-label-scan`.
|
||
|
||
## Stage 1 — classification (which product is this?)
|
||
|
||
**Primary: DINOv2 similarity search** (`method: "dinov2_similarity"`). At startup
|
||
the server loads `dinov2_vits14` **from `torch.hub` (network fetch on first run)**
|
||
plus `models/dinov2_index.pkl` — precomputed L2-normalized 384-dim embeddings of
|
||
all 118 reference photos across 16 SKU class folders. Per request: embed the query
|
||
image (resize 224², ImageNet normalization), dot-product against all reference
|
||
embeddings (= cosine similarity), then aggregate **per class = max similarity of
|
||
any reference photo in that class**. Classes sorted by similarity become
|
||
`all_probabilities`. Caveat: these "confidences" are cosine similarities, **not
|
||
probabilities** — they don't sum to 1 and are typically all high (0.4–0.9);
|
||
compare relatively, not against an absolute threshold.
|
||
|
||
**Fallback: YOLO classifier** (`method: "yolo_classifier"`) — only when DINOv2 is
|
||
unavailable (no index/model) or throws. A fine-tuned `yolo26n-cls` checkpoint;
|
||
its `all_probabilities` are real softmax probabilities. Weights are
|
||
**auto-discovered**: `CLASSIFIER_MODEL_PATH` env wins; otherwise the newest
|
||
`produk-pfm-classifier-26n-*e-*.pt` in `models/` by (date-in-filename, mtime) —
|
||
so retraining just drops a new dated file, no config change.
|
||
|
||
If both are unavailable, `classification` carries an `error` field instead.
|
||
|
||
## Stage 2 — OCR extraction (SKU, expiry date, product name)
|
||
|
||
PaddleOCR (`lang='en'`, textline orientation on) produces `rec_texts` lines +
|
||
`rec_polys` boxes. Three extractors run over the lines:
|
||
|
||
- **SKU** (`extract_sku`): first 8-digit number anywhere; else first 7–9 digit
|
||
number. (Primafood SKUs are 8 digits, printed near the label top.)
|
||
- **Expiry date** (`extract_expired_date`): each line is first noise-cleaned
|
||
(`clean_date_line`: `1)`→`0`, `()`→`0`, `B8/8B/88`→`BB` before digits, o→0,
|
||
I/l/|→1, S→5, Z→2, B→8 when digit-flanked, plus `012`/`112` month-misread
|
||
repairs), then a **6-level priority cascade** runs: (1) BB/EXP-keyword line
|
||
with compact `DDMMYYYY`; (2) keyword line with spaced `DD MM YYYY`; (3)
|
||
keyword + 6–8 digit run; (3.5) keyword line, lenient noisy match; (4) any line
|
||
spaced date; (5) any line compact `DDMMYYYY` — skipping lines that look like a
|
||
SKU-on-product-name; (6) legacy formats (slashes, `05 MAR 2027`). Recognized
|
||
keywords: `EXP`, `EXPIRED`, `TGL`, `EXPIRY`, `BBD`, `BEST BEFORE`, `BB`,
|
||
`BAIK DIGUNAKAN`. Output normalized to `DD/MM/YYYY`.
|
||
- **Product name** (`extract_product_name`): longest line containing a brand/
|
||
product keyword (FIESTA, CHAMP, OKEY, AKUMO, ASIMO, NUGGET, SOSIS, …) after
|
||
stripping SKU digits and date fragments; falls back to the classifier's
|
||
`top1_name`, then the longest non-numeric line, then `"Unknown Product"`.
|
||
|
||
Visualization artifacts built server-side: `vis_image_base64` (all OCR boxes
|
||
drawn teal `TEXT`, the expiry line amber `EXP`, on the orientation-corrected
|
||
image so boxes align), `expired_date_crop_base64` (padded crop of the expiry
|
||
line for eyeball verification — `find_expired_crop_index` prefers the box whose
|
||
digits actually contain the date), and `spotting_image_base64` (a second
|
||
pipeline call with `promptLabel: "spotting"`, no layout detection).
|
||
|
||
## Endpoint reference
|
||
|
||
| Endpoint | Where | Purpose |
|
||
|---|---|---|
|
||
| `POST /api/scan-pfm` | gateway | Main scan. Body `{image_base64}` (data-URL ok). Returns `{classification, ocr, possibleMatches[], layoutParsingResult}` |
|
||
| `POST http://paddleocr-pipeline-api:8120/classify-ocr` | classifier server | Internal. Body `{image_base64}`. Returns `{classification: {top1_name, top1_confidence, all_probabilities[], method}, ocr: {text_lines[], extracted_product_name, extracted_sku, extracted_expired_date, expired_line_index, expired_source_line, expired_date_crop_base64, vis_image_base64, spotting_image_base64}}` |
|
||
| `GET /api/produk-pfm` | gateway | Gallery: SKU folders under `public/produk-pfm/foto-kemasan-v2/` with image + thumb URLs |
|
||
| `GET/POST /api/manual-label-scan` | gateway | Ground-truth read/upsert to `sources/product_manual_labels.json` (host-visible via the `./backend/sources:/sources` mount) |
|
||
| `POST :8090/layout-parsing` | pipeline | Shared PaddleX pipeline; used here for Visual Grid images and (with `promptLabel: "spotting"`) the Spotting Grid |
|
||
| `/scan-pfm` | nginx :8000 | Proxies the page to Next.js :3000 |
|
||
|
||
`possibleMatches[]` items: `{no_sku, nama_item, score, yoloSimilarity, isBestMatch}` —
|
||
`score` currently equals `yoloSimilarity` (name-vs-name Levenshtein, 0..1).
|
||
|
||
## Model artifacts & retraining
|
||
|
||
| File (`pfm-web-app/public/produk-pfm/`) | What |
|
||
|---|---|
|
||
| `foto-kemasan-v2/<SKU or class>/…` | Reference photo dataset — 16 classes, 118 photos (2–16 each) |
|
||
| `models/dinov2_index.pkl` | DINOv2 embeddings + metadata (rebuild after adding photos) |
|
||
| `models/produk-pfm-classifier-26n-100e-2026-07-08.pt` / `.onnx` | Fine-tuned YOLO classifier (83.3% top-1 / 90% top-5 val on the thin dataset) |
|
||
| `index_dinov2.py` | Rebuilds the pickle index from `foto-kemasan-v2/` |
|
||
| `train_classifier.py` | Splits 80/20 into `yolo_dataset/`, fine-tunes `yolo26n-cls.pt` (default 100 epochs, `--imgsz 224`), writes a dated checkpoint |
|
||
|
||
**Retraining procedure (Windows host — bare-metal doesn't work here,
|
||
`paddlepaddle-gpu` wheels are Linux-only):** add photos to `foto-kemasan-v2/`,
|
||
`docker compose build pipeline-api` from the **repo root**, run a one-off
|
||
`docker run --gpus all` from that image with `models/` mounted **writable** (the
|
||
live service mounts it `:ro`), run `index_dinov2.py` then
|
||
`train_classifier.py train --imgsz 224`, then `docker compose restart
|
||
pipeline-api`. From Git Bash prefix `MSYS_NO_PATHCONV=1` or `/app/...` arguments
|
||
get mangled. Verify in `docker logs`: "DINOv2 index loaded with N reference
|
||
images", "Using classifier weights: <new dated file>". Full worked example:
|
||
`plans/next-enhancements.md` task 2.1.
|
||
|
||
## Accuracy regression harness
|
||
|
||
`backend/scripts/accuracy-check-scan.mts` — mirrors the DO-flow's
|
||
`pfm-web-app/scripts/accuracy-check.mts`. Hits the live `/api/scan-pfm` for
|
||
every labeled image in `sources/product_manual_labels.json`, checks 3 fields
|
||
(`no_sku`, `nama_item`, `expiry_date`) against ground truth, and splits into:
|
||
- **Training Set** — gallery photos under `foto-kemasan-v2/` (the classifier's
|
||
own reference images; scores here measure memorization, not generalization).
|
||
- **Validation Set** — flat filenames dropped into
|
||
`sources/product-test-images/` (a real held-out set; see that folder's
|
||
`README.md` for the drop-photo → label → re-run workflow via
|
||
`/manual-label-scan`).
|
||
|
||
Every run appends to `sources/product_accuracy_history.jsonl` and **auto-diffs
|
||
against the previous run**: the printed summary shows a Δ column per field per
|
||
split, flags field/image-level regressions and improvements, and reports
|
||
classifier method (`dinov2_similarity`/`yolo_classifier`) distribution +
|
||
average confidence as informational context (not scored pass/fail, since
|
||
DINOv2's "confidence" is a raw cosine similarity, not a calibrated
|
||
probability — see Stage 1 above). This is what makes it safe to tune
|
||
`classify_ocr_server.py` and immediately see whether a change helped or hurt.
|
||
|
||
```bash
|
||
node scripts/accuracy-check-scan.mts # from backend/
|
||
```
|
||
|
||
## Operational notes
|
||
|
||
- **Env vars**: `CLASSIFIER_SERVER_URL`, `PIPELINE_URL` (gateway, set in compose);
|
||
`CLASSIFIER_MODELS_DIR`, `CLASSIFIER_MODEL_PATH` (classifier server overrides).
|
||
The gateway's in-code default `PIPELINE_URL` (`localhost:7871`) is stale — the
|
||
compose env always overrides it in Docker.
|
||
- **Startup order/health**: the classifier server loads DINOv2 (torch.hub →
|
||
needs network/cache), YOLO, and PaddleOCR at import time; until done, :8120
|
||
refuses connections and `/api/scan-pfm` 500s. No healthcheck exists yet (plan
|
||
task 4.2 / 1.6).
|
||
- **GPU**: DINOv2 + YOLO + PaddleOCR share the container/GPU with the PaddleX
|
||
pipeline; all are small (ViT-S/14, nano YOLO) next to the vLLM server's
|
||
footprint, but they do add VRAM on the same `PIPELINE_DEVICE`.
|
||
- **Failure isolation**: layout-vis and spotting calls are best-effort
|
||
(`null`/absent on failure); classification and OCR errors surface as `error`
|
||
fields inside their sections rather than failing the whole scan.
|
||
|
||
## Known gaps & future recommendations
|
||
|
||
Tracked ones (see `plans/next-enhancements.md`):
|
||
- **Dataset thinness**: 2–16 photos/class caps both classifiers; every new real
|
||
photo (especially non-studio, in-warehouse shots) matters. The harness above
|
||
already reports gallery (training) vs. held-out (validation) accuracy
|
||
separately — but as of this writing `sources/product-test-images/` is empty,
|
||
so the Validation Set is still 0 images and every published number so far is
|
||
a training/memorization score. Dropping real photos there is the next step,
|
||
not yet done.
|
||
|
||
Additional recommendations (not yet tasks — promote via `e`/`n` when wanted):
|
||
1. ~~Use `extracted_sku` in match ranking.~~ **Done** — `product-scan.ts`'s
|
||
`classifyAndMatchProduct` already pins rank 1 to an exact `no_sku` match
|
||
(score forced to 1.0) before falling back to name similarity.
|
||
2. **Fuse DINOv2 and YOLO instead of primary/fallback** (e.g. agreement boosts
|
||
confidence; disagreement flags for review) — cheap, both already load.
|
||
3. **"Not a known product" handling**: DINOv2 always returns *some* class; add a
|
||
minimum-similarity threshold below which the response says unknown rather
|
||
than confidently misclassifying a foreign package.
|
||
4. **Pin the DINOv2 backbone offline** (vendor the weights or pre-bake the
|
||
torch.hub cache into the image) — startup currently depends on an internet
|
||
fetch on cold cache, bad for on-prem deploys.
|
||
5. **Batch/lot number extraction** — explicitly out of scope so far (plan §2
|
||
note); if requested, follow the expiry-date regex-cascade pattern.
|
||
6. **Mobile**: no web mobile page by design (task 2.2 cancelled) — real mobile
|
||
scanning should go through the Flutter app calling `POST /api/scan-pfm`
|
||
(would need an authenticated `/api/v1` variant; the classic route has no auth).
|