Files
pfm-ocr/backend/docs/scan-product.md
T

186 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Product Scan (scan-pfm) — How It Works
End-to-end reference for the Product/SKU scanning feature: a photo of a Primafood
product package goes in; the SKU class, product name, expiry date, and a ranked
SKU-master match list come out. Written 2026-07-08 against the live code. Related:
`plans/next-enhancements.md` §2 (build history) and §6 (ground-truth roadmap);
`docs/feature-list.md` tasks 2.1/2.3.
## High-level flow
```mermaid
flowchart LR
A[Browser: /scan-pfm page] -->|"POST /api/scan-pfm {image_base64}"| B[Next.js gateway<br/>pfm-web-app :3000]
B -->|"POST :8120/classify-ocr"| C[classify_ocr_server.py<br/>FastAPI, in pipeline-api]
C --> C1[1. DINOv2 similarity<br/>fallback: YOLO classifier]
C --> C2[2. PaddleOCR + regex<br/>SKU / expiry / name]
C -->|"POST localhost:8090/layout-parsing<br/>promptLabel: spotting"| D[PaddleX pipeline<br/>same container]
B -->|"POST :8090/layout-parsing"| D
B -->|"SELECT sku_master"| E[(Postgres)]
B -->|Levenshtein ranking| A
```
Two processes live in the `paddleocr-pipeline-api` container, both started by
`scripts/serve-pipeline.sh`: the PaddleX layout-parsing pipeline on **:8090**
(shared with the DO flow; VL recognition goes out to the vLLM server on :8118) and
`config/classify_ocr_server.py` on **:8120** (product scan only). The gateway
reaches them via Docker DNS (`CLASSIFIER_SERVER_URL`, `PIPELINE_URL` in root
`docker-compose.yml:87-88`); nginx (:8000) proxies `/scan-pfm` to the Next.js app.
## Request walkthrough
1. **Page** (`pfm-web-app/src/app/scan-pfm/page.tsx`, desktop-only test UI): pick a
sample from the gallery (`GET /api/produk-pfm`) or upload/rotate a photo (rotation
is done client-side on a canvas), then send it as a base64 data-URL.
2. **Gateway** (`api/scan-pfm/route.ts`):
- forwards `{image_base64}` to the classifier server (`/classify-ocr`);
- separately calls the layout-parsing pipeline with `useLayoutDetection: true`
for the Visual Grid tab's output images (failure here is non-fatal — logged,
`layoutParsingResult` returns `null`);
- loads the full `sku_master` table and ranks every SKU by **Levenshtein
similarity between `nama_item` and the classifier's `top1_name`**
(lowercased, alphanumerics only). Top 5 with score > 0.1 are returned;
rank 1 gets `isBestMatch: true`. Note: `ocr.extracted_sku` and
`ocr.extracted_product_name` are read but **not used** in this ranking —
see Future recommendations.
3. **Classifier server** (`config/classify_ocr_server.py`) does classification,
OCR extraction, and visualization — detailed below — and returns
`{classification, ocr}`.
4. **Page renders** four tabs: Summary (classification card + top-5 override
"Use" buttons + OCR fields + SKU matches), Visual Grid, Spotting Grid, Raw
Response (JSON). "Save Ground Truth" posts to `/api/manual-label-scan`.
## Stage 1 — classification (which product is this?)
**Primary: DINOv2 similarity search** (`method: "dinov2_similarity"`). At startup
the server loads `dinov2_vits14` **from `torch.hub` (network fetch on first run)**
plus `models/dinov2_index.pkl` — precomputed L2-normalized 384-dim embeddings of
all 118 reference photos across 16 SKU class folders. Per request: embed the query
image (resize 224², ImageNet normalization), dot-product against all reference
embeddings (= cosine similarity), then aggregate **per class = max similarity of
any reference photo in that class**. Classes sorted by similarity become
`all_probabilities`. Caveat: these "confidences" are cosine similarities, **not
probabilities** — they don't sum to 1 and are typically all high (0.4–0.9);
compare relatively, not against an absolute threshold.
**Fallback: YOLO classifier** (`method: "yolo_classifier"`) — only when DINOv2 is
unavailable (no index/model) or throws. A fine-tuned `yolo26n-cls` checkpoint;
its `all_probabilities` are real softmax probabilities. Weights are
**auto-discovered**: `CLASSIFIER_MODEL_PATH` env wins; otherwise the newest
`produk-pfm-classifier-26n-*e-*.pt` in `models/` by (date-in-filename, mtime) —
so retraining just drops a new dated file, no config change.
If both are unavailable, `classification` carries an `error` field instead.
## Stage 2 — OCR extraction (SKU, expiry date, product name)
PaddleOCR (`lang='en'`, textline orientation on) produces `rec_texts` lines +
`rec_polys` boxes. Three extractors run over the lines:
- **SKU** (`extract_sku`): first 8-digit number anywhere; else first 7–9 digit
number. (Primafood SKUs are 8 digits, printed near the label top.)
- **Expiry date** (`extract_expired_date`): each line is first noise-cleaned
(`clean_date_line`: `1)`→`0`, `()`→`0`, `B8/8B/88`→`BB` before digits, o→0,
I/l/|→1, S→5, Z→2, B→8 when digit-flanked, plus `012`/`112` month-misread
repairs), then a **6-level priority cascade** runs: (1) BB/EXP-keyword line
with compact `DDMMYYYY`; (2) keyword line with spaced `DD MM YYYY`; (3)
keyword + 6–8 digit run; (3.5) keyword line, lenient noisy match; (4) any line
spaced date; (5) any line compact `DDMMYYYY` — skipping lines that look like a
SKU-on-product-name; (6) legacy formats (slashes, `05 MAR 2027`). Recognized
keywords: `EXP`, `EXPIRED`, `TGL`, `EXPIRY`, `BBD`, `BEST BEFORE`, `BB`,
`BAIK DIGUNAKAN`. Output normalized to `DD/MM/YYYY`.
- **Product name** (`extract_product_name`): longest line containing a brand/
product keyword (FIESTA, CHAMP, OKEY, AKUMO, ASIMO, NUGGET, SOSIS, …) after
stripping SKU digits and date fragments; falls back to the classifier's
`top1_name`, then the longest non-numeric line, then `"Unknown Product"`.
Visualization artifacts built server-side: `vis_image_base64` (all OCR boxes
drawn teal `TEXT`, the expiry line amber `EXP`, on the orientation-corrected
image so boxes align), `expired_date_crop_base64` (padded crop of the expiry
line for eyeball verification — `find_expired_crop_index` prefers the box whose
digits actually contain the date), and `spotting_image_base64` (a second
pipeline call with `promptLabel: "spotting"`, no layout detection).
## Endpoint reference
| Endpoint | Where | Purpose |
|---|---|---|
| `POST /api/scan-pfm` | gateway | Main scan. Body `{image_base64}` (data-URL ok). Returns `{classification, ocr, possibleMatches[], layoutParsingResult}` |
| `POST http://paddleocr-pipeline-api:8120/classify-ocr` | classifier server | Internal. Body `{image_base64}`. Returns `{classification: {top1_name, top1_confidence, all_probabilities[], method}, ocr: {text_lines[], extracted_product_name, extracted_sku, extracted_expired_date, expired_line_index, expired_source_line, expired_date_crop_base64, vis_image_base64, spotting_image_base64}}` |
| `GET /api/produk-pfm` | gateway | Gallery: SKU folders under `public/produk-pfm/foto-kemasan-v2/` with image + thumb URLs |
| `GET/POST /api/manual-label-scan` | gateway | Ground-truth read/upsert to `sources/product_manual_labels.json` (host-visible via the `./backend/sources:/sources` mount) |
| `POST :8090/layout-parsing` | pipeline | Shared PaddleX pipeline; used here for Visual Grid images and (with `promptLabel: "spotting"`) the Spotting Grid |
| `/scan-pfm` | nginx :8000 | Proxies the page to Next.js :3000 |
`possibleMatches[]` items: `{no_sku, nama_item, score, yoloSimilarity, isBestMatch}` —
`score` currently equals `yoloSimilarity` (name-vs-name Levenshtein, 0..1).
## Model artifacts & retraining
| File (`pfm-web-app/public/produk-pfm/`) | What |
|---|---|
| `foto-kemasan-v2/<SKU or class>/…` | Reference photo dataset — 16 classes, 118 photos (2–16 each) |
| `models/dinov2_index.pkl` | DINOv2 embeddings + metadata (rebuild after adding photos) |
| `models/produk-pfm-classifier-26n-100e-2026-07-08.pt` / `.onnx` | Fine-tuned YOLO classifier (83.3% top-1 / 90% top-5 val on the thin dataset) |
| `index_dinov2.py` | Rebuilds the pickle index from `foto-kemasan-v2/` |
| `train_classifier.py` | Splits 80/20 into `yolo_dataset/`, fine-tunes `yolo26n-cls.pt` (default 100 epochs, `--imgsz 224`), writes a dated checkpoint |
**Retraining procedure (Windows host — bare-metal doesn't work here,
`paddlepaddle-gpu` wheels are Linux-only):** add photos to `foto-kemasan-v2/`,
`docker compose build pipeline-api` from the **repo root**, run a one-off
`docker run --gpus all` from that image with `models/` mounted **writable** (the
live service mounts it `:ro`), run `index_dinov2.py` then
`train_classifier.py train --imgsz 224`, then `docker compose restart
pipeline-api`. From Git Bash prefix `MSYS_NO_PATHCONV=1` or `/app/...` arguments
get mangled. Verify in `docker logs`: "DINOv2 index loaded with N reference
images", "Using classifier weights: <new dated file>". Full worked example:
`plans/next-enhancements.md` task 2.1.
## Operational notes
- **Env vars**: `CLASSIFIER_SERVER_URL`, `PIPELINE_URL` (gateway, set in compose);
`CLASSIFIER_MODELS_DIR`, `CLASSIFIER_MODEL_PATH` (classifier server overrides).
The gateway's in-code default `PIPELINE_URL` (`localhost:7871`) is stale — the
compose env always overrides it in Docker.
- **Startup order/health**: the classifier server loads DINOv2 (torch.hub →
needs network/cache), YOLO, and PaddleOCR at import time; until done, :8120
refuses connections and `/api/scan-pfm` 500s. No healthcheck exists yet (plan
task 4.2 / 1.6).
- **GPU**: DINOv2 + YOLO + PaddleOCR share the container/GPU with the PaddleX
pipeline; all are small (ViT-S/14, nano YOLO) next to the vLLM server's
footprint, but they do add VRAM on the same `PIPELINE_DEVICE`.
- **Failure isolation**: layout-vis and spotting calls are best-effort
(`null`/absent on failure); classification and OCR errors surface as `error`
fields inside their sections rather than failing the whole scan.
## Known gaps & future recommendations
Tracked ones (see `plans/next-enhancements.md`):
- **§6.1–6.3 ground truth**: today's "Save Ground Truth" stores the model's own
predictions (only the SKU is editable) and uploads get phantom
`uploaded-<timestamp>.jpg` keys with no image persisted; §6 plans the editable
annotation page, persisted uploads, and a scan accuracy harness.
- **Dataset thinness**: 2–16 photos/class caps both classifiers; every new real
photo (especially non-studio, in-warehouse shots) matters. The §6.3 harness
should report gallery vs. uploaded-photo accuracy separately — gallery photos
are training data, so scores on them measure memorization.
Additional recommendations (not yet tasks — promote via `e`/`n` when wanted):
1. **Use `extracted_sku` in match ranking.** An exact 8-digit SKU hit read off
the label is far stronger evidence than fuzzy name similarity, yet ranking
currently ignores it. Suggested: exact `no_sku` match pins rank 1; blend name
similarity for the rest.
2. **Fuse DINOv2 and YOLO instead of primary/fallback** (e.g. agreement boosts
confidence; disagreement flags for review) — cheap, both already load.
3. **"Not a known product" handling**: DINOv2 always returns *some* class; add a
minimum-similarity threshold below which the response says unknown rather
than confidently misclassifying a foreign package.
4. **Pin the DINOv2 backbone offline** (vendor the weights or pre-bake the
torch.hub cache into the image) — startup currently depends on an internet
fetch on cold cache, bad for on-prem deploys.
5. **Batch/lot number extraction** — explicitly out of scope so far (plan §2
note); if requested, follow the expiry-date regex-cascade pattern.
6. **Mobile**: no web mobile page by design (task 2.2 cancelled) — real mobile
scanning should go through the Flutter app calling `POST /api/scan-pfm`
(would need an authenticated `/api/v1` variant; the classic route has no auth).