docs: add scan-product reference and update backend/root plans for enhancements
This commit is contained in:
1 parent
e60ab63154
commit
9ff4a4a922
4 files changed
+475
-2
No files matched your search
@@ -0,0 +1,185 @@
|
||||
# Product Scan (scan-pfm) — How It Works
|
||||
|
||||
End-to-end reference for the Product/SKU scanning feature: a photo of a Primafood
|
||||
product package goes in; the SKU class, product name, expiry date, and a ranked
|
||||
SKU-master match list come out. Written 2026-07-08 against the live code. Related:
|
||||
`plans/next-enhancements.md` §2 (build history) and §6 (ground-truth roadmap);
|
||||
`docs/feature-list.md` tasks 2.1/2.3.
|
||||
|
||||
## High-level flow
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
A[Browser: /scan-pfm page] -->|"POST /api/scan-pfm {image_base64}"| B[Next.js gateway<br/>pfm-web-app :3000]
|
||||
B -->|"POST :8120/classify-ocr"| C[classify_ocr_server.py<br/>FastAPI, in pipeline-api]
|
||||
C --> C1[1. DINOv2 similarity<br/>fallback: YOLO classifier]
|
||||
C --> C2[2. PaddleOCR + regex<br/>SKU / expiry / name]
|
||||
C -->|"POST localhost:8090/layout-parsing<br/>promptLabel: spotting"| D[PaddleX pipeline<br/>same container]
|
||||
B -->|"POST :8090/layout-parsing"| D
|
||||
B -->|"SELECT sku_master"| E[(Postgres)]
|
||||
B -->|Levenshtein ranking| A
|
||||
```
|
||||
|
||||
Two processes live in the `paddleocr-pipeline-api` container, both started by
|
||||
`scripts/serve-pipeline.sh`: the PaddleX layout-parsing pipeline on **:8090**
|
||||
(shared with the DO flow; VL recognition goes out to the vLLM server on :8118) and
|
||||
`config/classify_ocr_server.py` on **:8120** (product scan only). The gateway
|
||||
reaches them via Docker DNS (`CLASSIFIER_SERVER_URL`, `PIPELINE_URL` in root
|
||||
`docker-compose.yml:87-88`); nginx (:8000) proxies `/scan-pfm` to the Next.js app.
|
||||
|
||||
## Request walkthrough
|
||||
|
||||
1. **Page** (`pfm-web-app/src/app/scan-pfm/page.tsx`, desktop-only test UI): pick a
|
||||
sample from the gallery (`GET /api/produk-pfm`) or upload/rotate a photo (rotation
|
||||
is done client-side on a canvas), then send it as a base64 data-URL.
|
||||
2. **Gateway** (`api/scan-pfm/route.ts`):
|
||||
- forwards `{image_base64}` to the classifier server (`/classify-ocr`);
|
||||
- separately calls the layout-parsing pipeline with `useLayoutDetection: true`
|
||||
for the Visual Grid tab's output images (failure here is non-fatal — logged,
|
||||
`layoutParsingResult` returns `null`);
|
||||
- loads the full `sku_master` table and ranks every SKU by **Levenshtein
|
||||
similarity between `nama_item` and the classifier's `top1_name`**
|
||||
(lowercased, alphanumerics only). Top 5 with score > 0.1 are returned;
|
||||
rank 1 gets `isBestMatch: true`. Note: `ocr.extracted_sku` and
|
||||
`ocr.extracted_product_name` are read but **not used** in this ranking —
|
||||
see Future recommendations.
|
||||
3. **Classifier server** (`config/classify_ocr_server.py`) does classification,
|
||||
OCR extraction, and visualization — detailed below — and returns
|
||||
`{classification, ocr}`.
|
||||
4. **Page renders** four tabs: Summary (classification card + top-5 override
|
||||
"Use" buttons + OCR fields + SKU matches), Visual Grid, Spotting Grid, Raw
|
||||
Response (JSON). "Save Ground Truth" posts to `/api/manual-label-scan`.
|
||||
|
||||
## Stage 1 — classification (which product is this?)
|
||||
|
||||
**Primary: DINOv2 similarity search** (`method: "dinov2_similarity"`). At startup
|
||||
the server loads `dinov2_vits14` **from `torch.hub` (network fetch on first run)**
|
||||
plus `models/dinov2_index.pkl` — precomputed L2-normalized 384-dim embeddings of
|
||||
all 118 reference photos across 16 SKU class folders. Per request: embed the query
|
||||
image (resize 224², ImageNet normalization), dot-product against all reference
|
||||
embeddings (= cosine similarity), then aggregate **per class = max similarity of
|
||||
any reference photo in that class**. Classes sorted by similarity become
|
||||
`all_probabilities`. Caveat: these "confidences" are cosine similarities, **not
|
||||
probabilities** — they don't sum to 1 and are typically all high (0.4–0.9);
|
||||
compare relatively, not against an absolute threshold.
|
||||
|
||||
**Fallback: YOLO classifier** (`method: "yolo_classifier"`) — only when DINOv2 is
|
||||
unavailable (no index/model) or throws. A fine-tuned `yolo26n-cls` checkpoint;
|
||||
its `all_probabilities` are real softmax probabilities. Weights are
|
||||
**auto-discovered**: `CLASSIFIER_MODEL_PATH` env wins; otherwise the newest
|
||||
`produk-pfm-classifier-26n-*e-*.pt` in `models/` by (date-in-filename, mtime) —
|
||||
so retraining just drops a new dated file, no config change.
|
||||
|
||||
If both are unavailable, `classification` carries an `error` field instead.
|
||||
|
||||
## Stage 2 — OCR extraction (SKU, expiry date, product name)
|
||||
|
||||
PaddleOCR (`lang='en'`, textline orientation on) produces `rec_texts` lines +
|
||||
`rec_polys` boxes. Three extractors run over the lines:
|
||||
|
||||
- **SKU** (`extract_sku`): first 8-digit number anywhere; else first 7–9 digit
|
||||
number. (Primafood SKUs are 8 digits, printed near the label top.)
|
||||
- **Expiry date** (`extract_expired_date`): each line is first noise-cleaned
|
||||
(`clean_date_line`: `1)`→`0`, `()`→`0`, `B8/8B/88`→`BB` before digits, o→0,
|
||||
I/l/|→1, S→5, Z→2, B→8 when digit-flanked, plus `012`/`112` month-misread
|
||||
repairs), then a **6-level priority cascade** runs: (1) BB/EXP-keyword line
|
||||
with compact `DDMMYYYY`; (2) keyword line with spaced `DD MM YYYY`; (3)
|
||||
keyword + 6–8 digit run; (3.5) keyword line, lenient noisy match; (4) any line
|
||||
spaced date; (5) any line compact `DDMMYYYY` — skipping lines that look like a
|
||||
SKU-on-product-name; (6) legacy formats (slashes, `05 MAR 2027`). Recognized
|
||||
keywords: `EXP`, `EXPIRED`, `TGL`, `EXPIRY`, `BBD`, `BEST BEFORE`, `BB`,
|
||||
`BAIK DIGUNAKAN`. Output normalized to `DD/MM/YYYY`.
|
||||
- **Product name** (`extract_product_name`): longest line containing a brand/
|
||||
product keyword (FIESTA, CHAMP, OKEY, AKUMO, ASIMO, NUGGET, SOSIS, …) after
|
||||
stripping SKU digits and date fragments; falls back to the classifier's
|
||||
`top1_name`, then the longest non-numeric line, then `"Unknown Product"`.
|
||||
|
||||
Visualization artifacts built server-side: `vis_image_base64` (all OCR boxes
|
||||
drawn teal `TEXT`, the expiry line amber `EXP`, on the orientation-corrected
|
||||
image so boxes align), `expired_date_crop_base64` (padded crop of the expiry
|
||||
line for eyeball verification — `find_expired_crop_index` prefers the box whose
|
||||
digits actually contain the date), and `spotting_image_base64` (a second
|
||||
pipeline call with `promptLabel: "spotting"`, no layout detection).
|
||||
|
||||
## Endpoint reference
|
||||
|
||||
| Endpoint | Where | Purpose |
|
||||
|---|---|---|
|
||||
| `POST /api/scan-pfm` | gateway | Main scan. Body `{image_base64}` (data-URL ok). Returns `{classification, ocr, possibleMatches[], layoutParsingResult}` |
|
||||
| `POST http://paddleocr-pipeline-api:8120/classify-ocr` | classifier server | Internal. Body `{image_base64}`. Returns `{classification: {top1_name, top1_confidence, all_probabilities[], method}, ocr: {text_lines[], extracted_product_name, extracted_sku, extracted_expired_date, expired_line_index, expired_source_line, expired_date_crop_base64, vis_image_base64, spotting_image_base64}}` |
|
||||
| `GET /api/produk-pfm` | gateway | Gallery: SKU folders under `public/produk-pfm/foto-kemasan-v2/` with image + thumb URLs |
|
||||
| `GET/POST /api/manual-label-scan` | gateway | Ground-truth read/upsert to `sources/product_manual_labels.json` (host-visible via the `./backend/sources:/sources` mount) |
|
||||
| `POST :8090/layout-parsing` | pipeline | Shared PaddleX pipeline; used here for Visual Grid images and (with `promptLabel: "spotting"`) the Spotting Grid |
|
||||
| `/scan-pfm` | nginx :8000 | Proxies the page to Next.js :3000 |
|
||||
|
||||
`possibleMatches[]` items: `{no_sku, nama_item, score, yoloSimilarity, isBestMatch}` —
|
||||
`score` currently equals `yoloSimilarity` (name-vs-name Levenshtein, 0..1).
|
||||
|
||||
## Model artifacts & retraining
|
||||
|
||||
| File (`pfm-web-app/public/produk-pfm/`) | What |
|
||||
|---|---|
|
||||
| `foto-kemasan-v2/<SKU or class>/…` | Reference photo dataset — 16 classes, 118 photos (2–16 each) |
|
||||
| `models/dinov2_index.pkl` | DINOv2 embeddings + metadata (rebuild after adding photos) |
|
||||
| `models/produk-pfm-classifier-26n-100e-2026-07-08.pt` / `.onnx` | Fine-tuned YOLO classifier (83.3% top-1 / 90% top-5 val on the thin dataset) |
|
||||
| `index_dinov2.py` | Rebuilds the pickle index from `foto-kemasan-v2/` |
|
||||
| `train_classifier.py` | Splits 80/20 into `yolo_dataset/`, fine-tunes `yolo26n-cls.pt` (default 100 epochs, `--imgsz 224`), writes a dated checkpoint |
|
||||
|
||||
**Retraining procedure (Windows host — bare-metal doesn't work here,
|
||||
`paddlepaddle-gpu` wheels are Linux-only):** add photos to `foto-kemasan-v2/`,
|
||||
`docker compose build pipeline-api` from the **repo root**, run a one-off
|
||||
`docker run --gpus all` from that image with `models/` mounted **writable** (the
|
||||
live service mounts it `:ro`), run `index_dinov2.py` then
|
||||
`train_classifier.py train --imgsz 224`, then `docker compose restart
|
||||
pipeline-api`. From Git Bash prefix `MSYS_NO_PATHCONV=1` or `/app/...` arguments
|
||||
get mangled. Verify in `docker logs`: "DINOv2 index loaded with N reference
|
||||
images", "Using classifier weights: <new dated file>". Full worked example:
|
||||
`plans/next-enhancements.md` task 2.1.
|
||||
|
||||
## Operational notes
|
||||
|
||||
- **Env vars**: `CLASSIFIER_SERVER_URL`, `PIPELINE_URL` (gateway, set in compose);
|
||||
`CLASSIFIER_MODELS_DIR`, `CLASSIFIER_MODEL_PATH` (classifier server overrides).
|
||||
The gateway's in-code default `PIPELINE_URL` (`localhost:7871`) is stale — the
|
||||
compose env always overrides it in Docker.
|
||||
- **Startup order/health**: the classifier server loads DINOv2 (torch.hub →
|
||||
needs network/cache), YOLO, and PaddleOCR at import time; until done, :8120
|
||||
refuses connections and `/api/scan-pfm` 500s. No healthcheck exists yet (plan
|
||||
task 4.2 / 1.6).
|
||||
- **GPU**: DINOv2 + YOLO + PaddleOCR share the container/GPU with the PaddleX
|
||||
pipeline; all are small (ViT-S/14, nano YOLO) next to the vLLM server's
|
||||
footprint, but they do add VRAM on the same `PIPELINE_DEVICE`.
|
||||
- **Failure isolation**: layout-vis and spotting calls are best-effort
|
||||
(`null`/absent on failure); classification and OCR errors surface as `error`
|
||||
fields inside their sections rather than failing the whole scan.
|
||||
|
||||
## Known gaps & future recommendations
|
||||
|
||||
Tracked ones (see `plans/next-enhancements.md`):
|
||||
- **§6.1–6.3 ground truth**: today's "Save Ground Truth" stores the model's own
|
||||
predictions (only the SKU is editable) and uploads get phantom
|
||||
`uploaded-<timestamp>.jpg` keys with no image persisted; §6 plans the editable
|
||||
annotation page, persisted uploads, and a scan accuracy harness.
|
||||
- **Dataset thinness**: 2–16 photos/class caps both classifiers; every new real
|
||||
photo (especially non-studio, in-warehouse shots) matters. The §6.3 harness
|
||||
should report gallery vs. uploaded-photo accuracy separately — gallery photos
|
||||
are training data, so scores on them measure memorization.
|
||||
|
||||
Additional recommendations (not yet tasks — promote via `e`/`n` when wanted):
|
||||
1. **Use `extracted_sku` in match ranking.** An exact 8-digit SKU hit read off
|
||||
the label is far stronger evidence than fuzzy name similarity, yet ranking
|
||||
currently ignores it. Suggested: exact `no_sku` match pins rank 1; blend name
|
||||
similarity for the rest.
|
||||
2. **Fuse DINOv2 and YOLO instead of primary/fallback** (e.g. agreement boosts
|
||||
confidence; disagreement flags for review) — cheap, both already load.
|
||||
3. **"Not a known product" handling**: DINOv2 always returns *some* class; add a
|
||||
minimum-similarity threshold below which the response says unknown rather
|
||||
than confidently misclassifying a foreign package.
|
||||
4. **Pin the DINOv2 backbone offline** (vendor the weights or pre-bake the
|
||||
torch.hub cache into the image) — startup currently depends on an internet
|
||||
fetch on cold cache, bad for on-prem deploys.
|
||||
5. **Batch/lot number extraction** — explicitly out of scope so far (plan §2
|
||||
note); if requested, follow the expiry-date regex-cascade pattern.
|
||||
6. **Mobile**: no web mobile page by design (task 2.2 cancelled) — real mobile
|
||||
scanning should go through the Flutter app calling `POST /api/scan-pfm`
|
||||
(would need an authenticated `/api/v1` variant; the classic route has no auth).
|
||||
Reference in new issue
Block a user