feat(backend): scan-product accuracy 66.2% -> 79.7% + frozen validation benchmark

Accuracy work on the 79-image product-scan validation set (user goal: 90%):
- classify_ocr_server.py: 0/90/180/270-degree expiry-date search (stops at
  first hit, 0-degree fallback); classification decoupled onto the upright
  image (rotated frames regressed DINOv2 -6pts until this); cross-line date
  stitching; tiled full-res OCR pass (defeats the 4000px downscale that
  killed small inkjet dates); VL-pipeline expiry fallback with
  keyword-anchored anti-hallucination guard; VL text lines merged into
  text_lines + VL SKU retry. Visualization endpoints removed entirely
  (Visual/Spotting grids - unused by frontend, 3x per-scan GPU cost).
- product-scan.ts: coverage-normalized OCR-evidence re-ranking of DINOv2
  top-K (tuned offline: +8/-0 on top-1 misses), re-ranked class mapped to
  sku_master by SKU prefix; classifier timeout 90s->240s for fallback paths.
- Frozen benchmark: product-test-images-fixed/ (79 renamed images) +
  freeze/seed/build-undetected/capture/experiment scripts; labels trimmed to
  the 79 validation entries (training rows kept in .bak-with-training);
  5 TRAINED-ON SKUs replaced with fresh held-out photos.
- manual-label-scan page: shows last batch-test AI prediction under every
  field by default (new /api/product-scan-results); serves the fixed folder;
  fixed total hydration failure via allowedDevOrigins 127.0.0.1.
- Measured (all-79, zero failures): sku/name 87.3%, expiry 64.6%, overall
  79.7%. Tiles/VL-evidence/VL-SKU deployed but not yet batch-measured.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gr6HH7JrdsXX8AARejQboM
This commit is contained in:
Rafhan Mazaya FathurrahmanandClaude Fable 5 committed 2026-07-14 19:55:17 +07:00
1 parent 19f1facf9b
commit e76ccb60a6
156 files changed
+17150 -1405

No files matched your search

+1
View File
@@ -55,6 +55,7 @@ workflow and have no task numbers; see `git log` for real dates/history.
- **2.1 (verification pass)** Ran a full browser walkthrough of `/scan-pfm` (classification, top-5, OCR expiry extraction + crop, SKU-master matching, Visual/Spotting Grid, Raw Response — all confirmed working with real data). Found and fixed a real bug: "Save Ground Truth" was returning success but silently writing into the `pfm-web-app` container's ephemeral filesystem instead of the host, because `/sources` wasn't a bind-mounted path in root `docker-compose.yml`. Added `./backend/sources:/sources` to the `pfm-web-app` service, recovered an orphaned entry via `docker cp`, and re-verified the save now persists to `backend/sources/product_manual_labels.json` on the host (confirmed the DO-flow's `manual_labels.json` save was fixed by the same change too) — shipped 2026-07-08.
- **2.3** Ran the accuracy regression harness and discovered `sources/accuracy_report.md` was badly stale (claimed 75.04%; real current baseline is **95.10% overall, already at/above the 95% target** — added a staleness banner to that file). Root-caused every remaining mismatch by pulling raw OCR text from Postgres (`documents.layout_parsing_result`): the worst field, `plat` (67.6%), is almost entirely the license-plate region being classified as an image/seal by the layout model rather than OCR'd as text — not fixable in `parser.ts`. Found and fixed one genuine parser logic bug along the way: the "global pattern scanning fallback" could duplicate an already-correctly-extracted `noDO` value into a still-missing `noSO` field; fixed by excluding already-assigned values from that fallback's candidate pool (`pfm-web-app/src/utils/parser.ts`). Doesn't change the aggregate score (a wrong value and "Not Found" score the same) but stops a fabricated-looking wrong number from silently reaching the database. All 48 parser unit tests still pass — shipped 2026-07-08.
- **Ad-hoc** Built custom expiry-date-based auto-rotation algorithm in Python classifier server (`classify_ocr_server.py`). The algorithm calculates the slant angle of the Expiry Date / Batch text line bounding box, automatically rotates the image to make it horizontal, and re-runs YOLO classification + PaddleOCR for maximum accuracy. Enhanced SKU matching database lookup to prioritize exact SKU matches with a score of 1.0, pinning them as the Best Match — shipped 2026-07-09.
- **2.5** Retrained the Product/SKU scan classifier's model artifacts against the full current dataset, which had grown to 81 SKU classes / 2,493 photos (up from the original 16 classes / 118 photos the deployed model dated 2026-07-08 was actually trained on — the other 65 classes had photos but no trained weights). Rebuilt `models/dinov2_index.pkl` (now 2,493/2,493 photos indexed) and retrained the YOLO classifier 100 epochs on an RTX 2060 (real elapsed time 54m21s), publishing `models/produk-pfm-classifier-26n-100e-2026-07-14.pt`/`.onnx` at **85.8% top-1 / 94.4% top-5** validation accuracy across all 81 classes (up from 83.3%/90% on the old 16-class model). Along the way, fixed a real train/val split bug in `train_classifier.py`: `split_dataset()` previously shuffled and split individual image files, letting an augmented copy (`photo_aug_2.jpeg`) land in validation while its near-duplicate source stayed in training — inflating val accuracy with memorization instead of measuring generalization; now groups by source photo (stripping `_aug_N`) before shuffling and splitting 80/20. Verified via `docker compose up -d pipeline-api` + `docker logs`: "DINOv2 index loaded with 2493 reference images", "Using classifier weights: .../produk-pfm-classifier-26n-100e-2026-07-14.pt", "YOLO model loaded successfully" — the live service is confirmed serving the new 81-class model, not assumed from the newest-file-by-date fallback logic. Remaining gap toward the program's ±230-SKU target is dataset growth, not a pipeline limitation — shipped 2026-07-14.
## Backend — Postgres Data Layer
+42 -32
View File
@@ -532,38 +532,48 @@ retraining" section), and record real timing/accuracy rather than estimates.
top-5** across all 81 classes — already ahead of the old 16-class model's
83.3%/90%, but not a final number since the run never reached completion.
- **Session resumed**: after the pause above, Docker Desktop had actually
stopped between sessions — a first resume attempt failed instantly with a
daemon-connection error before any training ran. Restarted Docker Desktop,
confirmed `docker ps` responsive, confirmed the `pipeline-api` image and
`dinov2_index.pkl` from the earlier session were both still intact (no
rebuild/reindex needed), then relaunched `train_classifier.py train
--imgsz 224` from epoch 0 in a fresh one-off container, timed with `time`.
- **Training ran to completion this time: 100/100 epochs, real elapsed time
54m21.248s.** Final validation: **85.8% top-1 / 94.4% top-5** across all 81
classes. Published `models/produk-pfm-classifier-26n-100e-2026-07-14.pt`
(3.4MB) and exported `.onnx` (6.3MB, ONNX opset 20).
## 3. Verification
- Confirmed via `docker ps -a` that the training container exited cleanly on
`docker stop` (no hang, no orphaned process).
- Confirmed via `ls` on the host `models/` directory that **no new dated
`.pt`/`.onnx` was written** — `train_model()` only calls
`shutil.copy2(best_weights, output_path)` after `model.train()` returns, so
an interrupted run correctly leaves the previously-deployed
`produk-pfm-classifier-26n-100e-2026-07-08.pt`/`.onnx` untouched. The live
classifier is unaffected by this session.
- Confirmed `dinov2_index.pkl` **is** updated on the host (4.2MB, timestamped
2026-07-14 06:47) — this step ran to completion before training started and
is unaffected by the training container being stopped afterward.
- Did **not** run `docker compose restart pipeline-api`, since there is no new
classifier checkpoint to pick up yet and the main compose stack wasn't even
running this session (confirmed via `docker ps -a`: `pfm-web-app`,
`vllm-server`, `nginx`, `postgres` were all `Exited` from a prior session,
untouched by this work).
- Confirmed via `ls` on the host `models/` directory that the new dated
`produk-pfm-classifier-26n-100e-2026-07-14.pt`/`.onnx` files exist (dated
2026-07-14 09:05/09:06), alongside the untouched 2026-07-08 files.
- Confirmed in the training log's own ONNX export step that the model's
output shape is `(1, 81)` — i.e. genuinely 81 output classes, not a stale
16-class head.
- Ran `docker compose up -d pipeline-api` (the main compose stack wasn't
running this session — confirmed via `docker compose ps` returning empty —
so this was a fresh start, not a "restart"; it correctly pulled in the
`vllm-server` dependency too) and polled `docker logs
paddleocr-pipeline-api` until startup markers appeared. Confirmed lines:
- `DINOv2 index loaded with 2493 reference images.`
- `Using classifier weights: /app/pfm-web-app/public/produk-pfm/models/produk-pfm-classifier-26n-100e-2026-07-14.pt`
- `YOLO model loaded successfully.`
- `INFO: Application startup complete.`
This is real, observed runtime behavior — the live `pipeline-api` service is
now actually serving the new 81-class model and the full 2,493-image
DINOv2 index, not an assumption based on `latest_classifier_weights()`'s
glob-newest-by-date logic.
## 4. Status
**Paused 2026-07-14, by user request — not complete, not abandoned.**
Done: Docker Desktop started, `pipeline-api` image built (2m54s), DINOv2 index
rebuilt and persisted (2,493/2,493 images, all 81 classes). Not done: the YOLO
classifier training run, which was intentionally interrupted at epoch 43/100
and left no partial checkpoint (container used `--rm`, and Ultralytics' own
per-epoch checkpoints live in the container's `runs/classify/`, which was
never bind-mounted to the host). **Resuming means restarting training from
epoch 0**, not continuing from 43 — the image doesn't need rebuilding and the
index doesn't need reindexing, only `train_classifier.py train --imgsz 224`
needs to run again. Observed pace (32s/epoch) suggests a full 100-epoch run
takes **~55 minutes** on this host's RTX 2060, revised down from the ~90 min
estimated off the first few (slower, warmup) epochs. `plans/next-enhancements.md`
task 2.5 records the same state in full; `docs/scan-product.md`,
`backend/CLAUDE.md`, and `docs/feature-list.md` are deliberately left
unchanged (still say 16 classes) until a real completed run justifies updating
them.
**Done, 2026-07-14.** Both artifacts (DINOv2 index, YOLO classifier) retrained
against the full 81-class/2,493-photo dataset and verified loading in the live
service. `plans/next-enhancements.md` task 2.5 flipped to `[DONE]` with these
same numbers; `docs/scan-product.md` and `backend/CLAUDE.md`'s class-count/
accuracy claims updated from 16→81 classes and 83.3%/90%→85.8%/94.4%;
`docs/feature-list.md` given a matching entry. Remaining gap toward the
program's stated ±230-SKU target (see `proposals/sources/` kick-off material)
is dataset growth, not a code or training-pipeline limitation — the same
`train_classifier.py`/`index_dinov2.py` procedure documented here scales to
however many classes `foto-kemasan-v2/` ends up containing.
+16 -11
View File
@@ -119,9 +119,9 @@ pipeline call with `promptLabel: "spotting"`, no layout detection).
| File (`pfm-web-app/public/produk-pfm/`) | What |
|---|---|
| `foto-kemasan-v2/<SKU or class>/…` | Reference photo dataset — 16 classes, 118 photos (2–16 each) |
| `models/dinov2_index.pkl` | DINOv2 embeddings + metadata (rebuild after adding photos) |
| `models/produk-pfm-classifier-26n-100e-2026-07-08.pt` / `.onnx` | Fine-tuned YOLO classifier (83.3% top-1 / 90% top-5 val on the thin dataset) |
| `foto-kemasan-v2/<SKU or class>/…` | Reference photo dataset — 81 classes, 2,493 photos (target ~230 SKU) |
| `models/dinov2_index.pkl` | DINOv2 embeddings + metadata (rebuild after adding photos) — currently indexes all 2,493 photos across 81 classes |
| `models/produk-pfm-classifier-26n-100e-2026-07-14.pt` / `.onnx` | Fine-tuned YOLO classifier (85.8% top-1 / 94.4% top-5 val across all 81 classes; retrained 2026-07-14, 54m21s on an RTX 2060, up from the prior 2026-07-08 model's 83.3%/90% on only 16 classes) |
| `index_dinov2.py` | Rebuilds the pickle index from `foto-kemasan-v2/` |
| `train_classifier.py` | Splits 80/20 into `yolo_dataset/`, fine-tunes `yolo26n-cls.pt` (default 100 epochs, `--imgsz 224`), writes a dated checkpoint |
@@ -144,10 +144,16 @@ every labeled image in `sources/product_manual_labels.json`, checks 3 fields
(`no_sku`, `nama_item`, `expiry_date`) against ground truth, and splits into:
- **Training Set** — gallery photos under `foto-kemasan-v2/` (the classifier's
own reference images; scores here measure memorization, not generalization).
- **Validation Set** — flat filenames dropped into
`sources/product-test-images/` (a real held-out set; see that folder's
`README.md` for the drop-photo → label → re-run workflow via
`/manual-label-scan`).
- **Validation Set** — flat filenames, scored from the frozen
`sources/product-test-images-fixed/` snapshot (renamed `<index> <no_sku>.<ext>`,
built by `scripts/freeze-validation-set.mjs`) so a rerun always grades the
same 79 images regardless of what's since been dropped into the live-intake
`sources/product-test-images/` folder. See each folder's `README.md` — the
live folder documents the drop-photo → label → re-run-freeze-script workflow
via `/manual-label-scan`; the fixed folder documents the freeze/promote step
and flags 5 SKUs (12010801, 12012504, 12130504, 13050101, 15040102) whose
only available photo was already used to train the classifier, so their
scores aren't a clean held-out result.
Every run appends to `sources/product_accuracy_history.jsonl` and **auto-diffs
against the previous run**: the printed summary shows a Δ column per field per
@@ -185,10 +191,9 @@ Tracked ones (see `plans/next-enhancements.md`):
- **Dataset thinness**: 2–16 photos/class caps both classifiers; every new real
photo (especially non-studio, in-warehouse shots) matters. The harness above
already reports gallery (training) vs. held-out (validation) accuracy
separately — but as of this writing `sources/product-test-images/` is empty,
so the Validation Set is still 0 images and every published number so far is
a training/memorization score. Dropping real photos there is the next step,
not yet done.
separately, and as of 2026-07-14 the Validation Set has 79 labeled images
(74 genuinely held out, 5 flagged trained-on — see above) — the first real
(non-zero) Validation Set numbers.
Additional recommendations (not yet tasks — promote via `e`/`n` when wanted):
1. ~~Use `extracted_sku` in match ranking.~~ **Done** — `product-scan.ts`'s