feat(backend): scan-product accuracy 66.2% -> 79.7% + frozen validation benchmark

Accuracy work on the 79-image product-scan validation set (user goal: 90%):
- classify_ocr_server.py: 0/90/180/270-degree expiry-date search (stops at
  first hit, 0-degree fallback); classification decoupled onto the upright
  image (rotated frames regressed DINOv2 -6pts until this); cross-line date
  stitching; tiled full-res OCR pass (defeats the 4000px downscale that
  killed small inkjet dates); VL-pipeline expiry fallback with
  keyword-anchored anti-hallucination guard; VL text lines merged into
  text_lines + VL SKU retry. Visualization endpoints removed entirely
  (Visual/Spotting grids - unused by frontend, 3x per-scan GPU cost).
- product-scan.ts: coverage-normalized OCR-evidence re-ranking of DINOv2
  top-K (tuned offline: +8/-0 on top-1 misses), re-ranked class mapped to
  sku_master by SKU prefix; classifier timeout 90s->240s for fallback paths.
- Frozen benchmark: product-test-images-fixed/ (79 renamed images) +
  freeze/seed/build-undetected/capture/experiment scripts; labels trimmed to
  the 79 validation entries (training rows kept in .bak-with-training);
  5 TRAINED-ON SKUs replaced with fresh held-out photos.
- manual-label-scan page: shows last batch-test AI prediction under every
  field by default (new /api/product-scan-results); serves the fixed folder;
  fixed total hydration failure via allowedDevOrigins 127.0.0.1.
- Measured (all-79, zero failures): sku/name 87.3%, expiry 64.6%, overall
  79.7%. Tiles/VL-evidence/VL-SKU deployed but not yet batch-measured.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gr6HH7JrdsXX8AARejQboM
This commit is contained in:
Rafhan Mazaya FathurrahmanandClaude Fable 5 committed 2026-07-14 19:55:17 +07:00
1 parent 19f1facf9b
commit e76ccb60a6
156 files changed
+17150 -1405

No files matched your search

+27 -33
View File
@@ -113,11 +113,11 @@ in either project (only SKU, product name, expiry date are extracted) — if
requested later, follow the same OCR-regex-cascade pattern already used for
expiry-date extraction.*
- **2.5** [IN PROGRESS 2026-07-14 — resumed, training run 2] **Retrain
classifier on the now-81-class dataset.** (Note: a first resume attempt
- **2.5** [DONE 2026-07-14] **Retrain classifier on the now-81-class
dataset.** (Note: a first resume attempt
failed instantly with a Docker daemon connection error — Docker Desktop had
stopped between sessions — before any training happened; restarted Docker
Desktop and relaunched. This is the actual second training attempt,
Desktop and relaunched. The successful run was the second attempt,
confirmed running via `docker ps`.) `foto-kemasan-v2/` grew from the 16
classes/118 photos the deployed model
(`produk-pfm-classifier-26n-100e-2026-07-08.pt`) was trained on to **81
@@ -138,36 +138,30 @@ expiry-date extraction.*
`docker compose restart pipeline-api` → verify via `docker logs` for
"DINOv2 index loaded with N reference images" and "Using classifier
weights: <new dated file>".
- **Status as of pause (2026-07-14)** — mixed state, read carefully before
resuming:
- ✅ `pipeline-api` image built (2m54s), bakes in the current 81-class
dataset.
- ✅ **`dinov2_index.pkl` already rebuilt and persisted to disk** —
"Success! Indexed 2493/2493 images" across all 81 classes. This artifact
is live on the host now (`models/dinov2_index.pkl`, 4.2MB, dated
2026-07-14) and does **not** need to be redone.
- ⏸️ **YOLO classifier training was started, then stopped by user request
at epoch 43/100 (~23 minutes in)** before it could write a new dated
checkpoint. `docker run` used `--rm` and the in-progress epoch
checkpoints live only in the container's own `runs/classify/` (not
bind-mounted), so **stopping the container discarded that partial
progress** — resuming means restarting from epoch 0, not continuing from
43. `models/` on the host still has only the original
`produk-pfm-classifier-26n-100e-2026-07-08.pt`/`.onnx` (16-class model) —
**the live/deployed classifier is unchanged**, still 16 classes.
- Observed pace before stopping: ~32s/epoch (43 epochs in 23m1s) → a full
100-epoch run should take **~55 minutes** on this host's RTX 2060 (6GB
VRAM), not the ~90 min extrapolated from the first few (slower, warmup)
epochs. At epoch 42 the in-progress run had already reached 84.3%
top-1 / 93.9% top-5 val accuracy across all 81 classes, ahead of the old
16-class model's 83.3%/90% — a promising sign for the eventual full run,
but not a final result since training didn't finish.
- **To resume**: image is already built and the DINOv2 index step can be
skipped — just re-run the one-off `train_classifier.py train --imgsz 224`
container, then `docker compose restart pipeline-api` and verify via
`docker logs`. Update the class count in `docs/scan-product.md`,
`CLAUDE.md`, and `docs/feature-list.md` (and flip this task to `[DONE]`)
only once that run actually completes with a final dated `.pt`/`.onnx`.
- **Final result (2026-07-14)** — both artifacts retrained and live:
- ✅ `dinov2_index.pkl` rebuilt and persisted to disk — "Success! Indexed
2493/2493 images" across all 81 classes (`models/dinov2_index.pkl`,
4.2MB, dated 2026-07-14).
- ✅ **YOLO classifier retrained to completion, 100/100 epochs, real
elapsed time 54m21s** (a first attempt was intentionally stopped by user
request at epoch 43/100 to pause the session; that partial progress was
discarded since `docker run --rm`'s in-container `runs/classify/`
checkpoints aren't bind-mounted, so the successful run below restarted
cleanly from epoch 0 rather than resuming from 43). Final validation:
**85.8% top-1 / 94.4% top-5** across all 81 classes — up from the old
16-class model's 83.3%/90%, now covering 5x the product classes.
Published artifacts: `models/produk-pfm-classifier-26n-100e-2026-07-14.pt`
(3.4MB) and matching `.onnx` (6.3MB, ONNX opset 20, output shape
confirmed `(1, 81)` — i.e. 81 output classes).
- ✅ Verified via `docker compose up -d pipeline-api` (main stack wasn't
running this session) + `docker logs paddleocr-pipeline-api`: "DINOv2
index loaded with 2493 reference images", "Using classifier weights:
/app/pfm-web-app/public/produk-pfm/models/produk-pfm-classifier-26n-100e-2026-07-14.pt",
"YOLO model loaded successfully", "Application startup complete" — the
live service is now serving the new 81-class model, not a code-review
assumption.
- Class-count claims updated in `docs/scan-product.md` and `../CLAUDE.md`
(16 → 81 classes); this task's `docs/feature-list.md` entry added.
## 3. Backend — Postgres Data Layer
`pfm-web-app/src/db/`