feat(backend): scan-product accuracy 66.2% -> 79.7% + frozen validation benchmark

Accuracy work on the 79-image product-scan validation set (user goal: 90%):
- classify_ocr_server.py: 0/90/180/270-degree expiry-date search (stops at
  first hit, 0-degree fallback); classification decoupled onto the upright
  image (rotated frames regressed DINOv2 -6pts until this); cross-line date
  stitching; tiled full-res OCR pass (defeats the 4000px downscale that
  killed small inkjet dates); VL-pipeline expiry fallback with
  keyword-anchored anti-hallucination guard; VL text lines merged into
  text_lines + VL SKU retry. Visualization endpoints removed entirely
  (Visual/Spotting grids - unused by frontend, 3x per-scan GPU cost).
- product-scan.ts: coverage-normalized OCR-evidence re-ranking of DINOv2
  top-K (tuned offline: +8/-0 on top-1 misses), re-ranked class mapped to
  sku_master by SKU prefix; classifier timeout 90s->240s for fallback paths.
- Frozen benchmark: product-test-images-fixed/ (79 renamed images) +
  freeze/seed/build-undetected/capture/experiment scripts; labels trimmed to
  the 79 validation entries (training rows kept in .bak-with-training);
  5 TRAINED-ON SKUs replaced with fresh held-out photos.
- manual-label-scan page: shows last batch-test AI prediction under every
  field by default (new /api/product-scan-results); serves the fixed folder;
  fixed total hydration failure via allowedDevOrigins 127.0.0.1.
- Measured (all-79, zero failures): sku/name 87.3%, expiry 64.6%, overall
  79.7%. Tiles/VL-evidence/VL-SKU deployed but not yet batch-measured.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gr6HH7JrdsXX8AARejQboM
This commit is contained in:
Rafhan Mazaya FathurrahmanandClaude Fable 5 committed 2026-07-14 19:55:17 +07:00
1 parent 19f1facf9b
commit e76ccb60a6
156 files changed
+17150 -1405

No files matched your search

+42 -32
View File
@@ -532,38 +532,48 @@ retraining" section), and record real timing/accuracy rather than estimates.
top-5** across all 81 classes — already ahead of the old 16-class model's
83.3%/90%, but not a final number since the run never reached completion.
- **Session resumed**: after the pause above, Docker Desktop had actually
stopped between sessions — a first resume attempt failed instantly with a
daemon-connection error before any training ran. Restarted Docker Desktop,
confirmed `docker ps` responsive, confirmed the `pipeline-api` image and
`dinov2_index.pkl` from the earlier session were both still intact (no
rebuild/reindex needed), then relaunched `train_classifier.py train
--imgsz 224` from epoch 0 in a fresh one-off container, timed with `time`.
- **Training ran to completion this time: 100/100 epochs, real elapsed time
54m21.248s.** Final validation: **85.8% top-1 / 94.4% top-5** across all 81
classes. Published `models/produk-pfm-classifier-26n-100e-2026-07-14.pt`
(3.4MB) and exported `.onnx` (6.3MB, ONNX opset 20).
## 3. Verification
- Confirmed via `docker ps -a` that the training container exited cleanly on
`docker stop` (no hang, no orphaned process).
- Confirmed via `ls` on the host `models/` directory that **no new dated
`.pt`/`.onnx` was written** — `train_model()` only calls
`shutil.copy2(best_weights, output_path)` after `model.train()` returns, so
an interrupted run correctly leaves the previously-deployed
`produk-pfm-classifier-26n-100e-2026-07-08.pt`/`.onnx` untouched. The live
classifier is unaffected by this session.
- Confirmed `dinov2_index.pkl` **is** updated on the host (4.2MB, timestamped
2026-07-14 06:47) — this step ran to completion before training started and
is unaffected by the training container being stopped afterward.
- Did **not** run `docker compose restart pipeline-api`, since there is no new
classifier checkpoint to pick up yet and the main compose stack wasn't even
running this session (confirmed via `docker ps -a`: `pfm-web-app`,
`vllm-server`, `nginx`, `postgres` were all `Exited` from a prior session,
untouched by this work).
- Confirmed via `ls` on the host `models/` directory that the new dated
`produk-pfm-classifier-26n-100e-2026-07-14.pt`/`.onnx` files exist (dated
2026-07-14 09:05/09:06), alongside the untouched 2026-07-08 files.
- Confirmed in the training log's own ONNX export step that the model's
output shape is `(1, 81)` — i.e. genuinely 81 output classes, not a stale
16-class head.
- Ran `docker compose up -d pipeline-api` (the main compose stack wasn't
running this session — confirmed via `docker compose ps` returning empty —
so this was a fresh start, not a "restart"; it correctly pulled in the
`vllm-server` dependency too) and polled `docker logs
paddleocr-pipeline-api` until startup markers appeared. Confirmed lines:
- `DINOv2 index loaded with 2493 reference images.`
- `Using classifier weights: /app/pfm-web-app/public/produk-pfm/models/produk-pfm-classifier-26n-100e-2026-07-14.pt`
- `YOLO model loaded successfully.`
- `INFO: Application startup complete.`
This is real, observed runtime behavior — the live `pipeline-api` service is
now actually serving the new 81-class model and the full 2,493-image
DINOv2 index, not an assumption based on `latest_classifier_weights()`'s
glob-newest-by-date logic.
## 4. Status
**Paused 2026-07-14, by user request — not complete, not abandoned.**
Done: Docker Desktop started, `pipeline-api` image built (2m54s), DINOv2 index
rebuilt and persisted (2,493/2,493 images, all 81 classes). Not done: the YOLO
classifier training run, which was intentionally interrupted at epoch 43/100
and left no partial checkpoint (container used `--rm`, and Ultralytics' own
per-epoch checkpoints live in the container's `runs/classify/`, which was
never bind-mounted to the host). **Resuming means restarting training from
epoch 0**, not continuing from 43 — the image doesn't need rebuilding and the
index doesn't need reindexing, only `train_classifier.py train --imgsz 224`
needs to run again. Observed pace (32s/epoch) suggests a full 100-epoch run
takes **~55 minutes** on this host's RTX 2060, revised down from the ~90 min
estimated off the first few (slower, warmup) epochs. `plans/next-enhancements.md`
task 2.5 records the same state in full; `docs/scan-product.md`,
`backend/CLAUDE.md`, and `docs/feature-list.md` are deliberately left
unchanged (still say 16 classes) until a real completed run justifies updating
them.
**Done, 2026-07-14.** Both artifacts (DINOv2 index, YOLO classifier) retrained
against the full 81-class/2,493-photo dataset and verified loading in the live
service. `plans/next-enhancements.md` task 2.5 flipped to `[DONE]` with these
same numbers; `docs/scan-product.md` and `backend/CLAUDE.md`'s class-count/
accuracy claims updated from 16→81 classes and 83.3%/90%→85.8%/94.4%;
`docs/feature-list.md` given a matching entry. Remaining gap toward the
program's stated ±230-SKU target (see `proposals/sources/` kick-off material)
is dataset growth, not a code or training-pipeline limitation — the same
`train_classifier.py`/`index_dinov2.py` procedure documented here scales to
however many classes `foto-kemasan-v2/` ends up containing.