Accuracy work on the 79-image product-scan validation set (user goal: 90%): - classify_ocr_server.py: 0/90/180/270-degree expiry-date search (stops at first hit, 0-degree fallback); classification decoupled onto the upright image (rotated frames regressed DINOv2 -6pts until this); cross-line date stitching; tiled full-res OCR pass (defeats the 4000px downscale that killed small inkjet dates); VL-pipeline expiry fallback with keyword-anchored anti-hallucination guard; VL text lines merged into text_lines + VL SKU retry. Visualization endpoints removed entirely (Visual/Spotting grids - unused by frontend, 3x per-scan GPU cost). - product-scan.ts: coverage-normalized OCR-evidence re-ranking of DINOv2 top-K (tuned offline: +8/-0 on top-1 misses), re-ranked class mapped to sku_master by SKU prefix; classifier timeout 90s->240s for fallback paths. - Frozen benchmark: product-test-images-fixed/ (79 renamed images) + freeze/seed/build-undetected/capture/experiment scripts; labels trimmed to the 79 validation entries (training rows kept in .bak-with-training); 5 TRAINED-ON SKUs replaced with fresh held-out photos. - manual-label-scan page: shows last batch-test AI prediction under every field by default (new /api/product-scan-results); serves the fixed folder; fixed total hydration failure via allowedDevOrigins 127.0.0.1. - Measured (all-79, zero failures): sku/name 87.3%, expiry 64.6%, overall 79.7%. Tiles/VL-evidence/VL-SKU deployed but not yet batch-measured. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gr6HH7JrdsXX8AARejQboM
9.9 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
What this repo is
app-pfm-ocr-v2/backend is the next-generation rewrite of ai-ocr-pfm-2026 — same underlying OCR infra (PaddleOCR-VL-1.6 on vLLM + a PaddlePaddle layout-parsing pipeline), same client (Charoen Pokphand/Primafood-branded frozen food products), but a reworked Next.js app (pfm-web-app/) and DB schema. If you need background on the shared OCR/vLLM infra (uv conventions, issue-recording workflow, GPU tuning), see AGENTS.md — it's carried over near-unchanged from the previous project.
The active plan for porting the Product/SKU-scanning feature lives in plans/next-enhancements.md §2 — read it before touching anything related to scan-pfm, produk-pfm, or the product classifier, since it records exactly what's done vs. still missing and the decisions already made about how to build it. (This used to be a separate next-implementation.md; that file was deleted 2026-07-08 once its content was folded into the plan for traceability with the rest of the e/n backlog.)
How this project differs from ai-ocr-pfm-2026
- DO-PFM UI is consolidated into a single page. Unlike the old project's per-route pages (
do-pfm/page.tsx,m-do-pfm/page.tsx), v2's entire upload/history/item-review flow lives in onepfm-web-app/src/app/page.tsx(client component, local state, no separate routes).nginx.confstill has/do-pfm//m-do-pfmlocation blocks left over from the old routing — these are currently dead (no matching Next.js route, would 404). - Standalone-purpose pages still get their own route folder, e.g.
pfm-web-app/src/app/manual-label/page.tsx— a self-contained ground-truth annotation tool (own header, own theme, no shared chrome with the root page) backed byapi/manual-label/route.tsandsources/manual_labels.json. This is the pattern to follow for any new single-purpose page (seeplans/next-enhancements.md§2 for the Product-scan pages, which follow it). - Real JWT auth, enforced on the production surface:
src/utils/auth.tssigns/verifies tokens (signAccountToken/verifyAccountToken/getAccountFromAuthHeader) against anaccountstable, each account bound to exactly onekode_toko(store) — the intent being that an account's own store is used on upload instead of relying on OCR-based store-text matching. Passwords are bcrypt-hashed (accounts.password, viabcryptjs— chosen over nativebcryptsince thepfm-web-appDocker stage isnode:20-slimwith no build toolchain for native addons;pfm-web-app/src/db/init.tshashes the seed and idempotently migrates any pre-existing plaintext rows on every startup). As of 2026-07-08,api/v1/documents/*(list, PUT-by-id, upload) reject requests with a missing/invalid token (401) — this is the real production surface, and the Flutter client already does a real login and attachesAuthorization: Bearer <token>to every request (lib/features/auth/auth_provider.dart+lib/core/network/api_client.dart). The classic routes (/api/upload,/api/scan-pfm,/api/parse,/api/history, etc.) and the root/scan-pfm/manual-labelpages deliberately do not check auth at all and never will unless that decision changes — they're dev-only web UI with no login screen, not part of the production surface (seeplans/next-enhancements.mdtask 1.3, cancelled, and 1.4, shipped instead). - Richer SKU master data:
pfm-web-app/import_sku.jsimports from a TSV with extended packaging columns (standar_jumlah,berat_kemasan,isi_outer_kg,isi_outer_pac,jenis_outer) added viaALTER TABLE sku_master ADD COLUMN IF NOT EXISTS, superseding the old project's bareno_sku/nama_itemseed list. - Accuracy regression harness (new, doesn't exist in the old project):
pfm-web-app/scripts/accuracy-check.mtshits the live/api/parseendpoint for every image insources/test-images/, diffs against hand-labeled ground truth insources/manual_labels.jsonat three post-processing stages (layer1RawRegex→layer2Sanitized→layer3Final— trace these stage names intoutils/parser.tsto see where each is produced), and appends run-over-run results tosources/accuracy_history.jsonl. Run this after touchingparser.tsto check for regressions:node pfm-web-app/scripts/accuracy-check.mts # reuse cached OCR (fast) node pfm-web-app/scripts/accuracy-check.mts --refresh-ocr # force fresh pipeline run node pfm-web-app/scripts/accuracy-check.mts --detail <filename> # full per-stage breakdown for one imagecompare_accuracy.py/compare_sources_accuracy.py/generate_excel.pyat the repo root build human-readable Excel/HTML comparison reports from the same data (sources/comparison_report.xlsx,sources/comparison_side_by_side.html) — these are analysis tooling, not part of the running app. api/vllm-proxy/[[...path]]/route.ts: a passthrough proxy to the vLLM server (paddleocr-vllm-server:8118) that logs every call vialogVllmCallToAll(utils/active-log.ts) — used for debugging/observability, not part of the OCR pipeline itself.docker-compose.override.ymlexposesdb(5432) andpipeline-api(8090) directly to the host for local dev — not present in the old project's compose setup.
Product/SKU scanning flow — status
How it works end-to-end (architecture, endpoints, classification/OCR internals, retraining): docs/scan-product.md. See plans/next-enhancements.md §2 (task 2.1) for full detail — kept there instead of a separate doc so status stays traceable against the rest of the e/n backlog. Feature-complete as of 2026-07-08: the backend (config/classify_ocr_server.py with DINOv2 similarity search + YOLO classifier fallback, api/scan-pfm/route.ts, api/produk-pfm/route.ts, DB schema), the reference photo dataset (pfm-web-app/public/produk-pfm/foto-kemasan-v2/, 81 SKU subfolders as of 2026-07-14, up from the original 16 — target ~230), the desktop frontend page (scan-pfm/page.tsx, full feature parity), and the trained model artifacts (models/dinov2_index.pkl — 2,493/2,493 photos indexed as of 2026-07-14; models/produk-pfm-classifier-26n-100e-2026-07-14.pt — 85.8% top-1 / 94.4% top-5 val accuracy across all 81 classes, retrained 2026-07-14 in 54m21s on an RTX 2060) all now exist and load cleanly on pipeline-api startup. No mobile web page is planned: scan-pfm/page.tsx is desktop-only, used to test the pipeline; real mobile product scanning goes through the Flutter app instead, so m-scan-pfm/page.tsx and its nginx.conf route are intentionally left unbuilt/dead (see plan task 2.2, cancelled 2026-07-08). Not yet done: an actual browser pass uploading a photo through /scan-pfm end-to-end (verified via container logs/model-loading so far, not a UI test).
Confidentiality
Same concerns as ai-ocr-pfm-2026 apply here, plus more surface area:
pfm-web-app/src/db/init.tsanddb/migrations/005_create_sku_master.sqlcontain the client's real product catalog and real vendor/customer identities, committed directly in source.sources/holds live business data:Rekap SKU Aktif CPI Cikande per April 2026 v2.xlsx,Tabel Toko Aktif Juni 2026.xlsx,toko_aktif.json,manual_labels.json,ai_results.json— real SKU/store master data and hand-labeled ground truth from real scanned documents, not fixtures.uploads/contains real scanned delivery-order photos and their OCR JSON output.- The
accountstable stores bcrypt-hashed passwords as of 2026-07-08 (see above) — still don't log or export its contents, and it's not wired into most routes yet (task 1.3), so don't treat it as a secure boundary for anything beyond theapi/v1/*REST layer.
Commands
Web app (pfm-web-app/):
npm run dev # next dev -H 0.0.0.0 (binds all interfaces — for LAN/tunnel access during mobile testing)
npm run build
npm run start
npm run lint
Accuracy regression check (see above) — run after any parser.ts change:
node pfm-web-app/scripts/accuracy-check.mts
pfm-web-app/src/utils/parser.test.ts — same standalone node:assert script as the old project, covering parseDOMetadata/sanitizeParsedMetadata. Run with a TS-capable runner, e.g. npx tsx pfm-web-app/src/utils/parser.test.ts.
Python services (uv-managed, same as ai-ocr-pfm-2026 — see AGENTS.md):
./scripts/install.sh # bootstrap .venv for vLLM server
./scripts/install-pipeline.sh # bootstrap .venv-api
./scripts/serve.sh # vLLM genai server on :8118
./scripts/serve-pipeline.sh # pipeline API on :8090 + classify_ocr_server.py on :8120
Full stack:
docker compose up -d --build
Agents Settings Kit (backend-scoped)
@AGENTS.md
AGENTS.md in this directory now has two parts: Part A is the pre-existing vLLM
service doc referenced above; Part B (appended 2026-07-08) is a backend-scoped
copy of the fhanyuh/agents-settings
e/enhance/n/next workflow, independent of the root-level copy that covers
the Flutter side (see root CLAUDE.md/AGENTS.md). Roles are in SKILLS.md (this
dir). The backlog and shipped-feature log live in plans/next-enhancements.md and
docs/feature-list.md (this dir) — these are backend-only and separate from the
root project's equivalents, which now only track Flutter work.
Claude-specific notes (same as root):
- Spawn the relevant
SKILLS.mdrole via theAgenttool for a fresh-context review/QA/architecture pass instead of continuing in the implementing context. - Use
AskUserQuestionfor the one-at-a-time clarification step (§B2a). - Use
EnterPlanModebefore writing code for anyn/nexttask that touches multiple files or has more than one reasonable implementation approach.