Files
pfm-ocr/backend/CLAUDE.md
T
Rafhan Mazaya FathurrahmanandClaude Fable 5 e76ccb60a6 feat(backend): scan-product accuracy 66.2% -> 79.7% + frozen validation benchmark
Accuracy work on the 79-image product-scan validation set (user goal: 90%):
- classify_ocr_server.py: 0/90/180/270-degree expiry-date search (stops at
  first hit, 0-degree fallback); classification decoupled onto the upright
  image (rotated frames regressed DINOv2 -6pts until this); cross-line date
  stitching; tiled full-res OCR pass (defeats the 4000px downscale that
  killed small inkjet dates); VL-pipeline expiry fallback with
  keyword-anchored anti-hallucination guard; VL text lines merged into
  text_lines + VL SKU retry. Visualization endpoints removed entirely
  (Visual/Spotting grids - unused by frontend, 3x per-scan GPU cost).
- product-scan.ts: coverage-normalized OCR-evidence re-ranking of DINOv2
  top-K (tuned offline: +8/-0 on top-1 misses), re-ranked class mapped to
  sku_master by SKU prefix; classifier timeout 90s->240s for fallback paths.
- Frozen benchmark: product-test-images-fixed/ (79 renamed images) +
  freeze/seed/build-undetected/capture/experiment scripts; labels trimmed to
  the 79 validation entries (training rows kept in .bak-with-training);
  5 TRAINED-ON SKUs replaced with fresh held-out photos.
- manual-label-scan page: shows last batch-test AI prediction under every
  field by default (new /api/product-scan-results); serves the fixed folder;
  fixed total hydration failure via allowedDevOrigins 127.0.0.1.
- Measured (all-79, zero failures): sku/name 87.3%, expiry 64.6%, overall
  79.7%. Tiles/VL-evidence/VL-SKU deployed but not yet batch-measured.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gr6HH7JrdsXX8AARejQboM
2026-07-14 19:55:17 +07:00

9.9 KiB

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

What this repo is

app-pfm-ocr-v2/backend is the next-generation rewrite of ai-ocr-pfm-2026 — same underlying OCR infra (PaddleOCR-VL-1.6 on vLLM + a PaddlePaddle layout-parsing pipeline), same client (Charoen Pokphand/Primafood-branded frozen food products), but a reworked Next.js app (pfm-web-app/) and DB schema. If you need background on the shared OCR/vLLM infra (uv conventions, issue-recording workflow, GPU tuning), see AGENTS.md — it's carried over near-unchanged from the previous project.

The active plan for porting the Product/SKU-scanning feature lives in plans/next-enhancements.md §2 — read it before touching anything related to scan-pfm, produk-pfm, or the product classifier, since it records exactly what's done vs. still missing and the decisions already made about how to build it. (This used to be a separate next-implementation.md; that file was deleted 2026-07-08 once its content was folded into the plan for traceability with the rest of the e/n backlog.)

How this project differs from ai-ocr-pfm-2026

  • DO-PFM UI is consolidated into a single page. Unlike the old project's per-route pages (do-pfm/page.tsx, m-do-pfm/page.tsx), v2's entire upload/history/item-review flow lives in one pfm-web-app/src/app/page.tsx (client component, local state, no separate routes). nginx.conf still has /do-pfm//m-do-pfm location blocks left over from the old routing — these are currently dead (no matching Next.js route, would 404).
  • Standalone-purpose pages still get their own route folder, e.g. pfm-web-app/src/app/manual-label/page.tsx — a self-contained ground-truth annotation tool (own header, own theme, no shared chrome with the root page) backed by api/manual-label/route.ts and sources/manual_labels.json. This is the pattern to follow for any new single-purpose page (see plans/next-enhancements.md §2 for the Product-scan pages, which follow it).
  • Real JWT auth, enforced on the production surface: src/utils/auth.ts signs/verifies tokens (signAccountToken/verifyAccountToken/getAccountFromAuthHeader) against an accounts table, each account bound to exactly one kode_toko (store) — the intent being that an account's own store is used on upload instead of relying on OCR-based store-text matching. Passwords are bcrypt-hashed (accounts.password, via bcryptjs — chosen over native bcrypt since the pfm-web-app Docker stage is node:20-slim with no build toolchain for native addons; pfm-web-app/src/db/init.ts hashes the seed and idempotently migrates any pre-existing plaintext rows on every startup). As of 2026-07-08, api/v1/documents/* (list, PUT-by-id, upload) reject requests with a missing/invalid token (401) — this is the real production surface, and the Flutter client already does a real login and attaches Authorization: Bearer <token> to every request (lib/features/auth/auth_provider.dart + lib/core/network/api_client.dart). The classic routes (/api/upload, /api/scan-pfm, /api/parse, /api/history, etc.) and the root/scan-pfm/manual-label pages deliberately do not check auth at all and never will unless that decision changes — they're dev-only web UI with no login screen, not part of the production surface (see plans/next-enhancements.md task 1.3, cancelled, and 1.4, shipped instead).
  • Richer SKU master data: pfm-web-app/import_sku.js imports from a TSV with extended packaging columns (standar_jumlah, berat_kemasan, isi_outer_kg, isi_outer_pac, jenis_outer) added via ALTER TABLE sku_master ADD COLUMN IF NOT EXISTS, superseding the old project's bare no_sku/nama_item seed list.
  • Accuracy regression harness (new, doesn't exist in the old project): pfm-web-app/scripts/accuracy-check.mts hits the live /api/parse endpoint for every image in sources/test-images/, diffs against hand-labeled ground truth in sources/manual_labels.json at three post-processing stages (layer1RawRegex → layer2Sanitized → layer3Final — trace these stage names into utils/parser.ts to see where each is produced), and appends run-over-run results to sources/accuracy_history.jsonl. Run this after touching parser.ts to check for regressions:
    node pfm-web-app/scripts/accuracy-check.mts                    # reuse cached OCR (fast)
    node pfm-web-app/scripts/accuracy-check.mts --refresh-ocr       # force fresh pipeline run
    node pfm-web-app/scripts/accuracy-check.mts --detail <filename> # full per-stage breakdown for one image
    
    compare_accuracy.py / compare_sources_accuracy.py / generate_excel.py at the repo root build human-readable Excel/HTML comparison reports from the same data (sources/comparison_report.xlsx, sources/comparison_side_by_side.html) — these are analysis tooling, not part of the running app.
  • api/vllm-proxy/[[...path]]/route.ts: a passthrough proxy to the vLLM server (paddleocr-vllm-server:8118) that logs every call via logVllmCallToAll (utils/active-log.ts) — used for debugging/observability, not part of the OCR pipeline itself.
  • docker-compose.override.yml exposes db (5432) and pipeline-api (8090) directly to the host for local dev — not present in the old project's compose setup.

Product/SKU scanning flow — status

How it works end-to-end (architecture, endpoints, classification/OCR internals, retraining): docs/scan-product.md. See plans/next-enhancements.md §2 (task 2.1) for full detail — kept there instead of a separate doc so status stays traceable against the rest of the e/n backlog. Feature-complete as of 2026-07-08: the backend (config/classify_ocr_server.py with DINOv2 similarity search + YOLO classifier fallback, api/scan-pfm/route.ts, api/produk-pfm/route.ts, DB schema), the reference photo dataset (pfm-web-app/public/produk-pfm/foto-kemasan-v2/, 81 SKU subfolders as of 2026-07-14, up from the original 16 — target ~230), the desktop frontend page (scan-pfm/page.tsx, full feature parity), and the trained model artifacts (models/dinov2_index.pkl — 2,493/2,493 photos indexed as of 2026-07-14; models/produk-pfm-classifier-26n-100e-2026-07-14.pt — 85.8% top-1 / 94.4% top-5 val accuracy across all 81 classes, retrained 2026-07-14 in 54m21s on an RTX 2060) all now exist and load cleanly on pipeline-api startup. No mobile web page is planned: scan-pfm/page.tsx is desktop-only, used to test the pipeline; real mobile product scanning goes through the Flutter app instead, so m-scan-pfm/page.tsx and its nginx.conf route are intentionally left unbuilt/dead (see plan task 2.2, cancelled 2026-07-08). Not yet done: an actual browser pass uploading a photo through /scan-pfm end-to-end (verified via container logs/model-loading so far, not a UI test).

Confidentiality

Same concerns as ai-ocr-pfm-2026 apply here, plus more surface area:

  • pfm-web-app/src/db/init.ts and db/migrations/005_create_sku_master.sql contain the client's real product catalog and real vendor/customer identities, committed directly in source.
  • sources/ holds live business data: Rekap SKU Aktif CPI Cikande per April 2026 v2.xlsx, Tabel Toko Aktif Juni 2026.xlsx, toko_aktif.json, manual_labels.json, ai_results.json — real SKU/store master data and hand-labeled ground truth from real scanned documents, not fixtures.
  • uploads/ contains real scanned delivery-order photos and their OCR JSON output.
  • The accounts table stores bcrypt-hashed passwords as of 2026-07-08 (see above) — still don't log or export its contents, and it's not wired into most routes yet (task 1.3), so don't treat it as a secure boundary for anything beyond the api/v1/* REST layer.

Commands

Web app (pfm-web-app/):

npm run dev      # next dev -H 0.0.0.0 (binds all interfaces — for LAN/tunnel access during mobile testing)
npm run build
npm run start
npm run lint

Accuracy regression check (see above) — run after any parser.ts change:

node pfm-web-app/scripts/accuracy-check.mts

pfm-web-app/src/utils/parser.test.ts — same standalone node:assert script as the old project, covering parseDOMetadata/sanitizeParsedMetadata. Run with a TS-capable runner, e.g. npx tsx pfm-web-app/src/utils/parser.test.ts.

Python services (uv-managed, same as ai-ocr-pfm-2026 — see AGENTS.md):

./scripts/install.sh            # bootstrap .venv for vLLM server
./scripts/install-pipeline.sh   # bootstrap .venv-api
./scripts/serve.sh              # vLLM genai server on :8118
./scripts/serve-pipeline.sh     # pipeline API on :8090 + classify_ocr_server.py on :8120

Full stack:

docker compose up -d --build

Agents Settings Kit (backend-scoped)

@AGENTS.md

AGENTS.md in this directory now has two parts: Part A is the pre-existing vLLM service doc referenced above; Part B (appended 2026-07-08) is a backend-scoped copy of the fhanyuh/agents-settings e/enhance/n/next workflow, independent of the root-level copy that covers the Flutter side (see root CLAUDE.md/AGENTS.md). Roles are in SKILLS.md (this dir). The backlog and shipped-feature log live in plans/next-enhancements.md and docs/feature-list.md (this dir) — these are backend-only and separate from the root project's equivalents, which now only track Flutter work.

Claude-specific notes (same as root):

  • Spawn the relevant SKILLS.md role via the Agent tool for a fresh-context review/QA/architecture pass instead of continuing in the implementing context.
  • Use AskUserQuestion for the one-at-a-time clarification step (§B2a).
  • Use EnterPlanMode before writing code for any n/next task that touches multiple files or has more than one reasonable implementation approach.