Files
pfm-ocr/backend/CLAUDE.md
T

9.7 KiB

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

What this repo is

app-pfm-ocr-v2/backend is the next-generation rewrite of ai-ocr-pfm-2026 — same underlying OCR infra (PaddleOCR-VL-1.6 on vLLM + a PaddlePaddle layout-parsing pipeline), same client (Charoen Pokphand/Primafood-branded frozen food products), but a reworked Next.js app (pfm-web-app/) and DB schema. If you need background on the shared OCR/vLLM infra (uv conventions, issue-recording workflow, GPU tuning), see AGENTS.md — it's carried over near-unchanged from the previous project.

The active plan for porting the Product/SKU-scanning feature lives in plans/next-enhancements.md §2 — read it before touching anything related to scan-pfm, produk-pfm, or the product classifier, since it records exactly what's done vs. still missing and the decisions already made about how to build it. (This used to be a separate next-implementation.md; that file was deleted 2026-07-08 once its content was folded into the plan for traceability with the rest of the e/n backlog.)

How this project differs from ai-ocr-pfm-2026

  • DO-PFM UI is consolidated into a single page. Unlike the old project's per-route pages (do-pfm/page.tsx, m-do-pfm/page.tsx), v2's entire upload/history/item-review flow lives in one pfm-web-app/src/app/page.tsx (client component, local state, no separate routes). nginx.conf still has /do-pfm//m-do-pfm location blocks left over from the old routing — these are currently dead (no matching Next.js route, would 404).
  • Standalone-purpose pages still get their own route folder, e.g. pfm-web-app/src/app/manual-label/page.tsx — a self-contained ground-truth annotation tool (own header, own theme, no shared chrome with the root page) backed by api/manual-label/route.ts and sources/manual_labels.json. This is the pattern to follow for any new single-purpose page (see plans/next-enhancements.md §2 for the Product-scan pages, which follow it).
  • Real JWT auth, enforced on the production surface: src/utils/auth.ts signs/verifies tokens (signAccountToken/verifyAccountToken/getAccountFromAuthHeader) against an accounts table, each account bound to exactly one kode_toko (store) — the intent being that an account's own store is used on upload instead of relying on OCR-based store-text matching. Passwords are bcrypt-hashed (accounts.password, via bcryptjs — chosen over native bcrypt since the pfm-web-app Docker stage is node:20-slim with no build toolchain for native addons; pfm-web-app/src/db/init.ts hashes the seed and idempotently migrates any pre-existing plaintext rows on every startup). As of 2026-07-08, api/v1/documents/* (list, PUT-by-id, upload) reject requests with a missing/invalid token (401) — this is the real production surface, and the Flutter client already does a real login and attaches Authorization: Bearer <token> to every request (lib/features/auth/auth_provider.dart + lib/core/network/api_client.dart). The classic routes (/api/upload, /api/scan-pfm, /api/parse, /api/history, etc.) and the root/scan-pfm/manual-label pages deliberately do not check auth at all and never will unless that decision changes — they're dev-only web UI with no login screen, not part of the production surface (see plans/next-enhancements.md task 1.3, cancelled, and 1.4, shipped instead).
  • Richer SKU master data: pfm-web-app/import_sku.js imports from a TSV with extended packaging columns (standar_jumlah, berat_kemasan, isi_outer_kg, isi_outer_pac, jenis_outer) added via ALTER TABLE sku_master ADD COLUMN IF NOT EXISTS, superseding the old project's bare no_sku/nama_item seed list.
  • Accuracy regression harness (new, doesn't exist in the old project): pfm-web-app/scripts/accuracy-check.mts hits the live /api/parse endpoint for every image in sources/test-images/, diffs against hand-labeled ground truth in sources/manual_labels.json at three post-processing stages (layer1RawRegex → layer2Sanitized → layer3Final — trace these stage names into utils/parser.ts to see where each is produced), and appends run-over-run results to sources/accuracy_history.jsonl. Run this after touching parser.ts to check for regressions:
    node pfm-web-app/scripts/accuracy-check.mts                    # reuse cached OCR (fast)
    node pfm-web-app/scripts/accuracy-check.mts --refresh-ocr       # force fresh pipeline run
    node pfm-web-app/scripts/accuracy-check.mts --detail <filename> # full per-stage breakdown for one image
    
    compare_accuracy.py / compare_sources_accuracy.py / generate_excel.py at the repo root build human-readable Excel/HTML comparison reports from the same data (sources/comparison_report.xlsx, sources/comparison_side_by_side.html) — these are analysis tooling, not part of the running app.
  • api/vllm-proxy/[[...path]]/route.ts: a passthrough proxy to the vLLM server (paddleocr-vllm-server:8118) that logs every call via logVllmCallToAll (utils/active-log.ts) — used for debugging/observability, not part of the OCR pipeline itself.
  • docker-compose.override.yml exposes db (5432) and pipeline-api (8090) directly to the host for local dev — not present in the old project's compose setup.

Product/SKU scanning flow — status

How it works end-to-end (architecture, endpoints, classification/OCR internals, retraining): docs/scan-product.md. See plans/next-enhancements.md §2 (task 2.1) for full detail — kept there instead of a separate doc so status stays traceable against the rest of the e/n backlog. Feature-complete as of 2026-07-08: the backend (config/classify_ocr_server.py with DINOv2 similarity search + YOLO classifier fallback, api/scan-pfm/route.ts, api/produk-pfm/route.ts, DB schema), the reference photo dataset (pfm-web-app/public/produk-pfm/foto-kemasan-v2/, 16 SKU subfolders), the desktop frontend page (scan-pfm/page.tsx, full feature parity), and the trained model artifacts (models/dinov2_index.pkl — 118/118 photos indexed; models/produk-pfm-classifier-26n-100e-2026-07-08.pt — 83.3% top-1 val accuracy on the current thin dataset) all now exist and load cleanly on pipeline-api startup. No mobile web page is planned: scan-pfm/page.tsx is desktop-only, used to test the pipeline; real mobile product scanning goes through the Flutter app instead, so m-scan-pfm/page.tsx and its nginx.conf route are intentionally left unbuilt/dead (see plan task 2.2, cancelled 2026-07-08). Not yet done: an actual browser pass uploading a photo through /scan-pfm end-to-end (verified via container logs/model-loading so far, not a UI test).

Confidentiality

Same concerns as ai-ocr-pfm-2026 apply here, plus more surface area:

  • pfm-web-app/src/db/init.ts and db/migrations/005_create_sku_master.sql contain the client's real product catalog and real vendor/customer identities, committed directly in source.
  • sources/ holds live business data: Rekap SKU Aktif CPI Cikande per April 2026 v2.xlsx, Tabel Toko Aktif Juni 2026.xlsx, toko_aktif.json, manual_labels.json, ai_results.json — real SKU/store master data and hand-labeled ground truth from real scanned documents, not fixtures.
  • uploads/ contains real scanned delivery-order photos and their OCR JSON output.
  • The accounts table stores bcrypt-hashed passwords as of 2026-07-08 (see above) — still don't log or export its contents, and it's not wired into most routes yet (task 1.3), so don't treat it as a secure boundary for anything beyond the api/v1/* REST layer.

Commands

Web app (pfm-web-app/):

npm run dev      # next dev -H 0.0.0.0 (binds all interfaces — for LAN/tunnel access during mobile testing)
npm run build
npm run start
npm run lint

Accuracy regression check (see above) — run after any parser.ts change:

node pfm-web-app/scripts/accuracy-check.mts

pfm-web-app/src/utils/parser.test.ts — same standalone node:assert script as the old project, covering parseDOMetadata/sanitizeParsedMetadata. Run with a TS-capable runner, e.g. npx tsx pfm-web-app/src/utils/parser.test.ts.

Python services (uv-managed, same as ai-ocr-pfm-2026 — see AGENTS.md):

./scripts/install.sh            # bootstrap .venv for vLLM server
./scripts/install-pipeline.sh   # bootstrap .venv-api
./scripts/serve.sh              # vLLM genai server on :8118
./scripts/serve-pipeline.sh     # pipeline API on :8090 + classify_ocr_server.py on :8120

Full stack:

docker compose up -d --build

Agents Settings Kit (backend-scoped)

@AGENTS.md

AGENTS.md in this directory now has two parts: Part A is the pre-existing vLLM service doc referenced above; Part B (appended 2026-07-08) is a backend-scoped copy of the fhanyuh/agents-settings e/enhance/n/next workflow, independent of the root-level copy that covers the Flutter side (see root CLAUDE.md/AGENTS.md). Roles are in SKILLS.md (this dir). The backlog and shipped-feature log live in plans/next-enhancements.md and docs/feature-list.md (this dir) — these are backend-only and separate from the root project's equivalents, which now only track Flutter work.

Claude-specific notes (same as root):

  • Spawn the relevant SKILLS.md role via the Agent tool for a fresh-context review/QA/architecture pass instead of continuing in the implementing context.
  • Use AskUserQuestion for the one-at-a-time clarification step (§B2a).
  • Use EnterPlanMode before writing code for any n/next task that touches multiple files or has more than one reasonable implementation approach.