Files
pfm-ocr/backend/AGENTS.md
T
Rafhan Mazaya FathurrahmanandClaude Sonnet 5 e60ab63154 Adopt agents-settings kit, ship Product/SKU scan models, harden auth, verify OCR accuracy
Backend (app-pfm-ocr-v2/backend):
- Product/SKU scan feature complete: trained DINOv2 index (118 reference
  photos, 16 SKU classes) and YOLO classifier (83.3% top-1 val accuracy),
  fixed scripts/install-pipeline.sh (was missing ultralytics/torch), fully
  browser-verified end-to-end on /scan-pfm. Mobile m-scan-pfm page cancelled
  (Flutter app handles mobile; web UI is desktop-only for pipeline testing).
- Fixed a real data-loss bug: Save Ground Truth (scan-pfm and the DO-flow's
  manual-label) was silently writing into the pfm-web-app container's
  ephemeral filesystem instead of the host, because /sources wasn't
  bind-mounted in docker-compose.yml. Added the mount, recovered an
  orphaned entry.
- accounts.password is now bcrypt-hashed (bcryptjs, idempotent migration
  in db/init.ts) instead of plaintext; login route compares hashes.
- /api/v1/documents/* (list, PUT, upload) now enforces real 401 auth,
  matching what the Flutter client already sends. The "classic" routes
  deliberately stay open — they're dev-only web UI with no login flow and
  won't exist in production.
- OCR accuracy investigated end-to-end: real baseline is 95.10% overall
  (target met; accuracy_report.md was stale at 75.04%, now flagged). Fixed
  one genuine parser.ts bug (SO/DO field duplication in the global fallback
  regex); remaining gaps are OCR/layout-model limitations, not parser bugs.
- Adopted a standalone copy of the fhanyuh/agents-settings e/n workflow
  scoped to backend/ (AGENTS.md Part A/B split, SKILLS.md, plans/, docs/),
  independent of the root copy which now covers Flutter only.
- next-implementation.md deleted; content folded into
  backend/plans/next-enhancements.md for traceability.

Root:
- Adopted fhanyuh/agents-settings kit (AGENTS.md, SKILLS.md, plans/,
  docs/feature-list.md), scoped to the Flutter app only.
- Pending documents queue now persists to Hive (lib/core/storage) instead
  of memory-only, surviving an app kill mid-upload.

Removed backend_backup/ (stale Express/Prisma prototype, superseded by
pfm-web-app) and the completed plans/next-enhancement-plan.md checklist.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-08 11:56:32 +07:00

11 KiB

AGENTS: PaddleOCR-VL-1.6 vLLM Service + Agents Settings Kit

This is the authoritative rules file for any AI coding agent (Claude Code, Cursor, GitHub Copilot, Aider, etc.) working inside backend/. Two unrelated concerns live here side by side: Part A is this repo's original vLLM/PaddleOCR service doc. Part B (appended 2026-07-08) is a backend-scoped copy of the fhanyuh/agents-settings e/enhance and n/next workflow — see root ../AGENTS.md for the same kit covering the Flutter side of this repo. The two copies are independent: this one's plans/next-enhancements.md and docs/feature-list.md only track backend work.


Part A — vLLM Service (PaddleOCR-VL-1.6)

This repository serves PaddleOCR-VL-1.6 as a dedicated VLM inference backend using vLLM. All Python workflows use uv (never bare pip or system Python). Full detail (client usage examples, tuning, troubleshooting, issue-file template) moved to docs/vllm-service.md 2026-07-08 to keep this file under the Part B kit's 256-line threshold (§3) — this section keeps only the essentials.

Architecture

Client (PaddleOCR pipeline)  -->  HTTP /v1  -->  paddleocr genai_server (vLLM backend)

This service exposes only the VLM stage. Clients connect with vl_rec_backend="vllm-server" and vl_rec_server_url="http://<host>:8118/v1".

Quick start

./scripts/install.sh   # 1) Create Python 3.12 venv and install dependencies
./scripts/serve.sh     # 2) Start the vLLM-backed genai server

Default endpoint: http://0.0.0.0:8118/v1. Never use python -m pip, pip install, or python -m venv directly in this repo — always uv sync / uv run / uv add.

Issue recording (always follow)

Every problem encountered during install, serve, debug, or client integration must be written to issues/{NN}-{slug}.md before moving on — even if resolved in the same session. Naming/template details: docs/vllm-service.md.

Environment variables

Copy .env.example to .env and adjust as needed:

Variable Default Description
GENAI_HOST 0.0.0.0 Bind address
GENAI_PORT 8118 Service port
GENAI_MODEL PaddleOCR-VL-1.6-0.9B Model name for genai_server
GENAI_BACKEND vllm Inference backend
VLLM_CONFIG config/vllm_config.yaml vLLM backend YAML config
CUDA_VISIBLE_DEVICES 1 (see .env.example) GPU index(es) to use

On dual-GPU hosts, pick the GPU with more free VRAM. If startup fails with a memory error, lower gpu-memory-utilization in config/vllm_config.yaml — see docs/vllm-service.md.

File map

Path Purpose
issues/ Recorded problems and fixes ({NN}-{slug}.md)
pyproject.toml uv project metadata and base dependencies
scripts/install.sh Bootstrap venv + vLLM server deps
scripts/serve.sh Start paddleocr genai_server
config/vllm_config.yaml vLLM backend tuning
.env.example Environment variable template
docs/vllm-service.md Full vLLM reference (client usage, tuning, troubleshooting)

Coding Guidelines (always follow)

We use the karpathy-guidelines skill to reduce common LLM coding mistakes:

  1. Think Before Coding: Explicitly state assumptions and surface tradeoffs instead of making silent choices.
  2. Simplicity First: Write the minimum amount of code to solve the problem with zero speculative configurations.
  3. Surgical Changes: Edit only what is required and match the existing coding style exactly.
  4. Goal-Driven Execution: Define verifiable success criteria and run automated tests/screenshots to confirm correctness.

Path Guidelines (always follow)

Never use full paths containing the user's logged-in name (e.g., /home/{uid}/path). Always use relative paths instead (e.g., . or ./path relative to the workspace root).

App Testing Guidelines (always follow)

When the user intentionally asks to test the app:

  • Use browser tools to test the app.
  • Take a screenshot for each sample image, each step, and each variant/option (if any), until the OCR result appears.
  • Save the screenshots in the /screenshots/ folder.
  • Follow the file naming convention: {2-digit-number}-{step#}-{variant_or_options_if_any}-{slug}.jpg (e.g., 01-step1-default-upload.jpg).

Part B — Agents Settings Kit (backend-scoped e/n workflow)

Backend-scoped copy of the fhanyuh/agents-settings kit, adopted 2026-07-08. Covers only backend/ modules (Next.js API Gateway, OCR Pipeline & Accuracy, Postgres Data Layer, DevOps/Docker) — Flutter modules are tracked by the separate copy at root ../AGENTS.md. ../CLAUDE.md (root) and CLAUDE.md (this dir) each import their own copy.

B0. Adopting Into an Existing Project

Already done for this repo (this split is that adoption, mirroring root's own §0 audit). Re-run "i"/"init" here to force a re-audit of backend/ specifically (e.g. after a large refactor).

B1. Trigger "e" or "enhance"

  • Read plans/next-enhancements.md (this dir) to understand current backend structure, history, and active tasks.
  • Overwrite or update the active tasks list inside it.
  • The plan must cover each backend section/module.
  • Define exactly 3 new enhancements per section, each with a unique number (e.g. 1.1), a clear functional description, and status [TODO].
  • Present the plan to the user in your final summary.

B2. Trigger "n", "next", or "n{x}"

  • Read plans/next-enhancements.md to check task status.
  • If all tasks are [DONE] (or none [TODO]), run "e"/"enhance" first.
  • Otherwise select the most impactful [TODO] task(s) by strategic value/impact — not just first-in-order. If {x} given, take the top {x} sequentially.

B2a. Clarify before building ("Grill Me" step)

Same rule as root AGENTS.md §2a: if scope/acceptance criteria are genuinely ambiguous, ask one question at a time (AskUserQuestion in Claude Code) until unambiguous, and record the resolved criteria as a 1-3 line note next to the task entry before writing code. Skip when the task is already unambiguous.

  • Implement the task(s) fully, applying the relevant role(s) from SKILLS.md (this dir).
  • On completion: flip status to [DONE] in plans/next-enhancements.md, document the feature in docs/feature-list.md (this dir) under the right section.
  • Verify build integrity: QA pass (golden path + edge cases + regression check on adjacent features — see backend/CLAUDE.md's accuracy regression harness for OCR/parser changes specifically) and Hardware/Compatibility pass (cross-platform, GPU/VRAM footprint under Local/on-prem deployment — see Part A above).
  • State which task(s) were completed and the exact route/endpoint/menu path to see the new feature.

B3. File Size & Refactoring Rules

Same 256-line threshold as root AGENTS.md §3, backend-wide. Applies to this file, SKILLS.md, and CLAUDE.md too — which is why Part A above was trimmed and linked out to docs/vllm-service.md rather than left inline.

B4. Roles

See SKILLS.md (this dir) — same 5 roles as root (Architect, Backend, Frontend, QA, Hardware/Compatibility), applied to backend surfaces only (API routes, OCR pipeline, DB layer, Docker/deploy).

B5. Mockup Data & Demo/Live Mode

Same as root AGENTS.md §5: mock data under /data/mockup/, a mock API layer mirroring the real backend contract, and a Demo/Live switcher. Not yet built for backend — see Adaptation Notes.

B6. Cloud vs Local (On-Premise)

Same as root AGENTS.md §6, applied to backend service endpoints (Next.js gateway, pipeline API, vLLM server, Postgres) rather than the Flutter client's API base URL.

B7. Ad-hoc Feature Requests

Direct feature requests not using "e"/"n": implement and document in docs/feature-list.md (this dir).

Adaptation Notes (backend, split from root 2026-07-08)

  • Origin: sections 5-8 of root plans/next-enhancements.md (Backend — Next.js API Gateway, Backend — OCR Pipeline & Accuracy, Backend — Postgres Data Layer, DevOps — Docker & Dev Tunnel) copied here as sections 1-4, statuses re-verified against the live code before the copy (not copied blind) — see task 7.1's withTransaction claim, task 5.1/5.2's dedup + timeout claims, and task 6.1's empty models/ claim, all confirmed still accurate as of 2026-07-08. The root copy is frozen/archival (see root AGENTS.md's "Scope: excludes backend/") rather than deleted, so this file — not the root one — is the single active source of truth going forward.
  • Real commands: npm run dev/build/lint in pfm-web-app/; accuracy regression harness node pfm-web-app/scripts/accuracy-check.mts; Python services via ./scripts/install.sh + ./scripts/serve.sh (this vLLM repo) and ./scripts/install-pipeline.sh + ./scripts/serve-pipeline.sh (pipeline API + classifier). Full stack: docker compose up -d --build from the repo root, not from inside backend/ (see root CLAUDE.md — two docker-compose.yml files exist and running from here risks container-name conflicts).
  • Pre-existing files over the 256-line threshold (§B3 debt, not a blocker — split only if/when touched): pfm-web-app/src/app/scan-pfm/page.tsx (1169), pfm-web-app/src/utils/parser.ts (908), pfm-web-app/src/app/page.tsx (737), config/classify_ocr_server.py (691), pfm-web-app/src/app/manual-label/page.tsx (612), pfm-web-app/src/app/api/parse/route.ts (604), pfm-web-app/src/db/init.ts (477), compare_sources_accuracy.py (451), pfm-web-app/src/utils/docker.ts (362), pfm-web-app/public/produk-pfm/train_classifier.py (351), compare_accuracy.py (308), pfm-web-app/src/app/api/arena/route.ts (265). This file itself (AGENTS.md) was at 237 lines pre-kit and would have exceeded 256 once Part B was appended — hence the split into docs/vllm-service.md.
  • No Demo/Live or Cloud/Local switch exists yet (§B5, §B6) for the backend either. docker-compose.override.yml exposing db/pipeline-api directly to the host is a local-dev convenience, not a Cloud/Local deployment switch.
  • Naming collision resolved by this split: AGENTS.md already existed in this directory (vLLM service doc, committed 2026-06-30, unrelated to this kit) before Part B was appended — unlike root, where AGENTS.md didn't previously exist. Don't assume backend's AGENTS.md is kit-only when reading it from another tool; Part A is unrelated, pre-existing content kept for a reason.
  • Pre-existing, unrelated governance files left as-is: .agents/AGENTS.md at the repo root (different path, OCR post-processing rules) and root plans/next-enhancement-plan.md (singular, [DONE] QA checklist) — neither is part of this kit; see root AGENTS.md's own Adaptation Notes.