Backend (app-pfm-ocr-v2/backend): - Product/SKU scan feature complete: trained DINOv2 index (118 reference photos, 16 SKU classes) and YOLO classifier (83.3% top-1 val accuracy), fixed scripts/install-pipeline.sh (was missing ultralytics/torch), fully browser-verified end-to-end on /scan-pfm. Mobile m-scan-pfm page cancelled (Flutter app handles mobile; web UI is desktop-only for pipeline testing). - Fixed a real data-loss bug: Save Ground Truth (scan-pfm and the DO-flow's manual-label) was silently writing into the pfm-web-app container's ephemeral filesystem instead of the host, because /sources wasn't bind-mounted in docker-compose.yml. Added the mount, recovered an orphaned entry. - accounts.password is now bcrypt-hashed (bcryptjs, idempotent migration in db/init.ts) instead of plaintext; login route compares hashes. - /api/v1/documents/* (list, PUT, upload) now enforces real 401 auth, matching what the Flutter client already sends. The "classic" routes deliberately stay open — they're dev-only web UI with no login flow and won't exist in production. - OCR accuracy investigated end-to-end: real baseline is 95.10% overall (target met; accuracy_report.md was stale at 75.04%, now flagged). Fixed one genuine parser.ts bug (SO/DO field duplication in the global fallback regex); remaining gaps are OCR/layout-model limitations, not parser bugs. - Adopted a standalone copy of the fhanyuh/agents-settings e/n workflow scoped to backend/ (AGENTS.md Part A/B split, SKILLS.md, plans/, docs/), independent of the root copy which now covers Flutter only. - next-implementation.md deleted; content folded into backend/plans/next-enhancements.md for traceability. Root: - Adopted fhanyuh/agents-settings kit (AGENTS.md, SKILLS.md, plans/, docs/feature-list.md), scoped to the Flutter app only. - Pending documents queue now persists to Hive (lib/core/storage) instead of memory-only, surviving an app kill mid-upload. Removed backend_backup/ (stale Express/Prisma prototype, superseded by pfm-web-app) and the completed plans/next-enhancement-plan.md checklist. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
11 KiB
AGENTS: PaddleOCR-VL-1.6 vLLM Service + Agents Settings Kit
This is the authoritative rules file for any AI coding agent (Claude Code, Cursor,
GitHub Copilot, Aider, etc.) working inside backend/. Two unrelated concerns live
here side by side: Part A is this repo's original vLLM/PaddleOCR service doc.
Part B (appended 2026-07-08) is a backend-scoped copy of the
fhanyuh/agents-settings e/enhance
and n/next workflow — see root ../AGENTS.md for the same kit covering the
Flutter side of this repo. The two copies are independent: this one's
plans/next-enhancements.md and docs/feature-list.md only track backend work.
Part A — vLLM Service (PaddleOCR-VL-1.6)
This repository serves PaddleOCR-VL-1.6 as a dedicated VLM inference backend using vLLM. All Python workflows use uv (never bare pip or system Python). Full detail (client usage examples, tuning, troubleshooting, issue-file template) moved to docs/vllm-service.md 2026-07-08 to keep this file under the Part B kit's 256-line threshold (§3) — this section keeps only the essentials.
Architecture
Client (PaddleOCR pipeline) --> HTTP /v1 --> paddleocr genai_server (vLLM backend)
This service exposes only the VLM stage. Clients connect with vl_rec_backend="vllm-server" and vl_rec_server_url="http://<host>:8118/v1".
Quick start
./scripts/install.sh # 1) Create Python 3.12 venv and install dependencies
./scripts/serve.sh # 2) Start the vLLM-backed genai server
Default endpoint: http://0.0.0.0:8118/v1. Never use python -m pip, pip install, or python -m venv directly in this repo — always uv sync / uv run / uv add.
Issue recording (always follow)
Every problem encountered during install, serve, debug, or client integration must be written to issues/{NN}-{slug}.md before moving on — even if resolved in the same session. Naming/template details: docs/vllm-service.md.
Environment variables
Copy .env.example to .env and adjust as needed:
| Variable | Default | Description |
|---|---|---|
GENAI_HOST |
0.0.0.0 |
Bind address |
GENAI_PORT |
8118 |
Service port |
GENAI_MODEL |
PaddleOCR-VL-1.6-0.9B |
Model name for genai_server |
GENAI_BACKEND |
vllm |
Inference backend |
VLLM_CONFIG |
config/vllm_config.yaml |
vLLM backend YAML config |
CUDA_VISIBLE_DEVICES |
1 (see .env.example) |
GPU index(es) to use |
On dual-GPU hosts, pick the GPU with more free VRAM. If startup fails with a memory error, lower gpu-memory-utilization in config/vllm_config.yaml — see docs/vllm-service.md.
File map
| Path | Purpose |
|---|---|
issues/ |
Recorded problems and fixes ({NN}-{slug}.md) |
pyproject.toml |
uv project metadata and base dependencies |
scripts/install.sh |
Bootstrap venv + vLLM server deps |
scripts/serve.sh |
Start paddleocr genai_server |
config/vllm_config.yaml |
vLLM backend tuning |
.env.example |
Environment variable template |
docs/vllm-service.md |
Full vLLM reference (client usage, tuning, troubleshooting) |
Coding Guidelines (always follow)
We use the karpathy-guidelines skill to reduce common LLM coding mistakes:
- Think Before Coding: Explicitly state assumptions and surface tradeoffs instead of making silent choices.
- Simplicity First: Write the minimum amount of code to solve the problem with zero speculative configurations.
- Surgical Changes: Edit only what is required and match the existing coding style exactly.
- Goal-Driven Execution: Define verifiable success criteria and run automated tests/screenshots to confirm correctness.
Path Guidelines (always follow)
Never use full paths containing the user's logged-in name (e.g., /home/{uid}/path). Always use relative paths instead (e.g., . or ./path relative to the workspace root).
App Testing Guidelines (always follow)
When the user intentionally asks to test the app:
- Use browser tools to test the app.
- Take a screenshot for each sample image, each step, and each variant/option (if any), until the OCR result appears.
- Save the screenshots in the
/screenshots/folder. - Follow the file naming convention:
{2-digit-number}-{step#}-{variant_or_options_if_any}-{slug}.jpg(e.g.,01-step1-default-upload.jpg).
Part B — Agents Settings Kit (backend-scoped e/n workflow)
Backend-scoped copy of the fhanyuh/agents-settings kit, adopted 2026-07-08. Covers only backend/ modules (Next.js API Gateway, OCR Pipeline & Accuracy, Postgres Data Layer, DevOps/Docker) — Flutter modules are tracked by the separate copy at root ../AGENTS.md. ../CLAUDE.md (root) and CLAUDE.md (this dir) each import their own copy.
B0. Adopting Into an Existing Project
Already done for this repo (this split is that adoption, mirroring root's own §0 audit). Re-run "i"/"init" here to force a re-audit of backend/ specifically (e.g. after a large refactor).
B1. Trigger "e" or "enhance"
- Read
plans/next-enhancements.md(this dir) to understand current backend structure, history, and active tasks. - Overwrite or update the active tasks list inside it.
- The plan must cover each backend section/module.
- Define exactly 3 new enhancements per section, each with a unique number (e.g.
1.1), a clear functional description, and status[TODO]. - Present the plan to the user in your final summary.
B2. Trigger "n", "next", or "n{x}"
- Read
plans/next-enhancements.mdto check task status. - If all tasks are
[DONE](or none[TODO]), run "e"/"enhance" first. - Otherwise select the most impactful
[TODO]task(s) by strategic value/impact — not just first-in-order. If{x}given, take the top{x}sequentially.
B2a. Clarify before building ("Grill Me" step)
Same rule as root AGENTS.md §2a: if scope/acceptance criteria are genuinely ambiguous, ask one question at a time (AskUserQuestion in Claude Code) until unambiguous, and record the resolved criteria as a 1-3 line note next to the task entry before writing code. Skip when the task is already unambiguous.
- Implement the task(s) fully, applying the relevant role(s) from
SKILLS.md(this dir). - On completion: flip status to
[DONE]inplans/next-enhancements.md, document the feature indocs/feature-list.md(this dir) under the right section. - Verify build integrity: QA pass (golden path + edge cases + regression check on adjacent features — see
backend/CLAUDE.md's accuracy regression harness for OCR/parser changes specifically) and Hardware/Compatibility pass (cross-platform, GPU/VRAM footprint under Local/on-prem deployment — see Part A above). - State which task(s) were completed and the exact route/endpoint/menu path to see the new feature.
B3. File Size & Refactoring Rules
Same 256-line threshold as root AGENTS.md §3, backend-wide. Applies to this file, SKILLS.md, and CLAUDE.md too — which is why Part A above was trimmed and linked out to docs/vllm-service.md rather than left inline.
B4. Roles
See SKILLS.md (this dir) — same 5 roles as root (Architect, Backend, Frontend, QA, Hardware/Compatibility), applied to backend surfaces only (API routes, OCR pipeline, DB layer, Docker/deploy).
B5. Mockup Data & Demo/Live Mode
Same as root AGENTS.md §5: mock data under /data/mockup/, a mock API layer mirroring the real backend contract, and a Demo/Live switcher. Not yet built for backend — see Adaptation Notes.
B6. Cloud vs Local (On-Premise)
Same as root AGENTS.md §6, applied to backend service endpoints (Next.js gateway, pipeline API, vLLM server, Postgres) rather than the Flutter client's API base URL.
B7. Ad-hoc Feature Requests
Direct feature requests not using "e"/"n": implement and document in docs/feature-list.md (this dir).
Adaptation Notes (backend, split from root 2026-07-08)
- Origin: sections 5-8 of root
plans/next-enhancements.md(Backend — Next.js API Gateway, Backend — OCR Pipeline & Accuracy, Backend — Postgres Data Layer, DevOps — Docker & Dev Tunnel) copied here as sections 1-4, statuses re-verified against the live code before the copy (not copied blind) — see task 7.1'swithTransactionclaim, task 5.1/5.2's dedup + timeout claims, and task 6.1's emptymodels/claim, all confirmed still accurate as of 2026-07-08. The root copy is frozen/archival (see rootAGENTS.md's "Scope: excludesbackend/") rather than deleted, so this file — not the root one — is the single active source of truth going forward. - Real commands:
npm run dev/build/lintinpfm-web-app/; accuracy regression harnessnode pfm-web-app/scripts/accuracy-check.mts; Python services via./scripts/install.sh+./scripts/serve.sh(this vLLM repo) and./scripts/install-pipeline.sh+./scripts/serve-pipeline.sh(pipeline API + classifier). Full stack:docker compose up -d --buildfrom the repo root, not from insidebackend/(see rootCLAUDE.md— twodocker-compose.ymlfiles exist and running from here risks container-name conflicts). - Pre-existing files over the 256-line threshold (§B3 debt, not a blocker — split
only if/when touched):
pfm-web-app/src/app/scan-pfm/page.tsx(1169),pfm-web-app/src/utils/parser.ts(908),pfm-web-app/src/app/page.tsx(737),config/classify_ocr_server.py(691),pfm-web-app/src/app/manual-label/page.tsx(612),pfm-web-app/src/app/api/parse/route.ts(604),pfm-web-app/src/db/init.ts(477),compare_sources_accuracy.py(451),pfm-web-app/src/utils/docker.ts(362),pfm-web-app/public/produk-pfm/train_classifier.py(351),compare_accuracy.py(308),pfm-web-app/src/app/api/arena/route.ts(265). This file itself (AGENTS.md) was at 237 lines pre-kit and would have exceeded 256 once Part B was appended — hence the split intodocs/vllm-service.md. - No Demo/Live or Cloud/Local switch exists yet (§B5, §B6) for the backend
either.
docker-compose.override.ymlexposingdb/pipeline-apidirectly to the host is a local-dev convenience, not a Cloud/Local deployment switch. - Naming collision resolved by this split:
AGENTS.mdalready existed in this directory (vLLM service doc, committed 2026-06-30, unrelated to this kit) before Part B was appended — unlike root, whereAGENTS.mddidn't previously exist. Don't assume backend'sAGENTS.mdis kit-only when reading it from another tool; Part A is unrelated, pre-existing content kept for a reason. - Pre-existing, unrelated governance files left as-is:
.agents/AGENTS.mdat the repo root (different path, OCR post-processing rules) and rootplans/next-enhancement-plan.md(singular,[DONE]QA checklist) — neither is part of this kit; see rootAGENTS.md's own Adaptation Notes.