No content changes: git diff --ignore-all-space over these files is empty. The churn came from editing on Windows against a repo checked out with LF.
12 KiB
AGENTS: PaddleOCR-VL-1.6 vLLM Service + Agents Settings Kit
This is the authoritative rules file for any AI coding agent (Claude Code, Cursor,
GitHub Copilot, Aider, etc.) working inside backend/. Two unrelated concerns live
here side by side: Part A is this repo's original vLLM/PaddleOCR service doc.
Part B (appended 2026-07-08) is a backend-scoped copy of the
fhanyuh/agents-settings e/enhance
and n/next workflow — see root ../AGENTS.md for the same kit covering the
Flutter side of this repo. The two copies are independent: this one's
plans/next-enhancements.md and docs/feature-list.md only track backend work.
Part A — vLLM Service (PaddleOCR-VL-1.6)
This repository serves PaddleOCR-VL-1.6 as a dedicated VLM inference backend using vLLM. All Python workflows use uv (never bare pip or system Python). Full detail (client usage examples, tuning, troubleshooting, issue-file template) moved to docs/vllm-service.md 2026-07-08 to keep this file under the Part B kit's 256-line threshold (§3) — this section keeps only the essentials.
Architecture
Client (PaddleOCR pipeline) --> HTTP /v1 --> paddleocr genai_server (vLLM backend)
This service exposes only the VLM stage. Clients connect with vl_rec_backend="vllm-server" and vl_rec_server_url="http://<host>:8118/v1".
Quick start
./scripts/install.sh # 1) Create Python 3.12 venv and install dependencies
./scripts/serve.sh # 2) Start the vLLM-backed genai server
Default endpoint: http://0.0.0.0:8118/v1. Never use python -m pip, pip install, or python -m venv directly in this repo — always uv sync / uv run / uv add.
Issue recording (always follow)
Every problem encountered during install, serve, debug, or client integration must be written to issues/{NN}-{slug}.md before moving on — even if resolved in the same session. Naming/template details: docs/vllm-service.md.
Environment variables
Copy .env.example to .env and adjust as needed:
| Variable | Default | Description |
|---|---|---|
GENAI_HOST |
0.0.0.0 |
Bind address |
GENAI_PORT |
8118 |
Service port |
GENAI_MODEL |
PaddleOCR-VL-1.6-0.9B |
Model name for genai_server |
GENAI_BACKEND |
vllm |
Inference backend |
VLLM_CONFIG |
config/vllm_config.yaml |
vLLM backend YAML config |
CUDA_VISIBLE_DEVICES |
1 (see .env.example) |
GPU index(es) to use |
On dual-GPU hosts, pick the GPU with more free VRAM. If startup fails with a memory error, lower gpu-memory-utilization in config/vllm_config.yaml — see docs/vllm-service.md.
File map
| Path | Purpose |
|---|---|
issues/ |
Recorded problems and fixes ({NN}-{slug}.md) |
pyproject.toml |
uv project metadata and base dependencies |
scripts/install.sh |
Bootstrap venv + vLLM server deps |
scripts/serve.sh |
Start paddleocr genai_server |
config/vllm_config.yaml |
vLLM backend tuning |
.env.example |
Environment variable template |
docs/vllm-service.md |
Full vLLM reference (client usage, tuning, troubleshooting) |
Coding Guidelines (always follow)
We use the karpathy-guidelines skill to reduce common LLM coding mistakes:
- Think Before Coding: Explicitly state assumptions and surface tradeoffs instead of making silent choices.
- Simplicity First: Write the minimum amount of code to solve the problem with zero speculative configurations.
- Surgical Changes: Edit only what is required and match the existing coding style exactly.
- Goal-Driven Execution: Define verifiable success criteria and run automated tests/screenshots to confirm correctness.
- SOLID Principles: Always design, implement, and refactor code adhering to SOLID programming principles (Single Responsibility, Open/Closed, Liskov Substitution, Interface Segregation, Dependency Inversion) to ensure modularity, scalability, and maintainability.
Path Guidelines (always follow)
Never use full paths containing the user's logged-in name (e.g., /home/{uid}/path). Always use relative paths instead (e.g., . or ./path relative to the workspace root).
App Testing Guidelines (always follow)
When the user intentionally asks to test the app:
- Use browser tools to test the app.
- Take a screenshot for each sample image, each step, and each variant/option (if any), until the OCR result appears.
- Save the screenshots in the
/screenshots/folder. - Follow the file naming convention:
{2-digit-number}-{step#}-{variant_or_options_if_any}-{slug}.jpg(e.g.,01-step1-default-upload.jpg).
Part B — Agents Settings Kit (backend-scoped e/n workflow)
Backend-scoped copy of the fhanyuh/agents-settings kit, adopted 2026-07-08. Covers only backend/ modules (Next.js API Gateway, OCR Pipeline & Accuracy, Postgres Data Layer, DevOps/Docker) — Flutter modules are tracked by the separate copy at root ../AGENTS.md. ../CLAUDE.md (root) and CLAUDE.md (this dir) each import their own copy.
B0. Adopting Into an Existing Project
Already done for this repo (this split is that adoption, mirroring root's own §0 audit). Re-run "i"/"init" here to force a re-audit of backend/ specifically (e.g. after a large refactor).
B1. Trigger "e" or "enhance"
- Read
plans/next-enhancements.md(this dir) to understand current backend structure, history, and active tasks. - Overwrite or update the active tasks list inside it.
- The plan must cover each backend section/module.
- Define exactly 3 new enhancements per section, each with a unique number (e.g.
1.1), a clear functional description, and status[TODO]. - Present the plan to the user in your final summary.
B2. Trigger "n", "next", or "n{x}"
- Read
plans/next-enhancements.mdto check task status. - If all tasks are
[DONE](or none[TODO]), run "e"/"enhance" first. - Otherwise select the most impactful
[TODO]task(s) by strategic value/impact — not just first-in-order. If{x}given, take the top{x}sequentially.
B2a. Clarify before building ("Grill Me" step)
Same rule as root AGENTS.md §2a: if scope/acceptance criteria are genuinely ambiguous, ask one question at a time (AskUserQuestion in Claude Code) until unambiguous, and record the resolved criteria as a 1-3 line note next to the task entry before writing code. Skip when the task is already unambiguous.
B2b. TDD Workflow (Test First)
-
Write Tests First: Before implementing the actual feature code for a task, write automated tests defining the expected behavior.
-
Iterate Until Green: Run the tests to confirm they fail, then write the implementation until all tests pass perfectly.
-
Browser Testing: If the enhancement involves web UI or visual components, use browser tools (e.g., Chrome) to test the app visually and functionally if necessary.
-
Implement the task(s) fully, applying the relevant role(s) from
SKILLS.md(this dir). -
On completion:
- Flip status to
[DONE]inplans/next-enhancements.md. - Document the feature in
docs/feature-list.md(this dir) under the right section. - Create an Iteration Log: Perform a code review and audit of the tasks just completed. Document this audit in
docs/iteration-log.md(or append to it) to ensure all functions work perfectly. - Update Documentation: Sync any architecture or workflow changes back to
CLAUDE.mdandSKILLS.mdto keep the agent instructions current.
- Flip status to
-
Verify build integrity: QA pass (golden path + edge cases + regression check on adjacent features — see
backend/CLAUDE.md's accuracy regression harness for OCR/parser changes specifically) and Hardware/Compatibility pass (cross-platform, GPU/VRAM footprint under Local/on-prem deployment — see Part A above). -
State which task(s) were completed and the exact route/endpoint/menu path to see the new feature.
B3. File Size & Refactoring Rules
Same 256-line threshold as root AGENTS.md §3, backend-wide. Applies to this file, SKILLS.md, and CLAUDE.md too — which is why Part A above was trimmed and linked out to docs/vllm-service.md rather than left inline.
B4. Roles
See SKILLS.md (this dir) — same 5 roles as root (Architect, Backend, Frontend, QA, Hardware/Compatibility), applied to backend surfaces only (API routes, OCR pipeline, DB layer, Docker/deploy).
B5. Mockup Data & Demo/Live Mode
Same as root AGENTS.md §5: mock data under /data/mockup/, a mock API layer mirroring the real backend contract, and a Demo/Live switcher. Not yet built for backend — see Adaptation Notes.
B6. Cloud vs Local (On-Premise)
Same as root AGENTS.md §6, applied to backend service endpoints (Next.js gateway, pipeline API, vLLM server, Postgres) rather than the Flutter client's API base URL.
B7. Ad-hoc Feature Requests
Direct feature requests not using "e"/"n": implement and document in docs/feature-list.md (this dir).
Adaptation Notes (backend, split from root 2026-07-08)
- Origin: sections 5-8 of root
plans/next-enhancements.md(Backend — Next.js API Gateway, Backend — OCR Pipeline & Accuracy, Backend — Postgres Data Layer, DevOps — Docker & Dev Tunnel) copied here as sections 1-4, statuses re-verified against the live code before the copy (not copied blind) — see task 7.1'swithTransactionclaim, task 5.1/5.2's dedup + timeout claims, and task 6.1's emptymodels/claim, all confirmed still accurate as of 2026-07-08. The root copy is frozen/archival (see rootAGENTS.md's "Scope: excludesbackend/") rather than deleted, so this file — not the root one — is the single active source of truth going forward. - Real commands:
npm run dev/build/lintinpfm-web-app/; accuracy regression harnessnode pfm-web-app/scripts/accuracy-check.mts; Python services via./scripts/install.sh+./scripts/serve.sh(this vLLM repo) and./scripts/install-pipeline.sh+./scripts/serve-pipeline.sh(pipeline API + classifier). Full stack:docker compose up -d --buildfrom the repo root, not from insidebackend/(see rootCLAUDE.md— twodocker-compose.ymlfiles exist and running from here risks container-name conflicts). - Pre-existing files over the 256-line threshold (§B3 debt, not a blocker — split
only if/when touched):
pfm-web-app/src/app/scan-pfm/page.tsx(1169),pfm-web-app/src/utils/parser.ts(908),pfm-web-app/src/app/page.tsx(737),config/classify_ocr_server.py(691),pfm-web-app/src/app/manual-label/page.tsx(612),pfm-web-app/src/app/api/parse/route.ts(604),pfm-web-app/src/db/init.ts(477),compare_sources_accuracy.py(451),pfm-web-app/src/utils/docker.ts(362),pfm-web-app/public/produk-pfm/train_classifier.py(351),compare_accuracy.py(308),pfm-web-app/src/app/api/arena/route.ts(265). This file itself (AGENTS.md) was at 237 lines pre-kit and would have exceeded 256 once Part B was appended — hence the split intodocs/vllm-service.md. - No Demo/Live or Cloud/Local switch exists yet (§B5, §B6) for the backend
either.
docker-compose.override.ymlexposingdb/pipeline-apidirectly to the host is a local-dev convenience, not a Cloud/Local deployment switch. - Naming collision resolved by this split:
AGENTS.mdalready existed in this directory (vLLM service doc, committed 2026-06-30, unrelated to this kit) before Part B was appended — unlike root, whereAGENTS.mddidn't previously exist. Don't assume backend'sAGENTS.mdis kit-only when reading it from another tool; Part A is unrelated, pre-existing content kept for a reason. - Pre-existing, unrelated governance files left as-is:
.agents/AGENTS.mdat the repo root (different path, OCR post-processing rules) and rootplans/next-enhancement-plan.md(singular,[DONE]QA checklist) — neither is part of this kit; see rootAGENTS.md's own Adaptation Notes.