# AGENTS: PaddleOCR-VL-1.6 vLLM Service + Agents Settings Kit This is the authoritative rules file for any AI coding agent (Claude Code, Cursor, GitHub Copilot, Aider, etc.) working inside `backend/`. Two unrelated concerns live here side by side: **Part A** is this repo's original vLLM/PaddleOCR service doc. **Part B** (appended 2026-07-08) is a **backend-scoped copy** of the [fhanyuh/agents-settings](https://github.com/fhanyuh/agents-settings) `e`/`enhance` and `n`/`next` workflow — see root `../AGENTS.md` for the same kit covering the Flutter side of this repo. The two copies are independent: this one's `plans/next-enhancements.md` and `docs/feature-list.md` only track backend work. --- # Part A — vLLM Service (PaddleOCR-VL-1.6) This repository serves **PaddleOCR-VL-1.6** as a dedicated VLM inference backend using **vLLM**. All Python workflows use **uv** (never bare `pip` or system Python). Full detail (client usage examples, tuning, troubleshooting, issue-file template) moved to [docs/vllm-service.md](docs/vllm-service.md) 2026-07-08 to keep this file under the Part B kit's 256-line threshold (§3) — this section keeps only the essentials. ## Architecture ``` Client (PaddleOCR pipeline) --> HTTP /v1 --> paddleocr genai_server (vLLM backend) ``` This service exposes only the VLM stage. Clients connect with `vl_rec_backend="vllm-server"` and `vl_rec_server_url="http://:8118/v1"`. ## Quick start ```bash ./scripts/install.sh # 1) Create Python 3.12 venv and install dependencies ./scripts/serve.sh # 2) Start the vLLM-backed genai server ``` Default endpoint: `http://0.0.0.0:8118/v1`. Never use `python -m pip`, `pip install`, or `python -m venv` directly in this repo — always `uv sync` / `uv run` / `uv add`. ## Issue recording (always follow) **Every problem encountered** during install, serve, debug, or client integration must be written to `issues/{NN}-{slug}.md` before moving on — even if resolved in the same session. Naming/template details: [docs/vllm-service.md](docs/vllm-service.md#issue-recording--naming-and-template). ## Environment variables Copy `.env.example` to `.env` and adjust as needed: | Variable | Default | Description | |----------|---------|-------------| | `GENAI_HOST` | `0.0.0.0` | Bind address | | `GENAI_PORT` | `8118` | Service port | | `GENAI_MODEL` | `PaddleOCR-VL-1.6-0.9B` | Model name for `genai_server` | | `GENAI_BACKEND` | `vllm` | Inference backend | | `VLLM_CONFIG` | `config/vllm_config.yaml` | vLLM backend YAML config | | `CUDA_VISIBLE_DEVICES` | `1` (see `.env.example`) | GPU index(es) to use | On dual-GPU hosts, pick the GPU with more free VRAM. If startup fails with a memory error, lower `gpu-memory-utilization` in `config/vllm_config.yaml` — see [docs/vllm-service.md](docs/vllm-service.md#gpu-memory-on-startup). ## File map | Path | Purpose | |------|---------| | `issues/` | Recorded problems and fixes (`{NN}-{slug}.md`) | | `pyproject.toml` | uv project metadata and base dependencies | | `scripts/install.sh` | Bootstrap venv + vLLM server deps | | `scripts/serve.sh` | Start `paddleocr genai_server` | | `config/vllm_config.yaml` | vLLM backend tuning | | `.env.example` | Environment variable template | | `docs/vllm-service.md` | Full vLLM reference (client usage, tuning, troubleshooting) | ## Coding Guidelines (always follow) We use the karpathy-guidelines skill to reduce common LLM coding mistakes: 1. **Think Before Coding**: Explicitly state assumptions and surface tradeoffs instead of making silent choices. 2. **Simplicity First**: Write the minimum amount of code to solve the problem with zero speculative configurations. 3. **Surgical Changes**: Edit only what is required and match the existing coding style exactly. 4. **Goal-Driven Execution**: Define verifiable success criteria and run automated tests/screenshots to confirm correctness. 5. **SOLID Principles**: Always design, implement, and refactor code adhering to SOLID programming principles (Single Responsibility, Open/Closed, Liskov Substitution, Interface Segregation, Dependency Inversion) to ensure modularity, scalability, and maintainability. ## Path Guidelines (always follow) Never use full paths containing the user's logged-in name (e.g., `/home/{uid}/path`). Always use relative paths instead (e.g., `.` or `./path` relative to the workspace root). ## App Testing Guidelines (always follow) When the user intentionally asks to test the app: - Use browser tools to test the app. - Take a screenshot for each sample image, each step, and each variant/option (if any), until the OCR result appears. - Save the screenshots in the `/screenshots/` folder. - Follow the file naming convention: `{2-digit-number}-{step#}-{variant_or_options_if_any}-{slug}.jpg` (e.g., `01-step1-default-upload.jpg`). --- # Part B — Agents Settings Kit (backend-scoped `e`/`n` workflow) Backend-scoped copy of the [fhanyuh/agents-settings](https://github.com/fhanyuh/agents-settings) kit, adopted 2026-07-08. Covers only `backend/` modules (Next.js API Gateway, OCR Pipeline & Accuracy, Postgres Data Layer, DevOps/Docker) — Flutter modules are tracked by the separate copy at root `../AGENTS.md`. `../CLAUDE.md` (root) and `CLAUDE.md` (this dir) each import their own copy. ## B0. Adopting Into an Existing Project Already done for this repo (this split *is* that adoption, mirroring root's own §0 audit). Re-run "i"/"init" here to force a re-audit of `backend/` specifically (e.g. after a large refactor). ## B1. Trigger "e" or "enhance" - Read `plans/next-enhancements.md` (this dir) to understand current backend structure, history, and active tasks. - Overwrite or update the active tasks list inside it. - The plan must cover each backend section/module. - Define **exactly 3 new enhancements per section**, each with a unique number (e.g. `1.1`), a clear functional description, and status `[TODO]`. - Present the plan to the user in your final summary. ## B2. Trigger "n", "next", or "n{x}" - Read `plans/next-enhancements.md` to check task status. - If all tasks are `[DONE]` (or none `[TODO]`), run **"e"/"enhance"** first. - Otherwise select the most impactful `[TODO]` task(s) by strategic value/impact — not just first-in-order. If `{x}` given, take the top `{x}` sequentially. ### B2a. Clarify before building ("Grill Me" step) Same rule as root AGENTS.md §2a: if scope/acceptance criteria are genuinely ambiguous, ask one question at a time (`AskUserQuestion` in Claude Code) until unambiguous, and record the resolved criteria as a 1-3 line note next to the task entry before writing code. Skip when the task is already unambiguous. ### B2b. TDD Workflow (Test First) - **Write Tests First**: Before implementing the actual feature code for a task, write automated tests defining the expected behavior. - **Iterate Until Green**: Run the tests to confirm they fail, then write the implementation until all tests pass perfectly. - **Browser Testing**: If the enhancement involves web UI or visual components, use browser tools (e.g., Chrome) to test the app visually and functionally if necessary. - Implement the task(s) fully, applying the relevant role(s) from `SKILLS.md` (this dir). - On completion: 1. Flip status to `[DONE]` in `plans/next-enhancements.md`. 2. Document the feature in `docs/feature-list.md` (this dir) under the right section. 3. **Create an Iteration Log**: Perform a code review and audit of the tasks just completed. Document this audit in `docs/iteration-log.md` (or append to it) to ensure all functions work perfectly. 4. **Update Documentation**: Sync any architecture or workflow changes back to `CLAUDE.md` and `SKILLS.md` to keep the agent instructions current. - **Verify build integrity**: QA pass (golden path + edge cases + regression check on adjacent features — see `backend/CLAUDE.md`'s accuracy regression harness for OCR/parser changes specifically) and Hardware/Compatibility pass (cross-platform, GPU/VRAM footprint under Local/on-prem deployment — see Part A above). - State which task(s) were completed and the exact route/endpoint/menu path to see the new feature. ## B3. File Size & Refactoring Rules Same 256-line threshold as root AGENTS.md §3, backend-wide. Applies to this file, `SKILLS.md`, and `CLAUDE.md` too — which is why Part A above was trimmed and linked out to `docs/vllm-service.md` rather than left inline. ## B4. Roles See `SKILLS.md` (this dir) — same 5 roles as root (Architect, Backend, Frontend, QA, Hardware/Compatibility), applied to backend surfaces only (API routes, OCR pipeline, DB layer, Docker/deploy). ## B5. Mockup Data & Demo/Live Mode Same as root AGENTS.md §5: mock data under `/data/mockup/`, a mock API layer mirroring the real backend contract, and a Demo/Live switcher. Not yet built for backend — see Adaptation Notes. ## B6. Cloud vs Local (On-Premise) Same as root AGENTS.md §6, applied to backend service endpoints (Next.js gateway, pipeline API, vLLM server, Postgres) rather than the Flutter client's API base URL. ## B7. Ad-hoc Feature Requests Direct feature requests not using "e"/"n": implement and document in `docs/feature-list.md` (this dir). ## Adaptation Notes (backend, split from root 2026-07-08) - **Origin**: sections 5-8 of root `plans/next-enhancements.md` (Backend — Next.js API Gateway, Backend — OCR Pipeline & Accuracy, Backend — Postgres Data Layer, DevOps — Docker & Dev Tunnel) copied here as sections 1-4, statuses re-verified against the live code before the copy (not copied blind) — see task 7.1's `withTransaction` claim, task 5.1/5.2's dedup + timeout claims, and task 6.1's empty `models/` claim, all confirmed still accurate as of 2026-07-08. The root copy is frozen/archival (see root `AGENTS.md`'s "Scope: excludes `backend/`") rather than deleted, so this file — not the root one — is the single active source of truth going forward. - **Real commands**: `npm run dev`/`build`/`lint` in `pfm-web-app/`; accuracy regression harness `node pfm-web-app/scripts/accuracy-check.mts`; Python services via `./scripts/install.sh` + `./scripts/serve.sh` (this vLLM repo) and `./scripts/install-pipeline.sh` + `./scripts/serve-pipeline.sh` (pipeline API + classifier). Full stack: `docker compose up -d --build` **from the repo root**, not from inside `backend/` (see root `CLAUDE.md` — two `docker-compose.yml` files exist and running from here risks container-name conflicts). - **Pre-existing files over the 256-line threshold** (§B3 debt, not a blocker — split only if/when touched): `pfm-web-app/src/app/scan-pfm/page.tsx` (1169), `pfm-web-app/src/utils/parser.ts` (908), `pfm-web-app/src/app/page.tsx` (737), `config/classify_ocr_server.py` (691), `pfm-web-app/src/app/manual-label/page.tsx` (612), `pfm-web-app/src/app/api/parse/route.ts` (604), `pfm-web-app/src/db/init.ts` (477), `compare_sources_accuracy.py` (451), `pfm-web-app/src/utils/docker.ts` (362), `pfm-web-app/public/produk-pfm/train_classifier.py` (351), `compare_accuracy.py` (308), `pfm-web-app/src/app/api/arena/route.ts` (265). This file itself (`AGENTS.md`) was at 237 lines pre-kit and would have exceeded 256 once Part B was appended — hence the split into `docs/vllm-service.md`. - **No Demo/Live or Cloud/Local switch exists yet** (§B5, §B6) for the backend either. `docker-compose.override.yml` exposing `db`/`pipeline-api` directly to the host is a local-dev convenience, not a Cloud/Local deployment switch. - **Naming collision resolved by this split**: `AGENTS.md` already existed in this directory (vLLM service doc, committed 2026-06-30, unrelated to this kit) before Part B was appended — unlike root, where `AGENTS.md` didn't previously exist. Don't assume backend's `AGENTS.md` is kit-only when reading it from another tool; Part A is unrelated, pre-existing content kept for a reason. - **Pre-existing, unrelated governance files left as-is**: `.agents/AGENTS.md` at the *repo root* (different path, OCR post-processing rules) and root `plans/next-enhancement-plan.md` (singular, `[DONE]` QA checklist) — neither is part of this kit; see root `AGENTS.md`'s own Adaptation Notes.