Files
fhanyuh caf8e98378 chore: normalize line endings (CRLF -> LF)
No content changes: git diff --ignore-all-space over these files is empty.
The churn came from editing on Windows against a repo checked out with LF.
2026-08-27 10:40:49 +07:00

12 KiB

AGENTS: PaddleOCR-VL-1.6 vLLM Service + Agents Settings Kit

This is the authoritative rules file for any AI coding agent (Claude Code, Cursor, GitHub Copilot, Aider, etc.) working inside backend/. Two unrelated concerns live here side by side: Part A is this repo's original vLLM/PaddleOCR service doc. Part B (appended 2026-07-08) is a backend-scoped copy of the fhanyuh/agents-settings e/enhance and n/next workflow — see root ../AGENTS.md for the same kit covering the Flutter side of this repo. The two copies are independent: this one's plans/next-enhancements.md and docs/feature-list.md only track backend work.


Part A — vLLM Service (PaddleOCR-VL-1.6)

This repository serves PaddleOCR-VL-1.6 as a dedicated VLM inference backend using vLLM. All Python workflows use uv (never bare pip or system Python). Full detail (client usage examples, tuning, troubleshooting, issue-file template) moved to docs/vllm-service.md 2026-07-08 to keep this file under the Part B kit's 256-line threshold (§3) — this section keeps only the essentials.

Architecture

Client (PaddleOCR pipeline)  -->  HTTP /v1  -->  paddleocr genai_server (vLLM backend)

This service exposes only the VLM stage. Clients connect with vl_rec_backend="vllm-server" and vl_rec_server_url="http://<host>:8118/v1".

Quick start

./scripts/install.sh   # 1) Create Python 3.12 venv and install dependencies
./scripts/serve.sh     # 2) Start the vLLM-backed genai server

Default endpoint: http://0.0.0.0:8118/v1. Never use python -m pip, pip install, or python -m venv directly in this repo — always uv sync / uv run / uv add.

Issue recording (always follow)

Every problem encountered during install, serve, debug, or client integration must be written to issues/{NN}-{slug}.md before moving on — even if resolved in the same session. Naming/template details: docs/vllm-service.md.

Environment variables

Copy .env.example to .env and adjust as needed:

Variable Default Description
GENAI_HOST 0.0.0.0 Bind address
GENAI_PORT 8118 Service port
GENAI_MODEL PaddleOCR-VL-1.6-0.9B Model name for genai_server
GENAI_BACKEND vllm Inference backend
VLLM_CONFIG config/vllm_config.yaml vLLM backend YAML config
CUDA_VISIBLE_DEVICES 1 (see .env.example) GPU index(es) to use

On dual-GPU hosts, pick the GPU with more free VRAM. If startup fails with a memory error, lower gpu-memory-utilization in config/vllm_config.yaml — see docs/vllm-service.md.

File map

Path Purpose
issues/ Recorded problems and fixes ({NN}-{slug}.md)
pyproject.toml uv project metadata and base dependencies
scripts/install.sh Bootstrap venv + vLLM server deps
scripts/serve.sh Start paddleocr genai_server
config/vllm_config.yaml vLLM backend tuning
.env.example Environment variable template
docs/vllm-service.md Full vLLM reference (client usage, tuning, troubleshooting)

Coding Guidelines (always follow)

We use the karpathy-guidelines skill to reduce common LLM coding mistakes:

  1. Think Before Coding: Explicitly state assumptions and surface tradeoffs instead of making silent choices.
  2. Simplicity First: Write the minimum amount of code to solve the problem with zero speculative configurations.
  3. Surgical Changes: Edit only what is required and match the existing coding style exactly.
  4. Goal-Driven Execution: Define verifiable success criteria and run automated tests/screenshots to confirm correctness.
  5. SOLID Principles: Always design, implement, and refactor code adhering to SOLID programming principles (Single Responsibility, Open/Closed, Liskov Substitution, Interface Segregation, Dependency Inversion) to ensure modularity, scalability, and maintainability.

Path Guidelines (always follow)

Never use full paths containing the user's logged-in name (e.g., /home/{uid}/path). Always use relative paths instead (e.g., . or ./path relative to the workspace root).

App Testing Guidelines (always follow)

When the user intentionally asks to test the app:

  • Use browser tools to test the app.
  • Take a screenshot for each sample image, each step, and each variant/option (if any), until the OCR result appears.
  • Save the screenshots in the /screenshots/ folder.
  • Follow the file naming convention: {2-digit-number}-{step#}-{variant_or_options_if_any}-{slug}.jpg (e.g., 01-step1-default-upload.jpg).

Part B — Agents Settings Kit (backend-scoped e/n workflow)

Backend-scoped copy of the fhanyuh/agents-settings kit, adopted 2026-07-08. Covers only backend/ modules (Next.js API Gateway, OCR Pipeline & Accuracy, Postgres Data Layer, DevOps/Docker) — Flutter modules are tracked by the separate copy at root ../AGENTS.md. ../CLAUDE.md (root) and CLAUDE.md (this dir) each import their own copy.

B0. Adopting Into an Existing Project

Already done for this repo (this split is that adoption, mirroring root's own §0 audit). Re-run "i"/"init" here to force a re-audit of backend/ specifically (e.g. after a large refactor).

B1. Trigger "e" or "enhance"

  • Read plans/next-enhancements.md (this dir) to understand current backend structure, history, and active tasks.
  • Overwrite or update the active tasks list inside it.
  • The plan must cover each backend section/module.
  • Define exactly 3 new enhancements per section, each with a unique number (e.g. 1.1), a clear functional description, and status [TODO].
  • Present the plan to the user in your final summary.

B2. Trigger "n", "next", or "n{x}"

  • Read plans/next-enhancements.md to check task status.
  • If all tasks are [DONE] (or none [TODO]), run "e"/"enhance" first.
  • Otherwise select the most impactful [TODO] task(s) by strategic value/impact — not just first-in-order. If {x} given, take the top {x} sequentially.

B2a. Clarify before building ("Grill Me" step)

Same rule as root AGENTS.md §2a: if scope/acceptance criteria are genuinely ambiguous, ask one question at a time (AskUserQuestion in Claude Code) until unambiguous, and record the resolved criteria as a 1-3 line note next to the task entry before writing code. Skip when the task is already unambiguous.

B2b. TDD Workflow (Test First)

  • Write Tests First: Before implementing the actual feature code for a task, write automated tests defining the expected behavior.

  • Iterate Until Green: Run the tests to confirm they fail, then write the implementation until all tests pass perfectly.

  • Browser Testing: If the enhancement involves web UI or visual components, use browser tools (e.g., Chrome) to test the app visually and functionally if necessary.

  • Implement the task(s) fully, applying the relevant role(s) from SKILLS.md (this dir).

  • On completion:

    1. Flip status to [DONE] in plans/next-enhancements.md.
    2. Document the feature in docs/feature-list.md (this dir) under the right section.
    3. Create an Iteration Log: Perform a code review and audit of the tasks just completed. Document this audit in docs/iteration-log.md (or append to it) to ensure all functions work perfectly.
    4. Update Documentation: Sync any architecture or workflow changes back to CLAUDE.md and SKILLS.md to keep the agent instructions current.
  • Verify build integrity: QA pass (golden path + edge cases + regression check on adjacent features — see backend/CLAUDE.md's accuracy regression harness for OCR/parser changes specifically) and Hardware/Compatibility pass (cross-platform, GPU/VRAM footprint under Local/on-prem deployment — see Part A above).

  • State which task(s) were completed and the exact route/endpoint/menu path to see the new feature.

B3. File Size & Refactoring Rules

Same 256-line threshold as root AGENTS.md §3, backend-wide. Applies to this file, SKILLS.md, and CLAUDE.md too — which is why Part A above was trimmed and linked out to docs/vllm-service.md rather than left inline.

B4. Roles

See SKILLS.md (this dir) — same 5 roles as root (Architect, Backend, Frontend, QA, Hardware/Compatibility), applied to backend surfaces only (API routes, OCR pipeline, DB layer, Docker/deploy).

B5. Mockup Data & Demo/Live Mode

Same as root AGENTS.md §5: mock data under /data/mockup/, a mock API layer mirroring the real backend contract, and a Demo/Live switcher. Not yet built for backend — see Adaptation Notes.

B6. Cloud vs Local (On-Premise)

Same as root AGENTS.md §6, applied to backend service endpoints (Next.js gateway, pipeline API, vLLM server, Postgres) rather than the Flutter client's API base URL.

B7. Ad-hoc Feature Requests

Direct feature requests not using "e"/"n": implement and document in docs/feature-list.md (this dir).

Adaptation Notes (backend, split from root 2026-07-08)

  • Origin: sections 5-8 of root plans/next-enhancements.md (Backend — Next.js API Gateway, Backend — OCR Pipeline & Accuracy, Backend — Postgres Data Layer, DevOps — Docker & Dev Tunnel) copied here as sections 1-4, statuses re-verified against the live code before the copy (not copied blind) — see task 7.1's withTransaction claim, task 5.1/5.2's dedup + timeout claims, and task 6.1's empty models/ claim, all confirmed still accurate as of 2026-07-08. The root copy is frozen/archival (see root AGENTS.md's "Scope: excludes backend/") rather than deleted, so this file — not the root one — is the single active source of truth going forward.
  • Real commands: npm run dev/build/lint in pfm-web-app/; accuracy regression harness node pfm-web-app/scripts/accuracy-check.mts; Python services via ./scripts/install.sh + ./scripts/serve.sh (this vLLM repo) and ./scripts/install-pipeline.sh + ./scripts/serve-pipeline.sh (pipeline API + classifier). Full stack: docker compose up -d --build from the repo root, not from inside backend/ (see root CLAUDE.md — two docker-compose.yml files exist and running from here risks container-name conflicts).
  • Pre-existing files over the 256-line threshold (§B3 debt, not a blocker — split only if/when touched): pfm-web-app/src/app/scan-pfm/page.tsx (1169), pfm-web-app/src/utils/parser.ts (908), pfm-web-app/src/app/page.tsx (737), config/classify_ocr_server.py (691), pfm-web-app/src/app/manual-label/page.tsx (612), pfm-web-app/src/app/api/parse/route.ts (604), pfm-web-app/src/db/init.ts (477), compare_sources_accuracy.py (451), pfm-web-app/src/utils/docker.ts (362), pfm-web-app/public/produk-pfm/train_classifier.py (351), compare_accuracy.py (308), pfm-web-app/src/app/api/arena/route.ts (265). This file itself (AGENTS.md) was at 237 lines pre-kit and would have exceeded 256 once Part B was appended — hence the split into docs/vllm-service.md.
  • No Demo/Live or Cloud/Local switch exists yet (§B5, §B6) for the backend either. docker-compose.override.yml exposing db/pipeline-api directly to the host is a local-dev convenience, not a Cloud/Local deployment switch.
  • Naming collision resolved by this split: AGENTS.md already existed in this directory (vLLM service doc, committed 2026-06-30, unrelated to this kit) before Part B was appended — unlike root, where AGENTS.md didn't previously exist. Don't assume backend's AGENTS.md is kit-only when reading it from another tool; Part A is unrelated, pre-existing content kept for a reason.
  • Pre-existing, unrelated governance files left as-is: .agents/AGENTS.md at the repo root (different path, OCR post-processing rules) and root plans/next-enhancement-plan.md (singular, [DONE] QA checklist) — neither is part of this kit; see root AGENTS.md's own Adaptation Notes.