No content changes: git diff --ignore-all-space over these files is empty. The churn came from editing on Windows against a repo checked out with LF.
189 lines
12 KiB
Markdown
189 lines
12 KiB
Markdown
# AGENTS: PaddleOCR-VL-1.6 vLLM Service + Agents Settings Kit
|
|
|
|
This is the authoritative rules file for any AI coding agent (Claude Code, Cursor,
|
|
GitHub Copilot, Aider, etc.) working inside `backend/`. Two unrelated concerns live
|
|
here side by side: **Part A** is this repo's original vLLM/PaddleOCR service doc.
|
|
**Part B** (appended 2026-07-08) is a **backend-scoped copy** of the
|
|
[fhanyuh/agents-settings](https://github.com/fhanyuh/agents-settings) `e`/`enhance`
|
|
and `n`/`next` workflow — see root `../AGENTS.md` for the same kit covering the
|
|
Flutter side of this repo. The two copies are independent: this one's
|
|
`plans/next-enhancements.md` and `docs/feature-list.md` only track backend work.
|
|
|
|
---
|
|
|
|
# Part A — vLLM Service (PaddleOCR-VL-1.6)
|
|
|
|
This repository serves **PaddleOCR-VL-1.6** as a dedicated VLM inference backend using **vLLM**. All Python workflows use **uv** (never bare `pip` or system Python). Full detail (client usage examples, tuning, troubleshooting, issue-file template) moved to [docs/vllm-service.md](docs/vllm-service.md) 2026-07-08 to keep this file under the Part B kit's 256-line threshold (§3) — this section keeps only the essentials.
|
|
|
|
## Architecture
|
|
|
|
```
|
|
Client (PaddleOCR pipeline) --> HTTP /v1 --> paddleocr genai_server (vLLM backend)
|
|
```
|
|
|
|
This service exposes only the VLM stage. Clients connect with `vl_rec_backend="vllm-server"` and `vl_rec_server_url="http://<host>:8118/v1"`.
|
|
|
|
## Quick start
|
|
|
|
```bash
|
|
./scripts/install.sh # 1) Create Python 3.12 venv and install dependencies
|
|
./scripts/serve.sh # 2) Start the vLLM-backed genai server
|
|
```
|
|
|
|
Default endpoint: `http://0.0.0.0:8118/v1`. Never use `python -m pip`, `pip install`, or `python -m venv` directly in this repo — always `uv sync` / `uv run` / `uv add`.
|
|
|
|
## Issue recording (always follow)
|
|
|
|
**Every problem encountered** during install, serve, debug, or client integration must be written to `issues/{NN}-{slug}.md` before moving on — even if resolved in the same session. Naming/template details: [docs/vllm-service.md](docs/vllm-service.md#issue-recording--naming-and-template).
|
|
|
|
## Environment variables
|
|
|
|
Copy `.env.example` to `.env` and adjust as needed:
|
|
|
|
| Variable | Default | Description |
|
|
|----------|---------|-------------|
|
|
| `GENAI_HOST` | `0.0.0.0` | Bind address |
|
|
| `GENAI_PORT` | `8118` | Service port |
|
|
| `GENAI_MODEL` | `PaddleOCR-VL-1.6-0.9B` | Model name for `genai_server` |
|
|
| `GENAI_BACKEND` | `vllm` | Inference backend |
|
|
| `VLLM_CONFIG` | `config/vllm_config.yaml` | vLLM backend YAML config |
|
|
| `CUDA_VISIBLE_DEVICES` | `1` (see `.env.example`) | GPU index(es) to use |
|
|
|
|
On dual-GPU hosts, pick the GPU with more free VRAM. If startup fails with a memory error, lower `gpu-memory-utilization` in `config/vllm_config.yaml` — see [docs/vllm-service.md](docs/vllm-service.md#gpu-memory-on-startup).
|
|
|
|
## File map
|
|
|
|
| Path | Purpose |
|
|
|------|---------|
|
|
| `issues/` | Recorded problems and fixes (`{NN}-{slug}.md`) |
|
|
| `pyproject.toml` | uv project metadata and base dependencies |
|
|
| `scripts/install.sh` | Bootstrap venv + vLLM server deps |
|
|
| `scripts/serve.sh` | Start `paddleocr genai_server` |
|
|
| `config/vllm_config.yaml` | vLLM backend tuning |
|
|
| `.env.example` | Environment variable template |
|
|
| `docs/vllm-service.md` | Full vLLM reference (client usage, tuning, troubleshooting) |
|
|
|
|
## Coding Guidelines (always follow)
|
|
|
|
We use the karpathy-guidelines skill to reduce common LLM coding mistakes:
|
|
1. **Think Before Coding**: Explicitly state assumptions and surface tradeoffs instead of making silent choices.
|
|
2. **Simplicity First**: Write the minimum amount of code to solve the problem with zero speculative configurations.
|
|
3. **Surgical Changes**: Edit only what is required and match the existing coding style exactly.
|
|
4. **Goal-Driven Execution**: Define verifiable success criteria and run automated tests/screenshots to confirm correctness.
|
|
5. **SOLID Principles**: Always design, implement, and refactor code adhering to SOLID programming principles (Single Responsibility, Open/Closed, Liskov Substitution, Interface Segregation, Dependency Inversion) to ensure modularity, scalability, and maintainability.
|
|
|
|
## Path Guidelines (always follow)
|
|
|
|
Never use full paths containing the user's logged-in name (e.g., `/home/{uid}/path`). Always use relative paths instead (e.g., `.` or `./path` relative to the workspace root).
|
|
|
|
## App Testing Guidelines (always follow)
|
|
|
|
When the user intentionally asks to test the app:
|
|
- Use browser tools to test the app.
|
|
- Take a screenshot for each sample image, each step, and each variant/option (if any), until the OCR result appears.
|
|
- Save the screenshots in the `/screenshots/` folder.
|
|
- Follow the file naming convention: `{2-digit-number}-{step#}-{variant_or_options_if_any}-{slug}.jpg` (e.g., `01-step1-default-upload.jpg`).
|
|
|
|
---
|
|
|
|
# Part B — Agents Settings Kit (backend-scoped `e`/`n` workflow)
|
|
|
|
Backend-scoped copy of the [fhanyuh/agents-settings](https://github.com/fhanyuh/agents-settings) kit, adopted 2026-07-08. Covers only `backend/` modules (Next.js API Gateway, OCR Pipeline & Accuracy, Postgres Data Layer, DevOps/Docker) — Flutter modules are tracked by the separate copy at root `../AGENTS.md`. `../CLAUDE.md` (root) and `CLAUDE.md` (this dir) each import their own copy.
|
|
|
|
## B0. Adopting Into an Existing Project
|
|
|
|
Already done for this repo (this split *is* that adoption, mirroring root's own §0 audit). Re-run "i"/"init" here to force a re-audit of `backend/` specifically (e.g. after a large refactor).
|
|
|
|
## B1. Trigger "e" or "enhance"
|
|
|
|
- Read `plans/next-enhancements.md` (this dir) to understand current backend structure, history, and active tasks.
|
|
- Overwrite or update the active tasks list inside it.
|
|
- The plan must cover each backend section/module.
|
|
- Define **exactly 3 new enhancements per section**, each with a unique number (e.g. `1.1`), a clear functional description, and status `[TODO]`.
|
|
- Present the plan to the user in your final summary.
|
|
|
|
## B2. Trigger "n", "next", or "n{x}"
|
|
|
|
- Read `plans/next-enhancements.md` to check task status.
|
|
- If all tasks are `[DONE]` (or none `[TODO]`), run **"e"/"enhance"** first.
|
|
- Otherwise select the most impactful `[TODO]` task(s) by strategic value/impact — not just first-in-order. If `{x}` given, take the top `{x}` sequentially.
|
|
|
|
### B2a. Clarify before building ("Grill Me" step)
|
|
|
|
Same rule as root AGENTS.md §2a: if scope/acceptance criteria are genuinely ambiguous, ask one question at a time (`AskUserQuestion` in Claude Code) until unambiguous, and record the resolved criteria as a 1-3 line note next to the task entry before writing code. Skip when the task is already unambiguous.
|
|
|
|
### B2b. TDD Workflow (Test First)
|
|
|
|
- **Write Tests First**: Before implementing the actual feature code for a task, write automated tests defining the expected behavior.
|
|
- **Iterate Until Green**: Run the tests to confirm they fail, then write the implementation until all tests pass perfectly.
|
|
- **Browser Testing**: If the enhancement involves web UI or visual components, use browser tools (e.g., Chrome) to test the app visually and functionally if necessary.
|
|
|
|
- Implement the task(s) fully, applying the relevant role(s) from `SKILLS.md` (this dir).
|
|
- On completion:
|
|
1. Flip status to `[DONE]` in `plans/next-enhancements.md`.
|
|
2. Document the feature in `docs/feature-list.md` (this dir) under the right section.
|
|
3. **Create an Iteration Log**: Perform a code review and audit of the tasks just completed. Document this audit in `docs/iteration-log.md` (or append to it) to ensure all functions work perfectly.
|
|
4. **Update Documentation**: Sync any architecture or workflow changes back to `CLAUDE.md` and `SKILLS.md` to keep the agent instructions current.
|
|
- **Verify build integrity**: QA pass (golden path + edge cases + regression check on adjacent features — see `backend/CLAUDE.md`'s accuracy regression harness for OCR/parser changes specifically) and Hardware/Compatibility pass (cross-platform, GPU/VRAM footprint under Local/on-prem deployment — see Part A above).
|
|
- State which task(s) were completed and the exact route/endpoint/menu path to see the new feature.
|
|
|
|
## B3. File Size & Refactoring Rules
|
|
|
|
Same 256-line threshold as root AGENTS.md §3, backend-wide. Applies to this file, `SKILLS.md`, and `CLAUDE.md` too — which is why Part A above was trimmed and linked out to `docs/vllm-service.md` rather than left inline.
|
|
|
|
## B4. Roles
|
|
|
|
See `SKILLS.md` (this dir) — same 5 roles as root (Architect, Backend, Frontend, QA, Hardware/Compatibility), applied to backend surfaces only (API routes, OCR pipeline, DB layer, Docker/deploy).
|
|
|
|
## B5. Mockup Data & Demo/Live Mode
|
|
|
|
Same as root AGENTS.md §5: mock data under `/data/mockup/`, a mock API layer mirroring the real backend contract, and a Demo/Live switcher. Not yet built for backend — see Adaptation Notes.
|
|
|
|
## B6. Cloud vs Local (On-Premise)
|
|
|
|
Same as root AGENTS.md §6, applied to backend service endpoints (Next.js gateway, pipeline API, vLLM server, Postgres) rather than the Flutter client's API base URL.
|
|
|
|
## B7. Ad-hoc Feature Requests
|
|
|
|
Direct feature requests not using "e"/"n": implement and document in `docs/feature-list.md` (this dir).
|
|
|
|
## Adaptation Notes (backend, split from root 2026-07-08)
|
|
|
|
- **Origin**: sections 5-8 of root `plans/next-enhancements.md` (Backend — Next.js API
|
|
Gateway, Backend — OCR Pipeline & Accuracy, Backend — Postgres Data Layer, DevOps —
|
|
Docker & Dev Tunnel) copied here as sections 1-4, statuses re-verified against the
|
|
live code before the copy (not copied blind) — see task 7.1's `withTransaction`
|
|
claim, task 5.1/5.2's dedup + timeout claims, and task 6.1's empty `models/` claim,
|
|
all confirmed still accurate as of 2026-07-08. The root copy is frozen/archival
|
|
(see root `AGENTS.md`'s "Scope: excludes `backend/`") rather than deleted, so this
|
|
file — not the root one — is the single active source of truth going forward.
|
|
- **Real commands**: `npm run dev`/`build`/`lint` in `pfm-web-app/`; accuracy
|
|
regression harness `node pfm-web-app/scripts/accuracy-check.mts`; Python services
|
|
via `./scripts/install.sh` + `./scripts/serve.sh` (this vLLM repo) and
|
|
`./scripts/install-pipeline.sh` + `./scripts/serve-pipeline.sh` (pipeline API +
|
|
classifier). Full stack: `docker compose up -d --build` **from the repo root**, not
|
|
from inside `backend/` (see root `CLAUDE.md` — two `docker-compose.yml` files
|
|
exist and running from here risks container-name conflicts).
|
|
- **Pre-existing files over the 256-line threshold** (§B3 debt, not a blocker — split
|
|
only if/when touched): `pfm-web-app/src/app/scan-pfm/page.tsx` (1169),
|
|
`pfm-web-app/src/utils/parser.ts` (908), `pfm-web-app/src/app/page.tsx` (737),
|
|
`config/classify_ocr_server.py` (691), `pfm-web-app/src/app/manual-label/page.tsx`
|
|
(612), `pfm-web-app/src/app/api/parse/route.ts` (604), `pfm-web-app/src/db/init.ts`
|
|
(477), `compare_sources_accuracy.py` (451), `pfm-web-app/src/utils/docker.ts` (362),
|
|
`pfm-web-app/public/produk-pfm/train_classifier.py` (351), `compare_accuracy.py`
|
|
(308), `pfm-web-app/src/app/api/arena/route.ts` (265). This file itself (`AGENTS.md`)
|
|
was at 237 lines pre-kit and would have exceeded 256 once Part B was appended —
|
|
hence the split into `docs/vllm-service.md`.
|
|
- **No Demo/Live or Cloud/Local switch exists yet** (§B5, §B6) for the backend
|
|
either. `docker-compose.override.yml` exposing `db`/`pipeline-api` directly to the
|
|
host is a local-dev convenience, not a Cloud/Local deployment switch.
|
|
- **Naming collision resolved by this split**: `AGENTS.md` already existed in this
|
|
directory (vLLM service doc, committed 2026-06-30, unrelated to this kit) before
|
|
Part B was appended — unlike root, where `AGENTS.md` didn't previously exist. Don't
|
|
assume backend's `AGENTS.md` is kit-only when reading it from another tool; Part A
|
|
is unrelated, pre-existing content kept for a reason.
|
|
- **Pre-existing, unrelated governance files left as-is**: `.agents/AGENTS.md` at the
|
|
*repo root* (different path, OCR post-processing rules) and root
|
|
`plans/next-enhancement-plan.md` (singular, `[DONE]` QA checklist) — neither is
|
|
part of this kit; see root `AGENTS.md`'s own Adaptation Notes.
|