Adopt agents-settings kit, ship Product/SKU scan models, harden auth, verify OCR accuracy
Backend (app-pfm-ocr-v2/backend): - Product/SKU scan feature complete: trained DINOv2 index (118 reference photos, 16 SKU classes) and YOLO classifier (83.3% top-1 val accuracy), fixed scripts/install-pipeline.sh (was missing ultralytics/torch), fully browser-verified end-to-end on /scan-pfm. Mobile m-scan-pfm page cancelled (Flutter app handles mobile; web UI is desktop-only for pipeline testing). - Fixed a real data-loss bug: Save Ground Truth (scan-pfm and the DO-flow's manual-label) was silently writing into the pfm-web-app container's ephemeral filesystem instead of the host, because /sources wasn't bind-mounted in docker-compose.yml. Added the mount, recovered an orphaned entry. - accounts.password is now bcrypt-hashed (bcryptjs, idempotent migration in db/init.ts) instead of plaintext; login route compares hashes. - /api/v1/documents/* (list, PUT, upload) now enforces real 401 auth, matching what the Flutter client already sends. The "classic" routes deliberately stay open — they're dev-only web UI with no login flow and won't exist in production. - OCR accuracy investigated end-to-end: real baseline is 95.10% overall (target met; accuracy_report.md was stale at 75.04%, now flagged). Fixed one genuine parser.ts bug (SO/DO field duplication in the global fallback regex); remaining gaps are OCR/layout-model limitations, not parser bugs. - Adopted a standalone copy of the fhanyuh/agents-settings e/n workflow scoped to backend/ (AGENTS.md Part A/B split, SKILLS.md, plans/, docs/), independent of the root copy which now covers Flutter only. - next-implementation.md deleted; content folded into backend/plans/next-enhancements.md for traceability. Root: - Adopted fhanyuh/agents-settings kit (AGENTS.md, SKILLS.md, plans/, docs/feature-list.md), scoped to the Flutter app only. - Pending documents queue now persists to Hive (lib/core/storage) instead of memory-only, surviving an app kill mid-upload. Removed backend_backup/ (stale Express/Prisma prototype, superseded by pfm-web-app) and the completed plans/next-enhancement-plan.md checklist. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
1 parent
3df9f6ec5d
commit
e60ab63154
129 files changed
+8520
-6684
No files matched your search
+115
-175
@@ -1,6 +1,19 @@
|
||||
# AGENTS: PaddleOCR-VL-1.6 vLLM Service
|
||||
# AGENTS: PaddleOCR-VL-1.6 vLLM Service + Agents Settings Kit
|
||||
|
||||
This repository serves **PaddleOCR-VL-1.6** as a dedicated VLM inference backend using **vLLM**. All Python workflows use **uv** (never bare `pip` or system Python).
|
||||
This is the authoritative rules file for any AI coding agent (Claude Code, Cursor,
|
||||
GitHub Copilot, Aider, etc.) working inside `backend/`. Two unrelated concerns live
|
||||
here side by side: **Part A** is this repo's original vLLM/PaddleOCR service doc.
|
||||
**Part B** (appended 2026-07-08) is a **backend-scoped copy** of the
|
||||
[fhanyuh/agents-settings](https://github.com/fhanyuh/agents-settings) `e`/`enhance`
|
||||
and `n`/`next` workflow — see root `../AGENTS.md` for the same kit covering the
|
||||
Flutter side of this repo. The two copies are independent: this one's
|
||||
`plans/next-enhancements.md` and `docs/feature-list.md` only track backend work.
|
||||
|
||||
---
|
||||
|
||||
# Part A — vLLM Service (PaddleOCR-VL-1.6)
|
||||
|
||||
This repository serves **PaddleOCR-VL-1.6** as a dedicated VLM inference backend using **vLLM**. All Python workflows use **uv** (never bare `pip` or system Python). Full detail (client usage examples, tuning, troubleshooting, issue-file template) moved to [docs/vllm-service.md](docs/vllm-service.md) 2026-07-08 to keep this file under the Part B kit's 256-line threshold (§3) — this section keeps only the essentials.
|
||||
|
||||
## Architecture
|
||||
|
||||
@@ -10,90 +23,18 @@ Client (PaddleOCR pipeline) --> HTTP /v1 --> paddleocr genai_server (vLLM ba
|
||||
|
||||
This service exposes only the VLM stage. Clients connect with `vl_rec_backend="vllm-server"` and `vl_rec_server_url="http://<host>:8118/v1"`.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- Linux with NVIDIA GPU (CC >= 8.0 recommended; CUDA 12.6+ driver support)
|
||||
- [uv](https://docs.astral.sh/uv/) installed (`uv --version`)
|
||||
- ~16 GB GPU VRAM for default settings (tune via `config/vllm_config.yaml`)
|
||||
|
||||
## Quick start
|
||||
|
||||
```bash
|
||||
cd <YOUR-WORKING-DIR>/ai-ocr-pfm-2026
|
||||
|
||||
# 1) Create Python 3.12 venv and install dependencies
|
||||
./scripts/install.sh
|
||||
|
||||
# 2) Start the vLLM-backed genai server
|
||||
./scripts/serve.sh
|
||||
./scripts/install.sh # 1) Create Python 3.12 venv and install dependencies
|
||||
./scripts/serve.sh # 2) Start the vLLM-backed genai server
|
||||
```
|
||||
|
||||
Default endpoint: `http://0.0.0.0:8118/v1`
|
||||
|
||||
## uv conventions (always follow)
|
||||
|
||||
| Task | Command |
|
||||
|------|---------|
|
||||
| Create/sync env | `uv sync` |
|
||||
| Run any Python | `uv run <command>` |
|
||||
| Add a package | `uv add <package>` |
|
||||
| Run server | `./scripts/serve.sh` or `uv run paddleocr genai_server ...` |
|
||||
|
||||
Never use `python -m pip`, `pip install`, or `python -m venv` directly in this repo.
|
||||
Default endpoint: `http://0.0.0.0:8118/v1`. Never use `python -m pip`, `pip install`, or `python -m venv` directly in this repo — always `uv sync` / `uv run` / `uv add`.
|
||||
|
||||
## Issue recording (always follow)
|
||||
|
||||
**Every problem encountered** during install, serve, debug, or client integration must be written to `issues/` before moving on — even if it was resolved in the same session.
|
||||
|
||||
### Naming
|
||||
|
||||
```
|
||||
issues/{NN}-{slug}.md
|
||||
```
|
||||
|
||||
| Part | Rule | Example |
|
||||
|------|------|---------|
|
||||
| `{NN}` | Two-digit running number (`01`, `02`, …). Increment from the highest existing file. | `03` |
|
||||
| `{slug}` | Lowercase kebab-case summary of the problem | `gpu-memory-startup-failure` |
|
||||
|
||||
Full example: `issues/04-gpu-memory-startup-failure.md`
|
||||
|
||||
### When to create a file
|
||||
|
||||
- Install or dependency errors (flash-attn, vLLM, uv conflicts)
|
||||
- Server startup or runtime failures (OOM, port bind, model load)
|
||||
- Client integration bugs or misconfiguration
|
||||
- Workarounds that took non-obvious steps to discover
|
||||
|
||||
Do **not** rely on chat history or inline comments alone — if it blocked progress, it belongs in `issues/`.
|
||||
|
||||
### File template
|
||||
|
||||
```markdown
|
||||
# Issue {NN}: {Short title}
|
||||
|
||||
## Problem
|
||||
What failed, with exact error message or symptom.
|
||||
|
||||
## Context
|
||||
Environment, command run, relevant config (`.env`, `config/vllm_config.yaml`).
|
||||
|
||||
## Solution
|
||||
What fixed it, or current workaround / open status.
|
||||
|
||||
## References
|
||||
Links, related issue files, or AGENTS.md sections.
|
||||
```
|
||||
|
||||
### Index
|
||||
|
||||
Check `issues/` for the next number:
|
||||
|
||||
```bash
|
||||
ls issues/*.md 2>/dev/null | sort
|
||||
```
|
||||
|
||||
See [issues/](issues/) for recorded problems and fixes from this project.
|
||||
**Every problem encountered** during install, serve, debug, or client integration must be written to `issues/{NN}-{slug}.md` before moving on — even if resolved in the same session. Naming/template details: [docs/vllm-service.md](docs/vllm-service.md#issue-recording--naming-and-template).
|
||||
|
||||
## Environment variables
|
||||
|
||||
@@ -108,97 +49,7 @@ Copy `.env.example` to `.env` and adjust as needed:
|
||||
| `VLLM_CONFIG` | `config/vllm_config.yaml` | vLLM backend YAML config |
|
||||
| `CUDA_VISIBLE_DEVICES` | `1` (see `.env.example`) | GPU index(es) to use |
|
||||
|
||||
On dual-GPU hosts, pick the GPU with more free VRAM. If startup fails with a memory error, lower `gpu-memory-utilization` in `config/vllm_config.yaml`.
|
||||
|
||||
## Client usage
|
||||
|
||||
After the server is running:
|
||||
|
||||
```bash
|
||||
# CLI
|
||||
uv run paddleocr doc_parser \
|
||||
--input https://paddle-model-ecology.bj.bcebos.com/paddlex/imgs/demo_image/paddleocr_vl_demo.png \
|
||||
--vl_rec_backend vllm-server \
|
||||
--vl_rec_server_url http://localhost:8118/v1
|
||||
```
|
||||
|
||||
```python
|
||||
from paddleocr import PaddleOCRVL
|
||||
|
||||
pipeline = PaddleOCRVL(
|
||||
vl_rec_backend="vllm-server",
|
||||
vl_rec_server_url="http://127.0.0.1:8118/v1",
|
||||
)
|
||||
output = pipeline.predict("path/to/image.png")
|
||||
```
|
||||
|
||||
Note: The full PaddleOCR-VL client should run in a **separate** environment if it needs PaddlePaddle GPU + Transformers. This repo is the isolated vLLM server only.
|
||||
|
||||
## Tuning vLLM
|
||||
|
||||
Edit `config/vllm_config.yaml`:
|
||||
|
||||
```yaml
|
||||
gpu-memory-utilization: 0.8
|
||||
max-num-seqs: 128
|
||||
```
|
||||
|
||||
Reference: [PaddleOCR-VL vLLM parameter tuning](https://www.paddleocr.ai/latest/en/version3.x/pipeline_usage/PaddleOCR-VL.html#331-server-side-parameter-adjustment)
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
See `issues/` for full write-ups. Quick pointers:
|
||||
|
||||
| Symptom | Issue file |
|
||||
|---------|------------|
|
||||
| `paddleocr install_genai_server_deps` / `No module named pip` | [01-genai-server-deps-pip-in-uv-venv.md](issues/01-genai-server-deps-pip-in-uv-venv.md) |
|
||||
| flash-attn wheel incompatible with Python version | [02-flash-attn-wheel-python-version-mismatch.md](issues/02-flash-attn-wheel-python-version-mismatch.md) |
|
||||
| `uv pip` targets wrong venv from another project | [03-active-virtual-env-from-other-project.md](issues/03-active-virtual-env-from-other-project.md) |
|
||||
| Free memory below `gpu-memory-utilization` on startup | [04-gpu-memory-startup-failure.md](issues/04-gpu-memory-startup-failure.md) |
|
||||
| `TokenizersBackend has no attribute all_special_tokens_extended` | [05-transformers-tokenizers-incompatibility.md](issues/05-transformers-tokenizers-incompatibility.md) |
|
||||
| Extracted images not shown in Gradio demo (raw base64 in markdown) | [06-extracted-images-raw-base64-not-displayed.md](issues/06-extracted-images-raw-base64-not-displayed.md) |
|
||||
|
||||
### flash-attn build failures
|
||||
|
||||
Install the prebuilt wheel after `uv sync` (see `scripts/install.sh`):
|
||||
|
||||
```bash
|
||||
FLASH_ATTN_WHEEL="https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.3.14/flash_attn-2.8.2+cu128torch2.8-cp312-cp312-linux_x86_64.whl" \
|
||||
./scripts/install.sh
|
||||
```
|
||||
|
||||
Pick the wheel matching your Python and CUDA versions from [flash-attention prebuild wheels](https://mjunya.com/flash-attention-prebuild-wheels/). Details: [02-flash-attn-wheel-python-version-mismatch.md](issues/02-flash-attn-wheel-python-version-mismatch.md).
|
||||
|
||||
Note: `paddleocr install_genai_server_deps` uses `pip` internally and is incompatible with uv-managed venvs. See [01-genai-server-deps-pip-in-uv-venv.md](issues/01-genai-server-deps-pip-in-uv-venv.md). This repo installs the vLLM stack via `uv sync` + `uv pip`.
|
||||
|
||||
### `TokenizersBackend has no attribute all_special_tokens_extended`
|
||||
|
||||
Pin transformers (already in `pyproject.toml`):
|
||||
|
||||
```bash
|
||||
uv pip install "transformers==4.57.6"
|
||||
```
|
||||
|
||||
See [05-transformers-tokenizers-incompatibility.md](issues/05-transformers-tokenizers-incompatibility.md).
|
||||
|
||||
### Do not install `paddlepaddle-gpu` in this venv
|
||||
|
||||
vLLM and PaddlePaddle GPU conflict. This server env uses `paddleocr[doc-parser]` without Paddle GPU.
|
||||
|
||||
### GPU memory on startup
|
||||
|
||||
If vLLM reports free memory below `gpu-memory-utilization`, either:
|
||||
|
||||
- Set `CUDA_VISIBLE_DEVICES` to a less-busy GPU
|
||||
- Lower `gpu-memory-utilization` in `config/vllm_config.yaml` (e.g. `0.75` or `0.7`)
|
||||
|
||||
See [04-gpu-memory-startup-failure.md](issues/04-gpu-memory-startup-failure.md).
|
||||
|
||||
### Health check
|
||||
|
||||
```bash
|
||||
curl -s http://localhost:8118/v1/models | jq .
|
||||
```
|
||||
On dual-GPU hosts, pick the GPU with more free VRAM. If startup fails with a memory error, lower `gpu-memory-utilization` in `config/vllm_config.yaml` — see [docs/vllm-service.md](docs/vllm-service.md#gpu-memory-on-startup).
|
||||
|
||||
## File map
|
||||
|
||||
@@ -210,15 +61,11 @@ curl -s http://localhost:8118/v1/models | jq .
|
||||
| `scripts/serve.sh` | Start `paddleocr genai_server` |
|
||||
| `config/vllm_config.yaml` | vLLM backend tuning |
|
||||
| `.env.example` | Environment variable template |
|
||||
|
||||
## References
|
||||
|
||||
- [PaddleOCR-VL usage tutorial](https://www.paddleocr.ai/latest/en/version3.x/pipeline_usage/PaddleOCR-VL.html)
|
||||
- [PaddleOCR genai_server FAQ](https://github.com/PaddlePaddle/PaddleOCR/discussions/16822)
|
||||
| `docs/vllm-service.md` | Full vLLM reference (client usage, tuning, troubleshooting) |
|
||||
|
||||
## Coding Guidelines (always follow)
|
||||
|
||||
We use the [karpathy-guidelines](file:///home/user/LABS/OCR/paddle-ocr-vl-1-6-using-vllm-2026/andrej-karpathy-skills/skills/karpathy-guidelines/SKILL.md) skill to reduce common LLM coding mistakes. Refer to [SKILL.md](file:///home/user/LABS/OCR/paddle-ocr-vl-1-6-using-vllm-2026/andrej-karpathy-skills/skills/karpathy-guidelines/SKILL.md) for details:
|
||||
We use the karpathy-guidelines skill to reduce common LLM coding mistakes:
|
||||
1. **Think Before Coding**: Explicitly state assumptions and surface tradeoffs instead of making silent choices.
|
||||
2. **Simplicity First**: Write the minimum amount of code to solve the problem with zero speculative configurations.
|
||||
3. **Surgical Changes**: Edit only what is required and match the existing coding style exactly.
|
||||
@@ -235,3 +82,96 @@ When the user intentionally asks to test the app:
|
||||
- Take a screenshot for each sample image, each step, and each variant/option (if any), until the OCR result appears.
|
||||
- Save the screenshots in the `/screenshots/` folder.
|
||||
- Follow the file naming convention: `{2-digit-number}-{step#}-{variant_or_options_if_any}-{slug}.jpg` (e.g., `01-step1-default-upload.jpg`).
|
||||
|
||||
---
|
||||
|
||||
# Part B — Agents Settings Kit (backend-scoped `e`/`n` workflow)
|
||||
|
||||
Backend-scoped copy of the [fhanyuh/agents-settings](https://github.com/fhanyuh/agents-settings) kit, adopted 2026-07-08. Covers only `backend/` modules (Next.js API Gateway, OCR Pipeline & Accuracy, Postgres Data Layer, DevOps/Docker) — Flutter modules are tracked by the separate copy at root `../AGENTS.md`. `../CLAUDE.md` (root) and `CLAUDE.md` (this dir) each import their own copy.
|
||||
|
||||
## B0. Adopting Into an Existing Project
|
||||
|
||||
Already done for this repo (this split *is* that adoption, mirroring root's own §0 audit). Re-run "i"/"init" here to force a re-audit of `backend/` specifically (e.g. after a large refactor).
|
||||
|
||||
## B1. Trigger "e" or "enhance"
|
||||
|
||||
- Read `plans/next-enhancements.md` (this dir) to understand current backend structure, history, and active tasks.
|
||||
- Overwrite or update the active tasks list inside it.
|
||||
- The plan must cover each backend section/module.
|
||||
- Define **exactly 3 new enhancements per section**, each with a unique number (e.g. `1.1`), a clear functional description, and status `[TODO]`.
|
||||
- Present the plan to the user in your final summary.
|
||||
|
||||
## B2. Trigger "n", "next", or "n{x}"
|
||||
|
||||
- Read `plans/next-enhancements.md` to check task status.
|
||||
- If all tasks are `[DONE]` (or none `[TODO]`), run **"e"/"enhance"** first.
|
||||
- Otherwise select the most impactful `[TODO]` task(s) by strategic value/impact — not just first-in-order. If `{x}` given, take the top `{x}` sequentially.
|
||||
|
||||
### B2a. Clarify before building ("Grill Me" step)
|
||||
|
||||
Same rule as root AGENTS.md §2a: if scope/acceptance criteria are genuinely ambiguous, ask one question at a time (`AskUserQuestion` in Claude Code) until unambiguous, and record the resolved criteria as a 1-3 line note next to the task entry before writing code. Skip when the task is already unambiguous.
|
||||
|
||||
- Implement the task(s) fully, applying the relevant role(s) from `SKILLS.md` (this dir).
|
||||
- On completion: flip status to `[DONE]` in `plans/next-enhancements.md`, document the feature in `docs/feature-list.md` (this dir) under the right section.
|
||||
- **Verify build integrity**: QA pass (golden path + edge cases + regression check on adjacent features — see `backend/CLAUDE.md`'s accuracy regression harness for OCR/parser changes specifically) and Hardware/Compatibility pass (cross-platform, GPU/VRAM footprint under Local/on-prem deployment — see Part A above).
|
||||
- State which task(s) were completed and the exact route/endpoint/menu path to see the new feature.
|
||||
|
||||
## B3. File Size & Refactoring Rules
|
||||
|
||||
Same 256-line threshold as root AGENTS.md §3, backend-wide. Applies to this file, `SKILLS.md`, and `CLAUDE.md` too — which is why Part A above was trimmed and linked out to `docs/vllm-service.md` rather than left inline.
|
||||
|
||||
## B4. Roles
|
||||
|
||||
See `SKILLS.md` (this dir) — same 5 roles as root (Architect, Backend, Frontend, QA, Hardware/Compatibility), applied to backend surfaces only (API routes, OCR pipeline, DB layer, Docker/deploy).
|
||||
|
||||
## B5. Mockup Data & Demo/Live Mode
|
||||
|
||||
Same as root AGENTS.md §5: mock data under `/data/mockup/`, a mock API layer mirroring the real backend contract, and a Demo/Live switcher. Not yet built for backend — see Adaptation Notes.
|
||||
|
||||
## B6. Cloud vs Local (On-Premise)
|
||||
|
||||
Same as root AGENTS.md §6, applied to backend service endpoints (Next.js gateway, pipeline API, vLLM server, Postgres) rather than the Flutter client's API base URL.
|
||||
|
||||
## B7. Ad-hoc Feature Requests
|
||||
|
||||
Direct feature requests not using "e"/"n": implement and document in `docs/feature-list.md` (this dir).
|
||||
|
||||
## Adaptation Notes (backend, split from root 2026-07-08)
|
||||
|
||||
- **Origin**: sections 5-8 of root `plans/next-enhancements.md` (Backend — Next.js API
|
||||
Gateway, Backend — OCR Pipeline & Accuracy, Backend — Postgres Data Layer, DevOps —
|
||||
Docker & Dev Tunnel) copied here as sections 1-4, statuses re-verified against the
|
||||
live code before the copy (not copied blind) — see task 7.1's `withTransaction`
|
||||
claim, task 5.1/5.2's dedup + timeout claims, and task 6.1's empty `models/` claim,
|
||||
all confirmed still accurate as of 2026-07-08. The root copy is frozen/archival
|
||||
(see root `AGENTS.md`'s "Scope: excludes `backend/`") rather than deleted, so this
|
||||
file — not the root one — is the single active source of truth going forward.
|
||||
- **Real commands**: `npm run dev`/`build`/`lint` in `pfm-web-app/`; accuracy
|
||||
regression harness `node pfm-web-app/scripts/accuracy-check.mts`; Python services
|
||||
via `./scripts/install.sh` + `./scripts/serve.sh` (this vLLM repo) and
|
||||
`./scripts/install-pipeline.sh` + `./scripts/serve-pipeline.sh` (pipeline API +
|
||||
classifier). Full stack: `docker compose up -d --build` **from the repo root**, not
|
||||
from inside `backend/` (see root `CLAUDE.md` — two `docker-compose.yml` files
|
||||
exist and running from here risks container-name conflicts).
|
||||
- **Pre-existing files over the 256-line threshold** (§B3 debt, not a blocker — split
|
||||
only if/when touched): `pfm-web-app/src/app/scan-pfm/page.tsx` (1169),
|
||||
`pfm-web-app/src/utils/parser.ts` (908), `pfm-web-app/src/app/page.tsx` (737),
|
||||
`config/classify_ocr_server.py` (691), `pfm-web-app/src/app/manual-label/page.tsx`
|
||||
(612), `pfm-web-app/src/app/api/parse/route.ts` (604), `pfm-web-app/src/db/init.ts`
|
||||
(477), `compare_sources_accuracy.py` (451), `pfm-web-app/src/utils/docker.ts` (362),
|
||||
`pfm-web-app/public/produk-pfm/train_classifier.py` (351), `compare_accuracy.py`
|
||||
(308), `pfm-web-app/src/app/api/arena/route.ts` (265). This file itself (`AGENTS.md`)
|
||||
was at 237 lines pre-kit and would have exceeded 256 once Part B was appended —
|
||||
hence the split into `docs/vllm-service.md`.
|
||||
- **No Demo/Live or Cloud/Local switch exists yet** (§B5, §B6) for the backend
|
||||
either. `docker-compose.override.yml` exposing `db`/`pipeline-api` directly to the
|
||||
host is a local-dev convenience, not a Cloud/Local deployment switch.
|
||||
- **Naming collision resolved by this split**: `AGENTS.md` already existed in this
|
||||
directory (vLLM service doc, committed 2026-06-30, unrelated to this kit) before
|
||||
Part B was appended — unlike root, where `AGENTS.md` didn't previously exist. Don't
|
||||
assume backend's `AGENTS.md` is kit-only when reading it from another tool; Part A
|
||||
is unrelated, pre-existing content kept for a reason.
|
||||
- **Pre-existing, unrelated governance files left as-is**: `.agents/AGENTS.md` at the
|
||||
*repo root* (different path, OCR post-processing rules) and root
|
||||
`plans/next-enhancement-plan.md` (singular, `[DONE]` QA checklist) — neither is
|
||||
part of this kit; see root `AGENTS.md`'s own Adaptation Notes.
|
||||
Reference in new issue
Block a user