Adopt agents-settings kit, ship Product/SKU scan models, harden auth, verify OCR accuracy

Backend (app-pfm-ocr-v2/backend):
- Product/SKU scan feature complete: trained DINOv2 index (118 reference
  photos, 16 SKU classes) and YOLO classifier (83.3% top-1 val accuracy),
  fixed scripts/install-pipeline.sh (was missing ultralytics/torch), fully
  browser-verified end-to-end on /scan-pfm. Mobile m-scan-pfm page cancelled
  (Flutter app handles mobile; web UI is desktop-only for pipeline testing).
- Fixed a real data-loss bug: Save Ground Truth (scan-pfm and the DO-flow's
  manual-label) was silently writing into the pfm-web-app container's
  ephemeral filesystem instead of the host, because /sources wasn't
  bind-mounted in docker-compose.yml. Added the mount, recovered an
  orphaned entry.
- accounts.password is now bcrypt-hashed (bcryptjs, idempotent migration
  in db/init.ts) instead of plaintext; login route compares hashes.
- /api/v1/documents/* (list, PUT, upload) now enforces real 401 auth,
  matching what the Flutter client already sends. The "classic" routes
  deliberately stay open — they're dev-only web UI with no login flow and
  won't exist in production.
- OCR accuracy investigated end-to-end: real baseline is 95.10% overall
  (target met; accuracy_report.md was stale at 75.04%, now flagged). Fixed
  one genuine parser.ts bug (SO/DO field duplication in the global fallback
  regex); remaining gaps are OCR/layout-model limitations, not parser bugs.
- Adopted a standalone copy of the fhanyuh/agents-settings e/n workflow
  scoped to backend/ (AGENTS.md Part A/B split, SKILLS.md, plans/, docs/),
  independent of the root copy which now covers Flutter only.
- next-implementation.md deleted; content folded into
  backend/plans/next-enhancements.md for traceability.

Root:
- Adopted fhanyuh/agents-settings kit (AGENTS.md, SKILLS.md, plans/,
  docs/feature-list.md), scoped to the Flutter app only.
- Pending documents queue now persists to Hive (lib/core/storage) instead
  of memory-only, surviving an app kill mid-upload.

Removed backend_backup/ (stale Express/Prisma prototype, superseded by
pfm-web-app) and the completed plans/next-enhancement-plan.md checklist.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
Rafhan Mazaya FathurrahmanandClaude Sonnet 5 committed 2026-07-08 11:56:32 +07:00
1 parent 3df9f6ec5d
commit e60ab63154
129 files changed
+8520 -6684

No files matched your search

+115 -175
View File
@@ -1,6 +1,19 @@
# AGENTS: PaddleOCR-VL-1.6 vLLM Service
# AGENTS: PaddleOCR-VL-1.6 vLLM Service + Agents Settings Kit
This repository serves **PaddleOCR-VL-1.6** as a dedicated VLM inference backend using **vLLM**. All Python workflows use **uv** (never bare `pip` or system Python).
This is the authoritative rules file for any AI coding agent (Claude Code, Cursor,
GitHub Copilot, Aider, etc.) working inside `backend/`. Two unrelated concerns live
here side by side: **Part A** is this repo's original vLLM/PaddleOCR service doc.
**Part B** (appended 2026-07-08) is a **backend-scoped copy** of the
[fhanyuh/agents-settings](https://github.com/fhanyuh/agents-settings) `e`/`enhance`
and `n`/`next` workflow — see root `../AGENTS.md` for the same kit covering the
Flutter side of this repo. The two copies are independent: this one's
`plans/next-enhancements.md` and `docs/feature-list.md` only track backend work.
---
# Part A — vLLM Service (PaddleOCR-VL-1.6)
This repository serves **PaddleOCR-VL-1.6** as a dedicated VLM inference backend using **vLLM**. All Python workflows use **uv** (never bare `pip` or system Python). Full detail (client usage examples, tuning, troubleshooting, issue-file template) moved to [docs/vllm-service.md](docs/vllm-service.md) 2026-07-08 to keep this file under the Part B kit's 256-line threshold (§3) — this section keeps only the essentials.
## Architecture
@@ -10,90 +23,18 @@ Client (PaddleOCR pipeline) --> HTTP /v1 --> paddleocr genai_server (vLLM ba
This service exposes only the VLM stage. Clients connect with `vl_rec_backend="vllm-server"` and `vl_rec_server_url="http://<host>:8118/v1"`.
## Prerequisites
- Linux with NVIDIA GPU (CC >= 8.0 recommended; CUDA 12.6+ driver support)
- [uv](https://docs.astral.sh/uv/) installed (`uv --version`)
- ~16 GB GPU VRAM for default settings (tune via `config/vllm_config.yaml`)
## Quick start
```bash
cd <YOUR-WORKING-DIR>/ai-ocr-pfm-2026
# 1) Create Python 3.12 venv and install dependencies
./scripts/install.sh
# 2) Start the vLLM-backed genai server
./scripts/serve.sh
./scripts/install.sh # 1) Create Python 3.12 venv and install dependencies
./scripts/serve.sh # 2) Start the vLLM-backed genai server
```
Default endpoint: `http://0.0.0.0:8118/v1`
## uv conventions (always follow)
| Task | Command |
|------|---------|
| Create/sync env | `uv sync` |
| Run any Python | `uv run <command>` |
| Add a package | `uv add <package>` |
| Run server | `./scripts/serve.sh` or `uv run paddleocr genai_server ...` |
Never use `python -m pip`, `pip install`, or `python -m venv` directly in this repo.
Default endpoint: `http://0.0.0.0:8118/v1`. Never use `python -m pip`, `pip install`, or `python -m venv` directly in this repo — always `uv sync` / `uv run` / `uv add`.
## Issue recording (always follow)
**Every problem encountered** during install, serve, debug, or client integration must be written to `issues/` before moving on — even if it was resolved in the same session.
### Naming
```
issues/{NN}-{slug}.md
```
| Part | Rule | Example |
|------|------|---------|
| `{NN}` | Two-digit running number (`01`, `02`, …). Increment from the highest existing file. | `03` |
| `{slug}` | Lowercase kebab-case summary of the problem | `gpu-memory-startup-failure` |
Full example: `issues/04-gpu-memory-startup-failure.md`
### When to create a file
- Install or dependency errors (flash-attn, vLLM, uv conflicts)
- Server startup or runtime failures (OOM, port bind, model load)
- Client integration bugs or misconfiguration
- Workarounds that took non-obvious steps to discover
Do **not** rely on chat history or inline comments alone — if it blocked progress, it belongs in `issues/`.
### File template
```markdown
# Issue {NN}: {Short title}
## Problem
What failed, with exact error message or symptom.
## Context
Environment, command run, relevant config (`.env`, `config/vllm_config.yaml`).
## Solution
What fixed it, or current workaround / open status.
## References
Links, related issue files, or AGENTS.md sections.
```
### Index
Check `issues/` for the next number:
```bash
ls issues/*.md 2>/dev/null | sort
```
See [issues/](issues/) for recorded problems and fixes from this project.
**Every problem encountered** during install, serve, debug, or client integration must be written to `issues/{NN}-{slug}.md` before moving on — even if resolved in the same session. Naming/template details: [docs/vllm-service.md](docs/vllm-service.md#issue-recording--naming-and-template).
## Environment variables
@@ -108,97 +49,7 @@ Copy `.env.example` to `.env` and adjust as needed:
| `VLLM_CONFIG` | `config/vllm_config.yaml` | vLLM backend YAML config |
| `CUDA_VISIBLE_DEVICES` | `1` (see `.env.example`) | GPU index(es) to use |
On dual-GPU hosts, pick the GPU with more free VRAM. If startup fails with a memory error, lower `gpu-memory-utilization` in `config/vllm_config.yaml`.
## Client usage
After the server is running:
```bash
# CLI
uv run paddleocr doc_parser \
--input https://paddle-model-ecology.bj.bcebos.com/paddlex/imgs/demo_image/paddleocr_vl_demo.png \
--vl_rec_backend vllm-server \
--vl_rec_server_url http://localhost:8118/v1
```
```python
from paddleocr import PaddleOCRVL
pipeline = PaddleOCRVL(
vl_rec_backend="vllm-server",
vl_rec_server_url="http://127.0.0.1:8118/v1",
)
output = pipeline.predict("path/to/image.png")
```
Note: The full PaddleOCR-VL client should run in a **separate** environment if it needs PaddlePaddle GPU + Transformers. This repo is the isolated vLLM server only.
## Tuning vLLM
Edit `config/vllm_config.yaml`:
```yaml
gpu-memory-utilization: 0.8
max-num-seqs: 128
```
Reference: [PaddleOCR-VL vLLM parameter tuning](https://www.paddleocr.ai/latest/en/version3.x/pipeline_usage/PaddleOCR-VL.html#331-server-side-parameter-adjustment)
## Troubleshooting
See `issues/` for full write-ups. Quick pointers:
| Symptom | Issue file |
|---------|------------|
| `paddleocr install_genai_server_deps` / `No module named pip` | [01-genai-server-deps-pip-in-uv-venv.md](issues/01-genai-server-deps-pip-in-uv-venv.md) |
| flash-attn wheel incompatible with Python version | [02-flash-attn-wheel-python-version-mismatch.md](issues/02-flash-attn-wheel-python-version-mismatch.md) |
| `uv pip` targets wrong venv from another project | [03-active-virtual-env-from-other-project.md](issues/03-active-virtual-env-from-other-project.md) |
| Free memory below `gpu-memory-utilization` on startup | [04-gpu-memory-startup-failure.md](issues/04-gpu-memory-startup-failure.md) |
| `TokenizersBackend has no attribute all_special_tokens_extended` | [05-transformers-tokenizers-incompatibility.md](issues/05-transformers-tokenizers-incompatibility.md) |
| Extracted images not shown in Gradio demo (raw base64 in markdown) | [06-extracted-images-raw-base64-not-displayed.md](issues/06-extracted-images-raw-base64-not-displayed.md) |
### flash-attn build failures
Install the prebuilt wheel after `uv sync` (see `scripts/install.sh`):
```bash
FLASH_ATTN_WHEEL="https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.3.14/flash_attn-2.8.2+cu128torch2.8-cp312-cp312-linux_x86_64.whl" \
./scripts/install.sh
```
Pick the wheel matching your Python and CUDA versions from [flash-attention prebuild wheels](https://mjunya.com/flash-attention-prebuild-wheels/). Details: [02-flash-attn-wheel-python-version-mismatch.md](issues/02-flash-attn-wheel-python-version-mismatch.md).
Note: `paddleocr install_genai_server_deps` uses `pip` internally and is incompatible with uv-managed venvs. See [01-genai-server-deps-pip-in-uv-venv.md](issues/01-genai-server-deps-pip-in-uv-venv.md). This repo installs the vLLM stack via `uv sync` + `uv pip`.
### `TokenizersBackend has no attribute all_special_tokens_extended`
Pin transformers (already in `pyproject.toml`):
```bash
uv pip install "transformers==4.57.6"
```
See [05-transformers-tokenizers-incompatibility.md](issues/05-transformers-tokenizers-incompatibility.md).
### Do not install `paddlepaddle-gpu` in this venv
vLLM and PaddlePaddle GPU conflict. This server env uses `paddleocr[doc-parser]` without Paddle GPU.
### GPU memory on startup
If vLLM reports free memory below `gpu-memory-utilization`, either:
- Set `CUDA_VISIBLE_DEVICES` to a less-busy GPU
- Lower `gpu-memory-utilization` in `config/vllm_config.yaml` (e.g. `0.75` or `0.7`)
See [04-gpu-memory-startup-failure.md](issues/04-gpu-memory-startup-failure.md).
### Health check
```bash
curl -s http://localhost:8118/v1/models | jq .
```
On dual-GPU hosts, pick the GPU with more free VRAM. If startup fails with a memory error, lower `gpu-memory-utilization` in `config/vllm_config.yaml` — see [docs/vllm-service.md](docs/vllm-service.md#gpu-memory-on-startup).
## File map
@@ -210,15 +61,11 @@ curl -s http://localhost:8118/v1/models | jq .
| `scripts/serve.sh` | Start `paddleocr genai_server` |
| `config/vllm_config.yaml` | vLLM backend tuning |
| `.env.example` | Environment variable template |
## References
- [PaddleOCR-VL usage tutorial](https://www.paddleocr.ai/latest/en/version3.x/pipeline_usage/PaddleOCR-VL.html)
- [PaddleOCR genai_server FAQ](https://github.com/PaddlePaddle/PaddleOCR/discussions/16822)
| `docs/vllm-service.md` | Full vLLM reference (client usage, tuning, troubleshooting) |
## Coding Guidelines (always follow)
We use the [karpathy-guidelines](file:///home/user/LABS/OCR/paddle-ocr-vl-1-6-using-vllm-2026/andrej-karpathy-skills/skills/karpathy-guidelines/SKILL.md) skill to reduce common LLM coding mistakes. Refer to [SKILL.md](file:///home/user/LABS/OCR/paddle-ocr-vl-1-6-using-vllm-2026/andrej-karpathy-skills/skills/karpathy-guidelines/SKILL.md) for details:
We use the karpathy-guidelines skill to reduce common LLM coding mistakes:
1. **Think Before Coding**: Explicitly state assumptions and surface tradeoffs instead of making silent choices.
2. **Simplicity First**: Write the minimum amount of code to solve the problem with zero speculative configurations.
3. **Surgical Changes**: Edit only what is required and match the existing coding style exactly.
@@ -235,3 +82,96 @@ When the user intentionally asks to test the app:
- Take a screenshot for each sample image, each step, and each variant/option (if any), until the OCR result appears.
- Save the screenshots in the `/screenshots/` folder.
- Follow the file naming convention: `{2-digit-number}-{step#}-{variant_or_options_if_any}-{slug}.jpg` (e.g., `01-step1-default-upload.jpg`).
---
# Part B — Agents Settings Kit (backend-scoped `e`/`n` workflow)
Backend-scoped copy of the [fhanyuh/agents-settings](https://github.com/fhanyuh/agents-settings) kit, adopted 2026-07-08. Covers only `backend/` modules (Next.js API Gateway, OCR Pipeline & Accuracy, Postgres Data Layer, DevOps/Docker) — Flutter modules are tracked by the separate copy at root `../AGENTS.md`. `../CLAUDE.md` (root) and `CLAUDE.md` (this dir) each import their own copy.
## B0. Adopting Into an Existing Project
Already done for this repo (this split *is* that adoption, mirroring root's own §0 audit). Re-run "i"/"init" here to force a re-audit of `backend/` specifically (e.g. after a large refactor).
## B1. Trigger "e" or "enhance"
- Read `plans/next-enhancements.md` (this dir) to understand current backend structure, history, and active tasks.
- Overwrite or update the active tasks list inside it.
- The plan must cover each backend section/module.
- Define **exactly 3 new enhancements per section**, each with a unique number (e.g. `1.1`), a clear functional description, and status `[TODO]`.
- Present the plan to the user in your final summary.
## B2. Trigger "n", "next", or "n{x}"
- Read `plans/next-enhancements.md` to check task status.
- If all tasks are `[DONE]` (or none `[TODO]`), run **"e"/"enhance"** first.
- Otherwise select the most impactful `[TODO]` task(s) by strategic value/impact — not just first-in-order. If `{x}` given, take the top `{x}` sequentially.
### B2a. Clarify before building ("Grill Me" step)
Same rule as root AGENTS.md §2a: if scope/acceptance criteria are genuinely ambiguous, ask one question at a time (`AskUserQuestion` in Claude Code) until unambiguous, and record the resolved criteria as a 1-3 line note next to the task entry before writing code. Skip when the task is already unambiguous.
- Implement the task(s) fully, applying the relevant role(s) from `SKILLS.md` (this dir).
- On completion: flip status to `[DONE]` in `plans/next-enhancements.md`, document the feature in `docs/feature-list.md` (this dir) under the right section.
- **Verify build integrity**: QA pass (golden path + edge cases + regression check on adjacent features — see `backend/CLAUDE.md`'s accuracy regression harness for OCR/parser changes specifically) and Hardware/Compatibility pass (cross-platform, GPU/VRAM footprint under Local/on-prem deployment — see Part A above).
- State which task(s) were completed and the exact route/endpoint/menu path to see the new feature.
## B3. File Size & Refactoring Rules
Same 256-line threshold as root AGENTS.md §3, backend-wide. Applies to this file, `SKILLS.md`, and `CLAUDE.md` too — which is why Part A above was trimmed and linked out to `docs/vllm-service.md` rather than left inline.
## B4. Roles
See `SKILLS.md` (this dir) — same 5 roles as root (Architect, Backend, Frontend, QA, Hardware/Compatibility), applied to backend surfaces only (API routes, OCR pipeline, DB layer, Docker/deploy).
## B5. Mockup Data & Demo/Live Mode
Same as root AGENTS.md §5: mock data under `/data/mockup/`, a mock API layer mirroring the real backend contract, and a Demo/Live switcher. Not yet built for backend — see Adaptation Notes.
## B6. Cloud vs Local (On-Premise)
Same as root AGENTS.md §6, applied to backend service endpoints (Next.js gateway, pipeline API, vLLM server, Postgres) rather than the Flutter client's API base URL.
## B7. Ad-hoc Feature Requests
Direct feature requests not using "e"/"n": implement and document in `docs/feature-list.md` (this dir).
## Adaptation Notes (backend, split from root 2026-07-08)
- **Origin**: sections 5-8 of root `plans/next-enhancements.md` (Backend — Next.js API
Gateway, Backend — OCR Pipeline & Accuracy, Backend — Postgres Data Layer, DevOps —
Docker & Dev Tunnel) copied here as sections 1-4, statuses re-verified against the
live code before the copy (not copied blind) — see task 7.1's `withTransaction`
claim, task 5.1/5.2's dedup + timeout claims, and task 6.1's empty `models/` claim,
all confirmed still accurate as of 2026-07-08. The root copy is frozen/archival
(see root `AGENTS.md`'s "Scope: excludes `backend/`") rather than deleted, so this
file — not the root one — is the single active source of truth going forward.
- **Real commands**: `npm run dev`/`build`/`lint` in `pfm-web-app/`; accuracy
regression harness `node pfm-web-app/scripts/accuracy-check.mts`; Python services
via `./scripts/install.sh` + `./scripts/serve.sh` (this vLLM repo) and
`./scripts/install-pipeline.sh` + `./scripts/serve-pipeline.sh` (pipeline API +
classifier). Full stack: `docker compose up -d --build` **from the repo root**, not
from inside `backend/` (see root `CLAUDE.md` — two `docker-compose.yml` files
exist and running from here risks container-name conflicts).
- **Pre-existing files over the 256-line threshold** (§B3 debt, not a blocker — split
only if/when touched): `pfm-web-app/src/app/scan-pfm/page.tsx` (1169),
`pfm-web-app/src/utils/parser.ts` (908), `pfm-web-app/src/app/page.tsx` (737),
`config/classify_ocr_server.py` (691), `pfm-web-app/src/app/manual-label/page.tsx`
(612), `pfm-web-app/src/app/api/parse/route.ts` (604), `pfm-web-app/src/db/init.ts`
(477), `compare_sources_accuracy.py` (451), `pfm-web-app/src/utils/docker.ts` (362),
`pfm-web-app/public/produk-pfm/train_classifier.py` (351), `compare_accuracy.py`
(308), `pfm-web-app/src/app/api/arena/route.ts` (265). This file itself (`AGENTS.md`)
was at 237 lines pre-kit and would have exceeded 256 once Part B was appended —
hence the split into `docs/vllm-service.md`.
- **No Demo/Live or Cloud/Local switch exists yet** (§B5, §B6) for the backend
either. `docker-compose.override.yml` exposing `db`/`pipeline-api` directly to the
host is a local-dev convenience, not a Cloud/Local deployment switch.
- **Naming collision resolved by this split**: `AGENTS.md` already existed in this
directory (vLLM service doc, committed 2026-06-30, unrelated to this kit) before
Part B was appended — unlike root, where `AGENTS.md` didn't previously exist. Don't
assume backend's `AGENTS.md` is kit-only when reading it from another tool; Part A
is unrelated, pre-existing content kept for a reason.
- **Pre-existing, unrelated governance files left as-is**: `.agents/AGENTS.md` at the
*repo root* (different path, OCR post-processing rules) and root
`plans/next-enhancement-plan.md` (singular, `[DONE]` QA checklist) — neither is
part of this kit; see root `AGENTS.md`'s own Adaptation Notes.