Files
pfm-ocr/backend/docs/vllm-service.md
T
Rafhan Mazaya FathurrahmanandClaude Sonnet 5 e60ab63154 Adopt agents-settings kit, ship Product/SKU scan models, harden auth, verify OCR accuracy
Backend (app-pfm-ocr-v2/backend):
- Product/SKU scan feature complete: trained DINOv2 index (118 reference
  photos, 16 SKU classes) and YOLO classifier (83.3% top-1 val accuracy),
  fixed scripts/install-pipeline.sh (was missing ultralytics/torch), fully
  browser-verified end-to-end on /scan-pfm. Mobile m-scan-pfm page cancelled
  (Flutter app handles mobile; web UI is desktop-only for pipeline testing).
- Fixed a real data-loss bug: Save Ground Truth (scan-pfm and the DO-flow's
  manual-label) was silently writing into the pfm-web-app container's
  ephemeral filesystem instead of the host, because /sources wasn't
  bind-mounted in docker-compose.yml. Added the mount, recovered an
  orphaned entry.
- accounts.password is now bcrypt-hashed (bcryptjs, idempotent migration
  in db/init.ts) instead of plaintext; login route compares hashes.
- /api/v1/documents/* (list, PUT, upload) now enforces real 401 auth,
  matching what the Flutter client already sends. The "classic" routes
  deliberately stay open — they're dev-only web UI with no login flow and
  won't exist in production.
- OCR accuracy investigated end-to-end: real baseline is 95.10% overall
  (target met; accuracy_report.md was stale at 75.04%, now flagged). Fixed
  one genuine parser.ts bug (SO/DO field duplication in the global fallback
  regex); remaining gaps are OCR/layout-model limitations, not parser bugs.
- Adopted a standalone copy of the fhanyuh/agents-settings e/n workflow
  scoped to backend/ (AGENTS.md Part A/B split, SKILLS.md, plans/, docs/),
  independent of the root copy which now covers Flutter only.
- next-implementation.md deleted; content folded into
  backend/plans/next-enhancements.md for traceability.

Root:
- Adopted fhanyuh/agents-settings kit (AGENTS.md, SKILLS.md, plans/,
  docs/feature-list.md), scoped to the Flutter app only.
- Pending documents queue now persists to Hive (lib/core/storage) instead
  of memory-only, surviving an app kill mid-upload.

Removed backend_backup/ (stale Express/Prisma prototype, superseded by
pfm-web-app) and the completed plans/next-enhancement-plan.md checklist.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-08 11:56:32 +07:00

4.9 KiB

vLLM Service — Full Reference

Detail split out of ../AGENTS.md (2026-07-08, to keep that file under the Agents Settings Kit's 256-line threshold once the e/n workflow was appended to it). AGENTS.md keeps the short version — architecture, quick start, the env var table, file map — and links here for everything else.

Issue recording — naming and template

issues/{NN}-{slug}.md
Part Rule Example
{NN} Two-digit running number (01, 02, …). Increment from the highest existing file. 03
{slug} Lowercase kebab-case summary of the problem gpu-memory-startup-failure

Full example: issues/04-gpu-memory-startup-failure.md

File template

# Issue {NN}: {Short title}

## Problem
What failed, with exact error message or symptom.

## Context
Environment, command run, relevant config (`.env`, `config/vllm_config.yaml`).

## Solution
What fixed it, or current workaround / open status.

## References
Links, related issue files, or AGENTS.md sections.

Check issues/ for the next number:

ls issues/*.md 2>/dev/null | sort

Client usage

After the server is running:

# CLI
uv run paddleocr doc_parser \
  --input https://paddle-model-ecology.bj.bcebos.com/paddlex/imgs/demo_image/paddleocr_vl_demo.png \
  --vl_rec_backend vllm-server \
  --vl_rec_server_url http://localhost:8118/v1
from paddleocr import PaddleOCRVL

pipeline = PaddleOCRVL(
    vl_rec_backend="vllm-server",
    vl_rec_server_url="http://127.0.0.1:8118/v1",
)
output = pipeline.predict("path/to/image.png")

Note: The full PaddleOCR-VL client should run in a separate environment if it needs PaddlePaddle GPU + Transformers. This repo is the isolated vLLM server only.

Tuning vLLM

Edit config/vllm_config.yaml:

gpu-memory-utilization: 0.8
max-num-seqs: 128

Reference: PaddleOCR-VL vLLM parameter tuning

Troubleshooting

See issues/ for full write-ups. Quick pointers:

Symptom Issue file
paddleocr install_genai_server_deps / No module named pip 01-genai-server-deps-pip-in-uv-venv.md
flash-attn wheel incompatible with Python version 02-flash-attn-wheel-python-version-mismatch.md
uv pip targets wrong venv from another project 03-active-virtual-env-from-other-project.md
Free memory below gpu-memory-utilization on startup 04-gpu-memory-startup-failure.md
TokenizersBackend has no attribute all_special_tokens_extended 05-transformers-tokenizers-incompatibility.md
Extracted images not shown in Gradio demo (raw base64 in markdown) 06-extracted-images-raw-base64-not-displayed.md

flash-attn build failures

Install the prebuilt wheel after uv sync (see scripts/install.sh):

FLASH_ATTN_WHEEL="https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.3.14/flash_attn-2.8.2+cu128torch2.8-cp312-cp312-linux_x86_64.whl" \
  ./scripts/install.sh

Pick the wheel matching your Python and CUDA versions from flash-attention prebuild wheels. Details: 02-flash-attn-wheel-python-version-mismatch.md.

Note: paddleocr install_genai_server_deps uses pip internally and is incompatible with uv-managed venvs. See 01-genai-server-deps-pip-in-uv-venv.md. This repo installs the vLLM stack via uv sync + uv pip.

TokenizersBackend has no attribute all_special_tokens_extended

Pin transformers (already in pyproject.toml):

uv pip install "transformers==4.57.6"

See 05-transformers-tokenizers-incompatibility.md.

Do not install paddlepaddle-gpu in this venv

vLLM and PaddlePaddle GPU conflict. This server env uses paddleocr[doc-parser] without Paddle GPU.

GPU memory on startup

If vLLM reports free memory below gpu-memory-utilization, either:

  • Set CUDA_VISIBLE_DEVICES to a less-busy GPU
  • Lower gpu-memory-utilization in config/vllm_config.yaml (e.g. 0.75 or 0.7)

See 04-gpu-memory-startup-failure.md.

Health check

curl -s http://localhost:8118/v1/models | jq .

References