Files
pfm-ocr/backend/AGENTS.md
T

9.0 KiB

AGENTS: PaddleOCR-VL-1.6 vLLM Service

This repository serves PaddleOCR-VL-1.6 as a dedicated VLM inference backend using vLLM. All Python workflows use uv (never bare pip or system Python).

Architecture

Client (PaddleOCR pipeline)  -->  HTTP /v1  -->  paddleocr genai_server (vLLM backend)

This service exposes only the VLM stage. Clients connect with vl_rec_backend="vllm-server" and vl_rec_server_url="http://<host>:8118/v1".

Prerequisites

  • Linux with NVIDIA GPU (CC >= 8.0 recommended; CUDA 12.6+ driver support)
  • uv installed (uv --version)
  • ~16 GB GPU VRAM for default settings (tune via config/vllm_config.yaml)

Quick start

cd <YOUR-WORKING-DIR>/ai-ocr-pfm-2026

# 1) Create Python 3.12 venv and install dependencies
./scripts/install.sh

# 2) Start the vLLM-backed genai server
./scripts/serve.sh

Default endpoint: http://0.0.0.0:8118/v1

uv conventions (always follow)

Task Command
Create/sync env uv sync
Run any Python uv run <command>
Add a package uv add <package>
Run server ./scripts/serve.sh or uv run paddleocr genai_server ...

Never use python -m pip, pip install, or python -m venv directly in this repo.

Issue recording (always follow)

Every problem encountered during install, serve, debug, or client integration must be written to issues/ before moving on — even if it was resolved in the same session.

Naming

issues/{NN}-{slug}.md
Part Rule Example
{NN} Two-digit running number (01, 02, …). Increment from the highest existing file. 03
{slug} Lowercase kebab-case summary of the problem gpu-memory-startup-failure

Full example: issues/04-gpu-memory-startup-failure.md

When to create a file

  • Install or dependency errors (flash-attn, vLLM, uv conflicts)
  • Server startup or runtime failures (OOM, port bind, model load)
  • Client integration bugs or misconfiguration
  • Workarounds that took non-obvious steps to discover

Do not rely on chat history or inline comments alone — if it blocked progress, it belongs in issues/.

File template

# Issue {NN}: {Short title}

## Problem
What failed, with exact error message or symptom.

## Context
Environment, command run, relevant config (`.env`, `config/vllm_config.yaml`).

## Solution
What fixed it, or current workaround / open status.

## References
Links, related issue files, or AGENTS.md sections.

Index

Check issues/ for the next number:

ls issues/*.md 2>/dev/null | sort

See issues/ for recorded problems and fixes from this project.

Environment variables

Copy .env.example to .env and adjust as needed:

Variable Default Description
GENAI_HOST 0.0.0.0 Bind address
GENAI_PORT 8118 Service port
GENAI_MODEL PaddleOCR-VL-1.6-0.9B Model name for genai_server
GENAI_BACKEND vllm Inference backend
VLLM_CONFIG config/vllm_config.yaml vLLM backend YAML config
CUDA_VISIBLE_DEVICES 1 (see .env.example) GPU index(es) to use

On dual-GPU hosts, pick the GPU with more free VRAM. If startup fails with a memory error, lower gpu-memory-utilization in config/vllm_config.yaml.

Client usage

After the server is running:

# CLI
uv run paddleocr doc_parser \
  --input https://paddle-model-ecology.bj.bcebos.com/paddlex/imgs/demo_image/paddleocr_vl_demo.png \
  --vl_rec_backend vllm-server \
  --vl_rec_server_url http://localhost:8118/v1
from paddleocr import PaddleOCRVL

pipeline = PaddleOCRVL(
    vl_rec_backend="vllm-server",
    vl_rec_server_url="http://127.0.0.1:8118/v1",
)
output = pipeline.predict("path/to/image.png")

Note: The full PaddleOCR-VL client should run in a separate environment if it needs PaddlePaddle GPU + Transformers. This repo is the isolated vLLM server only.

Tuning vLLM

Edit config/vllm_config.yaml:

gpu-memory-utilization: 0.8
max-num-seqs: 128

Reference: PaddleOCR-VL vLLM parameter tuning

Troubleshooting

See issues/ for full write-ups. Quick pointers:

Symptom Issue file
paddleocr install_genai_server_deps / No module named pip 01-genai-server-deps-pip-in-uv-venv.md
flash-attn wheel incompatible with Python version 02-flash-attn-wheel-python-version-mismatch.md
uv pip targets wrong venv from another project 03-active-virtual-env-from-other-project.md
Free memory below gpu-memory-utilization on startup 04-gpu-memory-startup-failure.md
TokenizersBackend has no attribute all_special_tokens_extended 05-transformers-tokenizers-incompatibility.md
Extracted images not shown in Gradio demo (raw base64 in markdown) 06-extracted-images-raw-base64-not-displayed.md

flash-attn build failures

Install the prebuilt wheel after uv sync (see scripts/install.sh):

FLASH_ATTN_WHEEL="https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.3.14/flash_attn-2.8.2+cu128torch2.8-cp312-cp312-linux_x86_64.whl" \
  ./scripts/install.sh

Pick the wheel matching your Python and CUDA versions from flash-attention prebuild wheels. Details: 02-flash-attn-wheel-python-version-mismatch.md.

Note: paddleocr install_genai_server_deps uses pip internally and is incompatible with uv-managed venvs. See 01-genai-server-deps-pip-in-uv-venv.md. This repo installs the vLLM stack via uv sync + uv pip.

TokenizersBackend has no attribute all_special_tokens_extended

Pin transformers (already in pyproject.toml):

uv pip install "transformers==4.57.6"

See 05-transformers-tokenizers-incompatibility.md.

Do not install paddlepaddle-gpu in this venv

vLLM and PaddlePaddle GPU conflict. This server env uses paddleocr[doc-parser] without Paddle GPU.

GPU memory on startup

If vLLM reports free memory below gpu-memory-utilization, either:

  • Set CUDA_VISIBLE_DEVICES to a less-busy GPU
  • Lower gpu-memory-utilization in config/vllm_config.yaml (e.g. 0.75 or 0.7)

See 04-gpu-memory-startup-failure.md.

Health check

curl -s http://localhost:8118/v1/models | jq .

File map

Path Purpose
issues/ Recorded problems and fixes ({NN}-{slug}.md)
pyproject.toml uv project metadata and base dependencies
scripts/install.sh Bootstrap venv + vLLM server deps
scripts/serve.sh Start paddleocr genai_server
config/vllm_config.yaml vLLM backend tuning
.env.example Environment variable template

References

Coding Guidelines (always follow)

We use the karpathy-guidelines skill to reduce common LLM coding mistakes. Refer to SKILL.md for details:

  1. Think Before Coding: Explicitly state assumptions and surface tradeoffs instead of making silent choices.
  2. Simplicity First: Write the minimum amount of code to solve the problem with zero speculative configurations.
  3. Surgical Changes: Edit only what is required and match the existing coding style exactly.
  4. Goal-Driven Execution: Define verifiable success criteria and run automated tests/screenshots to confirm correctness.

Path Guidelines (always follow)

Never use full paths containing the user's logged-in name (e.g., /home/{uid}/path). Always use relative paths instead (e.g., . or ./path relative to the workspace root).

App Testing Guidelines (always follow)

When the user intentionally asks to test the app:

  • Use browser tools to test the app.
  • Take a screenshot for each sample image, each step, and each variant/option (if any), until the OCR result appears.
  • Save the screenshots in the /screenshots/ folder.
  • Follow the file naming convention: {2-digit-number}-{step#}-{variant_or_options_if_any}-{slug}.jpg (e.g., 01-step1-default-upload.jpg).