9.0 KiB
AGENTS: PaddleOCR-VL-1.6 vLLM Service
This repository serves PaddleOCR-VL-1.6 as a dedicated VLM inference backend using vLLM. All Python workflows use uv (never bare pip or system Python).
Architecture
Client (PaddleOCR pipeline) --> HTTP /v1 --> paddleocr genai_server (vLLM backend)
This service exposes only the VLM stage. Clients connect with vl_rec_backend="vllm-server" and vl_rec_server_url="http://<host>:8118/v1".
Prerequisites
- Linux with NVIDIA GPU (CC >= 8.0 recommended; CUDA 12.6+ driver support)
- uv installed (
uv --version) - ~16 GB GPU VRAM for default settings (tune via
config/vllm_config.yaml)
Quick start
cd <YOUR-WORKING-DIR>/ai-ocr-pfm-2026
# 1) Create Python 3.12 venv and install dependencies
./scripts/install.sh
# 2) Start the vLLM-backed genai server
./scripts/serve.sh
Default endpoint: http://0.0.0.0:8118/v1
uv conventions (always follow)
| Task | Command |
|---|---|
| Create/sync env | uv sync |
| Run any Python | uv run <command> |
| Add a package | uv add <package> |
| Run server | ./scripts/serve.sh or uv run paddleocr genai_server ... |
Never use python -m pip, pip install, or python -m venv directly in this repo.
Issue recording (always follow)
Every problem encountered during install, serve, debug, or client integration must be written to issues/ before moving on — even if it was resolved in the same session.
Naming
issues/{NN}-{slug}.md
| Part | Rule | Example |
|---|---|---|
{NN} |
Two-digit running number (01, 02, …). Increment from the highest existing file. |
03 |
{slug} |
Lowercase kebab-case summary of the problem | gpu-memory-startup-failure |
Full example: issues/04-gpu-memory-startup-failure.md
When to create a file
- Install or dependency errors (flash-attn, vLLM, uv conflicts)
- Server startup or runtime failures (OOM, port bind, model load)
- Client integration bugs or misconfiguration
- Workarounds that took non-obvious steps to discover
Do not rely on chat history or inline comments alone — if it blocked progress, it belongs in issues/.
File template
# Issue {NN}: {Short title}
## Problem
What failed, with exact error message or symptom.
## Context
Environment, command run, relevant config (`.env`, `config/vllm_config.yaml`).
## Solution
What fixed it, or current workaround / open status.
## References
Links, related issue files, or AGENTS.md sections.
Index
Check issues/ for the next number:
ls issues/*.md 2>/dev/null | sort
See issues/ for recorded problems and fixes from this project.
Environment variables
Copy .env.example to .env and adjust as needed:
| Variable | Default | Description |
|---|---|---|
GENAI_HOST |
0.0.0.0 |
Bind address |
GENAI_PORT |
8118 |
Service port |
GENAI_MODEL |
PaddleOCR-VL-1.6-0.9B |
Model name for genai_server |
GENAI_BACKEND |
vllm |
Inference backend |
VLLM_CONFIG |
config/vllm_config.yaml |
vLLM backend YAML config |
CUDA_VISIBLE_DEVICES |
1 (see .env.example) |
GPU index(es) to use |
On dual-GPU hosts, pick the GPU with more free VRAM. If startup fails with a memory error, lower gpu-memory-utilization in config/vllm_config.yaml.
Client usage
After the server is running:
# CLI
uv run paddleocr doc_parser \
--input https://paddle-model-ecology.bj.bcebos.com/paddlex/imgs/demo_image/paddleocr_vl_demo.png \
--vl_rec_backend vllm-server \
--vl_rec_server_url http://localhost:8118/v1
from paddleocr import PaddleOCRVL
pipeline = PaddleOCRVL(
vl_rec_backend="vllm-server",
vl_rec_server_url="http://127.0.0.1:8118/v1",
)
output = pipeline.predict("path/to/image.png")
Note: The full PaddleOCR-VL client should run in a separate environment if it needs PaddlePaddle GPU + Transformers. This repo is the isolated vLLM server only.
Tuning vLLM
Edit config/vllm_config.yaml:
gpu-memory-utilization: 0.8
max-num-seqs: 128
Reference: PaddleOCR-VL vLLM parameter tuning
Troubleshooting
See issues/ for full write-ups. Quick pointers:
| Symptom | Issue file |
|---|---|
paddleocr install_genai_server_deps / No module named pip |
01-genai-server-deps-pip-in-uv-venv.md |
| flash-attn wheel incompatible with Python version | 02-flash-attn-wheel-python-version-mismatch.md |
uv pip targets wrong venv from another project |
03-active-virtual-env-from-other-project.md |
Free memory below gpu-memory-utilization on startup |
04-gpu-memory-startup-failure.md |
TokenizersBackend has no attribute all_special_tokens_extended |
05-transformers-tokenizers-incompatibility.md |
| Extracted images not shown in Gradio demo (raw base64 in markdown) | 06-extracted-images-raw-base64-not-displayed.md |
flash-attn build failures
Install the prebuilt wheel after uv sync (see scripts/install.sh):
FLASH_ATTN_WHEEL="https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.3.14/flash_attn-2.8.2+cu128torch2.8-cp312-cp312-linux_x86_64.whl" \
./scripts/install.sh
Pick the wheel matching your Python and CUDA versions from flash-attention prebuild wheels. Details: 02-flash-attn-wheel-python-version-mismatch.md.
Note: paddleocr install_genai_server_deps uses pip internally and is incompatible with uv-managed venvs. See 01-genai-server-deps-pip-in-uv-venv.md. This repo installs the vLLM stack via uv sync + uv pip.
TokenizersBackend has no attribute all_special_tokens_extended
Pin transformers (already in pyproject.toml):
uv pip install "transformers==4.57.6"
See 05-transformers-tokenizers-incompatibility.md.
Do not install paddlepaddle-gpu in this venv
vLLM and PaddlePaddle GPU conflict. This server env uses paddleocr[doc-parser] without Paddle GPU.
GPU memory on startup
If vLLM reports free memory below gpu-memory-utilization, either:
- Set
CUDA_VISIBLE_DEVICESto a less-busy GPU - Lower
gpu-memory-utilizationinconfig/vllm_config.yaml(e.g.0.75or0.7)
See 04-gpu-memory-startup-failure.md.
Health check
curl -s http://localhost:8118/v1/models | jq .
File map
| Path | Purpose |
|---|---|
issues/ |
Recorded problems and fixes ({NN}-{slug}.md) |
pyproject.toml |
uv project metadata and base dependencies |
scripts/install.sh |
Bootstrap venv + vLLM server deps |
scripts/serve.sh |
Start paddleocr genai_server |
config/vllm_config.yaml |
vLLM backend tuning |
.env.example |
Environment variable template |
References
Coding Guidelines (always follow)
We use the karpathy-guidelines skill to reduce common LLM coding mistakes. Refer to SKILL.md for details:
- Think Before Coding: Explicitly state assumptions and surface tradeoffs instead of making silent choices.
- Simplicity First: Write the minimum amount of code to solve the problem with zero speculative configurations.
- Surgical Changes: Edit only what is required and match the existing coding style exactly.
- Goal-Driven Execution: Define verifiable success criteria and run automated tests/screenshots to confirm correctness.
Path Guidelines (always follow)
Never use full paths containing the user's logged-in name (e.g., /home/{uid}/path). Always use relative paths instead (e.g., . or ./path relative to the workspace root).
App Testing Guidelines (always follow)
When the user intentionally asks to test the app:
- Use browser tools to test the app.
- Take a screenshot for each sample image, each step, and each variant/option (if any), until the OCR result appears.
- Save the screenshots in the
/screenshots/folder. - Follow the file naming convention:
{2-digit-number}-{step#}-{variant_or_options_if_any}-{slug}.jpg(e.g.,01-step1-default-upload.jpg).