Force-added: these paths are gitignored, so they stay ignored for new files unless added the same way. Committed on request so the working data is not lost during the migration off this machine.
PaddleOCR-VL-1.6 on vLLM
Local deployment of PaddleOCR-VL-1.6 using vLLM as the VLM inference backend. All Python workflows use uv.
Architecture
Gradio demo (7870)
│
▼
Pipeline API (8090) ── layout + preprocessing (PaddlePaddle GPU)
│
▼
vLLM genai server (8118) ── PaddleOCR-VL-1.6 VLM
| Service | Script | Default URL |
|---|---|---|
| vLLM VLM server | ./scripts/serve.sh |
http://127.0.0.1:8118/v1 |
| Full pipeline API | ./scripts/serve-pipeline.sh |
http://127.0.0.1:8090/layout-parsing |
| Online demo UI | ./scripts/run-demo.sh |
http://127.0.0.1:7870 |
The vLLM server exposes only the VLM stage. For HTTP document parsing (layout + OCR), run the pipeline API, which calls vLLM via config/pipeline_config_vllm.yaml.
Prerequisites
- Linux with NVIDIA GPU (CC ≥ 8.0 recommended; CUDA 12.6+ driver)
- uv installed
- ~16 GB GPU VRAM for default vLLM settings (tune in
config/vllm_config.yaml)
Quick start
git clone <repo-url> ai-ocr-pfm-2026
cd ai-ocr-pfm-2026
cp .env.example .env # adjust CUDA_VISIBLE_DEVICES if needed
# 1) Install vLLM server (.venv)
FLASH_ATTN_WHEEL="https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.3.14/flash_attn-2.8.2+cu128torch2.8-cp312-cp312-linux_x86_64.whl" \
./scripts/install.sh
# 2) Install pipeline API (.venv-api) — optional, needed for demo / full HTTP API
./scripts/install-pipeline.sh
Start services (three terminals, or background each):
./scripts/serve.sh # vLLM on :8118
./scripts/serve-pipeline.sh # pipeline on :8090
./scripts/run-demo.sh # Gradio on :7870
Health checks:
curl -s http://127.0.0.1:8118/v1/models | jq .
curl -s http://127.0.0.1:8090/health
curl -s -o /dev/null -w "%{http_code}\n" http://127.0.0.1:7870/
Configuration
Copy .env.example to .env:
| Variable | Default | Description |
|---|---|---|
GENAI_HOST |
0.0.0.0 |
vLLM bind address |
GENAI_PORT |
8118 |
vLLM port |
GENAI_MODEL |
PaddleOCR-VL-1.6-0.9B |
Model name |
VLLM_CONFIG |
config/vllm_config.yaml |
vLLM tuning |
CUDA_VISIBLE_DEVICES |
1 |
GPU for vLLM (use least-busy GPU) |
PIPELINE_PORT |
8090 |
Pipeline API port |
PIPELINE_DEVICE |
gpu:0 |
GPU for layout/preprocessing |
GRADIO_PORT |
7870 |
Demo UI port |
vLLM tuning (config/vllm_config.yaml):
gpu-memory-utilization: 0.75
max-num-seqs: 128
Client usage
Python (vLLM only)
from paddleocr import PaddleOCRVL
pipeline = PaddleOCRVL(
vl_rec_backend="vllm-server",
vl_rec_server_url="http://127.0.0.1:8118/v1",
)
output = pipeline.predict("path/to/image.png")
Run the client in a separate environment if it needs PaddlePaddle GPU alongside Transformers.
CLI
uv run paddleocr doc_parser \
--input demo.png \
--vl_rec_backend vllm-server \
--vl_rec_server_url http://127.0.0.1:8118/v1
HTTP (full pipeline)
curl -X POST http://127.0.0.1:8090/layout-parsing \
-H "Content-Type: application/json" \
-d '{"file":"<base64>", "fileType": 1, "useLayoutDetection": true}'
Project layout
config/
vllm_config.yaml # vLLM backend tuning
pipeline_config_vllm.yaml # pipeline → vLLM server URL
scripts/
install.sh # bootstrap .venv (vLLM)
install-pipeline.sh # bootstrap .venv-api (pipeline)
serve.sh # start vLLM genai server
serve-pipeline.sh # start pipeline API
run-demo.sh # start Gradio demo
PaddleOCR-VL-1.6_Online_Demo/ # bundled Hugging Face-style demo
issues/ # recorded problems and fixes
AGENTS.md # agent / contributor guide
Troubleshooting
See issues/ for detailed write-ups. Common fixes:
| Symptom | Fix |
|---|---|
| GPU OOM on vLLM startup | Lower gpu-memory-utilization or set CUDA_VISIBLE_DEVICES to a free GPU |
| flash-attn build failure | Use prebuilt wheel via FLASH_ATTN_WHEEL=... ./scripts/install.sh |
| Port 8080 in use | Pipeline defaults to 8090; demo defaults to 7870 |
Agent conventions and issue-recording rules: AGENTS.md.