4.7 KiB
4.7 KiB
PaddleOCR-VL-1.6 on vLLM
Local deployment of PaddleOCR-VL-1.6 using vLLM as the VLM inference backend. All Python workflows use uv.
Architecture
Gradio demo (7870)
│
▼
Pipeline API (8090) ── layout + preprocessing (PaddlePaddle GPU)
│
▼
vLLM genai server (8118) ── PaddleOCR-VL-1.6 VLM
| Service | Script | Default URL |
|---|---|---|
| vLLM VLM server | ./scripts/serve.sh |
http://127.0.0.1:8118/v1 |
| Full pipeline API | ./scripts/serve-pipeline.sh |
http://127.0.0.1:8090/layout-parsing |
| Online demo UI | ./scripts/run-demo.sh |
http://127.0.0.1:7870 |
The vLLM server exposes only the VLM stage. For HTTP document parsing (layout + OCR), run the pipeline API, which calls vLLM via config/pipeline_config_vllm.yaml.
Prerequisites
- Linux with NVIDIA GPU (CC ≥ 8.0 recommended; CUDA 12.6+ driver)
- uv installed
- ~16 GB GPU VRAM for default vLLM settings (tune in
config/vllm_config.yaml)
Quick start
git clone <repo-url> ai-ocr-pfm-2026
cd ai-ocr-pfm-2026
cp .env.example .env # adjust CUDA_VISIBLE_DEVICES if needed
# 1) Install vLLM server (.venv)
FLASH_ATTN_WHEEL="https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.3.14/flash_attn-2.8.2+cu128torch2.8-cp312-cp312-linux_x86_64.whl" \
./scripts/install.sh
# 2) Install pipeline API (.venv-api) — optional, needed for demo / full HTTP API
./scripts/install-pipeline.sh
Start services (three terminals, or background each):
./scripts/serve.sh # vLLM on :8118
./scripts/serve-pipeline.sh # pipeline on :8090
./scripts/run-demo.sh # Gradio on :7870
Health checks:
curl -s http://127.0.0.1:8118/v1/models | jq .
curl -s http://127.0.0.1:8090/health
curl -s -o /dev/null -w "%{http_code}\n" http://127.0.0.1:7870/
Configuration
Copy .env.example to .env:
| Variable | Default | Description |
|---|---|---|
GENAI_HOST |
0.0.0.0 |
vLLM bind address |
GENAI_PORT |
8118 |
vLLM port |
GENAI_MODEL |
PaddleOCR-VL-1.6-0.9B |
Model name |
VLLM_CONFIG |
config/vllm_config.yaml |
vLLM tuning |
CUDA_VISIBLE_DEVICES |
1 |
GPU for vLLM (use least-busy GPU) |
PIPELINE_PORT |
8090 |
Pipeline API port |
PIPELINE_DEVICE |
gpu:0 |
GPU for layout/preprocessing |
GRADIO_PORT |
7870 |
Demo UI port |
vLLM tuning (config/vllm_config.yaml):
gpu-memory-utilization: 0.75
max-num-seqs: 128
Client usage
Python (vLLM only)
from paddleocr import PaddleOCRVL
pipeline = PaddleOCRVL(
vl_rec_backend="vllm-server",
vl_rec_server_url="http://127.0.0.1:8118/v1",
)
output = pipeline.predict("path/to/image.png")
Run the client in a separate environment if it needs PaddlePaddle GPU alongside Transformers.
CLI
uv run paddleocr doc_parser \
--input demo.png \
--vl_rec_backend vllm-server \
--vl_rec_server_url http://127.0.0.1:8118/v1
HTTP (full pipeline)
curl -X POST http://127.0.0.1:8090/layout-parsing \
-H "Content-Type: application/json" \
-d '{"file":"<base64>", "fileType": 1, "useLayoutDetection": true}'
Project layout
config/
vllm_config.yaml # vLLM backend tuning
pipeline_config_vllm.yaml # pipeline → vLLM server URL
scripts/
install.sh # bootstrap .venv (vLLM)
install-pipeline.sh # bootstrap .venv-api (pipeline)
serve.sh # start vLLM genai server
serve-pipeline.sh # start pipeline API
run-demo.sh # start Gradio demo
PaddleOCR-VL-1.6_Online_Demo/ # bundled Hugging Face-style demo
issues/ # recorded problems and fixes
AGENTS.md # agent / contributor guide
Troubleshooting
See issues/ for detailed write-ups. Common fixes:
| Symptom | Fix |
|---|---|
| GPU OOM on vLLM startup | Lower gpu-memory-utilization or set CUDA_VISIBLE_DEVICES to a free GPU |
| flash-attn build failure | Use prebuilt wheel via FLASH_ATTN_WHEEL=... ./scripts/install.sh |
| Port 8080 in use | Pipeline defaults to 8090; demo defaults to 7870 |
Agent conventions and issue-recording rules: AGENTS.md.