Files
..

PaddleOCR-VL-1.6 on vLLM

Local deployment of PaddleOCR-VL-1.6 using vLLM as the VLM inference backend. All Python workflows use uv.

Architecture

Gradio demo (7870)
       │
       ▼
Pipeline API (8090)  ── layout + preprocessing (PaddlePaddle GPU)
       │
       ▼
vLLM genai server (8118)  ── PaddleOCR-VL-1.6 VLM
Service Script Default URL
vLLM VLM server ./scripts/serve.sh http://127.0.0.1:8118/v1
Full pipeline API ./scripts/serve-pipeline.sh http://127.0.0.1:8090/layout-parsing
Online demo UI ./scripts/run-demo.sh http://127.0.0.1:7870

The vLLM server exposes only the VLM stage. For HTTP document parsing (layout + OCR), run the pipeline API, which calls vLLM via config/pipeline_config_vllm.yaml.

Prerequisites

  • Linux with NVIDIA GPU (CC ≥ 8.0 recommended; CUDA 12.6+ driver)
  • uv installed
  • ~16 GB GPU VRAM for default vLLM settings (tune in config/vllm_config.yaml)

Quick start

git clone <repo-url> ai-ocr-pfm-2026
cd ai-ocr-pfm-2026

cp .env.example .env   # adjust CUDA_VISIBLE_DEVICES if needed

# 1) Install vLLM server (.venv)
FLASH_ATTN_WHEEL="https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.3.14/flash_attn-2.8.2+cu128torch2.8-cp312-cp312-linux_x86_64.whl" \
  ./scripts/install.sh

# 2) Install pipeline API (.venv-api) — optional, needed for demo / full HTTP API
./scripts/install-pipeline.sh

Start services (three terminals, or background each):

./scripts/serve.sh            # vLLM on :8118
./scripts/serve-pipeline.sh     # pipeline on :8090
./scripts/run-demo.sh           # Gradio on :7870

Health checks:

curl -s http://127.0.0.1:8118/v1/models | jq .
curl -s http://127.0.0.1:8090/health
curl -s -o /dev/null -w "%{http_code}\n" http://127.0.0.1:7870/

Configuration

Copy .env.example to .env:

Variable Default Description
GENAI_HOST 0.0.0.0 vLLM bind address
GENAI_PORT 8118 vLLM port
GENAI_MODEL PaddleOCR-VL-1.6-0.9B Model name
VLLM_CONFIG config/vllm_config.yaml vLLM tuning
CUDA_VISIBLE_DEVICES 1 GPU for vLLM (use least-busy GPU)
PIPELINE_PORT 8090 Pipeline API port
PIPELINE_DEVICE gpu:0 GPU for layout/preprocessing
GRADIO_PORT 7870 Demo UI port

vLLM tuning (config/vllm_config.yaml):

gpu-memory-utilization: 0.75
max-num-seqs: 128

Client usage

Python (vLLM only)

from paddleocr import PaddleOCRVL

pipeline = PaddleOCRVL(
    vl_rec_backend="vllm-server",
    vl_rec_server_url="http://127.0.0.1:8118/v1",
)
output = pipeline.predict("path/to/image.png")

Run the client in a separate environment if it needs PaddlePaddle GPU alongside Transformers.

CLI

uv run paddleocr doc_parser \
  --input demo.png \
  --vl_rec_backend vllm-server \
  --vl_rec_server_url http://127.0.0.1:8118/v1

HTTP (full pipeline)

curl -X POST http://127.0.0.1:8090/layout-parsing \
  -H "Content-Type: application/json" \
  -d '{"file":"<base64>", "fileType": 1, "useLayoutDetection": true}'

Project layout

config/
  vllm_config.yaml           # vLLM backend tuning
  pipeline_config_vllm.yaml  # pipeline → vLLM server URL
scripts/
  install.sh                 # bootstrap .venv (vLLM)
  install-pipeline.sh        # bootstrap .venv-api (pipeline)
  serve.sh                   # start vLLM genai server
  serve-pipeline.sh          # start pipeline API
  run-demo.sh                # start Gradio demo
PaddleOCR-VL-1.6_Online_Demo/  # bundled Hugging Face-style demo
issues/                      # recorded problems and fixes
AGENTS.md                    # agent / contributor guide

Troubleshooting

See issues/ for detailed write-ups. Common fixes:

Symptom Fix
GPU OOM on vLLM startup Lower gpu-memory-utilization or set CUDA_VISIBLE_DEVICES to a free GPU
flash-attn build failure Use prebuilt wheel via FLASH_ATTN_WHEEL=... ./scripts/install.sh
Port 8080 in use Pipeline defaults to 8090; demo defaults to 7870

Agent conventions and issue-recording rules: AGENTS.md.

References