Files
pfm-ocr/backend
fhanyuh e4dc24eaaa data: add DO photos, upload archive, and backend .env
Force-added: these paths are gitignored, so they stay ignored for new files
unless added the same way. Committed on request so the working data is not
lost during the migration off this machine.
2026-08-27 10:47:45 +07:00
..

PaddleOCR-VL-1.6 on vLLM

Local deployment of PaddleOCR-VL-1.6 using vLLM as the VLM inference backend. All Python workflows use uv.

Architecture

Gradio demo (7870)
       │
       ▼
Pipeline API (8090)  ── layout + preprocessing (PaddlePaddle GPU)
       │
       ▼
vLLM genai server (8118)  ── PaddleOCR-VL-1.6 VLM
Service Script Default URL
vLLM VLM server ./scripts/serve.sh http://127.0.0.1:8118/v1
Full pipeline API ./scripts/serve-pipeline.sh http://127.0.0.1:8090/layout-parsing
Online demo UI ./scripts/run-demo.sh http://127.0.0.1:7870

The vLLM server exposes only the VLM stage. For HTTP document parsing (layout + OCR), run the pipeline API, which calls vLLM via config/pipeline_config_vllm.yaml.

Prerequisites

  • Linux with NVIDIA GPU (CC ≥ 8.0 recommended; CUDA 12.6+ driver)
  • uv installed
  • ~16 GB GPU VRAM for default vLLM settings (tune in config/vllm_config.yaml)

Quick start

git clone <repo-url> ai-ocr-pfm-2026
cd ai-ocr-pfm-2026

cp .env.example .env   # adjust CUDA_VISIBLE_DEVICES if needed

# 1) Install vLLM server (.venv)
FLASH_ATTN_WHEEL="https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.3.14/flash_attn-2.8.2+cu128torch2.8-cp312-cp312-linux_x86_64.whl" \
  ./scripts/install.sh

# 2) Install pipeline API (.venv-api) — optional, needed for demo / full HTTP API
./scripts/install-pipeline.sh

Start services (three terminals, or background each):

./scripts/serve.sh            # vLLM on :8118
./scripts/serve-pipeline.sh     # pipeline on :8090
./scripts/run-demo.sh           # Gradio on :7870

Health checks:

curl -s http://127.0.0.1:8118/v1/models | jq .
curl -s http://127.0.0.1:8090/health
curl -s -o /dev/null -w "%{http_code}\n" http://127.0.0.1:7870/

Configuration

Copy .env.example to .env:

Variable Default Description
GENAI_HOST 0.0.0.0 vLLM bind address
GENAI_PORT 8118 vLLM port
GENAI_MODEL PaddleOCR-VL-1.6-0.9B Model name
VLLM_CONFIG config/vllm_config.yaml vLLM tuning
CUDA_VISIBLE_DEVICES 1 GPU for vLLM (use least-busy GPU)
PIPELINE_PORT 8090 Pipeline API port
PIPELINE_DEVICE gpu:0 GPU for layout/preprocessing
GRADIO_PORT 7870 Demo UI port

vLLM tuning (config/vllm_config.yaml):

gpu-memory-utilization: 0.75
max-num-seqs: 128

Client usage

Python (vLLM only)

from paddleocr import PaddleOCRVL

pipeline = PaddleOCRVL(
    vl_rec_backend="vllm-server",
    vl_rec_server_url="http://127.0.0.1:8118/v1",
)
output = pipeline.predict("path/to/image.png")

Run the client in a separate environment if it needs PaddlePaddle GPU alongside Transformers.

CLI

uv run paddleocr doc_parser \
  --input demo.png \
  --vl_rec_backend vllm-server \
  --vl_rec_server_url http://127.0.0.1:8118/v1

HTTP (full pipeline)

curl -X POST http://127.0.0.1:8090/layout-parsing \
  -H "Content-Type: application/json" \
  -d '{"file":"<base64>", "fileType": 1, "useLayoutDetection": true}'

Project layout

config/
  vllm_config.yaml           # vLLM backend tuning
  pipeline_config_vllm.yaml  # pipeline → vLLM server URL
scripts/
  install.sh                 # bootstrap .venv (vLLM)
  install-pipeline.sh        # bootstrap .venv-api (pipeline)
  serve.sh                   # start vLLM genai server
  serve-pipeline.sh          # start pipeline API
  run-demo.sh                # start Gradio demo
PaddleOCR-VL-1.6_Online_Demo/  # bundled Hugging Face-style demo
issues/                      # recorded problems and fixes
AGENTS.md                    # agent / contributor guide

Troubleshooting

See issues/ for detailed write-ups. Common fixes:

Symptom Fix
GPU OOM on vLLM startup Lower gpu-memory-utilization or set CUDA_VISIBLE_DEVICES to a free GPU
flash-attn build failure Use prebuilt wheel via FLASH_ATTN_WHEEL=... ./scripts/install.sh
Port 8080 in use Pipeline defaults to 8090; demo defaults to 7870

Agent conventions and issue-recording rules: AGENTS.md.

References