# PaddleOCR-VL-1.6 on vLLM Local deployment of [PaddleOCR-VL-1.6](https://www.paddleocr.ai/latest/en/version3.x/pipeline_usage/PaddleOCR-VL.html) using **vLLM** as the VLM inference backend. All Python workflows use **[uv](https://docs.astral.sh/uv/)**. ## Architecture ``` Gradio demo (7870) │ ▼ Pipeline API (8090) ── layout + preprocessing (PaddlePaddle GPU) │ ▼ vLLM genai server (8118) ── PaddleOCR-VL-1.6 VLM ``` | Service | Script | Default URL | |---------|--------|-------------| | vLLM VLM server | `./scripts/serve.sh` | `http://127.0.0.1:8118/v1` | | Full pipeline API | `./scripts/serve-pipeline.sh` | `http://127.0.0.1:8090/layout-parsing` | | Online demo UI | `./scripts/run-demo.sh` | `http://127.0.0.1:7870` | The vLLM server exposes only the VLM stage. For HTTP document parsing (layout + OCR), run the pipeline API, which calls vLLM via `config/pipeline_config_vllm.yaml`. ## Prerequisites - Linux with NVIDIA GPU (CC ≥ 8.0 recommended; CUDA 12.6+ driver) - [uv](https://docs.astral.sh/uv/) installed - ~16 GB GPU VRAM for default vLLM settings (tune in `config/vllm_config.yaml`) ## Quick start ```bash git clone ai-ocr-pfm-2026 cd ai-ocr-pfm-2026 cp .env.example .env # adjust CUDA_VISIBLE_DEVICES if needed # 1) Install vLLM server (.venv) FLASH_ATTN_WHEEL="https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.3.14/flash_attn-2.8.2+cu128torch2.8-cp312-cp312-linux_x86_64.whl" \ ./scripts/install.sh # 2) Install pipeline API (.venv-api) — optional, needed for demo / full HTTP API ./scripts/install-pipeline.sh ``` Start services (three terminals, or background each): ```bash ./scripts/serve.sh # vLLM on :8118 ./scripts/serve-pipeline.sh # pipeline on :8090 ./scripts/run-demo.sh # Gradio on :7870 ``` Health checks: ```bash curl -s http://127.0.0.1:8118/v1/models | jq . curl -s http://127.0.0.1:8090/health curl -s -o /dev/null -w "%{http_code}\n" http://127.0.0.1:7870/ ``` ## Configuration Copy `.env.example` to `.env`: | Variable | Default | Description | |----------|---------|-------------| | `GENAI_HOST` | `0.0.0.0` | vLLM bind address | | `GENAI_PORT` | `8118` | vLLM port | | `GENAI_MODEL` | `PaddleOCR-VL-1.6-0.9B` | Model name | | `VLLM_CONFIG` | `config/vllm_config.yaml` | vLLM tuning | | `CUDA_VISIBLE_DEVICES` | `1` | GPU for vLLM (use least-busy GPU) | | `PIPELINE_PORT` | `8090` | Pipeline API port | | `PIPELINE_DEVICE` | `gpu:0` | GPU for layout/preprocessing | | `GRADIO_PORT` | `7870` | Demo UI port | vLLM tuning (`config/vllm_config.yaml`): ```yaml gpu-memory-utilization: 0.75 max-num-seqs: 128 ``` ## Client usage ### Python (vLLM only) ```python from paddleocr import PaddleOCRVL pipeline = PaddleOCRVL( vl_rec_backend="vllm-server", vl_rec_server_url="http://127.0.0.1:8118/v1", ) output = pipeline.predict("path/to/image.png") ``` Run the client in a **separate** environment if it needs PaddlePaddle GPU alongside Transformers. ### CLI ```bash uv run paddleocr doc_parser \ --input demo.png \ --vl_rec_backend vllm-server \ --vl_rec_server_url http://127.0.0.1:8118/v1 ``` ### HTTP (full pipeline) ```bash curl -X POST http://127.0.0.1:8090/layout-parsing \ -H "Content-Type: application/json" \ -d '{"file":"", "fileType": 1, "useLayoutDetection": true}' ``` ## Project layout ``` config/ vllm_config.yaml # vLLM backend tuning pipeline_config_vllm.yaml # pipeline → vLLM server URL scripts/ install.sh # bootstrap .venv (vLLM) install-pipeline.sh # bootstrap .venv-api (pipeline) serve.sh # start vLLM genai server serve-pipeline.sh # start pipeline API run-demo.sh # start Gradio demo PaddleOCR-VL-1.6_Online_Demo/ # bundled Hugging Face-style demo issues/ # recorded problems and fixes AGENTS.md # agent / contributor guide ``` ## Troubleshooting See [issues/](issues/) for detailed write-ups. Common fixes: | Symptom | Fix | |---------|-----| | GPU OOM on vLLM startup | Lower `gpu-memory-utilization` or set `CUDA_VISIBLE_DEVICES` to a free GPU | | flash-attn build failure | Use prebuilt wheel via `FLASH_ATTN_WHEEL=... ./scripts/install.sh` | | Port 8080 in use | Pipeline defaults to **8090**; demo defaults to **7870** | Agent conventions and issue-recording rules: [AGENTS.md](AGENTS.md). ## References - [PaddleOCR-VL usage tutorial](https://www.paddleocr.ai/latest/en/version3.x/pipeline_usage/PaddleOCR-VL.html) - [PaddleOCR genai_server FAQ](https://github.com/PaddlePaddle/PaddleOCR/discussions/16822) - [flash-attention prebuild wheels](https://mjunya.com/flash-attention-prebuild-wheels/)