152 lines
4.7 KiB
Markdown
152 lines
4.7 KiB
Markdown
# PaddleOCR-VL-1.6 on vLLM
|
|
|
|
Local deployment of [PaddleOCR-VL-1.6](https://www.paddleocr.ai/latest/en/version3.x/pipeline_usage/PaddleOCR-VL.html) using **vLLM** as the VLM inference backend. All Python workflows use **[uv](https://docs.astral.sh/uv/)**.
|
|
|
|
## Architecture
|
|
|
|
```
|
|
Gradio demo (7870)
|
|
│
|
|
▼
|
|
Pipeline API (8090) ── layout + preprocessing (PaddlePaddle GPU)
|
|
│
|
|
▼
|
|
vLLM genai server (8118) ── PaddleOCR-VL-1.6 VLM
|
|
```
|
|
|
|
| Service | Script | Default URL |
|
|
|---------|--------|-------------|
|
|
| vLLM VLM server | `./scripts/serve.sh` | `http://127.0.0.1:8118/v1` |
|
|
| Full pipeline API | `./scripts/serve-pipeline.sh` | `http://127.0.0.1:8090/layout-parsing` |
|
|
| Online demo UI | `./scripts/run-demo.sh` | `http://127.0.0.1:7870` |
|
|
|
|
The vLLM server exposes only the VLM stage. For HTTP document parsing (layout + OCR), run the pipeline API, which calls vLLM via `config/pipeline_config_vllm.yaml`.
|
|
|
|
## Prerequisites
|
|
|
|
- Linux with NVIDIA GPU (CC ≥ 8.0 recommended; CUDA 12.6+ driver)
|
|
- [uv](https://docs.astral.sh/uv/) installed
|
|
- ~16 GB GPU VRAM for default vLLM settings (tune in `config/vllm_config.yaml`)
|
|
|
|
## Quick start
|
|
|
|
```bash
|
|
git clone <repo-url> ai-ocr-pfm-2026
|
|
cd ai-ocr-pfm-2026
|
|
|
|
cp .env.example .env # adjust CUDA_VISIBLE_DEVICES if needed
|
|
|
|
# 1) Install vLLM server (.venv)
|
|
FLASH_ATTN_WHEEL="https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.3.14/flash_attn-2.8.2+cu128torch2.8-cp312-cp312-linux_x86_64.whl" \
|
|
./scripts/install.sh
|
|
|
|
# 2) Install pipeline API (.venv-api) — optional, needed for demo / full HTTP API
|
|
./scripts/install-pipeline.sh
|
|
```
|
|
|
|
Start services (three terminals, or background each):
|
|
|
|
```bash
|
|
./scripts/serve.sh # vLLM on :8118
|
|
./scripts/serve-pipeline.sh # pipeline on :8090
|
|
./scripts/run-demo.sh # Gradio on :7870
|
|
```
|
|
|
|
Health checks:
|
|
|
|
```bash
|
|
curl -s http://127.0.0.1:8118/v1/models | jq .
|
|
curl -s http://127.0.0.1:8090/health
|
|
curl -s -o /dev/null -w "%{http_code}\n" http://127.0.0.1:7870/
|
|
```
|
|
|
|
## Configuration
|
|
|
|
Copy `.env.example` to `.env`:
|
|
|
|
| Variable | Default | Description |
|
|
|----------|---------|-------------|
|
|
| `GENAI_HOST` | `0.0.0.0` | vLLM bind address |
|
|
| `GENAI_PORT` | `8118` | vLLM port |
|
|
| `GENAI_MODEL` | `PaddleOCR-VL-1.6-0.9B` | Model name |
|
|
| `VLLM_CONFIG` | `config/vllm_config.yaml` | vLLM tuning |
|
|
| `CUDA_VISIBLE_DEVICES` | `1` | GPU for vLLM (use least-busy GPU) |
|
|
| `PIPELINE_PORT` | `8090` | Pipeline API port |
|
|
| `PIPELINE_DEVICE` | `gpu:0` | GPU for layout/preprocessing |
|
|
| `GRADIO_PORT` | `7870` | Demo UI port |
|
|
|
|
vLLM tuning (`config/vllm_config.yaml`):
|
|
|
|
```yaml
|
|
gpu-memory-utilization: 0.75
|
|
max-num-seqs: 128
|
|
```
|
|
|
|
## Client usage
|
|
|
|
### Python (vLLM only)
|
|
|
|
```python
|
|
from paddleocr import PaddleOCRVL
|
|
|
|
pipeline = PaddleOCRVL(
|
|
vl_rec_backend="vllm-server",
|
|
vl_rec_server_url="http://127.0.0.1:8118/v1",
|
|
)
|
|
output = pipeline.predict("path/to/image.png")
|
|
```
|
|
|
|
Run the client in a **separate** environment if it needs PaddlePaddle GPU alongside Transformers.
|
|
|
|
### CLI
|
|
|
|
```bash
|
|
uv run paddleocr doc_parser \
|
|
--input demo.png \
|
|
--vl_rec_backend vllm-server \
|
|
--vl_rec_server_url http://127.0.0.1:8118/v1
|
|
```
|
|
|
|
### HTTP (full pipeline)
|
|
|
|
```bash
|
|
curl -X POST http://127.0.0.1:8090/layout-parsing \
|
|
-H "Content-Type: application/json" \
|
|
-d '{"file":"<base64>", "fileType": 1, "useLayoutDetection": true}'
|
|
```
|
|
|
|
## Project layout
|
|
|
|
```
|
|
config/
|
|
vllm_config.yaml # vLLM backend tuning
|
|
pipeline_config_vllm.yaml # pipeline → vLLM server URL
|
|
scripts/
|
|
install.sh # bootstrap .venv (vLLM)
|
|
install-pipeline.sh # bootstrap .venv-api (pipeline)
|
|
serve.sh # start vLLM genai server
|
|
serve-pipeline.sh # start pipeline API
|
|
run-demo.sh # start Gradio demo
|
|
PaddleOCR-VL-1.6_Online_Demo/ # bundled Hugging Face-style demo
|
|
issues/ # recorded problems and fixes
|
|
AGENTS.md # agent / contributor guide
|
|
```
|
|
|
|
## Troubleshooting
|
|
|
|
See [issues/](issues/) for detailed write-ups. Common fixes:
|
|
|
|
| Symptom | Fix |
|
|
|---------|-----|
|
|
| GPU OOM on vLLM startup | Lower `gpu-memory-utilization` or set `CUDA_VISIBLE_DEVICES` to a free GPU |
|
|
| flash-attn build failure | Use prebuilt wheel via `FLASH_ATTN_WHEEL=... ./scripts/install.sh` |
|
|
| Port 8080 in use | Pipeline defaults to **8090**; demo defaults to **7870** |
|
|
|
|
Agent conventions and issue-recording rules: [AGENTS.md](AGENTS.md).
|
|
|
|
## References
|
|
|
|
- [PaddleOCR-VL usage tutorial](https://www.paddleocr.ai/latest/en/version3.x/pipeline_usage/PaddleOCR-VL.html)
|
|
- [PaddleOCR genai_server FAQ](https://github.com/PaddlePaddle/PaddleOCR/discussions/16822)
|
|
- [flash-attention prebuild wheels](https://mjunya.com/flash-attention-prebuild-wheels/)
|