feat: consolidate backend and docker-compose setup
This commit is contained in:
commit
ff3753a745
306 files changed
+35450
No files matched your search
@@ -0,0 +1,151 @@
|
||||
# PaddleOCR-VL-1.6 on vLLM
|
||||
|
||||
Local deployment of [PaddleOCR-VL-1.6](https://www.paddleocr.ai/latest/en/version3.x/pipeline_usage/PaddleOCR-VL.html) using **vLLM** as the VLM inference backend. All Python workflows use **[uv](https://docs.astral.sh/uv/)**.
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
Gradio demo (7870)
|
||||
│
|
||||
▼
|
||||
Pipeline API (8090) ── layout + preprocessing (PaddlePaddle GPU)
|
||||
│
|
||||
▼
|
||||
vLLM genai server (8118) ── PaddleOCR-VL-1.6 VLM
|
||||
```
|
||||
|
||||
| Service | Script | Default URL |
|
||||
|---------|--------|-------------|
|
||||
| vLLM VLM server | `./scripts/serve.sh` | `http://127.0.0.1:8118/v1` |
|
||||
| Full pipeline API | `./scripts/serve-pipeline.sh` | `http://127.0.0.1:8090/layout-parsing` |
|
||||
| Online demo UI | `./scripts/run-demo.sh` | `http://127.0.0.1:7870` |
|
||||
|
||||
The vLLM server exposes only the VLM stage. For HTTP document parsing (layout + OCR), run the pipeline API, which calls vLLM via `config/pipeline_config_vllm.yaml`.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- Linux with NVIDIA GPU (CC ≥ 8.0 recommended; CUDA 12.6+ driver)
|
||||
- [uv](https://docs.astral.sh/uv/) installed
|
||||
- ~16 GB GPU VRAM for default vLLM settings (tune in `config/vllm_config.yaml`)
|
||||
|
||||
## Quick start
|
||||
|
||||
```bash
|
||||
git clone <repo-url> ai-ocr-pfm-2026
|
||||
cd ai-ocr-pfm-2026
|
||||
|
||||
cp .env.example .env # adjust CUDA_VISIBLE_DEVICES if needed
|
||||
|
||||
# 1) Install vLLM server (.venv)
|
||||
FLASH_ATTN_WHEEL="https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.3.14/flash_attn-2.8.2+cu128torch2.8-cp312-cp312-linux_x86_64.whl" \
|
||||
./scripts/install.sh
|
||||
|
||||
# 2) Install pipeline API (.venv-api) — optional, needed for demo / full HTTP API
|
||||
./scripts/install-pipeline.sh
|
||||
```
|
||||
|
||||
Start services (three terminals, or background each):
|
||||
|
||||
```bash
|
||||
./scripts/serve.sh # vLLM on :8118
|
||||
./scripts/serve-pipeline.sh # pipeline on :8090
|
||||
./scripts/run-demo.sh # Gradio on :7870
|
||||
```
|
||||
|
||||
Health checks:
|
||||
|
||||
```bash
|
||||
curl -s http://127.0.0.1:8118/v1/models | jq .
|
||||
curl -s http://127.0.0.1:8090/health
|
||||
curl -s -o /dev/null -w "%{http_code}\n" http://127.0.0.1:7870/
|
||||
```
|
||||
|
||||
## Configuration
|
||||
|
||||
Copy `.env.example` to `.env`:
|
||||
|
||||
| Variable | Default | Description |
|
||||
|----------|---------|-------------|
|
||||
| `GENAI_HOST` | `0.0.0.0` | vLLM bind address |
|
||||
| `GENAI_PORT` | `8118` | vLLM port |
|
||||
| `GENAI_MODEL` | `PaddleOCR-VL-1.6-0.9B` | Model name |
|
||||
| `VLLM_CONFIG` | `config/vllm_config.yaml` | vLLM tuning |
|
||||
| `CUDA_VISIBLE_DEVICES` | `1` | GPU for vLLM (use least-busy GPU) |
|
||||
| `PIPELINE_PORT` | `8090` | Pipeline API port |
|
||||
| `PIPELINE_DEVICE` | `gpu:0` | GPU for layout/preprocessing |
|
||||
| `GRADIO_PORT` | `7870` | Demo UI port |
|
||||
|
||||
vLLM tuning (`config/vllm_config.yaml`):
|
||||
|
||||
```yaml
|
||||
gpu-memory-utilization: 0.75
|
||||
max-num-seqs: 128
|
||||
```
|
||||
|
||||
## Client usage
|
||||
|
||||
### Python (vLLM only)
|
||||
|
||||
```python
|
||||
from paddleocr import PaddleOCRVL
|
||||
|
||||
pipeline = PaddleOCRVL(
|
||||
vl_rec_backend="vllm-server",
|
||||
vl_rec_server_url="http://127.0.0.1:8118/v1",
|
||||
)
|
||||
output = pipeline.predict("path/to/image.png")
|
||||
```
|
||||
|
||||
Run the client in a **separate** environment if it needs PaddlePaddle GPU alongside Transformers.
|
||||
|
||||
### CLI
|
||||
|
||||
```bash
|
||||
uv run paddleocr doc_parser \
|
||||
--input demo.png \
|
||||
--vl_rec_backend vllm-server \
|
||||
--vl_rec_server_url http://127.0.0.1:8118/v1
|
||||
```
|
||||
|
||||
### HTTP (full pipeline)
|
||||
|
||||
```bash
|
||||
curl -X POST http://127.0.0.1:8090/layout-parsing \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"file":"<base64>", "fileType": 1, "useLayoutDetection": true}'
|
||||
```
|
||||
|
||||
## Project layout
|
||||
|
||||
```
|
||||
config/
|
||||
vllm_config.yaml # vLLM backend tuning
|
||||
pipeline_config_vllm.yaml # pipeline → vLLM server URL
|
||||
scripts/
|
||||
install.sh # bootstrap .venv (vLLM)
|
||||
install-pipeline.sh # bootstrap .venv-api (pipeline)
|
||||
serve.sh # start vLLM genai server
|
||||
serve-pipeline.sh # start pipeline API
|
||||
run-demo.sh # start Gradio demo
|
||||
PaddleOCR-VL-1.6_Online_Demo/ # bundled Hugging Face-style demo
|
||||
issues/ # recorded problems and fixes
|
||||
AGENTS.md # agent / contributor guide
|
||||
```
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
See [issues/](issues/) for detailed write-ups. Common fixes:
|
||||
|
||||
| Symptom | Fix |
|
||||
|---------|-----|
|
||||
| GPU OOM on vLLM startup | Lower `gpu-memory-utilization` or set `CUDA_VISIBLE_DEVICES` to a free GPU |
|
||||
| flash-attn build failure | Use prebuilt wheel via `FLASH_ATTN_WHEEL=... ./scripts/install.sh` |
|
||||
| Port 8080 in use | Pipeline defaults to **8090**; demo defaults to **7870** |
|
||||
|
||||
Agent conventions and issue-recording rules: [AGENTS.md](AGENTS.md).
|
||||
|
||||
## References
|
||||
|
||||
- [PaddleOCR-VL usage tutorial](https://www.paddleocr.ai/latest/en/version3.x/pipeline_usage/PaddleOCR-VL.html)
|
||||
- [PaddleOCR genai_server FAQ](https://github.com/PaddlePaddle/PaddleOCR/discussions/16822)
|
||||
- [flash-attention prebuild wheels](https://mjunya.com/flash-attention-prebuild-wheels/)
|
||||
Reference in new issue
Block a user