238 lines
9.0 KiB
Markdown
238 lines
9.0 KiB
Markdown
# AGENTS: PaddleOCR-VL-1.6 vLLM Service
|
|
|
|
This repository serves **PaddleOCR-VL-1.6** as a dedicated VLM inference backend using **vLLM**. All Python workflows use **uv** (never bare `pip` or system Python).
|
|
|
|
## Architecture
|
|
|
|
```
|
|
Client (PaddleOCR pipeline) --> HTTP /v1 --> paddleocr genai_server (vLLM backend)
|
|
```
|
|
|
|
This service exposes only the VLM stage. Clients connect with `vl_rec_backend="vllm-server"` and `vl_rec_server_url="http://<host>:8118/v1"`.
|
|
|
|
## Prerequisites
|
|
|
|
- Linux with NVIDIA GPU (CC >= 8.0 recommended; CUDA 12.6+ driver support)
|
|
- [uv](https://docs.astral.sh/uv/) installed (`uv --version`)
|
|
- ~16 GB GPU VRAM for default settings (tune via `config/vllm_config.yaml`)
|
|
|
|
## Quick start
|
|
|
|
```bash
|
|
cd <YOUR-WORKING-DIR>/ai-ocr-pfm-2026
|
|
|
|
# 1) Create Python 3.12 venv and install dependencies
|
|
./scripts/install.sh
|
|
|
|
# 2) Start the vLLM-backed genai server
|
|
./scripts/serve.sh
|
|
```
|
|
|
|
Default endpoint: `http://0.0.0.0:8118/v1`
|
|
|
|
## uv conventions (always follow)
|
|
|
|
| Task | Command |
|
|
|------|---------|
|
|
| Create/sync env | `uv sync` |
|
|
| Run any Python | `uv run <command>` |
|
|
| Add a package | `uv add <package>` |
|
|
| Run server | `./scripts/serve.sh` or `uv run paddleocr genai_server ...` |
|
|
|
|
Never use `python -m pip`, `pip install`, or `python -m venv` directly in this repo.
|
|
|
|
## Issue recording (always follow)
|
|
|
|
**Every problem encountered** during install, serve, debug, or client integration must be written to `issues/` before moving on — even if it was resolved in the same session.
|
|
|
|
### Naming
|
|
|
|
```
|
|
issues/{NN}-{slug}.md
|
|
```
|
|
|
|
| Part | Rule | Example |
|
|
|------|------|---------|
|
|
| `{NN}` | Two-digit running number (`01`, `02`, …). Increment from the highest existing file. | `03` |
|
|
| `{slug}` | Lowercase kebab-case summary of the problem | `gpu-memory-startup-failure` |
|
|
|
|
Full example: `issues/04-gpu-memory-startup-failure.md`
|
|
|
|
### When to create a file
|
|
|
|
- Install or dependency errors (flash-attn, vLLM, uv conflicts)
|
|
- Server startup or runtime failures (OOM, port bind, model load)
|
|
- Client integration bugs or misconfiguration
|
|
- Workarounds that took non-obvious steps to discover
|
|
|
|
Do **not** rely on chat history or inline comments alone — if it blocked progress, it belongs in `issues/`.
|
|
|
|
### File template
|
|
|
|
```markdown
|
|
# Issue {NN}: {Short title}
|
|
|
|
## Problem
|
|
What failed, with exact error message or symptom.
|
|
|
|
## Context
|
|
Environment, command run, relevant config (`.env`, `config/vllm_config.yaml`).
|
|
|
|
## Solution
|
|
What fixed it, or current workaround / open status.
|
|
|
|
## References
|
|
Links, related issue files, or AGENTS.md sections.
|
|
```
|
|
|
|
### Index
|
|
|
|
Check `issues/` for the next number:
|
|
|
|
```bash
|
|
ls issues/*.md 2>/dev/null | sort
|
|
```
|
|
|
|
See [issues/](issues/) for recorded problems and fixes from this project.
|
|
|
|
## Environment variables
|
|
|
|
Copy `.env.example` to `.env` and adjust as needed:
|
|
|
|
| Variable | Default | Description |
|
|
|----------|---------|-------------|
|
|
| `GENAI_HOST` | `0.0.0.0` | Bind address |
|
|
| `GENAI_PORT` | `8118` | Service port |
|
|
| `GENAI_MODEL` | `PaddleOCR-VL-1.6-0.9B` | Model name for `genai_server` |
|
|
| `GENAI_BACKEND` | `vllm` | Inference backend |
|
|
| `VLLM_CONFIG` | `config/vllm_config.yaml` | vLLM backend YAML config |
|
|
| `CUDA_VISIBLE_DEVICES` | `1` (see `.env.example`) | GPU index(es) to use |
|
|
|
|
On dual-GPU hosts, pick the GPU with more free VRAM. If startup fails with a memory error, lower `gpu-memory-utilization` in `config/vllm_config.yaml`.
|
|
|
|
## Client usage
|
|
|
|
After the server is running:
|
|
|
|
```bash
|
|
# CLI
|
|
uv run paddleocr doc_parser \
|
|
--input https://paddle-model-ecology.bj.bcebos.com/paddlex/imgs/demo_image/paddleocr_vl_demo.png \
|
|
--vl_rec_backend vllm-server \
|
|
--vl_rec_server_url http://localhost:8118/v1
|
|
```
|
|
|
|
```python
|
|
from paddleocr import PaddleOCRVL
|
|
|
|
pipeline = PaddleOCRVL(
|
|
vl_rec_backend="vllm-server",
|
|
vl_rec_server_url="http://127.0.0.1:8118/v1",
|
|
)
|
|
output = pipeline.predict("path/to/image.png")
|
|
```
|
|
|
|
Note: The full PaddleOCR-VL client should run in a **separate** environment if it needs PaddlePaddle GPU + Transformers. This repo is the isolated vLLM server only.
|
|
|
|
## Tuning vLLM
|
|
|
|
Edit `config/vllm_config.yaml`:
|
|
|
|
```yaml
|
|
gpu-memory-utilization: 0.8
|
|
max-num-seqs: 128
|
|
```
|
|
|
|
Reference: [PaddleOCR-VL vLLM parameter tuning](https://www.paddleocr.ai/latest/en/version3.x/pipeline_usage/PaddleOCR-VL.html#331-server-side-parameter-adjustment)
|
|
|
|
## Troubleshooting
|
|
|
|
See `issues/` for full write-ups. Quick pointers:
|
|
|
|
| Symptom | Issue file |
|
|
|---------|------------|
|
|
| `paddleocr install_genai_server_deps` / `No module named pip` | [01-genai-server-deps-pip-in-uv-venv.md](issues/01-genai-server-deps-pip-in-uv-venv.md) |
|
|
| flash-attn wheel incompatible with Python version | [02-flash-attn-wheel-python-version-mismatch.md](issues/02-flash-attn-wheel-python-version-mismatch.md) |
|
|
| `uv pip` targets wrong venv from another project | [03-active-virtual-env-from-other-project.md](issues/03-active-virtual-env-from-other-project.md) |
|
|
| Free memory below `gpu-memory-utilization` on startup | [04-gpu-memory-startup-failure.md](issues/04-gpu-memory-startup-failure.md) |
|
|
| `TokenizersBackend has no attribute all_special_tokens_extended` | [05-transformers-tokenizers-incompatibility.md](issues/05-transformers-tokenizers-incompatibility.md) |
|
|
| Extracted images not shown in Gradio demo (raw base64 in markdown) | [06-extracted-images-raw-base64-not-displayed.md](issues/06-extracted-images-raw-base64-not-displayed.md) |
|
|
|
|
### flash-attn build failures
|
|
|
|
Install the prebuilt wheel after `uv sync` (see `scripts/install.sh`):
|
|
|
|
```bash
|
|
FLASH_ATTN_WHEEL="https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.3.14/flash_attn-2.8.2+cu128torch2.8-cp312-cp312-linux_x86_64.whl" \
|
|
./scripts/install.sh
|
|
```
|
|
|
|
Pick the wheel matching your Python and CUDA versions from [flash-attention prebuild wheels](https://mjunya.com/flash-attention-prebuild-wheels/). Details: [02-flash-attn-wheel-python-version-mismatch.md](issues/02-flash-attn-wheel-python-version-mismatch.md).
|
|
|
|
Note: `paddleocr install_genai_server_deps` uses `pip` internally and is incompatible with uv-managed venvs. See [01-genai-server-deps-pip-in-uv-venv.md](issues/01-genai-server-deps-pip-in-uv-venv.md). This repo installs the vLLM stack via `uv sync` + `uv pip`.
|
|
|
|
### `TokenizersBackend has no attribute all_special_tokens_extended`
|
|
|
|
Pin transformers (already in `pyproject.toml`):
|
|
|
|
```bash
|
|
uv pip install "transformers==4.57.6"
|
|
```
|
|
|
|
See [05-transformers-tokenizers-incompatibility.md](issues/05-transformers-tokenizers-incompatibility.md).
|
|
|
|
### Do not install `paddlepaddle-gpu` in this venv
|
|
|
|
vLLM and PaddlePaddle GPU conflict. This server env uses `paddleocr[doc-parser]` without Paddle GPU.
|
|
|
|
### GPU memory on startup
|
|
|
|
If vLLM reports free memory below `gpu-memory-utilization`, either:
|
|
|
|
- Set `CUDA_VISIBLE_DEVICES` to a less-busy GPU
|
|
- Lower `gpu-memory-utilization` in `config/vllm_config.yaml` (e.g. `0.75` or `0.7`)
|
|
|
|
See [04-gpu-memory-startup-failure.md](issues/04-gpu-memory-startup-failure.md).
|
|
|
|
### Health check
|
|
|
|
```bash
|
|
curl -s http://localhost:8118/v1/models | jq .
|
|
```
|
|
|
|
## File map
|
|
|
|
| Path | Purpose |
|
|
|------|---------|
|
|
| `issues/` | Recorded problems and fixes (`{NN}-{slug}.md`) |
|
|
| `pyproject.toml` | uv project metadata and base dependencies |
|
|
| `scripts/install.sh` | Bootstrap venv + vLLM server deps |
|
|
| `scripts/serve.sh` | Start `paddleocr genai_server` |
|
|
| `config/vllm_config.yaml` | vLLM backend tuning |
|
|
| `.env.example` | Environment variable template |
|
|
|
|
## References
|
|
|
|
- [PaddleOCR-VL usage tutorial](https://www.paddleocr.ai/latest/en/version3.x/pipeline_usage/PaddleOCR-VL.html)
|
|
- [PaddleOCR genai_server FAQ](https://github.com/PaddlePaddle/PaddleOCR/discussions/16822)
|
|
|
|
## Coding Guidelines (always follow)
|
|
|
|
We use the [karpathy-guidelines](file:///home/user/LABS/OCR/paddle-ocr-vl-1-6-using-vllm-2026/andrej-karpathy-skills/skills/karpathy-guidelines/SKILL.md) skill to reduce common LLM coding mistakes. Refer to [SKILL.md](file:///home/user/LABS/OCR/paddle-ocr-vl-1-6-using-vllm-2026/andrej-karpathy-skills/skills/karpathy-guidelines/SKILL.md) for details:
|
|
1. **Think Before Coding**: Explicitly state assumptions and surface tradeoffs instead of making silent choices.
|
|
2. **Simplicity First**: Write the minimum amount of code to solve the problem with zero speculative configurations.
|
|
3. **Surgical Changes**: Edit only what is required and match the existing coding style exactly.
|
|
4. **Goal-Driven Execution**: Define verifiable success criteria and run automated tests/screenshots to confirm correctness.
|
|
|
|
## Path Guidelines (always follow)
|
|
|
|
Never use full paths containing the user's logged-in name (e.g., `/home/{uid}/path`). Always use relative paths instead (e.g., `.` or `./path` relative to the workspace root).
|
|
|
|
## App Testing Guidelines (always follow)
|
|
|
|
When the user intentionally asks to test the app:
|
|
- Use browser tools to test the app.
|
|
- Take a screenshot for each sample image, each step, and each variant/option (if any), until the OCR result appears.
|
|
- Save the screenshots in the `/screenshots/` folder.
|
|
- Follow the file naming convention: `{2-digit-number}-{step#}-{variant_or_options_if_any}-{slug}.jpg` (e.g., `01-step1-default-upload.jpg`).
|