chore: normalize line endings (CRLF -> LF)
No content changes: git diff --ignore-all-space over these files is empty. The churn came from editing on Windows against a repo checked out with LF.
This commit is contained in:
1 parent
15566a6951
commit
caf8e98378
315 files changed
+86950
-86950
No files matched your search
+138
-138
@@ -1,138 +1,138 @@
|
||||
# vLLM Service — Full Reference
|
||||
|
||||
Detail split out of `../AGENTS.md` (2026-07-08, to keep that file under the
|
||||
Agents Settings Kit's 256-line threshold once the `e`/`n` workflow was appended
|
||||
to it). `AGENTS.md` keeps the short version — architecture, quick start, the
|
||||
env var table, file map — and links here for everything else.
|
||||
|
||||
## Issue recording — naming and template
|
||||
|
||||
```
|
||||
issues/{NN}-{slug}.md
|
||||
```
|
||||
|
||||
| Part | Rule | Example |
|
||||
|------|------|---------|
|
||||
| `{NN}` | Two-digit running number (`01`, `02`, …). Increment from the highest existing file. | `03` |
|
||||
| `{slug}` | Lowercase kebab-case summary of the problem | `gpu-memory-startup-failure` |
|
||||
|
||||
Full example: `issues/04-gpu-memory-startup-failure.md`
|
||||
|
||||
### File template
|
||||
|
||||
```markdown
|
||||
# Issue {NN}: {Short title}
|
||||
|
||||
## Problem
|
||||
What failed, with exact error message or symptom.
|
||||
|
||||
## Context
|
||||
Environment, command run, relevant config (`.env`, `config/vllm_config.yaml`).
|
||||
|
||||
## Solution
|
||||
What fixed it, or current workaround / open status.
|
||||
|
||||
## References
|
||||
Links, related issue files, or AGENTS.md sections.
|
||||
```
|
||||
|
||||
Check `issues/` for the next number:
|
||||
|
||||
```bash
|
||||
ls issues/*.md 2>/dev/null | sort
|
||||
```
|
||||
|
||||
## Client usage
|
||||
|
||||
After the server is running:
|
||||
|
||||
```bash
|
||||
# CLI
|
||||
uv run paddleocr doc_parser \
|
||||
--input https://paddle-model-ecology.bj.bcebos.com/paddlex/imgs/demo_image/paddleocr_vl_demo.png \
|
||||
--vl_rec_backend vllm-server \
|
||||
--vl_rec_server_url http://localhost:8118/v1
|
||||
```
|
||||
|
||||
```python
|
||||
from paddleocr import PaddleOCRVL
|
||||
|
||||
pipeline = PaddleOCRVL(
|
||||
vl_rec_backend="vllm-server",
|
||||
vl_rec_server_url="http://127.0.0.1:8118/v1",
|
||||
)
|
||||
output = pipeline.predict("path/to/image.png")
|
||||
```
|
||||
|
||||
Note: The full PaddleOCR-VL client should run in a **separate** environment if it needs PaddlePaddle GPU + Transformers. This repo is the isolated vLLM server only.
|
||||
|
||||
## Tuning vLLM
|
||||
|
||||
Edit `config/vllm_config.yaml`:
|
||||
|
||||
```yaml
|
||||
gpu-memory-utilization: 0.8
|
||||
max-num-seqs: 128
|
||||
```
|
||||
|
||||
Reference: [PaddleOCR-VL vLLM parameter tuning](https://www.paddleocr.ai/latest/en/version3.x/pipeline_usage/PaddleOCR-VL.html#331-server-side-parameter-adjustment)
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
See `issues/` for full write-ups. Quick pointers:
|
||||
|
||||
| Symptom | Issue file |
|
||||
|---------|------------|
|
||||
| `paddleocr install_genai_server_deps` / `No module named pip` | [01-genai-server-deps-pip-in-uv-venv.md](../issues/01-genai-server-deps-pip-in-uv-venv.md) |
|
||||
| flash-attn wheel incompatible with Python version | [02-flash-attn-wheel-python-version-mismatch.md](../issues/02-flash-attn-wheel-python-version-mismatch.md) |
|
||||
| `uv pip` targets wrong venv from another project | [03-active-virtual-env-from-other-project.md](../issues/03-active-virtual-env-from-other-project.md) |
|
||||
| Free memory below `gpu-memory-utilization` on startup | [04-gpu-memory-startup-failure.md](../issues/04-gpu-memory-startup-failure.md) |
|
||||
| `TokenizersBackend has no attribute all_special_tokens_extended` | [05-transformers-tokenizers-incompatibility.md](../issues/05-transformers-tokenizers-incompatibility.md) |
|
||||
| Extracted images not shown in Gradio demo (raw base64 in markdown) | [06-extracted-images-raw-base64-not-displayed.md](../issues/06-extracted-images-raw-base64-not-displayed.md) |
|
||||
|
||||
### flash-attn build failures
|
||||
|
||||
Install the prebuilt wheel after `uv sync` (see `scripts/install.sh`):
|
||||
|
||||
```bash
|
||||
FLASH_ATTN_WHEEL="https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.3.14/flash_attn-2.8.2+cu128torch2.8-cp312-cp312-linux_x86_64.whl" \
|
||||
./scripts/install.sh
|
||||
```
|
||||
|
||||
Pick the wheel matching your Python and CUDA versions from [flash-attention prebuild wheels](https://mjunya.com/flash-attention-prebuild-wheels/). Details: [02-flash-attn-wheel-python-version-mismatch.md](../issues/02-flash-attn-wheel-python-version-mismatch.md).
|
||||
|
||||
Note: `paddleocr install_genai_server_deps` uses `pip` internally and is incompatible with uv-managed venvs. See [01-genai-server-deps-pip-in-uv-venv.md](../issues/01-genai-server-deps-pip-in-uv-venv.md). This repo installs the vLLM stack via `uv sync` + `uv pip`.
|
||||
|
||||
### `TokenizersBackend has no attribute all_special_tokens_extended`
|
||||
|
||||
Pin transformers (already in `pyproject.toml`):
|
||||
|
||||
```bash
|
||||
uv pip install "transformers==4.57.6"
|
||||
```
|
||||
|
||||
See [05-transformers-tokenizers-incompatibility.md](../issues/05-transformers-tokenizers-incompatibility.md).
|
||||
|
||||
### Do not install `paddlepaddle-gpu` in this venv
|
||||
|
||||
vLLM and PaddlePaddle GPU conflict. This server env uses `paddleocr[doc-parser]` without Paddle GPU.
|
||||
|
||||
### GPU memory on startup
|
||||
|
||||
If vLLM reports free memory below `gpu-memory-utilization`, either:
|
||||
|
||||
- Set `CUDA_VISIBLE_DEVICES` to a less-busy GPU
|
||||
- Lower `gpu-memory-utilization` in `config/vllm_config.yaml` (e.g. `0.75` or `0.7`)
|
||||
|
||||
See [04-gpu-memory-startup-failure.md](../issues/04-gpu-memory-startup-failure.md).
|
||||
|
||||
### Health check
|
||||
|
||||
```bash
|
||||
curl -s http://localhost:8118/v1/models | jq .
|
||||
```
|
||||
|
||||
## References
|
||||
|
||||
- [PaddleOCR-VL usage tutorial](https://www.paddleocr.ai/latest/en/version3.x/pipeline_usage/PaddleOCR-VL.html)
|
||||
- [PaddleOCR genai_server FAQ](https://github.com/PaddlePaddle/PaddleOCR/discussions/16822)
|
||||
# vLLM Service — Full Reference
|
||||
|
||||
Detail split out of `../AGENTS.md` (2026-07-08, to keep that file under the
|
||||
Agents Settings Kit's 256-line threshold once the `e`/`n` workflow was appended
|
||||
to it). `AGENTS.md` keeps the short version — architecture, quick start, the
|
||||
env var table, file map — and links here for everything else.
|
||||
|
||||
## Issue recording — naming and template
|
||||
|
||||
```
|
||||
issues/{NN}-{slug}.md
|
||||
```
|
||||
|
||||
| Part | Rule | Example |
|
||||
|------|------|---------|
|
||||
| `{NN}` | Two-digit running number (`01`, `02`, …). Increment from the highest existing file. | `03` |
|
||||
| `{slug}` | Lowercase kebab-case summary of the problem | `gpu-memory-startup-failure` |
|
||||
|
||||
Full example: `issues/04-gpu-memory-startup-failure.md`
|
||||
|
||||
### File template
|
||||
|
||||
```markdown
|
||||
# Issue {NN}: {Short title}
|
||||
|
||||
## Problem
|
||||
What failed, with exact error message or symptom.
|
||||
|
||||
## Context
|
||||
Environment, command run, relevant config (`.env`, `config/vllm_config.yaml`).
|
||||
|
||||
## Solution
|
||||
What fixed it, or current workaround / open status.
|
||||
|
||||
## References
|
||||
Links, related issue files, or AGENTS.md sections.
|
||||
```
|
||||
|
||||
Check `issues/` for the next number:
|
||||
|
||||
```bash
|
||||
ls issues/*.md 2>/dev/null | sort
|
||||
```
|
||||
|
||||
## Client usage
|
||||
|
||||
After the server is running:
|
||||
|
||||
```bash
|
||||
# CLI
|
||||
uv run paddleocr doc_parser \
|
||||
--input https://paddle-model-ecology.bj.bcebos.com/paddlex/imgs/demo_image/paddleocr_vl_demo.png \
|
||||
--vl_rec_backend vllm-server \
|
||||
--vl_rec_server_url http://localhost:8118/v1
|
||||
```
|
||||
|
||||
```python
|
||||
from paddleocr import PaddleOCRVL
|
||||
|
||||
pipeline = PaddleOCRVL(
|
||||
vl_rec_backend="vllm-server",
|
||||
vl_rec_server_url="http://127.0.0.1:8118/v1",
|
||||
)
|
||||
output = pipeline.predict("path/to/image.png")
|
||||
```
|
||||
|
||||
Note: The full PaddleOCR-VL client should run in a **separate** environment if it needs PaddlePaddle GPU + Transformers. This repo is the isolated vLLM server only.
|
||||
|
||||
## Tuning vLLM
|
||||
|
||||
Edit `config/vllm_config.yaml`:
|
||||
|
||||
```yaml
|
||||
gpu-memory-utilization: 0.8
|
||||
max-num-seqs: 128
|
||||
```
|
||||
|
||||
Reference: [PaddleOCR-VL vLLM parameter tuning](https://www.paddleocr.ai/latest/en/version3.x/pipeline_usage/PaddleOCR-VL.html#331-server-side-parameter-adjustment)
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
See `issues/` for full write-ups. Quick pointers:
|
||||
|
||||
| Symptom | Issue file |
|
||||
|---------|------------|
|
||||
| `paddleocr install_genai_server_deps` / `No module named pip` | [01-genai-server-deps-pip-in-uv-venv.md](../issues/01-genai-server-deps-pip-in-uv-venv.md) |
|
||||
| flash-attn wheel incompatible with Python version | [02-flash-attn-wheel-python-version-mismatch.md](../issues/02-flash-attn-wheel-python-version-mismatch.md) |
|
||||
| `uv pip` targets wrong venv from another project | [03-active-virtual-env-from-other-project.md](../issues/03-active-virtual-env-from-other-project.md) |
|
||||
| Free memory below `gpu-memory-utilization` on startup | [04-gpu-memory-startup-failure.md](../issues/04-gpu-memory-startup-failure.md) |
|
||||
| `TokenizersBackend has no attribute all_special_tokens_extended` | [05-transformers-tokenizers-incompatibility.md](../issues/05-transformers-tokenizers-incompatibility.md) |
|
||||
| Extracted images not shown in Gradio demo (raw base64 in markdown) | [06-extracted-images-raw-base64-not-displayed.md](../issues/06-extracted-images-raw-base64-not-displayed.md) |
|
||||
|
||||
### flash-attn build failures
|
||||
|
||||
Install the prebuilt wheel after `uv sync` (see `scripts/install.sh`):
|
||||
|
||||
```bash
|
||||
FLASH_ATTN_WHEEL="https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.3.14/flash_attn-2.8.2+cu128torch2.8-cp312-cp312-linux_x86_64.whl" \
|
||||
./scripts/install.sh
|
||||
```
|
||||
|
||||
Pick the wheel matching your Python and CUDA versions from [flash-attention prebuild wheels](https://mjunya.com/flash-attention-prebuild-wheels/). Details: [02-flash-attn-wheel-python-version-mismatch.md](../issues/02-flash-attn-wheel-python-version-mismatch.md).
|
||||
|
||||
Note: `paddleocr install_genai_server_deps` uses `pip` internally and is incompatible with uv-managed venvs. See [01-genai-server-deps-pip-in-uv-venv.md](../issues/01-genai-server-deps-pip-in-uv-venv.md). This repo installs the vLLM stack via `uv sync` + `uv pip`.
|
||||
|
||||
### `TokenizersBackend has no attribute all_special_tokens_extended`
|
||||
|
||||
Pin transformers (already in `pyproject.toml`):
|
||||
|
||||
```bash
|
||||
uv pip install "transformers==4.57.6"
|
||||
```
|
||||
|
||||
See [05-transformers-tokenizers-incompatibility.md](../issues/05-transformers-tokenizers-incompatibility.md).
|
||||
|
||||
### Do not install `paddlepaddle-gpu` in this venv
|
||||
|
||||
vLLM and PaddlePaddle GPU conflict. This server env uses `paddleocr[doc-parser]` without Paddle GPU.
|
||||
|
||||
### GPU memory on startup
|
||||
|
||||
If vLLM reports free memory below `gpu-memory-utilization`, either:
|
||||
|
||||
- Set `CUDA_VISIBLE_DEVICES` to a less-busy GPU
|
||||
- Lower `gpu-memory-utilization` in `config/vllm_config.yaml` (e.g. `0.75` or `0.7`)
|
||||
|
||||
See [04-gpu-memory-startup-failure.md](../issues/04-gpu-memory-startup-failure.md).
|
||||
|
||||
### Health check
|
||||
|
||||
```bash
|
||||
curl -s http://localhost:8118/v1/models | jq .
|
||||
```
|
||||
|
||||
## References
|
||||
|
||||
- [PaddleOCR-VL usage tutorial](https://www.paddleocr.ai/latest/en/version3.x/pipeline_usage/PaddleOCR-VL.html)
|
||||
- [PaddleOCR genai_server FAQ](https://github.com/PaddlePaddle/PaddleOCR/discussions/16822)
|
||||
Reference in new issue
Block a user