Grilled 2026-07-16 with the user; full decision record in docs/expiry-tracking-plan.md. Core reframe: expiry is captured once per batch at DO intake (staff-typed on the stock-entry confirmation page, from the physical packs), so the cashier scan only MATCHES OCR fragments against the 1-3 known in-stock batch dates instead of free-reading damaged dot-matrix prints (proven model-capability ceiling, 2026-07-15). Fallback: auto-FEFO + 'inferred' flag, zero cashier interaction. No cloud, ever. - docs/expiry-tracking-plan.md: architecture, matching algorithm spec (resolveExpiryFromEvidence), schema/API deltas, phases 1-3, testing plan - backend plans §13 (13.1-13.4): matcher util + offline tuning, route wiring + expiry_source provenance, multi-frame union, dot-matrix recognizer fine-tune - root plans §10 (10.1-10.3): cashier fast path, inferred badge + end-of-day review, burst capture for mounted camera - stock-feature-plan.md: extension note (batch dropdown becomes the manual-override path) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q8TumxFDnyVnfsR3mxPXfX
PaddleOCR-VL-1.6 on vLLM
Local deployment of PaddleOCR-VL-1.6 using vLLM as the VLM inference backend. All Python workflows use uv.
Architecture
Gradio demo (7870)
│
▼
Pipeline API (8090) ── layout + preprocessing (PaddlePaddle GPU)
│
▼
vLLM genai server (8118) ── PaddleOCR-VL-1.6 VLM
| Service | Script | Default URL |
|---|---|---|
| vLLM VLM server | ./scripts/serve.sh |
http://127.0.0.1:8118/v1 |
| Full pipeline API | ./scripts/serve-pipeline.sh |
http://127.0.0.1:8090/layout-parsing |
| Online demo UI | ./scripts/run-demo.sh |
http://127.0.0.1:7870 |
The vLLM server exposes only the VLM stage. For HTTP document parsing (layout + OCR), run the pipeline API, which calls vLLM via config/pipeline_config_vllm.yaml.
Prerequisites
- Linux with NVIDIA GPU (CC ≥ 8.0 recommended; CUDA 12.6+ driver)
- uv installed
- ~16 GB GPU VRAM for default vLLM settings (tune in
config/vllm_config.yaml)
Quick start
git clone <repo-url> ai-ocr-pfm-2026
cd ai-ocr-pfm-2026
cp .env.example .env # adjust CUDA_VISIBLE_DEVICES if needed
# 1) Install vLLM server (.venv)
FLASH_ATTN_WHEEL="https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.3.14/flash_attn-2.8.2+cu128torch2.8-cp312-cp312-linux_x86_64.whl" \
./scripts/install.sh
# 2) Install pipeline API (.venv-api) — optional, needed for demo / full HTTP API
./scripts/install-pipeline.sh
Start services (three terminals, or background each):
./scripts/serve.sh # vLLM on :8118
./scripts/serve-pipeline.sh # pipeline on :8090
./scripts/run-demo.sh # Gradio on :7870
Health checks:
curl -s http://127.0.0.1:8118/v1/models | jq .
curl -s http://127.0.0.1:8090/health
curl -s -o /dev/null -w "%{http_code}\n" http://127.0.0.1:7870/
Configuration
Copy .env.example to .env:
| Variable | Default | Description |
|---|---|---|
GENAI_HOST |
0.0.0.0 |
vLLM bind address |
GENAI_PORT |
8118 |
vLLM port |
GENAI_MODEL |
PaddleOCR-VL-1.6-0.9B |
Model name |
VLLM_CONFIG |
config/vllm_config.yaml |
vLLM tuning |
CUDA_VISIBLE_DEVICES |
1 |
GPU for vLLM (use least-busy GPU) |
PIPELINE_PORT |
8090 |
Pipeline API port |
PIPELINE_DEVICE |
gpu:0 |
GPU for layout/preprocessing |
GRADIO_PORT |
7870 |
Demo UI port |
vLLM tuning (config/vllm_config.yaml):
gpu-memory-utilization: 0.75
max-num-seqs: 128
Client usage
Python (vLLM only)
from paddleocr import PaddleOCRVL
pipeline = PaddleOCRVL(
vl_rec_backend="vllm-server",
vl_rec_server_url="http://127.0.0.1:8118/v1",
)
output = pipeline.predict("path/to/image.png")
Run the client in a separate environment if it needs PaddlePaddle GPU alongside Transformers.
CLI
uv run paddleocr doc_parser \
--input demo.png \
--vl_rec_backend vllm-server \
--vl_rec_server_url http://127.0.0.1:8118/v1
HTTP (full pipeline)
curl -X POST http://127.0.0.1:8090/layout-parsing \
-H "Content-Type: application/json" \
-d '{"file":"<base64>", "fileType": 1, "useLayoutDetection": true}'
Project layout
config/
vllm_config.yaml # vLLM backend tuning
pipeline_config_vllm.yaml # pipeline → vLLM server URL
scripts/
install.sh # bootstrap .venv (vLLM)
install-pipeline.sh # bootstrap .venv-api (pipeline)
serve.sh # start vLLM genai server
serve-pipeline.sh # start pipeline API
run-demo.sh # start Gradio demo
PaddleOCR-VL-1.6_Online_Demo/ # bundled Hugging Face-style demo
issues/ # recorded problems and fixes
AGENTS.md # agent / contributor guide
Troubleshooting
See issues/ for detailed write-ups. Common fixes:
| Symptom | Fix |
|---|---|
| GPU OOM on vLLM startup | Lower gpu-memory-utilization or set CUDA_VISIBLE_DEVICES to a free GPU |
| flash-attn build failure | Use prebuilt wheel via FLASH_ATTN_WHEEL=... ./scripts/install.sh |
| Port 8080 in use | Pipeline defaults to 8090; demo defaults to 7870 |
Agent conventions and issue-recording rules: AGENTS.md.