Replaces stale setup instructions (single hardcoded API URL, old migration story) with the dual-mode ngrok/LAN config, the demo/production compose override, and a "what to consider" section grounded in real issues hit this session: the duplicate backend/docker-compose.yml project-name collision, store_master shipping unseeded, and the debug-signed APK. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017eRAsLqN9Sg1b9YPz22Lxz
Prima Fresh Mart Scanner (app-pfm-ocr-v2)
An on-premise, GPU-accelerated OCR system that turns a driver's phone photo of a Delivery Order (DO) into structured, database-backed data — PO/SO/DO numbers, dates, store, and item lines — with a Flutter mobile client on one end and a Dockerized AI pipeline on the other.
No cloud OCR API is used. Everything (layout detection, text recognition, LLM-assisted structuring) runs on your own GPU.
Table of Contents
- What This Is
- Tech Stack
- Architecture
- How It Works
- Repository Structure
- Getting Started
- Running in Demo / Production Mode
- What to Consider Before You Install
- Testing & Tooling
- Known Limitations
What This Is
A driver or warehouse operator photographs a DO paper on the Flutter app. The app checks the photo isn't blurry, tags it with GPS, and uploads it. The backend runs the image through a GPU OCR pipeline (deskew → layout detection → text recognition → LLM structuring), cross-checks every item line against a master SKU/store database, and saves the result. The app polls for the result, the operator reviews/corrects it on-device, and can print or export a signed delivery receipt as a PDF.
Two consumers of the same backend exist:
- Flutter mobile app (
lib/) — the primary, field-facing client. - Next.js web pages (
backend/pfm-web-app/src/app/*.tsx) — internal tooling for manual labeling, accuracy comparison, and an OCR "arena" for engine comparison. Not part of the driver-facing product.
Tech Stack
| Layer | Technology | Role |
|---|---|---|
| Mobile client | Flutter (Riverpod, Dio, Hive, go_router) | Camera capture, blur detection, offline-tolerant upload queue, manual correction editor, native PDF/print |
| API gateway | Next.js (App Router, TypeScript) | Auth, upload handling, orchestrates the OCR pipeline, post-processing (regex extraction, fuzzy SKU/store matching), serves internal web tooling |
| Reverse proxy | Nginx | Single entry point (:8000) routing to the gateway, pipeline API, and vLLM server |
| OCR pipeline | PaddleOCR v6 + PP-DocLayoutV3 (FastAPI, GPU) | Auto-deskew/unwarp, layout segmentation, text detection & recognition |
| Structuring LLM | vLLM serving PaddleOCR-VL-1.6-0.9B | Reassembles OCR text fragments into coherent structured text |
| Database | PostgreSQL 15 | Documents, line items, SKU/vendor/customer/store master data |
| Orchestration | Docker Compose (multi-stage GPU Dockerfile) | Runs the whole backend as one stack |
| Product scan (secondary) | YOLO classifier + PaddleOCR (classify_ocr_server.py, :8120) |
Single-product photo → product match + expiry-date extraction |
Architecture
graph LR
subgraph Client["Flutter Mobile App"]
A[Camera + Blur Check]
end
subgraph Edge["Nginx :8000"]
N[Reverse Proxy]
end
subgraph Gateway["Next.js API Gateway :3000"]
G1["/api/v1/documents/upload"]
G2["/api/parse"]
G3["/api/v1/documents (poll/edit)"]
end
subgraph AI["GPU OCR Pipeline"]
P["Pipeline API :8090\n(deskew · layout · OCR)"]
V["vLLM Server :8118\n(PaddleOCR-VL-1.6)"]
end
DB[(PostgreSQL\ndocuments · ocr_items · sku_master · store_master)]
A -- "multipart POST" --> N --> G1
G1 -- "insert parsed=false" --> DB
G1 -- "trigger" --> G2
G2 -- "image" --> P
P <--> V
G2 -- "regex + fuzzy match\nparsed=true" --> DB
A -- "poll every 2s" --> N --> G3 --> DB
How It Works
sequenceDiagram
participant App as Flutter App
participant GW as Next.js Gateway
participant Pipe as Pipeline API + vLLM
participant DB as PostgreSQL
App->>App: Capture photo, check blur (Laplacian variance)
App->>GW: POST /documents/upload (image + GPS)
GW->>DB: Insert document (parsed=false)
GW->>Pipe: Forward image (deskew, layout, OCR)
Pipe-->>GW: Structured markdown text
GW->>GW: Regex extract (PO/SO/DO/date/plate)<br/>Fuzzy-match SKU & store master
GW->>DB: Update document (parsed=true) + items
loop every 2s, up to 2 min
App->>GW: GET /documents
GW-->>App: Parsed result once ready
end
App->>App: Operator reviews & corrects
App->>GW: PUT /documents/:id (final data)
App->>App: Generate & print delivery receipt PDF
For the full field-by-field extraction rules (fused-digit correction, date sanitization, the triple-check SKU matcher, fuzzy store resolution), see docs/workflow_detail_aplikasi.md and docs/regex_rules_example.md.
Repository Structure
.
├── backend/
│ ├── config/ # Pipeline & vLLM YAML configs
│ ├── db/migrations/ # SQL run once by Postgres on a brand-new volume
│ ├── pfm-web-app/ # Next.js API gateway + internal web tooling
│ │ ├── src/app/api/ # All backend routes (upload, parse, documents, auth, arena...)
│ │ ├── src/db/ # DB pool + init.ts (idempotent schema + seed data)
│ │ ├── src/utils/parser.ts # Regex extraction + sanitization rules
│ │ └── scripts/ # Accuracy-check tooling
│ ├── sources/ # Test images, manual labels, SKU/store reference data
│ ├── Dockerfile # Multi-stage GPU build (vllm-server, pipeline-api, pfm-web-app, gradio-ui)
│ ├── nginx.conf
│ └── docker-compose.yml # ⚠️ Legacy standalone copy — see "What to Consider" below
├── lib/ # Flutter app
│ ├── config/ # app_config.dart — theme + API base URL resolution
│ ├── core/ # Dio client, Hive storage, location, router
│ ├── features/ # auth, camera, documents, editor
│ └── models/
├── docs/ # Deep-dive workflow & extraction-rule docs
├── test/ # Flutter widget/unit tests
├── docker-compose.yml # ⭐ Canonical backend stack — run this one
├── docker-compose.demo.yml # Production-mode override (see below)
└── start-dev-tunnel.ps1 # Syncs LAN IP into app_config.dart + starts ngrok
Getting Started
Prerequisites
| Component | Requirement |
|---|---|
| Backend host | NVIDIA GPU, CUDA 12.6+ driver, ~8GB+ VRAM. Tested working on both native Linux and Windows + Docker Desktop with WSL2 GPU passthrough. |
| Docker | Docker Engine/Desktop with the NVIDIA Container Toolkit (docker info should list nvidia under Runtimes) |
| Disk space | 60GB+ free — the pipeline-api and vllm-server images alone are ~30GB each once built, plus model weight caches |
| Flutter | Flutter SDK >=3.2.0 <4.0.0, Android Studio/Xcode for device tooling |
| Node.js | v20+ (optional — only for running the parser unit tests or accuracy tooling outside Docker) |
1. Backend Setup
-
Configure environment variables (from the repo root):
cp backend/.env.example backend/.envSet
CUDA_VISIBLE_DEVICESto your GPU index, andAPP_PORTif8000is taken. -
Start the stack from the repo root (not
backend/— see the gotcha below):docker compose up --buildFirst build pulls/builds ~60GB of GPU images and downloads model weights — expect this to take a long time on the first run. Subsequent starts are fast.
-
Verify it's up:
curl http://localhost:8000/health # pipeline API health curl http://localhost:8000/v1/models # vLLM model list curl -X POST http://localhost:8000/api/v1/auth/login \ -H "Content-Type: application/json" -d '{"username":"admin","password":"password"}'
Database schema and reference data (vendor, customer, a starter SKU catalog) are created automatically on first request — backend/pfm-web-app/src/db/init.ts runs idempotent CREATE TABLE IF NOT EXISTS + seed statements every time the app starts, so there is no manual migration step. backend/db/migrations/*.sql also runs once via Postgres's own docker-entrypoint-initdb.d on a brand-new volume, but init.ts is what you should treat as the source of truth.
2. Frontend Setup
-
Install dependencies:
flutter pub get -
Point the app at your backend.
lib/config/app_config.dartresolves the API URL dynamically at startup: it first tries a public ngrok tunnel, and falls back to a hardcoded LAN URL if that's unreachable. Both need to match your actual machine:./start-dev-tunnel.ps1This detects your current LAN IP, patches
_lanBaseUrlinapp_config.dartfor you, and launchesngrokpointed at the fixed reserved domain in that same file. Run it again any time your IP changes (new network, DHCP renewal, etc.) — a stale IP here is the single most common reason "the app can't log in" after this backend is confirmed healthy. -
Run it:
flutter run -
Build a release APK when you need an installable build instead of a debug session:
flutter build apk --releaseOutput:
build/app/outputs/flutter-apk/app-release.apk. See the signing note below before distributing it.
Running in Demo / Production Mode
The default docker-compose.yml runs the gateway via npm run dev with the source bind-mounted in — good for iterating on parser.ts, but a single dev-server process, not what you want in front of a client with multiple people uploading at once. An opt-in override runs the real production build instead:
docker compose -f docker-compose.yml -f docker-compose.demo.yml up -d --build
This drops the dev bind-mount and runs npm start against the image's own npm run build output. Rebuild (--build) before every demo — this mode does not hot-reload code changes.
What to Consider Before You Install
-
Two
docker-compose.ymlfiles exist — always run from the repo root.backend/docker-compose.ymlis a near-duplicate, standalone copy of the same stack (project nameai-ocr-pfm-2026, originally a separate repo vendored intobackend/). Docker Compose names containers/volumes after the project name declared in whichever compose file you invoke first. If you ever rundocker compose upfrom insidebackend/, you'll get container-name conflicts against anything already started from the root — this isn't hypothetical, it happened during this project's own testing. Pick one (the root file) and stick to it. -
store_masterships empty. Schema is created automatically, but no store data is seeded — fuzzy store-matching (resolveStoreFromText) will not resolve any delivery address until you load real store data into that table yourself.backend/sources/toko_aktif.jsonlooks like the right source for this but isn't wired into an automatic import yet. -
First build is large and slow. The pipeline-api and vllm-server images are ~30GB each with GPU model weights. Budget real time and disk space for the first
docker compose up --build. -
Single GPU pipeline, no horizontal scaling. There's one
pipeline-apiand onevllm-servercontainer. Concurrent uploads queue behind the GPU; this is a real throughput ceiling worth load-testing before a multi-driver demo, not just a single-user smoke test. -
Auth is a demo stub, not real security.
/api/v1/auth/loginonly accepts a single hardcodedadmin/passwordpair and returns a fixed literal token string — no route actually verifies that token server-side, and every API route setsAccess-Control-Allow-Origin: *. Fine for a controlled LAN/demo deployment; do not expose this stack to the open internet as-is. -
The Android release build is debug-signed.
android/app/build.gradle.ktshas a// TODO: Add your own signing configand currently signs release builds with the debug key. Fine for internal install/testing, not for Play Store distribution. -
Mobile connectivity is two independent, manually-synced paths. The ngrok domain in
app_config.dartis fixed/reserved, but the ngrok process isn't started automatically — you (orstart-dev-tunnel.ps1) have to launch it. The LAN IP fallback is hardcoded and will silently go stale the moment your host machine's IP changes. If login/upload fails on a real device, check these two before anything else.
Testing & Tooling
| What | How |
|---|---|
| Parser regex/sanitization rules | npx tsx backend/pfm-web-app/src/utils/parser.test.ts — no Docker needed |
| OCR accuracy regression suite | backend/pfm-web-app/run_batch_test.js and backend/pfm-web-app/scripts/accuracy-check.mts, run against backend/sources/test-images + manual_labels.json |
| Flutter widget/unit tests | flutter test (covers blur detection, camera navigation, editor validation, geotagging, the pending-upload queue, and more — see test/) |
| Static analysis | flutter analyze lib |
Known Limitations
Tracked, not yet fixed — worth knowing before relying on this for unattended field use:
- The pending-upload queue is in-memory only; an OS-level app kill mid-upload loses that document with no trace or retry affordance.
- Item-row writes on document save aren't wrapped in a transaction — a failure mid-write can leave a document with its header saved but item rows silently missing.
See docs/ for the deeper workflow documentation these decisions were audited against.