Uses the logo already committed at assets/logo.png. Also recolours the accuracy badge to the brand green so the header reads as one block.
Prima Fresh Mart Scanner
Point a phone at a paper Delivery Order. Get a verified database row back.
A self-hosted, GPU-accelerated OCR pipeline that turns store-staff phone photos of Delivery Orders into structured, master-data-validated records — with a Flutter client on one end and Docker + PaddleOCR + vLLM + PostgreSQL on the other.
⚡ Quickstart · 🏗️ Architecture · 🔄 How It Works · 🎯 Accuracy · 🖼️ Screens · ⚠️ Gotchas · 📚 Docs · 🔧 Troubleshooting
📥 Download the handover pack
Every box except the phone and the browser is a container in one Docker Compose stack.
💡 Why This Exists
When a delivery arrives at a Prima Fresh Mart store, the store staff member on duty (petugas toko) — not the driver — has to transcribe a paper Delivery Order into the system. Done by hand it is slow, error-prone, and unverifiable after the fact.
This system replaces that with a photo, and adds three things a human typist cannot cheaply provide:
- No cloud OCR. Layout detection, text recognition, and LLM structuring all run on your own GPU. DO photos containing customer and pricing data never leave the premises.
- Master-data validation, not just transcription. Every item line is cross-checked against
sku_master; every store againststore_master. Raw regex off the OCR text scores 67.7% — the correction pass against master data is what lifts it to 89.4%. - A proof-of-receipt trail. GPS tag at capture, blur check before upload, SHA-256 dedup server-side, a named receiver on confirmation, and a printable PDF receipt — all bound to one account per store.
Two clients share one backend:
| Client | Path | Audience |
|---|---|---|
| Flutter mobile app | lib/ |
The product. Store staff at each location. |
| Next.js web pages | backend/pfm-web-app/src/app/ |
Internal tooling only — manual labeling, accuracy comparison, OCR engine "arena". Not store-facing. |
⚡ Quickstart
Prerequisites
| Component | Requirement |
|---|---|
| GPU host | NVIDIA GPU, CUDA 12.6+ driver, ~8 GB+ VRAM. Verified on native Linux and Windows + Docker Desktop with WSL2 GPU passthrough. |
| Docker | Engine/Desktop with NVIDIA Container Toolkit — docker info must list nvidia under Runtimes. |
| Disk | 60 GB+ free. pipeline-api and vllm-server images are ~30 GB each once built, before model weight caches. |
| Flutter | SDK >=3.2.0 <4.0.0, plus Android Studio / Xcode device tooling. |
| Node.js | v20+ — optional, only for running parser tests or accuracy tooling outside Docker. |
Important
Run every
docker composecommand from the repository root, never from insidebackend/. A second, legacy compose file lives there under a different project name and will collide with containers already started from the root. See Gotchas.
1 · Backend
# 1. Configure environment (from the repo root)
cp backend/.env.example backend/.env
# Edit backend/.env:
# APP_PORT=8000 host port exposed via Nginx
# CUDA_VISIBLE_DEVICES=0 your GPU index — check with nvidia-smi
# JWT_SECRET=<random-value> do NOT leave this as "change-me"
# 2. Build and start the whole stack
docker compose up --build
First build pulls/builds ~60 GB of GPU images and downloads model weights — expect a long wait once. Subsequent starts are fast.
Verify it came up:
docker compose ps # every service running / healthy
curl http://localhost:8000/health # pipeline API
curl http://localhost:8000/v1/models # vLLM model list
curl -X POST http://localhost:8000/api/v1/auth/login \
-H "Content-Type: application/json" \
-d '{"username":"admin","password":"password"}'
Services come up in stages: db and vllm-server first, then pipeline-api waits until healthy, then pfm-web-app. vllm-server loading model weights onto the GPU can take several minutes on a cold start — that is expected, not a hang.
The schema and reference data are created automatically — backend/pfm-web-app/src/db/init.ts runs idempotent CREATE TABLE IF NOT EXISTS + seed statements on every start, so there is no manual migration step. (backend/db/migrations/*.sql also runs once via Postgres's docker-entrypoint-initdb.d, but only on a brand-new volume; treat init.ts as the source of truth.)
2 · Flutter app
flutter pub get
flutter devices # confirm your phone/emulator is visible
flutter run
Point the app at your backend first. lib/config/app_config.dart resolves the API URL at every launch: it tries the LAN address first, then falls back to the ngrok tunnel — whichever answers with genuine backend JSON wins (a dead tunnel's error page is rejected, not accepted).
./start-dev-tunnel.ps1
That script detects your current LAN IP, patches _lanBaseUrl in app_config.dart, and launches ngrok against the reserved domain in that same file — pointed at port 8001, the narrow /api/v1-only door, so the internal web tooling is never exposed publicly.
| Running on | LAN address to use in app_config.dart |
|---|---|
| Physical device, same Wi-Fi | http://<host LAN IP>:8000/api/v1 |
| Android emulator | http://10.0.2.2:8000/api/v1 |
| iOS simulator | http://localhost:8000/api/v1 |
3 · Release APK
flutter build apk --release
# → build/app/outputs/flutter-apk/app-release.apk
Warning
Both backend URLs are compile-time constants. Every time the host LAN IP or ngrok domain changes, the APK must be rebuilt and redistributed to every store device. This is the single most common cause of "the app suddenly can't log in."
🏗️ Architecture
graph LR
subgraph Client["Flutter Mobile App"]
A[Camera + Blur Check]
end
subgraph Edge["Nginx :8000 / :8001"]
N[Reverse Proxy]
end
subgraph Gateway["Next.js API Gateway :3000"]
G1["/api/v1/documents/upload"]
G2["/api/parse"]
G3["/api/v1/documents (poll/edit)"]
end
subgraph AI["GPU OCR Pipeline"]
P["Pipeline API :8090<br/>deskew · layout · OCR"]
V["vLLM Server :8118<br/>PaddleOCR-VL-1.6"]
end
DB[(PostgreSQL :5432<br/>documents · ocr_items<br/>sku_master · store_master)]
FS[/backend/uploads/<br/>photo + JSON/]
A -- "multipart POST" --> N --> G1
G1 -- "write file" --> FS
G1 -- "insert parsed=false" --> DB
G1 -- "trigger" --> G2
G2 -- "image" --> P
P <--> V
G2 -- "regex + fuzzy match<br/>parsed=true" --> DB
A -- "poll every 2s" --> N --> G3 --> DB
Ports
Only two ports are exposed off the host. Everything else is container-internal.
| Port | Container | Serves | Exposure |
|---|---|---|---|
| 8000 | paddleocr-nginx |
Front door — full internal surface, including web tooling | 🌐 Public |
| 8001 | paddleocr-nginx |
Narrow door — /api/v1/* only, everything else 404 |
🌐 Public |
| 3000 | paddleocr-pfm-web-app |
Next.js gateway + admin pages | 🔒 Internal |
| 8090 | paddleocr-pipeline-api |
PaddleOCR pipeline, /layout-parsing |
🔒 Internal |
| 8118 | paddleocr-vllm-server |
vLLM serving PaddleOCR-VL-1.6-0.9B | 🔒 Internal |
| 8120 | paddleocr-pipeline-api |
Product-scan classifier, /classify-ocr |
🔒 Internal |
| 5432 | paddleocr-db |
PostgreSQL 15, database dopfm |
🔒 Internal |
Body-size limits are disabled (client_max_body_size 0) because DO photos are large, and proxy timeouts are raised to 300 s because a GPU OCR pass is slow. Both are set in backend/nginx.conf.
Stack
| Layer | Technology | Role |
|---|---|---|
| Mobile client | Flutter — Riverpod, Dio, Hive, go_router | Capture, blur detection, GPS, crash-tolerant upload queue, correction editor, native PDF/print |
| API gateway | Next.js 16 (App Router, TypeScript) | Auth, upload, pipeline orchestration, regex extraction, fuzzy SKU/store matching, internal tooling |
| Reverse proxy | Nginx | Two server blocks — full surface on :8000, API-only on :8001 |
| OCR pipeline | PaddleOCR v6 + PP-DocLayoutV3 (FastAPI, GPU) | Auto-deskew/unwarp, layout segmentation, text detection & recognition |
| Structuring LLM | vLLM + PaddleOCR-VL-1.6-0.9B | Reassembles OCR fragments into coherent structured text |
| Database | PostgreSQL 15 | Documents, line items, SKU / vendor / customer / store master data |
| Orchestration | Docker Compose (multi-stage GPU Dockerfile) | The entire backend as one stack |
| Product scan | DINOv2 similarity search (YOLO fallback) + PaddleOCR | Single-product photo → SKU via nearest-embedding lookup + expiry-date extraction |
🔄 How It Works
Authentication — one account, one store
Accounts are generated, not typed. init.ts seeds one store account per row in store_master (username = store code), plus one admin. The signed JWT carries kodeToko, which is why store name and address never have to be OCR'd off the photo — they come from whoever uploaded it.
Upload and OCR
Three properties worth knowing:
- Dedup is content-addressed. The server hashes the file (SHA-256) before writing anything. A retried upload after a perceived timeout returns the original document instead of creating a second row or re-running the GPU pass.
- Photos are files, not blobs. Images land in
backend/uploads/; only the filename goes in the database. Back up the database and that folder together — either one alone is useless. - New documents are born
confirmed = falseand stay invisible toGET /api/v1/documentsuntil the staff member saves their corrections viaPUT /api/v1/documents/:id. That is the confirmation gate.
The app polls every 2 s for up to 130 attempts (~4.3 minutes) before surfacing a timeout — see lib/features/documents/pending_documents_provider.dart. Server-side, /api/parse is bounded at 210 s, marks the document parsed with "Not Found" placeholders on pipeline failure so it never hangs forever, and records the reason in parse_error when the call itself dies.
For field-by-field extraction rules — fused-digit correction, date sanitization, the triple-check SKU matcher, fuzzy store resolution — see docs/workflow_detail_aplikasi.md and docs/regex_rules_example.md.
🎯 OCR Accuracy
Tracked field-by-field against 37 hand-labeled real DO photos (backend/sources/test-images + manual_labels.json) — not a vague "it works". Latest logged run in backend/sources/accuracy_history.jsonl:
89.4% exact-field-match · up from an 82.2% baseline · target 95%
The interesting part is where the accuracy comes from. Raw regex straight off the OCR text scores only 67.7%; the second correction pass — fuzzy matching against master data, unit standardization, format sanitization — is what closes the gap.
| Field | Raw regex | After sanitize + triple-check | What closes the gap |
|---|---|---|---|
| Customer (Kepada Yth) | 100% | 100% | — |
| Kode Barang (SKU) | 98.4% | 98.4% | Already reliable at the OCR layer |
| Item count | 13.5% | 97.3% | Table-noise rows filtered by the SKU/unit triple-check |
| Nama Barang | 0% | 94.4% | Corrected to master sku_master.nama_item on SKU match |
| Banyak (qty) | 77.6% | 95.2% | Unit standardized from sku_master.jenis_outer |
| Jumlah (total) | 76.8% | 94.4% | Unit standardized from sku_master.standar_jumlah |
| No. PO | 91.9% | 91.9% | Fused-digit correction already applied at regex layer |
| No. DO | 91.9% | 91.9% | — |
| Tanggal | 89.2% | 89.2% | — |
| No. SO | 86.5% | 86.5% | — |
| Store | 0% | 64.9% | Fuzzy token match against store_master |
| Plat Truk | 62.2% | 62.2% | Mostly genuine OCR misses on the truck line |
| Alamat | 0% | 37.8% | Canonicalized against customers where phrasing matches |
Known gaps toward the 95% target — and why photo SOP matters more than parsing
| Field | Ceiling | Why it is stuck |
|---|---|---|
| Alamat | 37.8% | The ground-truth labels themselves use two different phrasings for the same physical address across photo batches. Fixing this needs the ground truth unified, not more parsing logic. |
| Plat Truk | 62.2% | Genuine OCR misses — the truck line is often faint or absent in the photo. Not a parsing bug. |
| Store | 64.9% | Bounded by how populated store_master is. See Gotchas. |
A real share of the remaining error is capture technique, not code. Many photos in the current test set were shot with the DO paper resting on other papers instead of a plain flat surface. Auto-deskew estimates page tilt from the average angle of detected text blocks — overlapping edges and text from the sheet underneath corrupt that estimate, which is exactly the failure mode the unwarp-retry logic exists for.
In other words these numbers are a floor, not a ceiling. Tightening field SOP — DO paper alone, flat contrasting surface, decent light, squared to the camera — should raise them with no code change at all.
Reproduce it yourself with backend/pfm-web-app/scripts/accuracy-check.mts (npm run accuracy, or --refresh-ocr to force a real pipeline re-run instead of reusing cached OCR output).
🖼️ The App, Screen by Screen
All screenshots below are real captures from a device running against the live backend — nothing is mocked.
![]() Login — one account per store |
![]() Capture — DO mode |
![]() Blur check before upload |
![]() Queue — survives an app kill |
![]() Parsed — ready to confirm |
![]() Product scan mode |
For the complete walkthrough — every page, button, popup, and the algorithms behind them — see screenshots/v2/WORKFLOW.md and the companion deck Prima-Mart-Scanner-Workflow.pptx.
📂 Repository Structure
.
├── lib/ # ⭐ FLUTTER APP — the store-facing client
│ ├── config/ # app_config.dart — theme + API URL resolution
│ ├── core/ # Dio client, Hive storage, location, router
│ ├── features/ # auth · camera · documents · editor · splash
│ ├── data/ # master_sku.dart — SKU list compiled into the APK
│ └── models/
│
├── backend/ # ⭐ EVERYTHING THAT RUNS ON THE SERVER
│ ├── pfm-web-app/ # Next.js gateway + internal web tooling
│ │ ├── src/app/api/v1/ # API consumed by the mobile app
│ │ ├── src/app/api/ # Classic dev routes (no auth — internal only)
│ │ ├── src/app/admin/ # Master-data pages
│ │ ├── src/db/init.ts # Idempotent schema + account seeding
│ │ ├── src/utils/parser.ts # Regex extraction + sanitization rules
│ │ └── scripts/ # Accuracy-check tooling
│ ├── config/ # Pipeline & vLLM YAML configs
│ ├── db/migrations/ # SQL, run once on a brand-new Postgres volume
│ ├── sources/ # 🔒 Test images, manual labels, master reference data
│ ├── uploads/ # 🔒 Incoming DO photos + their OCR JSON
│ ├── nginx.conf # Routing for every port
│ ├── Dockerfile # Multi-stage GPU build (3 targets)
│ └── docker-compose.yml # ⚠️ Legacy standalone copy — do not use
│
├── docs/ # Deep-dive workflow & extraction-rule docs
│ └── handover/ # 📚 Handover deck + documents (see below)
├── screenshots/v2/ # Real-device walkthrough + slide deck
├── test/ # Flutter widget/unit tests
├── docker-compose.yml # ⭐ Canonical stack — run THIS one, from the repo root
├── docker-compose.demo.yml # Production-mode override
└── start-dev-tunnel.ps1 # Sync LAN IP into app_config.dart + start ngrok
🚀 Development vs. Demo Mode
Caution
You MUST use the production override for any client demo or field test. The default stack runs the gateway through
npm run devwith source bind-mounted for hot-reload — a single dev-server process with a known throughput ceiling that will bottleneck the moment several devices upload at once.
| Mode | Command | Behaviour |
|---|---|---|
| Development | docker compose up --build |
npm run dev, source bind-mounted, hot-reload. Edit parser.ts and see it live. |
| Demo / Production | docker compose -f docker-compose.yml -f docker-compose.demo.yml up -d --build |
npm start against the image's own build output. No hot-reload. |
--build is mandatory in demo mode, not optional — the override drops the source mount, so without a rebuild you are running whatever code was baked into the previous image.
⚠️ Gotchas — Read Before You Install
Deployment & configuration
-
Two
docker-compose.ymlfiles exist — always run from the repo root.backend/docker-compose.ymlis a near-duplicate standalone copy under project nameai-ocr-pfm-2026(it was a separate repo, vendored in). Compose names containers and volumes after whichever file you invoke, so running it from insidebackend/produces container-name conflicts against anything already up from the root. This is not hypothetical — it happened during this project's own testing. -
First build is large and slow. ~30 GB per GPU image, plus weights. Budget the disk and the time once.
-
Single GPU pipeline, no horizontal scaling. One
pipeline-api, onevllm-server. Concurrent uploads queue behind the GPU. Load-test this before a multi-store rollout — a single-user smoke test will not reveal the ceiling. -
docker compose down -vdestroys the database. The-vflag dropspaddleocr_pgdataand the model caches. There is no automatic backup.
Security posture
-
/api/v1/*enforces real auth; the classic dev routes deliberately do not./api/v1/auth/loginchecks a bcrypt hash against theaccountstable and signs a JWT;/api/v1/documents/*reject missing/invalid tokens with a real401. The Flutter app only ever uses this surface. The classic routes (/api/upload,/api/parse,/api/history) and thescan-pfm/manual-labelpages have no login flow and never will — they are dev tooling. Every route still setsAccess-Control-Allow-Origin: *: fine on a controlled LAN, not fine facing the open internet. -
Seeded passwords reset on every restart.
init.tsseeds accounts withON CONFLICT (username) DO UPDATE SET password, and it runs whenever the gateway first touches the database. Every backend restart therefore resets all account passwords back to the seeded default. There is no password-change flow yet. Fix this before any wide production rollout. -
JWT_SECREThas an insecure fallback. Leave it unset inbackend/.envand the code falls back to a constant written in the source, which would let anyone holding it mint tokens for any store. Set a real random value. -
The LAN leg is plain HTTP. The ngrok leg is HTTPS;
_lanBaseUrlis not. Bearer tokens and DO contents cross the store Wi-Fi unencrypted. -
The Android release build is debug-signed.
android/app/build.gradle.ktsstill carries// TODO: Add your own signing config. Fine for internal installs, not for Play Store distribution.
Data & master tables
-
store_masterships empty. The schema is created automatically but no store rows are seeded, soresolveStoreFromTextresolves nothing until you load real data.backend/sources/toko_aktif.jsonlooks like the right source but is not wired to an automatic import yet. -
backend/sources/andbackend/uploads/hold real client data — master SKU/vendor/customer records, genuine DO photographs, hand-labeled ground truth. Do not export, log, or forward them. See the Confidentiality section ofbackend/CLAUDE.md. -
Mobile connectivity is two independently-maintained paths. The ngrok domain is reserved but the ngrok process is not auto-started; the LAN IP is hardcoded and goes stale the moment DHCP moves. When a device cannot log in, check these two before anything else.
🧪 Testing & Tooling
| Target | Command |
|---|---|
| Parser regex / sanitization rules | npx tsx backend/pfm-web-app/src/utils/parser.test.ts — no Docker needed |
| OCR accuracy regression suite | backend/pfm-web-app/scripts/accuracy-check.mts and run_batch_test.js, against backend/sources/test-images + manual_labels.json |
| Flutter widget / unit tests | flutter test — blur detection, camera navigation, editor validation, geotagging, pending-upload queue |
| Single Flutter test file | flutter test test/blur_detector_test.dart |
| Static analysis | flutter analyze lib |
📚 Documentation
| Document | Format | What it covers |
|---|---|---|
| Ringkasan Ruang Lingkup | PDF 20 pp |
🇮🇩 Start here. Flow, folder map, ports, how to run each part and from which directory, how login and upload move data, and what is built versus what is not. |
| Handover Document | PDF 42 pp |
Full handover manual — operational runbook, troubleshooting guide, glossary, screenshot appendix. |
| Knowledge Transfer Deck | PDF 30 slides |
Walkthrough deck for a live handover session. Diagrams are native shapes, editable in Impress/Draw. |
| App Walkthrough | MD |
Screen-by-screen tour with the algorithm behind each one, captured against the live backend. |
| Extraction Workflow | MD |
Field-by-field extraction rules end to end. |
| Regex Rules | MD |
Worked examples of every pattern and sanitization step. |
| API Contract Map | MD |
Every endpoint, its payload, and known contract gaps. |
| Feature List · Iteration Log | MD |
What shipped, and the audit trail behind each change. |
📥 Direct downloads
Read-only PDFs, plus the editable LibreOffice sources if you need to change them.
| Document | Read / print | Edit |
|---|---|---|
| Ringkasan Ruang Lingkup 🇮🇩 20 pages — start here |
||
| Handover Document 42 pages — runbook, troubleshooting, glossary |
||
| Knowledge Transfer Deck 30 slides — for a live session |
Tip
To grab all three at once without cloning the whole repository (it is large — the photo archives dominate):
git clone --filter=blob:none --sparse https://github.com/DBS-Internship/pfm-ocr.git cd pfm-ocr && git sparse-checkout set docs/handover
The PDFs are generated, not hand-edited. Regenerate them from docs/handover/src/ — see that folder's README.
Working rules for anyone (human or AI agent) editing this repo live in CLAUDE.md and AGENTS.md for the Flutter app, and separately in backend/CLAUDE.md / backend/AGENTS.md for the backend. They are deliberately not interchangeable — do not apply one set to the other subtree.
🔧 Troubleshooting
The app cannot log in or upload
| Symptom | Cause & fix |
|---|---|
| Login times out on a real device | The compiled LAN IP is stale or ngrok is not running. Run ./start-dev-tunnel.ps1, then rebuild and reinstall the APK — the URL is a compile-time constant. |
| Works on emulator, not on phone | Emulator needs 10.0.2.2, a physical device needs the host's real LAN IP. They cannot share one value. |
401 on every request |
Token expired (30-day JWT), or the backend restarted and reset the seeded passwords. Log in again. |
| Login succeeds but uploads hang | Check docker compose logs -f pipeline-api and vllm-server. A cold GPU model load takes minutes. |
Backend will not start
| Symptom | Cause & fix |
|---|---|
could not select device driver |
NVIDIA Container Toolkit missing. docker info must list nvidia under Runtimes. |
| Container name conflicts | You ran docker compose from inside backend/. Stop everything, then run from the repo root only. |
pfm-web-app never starts |
It waits for pipeline-api to report healthy, which waits on vllm-server. Watch docker compose logs -f vllm-server — first-time weight download is slow. |
| Out of disk mid-build | Images are ~30 GB each. Free space and rebuild. |
OCR results are wrong or empty
| Symptom | Cause & fix |
|---|---|
| Store field always empty | store_master is unpopulated — see Gotchas. |
| Item names look like OCR noise | The SKU triple-check could not match sku_master. Confirm the SKU catalogue covers those products. |
| Whole document skewed / garbled | The DO was photographed on top of other papers, breaking deskew. Reshoot on a plain flat surface. |
| Document stuck "processing" | Poll gives up after ~4.3 min. Check parse_error on the row and the pipeline-api logs. |
📋 Known Limitations
Resolved, kept here for context:
Pending-upload queue was in-memory only— now persisted to a Hive box (lib/core/storage/local_storage.dart); an OS-level app kill mid-upload no longer loses the document, the queue reloads and resumes on next launch.Item-row writes on save were not transactional—PUT /documents/:idnow runs insidewithTransaction.
Still open, worth knowing before unattended field use:
- The DO review form fails silently on submit when a required field (e.g. driver name) is empty —
Form.validate()returns false and the confirm button simply returns, with no toast and no scroll-to-error explaining why nothing happened. - The blur check is advisory only — a photo flagged blurry can still be uploaded; the badge does not gate the button.
- SKUs are validated against a bundled list (
lib/data/master_sku.dart) rather than the live/api/v1/master/skusendpoint, so a newly added SKU fails client-side validation until the APK is rebuilt.
A fuller backlog — 51 open items grouped by impact — lives in plans/next-enhancements.md, and is summarized in plain language in the Ringkasan Ruang Lingkup PDF.





