fhanyuh d3d519dd04 docs(readme): add Prima Fresh Mart logo to the header
Uses the logo already committed at assets/logo.png. Also recolours the accuracy
badge to the brand green so the header reads as one block.
2026-08-27 15:24:45 +07:00

Prima Fresh Mart

Prima Fresh Mart Scanner

Point a phone at a paper Delivery Order. Get a verified database row back.

A self-hosted, GPU-accelerated OCR pipeline that turns store-staff phone photos of Delivery Orders into structured, master-data-validated records — with a Flutter client on one end and Docker + PaddleOCR + vLLM + PostgreSQL on the other.

Flutter Next.js PaddleOCR vLLM PostgreSQL Docker CUDA Accuracy Docs

⚡ Quickstart · 🏗️ Architecture · 🔄 How It Works · 🎯 Accuracy · 🖼️ Screens · ⚠️ Gotchas · 📚 Docs · 🔧 Troubleshooting

📥 Download the handover pack

Download Ringkasan Ruang Lingkup PDF Download Handover Document PDF Download Knowledge Transfer deck


System architecture: phone, Nginx, Next.js gateway, GPU OCR pipeline, PostgreSQL

Every box except the phone and the browser is a container in one Docker Compose stack.


💡 Why This Exists

When a delivery arrives at a Prima Fresh Mart store, the store staff member on duty (petugas toko) — not the driver — has to transcribe a paper Delivery Order into the system. Done by hand it is slow, error-prone, and unverifiable after the fact.

This system replaces that with a photo, and adds three things a human typist cannot cheaply provide:

  1. No cloud OCR. Layout detection, text recognition, and LLM structuring all run on your own GPU. DO photos containing customer and pricing data never leave the premises.
  2. Master-data validation, not just transcription. Every item line is cross-checked against sku_master; every store against store_master. Raw regex off the OCR text scores 67.7% — the correction pass against master data is what lifts it to 89.4%.
  3. A proof-of-receipt trail. GPS tag at capture, blur check before upload, SHA-256 dedup server-side, a named receiver on confirmation, and a printable PDF receipt — all bound to one account per store.

Two clients share one backend:

Client Path Audience
Flutter mobile app lib/ The product. Store staff at each location.
Next.js web pages backend/pfm-web-app/src/app/ Internal tooling only — manual labeling, accuracy comparison, OCR engine "arena". Not store-facing.

⚡ Quickstart

Prerequisites

Component Requirement
GPU host NVIDIA GPU, CUDA 12.6+ driver, ~8 GB+ VRAM. Verified on native Linux and Windows + Docker Desktop with WSL2 GPU passthrough.
Docker Engine/Desktop with NVIDIA Container Toolkit — docker info must list nvidia under Runtimes.
Disk 60 GB+ free. pipeline-api and vllm-server images are ~30 GB each once built, before model weight caches.
Flutter SDK >=3.2.0 <4.0.0, plus Android Studio / Xcode device tooling.
Node.js v20+ — optional, only for running parser tests or accuracy tooling outside Docker.

Important

Run every docker compose command from the repository root, never from inside backend/. A second, legacy compose file lives there under a different project name and will collide with containers already started from the root. See Gotchas.

1 · Backend

# 1. Configure environment (from the repo root)
cp backend/.env.example backend/.env
# Edit backend/.env:
#   APP_PORT=8000              host port exposed via Nginx
#   CUDA_VISIBLE_DEVICES=0     your GPU index — check with nvidia-smi
#   JWT_SECRET=<random-value>  do NOT leave this as "change-me"

# 2. Build and start the whole stack
docker compose up --build

First build pulls/builds ~60 GB of GPU images and downloads model weights — expect a long wait once. Subsequent starts are fast.

Verify it came up:

docker compose ps                                  # every service running / healthy
curl http://localhost:8000/health                  # pipeline API
curl http://localhost:8000/v1/models               # vLLM model list
curl -X POST http://localhost:8000/api/v1/auth/login \
     -H "Content-Type: application/json" \
     -d '{"username":"admin","password":"password"}'

Services come up in stages: db and vllm-server first, then pipeline-api waits until healthy, then pfm-web-app. vllm-server loading model weights onto the GPU can take several minutes on a cold start — that is expected, not a hang.

The schema and reference data are created automatically — backend/pfm-web-app/src/db/init.ts runs idempotent CREATE TABLE IF NOT EXISTS + seed statements on every start, so there is no manual migration step. (backend/db/migrations/*.sql also runs once via Postgres's docker-entrypoint-initdb.d, but only on a brand-new volume; treat init.ts as the source of truth.)

2 · Flutter app

flutter pub get
flutter devices                    # confirm your phone/emulator is visible
flutter run

Point the app at your backend first. lib/config/app_config.dart resolves the API URL at every launch: it tries the LAN address first, then falls back to the ngrok tunnel — whichever answers with genuine backend JSON wins (a dead tunnel's error page is rejected, not accepted).

./start-dev-tunnel.ps1

That script detects your current LAN IP, patches _lanBaseUrl in app_config.dart, and launches ngrok against the reserved domain in that same file — pointed at port 8001, the narrow /api/v1-only door, so the internal web tooling is never exposed publicly.

Running on LAN address to use in app_config.dart
Physical device, same Wi-Fi http://<host LAN IP>:8000/api/v1
Android emulator http://10.0.2.2:8000/api/v1
iOS simulator http://localhost:8000/api/v1

3 · Release APK

flutter build apk --release
# → build/app/outputs/flutter-apk/app-release.apk

Warning

Both backend URLs are compile-time constants. Every time the host LAN IP or ngrok domain changes, the APK must be rebuilt and redistributed to every store device. This is the single most common cause of "the app suddenly can't log in."


🏗️ Architecture

graph LR
    subgraph Client["Flutter Mobile App"]
        A[Camera + Blur Check]
    end

    subgraph Edge["Nginx :8000 / :8001"]
        N[Reverse Proxy]
    end

    subgraph Gateway["Next.js API Gateway :3000"]
        G1["/api/v1/documents/upload"]
        G2["/api/parse"]
        G3["/api/v1/documents (poll/edit)"]
    end

    subgraph AI["GPU OCR Pipeline"]
        P["Pipeline API :8090<br/>deskew · layout · OCR"]
        V["vLLM Server :8118<br/>PaddleOCR-VL-1.6"]
    end

    DB[(PostgreSQL :5432<br/>documents · ocr_items<br/>sku_master · store_master)]
    FS[/backend/uploads/<br/>photo + JSON/]

    A -- "multipart POST" --> N --> G1
    G1 -- "write file" --> FS
    G1 -- "insert parsed=false" --> DB
    G1 -- "trigger" --> G2
    G2 -- "image" --> P
    P <--> V
    G2 -- "regex + fuzzy match<br/>parsed=true" --> DB
    A -- "poll every 2s" --> N --> G3 --> DB

Ports

Only two ports are exposed off the host. Everything else is container-internal.

Port Container Serves Exposure
8000 paddleocr-nginx Front door — full internal surface, including web tooling 🌐 Public
8001 paddleocr-nginx Narrow door — /api/v1/* only, everything else 404 🌐 Public
3000 paddleocr-pfm-web-app Next.js gateway + admin pages 🔒 Internal
8090 paddleocr-pipeline-api PaddleOCR pipeline, /layout-parsing 🔒 Internal
8118 paddleocr-vllm-server vLLM serving PaddleOCR-VL-1.6-0.9B 🔒 Internal
8120 paddleocr-pipeline-api Product-scan classifier, /classify-ocr 🔒 Internal
5432 paddleocr-db PostgreSQL 15, database dopfm 🔒 Internal

Body-size limits are disabled (client_max_body_size 0) because DO photos are large, and proxy timeouts are raised to 300 s because a GPU OCR pass is slow. Both are set in backend/nginx.conf.

Stack

Layer Technology Role
Mobile client Flutter — Riverpod, Dio, Hive, go_router Capture, blur detection, GPS, crash-tolerant upload queue, correction editor, native PDF/print
API gateway Next.js 16 (App Router, TypeScript) Auth, upload, pipeline orchestration, regex extraction, fuzzy SKU/store matching, internal tooling
Reverse proxy Nginx Two server blocks — full surface on :8000, API-only on :8001
OCR pipeline PaddleOCR v6 + PP-DocLayoutV3 (FastAPI, GPU) Auto-deskew/unwarp, layout segmentation, text detection & recognition
Structuring LLM vLLM + PaddleOCR-VL-1.6-0.9B Reassembles OCR fragments into coherent structured text
Database PostgreSQL 15 Documents, line items, SKU / vendor / customer / store master data
Orchestration Docker Compose (multi-stage GPU Dockerfile) The entire backend as one stack
Product scan DINOv2 similarity search (YOLO fallback) + PaddleOCR Single-product photo → SKU via nearest-embedding lookup + expiry-date extraction

🔄 How It Works

Authentication — one account, one store

Accounts are generated, not typed. init.ts seeds one store account per row in store_master (username = store code), plus one admin. The signed JWT carries kodeToko, which is why store name and address never have to be OCR'd off the photo — they come from whoever uploaded it.

Upload and OCR

Three properties worth knowing:

  • Dedup is content-addressed. The server hashes the file (SHA-256) before writing anything. A retried upload after a perceived timeout returns the original document instead of creating a second row or re-running the GPU pass.
  • Photos are files, not blobs. Images land in backend/uploads/; only the filename goes in the database. Back up the database and that folder together — either one alone is useless.
  • New documents are born confirmed = false and stay invisible to GET /api/v1/documents until the staff member saves their corrections via PUT /api/v1/documents/:id. That is the confirmation gate.

The app polls every 2 s for up to 130 attempts (~4.3 minutes) before surfacing a timeout — see lib/features/documents/pending_documents_provider.dart. Server-side, /api/parse is bounded at 210 s, marks the document parsed with "Not Found" placeholders on pipeline failure so it never hangs forever, and records the reason in parse_error when the call itself dies.

For field-by-field extraction rules — fused-digit correction, date sanitization, the triple-check SKU matcher, fuzzy store resolution — see docs/workflow_detail_aplikasi.md and docs/regex_rules_example.md.


🎯 OCR Accuracy

Tracked field-by-field against 37 hand-labeled real DO photos (backend/sources/test-images + manual_labels.json) — not a vague "it works". Latest logged run in backend/sources/accuracy_history.jsonl:

89.4% exact-field-match  ·  up from an 82.2% baseline  ·  target 95%

The interesting part is where the accuracy comes from. Raw regex straight off the OCR text scores only 67.7%; the second correction pass — fuzzy matching against master data, unit standardization, format sanitization — is what closes the gap.

Field Raw regex After sanitize + triple-check What closes the gap
Customer (Kepada Yth) 100% 100% —
Kode Barang (SKU) 98.4% 98.4% Already reliable at the OCR layer
Item count 13.5% 97.3% Table-noise rows filtered by the SKU/unit triple-check
Nama Barang 0% 94.4% Corrected to master sku_master.nama_item on SKU match
Banyak (qty) 77.6% 95.2% Unit standardized from sku_master.jenis_outer
Jumlah (total) 76.8% 94.4% Unit standardized from sku_master.standar_jumlah
No. PO 91.9% 91.9% Fused-digit correction already applied at regex layer
No. DO 91.9% 91.9% —
Tanggal 89.2% 89.2% —
No. SO 86.5% 86.5% —
Store 0% 64.9% Fuzzy token match against store_master
Plat Truk 62.2% 62.2% Mostly genuine OCR misses on the truck line
Alamat 0% 37.8% Canonicalized against customers where phrasing matches
Known gaps toward the 95% target — and why photo SOP matters more than parsing
Field Ceiling Why it is stuck
Alamat 37.8% The ground-truth labels themselves use two different phrasings for the same physical address across photo batches. Fixing this needs the ground truth unified, not more parsing logic.
Plat Truk 62.2% Genuine OCR misses — the truck line is often faint or absent in the photo. Not a parsing bug.
Store 64.9% Bounded by how populated store_master is. See Gotchas.

A real share of the remaining error is capture technique, not code. Many photos in the current test set were shot with the DO paper resting on other papers instead of a plain flat surface. Auto-deskew estimates page tilt from the average angle of detected text blocks — overlapping edges and text from the sheet underneath corrupt that estimate, which is exactly the failure mode the unwarp-retry logic exists for.

In other words these numbers are a floor, not a ceiling. Tightening field SOP — DO paper alone, flat contrasting surface, decent light, squared to the camera — should raise them with no code change at all.

Reproduce it yourself with backend/pfm-web-app/scripts/accuracy-check.mts (npm run accuracy, or --refresh-ocr to force a real pipeline re-run instead of reusing cached OCR output).


🖼️ The App, Screen by Screen

All screenshots below are real captures from a device running against the live backend — nothing is mocked.


Login — one account per store

Capture — DO mode

Blur check before upload

Queue — survives an app kill

Parsed — ready to confirm

Product scan mode

For the complete walkthrough — every page, button, popup, and the algorithms behind them — see screenshots/v2/WORKFLOW.md and the companion deck Prima-Mart-Scanner-Workflow.pptx.


📂 Repository Structure

.
├── lib/                          # ⭐ FLUTTER APP — the store-facing client
│   ├── config/                   #    app_config.dart — theme + API URL resolution
│   ├── core/                     #    Dio client, Hive storage, location, router
│   ├── features/                 #    auth · camera · documents · editor · splash
│   ├── data/                     #    master_sku.dart — SKU list compiled into the APK
│   └── models/
│
├── backend/                      # ⭐ EVERYTHING THAT RUNS ON THE SERVER
│   ├── pfm-web-app/              #    Next.js gateway + internal web tooling
│   │   ├── src/app/api/v1/       #      API consumed by the mobile app
│   │   ├── src/app/api/          #      Classic dev routes (no auth — internal only)
│   │   ├── src/app/admin/        #      Master-data pages
│   │   ├── src/db/init.ts        #      Idempotent schema + account seeding
│   │   ├── src/utils/parser.ts   #      Regex extraction + sanitization rules
│   │   └── scripts/              #      Accuracy-check tooling
│   ├── config/                   #    Pipeline & vLLM YAML configs
│   ├── db/migrations/            #    SQL, run once on a brand-new Postgres volume
│   ├── sources/                  #    🔒 Test images, manual labels, master reference data
│   ├── uploads/                  #    🔒 Incoming DO photos + their OCR JSON
│   ├── nginx.conf                #    Routing for every port
│   ├── Dockerfile                #    Multi-stage GPU build (3 targets)
│   └── docker-compose.yml        #    ⚠️ Legacy standalone copy — do not use
│
├── docs/                         # Deep-dive workflow & extraction-rule docs
│   └── handover/                 # 📚 Handover deck + documents (see below)
├── screenshots/v2/               # Real-device walkthrough + slide deck
├── test/                         # Flutter widget/unit tests
├── docker-compose.yml            # ⭐ Canonical stack — run THIS one, from the repo root
├── docker-compose.demo.yml       # Production-mode override
└── start-dev-tunnel.ps1          # Sync LAN IP into app_config.dart + start ngrok

🚀 Development vs. Demo Mode

Caution

You MUST use the production override for any client demo or field test. The default stack runs the gateway through npm run dev with source bind-mounted for hot-reload — a single dev-server process with a known throughput ceiling that will bottleneck the moment several devices upload at once.

Mode Command Behaviour
Development docker compose up --build npm run dev, source bind-mounted, hot-reload. Edit parser.ts and see it live.
Demo / Production docker compose -f docker-compose.yml -f docker-compose.demo.yml up -d --build npm start against the image's own build output. No hot-reload.

--build is mandatory in demo mode, not optional — the override drops the source mount, so without a rebuild you are running whatever code was baked into the previous image.


⚠️ Gotchas — Read Before You Install

Deployment & configuration
  • Two docker-compose.yml files exist — always run from the repo root. backend/docker-compose.yml is a near-duplicate standalone copy under project name ai-ocr-pfm-2026 (it was a separate repo, vendored in). Compose names containers and volumes after whichever file you invoke, so running it from inside backend/ produces container-name conflicts against anything already up from the root. This is not hypothetical — it happened during this project's own testing.

  • First build is large and slow. ~30 GB per GPU image, plus weights. Budget the disk and the time once.

  • Single GPU pipeline, no horizontal scaling. One pipeline-api, one vllm-server. Concurrent uploads queue behind the GPU. Load-test this before a multi-store rollout — a single-user smoke test will not reveal the ceiling.

  • docker compose down -v destroys the database. The -v flag drops paddleocr_pgdata and the model caches. There is no automatic backup.

Security posture
  • /api/v1/* enforces real auth; the classic dev routes deliberately do not. /api/v1/auth/login checks a bcrypt hash against the accounts table and signs a JWT; /api/v1/documents/* reject missing/invalid tokens with a real 401. The Flutter app only ever uses this surface. The classic routes (/api/upload, /api/parse, /api/history) and the scan-pfm / manual-label pages have no login flow and never will — they are dev tooling. Every route still sets Access-Control-Allow-Origin: *: fine on a controlled LAN, not fine facing the open internet.

  • Seeded passwords reset on every restart. init.ts seeds accounts with ON CONFLICT (username) DO UPDATE SET password, and it runs whenever the gateway first touches the database. Every backend restart therefore resets all account passwords back to the seeded default. There is no password-change flow yet. Fix this before any wide production rollout.

  • JWT_SECRET has an insecure fallback. Leave it unset in backend/.env and the code falls back to a constant written in the source, which would let anyone holding it mint tokens for any store. Set a real random value.

  • The LAN leg is plain HTTP. The ngrok leg is HTTPS; _lanBaseUrl is not. Bearer tokens and DO contents cross the store Wi-Fi unencrypted.

  • The Android release build is debug-signed. android/app/build.gradle.kts still carries // TODO: Add your own signing config. Fine for internal installs, not for Play Store distribution.

Data & master tables
  • store_master ships empty. The schema is created automatically but no store rows are seeded, so resolveStoreFromText resolves nothing until you load real data. backend/sources/toko_aktif.json looks like the right source but is not wired to an automatic import yet.

  • backend/sources/ and backend/uploads/ hold real client data — master SKU/vendor/customer records, genuine DO photographs, hand-labeled ground truth. Do not export, log, or forward them. See the Confidentiality section of backend/CLAUDE.md.

  • Mobile connectivity is two independently-maintained paths. The ngrok domain is reserved but the ngrok process is not auto-started; the LAN IP is hardcoded and goes stale the moment DHCP moves. When a device cannot log in, check these two before anything else.


🧪 Testing & Tooling

Target Command
Parser regex / sanitization rules npx tsx backend/pfm-web-app/src/utils/parser.test.ts — no Docker needed
OCR accuracy regression suite backend/pfm-web-app/scripts/accuracy-check.mts and run_batch_test.js, against backend/sources/test-images + manual_labels.json
Flutter widget / unit tests flutter test — blur detection, camera navigation, editor validation, geotagging, pending-upload queue
Single Flutter test file flutter test test/blur_detector_test.dart
Static analysis flutter analyze lib

📚 Documentation

Document Format What it covers
Ringkasan Ruang Lingkup PDF 20 pp 🇮🇩 Start here. Flow, folder map, ports, how to run each part and from which directory, how login and upload move data, and what is built versus what is not.
Handover Document PDF 42 pp Full handover manual — operational runbook, troubleshooting guide, glossary, screenshot appendix.
Knowledge Transfer Deck PDF 30 slides Walkthrough deck for a live handover session. Diagrams are native shapes, editable in Impress/Draw.
App Walkthrough MD Screen-by-screen tour with the algorithm behind each one, captured against the live backend.
Extraction Workflow MD Field-by-field extraction rules end to end.
Regex Rules MD Worked examples of every pattern and sanitization step.
API Contract Map MD Every endpoint, its payload, and known contract gaps.
Feature List · Iteration Log MD What shipped, and the audit trail behind each change.

📥 Direct downloads

Read-only PDFs, plus the editable LibreOffice sources if you need to change them.

Document Read / print Edit
Ringkasan Ruang Lingkup
🇮🇩 20 pages — start here
Download PDF Download ODT
Handover Document
42 pages — runbook, troubleshooting, glossary
Download PDF Download ODT
Knowledge Transfer Deck
30 slides — for a live session
Download PDF Download ODP

Tip

To grab all three at once without cloning the whole repository (it is large — the photo archives dominate):

git clone --filter=blob:none --sparse https://github.com/DBS-Internship/pfm-ocr.git
cd pfm-ocr && git sparse-checkout set docs/handover

The PDFs are generated, not hand-edited. Regenerate them from docs/handover/src/ — see that folder's README.

Working rules for anyone (human or AI agent) editing this repo live in CLAUDE.md and AGENTS.md for the Flutter app, and separately in backend/CLAUDE.md / backend/AGENTS.md for the backend. They are deliberately not interchangeable — do not apply one set to the other subtree.


🔧 Troubleshooting

The app cannot log in or upload
Symptom Cause & fix
Login times out on a real device The compiled LAN IP is stale or ngrok is not running. Run ./start-dev-tunnel.ps1, then rebuild and reinstall the APK — the URL is a compile-time constant.
Works on emulator, not on phone Emulator needs 10.0.2.2, a physical device needs the host's real LAN IP. They cannot share one value.
401 on every request Token expired (30-day JWT), or the backend restarted and reset the seeded passwords. Log in again.
Login succeeds but uploads hang Check docker compose logs -f pipeline-api and vllm-server. A cold GPU model load takes minutes.
Backend will not start
Symptom Cause & fix
could not select device driver NVIDIA Container Toolkit missing. docker info must list nvidia under Runtimes.
Container name conflicts You ran docker compose from inside backend/. Stop everything, then run from the repo root only.
pfm-web-app never starts It waits for pipeline-api to report healthy, which waits on vllm-server. Watch docker compose logs -f vllm-server — first-time weight download is slow.
Out of disk mid-build Images are ~30 GB each. Free space and rebuild.
OCR results are wrong or empty
Symptom Cause & fix
Store field always empty store_master is unpopulated — see Gotchas.
Item names look like OCR noise The SKU triple-check could not match sku_master. Confirm the SKU catalogue covers those products.
Whole document skewed / garbled The DO was photographed on top of other papers, breaking deskew. Reshoot on a plain flat surface.
Document stuck "processing" Poll gives up after ~4.3 min. Check parse_error on the row and the pipeline-api logs.

📋 Known Limitations

Resolved, kept here for context:

  • Pending-upload queue was in-memory only — now persisted to a Hive box (lib/core/storage/local_storage.dart); an OS-level app kill mid-upload no longer loses the document, the queue reloads and resumes on next launch.
  • Item-row writes on save were not transactional — PUT /documents/:id now runs inside withTransaction.

Still open, worth knowing before unattended field use:

  • The DO review form fails silently on submit when a required field (e.g. driver name) is empty — Form.validate() returns false and the confirm button simply returns, with no toast and no scroll-to-error explaining why nothing happened.
  • The blur check is advisory only — a photo flagged blurry can still be uploaded; the badge does not gate the button.
  • SKUs are validated against a bundled list (lib/data/master_sku.dart) rather than the live /api/v1/master/skus endpoint, so a newly added SKU fails client-side validation until the APK is rebuilt.

A fuller backlog — 51 open items grouped by impact — lives in plans/next-enhancements.md, and is summarized in plain language in the Ringkasan Ruang Lingkup PDF.

S
Description
PFM OCR to scan Delivery Order (DO)
Readme
1.8 GiB
0 Stars 1 Watchers 0 Forks
Languages
TypeScript 42.2%
Dart 26.4%
Python 10.1%
HTML 8.4%
JavaScript 7%
Other 5.8%