Rafhan Mazaya FathurrahmanandClaude Sonnet 5 3df9f6ec5d Add OCR Accuracy section to README with field-level breakdown
Documents the 89.4% overall accuracy (up from 82.2% baseline, 95% target),
grounded in the actual logged run in accuracy_history.jsonl rather than a
vague claim, plus which fields the sanitize/triple-check layer rescues over
raw regex. Notes that current gaps are partly a photo-capture SOP floor
(DO paper photographed on top of other documents confuses auto-deskew),
not just a parsing/OCR ceiling.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eRAsLqN9Sg1b9YPz22Lxz
2026-07-05 00:35:16 +07:00

Prima Fresh Mart Scanner (app-pfm-ocr-v2)

An on-premise, GPU-accelerated OCR system that turns a driver's phone photo of a Delivery Order (DO) into structured, database-backed data — PO/SO/DO numbers, dates, store, and item lines — with a Flutter mobile client on one end and a Dockerized AI pipeline on the other.

No cloud OCR API is used. Everything (layout detection, text recognition, LLM-assisted structuring) runs on your own GPU.


Table of Contents


What This Is

A driver or warehouse operator photographs a DO paper on the Flutter app. The app checks the photo isn't blurry, tags it with GPS, and uploads it. The backend runs the image through a GPU OCR pipeline (deskew → layout detection → text recognition → LLM structuring), cross-checks every item line against a master SKU/store database, and saves the result. The app polls for the result, the operator reviews/corrects it on-device, and can print or export a signed delivery receipt as a PDF.

Two consumers of the same backend exist:

  • Flutter mobile app (lib/) — the primary, field-facing client.
  • Next.js web pages (backend/pfm-web-app/src/app/*.tsx) — internal tooling for manual labeling, accuracy comparison, and an OCR "arena" for engine comparison. Not part of the driver-facing product.

Tech Stack

Layer Technology Role
Mobile client Flutter (Riverpod, Dio, Hive, go_router) Camera capture, blur detection, offline-tolerant upload queue, manual correction editor, native PDF/print
API gateway Next.js (App Router, TypeScript) Auth, upload handling, orchestrates the OCR pipeline, post-processing (regex extraction, fuzzy SKU/store matching), serves internal web tooling
Reverse proxy Nginx Single entry point (:8000) routing to the gateway, pipeline API, and vLLM server
OCR pipeline PaddleOCR v6 + PP-DocLayoutV3 (FastAPI, GPU) Auto-deskew/unwarp, layout segmentation, text detection & recognition
Structuring LLM vLLM serving PaddleOCR-VL-1.6-0.9B Reassembles OCR text fragments into coherent structured text
Database PostgreSQL 15 Documents, line items, SKU/vendor/customer/store master data
Orchestration Docker Compose (multi-stage GPU Dockerfile) Runs the whole backend as one stack
Product scan (secondary) YOLO classifier + PaddleOCR (classify_ocr_server.py, :8120) Single-product photo → product match + expiry-date extraction

Architecture

graph LR
    subgraph Client["Flutter Mobile App"]
        A[Camera + Blur Check]
    end

    subgraph Edge["Nginx :8000"]
        N[Reverse Proxy]
    end

    subgraph Gateway["Next.js API Gateway :3000"]
        G1["/api/v1/documents/upload"]
        G2["/api/parse"]
        G3["/api/v1/documents (poll/edit)"]
    end

    subgraph AI["GPU OCR Pipeline"]
        P["Pipeline API :8090\n(deskew · layout · OCR)"]
        V["vLLM Server :8118\n(PaddleOCR-VL-1.6)"]
    end

    DB[(PostgreSQL\ndocuments · ocr_items · sku_master · store_master)]

    A -- "multipart POST" --> N --> G1
    G1 -- "insert parsed=false" --> DB
    G1 -- "trigger" --> G2
    G2 -- "image" --> P
    P <--> V
    G2 -- "regex + fuzzy match\nparsed=true" --> DB
    A -- "poll every 2s" --> N --> G3 --> DB

How It Works

sequenceDiagram
    participant App as Flutter App
    participant GW as Next.js Gateway
    participant Pipe as Pipeline API + vLLM
    participant DB as PostgreSQL

    App->>App: Capture photo, check blur (Laplacian variance)
    App->>GW: POST /documents/upload (image + GPS)
    GW->>DB: Insert document (parsed=false)
    GW->>Pipe: Forward image (deskew, layout, OCR)
    Pipe-->>GW: Structured markdown text
    GW->>GW: Regex extract (PO/SO/DO/date/plate)<br/>Fuzzy-match SKU & store master
    GW->>DB: Update document (parsed=true) + items
    loop every 2s, up to 2 min
        App->>GW: GET /documents
        GW-->>App: Parsed result once ready
    end
    App->>App: Operator reviews & corrects
    App->>GW: PUT /documents/:id (final data)
    App->>App: Generate & print delivery receipt PDF

For the full field-by-field extraction rules (fused-digit correction, date sanitization, the triple-check SKU matcher, fuzzy store resolution), see docs/workflow_detail_aplikasi.md and docs/regex_rules_example.md.


OCR Accuracy

Accuracy is tracked field-by-field against 37 hand-labeled real DO photos (backend/sources/test-images + manual_labels.json), not a single vague "it works" claim. The latest logged run (backend/sources/accuracy_history.jsonl):

Overall: 89.4% exact-field-match — up from an 82.2% baseline when this tracking tool was first built, against a 95% target.

The bigger story is where that accuracy comes from. Raw regex extraction straight off the OCR text is only 67.7% — the gain to 89.4% comes from a second correction pass: fuzzy SKU/store matching against master data, unit standardization, and format sanitization.

Field Raw regex After sanitize + triple-check What closes the gap
Customer (Kepada Yth) 100% 100% —
Kode Barang (SKU) 98.4% 98.4% Already reliable at the OCR layer
Item count 13.5% 97.3% Table-noise rows filtered by the SKU/unit triple-check
Nama Barang 0% 94.4% Corrected to master sku_master.nama_item on SKU match
Banyak (qty) 77.6% 95.2% Unit standardized from sku_master.jenis_outer
Jumlah (total) 76.8% 94.4% Unit standardized from sku_master.standar_jumlah
No. PO 91.9% 91.9% Fused-digit correction already applied at regex layer
No. DO 91.9% 91.9% —
Tanggal 89.2% 89.2% —
No. SO 86.5% 86.5% —
Store 0% 64.9% Fuzzy token match against store_master
Plat Truk 62.2% 62.2% Mostly genuine OCR misses on the truck-line, not a parsing gap
Alamat 0% 37.8% Canonicalized against customers table where phrasing matches

Known gaps toward the 95% target:

  • Alamat (37.8%, capped) — the ground-truth labels themselves use two different phrasings for the same physical address across photo batches; closing this needs the ground truth unified, not more parsing logic.
  • Plat (62.2%) — mostly genuine OCR misses (the truck line is often faint or absent in the photo) rather than a fixable parsing bug.
  • Store (64.9%) — bounded by the same store_master table noted in What to Consider — accuracy against it can only improve as far as that master data is populated.

A real factor behind these gaps is photo capture SOP, not just parsing or OCR quality. A meaningful share of the current test set was photographed with the DO paper placed on top of other papers/documents rather than a plain, flat surface. That confuses auto-deskew: the pipeline estimates page tilt from the average angle of detected text blocks, and overlapping paper edges/text from the sheet underneath make that estimate unreliable, which is exactly the failure mode behind the unwarp-retry logic described above. In practice this means accuracy here is a floor, not a ceiling — tightening the field SOP (DO paper alone, on a flat contrasting surface, reasonably well-lit and squared to the camera) should raise these numbers without any further code changes.

Reproduce this yourself with backend/pfm-web-app/scripts/accuracy-check.mts (npm run accuracy, or --refresh-ocr to force a real pipeline re-run instead of using cached OCR results) — see Testing & Tooling.


Repository Structure

.
├── backend/
│   ├── config/                 # Pipeline & vLLM YAML configs
│   ├── db/migrations/          # SQL run once by Postgres on a brand-new volume
│   ├── pfm-web-app/            # Next.js API gateway + internal web tooling
│   │   ├── src/app/api/        # All backend routes (upload, parse, documents, auth, arena...)
│   │   ├── src/db/             # DB pool + init.ts (idempotent schema + seed data)
│   │   ├── src/utils/parser.ts # Regex extraction + sanitization rules
│   │   └── scripts/            # Accuracy-check tooling
│   ├── sources/                # Test images, manual labels, SKU/store reference data
│   ├── Dockerfile              # Multi-stage GPU build (vllm-server, pipeline-api, pfm-web-app, gradio-ui)
│   ├── nginx.conf
│   └── docker-compose.yml      # ⚠️ Legacy standalone copy — see "What to Consider" below
├── lib/                         # Flutter app
│   ├── config/                 # app_config.dart — theme + API base URL resolution
│   ├── core/                   # Dio client, Hive storage, location, router
│   ├── features/               # auth, camera, documents, editor
│   └── models/
├── docs/                        # Deep-dive workflow & extraction-rule docs
├── test/                        # Flutter widget/unit tests
├── docker-compose.yml           # ⭐ Canonical backend stack — run this one
├── docker-compose.demo.yml      # Production-mode override (see below)
└── start-dev-tunnel.ps1         # Syncs LAN IP into app_config.dart + starts ngrok

Getting Started

Prerequisites

Component Requirement
Backend host NVIDIA GPU, CUDA 12.6+ driver, ~8GB+ VRAM. Tested working on both native Linux and Windows + Docker Desktop with WSL2 GPU passthrough.
Docker Docker Engine/Desktop with the NVIDIA Container Toolkit (docker info should list nvidia under Runtimes)
Disk space 60GB+ free — the pipeline-api and vllm-server images alone are ~30GB each once built, plus model weight caches
Flutter Flutter SDK >=3.2.0 <4.0.0, Android Studio/Xcode for device tooling
Node.js v20+ (optional — only for running the parser unit tests or accuracy tooling outside Docker)

1. Backend Setup

  1. Configure environment variables (from the repo root):

    cp backend/.env.example backend/.env
    

    Set CUDA_VISIBLE_DEVICES to your GPU index, and APP_PORT if 8000 is taken.

  2. Start the stack from the repo root (not backend/ — see the gotcha below):

    docker compose up --build
    

    First build pulls/builds ~60GB of GPU images and downloads model weights — expect this to take a long time on the first run. Subsequent starts are fast.

  3. Verify it's up:

    curl http://localhost:8000/health              # pipeline API health
    curl http://localhost:8000/v1/models            # vLLM model list
    curl -X POST http://localhost:8000/api/v1/auth/login \
         -H "Content-Type: application/json" -d '{"username":"admin","password":"password"}'
    

Database schema and reference data (vendor, customer, a starter SKU catalog) are created automatically on first request — backend/pfm-web-app/src/db/init.ts runs idempotent CREATE TABLE IF NOT EXISTS + seed statements every time the app starts, so there is no manual migration step. backend/db/migrations/*.sql also runs once via Postgres's own docker-entrypoint-initdb.d on a brand-new volume, but init.ts is what you should treat as the source of truth.

2. Frontend Setup

  1. Install dependencies:

    flutter pub get
    
  2. Point the app at your backend. lib/config/app_config.dart resolves the API URL dynamically at startup: it first tries a public ngrok tunnel, and falls back to a hardcoded LAN URL if that's unreachable. Both need to match your actual machine:

    ./start-dev-tunnel.ps1
    

    This detects your current LAN IP, patches _lanBaseUrl in app_config.dart for you, and launches ngrok pointed at the fixed reserved domain in that same file. Run it again any time your IP changes (new network, DHCP renewal, etc.) — a stale IP here is the single most common reason "the app can't log in" after this backend is confirmed healthy.

  3. Run it:

    flutter run
    
  4. Build a release APK when you need an installable build instead of a debug session:

    flutter build apk --release
    

    Output: build/app/outputs/flutter-apk/app-release.apk. See the signing note below before distributing it.


Running in Demo / Production Mode

The default docker-compose.yml runs the gateway via npm run dev with the source bind-mounted in — good for iterating on parser.ts, but a single dev-server process, not what you want in front of a client with multiple people uploading at once. An opt-in override runs the real production build instead:

docker compose -f docker-compose.yml -f docker-compose.demo.yml up -d --build

This drops the dev bind-mount and runs npm start against the image's own npm run build output. Rebuild (--build) before every demo — this mode does not hot-reload code changes.


What to Consider Before You Install

  • Two docker-compose.yml files exist — always run from the repo root. backend/docker-compose.yml is a near-duplicate, standalone copy of the same stack (project name ai-ocr-pfm-2026, originally a separate repo vendored into backend/). Docker Compose names containers/volumes after the project name declared in whichever compose file you invoke first. If you ever run docker compose up from inside backend/, you'll get container-name conflicts against anything already started from the root — this isn't hypothetical, it happened during this project's own testing. Pick one (the root file) and stick to it.

  • store_master ships empty. Schema is created automatically, but no store data is seeded — fuzzy store-matching (resolveStoreFromText) will not resolve any delivery address until you load real store data into that table yourself. backend/sources/toko_aktif.json looks like the right source for this but isn't wired into an automatic import yet.

  • First build is large and slow. The pipeline-api and vllm-server images are ~30GB each with GPU model weights. Budget real time and disk space for the first docker compose up --build.

  • Single GPU pipeline, no horizontal scaling. There's one pipeline-api and one vllm-server container. Concurrent uploads queue behind the GPU; this is a real throughput ceiling worth load-testing before a multi-driver demo, not just a single-user smoke test.

  • Auth is a demo stub, not real security. /api/v1/auth/login only accepts a single hardcoded admin/password pair and returns a fixed literal token string — no route actually verifies that token server-side, and every API route sets Access-Control-Allow-Origin: *. Fine for a controlled LAN/demo deployment; do not expose this stack to the open internet as-is.

  • The Android release build is debug-signed. android/app/build.gradle.kts has a // TODO: Add your own signing config and currently signs release builds with the debug key. Fine for internal install/testing, not for Play Store distribution.

  • Mobile connectivity is two independent, manually-synced paths. The ngrok domain in app_config.dart is fixed/reserved, but the ngrok process isn't started automatically — you (or start-dev-tunnel.ps1) have to launch it. The LAN IP fallback is hardcoded and will silently go stale the moment your host machine's IP changes. If login/upload fails on a real device, check these two before anything else.


Testing & Tooling

What How
Parser regex/sanitization rules npx tsx backend/pfm-web-app/src/utils/parser.test.ts — no Docker needed
OCR accuracy regression suite backend/pfm-web-app/run_batch_test.js and backend/pfm-web-app/scripts/accuracy-check.mts, run against backend/sources/test-images + manual_labels.json
Flutter widget/unit tests flutter test (covers blur detection, camera navigation, editor validation, geotagging, the pending-upload queue, and more — see test/)
Static analysis flutter analyze lib

Known Limitations

Tracked, not yet fixed — worth knowing before relying on this for unattended field use:

  • The pending-upload queue is in-memory only; an OS-level app kill mid-upload loses that document with no trace or retry affordance.
  • Item-row writes on document save aren't wrapped in a transaction — a failure mid-write can leave a document with its header saved but item rows silently missing.

See docs/ for the deeper workflow documentation these decisions were audited against.

S
Description
PFM OCR to scan Delivery Order (DO)
Readme
1.8 GiB
0 Stars 1 Watchers 0 Forks
Languages
TypeScript 42.2%
Dart 26.4%
Python 10.1%
HTML 8.4%
JavaScript 7%
Other 5.8%