Rafhan Mazaya FathurrahmanandClaude Fable 5 e76ccb60a6 feat(backend): scan-product accuracy 66.2% -> 79.7% + frozen validation benchmark
Accuracy work on the 79-image product-scan validation set (user goal: 90%):
- classify_ocr_server.py: 0/90/180/270-degree expiry-date search (stops at
  first hit, 0-degree fallback); classification decoupled onto the upright
  image (rotated frames regressed DINOv2 -6pts until this); cross-line date
  stitching; tiled full-res OCR pass (defeats the 4000px downscale that
  killed small inkjet dates); VL-pipeline expiry fallback with
  keyword-anchored anti-hallucination guard; VL text lines merged into
  text_lines + VL SKU retry. Visualization endpoints removed entirely
  (Visual/Spotting grids - unused by frontend, 3x per-scan GPU cost).
- product-scan.ts: coverage-normalized OCR-evidence re-ranking of DINOv2
  top-K (tuned offline: +8/-0 on top-1 misses), re-ranked class mapped to
  sku_master by SKU prefix; classifier timeout 90s->240s for fallback paths.
- Frozen benchmark: product-test-images-fixed/ (79 renamed images) +
  freeze/seed/build-undetected/capture/experiment scripts; labels trimmed to
  the 79 validation entries (training rows kept in .bak-with-training);
  5 TRAINED-ON SKUs replaced with fresh held-out photos.
- manual-label-scan page: shows last batch-test AI prediction under every
  field by default (new /api/product-scan-results); serves the fixed folder;
  fixed total hydration failure via allowedDevOrigins 127.0.0.1.
- Measured (all-79, zero failures): sku/name 87.3%, expiry 64.6%, overall
  79.7%. Tiles/VL-evidence/VL-SKU deployed but not yet batch-measured.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gr6HH7JrdsXX8AARejQboM
2026-07-14 19:55:17 +07:00

Prima Fresh Mart Scanner (app-pfm-ocr-v2)

An on-premise, GPU-accelerated OCR system that turns a Prima Fresh Mart store staff member's phone photo of a Delivery Order (DO) into structured, database-backed data — PO/SO/DO numbers, dates, store, and item lines — with a Flutter mobile client on one end and a Dockerized AI pipeline on the other.

No cloud OCR API is used. Everything (layout detection, text recognition, LLM-assisted structuring) runs on your own GPU.


Table of Contents


What This Is

A Prima Fresh Mart store staff member (petugas toko), on duty at the store — not the delivery driver — photographs the DO paper on the Flutter app when a delivery arrives. The app checks the photo isn't blurry, tags it with GPS, and uploads it. The backend runs the image through a GPU OCR pipeline (deskew → layout detection → text recognition → LLM structuring), cross-checks every item line against a master SKU/store database, and saves the result. The app polls for the result, the store staff reviews/corrects it on-device and confirms receipt (entering their own name as receiver), and can print or export a signed delivery receipt as a PDF. Each account is bound to exactly one store (role: store), so whoever is on duty there uses the same login.

Two consumers of the same backend exist:

  • Flutter mobile app (lib/) — the primary, store-facing client used by staff at each Prima Fresh Mart location.
  • Next.js web pages (backend/pfm-web-app/src/app/*.tsx) — internal tooling for manual labeling, accuracy comparison, and an OCR "arena" for engine comparison. Not part of the store-facing product.

Tech Stack

Layer Technology Role
Mobile client Flutter (Riverpod, Dio, Hive, go_router) Camera capture, blur detection, offline-tolerant upload queue, manual correction editor, native PDF/print
API gateway Next.js (App Router, TypeScript) Auth, upload handling, orchestrates the OCR pipeline, post-processing (regex extraction, fuzzy SKU/store matching), serves internal web tooling
Reverse proxy Nginx Single entry point (:8000) routing to the gateway, pipeline API, and vLLM server
OCR pipeline PaddleOCR v6 + PP-DocLayoutV3 (FastAPI, GPU) Auto-deskew/unwarp, layout segmentation, text detection & recognition
Structuring LLM vLLM serving PaddleOCR-VL-1.6-0.9B Reassembles OCR text fragments into coherent structured text
Database PostgreSQL 15 Documents, line items, SKU/vendor/customer/store master data
Orchestration Docker Compose (multi-stage GPU Dockerfile) Runs the whole backend as one stack
Product scan (secondary) DINOv2 similarity search (YOLO classifier fallback) + PaddleOCR (classify_ocr_server.py, :8120) Single-product photo → SKU match via nearest-embedding lookup against reference photos + expiry-date extraction

Architecture

graph LR
    subgraph Client["Flutter Mobile App"]
        A[Camera + Blur Check]
    end

    subgraph Edge["Nginx :8000"]
        N[Reverse Proxy]
    end

    subgraph Gateway["Next.js API Gateway :3000"]
        G1["/api/v1/documents/upload"]
        G2["/api/parse"]
        G3["/api/v1/documents (poll/edit)"]
    end

    subgraph AI["GPU OCR Pipeline"]
        P["Pipeline API :8090\n(deskew · layout · OCR)"]
        V["vLLM Server :8118\n(PaddleOCR-VL-1.6)"]
    end

    DB[(PostgreSQL\ndocuments · ocr_items · sku_master · store_master)]

    A -- "multipart POST" --> N --> G1
    G1 -- "insert parsed=false" --> DB
    G1 -- "trigger" --> G2
    G2 -- "image" --> P
    P <--> V
    G2 -- "regex + fuzzy match\nparsed=true" --> DB
    A -- "poll every 2s" --> N --> G3 --> DB

How It Works

sequenceDiagram
    participant App as Flutter App
    participant GW as Next.js Gateway
    participant Pipe as Pipeline API + vLLM
    participant DB as PostgreSQL

    App->>App: Capture photo, check blur (Laplacian variance)
    App->>GW: POST /documents/upload (image + GPS)
    GW->>DB: Insert document (parsed=false)
    GW->>Pipe: Forward image (deskew, layout, OCR)
    Pipe-->>GW: Structured markdown text
    GW->>GW: Regex extract (PO/SO/DO/date/plate)<br/>Fuzzy-match SKU & store master
    GW->>DB: Update document (parsed=true) + items
    loop every 2s, up to 2 min
        App->>GW: GET /documents
        GW-->>App: Parsed result once ready
    end
    App->>App: Operator reviews & corrects
    App->>App: Confirmation dialog (receiver name + consent checkbox)
    App->>GW: PUT /documents/:id (final data)
    App->>App: Generate & print delivery receipt PDF

The poll loop actually runs every 2 seconds for up to 130 attempts (~4.3 minutes) before giving up and surfacing a timeout error — see lib/features/documents/pending_documents_provider.dart.

For the full field-by-field extraction rules (fused-digit correction, date sanitization, the triple-check SKU matcher, fuzzy store resolution), see docs/workflow_detail_aplikasi.md and docs/regex_rules_example.md.

For a visual, screen-by-screen walkthrough of the Flutter app itself (every page, button, popup, and the algorithms behind them — blur detection, upload/poll/retry, DINOv2 product classification — all captured against the live backend), see screenshots/v2/WORKFLOW.md and the companion slide deck screenshots/v2/Prima-Mart-Scanner-Workflow.pptx.


OCR Accuracy

Accuracy is tracked field-by-field against 37 hand-labeled real DO photos (backend/sources/test-images + manual_labels.json), not a single vague "it works" claim. The latest logged run (backend/sources/accuracy_history.jsonl):

Overall: 89.4% exact-field-match — up from an 82.2% baseline when this tracking tool was first built, against a 95% target.

The bigger story is where that accuracy comes from. Raw regex extraction straight off the OCR text is only 67.7% — the gain to 89.4% comes from a second correction pass: fuzzy SKU/store matching against master data, unit standardization, and format sanitization.

Field Raw regex After sanitize + triple-check What closes the gap
Customer (Kepada Yth) 100% 100% —
Kode Barang (SKU) 98.4% 98.4% Already reliable at the OCR layer
Item count 13.5% 97.3% Table-noise rows filtered by the SKU/unit triple-check
Nama Barang 0% 94.4% Corrected to master sku_master.nama_item on SKU match
Banyak (qty) 77.6% 95.2% Unit standardized from sku_master.jenis_outer
Jumlah (total) 76.8% 94.4% Unit standardized from sku_master.standar_jumlah
No. PO 91.9% 91.9% Fused-digit correction already applied at regex layer
No. DO 91.9% 91.9% —
Tanggal 89.2% 89.2% —
No. SO 86.5% 86.5% —
Store 0% 64.9% Fuzzy token match against store_master
Plat Truk 62.2% 62.2% Mostly genuine OCR misses on the truck-line, not a parsing gap
Alamat 0% 37.8% Canonicalized against customers table where phrasing matches

Known gaps toward the 95% target:

  • Alamat (37.8%, capped) — the ground-truth labels themselves use two different phrasings for the same physical address across photo batches; closing this needs the ground truth unified, not more parsing logic.
  • Plat (62.2%) — mostly genuine OCR misses (the truck line is often faint or absent in the photo) rather than a fixable parsing bug.
  • Store (64.9%) — bounded by the same store_master table noted in What to Consider — accuracy against it can only improve as far as that master data is populated.

A real factor behind these gaps is photo capture SOP, not just parsing or OCR quality. A meaningful share of the current test set was photographed with the DO paper placed on top of other papers/documents rather than a plain, flat surface. That confuses auto-deskew: the pipeline estimates page tilt from the average angle of detected text blocks, and overlapping paper edges/text from the sheet underneath make that estimate unreliable, which is exactly the failure mode behind the unwarp-retry logic described above. In practice this means accuracy here is a floor, not a ceiling — tightening the field SOP (DO paper alone, on a flat contrasting surface, reasonably well-lit and squared to the camera) should raise these numbers without any further code changes.

Reproduce this yourself with backend/pfm-web-app/scripts/accuracy-check.mts (npm run accuracy, or --refresh-ocr to force a real pipeline re-run instead of using cached OCR results) — see Testing & Tooling.


Repository Structure

.
├── backend/
│   ├── config/                 # Pipeline & vLLM YAML configs
│   ├── db/migrations/          # SQL run once by Postgres on a brand-new volume
│   ├── pfm-web-app/            # Next.js API gateway + internal web tooling
│   │   ├── src/app/api/        # All backend routes (upload, parse, documents, auth, arena...)
│   │   ├── src/db/             # DB pool + init.ts (idempotent schema + seed data)
│   │   ├── src/utils/parser.ts # Regex extraction + sanitization rules
│   │   └── scripts/            # Accuracy-check tooling
│   ├── sources/                # Test images, manual labels, SKU/store reference data
│   ├── Dockerfile              # Multi-stage GPU build (vllm-server, pipeline-api, pfm-web-app, gradio-ui)
│   ├── nginx.conf
│   └── docker-compose.yml      # ⚠️ Legacy standalone copy — see "What to Consider" below
├── lib/                         # Flutter app
│   ├── config/                 # app_config.dart — theme + API base URL resolution
│   ├── core/                   # Dio client, Hive storage, location, router
│   ├── features/               # auth, camera, documents, editor
│   └── models/
├── docs/                        # Deep-dive workflow & extraction-rule docs
├── screenshots/v2/               # App walkthrough: WORKFLOW.md + slide deck, real-backend screenshots
├── test/                        # Flutter widget/unit tests
├── docker-compose.yml           # ⭐ Canonical backend stack — run this one
├── docker-compose.demo.yml      # Production-mode override (see below)
└── start-dev-tunnel.ps1         # Syncs LAN IP into app_config.dart + starts ngrok

Getting Started

Prerequisites

Component Requirement
Backend host NVIDIA GPU, CUDA 12.6+ driver, ~8GB+ VRAM. Tested working on both native Linux and Windows + Docker Desktop with WSL2 GPU passthrough.
Docker Docker Engine/Desktop with the NVIDIA Container Toolkit (docker info should list nvidia under Runtimes)
Disk space 60GB+ free — the pipeline-api and vllm-server images alone are ~30GB each once built, plus model weight caches
Flutter Flutter SDK >=3.2.0 <4.0.0, Android Studio/Xcode for device tooling
Node.js v20+ (optional — only for running the parser unit tests or accuracy tooling outside Docker)

1. Backend Setup

  1. Configure environment variables (from the repo root):

    cp backend/.env.example backend/.env
    

    Set CUDA_VISIBLE_DEVICES to your GPU index, and APP_PORT if 8000 is taken.

  2. Start the stack from the repo root (not backend/ — see the gotcha below):

    docker compose up --build
    

    First build pulls/builds ~60GB of GPU images and downloads model weights — expect this to take a long time on the first run. Subsequent starts are fast.

  3. Verify it's up:

    curl http://localhost:8000/health              # pipeline API health
    curl http://localhost:8000/v1/models            # vLLM model list
    curl -X POST http://localhost:8000/api/v1/auth/login \
         -H "Content-Type: application/json" -d '{"username":"admin","password":"password"}'
    

Database schema and reference data (vendor, customer, a starter SKU catalog) are created automatically on first request — backend/pfm-web-app/src/db/init.ts runs idempotent CREATE TABLE IF NOT EXISTS + seed statements every time the app starts, so there is no manual migration step. backend/db/migrations/*.sql also runs once via Postgres's own docker-entrypoint-initdb.d on a brand-new volume, but init.ts is what you should treat as the source of truth.

2. Frontend Setup

  1. Install dependencies:

    flutter pub get
    
  2. Point the app at your backend. lib/config/app_config.dart resolves the API URL dynamically at startup: it first tries a public ngrok tunnel, and falls back to a hardcoded LAN URL if that's unreachable. Both need to match your actual machine:

    ./start-dev-tunnel.ps1
    

    This detects your current LAN IP, patches _lanBaseUrl in app_config.dart for you, and launches ngrok pointed at the fixed reserved domain in that same file. Run it again any time your IP changes (new network, DHCP renewal, etc.) — a stale IP here is the single most common reason "the app can't log in" after this backend is confirmed healthy.

  3. Run it:

    flutter run
    
  4. Build a release APK when you need an installable build instead of a debug session:

    flutter build apk --release
    

    Output: build/app/outputs/flutter-apk/app-release.apk. See the signing note below before distributing it.


Running in Demo / Production Mode

POLICY: You MUST run the production mode override for any client demos or field testing.

The default docker-compose.yml runs the gateway via npm run dev with the source bind-mounted in. This is strictly for local development (enabling hot-reloading for iterating on parser.ts). It is a single dev-server process with a known throughput ceiling and will bottleneck if multiple people upload at once.

To run the production build instead, use this opt-in override:

docker compose -f docker-compose.yml -f docker-compose.demo.yml up -d --build

This drops the dev bind-mount and runs npm start against the image's own npm run build output. Rebuild (--build) before every demo — this mode does not hot-reload code changes.


What to Consider Before You Install

  • Two docker-compose.yml files exist — always run from the repo root. backend/docker-compose.yml is a near-duplicate, standalone copy of the same stack (project name ai-ocr-pfm-2026, originally a separate repo vendored into backend/). Docker Compose names containers/volumes after the project name declared in whichever compose file you invoke first. If you ever run docker compose up from inside backend/, you'll get container-name conflicts against anything already started from the root — this isn't hypothetical, it happened during this project's own testing. Pick one (the root file) and stick to it.

  • store_master ships empty. Schema is created automatically, but no store data is seeded — fuzzy store-matching (resolveStoreFromText) will not resolve any delivery address until you load real store data into that table yourself. backend/sources/toko_aktif.json looks like the right source for this but isn't wired into an automatic import yet.

  • First build is large and slow. The pipeline-api and vllm-server images are ~30GB each with GPU model weights. Budget real time and disk space for the first docker compose up --build.

  • Single GPU pipeline, no horizontal scaling. There's one pipeline-api and one vllm-server container. Concurrent uploads queue behind the GPU; this is a real throughput ceiling worth load-testing before a multi-store/multi-device demo, not just a single-user smoke test.

  • /api/v1/* enforces real auth; the classic dev routes deliberately don't. /api/v1/auth/login checks a bcrypt-hashed password against a real accounts table and signs a JWT; /api/v1/documents/* (list, upload, PUT-by-id) reject any request with a missing/invalid token with a real 401. The Flutter app always goes through this surface. The classic routes (/api/upload, /api/parse, /api/history, etc.) and the root/scan-pfm/manual-label web pages have no login flow and never will — they're dev-only internal tooling, not part of the store-facing product. Every API route still sets Access-Control-Allow-Origin: *, so this is fine for a controlled LAN/demo deployment but not for exposing the stack to the open internet as-is.

  • The Android release build is debug-signed. android/app/build.gradle.kts has a // TODO: Add your own signing config and currently signs release builds with the debug key. Fine for internal install/testing, not for Play Store distribution.

  • Mobile connectivity is two independent, manually-synced paths. The ngrok domain in app_config.dart is fixed/reserved, but the ngrok process isn't started automatically — you (or start-dev-tunnel.ps1) have to launch it. The LAN IP fallback is hardcoded and will silently go stale the moment your host machine's IP changes. If login/upload fails on a real device, check these two before anything else.


Testing & Tooling

What How
Parser regex/sanitization rules npx tsx backend/pfm-web-app/src/utils/parser.test.ts — no Docker needed
OCR accuracy regression suite backend/pfm-web-app/run_batch_test.js and backend/pfm-web-app/scripts/accuracy-check.mts, run against backend/sources/test-images + manual_labels.json
Flutter widget/unit tests flutter test (covers blur detection, camera navigation, editor validation, geotagging, the pending-upload queue, and more — see test/)
Static analysis flutter analyze lib

Known Limitations

Previously tracked here as open gaps, both now fixed and worth noting as resolved:

  • Pending-upload queue is in-memory only — it's now persisted to a local Hive box (lib/core/storage/local_storage.dart), so an OS-level app kill mid-upload no longer loses the document; the queue reloads and resumes on next launch.
  • Item-row writes on document save aren't wrapped in a transaction — PUT /documents/:id now runs inside withTransaction (backend/pfm-web-app/src/app/api/v1/documents/[id]/route.ts).

Still open, worth knowing before relying on this for unattended field use:

  • The DO Scan review form's validation fails silently on submit if a required field (e.g. driver name) is empty — Form.validate() returns false and the confirm button's onPressed just returns, with no toast or scroll-to-error to tell the operator why nothing happened.

See docs/ and screenshots/v2/WORKFLOW.md for the deeper workflow documentation these decisions were audited against.

S
Description
PFM OCR to scan Delivery Order (DO)
Readme
1.8 GiB
0 Stars 1 Watchers 0 Forks
Languages
TypeScript 42.2%
Dart 26.4%
Python 10.1%
HTML 8.4%
JavaScript 7%
Other 5.8%