docs(readme): restructure for clarity
Adds a badge/nav header, the four architecture and sequence diagrams from the handover set, a port table, a screenshot gallery, collapsible gotchas and troubleshooting sections, and a documentation index. Corrects two inaccuracies: endpoint resolution tries LAN first and ngrok second (not the reverse), and the poll loop runs ~4.3 minutes (not 2). Documents three security facts that were previously unrecorded: seeded passwords reset on every backend restart, JWT_SECRET falls back to a constant in source, and the LAN leg is plain HTTP.
This commit is contained in:
1 parent
0abe4e9729
commit
2e1cefe35d
1 file changed
+420
-224
@@ -1,56 +1,140 @@
|
||||
# Prima Fresh Mart Scanner (app-pfm-ocr-v2)
|
||||
<div align="center">
|
||||
|
||||
An on-premise, GPU-accelerated OCR system that turns a Prima Fresh Mart store staff member's phone photo of a **Delivery Order (DO)** into structured, database-backed data — PO/SO/DO numbers, dates, store, and item lines — with a Flutter mobile client on one end and a Dockerized AI pipeline on the other.
|
||||
# Prima Fresh Mart Scanner
|
||||
|
||||
No cloud OCR API is used. Everything (layout detection, text recognition, LLM-assisted structuring) runs on your own GPU.
|
||||
### Point a phone at a paper Delivery Order. Get a verified database row back.
|
||||
|
||||
**A self-hosted, GPU-accelerated OCR pipeline that turns store-staff phone photos of Delivery Orders into structured, master-data-validated records — with a Flutter client on one end and Docker + PaddleOCR + vLLM + PostgreSQL on the other.**
|
||||
|
||||
<p>
|
||||
<a href="https://flutter.dev/"><img alt="Flutter" src="https://img.shields.io/badge/Flutter-3.2+-02569B?logo=flutter&logoColor=white"></a>
|
||||
<a href="https://nextjs.org/"><img alt="Next.js" src="https://img.shields.io/badge/Next.js-16.2-000000?logo=nextdotjs&logoColor=white"></a>
|
||||
<a href="https://github.com/PaddlePaddle/PaddleOCR"><img alt="PaddleOCR" src="https://img.shields.io/badge/PaddleOCR-v6%20%2B%20PP--DocLayoutV3-0052CC"></a>
|
||||
<a href="https://github.com/vllm-project/vllm"><img alt="vLLM" src="https://img.shields.io/badge/vLLM-PaddleOCR--VL--1.6-FF6F00"></a>
|
||||
<a href="https://www.postgresql.org/"><img alt="PostgreSQL" src="https://img.shields.io/badge/PostgreSQL-15-4169E1?logo=postgresql&logoColor=white"></a>
|
||||
<a href="https://www.docker.com/"><img alt="Docker" src="https://img.shields.io/badge/Docker-Compose-2496ED?logo=docker&logoColor=white"></a>
|
||||
<a href="https://developer.nvidia.com/cuda-toolkit"><img alt="CUDA" src="https://img.shields.io/badge/CUDA-12.6+-76B900?logo=nvidia&logoColor=white"></a>
|
||||
<img alt="Accuracy" src="https://img.shields.io/badge/field%20accuracy-89.4%25-success">
|
||||
<a href="docs/handover/PFM-Scanner-Handover-Document.pdf"><img alt="Docs" src="https://img.shields.io/badge/Docs-PDF%20Handover-red?logo=adobe-acrobat-reader&logoColor=white"></a>
|
||||
</p>
|
||||
|
||||
[⚡ Quickstart](#quickstart) · [🏗️ Architecture](#architecture) · [🔄 How It Works](#how-it-works) · [🎯 Accuracy](#accuracy) · [🖼️ Screens](#screens) · [⚠️ Gotchas](#gotchas) · [📚 Docs](#documentation) · [🔧 Troubleshooting](#troubleshooting)
|
||||
|
||||
<br>
|
||||
|
||||
<a href="docs/handover/src/img/ring-alur.png">
|
||||
<img src="docs/handover/src/img/ring-alur.png" width="100%" alt="System architecture: phone, Nginx, Next.js gateway, GPU OCR pipeline, PostgreSQL" />
|
||||
</a>
|
||||
|
||||
*Every box except the phone and the browser is a container in one Docker Compose stack.*
|
||||
|
||||
</div>
|
||||
|
||||
---
|
||||
|
||||
## Table of Contents
|
||||
<a id="why"></a>
|
||||
## 💡 Why This Exists
|
||||
|
||||
- [What This Is](#what-this-is)
|
||||
- [Tech Stack](#tech-stack)
|
||||
- [Architecture](#architecture)
|
||||
- [How It Works](#how-it-works)
|
||||
- [OCR Accuracy](#ocr-accuracy)
|
||||
- [Repository Structure](#repository-structure)
|
||||
- [Getting Started](#getting-started)
|
||||
- [Prerequisites](#prerequisites)
|
||||
- [1. Backend Setup](#1-backend-setup)
|
||||
- [2. Frontend Setup](#2-frontend-setup)
|
||||
- [Running in Demo / Production Mode](#running-in-demo--production-mode)
|
||||
- [What to Consider Before You Install](#what-to-consider-before-you-install)
|
||||
- [Testing & Tooling](#testing--tooling)
|
||||
- [Known Limitations](#known-limitations)
|
||||
When a delivery arrives at a Prima Fresh Mart store, the **store staff member on duty (petugas toko)** — not the driver — has to transcribe a paper Delivery Order into the system. Done by hand it is slow, error-prone, and unverifiable after the fact.
|
||||
|
||||
---
|
||||
This system replaces that with a photo, and adds three things a human typist cannot cheaply provide:
|
||||
|
||||
## What This Is
|
||||
1. **No cloud OCR.** Layout detection, text recognition, and LLM structuring all run on your own GPU. DO photos containing customer and pricing data never leave the premises.
|
||||
2. **Master-data validation, not just transcription.** Every item line is cross-checked against `sku_master`; every store against `store_master`. Raw regex off the OCR text scores **67.7%** — the correction pass against master data is what lifts it to **89.4%**.
|
||||
3. **A proof-of-receipt trail.** GPS tag at capture, blur check before upload, SHA-256 dedup server-side, a named receiver on confirmation, and a printable PDF receipt — all bound to one account per store.
|
||||
|
||||
A Prima Fresh Mart **store staff member (petugas toko)**, on duty at the store — not the delivery driver — photographs the DO paper on the Flutter app when a delivery arrives. The app checks the photo isn't blurry, tags it with GPS, and uploads it. The backend runs the image through a GPU OCR pipeline (deskew → layout detection → text recognition → LLM structuring), cross-checks every item line against a master SKU/store database, and saves the result. The app polls for the result, the store staff reviews/corrects it on-device and confirms receipt (entering their own name as receiver), and can print or export a signed delivery receipt as a PDF. Each account is bound to exactly one store (`role: store`), so whoever is on duty there uses the same login.
|
||||
**Two clients share one backend:**
|
||||
|
||||
Two consumers of the same backend exist:
|
||||
- **Flutter mobile app** (`lib/`) — the primary, store-facing client used by staff at each Prima Fresh Mart location.
|
||||
- **Next.js web pages** (`backend/pfm-web-app/src/app/*.tsx`) — internal tooling for manual labeling, accuracy comparison, and an OCR "arena" for engine comparison. Not part of the store-facing product.
|
||||
|
||||
---
|
||||
|
||||
## Tech Stack
|
||||
|
||||
| Layer | Technology | Role |
|
||||
| Client | Path | Audience |
|
||||
|---|---|---|
|
||||
| Mobile client | **Flutter** (Riverpod, Dio, Hive, go_router) | Camera capture, blur detection, offline-tolerant upload queue, manual correction editor, native PDF/print |
|
||||
| API gateway | **Next.js** (App Router, TypeScript) | Auth, upload handling, orchestrates the OCR pipeline, post-processing (regex extraction, fuzzy SKU/store matching), serves internal web tooling |
|
||||
| Reverse proxy | **Nginx** | Single entry point (`:8000`) routing to the gateway, pipeline API, and vLLM server |
|
||||
| OCR pipeline | **PaddleOCR v6 + PP-DocLayoutV3** (FastAPI, GPU) | Auto-deskew/unwarp, layout segmentation, text detection & recognition |
|
||||
| Structuring LLM | **vLLM serving PaddleOCR-VL-1.6-0.9B** | Reassembles OCR text fragments into coherent structured text |
|
||||
| Database | **PostgreSQL 15** | Documents, line items, SKU/vendor/customer/store master data |
|
||||
| Orchestration | **Docker Compose** (multi-stage GPU Dockerfile) | Runs the whole backend as one stack |
|
||||
| Product scan (secondary) | **DINOv2 similarity search** (YOLO classifier fallback) **+ PaddleOCR** (`classify_ocr_server.py`, `:8120`) | Single-product photo → SKU match via nearest-embedding lookup against reference photos + expiry-date extraction |
|
||||
| **Flutter mobile app** | `lib/` | The product. Store staff at each location. |
|
||||
| **Next.js web pages** | `backend/pfm-web-app/src/app/` | Internal tooling only — manual labeling, accuracy comparison, OCR engine "arena". Not store-facing. |
|
||||
|
||||
---
|
||||
|
||||
## Architecture
|
||||
<a id="quickstart"></a>
|
||||
## ⚡ Quickstart
|
||||
|
||||
### Prerequisites
|
||||
|
||||
| Component | Requirement |
|
||||
|---|---|
|
||||
| **GPU host** | NVIDIA GPU, CUDA 12.6+ driver, ~8 GB+ VRAM. Verified on native Linux **and** Windows + Docker Desktop with WSL2 GPU passthrough. |
|
||||
| **Docker** | Engine/Desktop with NVIDIA Container Toolkit — `docker info` must list `nvidia` under Runtimes. |
|
||||
| **Disk** | **60 GB+ free.** `pipeline-api` and `vllm-server` images are ~30 GB each once built, before model weight caches. |
|
||||
| **Flutter** | SDK `>=3.2.0 <4.0.0`, plus Android Studio / Xcode device tooling. |
|
||||
| **Node.js** | v20+ — optional, only for running parser tests or accuracy tooling outside Docker. |
|
||||
|
||||
> [!IMPORTANT]
|
||||
> Run **every** `docker compose` command from the **repository root**, never from inside `backend/`. A second, legacy compose file lives there under a different project name and will collide with containers already started from the root. See [Gotchas](#gotchas).
|
||||
|
||||
### 1 · Backend
|
||||
|
||||
```bash
|
||||
# 1. Configure environment (from the repo root)
|
||||
cp backend/.env.example backend/.env
|
||||
# Edit backend/.env:
|
||||
# APP_PORT=8000 host port exposed via Nginx
|
||||
# CUDA_VISIBLE_DEVICES=0 your GPU index — check with nvidia-smi
|
||||
# JWT_SECRET=<random-value> do NOT leave this as "change-me"
|
||||
|
||||
# 2. Build and start the whole stack
|
||||
docker compose up --build
|
||||
```
|
||||
|
||||
First build pulls/builds ~60 GB of GPU images and downloads model weights — expect a long wait once. Subsequent starts are fast.
|
||||
|
||||
**Verify it came up:**
|
||||
|
||||
```bash
|
||||
docker compose ps # every service running / healthy
|
||||
curl http://localhost:8000/health # pipeline API
|
||||
curl http://localhost:8000/v1/models # vLLM model list
|
||||
curl -X POST http://localhost:8000/api/v1/auth/login \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"username":"admin","password":"password"}'
|
||||
```
|
||||
|
||||
Services come up in stages: `db` and `vllm-server` first, then `pipeline-api` waits until healthy, then `pfm-web-app`. `vllm-server` loading model weights onto the GPU can take several minutes on a cold start — that is expected, not a hang.
|
||||
|
||||
The schema and reference data are created **automatically** — `backend/pfm-web-app/src/db/init.ts` runs idempotent `CREATE TABLE IF NOT EXISTS` + seed statements on every start, so there is no manual migration step. (`backend/db/migrations/*.sql` also runs once via Postgres's `docker-entrypoint-initdb.d`, but only on a brand-new volume; treat `init.ts` as the source of truth.)
|
||||
|
||||
### 2 · Flutter app
|
||||
|
||||
```bash
|
||||
flutter pub get
|
||||
flutter devices # confirm your phone/emulator is visible
|
||||
flutter run
|
||||
```
|
||||
|
||||
**Point the app at your backend first.** `lib/config/app_config.dart` resolves the API URL at every launch: it tries the **LAN address first**, then falls back to the **ngrok tunnel** — whichever answers with genuine backend JSON wins (a dead tunnel's error page is rejected, not accepted).
|
||||
|
||||
```bash
|
||||
./start-dev-tunnel.ps1
|
||||
```
|
||||
|
||||
That script detects your current LAN IP, patches `_lanBaseUrl` in `app_config.dart`, and launches `ngrok` against the reserved domain in that same file — pointed at **port 8001**, the narrow `/api/v1`-only door, so the internal web tooling is never exposed publicly.
|
||||
|
||||
| Running on | LAN address to use in `app_config.dart` |
|
||||
|---|---|
|
||||
| Physical device, same Wi-Fi | `http://<host LAN IP>:8000/api/v1` |
|
||||
| Android emulator | `http://10.0.2.2:8000/api/v1` |
|
||||
| iOS simulator | `http://localhost:8000/api/v1` |
|
||||
|
||||
### 3 · Release APK
|
||||
|
||||
```bash
|
||||
flutter build apk --release
|
||||
# → build/app/outputs/flutter-apk/app-release.apk
|
||||
```
|
||||
|
||||
> [!WARNING]
|
||||
> Both backend URLs are **compile-time constants**. Every time the host LAN IP or ngrok domain changes, the APK must be rebuilt and redistributed to every store device. This is the single most common cause of "the app suddenly can't log in."
|
||||
|
||||
---
|
||||
|
||||
<a id="architecture"></a>
|
||||
## 🏗️ Architecture
|
||||
|
||||
```mermaid
|
||||
graph LR
|
||||
@@ -58,7 +142,7 @@ graph LR
|
||||
A[Camera + Blur Check]
|
||||
end
|
||||
|
||||
subgraph Edge["Nginx :8000"]
|
||||
subgraph Edge["Nginx :8000 / :8001"]
|
||||
N[Reverse Proxy]
|
||||
end
|
||||
|
||||
@@ -69,237 +153,349 @@ graph LR
|
||||
end
|
||||
|
||||
subgraph AI["GPU OCR Pipeline"]
|
||||
P["Pipeline API :8090\n(deskew · layout · OCR)"]
|
||||
V["vLLM Server :8118\n(PaddleOCR-VL-1.6)"]
|
||||
P["Pipeline API :8090<br/>deskew · layout · OCR"]
|
||||
V["vLLM Server :8118<br/>PaddleOCR-VL-1.6"]
|
||||
end
|
||||
|
||||
DB[(PostgreSQL\ndocuments · ocr_items · sku_master · store_master)]
|
||||
DB[(PostgreSQL :5432<br/>documents · ocr_items<br/>sku_master · store_master)]
|
||||
FS[/backend/uploads/<br/>photo + JSON/]
|
||||
|
||||
A -- "multipart POST" --> N --> G1
|
||||
G1 -- "write file" --> FS
|
||||
G1 -- "insert parsed=false" --> DB
|
||||
G1 -- "trigger" --> G2
|
||||
G2 -- "image" --> P
|
||||
P <--> V
|
||||
G2 -- "regex + fuzzy match\nparsed=true" --> DB
|
||||
G2 -- "regex + fuzzy match<br/>parsed=true" --> DB
|
||||
A -- "poll every 2s" --> N --> G3 --> DB
|
||||
```
|
||||
|
||||
---
|
||||
### Ports
|
||||
|
||||
## How It Works
|
||||
Only **two** ports are exposed off the host. Everything else is container-internal.
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant App as Flutter App
|
||||
participant GW as Next.js Gateway
|
||||
participant Pipe as Pipeline API + vLLM
|
||||
participant DB as PostgreSQL
|
||||
| Port | Container | Serves | Exposure |
|
||||
|:---:|---|---|:---:|
|
||||
| **8000** | `paddleocr-nginx` | Front door — full internal surface, including web tooling | 🌐 Public |
|
||||
| **8001** | `paddleocr-nginx` | Narrow door — `/api/v1/*` only, everything else `404` | 🌐 Public |
|
||||
| 3000 | `paddleocr-pfm-web-app` | Next.js gateway + admin pages | 🔒 Internal |
|
||||
| 8090 | `paddleocr-pipeline-api` | PaddleOCR pipeline, `/layout-parsing` | 🔒 Internal |
|
||||
| 8118 | `paddleocr-vllm-server` | vLLM serving PaddleOCR-VL-1.6-0.9B | 🔒 Internal |
|
||||
| 8120 | `paddleocr-pipeline-api` | Product-scan classifier, `/classify-ocr` | 🔒 Internal |
|
||||
| 5432 | `paddleocr-db` | PostgreSQL 15, database `dopfm` | 🔒 Internal |
|
||||
|
||||
App->>App: Capture photo, check blur (Laplacian variance)
|
||||
App->>GW: POST /documents/upload (image + GPS)
|
||||
GW->>DB: Insert document (parsed=false)
|
||||
GW->>Pipe: Forward image (deskew, layout, OCR)
|
||||
Pipe-->>GW: Structured markdown text
|
||||
GW->>GW: Regex extract (PO/SO/DO/date/plate)<br/>Fuzzy-match SKU & store master
|
||||
GW->>DB: Update document (parsed=true) + items
|
||||
loop every 2s, up to 2 min
|
||||
App->>GW: GET /documents
|
||||
GW-->>App: Parsed result once ready
|
||||
end
|
||||
App->>App: Operator reviews & corrects
|
||||
App->>App: Confirmation dialog (receiver name + consent checkbox)
|
||||
App->>GW: PUT /documents/:id (final data)
|
||||
App->>App: Generate & print delivery receipt PDF
|
||||
```
|
||||
<div align="center">
|
||||
<a href="docs/handover/src/img/ring-nginx.png">
|
||||
<img src="docs/handover/src/img/ring-nginx.png" width="92%" alt="Nginx routing: port 8000 full surface, port 8001 API-v1 only" />
|
||||
</a>
|
||||
</div>
|
||||
|
||||
The poll loop actually runs every 2 seconds for up to 130 attempts (~4.3 minutes) before giving up and surfacing a timeout error — see `lib/features/documents/pending_documents_provider.dart`.
|
||||
Body-size limits are disabled (`client_max_body_size 0`) because DO photos are large, and proxy timeouts are raised to 300 s because a GPU OCR pass is slow. Both are set in [`backend/nginx.conf`](backend/nginx.conf).
|
||||
|
||||
For the full field-by-field extraction rules (fused-digit correction, date sanitization, the triple-check SKU matcher, fuzzy store resolution), see [`docs/workflow_detail_aplikasi.md`](docs/workflow_detail_aplikasi.md) and [`docs/regex_rules_example.md`](docs/regex_rules_example.md).
|
||||
### Stack
|
||||
|
||||
For a visual, screen-by-screen walkthrough of the Flutter app itself (every page, button, popup, and the algorithms behind them — blur detection, upload/poll/retry, DINOv2 product classification — all captured against the live backend), see [`screenshots/v2/WORKFLOW.md`](screenshots/v2/WORKFLOW.md) and the companion slide deck [`screenshots/v2/Prima-Mart-Scanner-Workflow.pptx`](screenshots/v2/Prima-Mart-Scanner-Workflow.pptx).
|
||||
| Layer | Technology | Role |
|
||||
|---|---|---|
|
||||
| Mobile client | **Flutter** — Riverpod, Dio, Hive, go_router | Capture, blur detection, GPS, crash-tolerant upload queue, correction editor, native PDF/print |
|
||||
| API gateway | **Next.js 16** (App Router, TypeScript) | Auth, upload, pipeline orchestration, regex extraction, fuzzy SKU/store matching, internal tooling |
|
||||
| Reverse proxy | **Nginx** | Two server blocks — full surface on `:8000`, API-only on `:8001` |
|
||||
| OCR pipeline | **PaddleOCR v6 + PP-DocLayoutV3** (FastAPI, GPU) | Auto-deskew/unwarp, layout segmentation, text detection & recognition |
|
||||
| Structuring LLM | **vLLM + PaddleOCR-VL-1.6-0.9B** | Reassembles OCR fragments into coherent structured text |
|
||||
| Database | **PostgreSQL 15** | Documents, line items, SKU / vendor / customer / store master data |
|
||||
| Orchestration | **Docker Compose** (multi-stage GPU Dockerfile) | The entire backend as one stack |
|
||||
| Product scan | **DINOv2 similarity search** (YOLO fallback) **+ PaddleOCR** | Single-product photo → SKU via nearest-embedding lookup + expiry-date extraction |
|
||||
|
||||
---
|
||||
|
||||
## OCR Accuracy
|
||||
<a id="how-it-works"></a>
|
||||
## 🔄 How It Works
|
||||
|
||||
Accuracy is tracked field-by-field against 37 hand-labeled real DO photos (`backend/sources/test-images` + `manual_labels.json`), not a single vague "it works" claim. The latest logged run (`backend/sources/accuracy_history.jsonl`):
|
||||
### Authentication — one account, one store
|
||||
|
||||
**Overall: 89.4%** exact-field-match — up from an 82.2% baseline when this tracking tool was first built, against a 95% target.
|
||||
<div align="center">
|
||||
<a href="docs/handover/src/img/ring-login.png">
|
||||
<img src="docs/handover/src/img/ring-login.png" width="96%" alt="Login sequence: app to gateway to PostgreSQL, bcrypt check, JWT issue" />
|
||||
</a>
|
||||
</div>
|
||||
|
||||
The bigger story is *where* that accuracy comes from. Raw regex extraction straight off the OCR text is only **67.7%** — the gain to 89.4% comes from a second correction pass: fuzzy SKU/store matching against master data, unit standardization, and format sanitization.
|
||||
Accounts are **generated, not typed**. `init.ts` seeds one `store` account per row in `store_master` (username = store code), plus one `admin`. The signed JWT carries `kodeToko`, which is why store name and address never have to be OCR'd off the photo — they come from whoever uploaded it.
|
||||
|
||||
### Upload and OCR
|
||||
|
||||
<div align="center">
|
||||
<a href="docs/handover/src/img/ring-upload.png">
|
||||
<img src="docs/handover/src/img/ring-upload.png" width="100%" alt="Upload sequence across app, gateway, uploads folder, GPU pipeline, and database" />
|
||||
</a>
|
||||
</div>
|
||||
|
||||
Three properties worth knowing:
|
||||
|
||||
- **Dedup is content-addressed.** The server hashes the file (SHA-256) before writing anything. A retried upload after a perceived timeout returns the *original* document instead of creating a second row or re-running the GPU pass.
|
||||
- **Photos are files, not blobs.** Images land in `backend/uploads/`; only the filename goes in the database. **Back up the database and that folder together** — either one alone is useless.
|
||||
- **New documents are born `confirmed = false`** and stay invisible to `GET /api/v1/documents` until the staff member saves their corrections via `PUT /api/v1/documents/:id`. That is the confirmation gate.
|
||||
|
||||
The app polls every 2 s for up to **130 attempts (~4.3 minutes)** before surfacing a timeout — see `lib/features/documents/pending_documents_provider.dart`. Server-side, `/api/parse` is bounded at 210 s, marks the document parsed with `"Not Found"` placeholders on pipeline failure so it never hangs forever, and records the reason in `parse_error` when the call itself dies.
|
||||
|
||||
For field-by-field extraction rules — fused-digit correction, date sanitization, the triple-check SKU matcher, fuzzy store resolution — see [`docs/workflow_detail_aplikasi.md`](docs/workflow_detail_aplikasi.md) and [`docs/regex_rules_example.md`](docs/regex_rules_example.md).
|
||||
|
||||
---
|
||||
|
||||
<a id="accuracy"></a>
|
||||
## 🎯 OCR Accuracy
|
||||
|
||||
Tracked field-by-field against **37 hand-labeled real DO photos** (`backend/sources/test-images` + `manual_labels.json`) — not a vague "it works". Latest logged run in `backend/sources/accuracy_history.jsonl`:
|
||||
|
||||
<div align="center">
|
||||
|
||||
### **89.4%** exact-field-match · up from an **82.2%** baseline · target **95%**
|
||||
|
||||
</div>
|
||||
|
||||
The interesting part is *where* the accuracy comes from. Raw regex straight off the OCR text scores only **67.7%**; the second correction pass — fuzzy matching against master data, unit standardization, format sanitization — is what closes the gap.
|
||||
|
||||
| Field | Raw regex | After sanitize + triple-check | What closes the gap |
|
||||
|---|---:|---:|---|
|
||||
| Customer (Kepada Yth) | 100% | 100% | — |
|
||||
| Kode Barang (SKU) | 98.4% | 98.4% | Already reliable at the OCR layer |
|
||||
| Item count | 13.5% | 97.3% | Table-noise rows filtered by the SKU/unit triple-check |
|
||||
| Nama Barang | 0% | 94.4% | Corrected to master `sku_master.nama_item` on SKU match |
|
||||
| Banyak (qty) | 77.6% | 95.2% | Unit standardized from `sku_master.jenis_outer` |
|
||||
| Jumlah (total) | 76.8% | 94.4% | Unit standardized from `sku_master.standar_jumlah` |
|
||||
| No. PO | 91.9% | 91.9% | Fused-digit correction already applied at regex layer |
|
||||
| No. DO | 91.9% | 91.9% | — |
|
||||
| Tanggal | 89.2% | 89.2% | — |
|
||||
| No. SO | 86.5% | 86.5% | — |
|
||||
| Store | 0% | 64.9% | Fuzzy token match against `store_master` |
|
||||
| Plat Truk | 62.2% | 62.2% | Mostly genuine OCR misses on the truck-line, not a parsing gap |
|
||||
| Alamat | 0% | 37.8% | Canonicalized against `customers` table where phrasing matches |
|
||||
| Customer (Kepada Yth) | 100% | **100%** | — |
|
||||
| Kode Barang (SKU) | 98.4% | **98.4%** | Already reliable at the OCR layer |
|
||||
| Item count | 13.5% | **97.3%** | Table-noise rows filtered by the SKU/unit triple-check |
|
||||
| Nama Barang | 0% | **94.4%** | Corrected to master `sku_master.nama_item` on SKU match |
|
||||
| Banyak (qty) | 77.6% | **95.2%** | Unit standardized from `sku_master.jenis_outer` |
|
||||
| Jumlah (total) | 76.8% | **94.4%** | Unit standardized from `sku_master.standar_jumlah` |
|
||||
| No. PO | 91.9% | **91.9%** | Fused-digit correction already applied at regex layer |
|
||||
| No. DO | 91.9% | **91.9%** | — |
|
||||
| Tanggal | 89.2% | **89.2%** | — |
|
||||
| No. SO | 86.5% | **86.5%** | — |
|
||||
| Store | 0% | **64.9%** | Fuzzy token match against `store_master` |
|
||||
| Plat Truk | 62.2% | **62.2%** | Mostly genuine OCR misses on the truck line |
|
||||
| Alamat | 0% | **37.8%** | Canonicalized against `customers` where phrasing matches |
|
||||
|
||||
**Known gaps toward the 95% target:**
|
||||
- **Alamat (37.8%, capped)** — the ground-truth labels themselves use two different phrasings for the same physical address across photo batches; closing this needs the ground truth unified, not more parsing logic.
|
||||
- **Plat (62.2%)** — mostly genuine OCR misses (the truck line is often faint or absent in the photo) rather than a fixable parsing bug.
|
||||
- **Store (64.9%)** — bounded by the same `store_master` table noted in [What to Consider](#what-to-consider-before-you-install) — accuracy against it can only improve as far as that master data is populated.
|
||||
<details>
|
||||
<summary><b>Known gaps toward the 95% target — and why photo SOP matters more than parsing</b></summary>
|
||||
|
||||
**A real factor behind these gaps is photo capture SOP, not just parsing or OCR quality.** A meaningful share of the current test set was photographed with the DO paper placed on top of other papers/documents rather than a plain, flat surface. That confuses auto-deskew: the pipeline estimates page tilt from the average angle of detected text blocks, and overlapping paper edges/text from the sheet underneath make that estimate unreliable, which is exactly the failure mode behind the unwarp-retry logic described above. In practice this means accuracy here is a floor, not a ceiling — tightening the field SOP (DO paper alone, on a flat contrasting surface, reasonably well-lit and squared to the camera) should raise these numbers without any further code changes.
|
||||
<br>
|
||||
|
||||
Reproduce this yourself with `backend/pfm-web-app/scripts/accuracy-check.mts` (`npm run accuracy`, or `--refresh-ocr` to force a real pipeline re-run instead of using cached OCR results) — see [Testing & Tooling](#testing--tooling).
|
||||
| Field | Ceiling | Why it is stuck |
|
||||
|---|:---:|---|
|
||||
| **Alamat** | 37.8% | The ground-truth labels themselves use two different phrasings for the same physical address across photo batches. Fixing this needs the ground truth unified, not more parsing logic. |
|
||||
| **Plat Truk** | 62.2% | Genuine OCR misses — the truck line is often faint or absent in the photo. Not a parsing bug. |
|
||||
| **Store** | 64.9% | Bounded by how populated `store_master` is. See [Gotchas](#gotchas). |
|
||||
|
||||
**A real share of the remaining error is capture technique, not code.** Many photos in the current test set were shot with the DO paper resting on *other papers* instead of a plain flat surface. Auto-deskew estimates page tilt from the average angle of detected text blocks — overlapping edges and text from the sheet underneath corrupt that estimate, which is exactly the failure mode the unwarp-retry logic exists for.
|
||||
|
||||
In other words these numbers are a **floor, not a ceiling**. Tightening field SOP — DO paper alone, flat contrasting surface, decent light, squared to the camera — should raise them with no code change at all.
|
||||
|
||||
</details>
|
||||
|
||||
Reproduce it yourself with [`backend/pfm-web-app/scripts/accuracy-check.mts`](backend/pfm-web-app/scripts/accuracy-check.mts) (`npm run accuracy`, or `--refresh-ocr` to force a real pipeline re-run instead of reusing cached OCR output).
|
||||
|
||||
---
|
||||
|
||||
## Repository Structure
|
||||
<a id="screens"></a>
|
||||
## 🖼️ The App, Screen by Screen
|
||||
|
||||
All screenshots below are real captures from a device running against the live backend — nothing is mocked.
|
||||
|
||||
| | | |
|
||||
|:---:|:---:|:---:|
|
||||
| <img src="screenshots/v2/02b-login-filled.png" width="230"><br>**Login** — one account per store | <img src="screenshots/v2/04b-camera-home-do-mode.png" width="230"><br>**Capture** — DO mode | <img src="screenshots/v2/10-image-preview-blur-check.png" width="230"><br>**Blur check** before upload |
|
||||
| <img src="screenshots/v2/11-history-pending-upload.png" width="230"><br>**Queue** — survives an app kill | <img src="screenshots/v2/12-document-parsed-ready-to-confirm.png" width="230"><br>**Parsed** — ready to confirm | <img src="screenshots/v2/09-camera-home-product-mode.png" width="230"><br>**Product scan** mode |
|
||||
|
||||
For the complete walkthrough — every page, button, popup, and the algorithms behind them — see [`screenshots/v2/WORKFLOW.md`](screenshots/v2/WORKFLOW.md) and the companion deck [`Prima-Mart-Scanner-Workflow.pptx`](screenshots/v2/Prima-Mart-Scanner-Workflow.pptx).
|
||||
|
||||
---
|
||||
|
||||
<a id="structure"></a>
|
||||
## 📂 Repository Structure
|
||||
|
||||
```
|
||||
.
|
||||
├── backend/
|
||||
│ ├── config/ # Pipeline & vLLM YAML configs
|
||||
│ ├── db/migrations/ # SQL run once by Postgres on a brand-new volume
|
||||
│ ├── pfm-web-app/ # Next.js API gateway + internal web tooling
|
||||
│ │ ├── src/app/api/ # All backend routes (upload, parse, documents, auth, arena...)
|
||||
│ │ ├── src/db/ # DB pool + init.ts (idempotent schema + seed data)
|
||||
│ │ ├── src/utils/parser.ts # Regex extraction + sanitization rules
|
||||
│ │ └── scripts/ # Accuracy-check tooling
|
||||
│ ├── sources/ # Test images, manual labels, SKU/store reference data
|
||||
│ ├── Dockerfile # Multi-stage GPU build (vllm-server, pipeline-api, pfm-web-app, gradio-ui)
|
||||
│ ├── nginx.conf
|
||||
│ └── docker-compose.yml # ⚠️ Legacy standalone copy — see "What to Consider" below
|
||||
├── lib/ # Flutter app
|
||||
│ ├── config/ # app_config.dart — theme + API base URL resolution
|
||||
│ ├── core/ # Dio client, Hive storage, location, router
|
||||
│ ├── features/ # auth, camera, documents, editor
|
||||
├── lib/ # ⭐ FLUTTER APP — the store-facing client
|
||||
│ ├── config/ # app_config.dart — theme + API URL resolution
|
||||
│ ├── core/ # Dio client, Hive storage, location, router
|
||||
│ ├── features/ # auth · camera · documents · editor · splash
|
||||
│ ├── data/ # master_sku.dart — SKU list compiled into the APK
|
||||
│ └── models/
|
||||
├── docs/ # Deep-dive workflow & extraction-rule docs
|
||||
├── screenshots/v2/ # App walkthrough: WORKFLOW.md + slide deck, real-backend screenshots
|
||||
├── test/ # Flutter widget/unit tests
|
||||
├── docker-compose.yml # ⭐ Canonical backend stack — run this one
|
||||
├── docker-compose.demo.yml # Production-mode override (see below)
|
||||
└── start-dev-tunnel.ps1 # Syncs LAN IP into app_config.dart + starts ngrok
|
||||
│
|
||||
├── backend/ # ⭐ EVERYTHING THAT RUNS ON THE SERVER
|
||||
│ ├── pfm-web-app/ # Next.js gateway + internal web tooling
|
||||
│ │ ├── src/app/api/v1/ # API consumed by the mobile app
|
||||
│ │ ├── src/app/api/ # Classic dev routes (no auth — internal only)
|
||||
│ │ ├── src/app/admin/ # Master-data pages
|
||||
│ │ ├── src/db/init.ts # Idempotent schema + account seeding
|
||||
│ │ ├── src/utils/parser.ts # Regex extraction + sanitization rules
|
||||
│ │ └── scripts/ # Accuracy-check tooling
|
||||
│ ├── config/ # Pipeline & vLLM YAML configs
|
||||
│ ├── db/migrations/ # SQL, run once on a brand-new Postgres volume
|
||||
│ ├── sources/ # 🔒 Test images, manual labels, master reference data
|
||||
│ ├── uploads/ # 🔒 Incoming DO photos + their OCR JSON
|
||||
│ ├── nginx.conf # Routing for every port
|
||||
│ ├── Dockerfile # Multi-stage GPU build (3 targets)
|
||||
│ └── docker-compose.yml # ⚠️ Legacy standalone copy — do not use
|
||||
│
|
||||
├── docs/ # Deep-dive workflow & extraction-rule docs
|
||||
│ └── handover/ # 📚 Handover deck + documents (see below)
|
||||
├── screenshots/v2/ # Real-device walkthrough + slide deck
|
||||
├── test/ # Flutter widget/unit tests
|
||||
├── docker-compose.yml # ⭐ Canonical stack — run THIS one, from the repo root
|
||||
├── docker-compose.demo.yml # Production-mode override
|
||||
└── start-dev-tunnel.ps1 # Sync LAN IP into app_config.dart + start ngrok
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Getting Started
|
||||
<a id="demo-mode"></a>
|
||||
## 🚀 Development vs. Demo Mode
|
||||
|
||||
### Prerequisites
|
||||
> [!CAUTION]
|
||||
> **You MUST use the production override for any client demo or field test.** The default stack runs the gateway through `npm run dev` with source bind-mounted for hot-reload — a single dev-server process with a known throughput ceiling that will bottleneck the moment several devices upload at once.
|
||||
|
||||
| Component | Requirement |
|
||||
| Mode | Command | Behaviour |
|
||||
|---|---|---|
|
||||
| **Development** | `docker compose up --build` | `npm run dev`, source bind-mounted, hot-reload. Edit `parser.ts` and see it live. |
|
||||
| **Demo / Production** | `docker compose -f docker-compose.yml -f docker-compose.demo.yml up -d --build` | `npm start` against the image's own build output. **No hot-reload.** |
|
||||
|
||||
`--build` is mandatory in demo mode, not optional — the override drops the source mount, so without a rebuild you are running whatever code was baked into the previous image.
|
||||
|
||||
---
|
||||
|
||||
<a id="gotchas"></a>
|
||||
## ⚠️ Gotchas — Read Before You Install
|
||||
|
||||
<details open>
|
||||
<summary><b>Deployment & configuration</b></summary>
|
||||
|
||||
<br>
|
||||
|
||||
- **Two `docker-compose.yml` files exist — always run from the repo root.** `backend/docker-compose.yml` is a near-duplicate standalone copy under project name `ai-ocr-pfm-2026` (it was a separate repo, vendored in). Compose names containers and volumes after whichever file you invoke, so running it from inside `backend/` produces container-name conflicts against anything already up from the root. This is not hypothetical — it happened during this project's own testing.
|
||||
|
||||
- **First build is large and slow.** ~30 GB per GPU image, plus weights. Budget the disk and the time once.
|
||||
|
||||
- **Single GPU pipeline, no horizontal scaling.** One `pipeline-api`, one `vllm-server`. Concurrent uploads queue behind the GPU. Load-test this before a multi-store rollout — a single-user smoke test will not reveal the ceiling.
|
||||
|
||||
- **`docker compose down -v` destroys the database.** The `-v` flag drops `paddleocr_pgdata` and the model caches. There is no automatic backup.
|
||||
|
||||
</details>
|
||||
|
||||
<details open>
|
||||
<summary><b>Security posture</b></summary>
|
||||
|
||||
<br>
|
||||
|
||||
- **`/api/v1/*` enforces real auth; the classic dev routes deliberately do not.** `/api/v1/auth/login` checks a bcrypt hash against the `accounts` table and signs a JWT; `/api/v1/documents/*` reject missing/invalid tokens with a real `401`. The Flutter app only ever uses this surface. The classic routes (`/api/upload`, `/api/parse`, `/api/history`) and the `scan-pfm` / `manual-label` pages have no login flow and never will — they are dev tooling. Every route still sets `Access-Control-Allow-Origin: *`: fine on a controlled LAN, not fine facing the open internet.
|
||||
|
||||
- **Seeded passwords reset on every restart.** `init.ts` seeds accounts with `ON CONFLICT (username) DO UPDATE SET password`, and it runs whenever the gateway first touches the database. Every backend restart therefore resets **all** account passwords back to the seeded default. There is no password-change flow yet. Fix this before any wide production rollout.
|
||||
|
||||
- **`JWT_SECRET` has an insecure fallback.** Leave it unset in `backend/.env` and the code falls back to a constant written in the source, which would let anyone holding it mint tokens for any store. Set a real random value.
|
||||
|
||||
- **The LAN leg is plain HTTP.** The ngrok leg is HTTPS; `_lanBaseUrl` is not. Bearer tokens and DO contents cross the store Wi-Fi unencrypted.
|
||||
|
||||
- **The Android release build is debug-signed.** `android/app/build.gradle.kts` still carries `// TODO: Add your own signing config`. Fine for internal installs, not for Play Store distribution.
|
||||
|
||||
</details>
|
||||
|
||||
<details open>
|
||||
<summary><b>Data & master tables</b></summary>
|
||||
|
||||
<br>
|
||||
|
||||
- **`store_master` ships empty.** The schema is created automatically but no store rows are seeded, so `resolveStoreFromText` resolves nothing until you load real data. `backend/sources/toko_aktif.json` looks like the right source but is not wired to an automatic import yet.
|
||||
|
||||
- **`backend/sources/` and `backend/uploads/` hold real client data** — master SKU/vendor/customer records, genuine DO photographs, hand-labeled ground truth. Do not export, log, or forward them. See the Confidentiality section of [`backend/CLAUDE.md`](backend/CLAUDE.md).
|
||||
|
||||
- **Mobile connectivity is two independently-maintained paths.** The ngrok domain is reserved but the ngrok *process* is not auto-started; the LAN IP is hardcoded and goes stale the moment DHCP moves. When a device cannot log in, check these two before anything else.
|
||||
|
||||
</details>
|
||||
|
||||
---
|
||||
|
||||
<a id="testing"></a>
|
||||
## 🧪 Testing & Tooling
|
||||
|
||||
| Target | Command |
|
||||
|---|---|
|
||||
| Backend host | NVIDIA GPU, CUDA 12.6+ driver, ~8GB+ VRAM. Tested working on both native Linux and **Windows + Docker Desktop with WSL2 GPU passthrough**. |
|
||||
| Docker | Docker Engine/Desktop with the NVIDIA Container Toolkit (`docker info` should list `nvidia` under Runtimes) |
|
||||
| Disk space | **60GB+ free** — the pipeline-api and vllm-server images alone are ~30GB each once built, plus model weight caches |
|
||||
| Flutter | Flutter SDK `>=3.2.0 <4.0.0`, Android Studio/Xcode for device tooling |
|
||||
| Node.js | v20+ (optional — only for running the parser unit tests or accuracy tooling outside Docker) |
|
||||
|
||||
### 1. Backend Setup
|
||||
|
||||
1. **Configure environment variables** (from the repo root):
|
||||
```bash
|
||||
cp backend/.env.example backend/.env
|
||||
```
|
||||
Set `CUDA_VISIBLE_DEVICES` to your GPU index, and `APP_PORT` if `8000` is taken.
|
||||
|
||||
2. **Start the stack from the repo root** (not `backend/` — see the gotcha below):
|
||||
```bash
|
||||
docker compose up --build
|
||||
```
|
||||
First build pulls/builds ~60GB of GPU images and downloads model weights — expect this to take a long time on the first run. Subsequent starts are fast.
|
||||
|
||||
3. **Verify it's up**:
|
||||
```bash
|
||||
curl http://localhost:8000/health # pipeline API health
|
||||
curl http://localhost:8000/v1/models # vLLM model list
|
||||
curl -X POST http://localhost:8000/api/v1/auth/login \
|
||||
-H "Content-Type: application/json" -d '{"username":"admin","password":"password"}'
|
||||
```
|
||||
|
||||
Database schema and reference data (vendor, customer, a starter SKU catalog) are created **automatically** on first request — `backend/pfm-web-app/src/db/init.ts` runs idempotent `CREATE TABLE IF NOT EXISTS` + seed statements every time the app starts, so there is no manual migration step. `backend/db/migrations/*.sql` also runs once via Postgres's own `docker-entrypoint-initdb.d` on a brand-new volume, but `init.ts` is what you should treat as the source of truth.
|
||||
|
||||
### 2. Frontend Setup
|
||||
|
||||
1. **Install dependencies**:
|
||||
```bash
|
||||
flutter pub get
|
||||
```
|
||||
|
||||
2. **Point the app at your backend.** `lib/config/app_config.dart` resolves the API URL dynamically at startup: it first tries a public ngrok tunnel, and falls back to a hardcoded LAN URL if that's unreachable. Both need to match your actual machine:
|
||||
```bash
|
||||
./start-dev-tunnel.ps1
|
||||
```
|
||||
This detects your current LAN IP, patches `_lanBaseUrl` in `app_config.dart` for you, and launches `ngrok` pointed at the fixed reserved domain in that same file. Run it again any time your IP changes (new network, DHCP renewal, etc.) — a stale IP here is the single most common reason "the app can't log in" after this backend is confirmed healthy.
|
||||
|
||||
3. **Run it**:
|
||||
```bash
|
||||
flutter run
|
||||
```
|
||||
|
||||
4. **Build a release APK** when you need an installable build instead of a debug session:
|
||||
```bash
|
||||
flutter build apk --release
|
||||
```
|
||||
Output: `build/app/outputs/flutter-apk/app-release.apk`. See the signing note below before distributing it.
|
||||
|
||||
---
|
||||
|
||||
## Running in Demo / Production Mode
|
||||
|
||||
**POLICY: You MUST run the production mode override for any client demos or field testing.**
|
||||
|
||||
The default `docker-compose.yml` runs the gateway via `npm run dev` with the source bind-mounted in. This is strictly for local development (enabling hot-reloading for iterating on `parser.ts`). It is a single dev-server process with a known throughput ceiling and will bottleneck if multiple people upload at once.
|
||||
|
||||
To run the production build instead, use this opt-in override:
|
||||
|
||||
```bash
|
||||
docker compose -f docker-compose.yml -f docker-compose.demo.yml up -d --build
|
||||
```
|
||||
|
||||
This drops the dev bind-mount and runs `npm start` against the image's own `npm run build` output. **Rebuild (`--build`) before every demo** — this mode does not hot-reload code changes.
|
||||
|
||||
---
|
||||
|
||||
## What to Consider Before You Install
|
||||
|
||||
- **Two `docker-compose.yml` files exist — always run from the repo root.** `backend/docker-compose.yml` is a near-duplicate, standalone copy of the same stack (project name `ai-ocr-pfm-2026`, originally a separate repo vendored into `backend/`). Docker Compose names containers/volumes after the project name declared in whichever compose file you invoke first. If you ever run `docker compose up` from inside `backend/`, you'll get container-name conflicts against anything already started from the root — this isn't hypothetical, it happened during this project's own testing. Pick one (the root file) and stick to it.
|
||||
|
||||
- **`store_master` ships empty.** Schema is created automatically, but no store data is seeded — fuzzy store-matching (`resolveStoreFromText`) will not resolve any delivery address until you load real store data into that table yourself. `backend/sources/toko_aktif.json` looks like the right source for this but isn't wired into an automatic import yet.
|
||||
|
||||
- **First build is large and slow.** The pipeline-api and vllm-server images are ~30GB each with GPU model weights. Budget real time and disk space for the first `docker compose up --build`.
|
||||
|
||||
- **Single GPU pipeline, no horizontal scaling.** There's one `pipeline-api` and one `vllm-server` container. Concurrent uploads queue behind the GPU; this is a real throughput ceiling worth load-testing before a multi-store/multi-device demo, not just a single-user smoke test.
|
||||
|
||||
- **`/api/v1/*` enforces real auth; the classic dev routes deliberately don't.** `/api/v1/auth/login` checks a bcrypt-hashed password against a real `accounts` table and signs a JWT; `/api/v1/documents/*` (list, upload, PUT-by-id) reject any request with a missing/invalid token with a real `401`. The Flutter app always goes through this surface. The **classic routes** (`/api/upload`, `/api/parse`, `/api/history`, etc.) and the root/`scan-pfm`/`manual-label` web pages have no login flow and never will — they're dev-only internal tooling, not part of the store-facing product. Every API route still sets `Access-Control-Allow-Origin: *`, so this is fine for a controlled LAN/demo deployment but not for exposing the stack to the open internet as-is.
|
||||
|
||||
- **The Android release build is debug-signed.** `android/app/build.gradle.kts` has a `// TODO: Add your own signing config` and currently signs release builds with the debug key. Fine for internal install/testing, not for Play Store distribution.
|
||||
|
||||
- **Mobile connectivity is two independent, manually-synced paths.** The ngrok domain in `app_config.dart` is fixed/reserved, but the ngrok *process* isn't started automatically — you (or `start-dev-tunnel.ps1`) have to launch it. The LAN IP fallback is hardcoded and will silently go stale the moment your host machine's IP changes. If login/upload fails on a real device, check these two before anything else.
|
||||
|
||||
---
|
||||
|
||||
## Testing & Tooling
|
||||
|
||||
| What | How |
|
||||
|---|---|
|
||||
| Parser regex/sanitization rules | `npx tsx backend/pfm-web-app/src/utils/parser.test.ts` — no Docker needed |
|
||||
| OCR accuracy regression suite | `backend/pfm-web-app/run_batch_test.js` and `backend/pfm-web-app/scripts/accuracy-check.mts`, run against `backend/sources/test-images` + `manual_labels.json` |
|
||||
| Flutter widget/unit tests | `flutter test` (covers blur detection, camera navigation, editor validation, geotagging, the pending-upload queue, and more — see `test/`) |
|
||||
| Parser regex / sanitization rules | `npx tsx backend/pfm-web-app/src/utils/parser.test.ts` — no Docker needed |
|
||||
| OCR accuracy regression suite | `backend/pfm-web-app/scripts/accuracy-check.mts` and `run_batch_test.js`, against `backend/sources/test-images` + `manual_labels.json` |
|
||||
| Flutter widget / unit tests | `flutter test` — blur detection, camera navigation, editor validation, geotagging, pending-upload queue |
|
||||
| Single Flutter test file | `flutter test test/blur_detector_test.dart` |
|
||||
| Static analysis | `flutter analyze lib` |
|
||||
|
||||
---
|
||||
|
||||
## Known Limitations
|
||||
<a id="documentation"></a>
|
||||
## 📚 Documentation
|
||||
|
||||
Previously tracked here as open gaps, both now fixed and worth noting as resolved:
|
||||
- ~~Pending-upload queue is in-memory only~~ — it's now persisted to a local Hive box (`lib/core/storage/local_storage.dart`), so an OS-level app kill mid-upload no longer loses the document; the queue reloads and resumes on next launch.
|
||||
- ~~Item-row writes on document save aren't wrapped in a transaction~~ — `PUT /documents/:id` now runs inside `withTransaction` (`backend/pfm-web-app/src/app/api/v1/documents/[id]/route.ts`).
|
||||
| Document | Format | What it covers |
|
||||
|---|:---:|---|
|
||||
| [**Ringkasan Ruang Lingkup**](docs/handover/PFM-Scanner-Ringkasan-Scope.pdf) | `PDF` 20 pp | 🇮🇩 **Start here.** Flow, folder map, ports, how to run each part and from which directory, how login and upload move data, and what is built versus what is not. |
|
||||
| [**Handover Document**](docs/handover/PFM-Scanner-Handover-Document.pdf) | `PDF` 42 pp | Full handover manual — operational runbook, troubleshooting guide, glossary, screenshot appendix. |
|
||||
| [**Knowledge Transfer Deck**](docs/handover/PFM-Scanner-Knowledge-Transfer.pdf) | `PDF` 30 slides | Walkthrough deck for a live handover session. Diagrams are native shapes, editable in Impress/Draw. |
|
||||
| [**App Walkthrough**](screenshots/v2/WORKFLOW.md) | `MD` | Screen-by-screen tour with the algorithm behind each one, captured against the live backend. |
|
||||
| [**Extraction Workflow**](docs/workflow_detail_aplikasi.md) | `MD` | Field-by-field extraction rules end to end. |
|
||||
| [**Regex Rules**](docs/regex_rules_example.md) | `MD` | Worked examples of every pattern and sanitization step. |
|
||||
| [**API Contract Map**](docs/api-contract-map.md) | `MD` | Every endpoint, its payload, and known contract gaps. |
|
||||
| [**Feature List**](docs/feature-list.md) · [**Iteration Log**](docs/iteration-log.md) | `MD` | What shipped, and the audit trail behind each change. |
|
||||
|
||||
Still open, worth knowing before relying on this for unattended field use:
|
||||
- The **DO Scan review form's** validation fails silently on submit if a required field (e.g. driver name) is empty — `Form.validate()` returns false and the confirm button's `onPressed` just returns, with no toast or scroll-to-error to tell the operator why nothing happened.
|
||||
Working rules for anyone (human or AI agent) editing this repo live in [`CLAUDE.md`](CLAUDE.md) and [`AGENTS.md`](AGENTS.md) for the Flutter app, and separately in [`backend/CLAUDE.md`](backend/CLAUDE.md) / [`backend/AGENTS.md`](backend/AGENTS.md) for the backend. **They are deliberately not interchangeable** — do not apply one set to the other subtree.
|
||||
|
||||
See `docs/` and `screenshots/v2/WORKFLOW.md` for the deeper workflow documentation these decisions were audited against.
|
||||
---
|
||||
|
||||
<a id="troubleshooting"></a>
|
||||
## 🔧 Troubleshooting
|
||||
|
||||
<details>
|
||||
<summary><b>The app cannot log in or upload</b></summary>
|
||||
|
||||
<br>
|
||||
|
||||
| Symptom | Cause & fix |
|
||||
|---|---|
|
||||
| Login times out on a real device | The compiled LAN IP is stale or ngrok is not running. Run `./start-dev-tunnel.ps1`, then **rebuild and reinstall the APK** — the URL is a compile-time constant. |
|
||||
| Works on emulator, not on phone | Emulator needs `10.0.2.2`, a physical device needs the host's real LAN IP. They cannot share one value. |
|
||||
| `401` on every request | Token expired (30-day JWT), or the backend restarted and reset the seeded passwords. Log in again. |
|
||||
| Login succeeds but uploads hang | Check `docker compose logs -f pipeline-api` and `vllm-server`. A cold GPU model load takes minutes. |
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>Backend will not start</b></summary>
|
||||
|
||||
<br>
|
||||
|
||||
| Symptom | Cause & fix |
|
||||
|---|---|
|
||||
| `could not select device driver` | NVIDIA Container Toolkit missing. `docker info` must list `nvidia` under Runtimes. |
|
||||
| Container name conflicts | You ran `docker compose` from inside `backend/`. Stop everything, then run from the repo root only. |
|
||||
| `pfm-web-app` never starts | It waits for `pipeline-api` to report healthy, which waits on `vllm-server`. Watch `docker compose logs -f vllm-server` — first-time weight download is slow. |
|
||||
| Out of disk mid-build | Images are ~30 GB each. Free space and rebuild. |
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>OCR results are wrong or empty</b></summary>
|
||||
|
||||
<br>
|
||||
|
||||
| Symptom | Cause & fix |
|
||||
|---|---|
|
||||
| Store field always empty | `store_master` is unpopulated — see [Gotchas](#gotchas). |
|
||||
| Item names look like OCR noise | The SKU triple-check could not match `sku_master`. Confirm the SKU catalogue covers those products. |
|
||||
| Whole document skewed / garbled | The DO was photographed on top of other papers, breaking deskew. Reshoot on a plain flat surface. |
|
||||
| Document stuck "processing" | Poll gives up after ~4.3 min. Check `parse_error` on the row and the `pipeline-api` logs. |
|
||||
|
||||
</details>
|
||||
|
||||
---
|
||||
|
||||
<a id="limitations"></a>
|
||||
## 📋 Known Limitations
|
||||
|
||||
**Resolved, kept here for context:**
|
||||
- ~~Pending-upload queue was in-memory only~~ — now persisted to a Hive box (`lib/core/storage/local_storage.dart`); an OS-level app kill mid-upload no longer loses the document, the queue reloads and resumes on next launch.
|
||||
- ~~Item-row writes on save were not transactional~~ — `PUT /documents/:id` now runs inside `withTransaction`.
|
||||
|
||||
**Still open, worth knowing before unattended field use:**
|
||||
- The **DO review form fails silently on submit** when a required field (e.g. driver name) is empty — `Form.validate()` returns false and the confirm button simply returns, with no toast and no scroll-to-error explaining why nothing happened.
|
||||
- The **blur check is advisory only** — a photo flagged blurry can still be uploaded; the badge does not gate the button.
|
||||
- **SKUs are validated against a bundled list** (`lib/data/master_sku.dart`) rather than the live `/api/v1/master/skus` endpoint, so a newly added SKU fails client-side validation until the APK is rebuilt.
|
||||
|
||||
A fuller backlog — 51 open items grouped by impact — lives in [`plans/next-enhancements.md`](plans/next-enhancements.md), and is summarized in plain language in the [Ringkasan Ruang Lingkup](docs/handover/PFM-Scanner-Ringkasan-Scope.pdf) PDF.
|
||||
Reference in new issue
Block a user