docs: add scan-product reference and update backend/root plans for enhancements

This commit is contained in:
Rafhan Mazaya Fathurrahman committed 2026-07-08 14:50:32 +07:00
1 parent e60ab63154
commit 9ff4a4a922
4 files changed
+475 -2

No files matched your search

+1 -1
View File
@@ -26,7 +26,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co
## Product/SKU scanning flow — status
See [`plans/next-enhancements.md`](plans/next-enhancements.md) §2 (task 2.1) for full detail — kept there instead of a separate doc so status stays traceable against the rest of the `e`/`n` backlog. **Feature-complete as of 2026-07-08**: the backend (`config/classify_ocr_server.py` with DINOv2 similarity search + YOLO classifier fallback, `api/scan-pfm/route.ts`, `api/produk-pfm/route.ts`, DB schema), the reference photo dataset (`pfm-web-app/public/produk-pfm/foto-kemasan-v2/`, 16 SKU subfolders), the desktop frontend page (`scan-pfm/page.tsx`, full feature parity), and the trained model artifacts (`models/dinov2_index.pkl` — 118/118 photos indexed; `models/produk-pfm-classifier-26n-100e-2026-07-08.pt` — 83.3% top-1 val accuracy on the current thin dataset) all now exist and load cleanly on `pipeline-api` startup. **No mobile web page is planned**: `scan-pfm/page.tsx` is desktop-only, used to test the pipeline; real mobile product scanning goes through the Flutter app instead, so `m-scan-pfm/page.tsx` and its `nginx.conf` route are intentionally left unbuilt/dead (see plan task 2.2, cancelled 2026-07-08). Not yet done: an actual browser pass uploading a photo through `/scan-pfm` end-to-end (verified via container logs/model-loading so far, not a UI test).
**How it works end-to-end** (architecture, endpoints, classification/OCR internals, retraining): [`docs/scan-product.md`](docs/scan-product.md). See [`plans/next-enhancements.md`](plans/next-enhancements.md) §2 (task 2.1) for full detail — kept there instead of a separate doc so status stays traceable against the rest of the `e`/`n` backlog. **Feature-complete as of 2026-07-08**: the backend (`config/classify_ocr_server.py` with DINOv2 similarity search + YOLO classifier fallback, `api/scan-pfm/route.ts`, `api/produk-pfm/route.ts`, DB schema), the reference photo dataset (`pfm-web-app/public/produk-pfm/foto-kemasan-v2/`, 16 SKU subfolders), the desktop frontend page (`scan-pfm/page.tsx`, full feature parity), and the trained model artifacts (`models/dinov2_index.pkl` — 118/118 photos indexed; `models/produk-pfm-classifier-26n-100e-2026-07-08.pt` — 83.3% top-1 val accuracy on the current thin dataset) all now exist and load cleanly on `pipeline-api` startup. **No mobile web page is planned**: `scan-pfm/page.tsx` is desktop-only, used to test the pipeline; real mobile product scanning goes through the Flutter app instead, so `m-scan-pfm/page.tsx` and its `nginx.conf` route are intentionally left unbuilt/dead (see plan task 2.2, cancelled 2026-07-08). Not yet done: an actual browser pass uploading a photo through `/scan-pfm` end-to-end (verified via container logs/model-loading so far, not a UI test).
## Confidentiality
+185
View File
@@ -0,0 +1,185 @@
# Product Scan (scan-pfm) — How It Works
End-to-end reference for the Product/SKU scanning feature: a photo of a Primafood
product package goes in; the SKU class, product name, expiry date, and a ranked
SKU-master match list come out. Written 2026-07-08 against the live code. Related:
`plans/next-enhancements.md` §2 (build history) and §6 (ground-truth roadmap);
`docs/feature-list.md` tasks 2.1/2.3.
## High-level flow
```mermaid
flowchart LR
A[Browser: /scan-pfm page] -->|"POST /api/scan-pfm {image_base64}"| B[Next.js gateway<br/>pfm-web-app :3000]
B -->|"POST :8120/classify-ocr"| C[classify_ocr_server.py<br/>FastAPI, in pipeline-api]
C --> C1[1. DINOv2 similarity<br/>fallback: YOLO classifier]
C --> C2[2. PaddleOCR + regex<br/>SKU / expiry / name]
C -->|"POST localhost:8090/layout-parsing<br/>promptLabel: spotting"| D[PaddleX pipeline<br/>same container]
B -->|"POST :8090/layout-parsing"| D
B -->|"SELECT sku_master"| E[(Postgres)]
B -->|Levenshtein ranking| A
```
Two processes live in the `paddleocr-pipeline-api` container, both started by
`scripts/serve-pipeline.sh`: the PaddleX layout-parsing pipeline on **:8090**
(shared with the DO flow; VL recognition goes out to the vLLM server on :8118) and
`config/classify_ocr_server.py` on **:8120** (product scan only). The gateway
reaches them via Docker DNS (`CLASSIFIER_SERVER_URL`, `PIPELINE_URL` in root
`docker-compose.yml:87-88`); nginx (:8000) proxies `/scan-pfm` to the Next.js app.
## Request walkthrough
1. **Page** (`pfm-web-app/src/app/scan-pfm/page.tsx`, desktop-only test UI): pick a
sample from the gallery (`GET /api/produk-pfm`) or upload/rotate a photo (rotation
is done client-side on a canvas), then send it as a base64 data-URL.
2. **Gateway** (`api/scan-pfm/route.ts`):
- forwards `{image_base64}` to the classifier server (`/classify-ocr`);
- separately calls the layout-parsing pipeline with `useLayoutDetection: true`
for the Visual Grid tab's output images (failure here is non-fatal — logged,
`layoutParsingResult` returns `null`);
- loads the full `sku_master` table and ranks every SKU by **Levenshtein
similarity between `nama_item` and the classifier's `top1_name`**
(lowercased, alphanumerics only). Top 5 with score > 0.1 are returned;
rank 1 gets `isBestMatch: true`. Note: `ocr.extracted_sku` and
`ocr.extracted_product_name` are read but **not used** in this ranking —
see Future recommendations.
3. **Classifier server** (`config/classify_ocr_server.py`) does classification,
OCR extraction, and visualization — detailed below — and returns
`{classification, ocr}`.
4. **Page renders** four tabs: Summary (classification card + top-5 override
"Use" buttons + OCR fields + SKU matches), Visual Grid, Spotting Grid, Raw
Response (JSON). "Save Ground Truth" posts to `/api/manual-label-scan`.
## Stage 1 — classification (which product is this?)
**Primary: DINOv2 similarity search** (`method: "dinov2_similarity"`). At startup
the server loads `dinov2_vits14` **from `torch.hub` (network fetch on first run)**
plus `models/dinov2_index.pkl` — precomputed L2-normalized 384-dim embeddings of
all 118 reference photos across 16 SKU class folders. Per request: embed the query
image (resize 224², ImageNet normalization), dot-product against all reference
embeddings (= cosine similarity), then aggregate **per class = max similarity of
any reference photo in that class**. Classes sorted by similarity become
`all_probabilities`. Caveat: these "confidences" are cosine similarities, **not
probabilities** — they don't sum to 1 and are typically all high (0.4–0.9);
compare relatively, not against an absolute threshold.
**Fallback: YOLO classifier** (`method: "yolo_classifier"`) — only when DINOv2 is
unavailable (no index/model) or throws. A fine-tuned `yolo26n-cls` checkpoint;
its `all_probabilities` are real softmax probabilities. Weights are
**auto-discovered**: `CLASSIFIER_MODEL_PATH` env wins; otherwise the newest
`produk-pfm-classifier-26n-*e-*.pt` in `models/` by (date-in-filename, mtime) —
so retraining just drops a new dated file, no config change.
If both are unavailable, `classification` carries an `error` field instead.
## Stage 2 — OCR extraction (SKU, expiry date, product name)
PaddleOCR (`lang='en'`, textline orientation on) produces `rec_texts` lines +
`rec_polys` boxes. Three extractors run over the lines:
- **SKU** (`extract_sku`): first 8-digit number anywhere; else first 7–9 digit
number. (Primafood SKUs are 8 digits, printed near the label top.)
- **Expiry date** (`extract_expired_date`): each line is first noise-cleaned
(`clean_date_line`: `1)`→`0`, `()`→`0`, `B8/8B/88`→`BB` before digits, o→0,
I/l/|→1, S→5, Z→2, B→8 when digit-flanked, plus `012`/`112` month-misread
repairs), then a **6-level priority cascade** runs: (1) BB/EXP-keyword line
with compact `DDMMYYYY`; (2) keyword line with spaced `DD MM YYYY`; (3)
keyword + 6–8 digit run; (3.5) keyword line, lenient noisy match; (4) any line
spaced date; (5) any line compact `DDMMYYYY` — skipping lines that look like a
SKU-on-product-name; (6) legacy formats (slashes, `05 MAR 2027`). Recognized
keywords: `EXP`, `EXPIRED`, `TGL`, `EXPIRY`, `BBD`, `BEST BEFORE`, `BB`,
`BAIK DIGUNAKAN`. Output normalized to `DD/MM/YYYY`.
- **Product name** (`extract_product_name`): longest line containing a brand/
product keyword (FIESTA, CHAMP, OKEY, AKUMO, ASIMO, NUGGET, SOSIS, …) after
stripping SKU digits and date fragments; falls back to the classifier's
`top1_name`, then the longest non-numeric line, then `"Unknown Product"`.
Visualization artifacts built server-side: `vis_image_base64` (all OCR boxes
drawn teal `TEXT`, the expiry line amber `EXP`, on the orientation-corrected
image so boxes align), `expired_date_crop_base64` (padded crop of the expiry
line for eyeball verification — `find_expired_crop_index` prefers the box whose
digits actually contain the date), and `spotting_image_base64` (a second
pipeline call with `promptLabel: "spotting"`, no layout detection).
## Endpoint reference
| Endpoint | Where | Purpose |
|---|---|---|
| `POST /api/scan-pfm` | gateway | Main scan. Body `{image_base64}` (data-URL ok). Returns `{classification, ocr, possibleMatches[], layoutParsingResult}` |
| `POST http://paddleocr-pipeline-api:8120/classify-ocr` | classifier server | Internal. Body `{image_base64}`. Returns `{classification: {top1_name, top1_confidence, all_probabilities[], method}, ocr: {text_lines[], extracted_product_name, extracted_sku, extracted_expired_date, expired_line_index, expired_source_line, expired_date_crop_base64, vis_image_base64, spotting_image_base64}}` |
| `GET /api/produk-pfm` | gateway | Gallery: SKU folders under `public/produk-pfm/foto-kemasan-v2/` with image + thumb URLs |
| `GET/POST /api/manual-label-scan` | gateway | Ground-truth read/upsert to `sources/product_manual_labels.json` (host-visible via the `./backend/sources:/sources` mount) |
| `POST :8090/layout-parsing` | pipeline | Shared PaddleX pipeline; used here for Visual Grid images and (with `promptLabel: "spotting"`) the Spotting Grid |
| `/scan-pfm` | nginx :8000 | Proxies the page to Next.js :3000 |
`possibleMatches[]` items: `{no_sku, nama_item, score, yoloSimilarity, isBestMatch}` —
`score` currently equals `yoloSimilarity` (name-vs-name Levenshtein, 0..1).
## Model artifacts & retraining
| File (`pfm-web-app/public/produk-pfm/`) | What |
|---|---|
| `foto-kemasan-v2/<SKU or class>/…` | Reference photo dataset — 16 classes, 118 photos (2–16 each) |
| `models/dinov2_index.pkl` | DINOv2 embeddings + metadata (rebuild after adding photos) |
| `models/produk-pfm-classifier-26n-100e-2026-07-08.pt` / `.onnx` | Fine-tuned YOLO classifier (83.3% top-1 / 90% top-5 val on the thin dataset) |
| `index_dinov2.py` | Rebuilds the pickle index from `foto-kemasan-v2/` |
| `train_classifier.py` | Splits 80/20 into `yolo_dataset/`, fine-tunes `yolo26n-cls.pt` (default 100 epochs, `--imgsz 224`), writes a dated checkpoint |
**Retraining procedure (Windows host — bare-metal doesn't work here,
`paddlepaddle-gpu` wheels are Linux-only):** add photos to `foto-kemasan-v2/`,
`docker compose build pipeline-api` from the **repo root**, run a one-off
`docker run --gpus all` from that image with `models/` mounted **writable** (the
live service mounts it `:ro`), run `index_dinov2.py` then
`train_classifier.py train --imgsz 224`, then `docker compose restart
pipeline-api`. From Git Bash prefix `MSYS_NO_PATHCONV=1` or `/app/...` arguments
get mangled. Verify in `docker logs`: "DINOv2 index loaded with N reference
images", "Using classifier weights: <new dated file>". Full worked example:
`plans/next-enhancements.md` task 2.1.
## Operational notes
- **Env vars**: `CLASSIFIER_SERVER_URL`, `PIPELINE_URL` (gateway, set in compose);
`CLASSIFIER_MODELS_DIR`, `CLASSIFIER_MODEL_PATH` (classifier server overrides).
The gateway's in-code default `PIPELINE_URL` (`localhost:7871`) is stale — the
compose env always overrides it in Docker.
- **Startup order/health**: the classifier server loads DINOv2 (torch.hub →
needs network/cache), YOLO, and PaddleOCR at import time; until done, :8120
refuses connections and `/api/scan-pfm` 500s. No healthcheck exists yet (plan
task 4.2 / 1.6).
- **GPU**: DINOv2 + YOLO + PaddleOCR share the container/GPU with the PaddleX
pipeline; all are small (ViT-S/14, nano YOLO) next to the vLLM server's
footprint, but they do add VRAM on the same `PIPELINE_DEVICE`.
- **Failure isolation**: layout-vis and spotting calls are best-effort
(`null`/absent on failure); classification and OCR errors surface as `error`
fields inside their sections rather than failing the whole scan.
## Known gaps & future recommendations
Tracked ones (see `plans/next-enhancements.md`):
- **§6.1–6.3 ground truth**: today's "Save Ground Truth" stores the model's own
predictions (only the SKU is editable) and uploads get phantom
`uploaded-<timestamp>.jpg` keys with no image persisted; §6 plans the editable
annotation page, persisted uploads, and a scan accuracy harness.
- **Dataset thinness**: 2–16 photos/class caps both classifiers; every new real
photo (especially non-studio, in-warehouse shots) matters. The §6.3 harness
should report gallery vs. uploaded-photo accuracy separately — gallery photos
are training data, so scores on them measure memorization.
Additional recommendations (not yet tasks — promote via `e`/`n` when wanted):
1. **Use `extracted_sku` in match ranking.** An exact 8-digit SKU hit read off
the label is far stronger evidence than fuzzy name similarity, yet ranking
currently ignores it. Suggested: exact `no_sku` match pins rank 1; blend name
similarity for the rest.
2. **Fuse DINOv2 and YOLO instead of primary/fallback** (e.g. agreement boosts
confidence; disagreement flags for review) — cheap, both already load.
3. **"Not a known product" handling**: DINOv2 always returns *some* class; add a
minimum-similarity threshold below which the response says unknown rather
than confidently misclassifying a foreign package.
4. **Pin the DINOv2 backbone offline** (vendor the weights or pre-bake the
torch.hub cache into the image) — startup currently depends on an internet
fetch on cold cache, bad for on-prem deploys.
5. **Batch/lot number extraction** — explicitly out of scope so far (plan §2
note); if requested, follow the expiry-date regex-cascade pattern.
6. **Mobile**: no web mobile page by design (task 2.2 cancelled) — real mobile
scanning should go through the Flutter app calling `POST /api/scan-pfm`
(would need an authenticated `/api/v1` variant; the classic route has no auth).
+242
View File
@@ -53,6 +53,22 @@ flips to `[DONE]` and the feature is logged in
- **1.1** [DONE] ~~Make `documents/upload/route.ts` return `201` immediately...~~ Investigated 2026-07-08: the `file_hash` dedup check is already shipped pre-existing in `v1/documents/upload/route.ts` (lines 58-89) — a duplicate upload returns the existing document instead of re-inserting/re-parsing. The "return 201 immediately, don't await OCR" half turned out to be a deliberate, already-documented tradeoff (see the route's own comment): the pipeline call stays synchronous but is now bounded by `AbortSignal.timeout(210_000)`. Not changed further this round since it's an intentional decision, not an oversight.
- **1.2** [DONE] Investigated 2026-07-08: already shipped pre-existing. Both hops of the upload → `/api/parse` → pipeline-API chain already have `AbortSignal.timeout` (210s and `PIPELINE_TIMEOUT_MS`=90s respectively), with comments cross-referencing each other.
- **1.3** [CANCELLED 2026-07-08] ~~Wire the existing JWT auth onto the classic API routes (`/api/upload`, `/api/scan-pfm`, `/api/parse`, `/api/history`, etc.)~~ — investigated after 3.2 unblocked it, found the premise was stale: those "classic" routes are used exclusively by the dev-only web UI (root DO-PFM page, `scan-pfm`, `manual-label`), which has no login screen and never sends a token — the user confirmed this web UI won't exist in production, so enforcing auth there would break it today for zero production benefit, and optional identification there is a no-op (nobody ever sends a token). Same reasoning as 2.2's cancellation. See **1.4** below for what was actually shipped instead.
- **1.5** [TODO] Per-account data scoping on `/api/v1/documents/*` — deferred from
task 1.4, which added authentication but not authorization: `GET /api/v1/documents`
still returns *all* non-sample parsed documents to any authenticated account, and
`PUT /api/v1/documents/:id` doesn't check the document belongs to the caller's
`kode_toko`. Decide the scoping rule (per-account vs. per-store) with the user
first — it's a behavior change for the Flutter document list, not just a filter.
- **1.6** [TODO] **Lightweight `GET /api/v1/health` endpoint** (added 2026-07-08,
dual-endpoint `e` run). No health route exists anywhere in the API today. One
cheap, unauthenticated JSON endpoint (no secrets in the body; optionally include
db/pipeline readiness booleans) serves three existing consumers at once: the
Flutter reachability probe (currently abuses `POST /auth/login` with an empty
body — see root plan task 5.3), the docker-compose healthcheck task 4.2 needs a
target for, and the tunnel-reachability verification in task 4.3. Must stay
unauthenticated (it's the thing that decides whether auth'd calls are even
attempted) and must return non-2xx or distinct flags when Postgres/pipeline
aren't ready, so 4.2's gate is meaningful.
- **1.4** [DONE 2026-07-08] Enforced real 401 auth on `/api/v1/documents/*` instead — this is the actual production API surface, already fully supported by the Flutter client (`lib/features/auth/auth_provider.dart` does a real login and stores the JWT; `lib/core/network/api_client.dart`'s interceptor already attaches `Authorization: Bearer <token>` to every request — root `CLAUDE.md`'s "demo-stub, no real server-side auth" claim for Flutter was itself stale). Before this, `v1/documents/route.ts` (list) and `v1/documents/[id]/route.ts` (PUT) didn't check auth at all, and `v1/documents/upload/route.ts` only optionally read the token (never rejected a missing one). Added `getAccountFromAuthHeader()` (`utils/auth.ts`, pre-existing helper) + a `401` guard to all three; `OPTIONS` (CORS preflight) on all three left untouched. Did **not** add per-account data filtering to the document list (still returns all non-sample parsed documents regardless of uploader — that's a feature change, not an auth fix; flagged as a possible future task). Did **not** add Flutter-side 401→auto-logout handling (`api_client.dart` has no interceptor for it; a 401 surfaces as a normal `ApiException` today) — a Flutter-side follow-up, not a backend task.
- Verified end-to-end via `curl`: all three endpoints return 401 with no token; after `POST /api/v1/auth/login` with `admin`/`password` to get a real token, the same three endpoints succeed with `Authorization: Bearer <token>` (`GET` 200 with real data, `PUT` reaches its normal 404-for-bad-id logic, `POST upload` reaches its normal downstream logic) — confirming the auth check itself works without touching any other behavior. `OPTIONS` on all three still returns 204 unauthenticated.
@@ -90,6 +106,10 @@ truth for it going forward):
(`src/app/produk-pfm/page.tsx` doesn't exist) — dead route, same class of issue as
scan-pfm/m-scan-pfm were. Not turned into a task below since it wasn't part of the
original decision record; flagging for a future `e`/`enhance` pass to pick up.
- **Correction (2026-07-08 audit)**: only *half* dead. `api/produk-pfm/route.ts`
is live — `scan-pfm/page.tsx:159` fetches it for the sample-product gallery.
Only the nginx `location /produk-pfm` *page* proxy block points at a
nonexistent page. Don't remove the API route; see task 4.4.
- **2.1** [DONE 2026-07-08] Built the initial model artifacts. Bare-metal
(`./scripts/install-pipeline.sh`) doesn't work in this Windows dev environment —
@@ -218,6 +238,13 @@ truth for it going forward):
PaddleOCR-VL prompt/model tuning) or expanding the test set to catch different
failure modes, not more regex tweaks.
- **2.4** [TODO] Human review of `do-008.jpg`'s ground truth in
`sources/manual_labels.json` — flagged in task 2.3 as a suspected labeling error
(its `noPO`/`noSO`/`noDO` OCR cleanly but are entirely different digit sequences
from the labels, not plausible misreads) but never turned into a task. Verify
against the source photo with the client/labeler; if the labels are wrong, fix
them and re-run the harness (aggregate accuracy should tick up).
*Not in scope for either 2.1/2.2 (per next-implementation.md's own note, still true
2026-07-08): batch/lot number extraction doesn't exist in `classify_ocr_server.py`
in either project (only SKU, product name, expiry date are extracted) — if
@@ -237,6 +264,221 @@ expiry-date extraction.*
- **4.1** [TODO] Decide and document (in root `README.md`/`CLAUDE.md`) a deliberate policy for when the PoC/demo runs `docker-compose.demo.yml` (production build) vs. the dev-mode default — `docker-compose.yml` alone runs `npm run dev`, a known throughput ceiling already flagged in the reliability audit.
- **4.2** [TODO] Add a startup healthcheck/readiness gate for `pipeline-api`/vLLM in `docker-compose.yml` so the Next.js gateway doesn't accept uploads before the GPU pipeline is actually ready to serve them.
- **4.3** [TODO] Extend `start-dev-tunnel.ps1` to verify the ngrok tunnel/LAN IP is actually reachable (not just started) before reporting success, since a stale tunnel is the documented first failure point for login/upload from the Flutter app.
- **4.4** [TODO] Prune/annotate the dead `nginx.conf` location blocks: `/do-pfm` and
`/m-do-pfm` (dead since v2 consolidated the DO-PFM UI into the root page — already
documented as dead in root `CLAUDE.md`), the `/produk-pfm` page proxy (no
`src/app/produk-pfm/page.tsx` exists — but keep `api/produk-pfm/route.ts`, it's
live, used by `scan-pfm/page.tsx:159`; see the corrected §2 note), and
`/m-scan-pfm` (page intentionally unbuilt per cancelled task 2.2 — removing the
proxy block doesn't resurrect that task, it just stops nginx advertising a 404).
- **4.5** [TODO] **Lock down the publicly tunneled surface** (added 2026-07-08,
dual-endpoint `e` run). `start-dev-tunnel.ps1` tunnels **all of
`localhost:8000`** (the whole nginx gateway) to a stable, reserved ngrok domain —
which makes the deliberately unauthenticated classic routes (`/api/upload`,
`/api/parse`, `/api/history`, ...) and the dev pages (root DO-PFM, `/scan-pfm`,
`/manual-label`) internet-reachable. Task 1.3's "classic routes never get auth"
decision was made in an on-prem/LAN context and stands — so restrict at the
edge instead of adding auth: either (a) tunnel a separate nginx server
block/port that proxies **only `/api/v1/*`** (the authenticated production
surface the Flutter app actually uses — see `app_config.dart`, both endpoints
end in `/api/v1`), or (b) an ngrok traffic policy (IP allowlist / basic auth)
covering everything except `/api/v1/*`. Decide (a) vs (b) with the user at
pickup; (a) is self-contained in `nginx.conf` + the script and doesn't depend
on ngrok plan features. Note `GET /api/v1/health` (task 1.6) must remain
reachable through whichever restriction ships — it's the fallback probe target.
## 5. Docs & Workflow Integrity
`CLAUDE.md`, `SKILLS.md`, `AGENTS.md`, this file — added 2026-07-08 after an audit
cross-checking the kit docs' claims against live code. All *code-level* claims in
this file's §1-4 re-verified accurate (auth guards, bcrypt, `withTransaction`,
parser-fallback fix, dedup, timeouts, absent index/healthcheck all match the code);
the drift found is in the guidance docs the SKILLS.md roles rely on.
- **5.1** [TODO] Fix three stale doc claims that could steer a future agent wrong:
(a) `SKILLS.md` §4 (QA role) says the accuracy baseline is "~89.4% overall,
target 95% — see backend `CLAUDE.md`" — the verified baseline is **95.10%, target
already met** (task 2.3), and backend `CLAUDE.md` states no number at all, so the
pointer dangles; point it at `sources/accuracy_history.jsonl`'s latest entry as
the living source instead of hardcoding a snapshot. Risk if unfixed: QA passes a
real regression (e.g. 91%) as "above baseline". (b) backend `CLAUDE.md`'s
Commands section still says `install-pipeline.sh` is "currently missing
ultralytics/torch" — task 2.1 added `ultralytics>=8.0` (torch arrives as its
dependency). (c) root `CLAUDE.md` still says Flutter auth "talks to a demo-stub
backend (... no real server-side token verification)" — stale since tasks 1.4 +
3.2 (real JWT verification with 401 enforcement, bcrypt-hashed passwords); the
single-seeded-account part is still true. Task 1.4's own writeup flagged this but
the doc was never corrected — see 5.3 for the process fix. (d) (found 2026-07-08,
dual-endpoint `e` run) root `CLAUDE.md` describes the API base URL resolution
**backwards**: it says the app "tries a fixed ngrok domain first, falling back to
a hardcoded LAN IP" — the code (`lib/config/app_config.dart:32-49`) probes the
LAN URL first and falls back to ngrok, exactly the local-first behavior wanted;
fix the description when touching that doc.
- **5.2** [TODO] This file has crossed the §B3 256-line threshold. Archive the
verbose `[DONE]`/`[CANCELLED]` task bodies (they're already duplicated in
`docs/feature-list.md`) down to one-line stubs pointing there, keeping full text
only for `[TODO]` tasks and still-load-bearing decision records.
- **5.3** [TODO] Amend `AGENTS.md` Part B §B2's completion checklist with a
doc-sync step: when an `n` task invalidates a claim in `CLAUDE.md`/`SKILLS.md`/
`README.md`, correct that doc *in the same task* rather than noting the staleness
only in the task writeup — that pattern (see 5.1a/5.1c) is exactly how this
section's drift accumulated.
## 6. Product Scan — Ground Truth Annotation & Accuracy
`scan-pfm/page.tsx`, `api/manual-label-scan/route.ts`, `sources/product_manual_labels.json`
Added 2026-07-08 via a user-directed `e` run: build an **editable ground-truth
surface for product scans**, mirroring what the DO flow already has in
`manual-label/page.tsx` + `api/manual-label/route.ts` + `sources/manual_labels.json`.
**Current-state audit (verified against code, not assumed):**
- What exists: `api/manual-label-scan/route.ts` (GET by exact filename / POST
upsert, schema `{filename, no_sku, nama_item, expiry_date, top1_confidence,
notes, saved_at}`, own `normalizeDateString`), a "Save Ground Truth" button in
`scan-pfm/page.tsx` (`handleSaveGroundTruth`, line 333), and the
`./backend/sources:/sources` mount that makes saves host-visible (fixed in 2.1).
- **Gap (a) — prediction saved as truth**: in the scan-pfm save, only `no_sku` is
correctable (the top-5 "Use" override → `editedSku`); `nama_item` is hardcoded to
the classifier's `top1_name` and `expiry_date` to the OCR output, `notes` always
`""`. Ground truth that *is* the model's own prediction makes any future accuracy
measurement against it trivially inflated — this is the core defect the user's
request fixes.
- **Gap (b) — phantom image keys**: uploaded (non-gallery) photos are keyed
`uploaded-${Date.now()}.jpg` and the image bytes are never persisted anywhere
(data-URL in the browser only) — those label entries reference files that don't
exist on disk, unusable for re-evaluation.
- **Gap (c) — no browse/edit**: GET requires an exact filename; there's no list
endpoint, no page to review/correct/delete existing entries (the DO side has a
full annotation page for this).
- **Gap (d) — no consumer**: nothing plays the role `accuracy-check.mts` plays for
the DO parser; the labels currently gate nothing.
- **6.1** [TODO] **Standalone annotation page `manual-label-scan/page.tsx`** —
follow the documented standalone-route pattern (`CLAUDE.md`: own header/theme, no
shared chrome), named to match its existing API route. Detail:
- **Image browser**: enumerate the reference dataset via the existing
`api/produk-pfm` gallery route (16 SKU folders under
`public/produk-pfm/foto-kemasan-v2/`), plus persisted uploaded scans once 6.2
lands. Prev/next navigation and a per-image labeled/unlabeled badge (from 6.2's
list endpoint) so the annotator can see coverage at a glance — same UX shape as
`manual-label/page.tsx`'s file list.
- **Editable fields (all of them, unlike the scan-pfm quick-save)**: `no_sku`
(autocomplete against the existing `/api/skus` master route — folder names in
foto-kemasan-v2 can prefill it since the dataset is organized by SKU),
`nama_item` (autofilled from SKU master on SKU select, still overridable),
`expiry_date` (reuse `normalizeDateString` — lift it out of the route file into
a shared util rather than copy-pasting), `notes`.
- **AI-vs-manual reference**: a "Scan with AI" action calls `/api/scan-pfm` for
the current image and shows top-1 class + confidence and OCR expiry as captions
next to each field (the `AiNote` pattern from `manual-label/page.tsx`) —
fill-blanks-only, never overwriting a manually corrected value (same merge rule
`api/manual-label`'s GET already implements for the DO side).
- Load/save through `api/manual-label-scan` as extended by 6.2. Keep the page
under the §B3 256-line rule by splitting components from the start — do *not*
clone `manual-label/page.tsx`'s structure wholesale (it's 612 lines, listed §B3
debt). Implementation note: `pfm-web-app/AGENTS.md` warns this Next.js version
differs from training data — read `node_modules/next/dist/docs/` before coding.
- **6.2** [TODO] **API + storage groundwork** (do this first; 6.1 builds on it):
- Extend `api/manual-label-scan`: GET without `filename` returns **all entries**
(list mode), and add DELETE by filename — needed for browse/cleanup. Additive,
schema-compatible changes only; existing entries keep working.
- **Persist uploaded scan photos**: when saving ground truth for a non-gallery
image, write the image bytes to `sources/product-test-images/` (host-visible
via the existing `/sources` mount, sibling of the DO flow's
`sources/test-images/`) under a stable content-hash filename, and key the label
on that — kills the `uploaded-${Date.now()}.jpg` phantom keys (gap b). Ask the
user what to do with already-saved phantom entries (delete vs. keep flagged).
- **Make the scan-pfm quick-save honest**: turn the Save Ground Truth panel's
`nama_item`/`expiry_date`/`notes` into editable inputs prefilled with the AI
values (gap a), so the fast path saves *reviewed* truth; `top1_confidence`
stays as AI metadata. Optionally record `source: "scan-pfm" | "annotation-page"`
per entry for provenance.
- **6.3** [TODO] **Product-scan accuracy harness** — the payoff that makes the
labels load-bearing, mirroring `accuracy-check.mts`: for every
`product_manual_labels.json` entry whose image exists on disk, run the
classify+OCR pipeline and diff `no_sku` (top-1 exact match + top-5 hit rate),
`nama_item`, and `expiry_date` against the label; print per-field results and
append run history to `sources/product_accuracy_history.jsonl`. Becomes the §B2
QA gate for any classifier retrain or `classify_ocr_server.py` change, same role
the DO harness plays for `parser.ts`. **Caveat to bake into the report output**:
foto-kemasan-v2 images are also the classifier's training data (83.3% top-1 val
on the thin dataset), so scores on them measure memorization — the meaningful
eval split is the persisted *uploaded* scans from 6.2; report the two populations
separately.
*Suggested order: 6.2 → 6.1 → 6.3 (storage/API first, page on top, harness once
labels exist in volume).*
## 7. Auth — Store Accounts & Profile-Sourced Metadata
`pfm-web-app/src/db/init.ts`, `api/v1/auth/*`, `api/parse/route.ts`, `sources/toko_aktif.json`
Added 2026-07-08 via a user-directed `e` run: one login account per store
(username = `kode_toko`, default password `"123"`), each carrying its store
profile (nama toko, kode toko, alamat), with the profile's store name + address
used as the parse response's values instead of OCR.
**Current-state audit (verified against code and the live DB, not assumed):**
- **The "don't OCR store/alamat" half is already live.** `v1/documents/upload`
passes `account?.kodeToko` into `/api/parse` (upload `route.ts:125`), and parse
stamps `metadata.orderUntuk`/`alamat` straight from `store_master` when a
`kodeToko` is present — OCR-based store matching (`resolveStoreFromText`) is
only the no-account fallback (`parse/route.ts:392-410`). The Flutter app reads
these same metadata fields from `GET /documents`, so **no response-shape or
Flutter change is needed** for that half.
- **What's missing is the accounts**: the live DB has **779 `store_master` rows
but exactly 1 account** (`admin`), so every real upload today authenticates as
`admin`/`WH_JOFFICE` and the per-store path never fires for actual stores.
- **Bootstrap gap**: no code anywhere populates `store_master` — the 779 rows
exist only in the live DB volume (imported out-of-band). `sources/toko_aktif.json`
(`{namaToko, kodeToko, alamat}`, same 3 fields) is the obvious source. On a
truly fresh DB, `init.ts` would even fail its own `admin` seed — the
`accounts.kode_toko → store_master(kode_toko)` FK can't resolve `WH_JOFFICE`
when `store_master` is empty.
- **7.1** [TODO] **Seed one account per store** in `db/init.ts`: username =
`kode_toko`, password = bcrypt of `"123"`. Implementation constraints:
- **One precomputed hash, one statement**: hash `"123"` once and use a single
`INSERT INTO accounts (username, password, kode_toko) SELECT kode_toko, $1,
kode_toko FROM store_master ON CONFLICT (username) DO NOTHING` — per-row
`bcrypt.hashSync` at ~50-100ms × 779 would add a minute+ to *every* container
startup. Sharing one hash across accounts with the same default password is
an accepted trade-off here.
- Idempotent by construction: re-runs are no-ops, an account whose password was
later changed is never clobbered, and stores added to `store_master` later
get their account on the next startup automatically.
- **Schema decision — don't duplicate `nama_toko`/`alamat` into `accounts`**:
the profile is `accounts.kode_toko JOIN store_master` (single source of
truth, no drift when a store's address changes). The user's requested fields
all exist across that join already.
- Suggested additions (user asked "tambahkan jika ada yang kurang"): a `role`
column (`'admin' | 'store'`) so the `admin` account is distinguishable from
store accounts, and `is_active BOOLEAN` for offboarding a store without
deleting its history. Confirm at pickup (§B2a).
- **Documented trade-off** for `docs/feature-list.md`: a shared default
password `"123"` on 779 internet-reachable-via-ngrok accounts is a client
decision for field simplicity — record it as accepted risk and
cross-reference task 4.5 (public tunnel lockdown), which becomes more
important once these accounts exist.
- **7.2** [TODO] **Return the store profile at login**: extend
`POST /api/v1/auth/login`'s response `data` from `{token}` to
`{token, profile: {username, kodeToko, namaToko, alamat}}` (JOIN
`store_master`; additive, non-breaking). Optionally add `GET /api/v1/auth/me`
(token → same profile) so the app can re-fetch without re-login. Flutter-side
display of the profile (e.g. store name in the drawer) is root-kit scope — note
it there if wanted, the backend contract just has to expose the data.
- **7.3** [TODO] **Reproducible `store_master` bootstrap**: an idempotent import
in `db/init.ts` (or a script it calls) from `sources/toko_aktif.json` →
`store_master` (`ON CONFLICT (kode_toko) DO NOTHING`; decide at pickup whether
to also `UPDATE` changed names/addresses), ordered **before** the accounts
seeding (7.1) and the `admin` seed so a fresh DB initializes cleanly end-to-end.
Verification (per the test-every-task rule): wipe to a fresh DB volume in a
throwaway compose project, boot, confirm 779 stores + 780 accounts, then log in
as a real `kode_toko`/`"123"`, upload a test image, and confirm the parsed
response's `orderUntuk`/`alamat` match that store's master row (not the OCR
text) — this also functions as the end-to-end proof for 7.1/7.2 and the
already-shipped parse path.
*Suggested order: 7.3 → 7.1 → 7.2 (bootstrap first — account seeding FK-depends
on it; login profile last, it's additive).*
---
+47 -1
View File
@@ -74,8 +74,54 @@ When complete, the status flips to `[DONE]` and the feature is logged in
- **4.2** [TODO] Block save when a line item's SKU isn't in the master registry, instead of only relabeling it for display. The SKU listener in `_addItem` (`editor_screen.dart:108-115`) sets the item name to "SKU Tidak Terdaftar" for an unrecognized SKU but doesn't stop form submission, so a document with an unregistered/mistyped SKU can still be saved and its receipt printed.
- **4.3** [TODO] Include the captured GPS coordinates on the printed PDF receipt. `PdfService.generateAndPrintReceipt` (`pdf_service.dart:36-52`) prints header/shipment/item fields but never includes `document.latitude`/`longitude`, even though the editor captures and displays them (`_latitudeCtrl`/`_longitudeCtrl`) — the geotag exists in the data model but isn't part of the audit-trail document a store keeps.
### 5. Flutter — Connectivity & Endpoint Resolution
`lib/config/app_config.dart`, `lib/core/network/api_client.dart`
Added 2026-07-08 via a user-directed `e` run ("dual endpoint: local first, public
ngrok fallback"). **Current-state audit**: the requested dual-endpoint fallback
*already exists at startup* — `AppConfig.initializeApiBaseUrl()`
(`app_config.dart:32-49`) probes the LAN URL first, falls back to the reserved
ngrok domain, and validates each probe is a *real* backend (checks the
`ngrok-error-code` header, JSON content-type, and 502/503/504) with the
`ngrok-skip-browser-warning` header set. (Root `CLAUDE.md` describes this order
backwards — tracked as a doc fix in backend task 5.1d.) These tasks close what's
actually missing:
- **5.1** [TODO] **Mid-session endpoint failover.** Resolution runs exactly once at
startup, and `ApiClient` freezes `baseUrl` at construction of a singleton
(`api_client.dart:13`, `apiClientProvider`) — a phone that resolves LAN on Wi-Fi
and then leaves the building fails every subsequent call with no path back to
the ngrok endpoint (and vice versa) until an app restart. Add a Dio interceptor
that, on *connectivity-class* failures only (`connectionTimeout`/
`connectionError` — not HTTP-level errors), re-runs endpoint resolution and
retries the request once against the newly resolved endpoint. Implementation
notes: the frozen-at-construction `baseUrl` must become dynamic (set
`_dio.options.baseUrl` on re-resolution, or read `AppConfig.apiBaseUrl` in the
existing `onRequest` interceptor); retry-once is safe for the upload path
because the backend dedups by `file_hash` (backend task 1.1), but audit other
POST/PUT call sites before blanket-retrying. Coordinate with 3.3 (explicit Dio
timeouts) — a hung connection must fail fast enough for failover to matter.
- **5.2** [TODO] **Endpoint status visibility + manual re-probe.** Show which
endpoint the app is on (LAN / Public / unreachable) as a small persistent
indicator (camera screen or drawer) with a tap-to-re-probe action, so a driver
or tester can see and fix "wrong/stale endpoint" in the field without reading
logs — a stale tunnel is the documented first failure point for login/upload
(root `CLAUDE.md`). Re-probe reuses 5.1's resolution path.
- **5.3** [TODO] **Make both endpoint URLs configurable without a code edit.**
`_lanBaseUrl` and `_ngrokBaseUrl` are compile-time consts — the LAN one is
regex-patched by `start-dev-tunnel.ps1` (which breaks silently if the const is
renamed/moved; the script warns but the app still ships the stale IP), the ngrok
one requires a manual source edit if the reserved domain ever changes. Add a
runtime override (e.g. long-press-hidden settings sheet writing to
`SharedPreferences`, seeded from the compiled defaults) so a field device can be
repointed without rebuilding the APK. Keep the script's patch working (or teach
it to fail loudly — backend task 4.3 covers verifying the tunnel end). Once
backend task 1.6 ships `GET /api/v1/health`, switch `_isBackendReachable`'s
probe to it — probing `POST /auth/login` with an empty body works but couples
reachability to the login route's error shape.
---
*Sections 1-4 (Flutter) are the only sections this file tracks. Backend
*Sections 1-5 (Flutter) are the only sections this file tracks. Backend
enhancements (formerly sections 5-8 here, removed 2026-07-08) now live
exclusively in [backend/plans/next-enhancements.md](../backend/plans/next-enhancements.md).*