Files
pfm-ocr/backend/plans/next-enhancements.md
T
Rafhan Mazaya FathurrahmanandClaude Sonnet 5 e60ab63154 Adopt agents-settings kit, ship Product/SKU scan models, harden auth, verify OCR accuracy
Backend (app-pfm-ocr-v2/backend):
- Product/SKU scan feature complete: trained DINOv2 index (118 reference
  photos, 16 SKU classes) and YOLO classifier (83.3% top-1 val accuracy),
  fixed scripts/install-pipeline.sh (was missing ultralytics/torch), fully
  browser-verified end-to-end on /scan-pfm. Mobile m-scan-pfm page cancelled
  (Flutter app handles mobile; web UI is desktop-only for pipeline testing).
- Fixed a real data-loss bug: Save Ground Truth (scan-pfm and the DO-flow's
  manual-label) was silently writing into the pfm-web-app container's
  ephemeral filesystem instead of the host, because /sources wasn't
  bind-mounted in docker-compose.yml. Added the mount, recovered an
  orphaned entry.
- accounts.password is now bcrypt-hashed (bcryptjs, idempotent migration
  in db/init.ts) instead of plaintext; login route compares hashes.
- /api/v1/documents/* (list, PUT, upload) now enforces real 401 auth,
  matching what the Flutter client already sends. The "classic" routes
  deliberately stay open — they're dev-only web UI with no login flow and
  won't exist in production.
- OCR accuracy investigated end-to-end: real baseline is 95.10% overall
  (target met; accuracy_report.md was stale at 75.04%, now flagged). Fixed
  one genuine parser.ts bug (SO/DO field duplication in the global fallback
  regex); remaining gaps are OCR/layout-model limitations, not parser bugs.
- Adopted a standalone copy of the fhanyuh/agents-settings e/n workflow
  scoped to backend/ (AGENTS.md Part A/B split, SKILLS.md, plans/, docs/),
  independent of the root copy which now covers Flutter only.
- next-implementation.md deleted; content folded into
  backend/plans/next-enhancements.md for traceability.

Root:
- Adopted fhanyuh/agents-settings kit (AGENTS.md, SKILLS.md, plans/,
  docs/feature-list.md), scoped to the Flutter app only.
- Pending documents queue now persists to Hive (lib/core/storage) instead
  of memory-only, surviving an app kill mid-upload.

Removed backend_backup/ (stale Express/Prisma prototype, superseded by
pfm-web-app) and the completed plans/next-enhancement-plan.md checklist.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-08 11:56:32 +07:00

22 KiB

Next Enhancements (backend)

Working backlog driven by the e/enhance and n/next triggers defined in AGENTS.md Part B. Split out 2026-07-08 from root plans/next-enhancements.md sections 5-8, renumbered 1-4 — this file now owns all backend enhancement tracking going forward; the root file only tracks Flutter sections from this point on. Statuses below were re-verified against the live code at split time (not copied blind):

  • 1.1 file_hash dedup — confirmed present, v1/documents/upload/route.ts:59,98.
  • 1.2 AbortSignal.timeout on both hops — confirmed present in parse/route.ts and v1/documents/upload/route.ts.
  • 1.3 classic routes skipping auth — confirmed still true at split time; later cancelled as out-of-scope (dev-only web UI, no production benefit) — see §1 below, 1.4 shipped instead.
  • 3.1 withTransaction — confirmed present and used in db/index.ts, api/parse/route.ts, api/v1/documents/[id]/route.ts.
  • 2.1 pfm-web-app/public/produk-pfm/models/ — corrected 2026-07-08 (was checked against the wrong path, backend/models, in an earlier pass this same day): the directory does exist. Later the same day, task 2.1 itself was completed — see §2 below — so the artifacts now exist too.
  • 2.2 m-scan-pfm/page.tsx — confirmed still missing; desktop scan-pfm/page.tsx confirmed present, and now confirmed full feature parity (confidence bar, top-5 candidate list, expiry-date crop preview all present, 1169 lines) — see §2 below, migrated from next-implementation.md (deleted 2026-07-08, content now lives here).
  • 3.2 password hashing — confirmed plaintext at split time (2026-07-08 morning); fixed later the same day, see §3 below.
  • 3.3 unique index on documents.file_hash — confirmed still absent (only filename has a UNIQUE constraint in db/init.ts).
  • 4.2 healthcheck for pipeline-api/vLLM — confirmed no healthcheck block in root docker-compose.yml.

Format

Tasks are grouped under a numbered section per backend module. Each section gets exactly 3 tasks:

## 1. <Section / Module Name>

- **1.1** [TODO] <clear, specific description of the functional change>
- **1.2** [TODO] <...>
- **1.3** [TODO] <...>

When a task is picked up via n/next, its clarified acceptance criteria (from AGENTS.md Part B §B2a) are appended directly under it. When complete, the status flips to [DONE] and the feature is logged in docs/feature-list.md (this dir).


1. Backend — Next.js API Gateway

pfm-web-app/src/app/api/

  • 1.1 [DONE] Make documents/upload/route.ts return 201 immediately... Investigated 2026-07-08: the file_hash dedup check is already shipped pre-existing in v1/documents/upload/route.ts (lines 58-89) — a duplicate upload returns the existing document instead of re-inserting/re-parsing. The "return 201 immediately, don't await OCR" half turned out to be a deliberate, already-documented tradeoff (see the route's own comment): the pipeline call stays synchronous but is now bounded by AbortSignal.timeout(210_000). Not changed further this round since it's an intentional decision, not an oversight.
  • 1.2 [DONE] Investigated 2026-07-08: already shipped pre-existing. Both hops of the upload → /api/parse → pipeline-API chain already have AbortSignal.timeout (210s and PIPELINE_TIMEOUT_MS=90s respectively), with comments cross-referencing each other.
  • 1.3 [CANCELLED 2026-07-08] Wire the existing JWT auth onto the classic API routes (/api/upload, /api/scan-pfm, /api/parse, /api/history, etc.) — investigated after 3.2 unblocked it, found the premise was stale: those "classic" routes are used exclusively by the dev-only web UI (root DO-PFM page, scan-pfm, manual-label), which has no login screen and never sends a token — the user confirmed this web UI won't exist in production, so enforcing auth there would break it today for zero production benefit, and optional identification there is a no-op (nobody ever sends a token). Same reasoning as 2.2's cancellation. See 1.4 below for what was actually shipped instead.
  • 1.4 [DONE 2026-07-08] Enforced real 401 auth on /api/v1/documents/* instead — this is the actual production API surface, already fully supported by the Flutter client (lib/features/auth/auth_provider.dart does a real login and stores the JWT; lib/core/network/api_client.dart's interceptor already attaches Authorization: Bearer <token> to every request — root CLAUDE.md's "demo-stub, no real server-side auth" claim for Flutter was itself stale). Before this, v1/documents/route.ts (list) and v1/documents/[id]/route.ts (PUT) didn't check auth at all, and v1/documents/upload/route.ts only optionally read the token (never rejected a missing one). Added getAccountFromAuthHeader() (utils/auth.ts, pre-existing helper) + a 401 guard to all three; OPTIONS (CORS preflight) on all three left untouched. Did not add per-account data filtering to the document list (still returns all non-sample parsed documents regardless of uploader — that's a feature change, not an auth fix; flagged as a possible future task). Did not add Flutter-side 401→auto-logout handling (api_client.dart has no interceptor for it; a 401 surfaces as a normal ApiException today) — a Flutter-side follow-up, not a backend task.
    • Verified end-to-end via curl: all three endpoints return 401 with no token; after POST /api/v1/auth/login with admin/password to get a real token, the same three endpoints succeed with Authorization: Bearer <token> (GET 200 with real data, PUT reaches its normal 404-for-bad-id logic, POST upload reaches its normal downstream logic) — confirming the auth check itself works without touching any other behavior. OPTIONS on all three still returns 204 unauthenticated.

2. Backend — OCR Pipeline & Accuracy

config/, pfm-web-app/src/utils/parser.ts, accuracy regression harness (see CLAUDE.md)

Product/SKU scan sub-feature — decisions & state (migrated 2026-07-08 from next-implementation.md, which is now deleted; this section is the sole source of truth for it going forward):

  • Decisions made: new pages are standalone routes (scan-pfm — done), following the same self-contained pattern as manual-label/page.tsx — own header/theme, no shared chrome with the root DO-PFM page. Build to full feature parity with the old ai-ocr-pfm-2026 pages (confidence bars, top-5 SKU candidates, visual OCR overlay, expiry-date crop preview), not a lean MVP.

  • Scope correction 2026-07-08 (later): m-scan-pfm/page.tsx (mobile web page) is not needed — user clarified the web scan-pfm page is desktop-only, used for testing the pipeline, not a production mobile surface. Real mobile scanning is already handled by the Flutter app instead. Focus for this sub-feature going forward is backend services (model artifacts, pipeline correctness), not any additional frontend. See 2.2 below — cancelled, not reassigned.

  • Already shipped, re-verified 2026-07-08 (all confirmed present, not assumed): config/classify_ocr_server.py (DINOv2 + YOLO classify, SKU/expiry/name OCR), pfm-web-app/public/produk-pfm/index_dinov2.py + train_classifier.py, api/scan-pfm/route.ts (Levenshtein match against sku_master), api/produk-pfm/route.ts, DB schema (sku_master/store_master/arena_runs in db/init.ts), scripts/serve-pipeline.sh wiring, Dockerfile's ultralytics install in the pipeline-api stage, nginx.conf routes for /scan-pfm, /m-scan-pfm, /produk-pfm.

  • Stale claims corrected 2026-07-08 (next-implementation.md said these didn't exist; they now do): pfm-web-app/public/produk-pfm/foto-kemasan-v2/ dataset — 16 SKU subfolders now present (was "❌ does not exist"); scan-pfm/page.tsx — now exists at full feature parity, 1169 lines (was "❌ does not exist").

  • Newly observed, not in original doc's scope: /produk-pfm has an nginx proxy block and an API route (api/produk-pfm/route.ts) but no matching frontend page (src/app/produk-pfm/page.tsx doesn't exist) — dead route, same class of issue as scan-pfm/m-scan-pfm were. Not turned into a task below since it wasn't part of the original decision record; flagging for a future e/enhance pass to pick up.

  • 2.1 [DONE 2026-07-08] Built the initial model artifacts. Bare-metal (./scripts/install-pipeline.sh) doesn't work in this Windows dev environment — paddlepaddle-gpu's wheel index is Linux-only — so this ran entirely through Docker instead:

    1. Added ultralytics>=8.0 to scripts/install-pipeline.sh (was missing there even though the Dockerfile pipeline-api stage already installed it — fixed for bare-metal parity, though bare-metal training itself still isn't viable on Windows).
    2. docker compose build pipeline-api from repo root — cached layers reused, only the final COPY . /app re-ran, picking up the (until-then-unbaked) 16-folder dataset.
    3. Ran a one-off docker run --gpus all from the fresh image with models/ mounted writable (not the live service's :ro mount) — the live pipeline-api container's copy was stale (image predated the dataset) and read-only either way, so training couldn't happen through it directly. index_dinov2.py → indexed 118/118 images across all 16 classes → models/dinov2_index.pkl (196KB). train_classifier.py train --imgsz 224 → 100 epochs (~1.5 min), 83.3% top-1 / 90% top-5 validation accuracy (thin dataset, 2-16 images/class — accuracy will improve as more reference photos are added) → models/produk-pfm-classifier-26n-100e-2026-07-08.pt (3.2MB) + .onnx export (5.9MB).
    4. docker compose restart pipeline-api (bind-mounted models/ is live, no rebuild needed for the running service) — confirmed via docker logs: "DINOv2 model loaded successfully" and "YOLO model loaded successfully... Using classifier weights: produk-pfm-classifier-26n-100e-2026-07-08.pt".
    • Note on Git Bash + Docker on Windows: docker run ... bash -c "/app/..." fails with exit 127 ("No such file or directory") because MSYS path-conversion mangles the leading /app/... argument into a Windows path — fix is MSYS_NO_PATHCONV=1 before the docker run/docker exec invocation.

    Full browser verification, 2026-07-08 (later same day) — ran /scan-pfm live (http://localhost:3000/scan-pfm, plus confirmed http://localhost:8000/scan-pfm through nginx) via Claude-in-Chrome, selecting the 16-image "FIESTA STIKIE" sample and clicking through the full flow, not just checking container logs:

    • AI Classification: 100.00% confidence, correct class, "Model Active" — confirms the freshly-trained YOLO weights are actually being used, not a fallback.
    • "View Top 5 Predicted Probabilities" — expands correctly with 5 ranked candidates + "Use" override buttons.
    • AI OCR Text Extraction — expiry date 26/02/2027 extracted correctly with a label-crop preview image.
    • "Possible Match from SKU Master" — Levenshtein-ranked list, correct best match at 71% similarity.
    • Visual Grid tab, Spotting Grid tab, Raw Response (JSON) tab — all render correctly with real data.
    • Bug found and fixed: "Save Ground Truth" returned HTTP 200 and looked successful in the UI, but api/manual-label-scan/route.ts writes to path.join(process.cwd(), "..", "sources", "product_manual_labels.json") — inside the pfm-web-app container this resolves to /sources/..., which was not a bind-mounted path in root docker-compose.yml (only /app, /uploads, and the Docker socket were mounted for that service). Every save was landing in the container's ephemeral writable layer, invisible to the host, to git, and to any other tooling — confirmed by finding an orphaned entry from 2026-07-07 (an earlier, unrelated test) sitting in the container with no trace on disk. Fix: added ./backend/sources:/sources to the pfm-web-app service's volumes in root docker-compose.yml, recovered the orphaned entry via docker cp before recreating the container, then docker compose up -d pfm-web-app (recreated both pfm-web-app and, as its dependency, pipeline-api — re-verified both models reloaded correctly after). Re-ran the full scan + save and confirmed the entry now lands directly in backend/sources/product_manual_labels.json on the host.
    • Confirmed (2026-07-08, follow-up check) the same bug pattern in api/manual-label/route.ts (DO-flow ground-truth save, sources/manual_labels.json, identical process.cwd()-relative path) is now also fixed by the same volume mount — verified md5sum matches exactly between the container's /sources/manual_labels.json and the host's backend/sources/manual_labels.json, with a fresh host mtime. No separate fix needed; the one docker-compose.yml change covers both routes.
    • Acceptance met: models/dinov2_index.pkl exists, "Indexed 118/118 images" log confirmed; both models load cleanly on pipeline-api restart per container logs. End-to-end /scan-pfm UI test (upload a real photo, confirm live top-5 result) — done, see the full browser verification note above (2026-07-08, later still).
  • 2.2 [CANCELLED 2026-07-08] Port m-scan-pfm/page.tsx (mobile) from ai-ocr-pfm-2026 into v2 — user decided this isn't needed: the web scan-pfm page is desktop-only (used to test the pipeline), and real mobile product scanning is handled by the Flutter app, not a web mobile page. The /m-scan-pfm nginx block stays dead intentionally — do not resurrect this task without an explicit ask. docs/feature-list.md will not get an entry for it.

  • 2.3 [DONE 2026-07-08] Ran the accuracy regression harness. Key finding: sources/accuracy_report.md was badly stale (said 75.04% overall) — the real current baseline (confirmed via a fresh harness run at commit 3df9f6e, matching sources/accuracy_history.jsonl's latest entry exactly) is 95.10% overall, already at/above the 95% target. Added a staleness banner to accuracy_report.md pointing here instead of leaving it to mislead future work.

    • Investigated the worst field, plat (67.6%, unchanged between L1/L3 — no sanitization ever touches it) by pulling raw OCR text for every mismatching image from documents.layout_parsing_result in Postgres. Of 13 mismatches: 11 are the plate region getting classified by the layout model as an image_box/ seal_box (never becomes text at all — e.g. Truck No. cell literally contains an <img> tag, not digits) or the suffix OCR'd as digits instead of letters ("B 9536 000" vs ground truth "B 9536 UCO" — 0 for O/C). Not fixable in parser.ts — there's no text to parse in most cases; the remaining 1-2 are single-character OCR misreads ("VOC" vs "VCC") too fragile to correct generically without risking false positives elsewhere.
    • Cross-checked all other mismatching fields (noSO/noPO/noDO/tanggal, 12 mismatches total via --dump-json) the same way. Same story: mostly garbled/missing OCR text (image regions, blank captures) or single-digit OCR noise (191848878 vs 1691848878 — one digit dropped; PO/26/0000145236 vs ...143236 — one digit substituted) — not reliably parser-fixable without risking corrupting currently-correct extractions.
    • do-008.jpg looks like a ground-truth labeling error, not a parser or OCR bug — its noPO/noSO/noDO are all cleanly, unambiguously OCR'd (no visual noise) but are completely different digit sequences from manual_labels.json's values, not a plausible misread of them. Flagged here for human review rather than "fixed" by bending the parser to match a ground-truth value that's plausibly just wrong for this image.
    • Found and fixed one real parser logic bug (not an OCR-quality issue): in do-026.jpg, the SO field OCR'd as garbage ("No. SO : 06 No. 2026", date text bleeding into the SO cell) so extraction correctly fell through to "Not Found" — but the global fallback (parser.ts, "Global pattern scanning fallback", \b16\d{8}\b scan) then backfilled it with the only 16xxxxxxxx-shaped number in the whole document, which was actually the already-correctly-extracted noDO value — silently duplicating a DO number into the SO field. Fixed by excluding values already assigned to noSO/noDO/noPO from that global backfill pool. Confirmed via a direct parseDOMetadata() unit call and re-running the harness: noSO now correctly shows "Not Found" instead of the fabricated duplicate. This doesn't move the overall percentage (a wrong value and "Not Found" are both mismatches against ground truth) — it's a correctness fix, not an accuracy-score fix: previously a plausible-looking wrong number could reach the database unreviewed; now it's honestly flagged as missing instead.
    • All 48 parser.test.ts unit tests still pass; full harness re-run shows no regressions (plat/noSO/noPO/noDO/tanggal/etc. all unchanged except the corrected do-026.jpg noSO value described above).
    • Conclusion: the 95% target is already met at the aggregate level, and the remaining gap is now predominantly an OCR/layout-model accuracy ceiling (garbled text, misclassified stamp/seal regions), not a parser.ts logic gap — further parser.ts tuning on the current 37-image set has low remaining ROI. Future accuracy work should target the OCR/layout pipeline itself (config/, PaddleOCR-VL prompt/model tuning) or expanding the test set to catch different failure modes, not more regex tweaks.

Not in scope for either 2.1/2.2 (per next-implementation.md's own note, still true 2026-07-08): batch/lot number extraction doesn't exist in classify_ocr_server.py in either project (only SKU, product name, expiry date are extracted) — if requested later, follow the same OCR-regex-cascade pattern already used for expiry-date extraction.

3. Backend — Postgres Data Layer

pfm-web-app/src/db/

  • 3.1 [DONE] Wrapped the ocr_items delete-then-reinsert logic in both /api/parse and the /api/v1/documents/[id] PUT route inside a single DB transaction (withTransaction helper in db/index.ts), so a mid-loop failure can no longer leave a document with a correct header but silently missing items — it now rolls back to the previous item set instead. Completed 2026-07-08.
  • 3.2 [DONE 2026-07-08] Hashed accounts.password with bcryptjs (not native bcrypt — the pfm-web-app Docker stage is node:20-slim with no build toolchain, so a native addon would break the image build). db/init.ts now hashes the seed password and idempotently migrates any existing plaintext rows (WHERE password NOT LIKE '$2%') on every startup — safe to re-run, no double-hashing. api/v1/auth/login/route.ts now fetches by username only and compares with bcrypt.compareSync, plus a small added guard for missing username/password (previously would have fallen through to a DB query with undefined; now a clean 401). Installed via docker exec paddleocr-pfm-web-app npm install bcryptjs (container has its own anonymous node_modules volume, not host-synced) + docker compose restart pfm-web-app. Verified: SELECT password FROM accounts shows a $2b$10$... hash, curl login with the correct password succeeds (200 + token), wrong password and missing password both return a clean 401 (not a 500). This unblocks task 1.3.
  • 3.3 [TODO] Add a unique index on documents.file_hash so the dedup check from task 1.1 can do an efficient existing-document lookup before insert instead of a full scan.

4. DevOps — Docker & Dev Tunnel

../docker-compose.yml, ../docker-compose.demo.yml, ../start-dev-tunnel.ps1

  • 4.1 [TODO] Decide and document (in root README.md/CLAUDE.md) a deliberate policy for when the PoC/demo runs docker-compose.demo.yml (production build) vs. the dev-mode default — docker-compose.yml alone runs npm run dev, a known throughput ceiling already flagged in the reliability audit.
  • 4.2 [TODO] Add a startup healthcheck/readiness gate for pipeline-api/vLLM in docker-compose.yml so the Next.js gateway doesn't accept uploads before the GPU pipeline is actually ready to serve them.
  • 4.3 [TODO] Extend start-dev-tunnel.ps1 to verify the ngrok tunnel/LAN IP is actually reachable (not just started) before reporting success, since a stale tunnel is the documented first failure point for login/upload from the Flutter app.

Sections 1-4 migrated 2026-07-08 from root plans/next-enhancements.md sections 5-8 (originally populated the same day via a backend-scoped e/enhance run, grounded in a prior end-to-end reliability audit and, for §2's Product/SKU scan sub-feature, next-implementation.md's status table). This file is the sole home for backend enhancement tasks — the root copy's sections 5-8 were removed outright from plans/next-enhancements.md (not just frozen) per root AGENTS.md's "Scope: excludes backend/" note; the already-shipped 3.1/7.1 DB-transaction fix stays recorded in root docs/feature-list.md regardless, since that's a shipped-feature log, not a backlog.

next-implementation.md itself was deleted 2026-07-08 once its full content (decisions, verified state, action steps, verification checklist) was folded into §2 above with corrected statuses — it's no longer a separate source of truth.