Backend (app-pfm-ocr-v2/backend): - Product/SKU scan feature complete: trained DINOv2 index (118 reference photos, 16 SKU classes) and YOLO classifier (83.3% top-1 val accuracy), fixed scripts/install-pipeline.sh (was missing ultralytics/torch), fully browser-verified end-to-end on /scan-pfm. Mobile m-scan-pfm page cancelled (Flutter app handles mobile; web UI is desktop-only for pipeline testing). - Fixed a real data-loss bug: Save Ground Truth (scan-pfm and the DO-flow's manual-label) was silently writing into the pfm-web-app container's ephemeral filesystem instead of the host, because /sources wasn't bind-mounted in docker-compose.yml. Added the mount, recovered an orphaned entry. - accounts.password is now bcrypt-hashed (bcryptjs, idempotent migration in db/init.ts) instead of plaintext; login route compares hashes. - /api/v1/documents/* (list, PUT, upload) now enforces real 401 auth, matching what the Flutter client already sends. The "classic" routes deliberately stay open — they're dev-only web UI with no login flow and won't exist in production. - OCR accuracy investigated end-to-end: real baseline is 95.10% overall (target met; accuracy_report.md was stale at 75.04%, now flagged). Fixed one genuine parser.ts bug (SO/DO field duplication in the global fallback regex); remaining gaps are OCR/layout-model limitations, not parser bugs. - Adopted a standalone copy of the fhanyuh/agents-settings e/n workflow scoped to backend/ (AGENTS.md Part A/B split, SKILLS.md, plans/, docs/), independent of the root copy which now covers Flutter only. - next-implementation.md deleted; content folded into backend/plans/next-enhancements.md for traceability. Root: - Adopted fhanyuh/agents-settings kit (AGENTS.md, SKILLS.md, plans/, docs/feature-list.md), scoped to the Flutter app only. - Pending documents queue now persists to Hive (lib/core/storage) instead of memory-only, surviving an app kill mid-upload. Removed backend_backup/ (stale Express/Prisma prototype, superseded by pfm-web-app) and the completed plans/next-enhancement-plan.md checklist. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
22 KiB
Next Enhancements (backend)
Working backlog driven by the e/enhance and n/next triggers defined in
AGENTS.md Part B. Split out 2026-07-08 from root
plans/next-enhancements.md sections 5-8, renumbered 1-4 — this file now owns all
backend enhancement tracking going forward; the root file only tracks Flutter
sections from this point on. Statuses below were re-verified against the live code
at split time (not copied blind):
- 1.1
file_hashdedup — confirmed present,v1/documents/upload/route.ts:59,98. - 1.2
AbortSignal.timeouton both hops — confirmed present inparse/route.tsandv1/documents/upload/route.ts. - 1.3 classic routes skipping auth — confirmed still true at split time; later cancelled as out-of-scope (dev-only web UI, no production benefit) — see §1 below, 1.4 shipped instead.
- 3.1
withTransaction— confirmed present and used indb/index.ts,api/parse/route.ts,api/v1/documents/[id]/route.ts. - 2.1
pfm-web-app/public/produk-pfm/models/— corrected 2026-07-08 (was checked against the wrong path,backend/models, in an earlier pass this same day): the directory does exist. Later the same day, task 2.1 itself was completed — see §2 below — so the artifacts now exist too. - 2.2
m-scan-pfm/page.tsx— confirmed still missing; desktopscan-pfm/page.tsxconfirmed present, and now confirmed full feature parity (confidence bar, top-5 candidate list, expiry-date crop preview all present, 1169 lines) — see §2 below, migrated fromnext-implementation.md(deleted 2026-07-08, content now lives here). - 3.2 password hashing — confirmed plaintext at split time (2026-07-08 morning); fixed later the same day, see §3 below.
- 3.3 unique index on
documents.file_hash— confirmed still absent (onlyfilenamehas a UNIQUE constraint indb/init.ts). - 4.2 healthcheck for
pipeline-api/vLLM — confirmed nohealthcheckblock in rootdocker-compose.yml.
Format
Tasks are grouped under a numbered section per backend module. Each section gets exactly 3 tasks:
## 1. <Section / Module Name>
- **1.1** [TODO] <clear, specific description of the functional change>
- **1.2** [TODO] <...>
- **1.3** [TODO] <...>
When a task is picked up via n/next, its clarified acceptance criteria (from
AGENTS.md Part B §B2a) are appended directly under it. When complete, the status
flips to [DONE] and the feature is logged in
docs/feature-list.md (this dir).
1. Backend — Next.js API Gateway
pfm-web-app/src/app/api/
- 1.1 [DONE]
MakeInvestigated 2026-07-08: thedocuments/upload/route.tsreturn201immediately...file_hashdedup check is already shipped pre-existing inv1/documents/upload/route.ts(lines 58-89) — a duplicate upload returns the existing document instead of re-inserting/re-parsing. The "return 201 immediately, don't await OCR" half turned out to be a deliberate, already-documented tradeoff (see the route's own comment): the pipeline call stays synchronous but is now bounded byAbortSignal.timeout(210_000). Not changed further this round since it's an intentional decision, not an oversight. - 1.2 [DONE] Investigated 2026-07-08: already shipped pre-existing. Both hops of the upload →
/api/parse→ pipeline-API chain already haveAbortSignal.timeout(210s andPIPELINE_TIMEOUT_MS=90s respectively), with comments cross-referencing each other. - 1.3 [CANCELLED 2026-07-08]
Wire the existing JWT auth onto the classic API routes (— investigated after 3.2 unblocked it, found the premise was stale: those "classic" routes are used exclusively by the dev-only web UI (root DO-PFM page,/api/upload,/api/scan-pfm,/api/parse,/api/history, etc.)scan-pfm,manual-label), which has no login screen and never sends a token — the user confirmed this web UI won't exist in production, so enforcing auth there would break it today for zero production benefit, and optional identification there is a no-op (nobody ever sends a token). Same reasoning as 2.2's cancellation. See 1.4 below for what was actually shipped instead. - 1.4 [DONE 2026-07-08] Enforced real 401 auth on
/api/v1/documents/*instead — this is the actual production API surface, already fully supported by the Flutter client (lib/features/auth/auth_provider.dartdoes a real login and stores the JWT;lib/core/network/api_client.dart's interceptor already attachesAuthorization: Bearer <token>to every request — rootCLAUDE.md's "demo-stub, no real server-side auth" claim for Flutter was itself stale). Before this,v1/documents/route.ts(list) andv1/documents/[id]/route.ts(PUT) didn't check auth at all, andv1/documents/upload/route.tsonly optionally read the token (never rejected a missing one). AddedgetAccountFromAuthHeader()(utils/auth.ts, pre-existing helper) + a401guard to all three;OPTIONS(CORS preflight) on all three left untouched. Did not add per-account data filtering to the document list (still returns all non-sample parsed documents regardless of uploader — that's a feature change, not an auth fix; flagged as a possible future task). Did not add Flutter-side 401→auto-logout handling (api_client.darthas no interceptor for it; a 401 surfaces as a normalApiExceptiontoday) — a Flutter-side follow-up, not a backend task.- Verified end-to-end via
curl: all three endpoints return 401 with no token; afterPOST /api/v1/auth/loginwithadmin/passwordto get a real token, the same three endpoints succeed withAuthorization: Bearer <token>(GET200 with real data,PUTreaches its normal 404-for-bad-id logic,POST uploadreaches its normal downstream logic) — confirming the auth check itself works without touching any other behavior.OPTIONSon all three still returns 204 unauthenticated.
- Verified end-to-end via
2. Backend — OCR Pipeline & Accuracy
config/, pfm-web-app/src/utils/parser.ts, accuracy regression harness (see CLAUDE.md)
Product/SKU scan sub-feature — decisions & state (migrated 2026-07-08 from
next-implementation.md, which is now deleted; this section is the sole source of
truth for it going forward):
-
Decisions made: new pages are standalone routes (
scan-pfm— done), following the same self-contained pattern asmanual-label/page.tsx— own header/theme, no shared chrome with the root DO-PFM page. Build to full feature parity with the oldai-ocr-pfm-2026pages (confidence bars, top-5 SKU candidates, visual OCR overlay, expiry-date crop preview), not a lean MVP. -
Scope correction 2026-07-08 (later):
m-scan-pfm/page.tsx(mobile web page) is not needed — user clarified the webscan-pfmpage is desktop-only, used for testing the pipeline, not a production mobile surface. Real mobile scanning is already handled by the Flutter app instead. Focus for this sub-feature going forward is backend services (model artifacts, pipeline correctness), not any additional frontend. See 2.2 below — cancelled, not reassigned. -
Already shipped, re-verified 2026-07-08 (all confirmed present, not assumed):
config/classify_ocr_server.py(DINOv2 + YOLO classify, SKU/expiry/name OCR),pfm-web-app/public/produk-pfm/index_dinov2.py+train_classifier.py,api/scan-pfm/route.ts(Levenshtein match againstsku_master),api/produk-pfm/route.ts, DB schema (sku_master/store_master/arena_runsindb/init.ts),scripts/serve-pipeline.shwiring,Dockerfile'sultralyticsinstall in the pipeline-api stage,nginx.confroutes for/scan-pfm,/m-scan-pfm,/produk-pfm. -
Stale claims corrected 2026-07-08 (next-implementation.md said these didn't exist; they now do):
pfm-web-app/public/produk-pfm/foto-kemasan-v2/dataset — 16 SKU subfolders now present (was "❌ does not exist");scan-pfm/page.tsx— now exists at full feature parity, 1169 lines (was "❌ does not exist"). -
Newly observed, not in original doc's scope:
/produk-pfmhas an nginx proxy block and an API route (api/produk-pfm/route.ts) but no matching frontend page (src/app/produk-pfm/page.tsxdoesn't exist) — dead route, same class of issue as scan-pfm/m-scan-pfm were. Not turned into a task below since it wasn't part of the original decision record; flagging for a futuree/enhancepass to pick up. -
2.1 [DONE 2026-07-08] Built the initial model artifacts. Bare-metal (
./scripts/install-pipeline.sh) doesn't work in this Windows dev environment —paddlepaddle-gpu's wheel index is Linux-only — so this ran entirely through Docker instead:- Added
ultralytics>=8.0toscripts/install-pipeline.sh(was missing there even though theDockerfilepipeline-api stage already installed it — fixed for bare-metal parity, though bare-metal training itself still isn't viable on Windows). docker compose build pipeline-apifrom repo root — cached layers reused, only the finalCOPY . /appre-ran, picking up the (until-then-unbaked) 16-folder dataset.- Ran a one-off
docker run --gpus allfrom the fresh image withmodels/mounted writable (not the live service's:romount) — the livepipeline-apicontainer's copy was stale (image predated the dataset) and read-only either way, so training couldn't happen through it directly.index_dinov2.py→ indexed 118/118 images across all 16 classes →models/dinov2_index.pkl(196KB).train_classifier.py train --imgsz 224→ 100 epochs (~1.5 min), 83.3% top-1 / 90% top-5 validation accuracy (thin dataset, 2-16 images/class — accuracy will improve as more reference photos are added) →models/produk-pfm-classifier-26n-100e-2026-07-08.pt(3.2MB) +.onnxexport (5.9MB). docker compose restart pipeline-api(bind-mountedmodels/is live, no rebuild needed for the running service) — confirmed viadocker logs: "DINOv2 model loaded successfully" and "YOLO model loaded successfully... Using classifier weights: produk-pfm-classifier-26n-100e-2026-07-08.pt".
- Note on Git Bash + Docker on Windows:
docker run ... bash -c "/app/..."fails with exit 127 ("No such file or directory") because MSYS path-conversion mangles the leading/app/...argument into a Windows path — fix isMSYS_NO_PATHCONV=1before thedocker run/docker execinvocation.
Full browser verification, 2026-07-08 (later same day) — ran
/scan-pfmlive (http://localhost:3000/scan-pfm, plus confirmedhttp://localhost:8000/scan-pfmthrough nginx) via Claude-in-Chrome, selecting the 16-image "FIESTA STIKIE" sample and clicking through the full flow, not just checking container logs:- AI Classification: 100.00% confidence, correct class, "Model Active" — confirms the freshly-trained YOLO weights are actually being used, not a fallback.
- "View Top 5 Predicted Probabilities" — expands correctly with 5 ranked candidates + "Use" override buttons.
- AI OCR Text Extraction — expiry date
26/02/2027extracted correctly with a label-crop preview image. - "Possible Match from SKU Master" — Levenshtein-ranked list, correct best match at 71% similarity.
- Visual Grid tab, Spotting Grid tab, Raw Response (JSON) tab — all render correctly with real data.
- Bug found and fixed: "Save Ground Truth" returned HTTP 200 and looked
successful in the UI, but
api/manual-label-scan/route.tswrites topath.join(process.cwd(), "..", "sources", "product_manual_labels.json")— inside thepfm-web-appcontainer this resolves to/sources/..., which was not a bind-mounted path in rootdocker-compose.yml(only/app,/uploads, and the Docker socket were mounted for that service). Every save was landing in the container's ephemeral writable layer, invisible to the host, to git, and to any other tooling — confirmed by finding an orphaned entry from 2026-07-07 (an earlier, unrelated test) sitting in the container with no trace on disk. Fix: added./backend/sources:/sourcesto thepfm-web-appservice's volumes in rootdocker-compose.yml, recovered the orphaned entry viadocker cpbefore recreating the container, thendocker compose up -d pfm-web-app(recreated bothpfm-web-appand, as its dependency,pipeline-api— re-verified both models reloaded correctly after). Re-ran the full scan + save and confirmed the entry now lands directly inbackend/sources/product_manual_labels.jsonon the host. - Confirmed (2026-07-08, follow-up check) the same bug pattern in
api/manual-label/route.ts(DO-flow ground-truth save,sources/manual_labels.json, identicalprocess.cwd()-relative path) is now also fixed by the same volume mount — verifiedmd5summatches exactly between the container's/sources/manual_labels.jsonand the host'sbackend/sources/manual_labels.json, with a fresh host mtime. No separate fix needed; the onedocker-compose.ymlchange covers both routes. - Acceptance met:
models/dinov2_index.pklexists, "Indexed 118/118 images" log confirmed; both models load cleanly onpipeline-apirestart per container logs. End-to-end/scan-pfmUI test (upload a real photo, confirm live top-5 result) — done, see the full browser verification note above (2026-07-08, later still).
- Added
-
2.2 [CANCELLED 2026-07-08]
Port— user decided this isn't needed: the webm-scan-pfm/page.tsx(mobile) fromai-ocr-pfm-2026into v2scan-pfmpage is desktop-only (used to test the pipeline), and real mobile product scanning is handled by the Flutter app, not a web mobile page. The/m-scan-pfmnginx block stays dead intentionally — do not resurrect this task without an explicit ask.docs/feature-list.mdwill not get an entry for it. -
2.3 [DONE 2026-07-08] Ran the accuracy regression harness. Key finding:
sources/accuracy_report.mdwas badly stale (said 75.04% overall) — the real current baseline (confirmed via a fresh harness run at commit3df9f6e, matchingsources/accuracy_history.jsonl's latest entry exactly) is 95.10% overall, already at/above the 95% target. Added a staleness banner toaccuracy_report.mdpointing here instead of leaving it to mislead future work.- Investigated the worst field,
plat(67.6%, unchanged between L1/L3 — no sanitization ever touches it) by pulling raw OCR text for every mismatching image fromdocuments.layout_parsing_resultin Postgres. Of 13 mismatches: 11 are the plate region getting classified by the layout model as animage_box/seal_box(never becomes text at all — e.g.Truck No.cell literally contains an<img>tag, not digits) or the suffix OCR'd as digits instead of letters ("B 9536 000"vs ground truth"B 9536 UCO"—0forO/C). Not fixable inparser.ts— there's no text to parse in most cases; the remaining 1-2 are single-character OCR misreads ("VOC"vs"VCC") too fragile to correct generically without risking false positives elsewhere. - Cross-checked all other mismatching fields (
noSO/noPO/noDO/tanggal, 12 mismatches total via--dump-json) the same way. Same story: mostly garbled/missing OCR text (image regions, blank captures) or single-digit OCR noise (191848878vs1691848878— one digit dropped;PO/26/0000145236vs...143236— one digit substituted) — not reliably parser-fixable without risking corrupting currently-correct extractions. do-008.jpglooks like a ground-truth labeling error, not a parser or OCR bug — itsnoPO/noSO/noDOare all cleanly, unambiguously OCR'd (no visual noise) but are completely different digit sequences frommanual_labels.json's values, not a plausible misread of them. Flagged here for human review rather than "fixed" by bending the parser to match a ground-truth value that's plausibly just wrong for this image.- Found and fixed one real parser logic bug (not an OCR-quality issue): in
do-026.jpg, the SO field OCR'd as garbage ("No. SO : 06 No. 2026", date text bleeding into the SO cell) so extraction correctly fell through to"Not Found"— but the global fallback (parser.ts, "Global pattern scanning fallback",\b16\d{8}\bscan) then backfilled it with the only16xxxxxxxx-shaped number in the whole document, which was actually the already-correctly-extractednoDOvalue — silently duplicating a DO number into the SO field. Fixed by excluding values already assigned tonoSO/noDO/noPOfrom that global backfill pool. Confirmed via a directparseDOMetadata()unit call and re-running the harness:noSOnow correctly shows"Not Found"instead of the fabricated duplicate. This doesn't move the overall percentage (a wrong value and "Not Found" are both mismatches against ground truth) — it's a correctness fix, not an accuracy-score fix: previously a plausible-looking wrong number could reach the database unreviewed; now it's honestly flagged as missing instead. - All 48
parser.test.tsunit tests still pass; full harness re-run shows no regressions (plat/noSO/noPO/noDO/tanggal/etc. all unchanged except the correcteddo-026.jpgnoSOvalue described above). - Conclusion: the 95% target is already met at the aggregate level, and the
remaining gap is now predominantly an OCR/layout-model accuracy ceiling (garbled
text, misclassified stamp/seal regions), not a
parser.tslogic gap — furtherparser.tstuning on the current 37-image set has low remaining ROI. Future accuracy work should target the OCR/layout pipeline itself (config/, PaddleOCR-VL prompt/model tuning) or expanding the test set to catch different failure modes, not more regex tweaks.
- Investigated the worst field,
Not in scope for either 2.1/2.2 (per next-implementation.md's own note, still true
2026-07-08): batch/lot number extraction doesn't exist in classify_ocr_server.py
in either project (only SKU, product name, expiry date are extracted) — if
requested later, follow the same OCR-regex-cascade pattern already used for
expiry-date extraction.
3. Backend — Postgres Data Layer
pfm-web-app/src/db/
- 3.1 [DONE] Wrapped the
ocr_itemsdelete-then-reinsert logic in both/api/parseand the/api/v1/documents/[id]PUT route inside a single DB transaction (withTransactionhelper indb/index.ts), so a mid-loop failure can no longer leave a document with a correct header but silently missing items — it now rolls back to the previous item set instead. Completed 2026-07-08. - 3.2 [DONE 2026-07-08] Hashed
accounts.passwordwithbcryptjs(not nativebcrypt— thepfm-web-appDocker stage isnode:20-slimwith no build toolchain, so a native addon would break the image build).db/init.tsnow hashes the seed password and idempotently migrates any existing plaintext rows (WHERE password NOT LIKE '$2%') on every startup — safe to re-run, no double-hashing.api/v1/auth/login/route.tsnow fetches by username only and compares withbcrypt.compareSync, plus a small added guard for missing username/password (previously would have fallen through to a DB query withundefined; now a clean 401). Installed viadocker exec paddleocr-pfm-web-app npm install bcryptjs(container has its own anonymousnode_modulesvolume, not host-synced) +docker compose restart pfm-web-app. Verified:SELECT password FROM accountsshows a$2b$10$...hash,curllogin with the correct password succeeds (200 + token), wrong password and missing password both return a clean 401 (not a 500). This unblocks task 1.3. - 3.3 [TODO] Add a unique index on
documents.file_hashso the dedup check from task 1.1 can do an efficient existing-document lookup before insert instead of a full scan.
4. DevOps — Docker & Dev Tunnel
../docker-compose.yml, ../docker-compose.demo.yml, ../start-dev-tunnel.ps1
- 4.1 [TODO] Decide and document (in root
README.md/CLAUDE.md) a deliberate policy for when the PoC/demo runsdocker-compose.demo.yml(production build) vs. the dev-mode default —docker-compose.ymlalone runsnpm run dev, a known throughput ceiling already flagged in the reliability audit. - 4.2 [TODO] Add a startup healthcheck/readiness gate for
pipeline-api/vLLM indocker-compose.ymlso the Next.js gateway doesn't accept uploads before the GPU pipeline is actually ready to serve them. - 4.3 [TODO] Extend
start-dev-tunnel.ps1to verify the ngrok tunnel/LAN IP is actually reachable (not just started) before reporting success, since a stale tunnel is the documented first failure point for login/upload from the Flutter app.
Sections 1-4 migrated 2026-07-08 from root plans/next-enhancements.md sections
5-8 (originally populated the same day via a backend-scoped e/enhance run,
grounded in a prior end-to-end reliability audit and, for §2's Product/SKU scan
sub-feature, next-implementation.md's status table). This file is the sole home
for backend enhancement tasks — the root copy's sections 5-8 were removed outright
from plans/next-enhancements.md (not just frozen) per root AGENTS.md's "Scope:
excludes backend/" note; the already-shipped 3.1/7.1 DB-transaction fix stays
recorded in root docs/feature-list.md regardless, since that's a shipped-feature
log, not a backlog.
next-implementation.md itself was deleted 2026-07-08 once its full content
(decisions, verified state, action steps, verification checklist) was folded into
§2 above with corrected statuses — it's no longer a separate source of truth.