Adopt agents-settings kit, ship Product/SKU scan models, harden auth, verify OCR accuracy

Backend (app-pfm-ocr-v2/backend):
- Product/SKU scan feature complete: trained DINOv2 index (118 reference
  photos, 16 SKU classes) and YOLO classifier (83.3% top-1 val accuracy),
  fixed scripts/install-pipeline.sh (was missing ultralytics/torch), fully
  browser-verified end-to-end on /scan-pfm. Mobile m-scan-pfm page cancelled
  (Flutter app handles mobile; web UI is desktop-only for pipeline testing).
- Fixed a real data-loss bug: Save Ground Truth (scan-pfm and the DO-flow's
  manual-label) was silently writing into the pfm-web-app container's
  ephemeral filesystem instead of the host, because /sources wasn't
  bind-mounted in docker-compose.yml. Added the mount, recovered an
  orphaned entry.
- accounts.password is now bcrypt-hashed (bcryptjs, idempotent migration
  in db/init.ts) instead of plaintext; login route compares hashes.
- /api/v1/documents/* (list, PUT, upload) now enforces real 401 auth,
  matching what the Flutter client already sends. The "classic" routes
  deliberately stay open — they're dev-only web UI with no login flow and
  won't exist in production.
- OCR accuracy investigated end-to-end: real baseline is 95.10% overall
  (target met; accuracy_report.md was stale at 75.04%, now flagged). Fixed
  one genuine parser.ts bug (SO/DO field duplication in the global fallback
  regex); remaining gaps are OCR/layout-model limitations, not parser bugs.
- Adopted a standalone copy of the fhanyuh/agents-settings e/n workflow
  scoped to backend/ (AGENTS.md Part A/B split, SKILLS.md, plans/, docs/),
  independent of the root copy which now covers Flutter only.
- next-implementation.md deleted; content folded into
  backend/plans/next-enhancements.md for traceability.

Root:
- Adopted fhanyuh/agents-settings kit (AGENTS.md, SKILLS.md, plans/,
  docs/feature-list.md), scoped to the Flutter app only.
- Pending documents queue now persists to Hive (lib/core/storage) instead
  of memory-only, surviving an app kill mid-upload.

Removed backend_backup/ (stale Express/Prisma prototype, superseded by
pfm-web-app) and the completed plans/next-enhancement-plan.md checklist.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
Rafhan Mazaya FathurrahmanandClaude Sonnet 5 committed 2026-07-08 11:56:32 +07:00
1 parent 3df9f6ec5d
commit e60ab63154
129 files changed
+8520 -6684

No files matched your search

+255
View File
@@ -0,0 +1,255 @@
# Next Enhancements (backend)
Working backlog driven by the `e`/`enhance` and `n`/`next` triggers defined in
[AGENTS.md](../AGENTS.md) Part B. Split out 2026-07-08 from root
`plans/next-enhancements.md` sections 5-8, renumbered 1-4 — this file now owns all
backend enhancement tracking going forward; the root file only tracks Flutter
sections from this point on. Statuses below were re-verified against the live code
at split time (not copied blind):
- **1.1** `file_hash` dedup — confirmed present, `v1/documents/upload/route.ts:59,98`.
- **1.2** `AbortSignal.timeout` on both hops — confirmed present in `parse/route.ts`
and `v1/documents/upload/route.ts`.
- **1.3** classic routes skipping auth — confirmed still true at split time; later cancelled as out-of-scope (dev-only web UI, no production benefit) — see §1 below, 1.4 shipped instead.
- **3.1** `withTransaction` — confirmed present and used in `db/index.ts`,
`api/parse/route.ts`, `api/v1/documents/[id]/route.ts`.
- **2.1** `pfm-web-app/public/produk-pfm/models/` — corrected 2026-07-08 (was
checked against the wrong path, `backend/models`, in an earlier pass this same
day): the directory does exist. Later the same day, task 2.1 itself was completed
— see §2 below — so the artifacts now exist too.
- **2.2** `m-scan-pfm/page.tsx` — confirmed still missing; desktop `scan-pfm/page.tsx`
confirmed present, **and now confirmed full feature parity** (confidence bar,
top-5 candidate list, expiry-date crop preview all present, 1169 lines) — see §2
below, migrated from `next-implementation.md` (deleted 2026-07-08, content now
lives here).
- **3.2** password hashing — confirmed plaintext at split time (2026-07-08 morning); fixed later the same day, see §3 below.
- **3.3** unique index on `documents.file_hash` — confirmed still absent (only
`filename` has a UNIQUE constraint in `db/init.ts`).
- **4.2** healthcheck for `pipeline-api`/vLLM — confirmed no `healthcheck` block in
root `docker-compose.yml`.
## Format
Tasks are grouped under a numbered section per backend module. Each section gets
exactly 3 tasks:
```
## 1. <Section / Module Name>
- **1.1** [TODO] <clear, specific description of the functional change>
- **1.2** [TODO] <...>
- **1.3** [TODO] <...>
```
When a task is picked up via `n`/`next`, its clarified acceptance criteria (from
`AGENTS.md` Part B §B2a) are appended directly under it. When complete, the status
flips to `[DONE]` and the feature is logged in
[docs/feature-list.md](../docs/feature-list.md) (this dir).
---
## 1. Backend — Next.js API Gateway
`pfm-web-app/src/app/api/`
- **1.1** [DONE] ~~Make `documents/upload/route.ts` return `201` immediately...~~ Investigated 2026-07-08: the `file_hash` dedup check is already shipped pre-existing in `v1/documents/upload/route.ts` (lines 58-89) — a duplicate upload returns the existing document instead of re-inserting/re-parsing. The "return 201 immediately, don't await OCR" half turned out to be a deliberate, already-documented tradeoff (see the route's own comment): the pipeline call stays synchronous but is now bounded by `AbortSignal.timeout(210_000)`. Not changed further this round since it's an intentional decision, not an oversight.
- **1.2** [DONE] Investigated 2026-07-08: already shipped pre-existing. Both hops of the upload → `/api/parse` → pipeline-API chain already have `AbortSignal.timeout` (210s and `PIPELINE_TIMEOUT_MS`=90s respectively), with comments cross-referencing each other.
- **1.3** [CANCELLED 2026-07-08] ~~Wire the existing JWT auth onto the classic API routes (`/api/upload`, `/api/scan-pfm`, `/api/parse`, `/api/history`, etc.)~~ — investigated after 3.2 unblocked it, found the premise was stale: those "classic" routes are used exclusively by the dev-only web UI (root DO-PFM page, `scan-pfm`, `manual-label`), which has no login screen and never sends a token — the user confirmed this web UI won't exist in production, so enforcing auth there would break it today for zero production benefit, and optional identification there is a no-op (nobody ever sends a token). Same reasoning as 2.2's cancellation. See **1.4** below for what was actually shipped instead.
- **1.4** [DONE 2026-07-08] Enforced real 401 auth on `/api/v1/documents/*` instead — this is the actual production API surface, already fully supported by the Flutter client (`lib/features/auth/auth_provider.dart` does a real login and stores the JWT; `lib/core/network/api_client.dart`'s interceptor already attaches `Authorization: Bearer <token>` to every request — root `CLAUDE.md`'s "demo-stub, no real server-side auth" claim for Flutter was itself stale). Before this, `v1/documents/route.ts` (list) and `v1/documents/[id]/route.ts` (PUT) didn't check auth at all, and `v1/documents/upload/route.ts` only optionally read the token (never rejected a missing one). Added `getAccountFromAuthHeader()` (`utils/auth.ts`, pre-existing helper) + a `401` guard to all three; `OPTIONS` (CORS preflight) on all three left untouched. Did **not** add per-account data filtering to the document list (still returns all non-sample parsed documents regardless of uploader — that's a feature change, not an auth fix; flagged as a possible future task). Did **not** add Flutter-side 401→auto-logout handling (`api_client.dart` has no interceptor for it; a 401 surfaces as a normal `ApiException` today) — a Flutter-side follow-up, not a backend task.
- Verified end-to-end via `curl`: all three endpoints return 401 with no token; after `POST /api/v1/auth/login` with `admin`/`password` to get a real token, the same three endpoints succeed with `Authorization: Bearer <token>` (`GET` 200 with real data, `PUT` reaches its normal 404-for-bad-id logic, `POST upload` reaches its normal downstream logic) — confirming the auth check itself works without touching any other behavior. `OPTIONS` on all three still returns 204 unauthenticated.
## 2. Backend — OCR Pipeline & Accuracy
`config/`, `pfm-web-app/src/utils/parser.ts`, accuracy regression harness (see `CLAUDE.md`)
**Product/SKU scan sub-feature — decisions & state** (migrated 2026-07-08 from
`next-implementation.md`, which is now deleted; this section is the sole source of
truth for it going forward):
- **Decisions made**: new pages are **standalone routes** (`scan-pfm` — done), following the same self-contained pattern as `manual-label/page.tsx` — own
header/theme, no shared chrome with the root DO-PFM page. Build to **full feature
parity** with the old `ai-ocr-pfm-2026` pages (confidence bars, top-5 SKU
candidates, visual OCR overlay, expiry-date crop preview), not a lean MVP.
- **Scope correction 2026-07-08 (later)**: `m-scan-pfm/page.tsx` (mobile web page)
is **not needed** — user clarified the web `scan-pfm` page is desktop-only,
used for testing the pipeline, not a production mobile surface. Real mobile
scanning is already handled by the Flutter app instead. Focus for this
sub-feature going forward is **backend services** (model artifacts, pipeline
correctness), not any additional frontend. See 2.2 below — cancelled, not
reassigned.
- **Already shipped, re-verified 2026-07-08** (all confirmed present, not assumed):
`config/classify_ocr_server.py` (DINOv2 + YOLO classify, SKU/expiry/name OCR),
`pfm-web-app/public/produk-pfm/index_dinov2.py` + `train_classifier.py`,
`api/scan-pfm/route.ts` (Levenshtein match against `sku_master`),
`api/produk-pfm/route.ts`, DB schema (`sku_master`/`store_master`/`arena_runs` in
`db/init.ts`), `scripts/serve-pipeline.sh` wiring, `Dockerfile`'s `ultralytics`
install in the pipeline-api stage, `nginx.conf` routes for `/scan-pfm`,
`/m-scan-pfm`, `/produk-pfm`.
- **Stale claims corrected 2026-07-08** (next-implementation.md said these didn't
exist; they now do): `pfm-web-app/public/produk-pfm/foto-kemasan-v2/` dataset —
16 SKU subfolders now present (was "❌ does not exist"); `scan-pfm/page.tsx` — now
exists at full feature parity, 1169 lines (was "❌ does not exist").
- **Newly observed, not in original doc's scope**: `/produk-pfm` has an nginx proxy
block and an API route (`api/produk-pfm/route.ts`) but no matching frontend page
(`src/app/produk-pfm/page.tsx` doesn't exist) — dead route, same class of issue as
scan-pfm/m-scan-pfm were. Not turned into a task below since it wasn't part of the
original decision record; flagging for a future `e`/`enhance` pass to pick up.
- **2.1** [DONE 2026-07-08] Built the initial model artifacts. Bare-metal
(`./scripts/install-pipeline.sh`) doesn't work in this Windows dev environment —
`paddlepaddle-gpu`'s wheel index is Linux-only — so this ran entirely through
Docker instead:
1. Added `ultralytics>=8.0` to `scripts/install-pipeline.sh` (was missing there
even though the `Dockerfile` pipeline-api stage already installed it — fixed
for bare-metal parity, though bare-metal training itself still isn't viable on
Windows).
2. `docker compose build pipeline-api` from repo root — cached layers reused, only
the final `COPY . /app` re-ran, picking up the (until-then-unbaked) 16-folder
dataset.
3. Ran a one-off `docker run --gpus all` from the fresh image with `models/`
mounted **writable** (not the live service's `:ro` mount) — the live
`pipeline-api` container's copy was stale (image predated the dataset) and
read-only either way, so training couldn't happen through it directly.
`index_dinov2.py` → indexed 118/118 images across all 16 classes →
`models/dinov2_index.pkl` (196KB). `train_classifier.py train --imgsz 224` → 100
epochs (~1.5 min), **83.3% top-1 / 90% top-5 validation accuracy** (thin dataset,
2-16 images/class — accuracy will improve as more reference photos are added) →
`models/produk-pfm-classifier-26n-100e-2026-07-08.pt` (3.2MB) + `.onnx` export
(5.9MB).
4. `docker compose restart pipeline-api` (bind-mounted `models/` is live, no
rebuild needed for the running service) — confirmed via `docker logs`: "DINOv2
model loaded successfully" and "YOLO model loaded successfully... Using
classifier weights: produk-pfm-classifier-26n-100e-2026-07-08.pt".
- Note on Git Bash + Docker on Windows: `docker run ... bash -c "/app/..."` fails
with exit 127 ("No such file or directory") because MSYS path-conversion mangles
the leading `/app/...` argument into a Windows path — fix is
`MSYS_NO_PATHCONV=1` before the `docker run`/`docker exec` invocation.
**Full browser verification, 2026-07-08 (later same day)** — ran `/scan-pfm` live
(`http://localhost:3000/scan-pfm`, plus confirmed `http://localhost:8000/scan-pfm`
through nginx) via Claude-in-Chrome, selecting the 16-image "FIESTA STIKIE"
sample and clicking through the full flow, not just checking container logs:
- AI Classification: 100.00% confidence, correct class, "Model Active" —
confirms the freshly-trained YOLO weights are actually being used, not a
fallback.
- "View Top 5 Predicted Probabilities" — expands correctly with 5 ranked
candidates + "Use" override buttons.
- AI OCR Text Extraction — expiry date `26/02/2027` extracted correctly with a
label-crop preview image.
- "Possible Match from SKU Master" — Levenshtein-ranked list, correct best match
at 71% similarity.
- Visual Grid tab, Spotting Grid tab, Raw Response (JSON) tab — all render
correctly with real data.
- **Bug found and fixed**: "Save Ground Truth" returned HTTP 200 and looked
successful in the UI, but `api/manual-label-scan/route.ts` writes to
`path.join(process.cwd(), "..", "sources", "product_manual_labels.json")` —
inside the `pfm-web-app` container this resolves to `/sources/...`, which
was **not a bind-mounted path** in root `docker-compose.yml` (only `/app`,
`/uploads`, and the Docker socket were mounted for that service). Every save
was landing in the container's ephemeral writable layer, invisible to the host,
to git, and to any other tooling — confirmed by finding an orphaned entry from
2026-07-07 (an earlier, unrelated test) sitting in the container with no trace
on disk. **Fix**: added `./backend/sources:/sources` to the `pfm-web-app`
service's volumes in root `docker-compose.yml`, recovered the orphaned entry
via `docker cp` before recreating the container, then `docker compose up -d
pfm-web-app` (recreated both `pfm-web-app` and, as its dependency,
`pipeline-api` — re-verified both models reloaded correctly after). Re-ran the
full scan + save and confirmed the entry now lands directly in
`backend/sources/product_manual_labels.json` on the host.
- **Confirmed (2026-07-08, follow-up check) the same bug pattern in
`api/manual-label/route.ts`** (DO-flow ground-truth save,
`sources/manual_labels.json`, identical `process.cwd()`-relative path) is now
also fixed by the same volume mount — verified `md5sum` matches exactly between
the container's `/sources/manual_labels.json` and the host's
`backend/sources/manual_labels.json`, with a fresh host mtime. No separate fix
needed; the one `docker-compose.yml` change covers both routes.
- Acceptance met: `models/dinov2_index.pkl` exists, "Indexed 118/118 images" log
confirmed; both models load cleanly on `pipeline-api` restart per container logs.
End-to-end `/scan-pfm` UI test (upload a real photo, confirm live top-5 result)
— done, see the full browser verification note above (2026-07-08, later still).
- **2.2** [CANCELLED 2026-07-08] ~~Port `m-scan-pfm/page.tsx` (mobile) from `ai-ocr-pfm-2026` into v2~~ — user decided this isn't needed: the web `scan-pfm` page is desktop-only (used to test the pipeline), and real mobile product scanning is handled by the Flutter app, not a web mobile page. The `/m-scan-pfm` nginx block stays dead intentionally — do not resurrect this task without an explicit ask. `docs/feature-list.md` will not get an entry for it.
- **2.3** [DONE 2026-07-08] Ran the accuracy regression harness. **Key finding:
`sources/accuracy_report.md` was badly stale** (said 75.04% overall) — the real
current baseline (confirmed via a fresh harness run at commit `3df9f6e`, matching
`sources/accuracy_history.jsonl`'s latest entry exactly) is **95.10% overall,
already at/above the 95% target**. Added a staleness banner to
`accuracy_report.md` pointing here instead of leaving it to mislead future work.
- **Investigated the worst field, `plat` (67.6%, unchanged between L1/L3 — no
sanitization ever touches it)** by pulling raw OCR text for every mismatching
image from `documents.layout_parsing_result` in Postgres. Of 13 mismatches: 11
are the plate region getting classified by the layout model as an `image_box`/
`seal_box` (never becomes text at all — e.g. `Truck No.` cell literally contains
an `<img>` tag, not digits) or the suffix OCR'd as digits instead of letters
(`"B 9536 000"` vs ground truth `"B 9536 UCO"` — `0` for `O`/`C`). **Not
fixable in `parser.ts`** — there's no text to parse in most cases; the
remaining 1-2 are single-character OCR misreads (`"VOC"` vs `"VCC"`) too
fragile to correct generically without risking false positives elsewhere.
- Cross-checked all other mismatching fields (`noSO`/`noPO`/`noDO`/`tanggal`,
12 mismatches total via `--dump-json`) the same way. Same story: mostly
garbled/missing OCR text (image regions, blank captures) or single-digit OCR
noise (`191848878` vs `1691848878` — one digit dropped;
`PO/26/0000145236` vs `...143236` — one digit substituted) — not reliably
parser-fixable without risking corrupting currently-correct extractions.
- **`do-008.jpg` looks like a ground-truth labeling error, not a parser or OCR
bug** — its `noPO`/`noSO`/`noDO` are all cleanly, unambiguously OCR'd (no
visual noise) but are completely different digit sequences from
`manual_labels.json`'s values, not a plausible misread of them. Flagged here
for human review rather than "fixed" by bending the parser to match a
ground-truth value that's plausibly just wrong for this image.
- **Found and fixed one real parser logic bug** (not an OCR-quality issue): in
`do-026.jpg`, the SO field OCR'd as garbage (`"No. SO : 06 No. 2026"`, date text
bleeding into the SO cell) so extraction correctly fell through to
`"Not Found"` — but the *global* fallback (`parser.ts`, "Global pattern
scanning fallback", `\b16\d{8}\b` scan) then backfilled it with the only
`16xxxxxxxx`-shaped number in the whole document, which was actually the
*already-correctly-extracted* `noDO` value — silently duplicating a DO number
into the SO field. Fixed by excluding values already assigned to
`noSO`/`noDO`/`noPO` from that global backfill pool. Confirmed via a direct
`parseDOMetadata()` unit call and re-running the harness:
`noSO` now correctly shows `"Not Found"` instead of the fabricated duplicate.
**This doesn't move the overall percentage** (a wrong value and "Not Found"
are both mismatches against ground truth) — it's a correctness fix, not an
accuracy-score fix: previously a plausible-looking *wrong* number could reach
the database unreviewed; now it's honestly flagged as missing instead.
- All 48 `parser.test.ts` unit tests still pass; full harness re-run shows no
regressions (`plat`/`noSO`/`noPO`/`noDO`/`tanggal`/etc. all unchanged except
the corrected `do-026.jpg` `noSO` value described above).
- **Conclusion**: the 95% target is already met at the aggregate level, and the
remaining gap is now predominantly an OCR/layout-model accuracy ceiling (garbled
text, misclassified stamp/seal regions), not a `parser.ts` logic gap — further
`parser.ts` tuning on the current 37-image set has low remaining ROI. Future
accuracy work should target the OCR/layout pipeline itself (`config/`,
PaddleOCR-VL prompt/model tuning) or expanding the test set to catch different
failure modes, not more regex tweaks.
*Not in scope for either 2.1/2.2 (per next-implementation.md's own note, still true
2026-07-08): batch/lot number extraction doesn't exist in `classify_ocr_server.py`
in either project (only SKU, product name, expiry date are extracted) — if
requested later, follow the same OCR-regex-cascade pattern already used for
expiry-date extraction.*
## 3. Backend — Postgres Data Layer
`pfm-web-app/src/db/`
- **3.1** [DONE] Wrapped the `ocr_items` delete-then-reinsert logic in both `/api/parse` and the `/api/v1/documents/[id]` PUT route inside a single DB transaction (`withTransaction` helper in `db/index.ts`), so a mid-loop failure can no longer leave a document with a correct header but silently missing items — it now rolls back to the previous item set instead. Completed 2026-07-08.
- **3.2** [DONE 2026-07-08] Hashed `accounts.password` with `bcryptjs` (not native `bcrypt` — the `pfm-web-app` Docker stage is `node:20-slim` with no build toolchain, so a native addon would break the image build). `db/init.ts` now hashes the seed password and idempotently migrates any existing plaintext rows (`WHERE password NOT LIKE '$2%'`) on every startup — safe to re-run, no double-hashing. `api/v1/auth/login/route.ts` now fetches by username only and compares with `bcrypt.compareSync`, plus a small added guard for missing username/password (previously would have fallen through to a DB query with `undefined`; now a clean 401). Installed via `docker exec paddleocr-pfm-web-app npm install bcryptjs` (container has its own anonymous `node_modules` volume, not host-synced) + `docker compose restart pfm-web-app`. Verified: `SELECT password FROM accounts` shows a `$2b$10$...` hash, `curl` login with the correct password succeeds (200 + token), wrong password and missing password both return a clean 401 (not a 500). This unblocks task 1.3.
- **3.3** [TODO] Add a unique index on `documents.file_hash` so the dedup check from task 1.1 can do an efficient existing-document lookup before insert instead of a full scan.
## 4. DevOps — Docker & Dev Tunnel
`../docker-compose.yml`, `../docker-compose.demo.yml`, `../start-dev-tunnel.ps1`
- **4.1** [TODO] Decide and document (in root `README.md`/`CLAUDE.md`) a deliberate policy for when the PoC/demo runs `docker-compose.demo.yml` (production build) vs. the dev-mode default — `docker-compose.yml` alone runs `npm run dev`, a known throughput ceiling already flagged in the reliability audit.
- **4.2** [TODO] Add a startup healthcheck/readiness gate for `pipeline-api`/vLLM in `docker-compose.yml` so the Next.js gateway doesn't accept uploads before the GPU pipeline is actually ready to serve them.
- **4.3** [TODO] Extend `start-dev-tunnel.ps1` to verify the ngrok tunnel/LAN IP is actually reachable (not just started) before reporting success, since a stale tunnel is the documented first failure point for login/upload from the Flutter app.
---
*Sections 1-4 migrated 2026-07-08 from root `plans/next-enhancements.md` sections
5-8 (originally populated the same day via a backend-scoped `e`/`enhance` run,
grounded in a prior end-to-end reliability audit and, for §2's Product/SKU scan
sub-feature, `next-implementation.md`'s status table). This file is the sole home
for backend enhancement tasks — the root copy's sections 5-8 were removed outright
from `plans/next-enhancements.md` (not just frozen) per root `AGENTS.md`'s "Scope:
excludes `backend/`" note; the already-shipped 3.1/7.1 DB-transaction fix stays
recorded in root `docs/feature-list.md` regardless, since that's a shipped-feature
log, not a backlog.*
*`next-implementation.md` itself was deleted 2026-07-08 once its full content
(decisions, verified state, action steps, verification checklist) was folded into
§2 above with corrected statuses — it's no longer a separate source of truth.*