Files
pfm-ocr/backend/plans/next-enhancements.md
T
Rafhan Mazaya FathurrahmanandClaude Fable 5 e6daa9b053 docs(plan): per-sale expiry tracking design — batch registry + candidate matching
Grilled 2026-07-16 with the user; full decision record in
docs/expiry-tracking-plan.md. Core reframe: expiry is captured once per
batch at DO intake (staff-typed on the stock-entry confirmation page, from
the physical packs), so the cashier scan only MATCHES OCR fragments against
the 1-3 known in-stock batch dates instead of free-reading damaged
dot-matrix prints (proven model-capability ceiling, 2026-07-15). Fallback:
auto-FEFO + 'inferred' flag, zero cashier interaction. No cloud, ever.

- docs/expiry-tracking-plan.md: architecture, matching algorithm spec
  (resolveExpiryFromEvidence), schema/API deltas, phases 1-3, testing plan
- backend plans §13 (13.1-13.4): matcher util + offline tuning, route
  wiring + expiry_source provenance, multi-frame union, dot-matrix
  recognizer fine-tune
- root plans §10 (10.1-10.3): cashier fast path, inferred badge +
  end-of-day review, burst capture for mounted camera
- stock-feature-plan.md: extension note (batch dropdown becomes the
  manual-override path)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q8TumxFDnyVnfsR3mxPXfX
2026-07-16 17:00:36 +07:00

523 lines
36 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Next Enhancements (backend)
Working backlog driven by the `e`/`enhance` and `n`/`next` triggers defined in
[AGENTS.md](../AGENTS.md) Part B. Split out 2026-07-08 from root
`plans/next-enhancements.md` sections 5-8, renumbered 1-4 — this file now owns all
backend enhancement tracking going forward; the root file only tracks Flutter
sections from this point on. Statuses below were re-verified against the live code
at split time (not copied blind):
- **1.1** `file_hash` dedup — confirmed present, `v1/documents/upload/route.ts:59,98`.
- **1.2** `AbortSignal.timeout` on both hops — confirmed present in `parse/route.ts`
and `v1/documents/upload/route.ts`.
- **1.3** classic routes skipping auth — confirmed still true at split time; later cancelled as out-of-scope (dev-only web UI, no production benefit) — see §1 below, 1.4 shipped instead.
- **3.1** `withTransaction` — confirmed present and used in `db/index.ts`,
`api/parse/route.ts`, `api/v1/documents/[id]/route.ts`.
- **2.1** `pfm-web-app/public/produk-pfm/models/` — corrected 2026-07-08 (was
checked against the wrong path, `backend/models`, in an earlier pass this same
day): the directory does exist. Later the same day, task 2.1 itself was completed
— see §2 below — so the artifacts now exist too.
- **2.2** `m-scan-pfm/page.tsx` — confirmed still missing; desktop `scan-pfm/page.tsx`
confirmed present, **and now confirmed full feature parity** (confidence bar,
top-5 candidate list, expiry-date crop preview all present, 1169 lines) — see §2
below, migrated from `next-implementation.md` (deleted 2026-07-08, content now
lives here).
- **3.2** password hashing — confirmed plaintext at split time (2026-07-08 morning); fixed later the same day, see §3 below.
- **3.3** unique index on `documents.file_hash` — confirmed still absent (only
`filename` has a UNIQUE constraint in `db/init.ts`).
- **4.2** healthcheck for `pipeline-api`/vLLM — confirmed no `healthcheck` block in
root `docker-compose.yml`.
## Format
Tasks are grouped under a numbered section per backend module. Each section gets
exactly 3 tasks:
```
## 1. <Section / Module Name>
- **1.1** [TODO] <clear, specific description of the functional change>
- **1.2** [TODO] <...>
- **1.3** [TODO] <...>
```
When a task is picked up via `n`/`next`, its clarified acceptance criteria (from
`AGENTS.md` Part B §B2a) are appended directly under it. When complete, the status
flips to `[DONE]` and the feature is logged in
[docs/feature-list.md](../docs/feature-list.md) (this dir).
---
## 1. Backend — Next.js API Gateway
`pfm-web-app/src/app/api/`
- **1.1** [DONE] File hash dedup already implemented. (See docs/feature-list.md)
- **1.2** [DONE] AbortSignal timeout already implemented. (See docs/feature-list.md)
- **1.3** [CANCELLED 2026-07-08] Wire JWT auth to classic routes (not needed for dev UI). (See docs/feature-list.md)
- **1.5** [DONE 2026-07-08] Per-account data scoping on `/api/v1/documents/*`. (See docs/feature-list.md)
- **1.6** [DONE 2026-07-08] Lightweight `GET /api/v1/health` endpoint. (See docs/feature-list.md)
- **1.4** [DONE 2026-07-08] Enforced real 401 auth on `/api/v1/documents/*`. (See docs/feature-list.md)
## 2. Backend — OCR Pipeline & Accuracy
`config/`, `pfm-web-app/src/utils/parser.ts`, accuracy regression harness (see `CLAUDE.md`)
**Product/SKU scan sub-feature — decisions & state** (migrated 2026-07-08 from
`next-implementation.md`, which is now deleted; this section is the sole source of
truth for it going forward):
- **Decisions made**: new pages are **standalone routes** (`scan-pfm` — done), following the same self-contained pattern as `manual-label/page.tsx` — own
header/theme, no shared chrome with the root DO-PFM page. Build to **full feature
parity** with the old `ai-ocr-pfm-2026` pages (confidence bars, top-5 SKU
candidates, visual OCR overlay, expiry-date crop preview), not a lean MVP.
- **Scope correction 2026-07-08 (later)**: `m-scan-pfm/page.tsx` (mobile web page)
is **not needed** — user clarified the web `scan-pfm` page is desktop-only,
used for testing the pipeline, not a production mobile surface. Real mobile
scanning is already handled by the Flutter app instead. Focus for this
sub-feature going forward is **backend services** (model artifacts, pipeline
correctness), not any additional frontend. See 2.2 below — cancelled, not
reassigned.
- **Already shipped, re-verified 2026-07-08** (all confirmed present, not assumed):
`config/classify_ocr_server.py` (DINOv2 + YOLO classify, SKU/expiry/name OCR),
`pfm-web-app/public/produk-pfm/index_dinov2.py` + `train_classifier.py`,
`api/scan-pfm/route.ts` (Levenshtein match against `sku_master`),
`api/produk-pfm/route.ts`, DB schema (`sku_master`/`store_master`/`arena_runs` in
`db/init.ts`), `scripts/serve-pipeline.sh` wiring, `Dockerfile`'s `ultralytics`
install in the pipeline-api stage, `nginx.conf` routes for `/scan-pfm`,
`/m-scan-pfm`, `/produk-pfm`.
- **Stale claims corrected 2026-07-08** (next-implementation.md said these didn't
exist; they now do): `pfm-web-app/public/produk-pfm/foto-kemasan-v2/` dataset —
16 SKU subfolders now present (was "❌ does not exist"); `scan-pfm/page.tsx` — now
exists at full feature parity, 1169 lines (was "❌ does not exist").
- **Newly observed, not in original doc's scope**: `/produk-pfm` has an nginx proxy
block and an API route (`api/produk-pfm/route.ts`) but no matching frontend page
(`src/app/produk-pfm/page.tsx` doesn't exist) — dead route, same class of issue as
scan-pfm/m-scan-pfm were. Not turned into a task below since it wasn't part of the
original decision record; flagging for a future `e`/`enhance` pass to pick up.
- **Correction (2026-07-08 audit)**: only *half* dead. `api/produk-pfm/route.ts`
is live — `scan-pfm/page.tsx:159` fetches it for the sample-product gallery.
Only the nginx `location /produk-pfm` *page* proxy block points at a
nonexistent page. Don't remove the API route; see task 4.4.
- **2.1** [DONE 2026-07-08] Built the initial model artifacts (DINOv2 + YOLO) and fixed volume mount issue. (See docs/feature-list.md)
- **2.2** [CANCELLED 2026-07-08] Port `m-scan-pfm/page.tsx` (mobile) from `ai-ocr-pfm-2026` into v2 — not needed, Flutter app handles mobile scanning.
- **2.3** [DONE 2026-07-08] Ran accuracy regression harness, target 95% already met. Fixed one parser logic bug. (See docs/feature-list.md)
- **2.4** [TODO] Human review of `do-008.jpg`'s ground truth in
`sources/manual_labels.json` — flagged in task 2.3 as a suspected labeling error
(its `noPO`/`noSO`/`noDO` OCR cleanly but are entirely different digit sequences
from the labels, not plausible misreads) but never turned into a task. Verify
against the source photo with the client/labeler; if the labels are wrong, fix
them and re-run the harness (aggregate accuracy should tick up).
*Not in scope for either 2.1/2.2 (per next-implementation.md's own note, still true
2026-07-08): batch/lot number extraction doesn't exist in `classify_ocr_server.py`
in either project (only SKU, product name, expiry date are extracted) — if
requested later, follow the same OCR-regex-cascade pattern already used for
expiry-date extraction.*
- **2.5** [DONE 2026-07-14] **Retrain classifier on the now-81-class
dataset.** (Note: a first resume attempt
failed instantly with a Docker daemon connection error — Docker Desktop had
stopped between sessions — before any training happened; restarted Docker
Desktop and relaunched. The successful run was the second attempt,
confirmed running via `docker ps`.) `foto-kemasan-v2/` grew from the 16
classes/118 photos the deployed model
(`produk-pfm-classifier-26n-100e-2026-07-08.pt`) was trained on to **81
classes / 2,493 photos** — the other 65 classes were never included in any
training run.
- **Goal**: retrain both artifacts (`dinov2_index.pkl` similarity index and the
YOLO classifier) against the full current dataset so the deployed model
actually recognizes all 81 SKU folders, not just the original 16.
- **Agreed procedure** (per `docs/scan-product.md`'s documented retraining
steps — training must run via Docker, not bare-metal Windows, since
`paddlepaddle-gpu` wheels are Linux-only): from repo root,
`docker compose build pipeline-api` (bakes in the current dataset) → one-off
`docker run --gpus all` with `models/` mounted **writable** (the live
compose service mounts it `:ro`) → `index_dinov2.py` (rebuilds the DINOv2
index) → `train_classifier.py train --imgsz 224` (its own `split_dataset()`
does an 80/20 split grouped by source photo and shuffled — not a naive
first-N-files split, so augmented copies always land with their source) →
`docker compose restart pipeline-api` → verify via `docker logs` for
"DINOv2 index loaded with N reference images" and "Using classifier
weights: <new dated file>".
- **Final result (2026-07-14)** — both artifacts retrained and live:
- ✅ `dinov2_index.pkl` rebuilt and persisted to disk — "Success! Indexed
2493/2493 images" across all 81 classes (`models/dinov2_index.pkl`,
4.2MB, dated 2026-07-14).
- ✅ **YOLO classifier retrained to completion, 100/100 epochs, real
elapsed time 54m21s** (a first attempt was intentionally stopped by user
request at epoch 43/100 to pause the session; that partial progress was
discarded since `docker run --rm`'s in-container `runs/classify/`
checkpoints aren't bind-mounted, so the successful run below restarted
cleanly from epoch 0 rather than resuming from 43). Final validation:
**85.8% top-1 / 94.4% top-5** across all 81 classes — up from the old
16-class model's 83.3%/90%, now covering 5x the product classes.
Published artifacts: `models/produk-pfm-classifier-26n-100e-2026-07-14.pt`
(3.4MB) and matching `.onnx` (6.3MB, ONNX opset 20, output shape
confirmed `(1, 81)` — i.e. 81 output classes).
- ✅ Verified via `docker compose up -d pipeline-api` (main stack wasn't
running this session) + `docker logs paddleocr-pipeline-api`: "DINOv2
index loaded with 2493 reference images", "Using classifier weights:
/app/pfm-web-app/public/produk-pfm/models/produk-pfm-classifier-26n-100e-2026-07-14.pt",
"YOLO model loaded successfully", "Application startup complete" — the
live service is now serving the new 81-class model, not a code-review
assumption.
- Class-count claims updated in `docs/scan-product.md` and `../CLAUDE.md`
(16 → 81 classes); this task's `docs/feature-list.md` entry added.
## 3. Backend — Postgres Data Layer
`pfm-web-app/src/db/`
- **3.1** [DONE] Wrapped ocr_items updates in DB transactions. (See docs/feature-list.md)
- **3.2** [DONE 2026-07-08] Hashed accounts.password with bcryptjs. (See docs/feature-list.md)
- **3.3** [DONE 2026-07-08] Added an index on `documents.file_hash` to speed up the dedup lookup, and scoped the dedup logic itself by `kode_toko` so stores can't leak duplicates to each other. (See docs/feature-list.md)
## 4. DevOps — Docker & Dev Tunnel
`../docker-compose.yml`, `../docker-compose.demo.yml`, `../start-dev-tunnel.ps1`
- **4.1** [DONE 2026-07-08] Formalized the Docker Compose usage policy in `README.md` and `CLAUDE.md`, making it explicit that the production override must be used for field testing. (See docs/feature-list.md)
- **4.2** [DONE 2026-07-08] Add a startup healthcheck/readiness gate for `pipeline-api`/vLLM in `docker-compose.yml` so the Next.js gateway doesn't accept uploads before the GPU pipeline is actually ready to serve them. (See docs/feature-list.md)
- **4.3** [DONE 2026-07-08] Extend `start-dev-tunnel.ps1` to verify the ngrok tunnel/LAN IP is actually reachable (not just started) before reporting success, since a stale tunnel is the documented first failure point for login/upload from the Flutter app. (See docs/feature-list.md)
- **4.4** [DONE 2026-07-08] Prune/annotate the dead `nginx.conf` location blocks: `/do-pfm` and `/m-do-pfm` (dead since v2 consolidated the DO-PFM UI into the root page — already documented as dead in root `CLAUDE.md`), the `/produk-pfm` page proxy (no `src/app/produk-pfm/page.tsx` exists — but keep `api/produk-pfm/route.ts`, it's live, used by `scan-pfm/page.tsx:159`; see the corrected §2 note), and `/m-scan-pfm` (page intentionally unbuilt per cancelled task 2.2 — removing the proxy block doesn't resurrect that task, it just stops nginx advertising a 404). (See docs/feature-list.md)
- **4.5** [DONE 2026-07-08] **Lock down the publicly tunneled surface**. `start-dev-tunnel.ps1` tunnels **all of `localhost:8000`** (the whole nginx gateway) to a stable, reserved ngrok domain — which makes the deliberately unauthenticated classic routes (`/api/upload`, `/api/parse`, `/api/history`, ...) and the dev pages (root DO-PFM, `/scan-pfm`, `/manual-label`) internet-reachable. Task 1.3's "classic routes never get auth" decision was made in an on-prem/LAN context and stands — so restrict at the edge instead of adding auth: either (a) tunnel a separate nginx server block/port that proxies **only `/api/v1/*`** (the authenticated production surface the Flutter app actually uses — see `app_config.dart`, both endpoints end in `/api/v1`), or (b) an ngrok traffic policy (IP allowlist / basic auth) covering everything except `/api/v1/*`. Decide (a) vs (b) with the user at pickup; (a) is self-contained in `nginx.conf` + the script and doesn't depend on ngrok plan features. Note `GET /api/v1/health` (task 1.6) must remain reachable through whichever restriction ships — it's the fallback probe target. (See docs/feature-list.md)
## 5. Docs & Workflow Integrity
`CLAUDE.md`, `SKILLS.md`, `AGENTS.md`, this file — added 2026-07-08 after an audit
cross-checking the kit docs' claims against live code. All *code-level* claims in
this file's §1-4 re-verified accurate (auth guards, bcrypt, `withTransaction`,
parser-fallback fix, dedup, timeouts, absent index/healthcheck all match the code);
the drift found is in the guidance docs the SKILLS.md roles rely on.
- **5.1** [DONE 2026-07-08] Fixed stale doc claims in SKILLS.md and CLAUDE.md. (See docs/feature-list.md)
- **5.2** [DONE 2026-07-08] Shrunk [DONE] tasks in plans/next-enhancements.md. (See docs/feature-list.md)
- **5.3** [DONE 2026-07-08] Amended AGENTS.md completion checklist with doc-sync step. (See docs/feature-list.md)
## 6. Product Scan — Ground Truth Annotation & Accuracy
`scan-pfm/page.tsx`, `api/manual-label-scan/route.ts`, `sources/product_manual_labels.json`
Added 2026-07-08 via a user-directed `e` run: build an **editable ground-truth
surface for product scans**, mirroring what the DO flow already has in
`manual-label/page.tsx` + `api/manual-label/route.ts` + `sources/manual_labels.json`.
**Current-state audit (verified against code, not assumed):**
- What exists: `api/manual-label-scan/route.ts` (GET by exact filename / POST
upsert, schema `{filename, no_sku, nama_item, expiry_date, top1_confidence,
notes, saved_at}`, own `normalizeDateString`), a "Save Ground Truth" button in
`scan-pfm/page.tsx` (`handleSaveGroundTruth`, line 333), and the
`./backend/sources:/sources` mount that makes saves host-visible (fixed in 2.1).
- **Gap (a) — prediction saved as truth**: in the scan-pfm save, only `no_sku` is
correctable (the top-5 "Use" override → `editedSku`); `nama_item` is hardcoded to
the classifier's `top1_name` and `expiry_date` to the OCR output, `notes` always
`""`. Ground truth that *is* the model's own prediction makes any future accuracy
measurement against it trivially inflated — this is the core defect the user's
request fixes.
- **Gap (b) — phantom image keys**: uploaded (non-gallery) photos are keyed
`uploaded-${Date.now()}.jpg` and the image bytes are never persisted anywhere
(data-URL in the browser only) — those label entries reference files that don't
exist on disk, unusable for re-evaluation.
- **Gap (c) — no browse/edit**: GET requires an exact filename; there's no list
endpoint, no page to review/correct/delete existing entries (the DO side has a
full annotation page for this).
- **Gap (d) — no consumer**: nothing plays the role `accuracy-check.mts` plays for
the DO parser; the labels currently gate nothing.
- **6.1** [DONE 2026-07-08] Built standalone annotation page manual-label-scan/page.tsx. (See docs/feature-list.md)
## 8. Master Data Management
`api/v1/master/stores/route.ts`, `api/v1/master/skus/route.ts`, `admin/master-data/page.tsx`
- **8.1** [DONE 2026-07-08] Build full CRUD API routes for `store_master` and `sku_master` to allow dynamic updating of reference data. (See docs/feature-list.md)
- **8.2** [DONE 2026-07-08] Implement auto-provisioning of login accounts with securely hashed passwords whenever a new store is created via the API. (See docs/feature-list.md)
- **8.3** [DONE 2026-07-08] Build an `/admin/master-data` web UI to visually manage both SKUs and Stores. (See docs/feature-list.md)
- **6.2** [DONE 2026-07-08] API + storage groundwork for scan annotation. (See docs/feature-list.md)
- **6.3** [DONE 2026-07-08] Product-scan accuracy harness created. (See docs/feature-list.md)
- **6.4** [DONE 2026-07-13] Auto-diff-vs-previous-run reporting (ported from the
DO-flow's `accuracy-check.mts`) plus classifier method/confidence tracking
added to `accuracy-check-scan.mts`; created the missing
`sources/product-test-images/` validation-photo folder. User-directed `n`
request: wanted to tune the scan algorithm and see improvement/regression
automatically instead of eyeballing two flat runs. (See docs/feature-list.md)
*Suggested order: 6.2 → 6.1 → 6.3 → 6.4 (storage/API first, page on top, harness
once labels exist in volume, diffing once the harness has history to diff against).*
## 7. Auth — Store Accounts & Profile-Sourced Metadata
`pfm-web-app/src/db/init.ts`, `api/v1/auth/*`, `api/parse/route.ts`, `sources/toko_aktif.json`
Added 2026-07-08 via a user-directed `e` run: one login account per store
(username = `kode_toko`, default password `"password"`), each carrying its store
profile (nama toko, kode toko, alamat), with the profile's store name + address
used as the parse response's values instead of OCR.
**Current-state audit (verified against code and the live DB, not assumed):**
- **The "don't OCR store/alamat" half is already live.** `v1/documents/upload`
passes `account?.kodeToko` into `/api/parse` (upload `route.ts:125`), and parse
stamps `metadata.orderUntuk`/`alamat` straight from `store_master` when a
`kodeToko` is present — OCR-based store matching (`resolveStoreFromText`) is
only the no-account fallback (`parse/route.ts:392-410`). The Flutter app reads
these same metadata fields from `GET /documents`, so **no response-shape or
Flutter change is needed** for that half.
- **What's missing is the accounts**: the live DB has **779 `store_master` rows
but exactly 1 account** (`admin`), so every real upload today authenticates as
`admin`/`WH_JOFFICE` and the per-store path never fires for actual stores.
- **Bootstrap gap**: no code anywhere populates `store_master` — the 779 rows
exist only in the live DB volume (imported out-of-band). `sources/toko_aktif.json`
(`{namaToko, kodeToko, alamat}`, same 3 fields) is the obvious source. On a
truly fresh DB, `init.ts` would even fail its own `admin` seed — the
`accounts.kode_toko → store_master(kode_toko)` FK can't resolve `WH_JOFFICE`
when `store_master` is empty.
- **7.1** [DONE 2026-07-08] Seeded one account per store in `db/init.ts`. (See docs/feature-list.md)
- **7.2** [DONE 2026-07-08] Returned the store profile at login and added `/me` endpoint. (See docs/feature-list.md)
- **7.3** [DONE 2026-07-08] Built reproducible `store_master` bootstrap in `db/init.ts`. (See docs/feature-list.md)
*Suggested order: 7.3 → 7.1 → 7.2 (bootstrap first — account seeding FK-depends
on it; login profile last, it's additive).*
## 9. Flutter Client Contract — v1 Surface Completion
`api/v1/documents/`, `api/v1/master/skus/`, `api/parse/route.ts`, `db/init.ts`
Added 2026-07-10 via a user-directed, explicitly backend-scoped `e` run auditing
the full Flutter↔backend request/response contract. **Context doc:
[../../docs/api-contract-map.md](../../docs/api-contract-map.md)** (repo-root
`docs/`) — endpoint inventory, envelopes, lifecycle, and gap IDs (G1-G10) cited
below. These are the *server* halves; the Flutter halves are root
`plans/next-enhancements.md` §6-7 and consume these, so this section ships first.
Keep the v1 envelope (`{status, data}` / `api-error.ts`) on everything new.
- **9.1** [DONE 2026-07-10] `GET /api/v1/documents/:id` with `parseStatus`/`docType`, `scan_mode` persistence, dedup-stub fix. (See docs/feature-list.md)
- **9.2** [DONE 2026-07-10] Relaxed `GET /api/v1/master/skus` to any authenticated
account (writes stay admin-only); no response-shape change. Picked up via
explicit `n{9.2}` request; user chose "relax existing endpoint" over "add a
new one" when asked. (See docs/feature-list.md)
- **9.3** [DONE 2026-07-10] Authenticated `POST /api/v1/scan-product`, shared classify+match util. (See docs/feature-list.md)
*Section 9 is now fully `[DONE]`. With 9.1-9.3 all shipped, every backend
blocker behind Flutter root `plans/next-enhancements.md` §7.1 (moving the
product editor onto the v1 surface) is cleared. §7.2 (eliminating the
duplicate classification pass) is unblocked in principle but still needs its
own grill-me decision on the client side about which single pass to keep.*
## 10. Backend — Document Confirmation Gate & Data Hygiene
`pfm-web-app/src/db/init.ts`, `app/api/v1/documents/`, `app/api/parse/route.ts`, `utils/document-mapper.ts`
Added 2026-07-10 from user testing feedback on the release APK
([`twinkly-riding-mitten.md`](C:/Users/rafha/.claude/plans/twinkly-riding-mitten.md)).
Backend counterpart to Flutter root `plans/next-enhancements.md` §8.
Root-cause documentation in [../../docs/api-contract-map.md](../../docs/api-contract-map.md)
**G11** (no draft/confirmed distinction) and **G12** (fabricated PO/SO/DO).
**Ship §10 before Flutter §8.2** — Flutter's model change consumes the new
`confirmed` field this section adds. Keep the v1 envelope (`{status, data}` /
`api-error.ts`) on everything new.
- **10.1** [DONE 2026-07-10] **Confirmation-gated document list visibility (Task B backend
half).** Three coordinated changes, one migration:
1. `db/init.ts` — `ALTER TABLE documents ADD COLUMN IF NOT EXISTS confirmed
BOOLEAN NOT NULL DEFAULT true;` (`DEFAULT true` grandfathers every
pre-existing row — today's history stays visible after migration).
2. `v1/documents/upload/route.ts` — add `confirmed = false` to the INSERT
column list for every new upload (dedup-hit branch unchanged — reflects
whatever `confirmed` state the original row already has).
3. `utils/document-mapper.ts` — add `confirmed: boolean` to `DocumentRow`
interface; return `confirmed: doc.confirmed` from `mapDocumentRow()`;
update the three SELECT statements that build a `DocumentRow` (list route,
`[id]` GET, upload dedup-hit SELECT) to include the `confirmed` column.
4. `v1/documents/route.ts` (list) — add `AND confirmed = true` to WHERE
clause, unconditionally for every account including admin (per clarified
answer — existing `kode_toko` scoping for non-admins is untouched).
5. `v1/documents/[id]/route.ts` — PUT handler: add `confirmed = true` to the
UPDATE SET. GET handler: no filter change (poller must keep seeing
pending/unconfirmed docs); just receives `confirmed` via mapper update.
- **Note on `parse/route.ts`'s own INSERT...ON CONFLICT statements (both DO
and Product branches)**: deliberately left untouched for `confirmed` —
in the real mobile flow, `upload/route.ts`'s INSERT always runs first
(explicit `confirmed = false`), so `parse/route.ts`'s upsert always hits
the `ON CONFLICT DO UPDATE` branch; since that branch's `SET` clause
doesn't mention `confirmed`, Postgres leaves the existing value untouched
— exactly the desired behavior (never regress an already-confirmed
document, never reset the pending flag mid-parse). Omission was verified
to be correct, not an oversight.
- Live verification (all 5 steps passed against the running Docker stack):
(a) uploaded a real DO photo as store `WH_JCIBBR1`, did not PUT — absent
from that store's `GET /documents` (count stayed at 3, new id 3400 not
present) while `GET /documents/3400` still returned `parseStatus: "done"`,
`confirmed: false`; (b) PUT (confirm) — doc count became 4, id 3400 present
with the real submitted `namaPenerima`, DB row's `confirmed` flipped to
`true`; (c) covered by (a)'s single-doc check; (d) `admin`'s `GET
/documents` also excluded the unconfirmed doc (13, unchanged) before the
PUT; (e) all 13 pre-existing rows carried `confirmed = true` after the
migration ran (`ALTER TABLE` executed on container restart, verified via
`\d documents` + a `count(*)` query — 13 confirmed, 0 unconfirmed
pre-restart).
- **10.2** [DONE 2026-07-10] **Remove fabricated Product Scan PO/SO/DO placeholders (Task
C, bundled with 10.1 — same file `parse/route.ts`).** Replaced hardcoded
`noPO: "PO-PRODUCT-001"`, `noSO: "1002003004"`, `noDO: "DO-PRODUCT-999"` —
both the flat keys and the mirrored `header.no_po`/`no_so`/`no_do`
sub-object — with empty strings `""`. Scope stayed narrow to exactly these
three fields; `nama_driver: "PRODUCT SCAN"` and `nama_penerima: "STORE STAFF"`
were left untouched (deliberate fixed convention, not a fabricated document
number that could mislead someone reading raw data — different failure mode
from G7's fake SKU/date data). Not user-facing: verified `pdf_service.dart`
and `product_editor_submit_logic.dart` neither reads nor displays these values.
- Verification: uploaded a fresh Product Scan as `WH_JCIBBR1` (doc id 3402),
read the raw unconfirmed `GET /documents/3402` response — `header.no_po`/
`no_so`/`no_do` all returned `""`, not the old fabricated strings.
## 11. Backend — Single-Pass Product Classification
`app/api/parse/route.ts`, `utils/product-scan.ts`, `utils/document-mapper.ts`
Added 2026-07-10 from user feedback ("kenapa harus dilakukan dua kali... GPU
tidak 2x kerja") after noticing Product Scan's editor took visibly longer to
open than DO Scan's. Root-caused as gap **G3** in
[../../docs/api-contract-map.md](../../docs/api-contract-map.md) (deferred
there pending exactly this client-side decision). Backend counterpart to
Flutter root `plans/next-enhancements.md` §7.2.
- **11.1** [DONE 2026-07-10] **Run the classify+match pipeline once per
photo, and persist the full result.** `api/parse/route.ts`'s Product branch
had its own separate, poorer inline `fetch` to the classifier that only
kept `top1_name`/`extracted_sku` — the Flutter editor then had to re-run
the *entire* pipeline a second time via `POST /api/v1/scan-product` (task
9.3) just to get the top-5 candidate list and OCR-extracted expiry date.
Replaced the inline fetch with a call to the same shared
`classifyAndMatchProduct()` (`utils/product-scan.ts`) that route already
uses — one GPU call, richer result, `b64` reused from the DO path's own
computation (not recomputed). The richer result is persisted under a new
`metadata.productScan` JSONB key (`possibleMatches` + `extractedExpiryDate`
— no schema migration needed, same pattern as `header`/`shipment`
coexisting in that column) and surfaced by `document-mapper.ts` as a
top-level `productScan` field on every GET response (list, by-id, upload
dedup). `ocr_items` still stores only the single best-match row, unchanged.
- **Regression caught and fixed during implementation**: delegating to
`classifyAndMatchProduct()` silently dropped the 90s
`AbortSignal.timeout` the old inline fetch had (a wedged GPU container
would otherwise hang past the intended fail-fast bound). Added the same
`PIPELINE_TIMEOUT_MS = 90_000` bound directly inside
`classifyAndMatchProduct()` itself, so both callers (`parse/route.ts` and
the live `POST /api/v1/scan-product` route, which never had this bound
either) are protected, not just the one this task touched.
- Verification: uploaded a genuinely fresh image/store combination
(`do-015.jpg` as `WH_JAFATAH`, never uploaded before, so this is
provably a real classify pass and not a dedup hit) — took 9s (one GPU
pass), response was `"Document uploaded successfully"` (not the dedup
branch). Immediate `GET /documents/:id` returned `productScan` with 5
real `possibleMatches` (real SKUs/names/scores from `sku_master`) and
the OCR-extracted expiry date — before any editor interaction. See
root `docs/iteration-log.md` for the Flutter-side verification that the
editor renders this without a second network call.
## 12. Backend — Stock Management
`src/db/init-stock.ts`, `src/app/api/v1/stock/`, `src/utils/stock-*.ts`,
`src/app/api/parse/route.ts`, `src/app/api/v1/documents/[id]/route.ts`,
`src/app/admin/master-data/`
Added 2026-07-10, backend counterpart to root `plans/next-enhancements.md` §9
(Flutter Stocks Menu & DO-to-Stock Flow) — both sections originated from the same
user-directed, extensively grilled ad-hoc feature request (not an `e`/`enhance`
section — see `AGENTS.md` Part B7). **Read
[../../docs/stock-feature-plan.md](../../docs/stock-feature-plan.md) first** — full
schema, API contracts, and sequencing for both sides. **Status: planned, not yet
implemented** — no code for this feature exists in the codebase yet.
- **12.1** [TODO] **Stock schema + core CRUD.** New `src/db/init-stock.ts`
(`stock_batches` — unique per `(kode_toko, no_sku, batch_code, expiry_date)`,
tracks both outer and inner qty; `stock_movements` — append-only audit log,
`intake`/`decrement`/`adjustment`/`manual_seed`), wired into `init.ts`. New
`src/utils/stock-mapper.ts`, `src/utils/stock-movement.ts` (`recordStockMovement`
only, for this task). New routes: `src/app/api/v1/stock/route.ts` (GET summary
per SKU, POST create/merge-by-unique-key), `stock/[noSku]/route.ts` (GET batch
detail), `stock/batches/[id]/route.ts` (PUT edit, logged as `adjustment`). No
DELETE route — batches are edit-only, never removed.
- **12.2** [TODO] **Product Scan decrement hook + in-stock candidate filter.**
Extend `stock-movement.ts` with `decrementBatchForProductScan` (row-locked,
allowed to go negative, logged as `decrement`); wire into
`documents/[id]/route.ts`'s existing PUT transaction, gated on `scan_mode ===
'Product'` and a new `stock_batch_id` payload field — an invalid/cross-store
batch id fails the **whole confirm** (400), never a silent skip (per user's
explicit answer during grilling). New `src/utils/stock-lookup.ts`
(`getInStockSkuSet`/`filterMatchesByStock`), applied to `possibleMatches` in
both `api/parse/route.ts`'s Product branch and `api/v1/scan-product/route.ts` —
**do not change `classifyAndMatchProduct()`'s signature**, it's shared with the
anonymous store-agnostic desktop dev route; filter at the two authenticated call
sites instead. Blocked on 12.1.
- **12.3** [TODO] **Read-only admin Stock view.** `admin/master-data/page.tsx`
(419 lines, already over the 256-line threshold) split into `page.tsx` (shell)
+ extracted `StoreManager.tsx` + `SkuManager.tsx` (pure extraction, no behavior
change) + new `StockManager.tsx` (all-stores table via `GET /api/v1/stock` with
no `kode_toko` param as admin; row click drills into batch detail). Blocked on
12.1.
*Suggested order: 12.1 → 12.2 (needs 12.1's tables/movement helper) and 12.3
(needs 12.1's summary route) — 12.2/12.3 are independent of each other.*
## 13. Backend — Per-Sale Expiry Resolution (Candidate Matching)
`src/utils/expiry-matcher.ts`, `src/app/api/parse/route.ts`,
`src/app/api/v1/scan-product/route.ts`, `src/app/api/v1/documents/`,
`config/classify_ocr_server.py`
Added 2026-07-16 from a user-directed grilling session (ad-hoc feature per
`AGENTS.md` Part B7, like §12). **Read
[../../docs/expiry-tracking-plan.md](../../docs/expiry-tracking-plan.md) first**
— full design, confirmed decisions (full automation at cashier, no cloud ever,
auto-FEFO + `inferred` flag as the only fallback), matching algorithm spec, and
phase plan. Core idea: expiry is captured once per batch at DO intake
(staff-typed on the stock-entry page, §12/root §9), so the cashier scan only has
to **match** OCR fragments against 1–3 known candidate dates — never free-read a
damaged dot-matrix print under time pressure. **Blocked on 12.1 + 12.2** (needs
`stock_batches` + the in-stock candidate filter). Flutter counterpart: root
`plans/next-enhancements.md` §10.
- **13.1** [TODO] **`resolveExpiryFromEvidence()` matcher util + offline tuning
harness.** Pure TS util implementing the 4-stage resolution
(`single_batch` → `matched_exact` → `matched_fragment` → `inferred_fefo`)
with digit-confusion-aware fuzzy scoring of candidate date print-forms (and
batch codes) against the scan's OCR `text_lines`; thresholds
(`SCORE_MIN`/`MARGIN_MIN`) tuned offline by replaying the 79 frozen-benchmark
line-sets against synthetic candidate sets built from ground-truth labels —
tune for **zero wrong-candidate picks** (flagged FEFO beats a confident wrong
match). Includes verifying the classify server response actually carries
`text_lines` to the gateway (add to payload if not — small
`classify_ocr_server.py` change). Unit tests + harness script committed.
- **13.2** [TODO] **Wire resolution into both scan routes + persist
provenance.** `documents.expiry_source VARCHAR(20)` CHECK
(`single_batch|matched_exact|matched_fragment|inferred_fefo|manual`) +
`expiry_match_score REAL NULL`; `parse/route.ts` Product branch and
`v1/scan-product/route.ts` call the matcher after the §12.2 in-stock filter
and include `resolvedBatch`/`source`/`score` in `metadata.productScan` and
the response (no second GPU call — same pattern as 11.1); PUT persists
`expiry_source` (client override ⇒ `'manual'`); documents list gains an
`?expiry_source=` filter for the end-of-day review of `inferred_fefo` sales.
Blocked on 13.1.
- **13.3** [TODO] **Phase 2 — multi-frame evidence union.** Accept burst
uploads (N frames per scan) on the product-scan path; classify on the best
frame, union OCR `text_lines` across all frames before matching (glare moves
between frames — fragments accumulate). Pairs with a mounted camera at the
cashier (root §10.3). Blocked on 13.2.
- **13.4** [TODO] **Phase 3 — on-prem dot-matrix recognizer.** Synthetic
dot-matrix/inkjet date-crop generator (dot dropout, scratch, fade, curvature,
glare augmentation) + real-data flywheel (harvest scan crops weakly labeled
by the batch registry's staff-typed expiry values); fine-tune a small rec
model on the RTX 2060; deploy as an additional reader in
`classify_ocr_server.py` feeding the same matcher. Goal: shrink the
`inferred_fefo` residue. **No cloud — hard constraint.** Blocked on 13.2;
independent of 13.3.
*Suggested order: 13.1 → 13.2 → (13.3 and/or 13.4 as needed once Phase-1
accuracy is measured in the field).*
---
*Sections 1-4 migrated 2026-07-08 from root `plans/next-enhancements.md` sections
5-8 (originally populated the same day via a backend-scoped `e`/`enhance` run,
grounded in a prior end-to-end reliability audit and, for §2's Product/SKU scan
sub-feature, `next-implementation.md`'s status table). This file is the sole home
for backend enhancement tasks — the root copy's sections 5-8 were removed outright
from `plans/next-enhancements.md` (not just frozen) per root `AGENTS.md`'s "Scope:
excludes `backend/`" note; the already-shipped 3.1/7.1 DB-transaction fix stays
recorded in root `docs/feature-list.md` regardless, since that's a shipped-feature
log, not a backlog.*
*`next-implementation.md` itself was deleted 2026-07-08 once its full content
(decisions, verified state, action steps, verification checklist) was folded into
§2 above with corrected statuses — it's no longer a separate source of truth.*