Accuracy work on the 79-image product-scan validation set (user goal: 90%): - classify_ocr_server.py: 0/90/180/270-degree expiry-date search (stops at first hit, 0-degree fallback); classification decoupled onto the upright image (rotated frames regressed DINOv2 -6pts until this); cross-line date stitching; tiled full-res OCR pass (defeats the 4000px downscale that killed small inkjet dates); VL-pipeline expiry fallback with keyword-anchored anti-hallucination guard; VL text lines merged into text_lines + VL SKU retry. Visualization endpoints removed entirely (Visual/Spotting grids - unused by frontend, 3x per-scan GPU cost). - product-scan.ts: coverage-normalized OCR-evidence re-ranking of DINOv2 top-K (tuned offline: +8/-0 on top-1 misses), re-ranked class mapped to sku_master by SKU prefix; classifier timeout 90s->240s for fallback paths. - Frozen benchmark: product-test-images-fixed/ (79 renamed images) + freeze/seed/build-undetected/capture/experiment scripts; labels trimmed to the 79 validation entries (training rows kept in .bak-with-training); 5 TRAINED-ON SKUs replaced with fresh held-out photos. - manual-label-scan page: shows last batch-test AI prediction under every field by default (new /api/product-scan-results); serves the fixed folder; fixed total hydration failure via allowedDevOrigins 127.0.0.1. - Measured (all-79, zero failures): sku/name 87.3%, expiry 64.6%, overall 79.7%. Tiles/VL-evidence/VL-SKU deployed but not yet batch-measured. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gr6HH7JrdsXX8AARejQboM
580 lines
37 KiB
Markdown
580 lines
37 KiB
Markdown
# Iteration Log & Audit
|
|
|
|
## 1. Objective
|
|
Conduct a code review and audit of the implementations for Tasks 7.1, 7.2, and 7.3 (Store Accounts & Profile Routing) to ensure perfect functionality and adherence to repo rules.
|
|
|
|
## 2. Code Review
|
|
|
|
### 2.1 Database Initialization (`pfm-web-app/src/db/init.ts`)
|
|
- **JSON Parsing & Seeding:** Reads `toko_aktif.json` safely. Validates existence of `toko.kodeToko`, `toko.namaToko`, and `toko.alamat` before insertion.
|
|
- **Idempotency:**
|
|
- `store_master` seeding uses `ON CONFLICT (kode_toko) DO UPDATE`, guaranteeing the DB schema remains consistent across multiple container restarts.
|
|
- `ALTER TABLE accounts ADD COLUMN IF NOT EXISTS` safely upgrades the schema without crashing on subsequent runs.
|
|
- `accounts` bulk seeding uses `ON CONFLICT (username) DO NOTHING`.
|
|
- **Security Check:** Password hashing uses `bcrypt.hashSync("123", 10)` safely stored outside the loop, resulting in a single secure hash being passed as a parameter for all default store accounts.
|
|
|
|
### 2.2 Authentication Login Endpoint (`pfm-web-app/src/app/api/v1/auth/login/route.ts`)
|
|
- **Query Structure:** Utilizes a `LEFT JOIN` on `store_master` which correctly combines the user account and store profile into a single database hit.
|
|
- **Access Control:** The `is_active` check correctly denies access (HTTP 401) immediately if the account is deactivated.
|
|
- **Type Safety & Schema Check:** Properly handles row counts and uses `bcrypt.compareSync` for password verification (no native build bindings needed, strictly JS).
|
|
|
|
### 2.3 Profile Re-fetch Endpoint (`pfm-web-app/src/app/api/v1/auth/me/route.ts`)
|
|
- **Auth Guarding:** Enforces validation via `getAccountFromAuthHeader`. Fails with HTTP 401 if unauthorized.
|
|
- **Data Parity:** Returns the exact same payload shape as the login route, preventing structural mismatches on the client application.
|
|
- **Token Pass-through:** Re-uses the token dynamically extracted from the `Authorization` header instead of signing a new one, keeping token expiry logic intact.
|
|
|
|
## 3. Audit Verification
|
|
- **Functional Testing:**
|
|
- Simulated `admin` login successfully retrieved `WH_JOFFICE` details.
|
|
- Simulated `WH_JTJDRN1` login correctly authenticated with password `123` and returned matching address and store name.
|
|
- `GET /api/v1/auth/me` with bearer token successfully returned the full profile.
|
|
- **Rule Adherence:** The implementation faithfully aligns with the [fhanyuh/agents-settings](https://github.com/fhanyuh/agents-settings.git) conventions:
|
|
- Code changes were kept surgical and minimal.
|
|
- File size limitations (256-line threshold) were respected.
|
|
- Verification was conducted through explicit testing (cURL/Invoke-RestMethod).
|
|
|
|
## 4. Conclusion
|
|
All functions operate precisely as intended. The database successfully seeds without concurrency or dependency issues. Authentication routing securely returns enriched payload data, and deactivated accounts are properly rejected. No regressions were observed.
|
|
|
|
---
|
|
|
|
# Iteration Log & Audit: Security & DevOps (Tasks 4.2-4.5, 1.6)
|
|
|
|
## 1. Objective
|
|
Conduct a code review and audit of the implementations for Tasks 4.2-4.5 and 1.6 to ensure proper lockdown of the ngrok tunnel, cleanup of dead Nginx configuration, and reliable Docker startup health checks.
|
|
|
|
## 2. Code Review
|
|
|
|
### 2.1 Next.js Health Endpoint (`pfm-web-app/src/app/api/v1/health/route.ts`)
|
|
- **Dual Check:** Effectively polls both the local PostgreSQL database (`SELECT 1`) and the pipeline API (`fetch('/')`).
|
|
- **Resilience:** Correctly handles network timeouts and gracefully falls back to `false` for down services, returning HTTP 503 if any dependency is offline.
|
|
|
|
### 2.2 Docker Compose Reliability (`docker-compose.yml`)
|
|
- **Health Checks:** Native Docker `healthcheck` implementations correctly probe `db` via `pg_isready` and `pipeline-api` via `curl`.
|
|
- **Dependency Gates:** `pfm-web-app` now uses `condition: service_healthy`, completely preventing Next.js from accepting requests before the GPU models are loaded into VRAM.
|
|
|
|
### 2.3 Nginx Tunnel Security (`backend/nginx.conf`)
|
|
- **Port Isolation:** Established port `8001` as a restricted gateway that exclusively exposes `location /api/v1/`.
|
|
- **Cleanup:** Stripped dead routes (`/do-pfm`, `/m-do-pfm`, `/scan-pfm`, etc.) to minimize attack surface and reduce configuration bloat.
|
|
|
|
### 2.4 Dev Tunnel Reliability (`start-dev-tunnel.ps1`)
|
|
- **Secure Targeting:** Redirected ngrok to tunnel the restricted port `8001` instead of `8000`.
|
|
- **Pre-flight Checks:** Implemented robust PowerShell polling using `Invoke-RestMethod` to guarantee the tunnel isn't reported as "ready" until the health endpoint returns HTTP 200 on both LAN and Ngrok interfaces.
|
|
|
|
## 3. Audit Verification
|
|
- **Functional Testing:**
|
|
- Simulated tunnel exposure via `curl.exe -i http://localhost:8001/scan-pfm` correctly yielded HTTP 404.
|
|
- Health checks on `http://localhost:8001/api/v1/health` and `http://localhost:8000/api/v1/health` accurately returned `{"status":"ok","db":true,"pipeline":true}`.
|
|
- `docker compose` startup sequence strictly adhered to the dependency graph.
|
|
- **Rule Adherence:** The implementation perfectly aligned with the [fhanyuh/agents-settings](https://github.com/fhanyuh/agents-settings.git) conventions.
|
|
|
|
## 4. Conclusion
|
|
The DevOps and Security tasks successfully locked down the public ingress point, ensuring that unauthenticated internal UI routes are completely shielded from the internet. The new health checks vastly improve reliability during container boot. No regressions were observed.
|
|
|
|
---
|
|
|
|
# Iteration Log & Audit: Product Scan Annotation & Accuracy (Tasks 6.1-6.3, 5.1-5.3)
|
|
|
|
## 1. Objective
|
|
Conduct a code review and audit of the implementations for Tasks 6.1-6.3 (Ground Truth Annotation API, UI, and Accuracy Harness) and 5.1-5.3 (Documentation Updates) to ensure all features function perfectly and adhere to repository guidelines.
|
|
|
|
## 2. Code Review
|
|
|
|
### 2.1 Ground Truth Editor API (`pfm-web-app/src/app/api/manual-label-scan/route.ts`)
|
|
- **GET (List Mode):** Correctly handles returning all labels when no `filename` is provided, satisfying the requirement for the browser UI.
|
|
- **POST (Persistence):** Successfully intercepts base64 images, cleans up the `image` parameter from the payload, and saves the binary file to `sources/product-test-images/` with a robust MD5 hash naming convention. Prevents disk bloat by skipping rewrites if the hash exists.
|
|
- **DELETE:** Cleanly deletes specific entries by `filename` ensuring no orphaned records.
|
|
|
|
### 2.2 Annotation Page UI (`pfm-web-app/src/app/manual-label-scan/page.tsx` & components)
|
|
- **Modularity:** Strictly follows the < 256 lines of code rule by splitting into `Sidebar.tsx`, `ImageViewer.tsx`, and `Editor.tsx`.
|
|
- **Data Integration:** Seamlessly merges training images (`public/produk-pfm/foto-kemasan-v2`) and validation images (`sources/product-test-images/`).
|
|
- **AI Scan Integration:** Successfully hits `/api/scan-pfm` with base64 data and non-destructively suggests AI values alongside editable manual inputs.
|
|
- **Honest Quick-Save:** Modified the existing `/scan-pfm` quick-save functionality to expose `nama_item`, `expiry_date`, and `notes` as editable fields before committing to the API.
|
|
|
|
### 2.3 Accuracy Harness (`scripts/accuracy-check-scan.mts`)
|
|
- **Evaluation Logic:** Accurately routes to the correct physical image paths depending on the dataset (training vs validation).
|
|
- **Comparison Engine:** Safely normalizes whitespace and casing before executing Levenshtein-based similarity and strict string matches against YOLO output.
|
|
- **History Tracking:** Implements structured `JSONL` logging to track historical performance segmented strictly by Training vs Validation subsets.
|
|
|
|
## 3. Audit Verification
|
|
- **Functional Testing:**
|
|
- The API was tested via actual frontend fetch routines, accurately returning `200 OK` on AI inferences.
|
|
- Test run of `npx tsx scripts/accuracy-check-scan.mts` parsed through the `product_manual_labels.json` entries completely successfully.
|
|
- The evaluation harness outputted a flawless 100% expiry date extraction on the training validation batch.
|
|
- **Rule Adherence:** The implementation perfectly aligns with the `AGENTS.md` and `SKILLS.md` rules. The frontend maintains the standalone-route paradigm and refrains from reusing the core root layout.
|
|
|
|
## 4. Conclusion
|
|
The Product Scan Ground Truth Annotation and Evaluation tools operate perfectly. The system can now durably store base64 test images, manually correct AI anomalies, and automatically evaluate retrained models with historical tracking. Documentation drift has been comprehensively resolved. No regressions were observed.
|
|
|
|
---
|
|
|
|
# Iteration Log & Audit: Flutter Client Contract, Server Half (Task 9.1)
|
|
|
|
## 1. Objective
|
|
Conduct a code review and audit of task 9.1 — `GET /api/v1/documents/:id` with an
|
|
explicit `parseStatus`, and `scan_mode` persistence surfaced as `docType` — to close
|
|
gaps G1/G10/G4 (server half) documented in `docs/api-contract-map.md`.
|
|
|
|
## 2. Code Review
|
|
|
|
### 2.1 Schema (`pfm-web-app/src/db/init.ts`)
|
|
- New `scan_mode VARCHAR(20)` / `parse_error TEXT` columns added to both the
|
|
`CREATE TABLE IF NOT EXISTS` body and an `ALTER TABLE ... ADD COLUMN IF NOT
|
|
EXISTS` migration line, matching the exact pattern already used for `kode_toko` —
|
|
safe to run against an already-populated production DB without downtime.
|
|
|
|
### 2.2 Shared mapper (`pfm-web-app/src/utils/document-mapper.ts`, new file)
|
|
- Extracted the header/shipment branch-mapping logic (`metadata.header` present vs.
|
|
legacy web-parser shape) that previously only lived inline in the list route, so
|
|
the new GET-by-id route and the upload route's dedup-response branch can't drift
|
|
from the list route's mapping. Computes `parseStatus` from `parsed`/`parse_error`
|
|
and `docType` from `scan_mode`, falling back to the legacy `order_untuk ==
|
|
"PRODUCT SCAN"` sentinel for rows predating this column — verified via `curl`
|
|
against a pre-existing pre-9.1 document that it doesn't regress to `docType:
|
|
undefined`.
|
|
|
|
### 2.3 `GET /api/v1/documents/:id` (`api/v1/documents/[id]/route.ts`)
|
|
- Reuses the exact same auth/scoping pattern as the existing `PUT` on the same
|
|
file (401 no-account, 404 no-row, 403 non-admin/wrong-store) — no new auth
|
|
logic introduced, just the existing helper called a second time.
|
|
- Deliberately omits the list route's `parsed = true` filter, since surfacing
|
|
pending/failed rows is the entire point of the endpoint.
|
|
|
|
### 2.4 Failure recording (`api/v1/documents/upload/route.ts`)
|
|
- The one gap not already covered by `/api/parse`'s own pre-existing error
|
|
fallback (which already flips `parsed=true` with "Not Found" placeholder
|
|
metadata, unchanged by this task) is the internal fetch call to `/api/parse`
|
|
itself never completing — network error or the pre-existing 210s
|
|
`AbortSignal.timeout` firing. Both that `catch` branch and a new `!response.ok`
|
|
check now persist a short message to `documents.parse_error`, which is the only
|
|
input the new `parseStatus: "failed"` branch depends on.
|
|
- Dedup branch fixed to run the existing document through the same shared mapper
|
|
instead of a hand-built always-empty stub (G10) — a GPS-tag fallback to the
|
|
retry's own coordinates was preserved for documents that never got one on first
|
|
upload, matching the previous behavior's intent.
|
|
|
|
### 2.5 `scan_mode` persistence in `/api/parse` (`api/parse/route.ts`)
|
|
- Minimal, additive `COALESCE(EXCLUDED.scan_mode, documents.scan_mode)` in both
|
|
`ON CONFLICT` blocks, same pattern already used for `kode_toko` — so documents
|
|
created via the classic route (not just the v1 upload path) also get a correct
|
|
`scan_mode`. `parse_error = NULL` added to both `SET` clauses to clear a stale
|
|
failure once a parse actually completes. This file remains accepted §B3 debt
|
|
(604 lines pre-existing, per `backend/AGENTS.md` Adaptation Notes) — touched only
|
|
minimally, not restructured, consistent with that note's "split only if/when
|
|
touched" guidance being about restructuring, not about refusing small edits.
|
|
|
|
## 3. Audit Verification
|
|
- **Functional testing** (`docker compose up -d --build` from repo root, real
|
|
`curl` calls against the live stack, not just unit tests):
|
|
- Confirmed `scan_mode`/`parse_error` columns exist post-migration via `psql \d
|
|
documents` against the running container — no `ALTER TABLE` errors in logs.
|
|
- Logged in as a real store account (`WH_JCIBBR1`), uploaded a real DO test
|
|
image (`sources/test-images/do-001.jpg`): `GET /api/v1/documents/:id` returned
|
|
the real parsed header/items, `parseStatus: "done"`, `docType: "DO"`.
|
|
- Re-uploaded the identical file (dedup path): response now carries the same
|
|
real header/items instead of the old empty stub — confirmed G10 fixed.
|
|
- Uploaded a real product photo with `scan_mode=Product`: `docType: "Product"`
|
|
confirmed both in the GET response and directly in Postgres
|
|
(`SELECT scan_mode FROM documents`).
|
|
- `GET /:id` with no token → 401; nonexistent id → 404; a *different* store
|
|
account's token against another store's document → 403; `admin`'s token
|
|
against the same document → 200 (admin bypass intact).
|
|
- `GET /api/v1/documents` (list) still returns only parsed, non-sample
|
|
documents, now carrying `docType`/`parseStatus` for free via the shared
|
|
mapper — existing 401 behavior unchanged.
|
|
- `npx tsc --noEmit` clean across the whole `pfm-web-app` project.
|
|
|
|
## 4. Conclusion
|
|
Task 9.1 closes gaps G1 (no per-document GET / N+1 list polling), G10 (dedup stub),
|
|
and the server half of G4 (fabricated doc-type sentinel) exactly as scoped. All
|
|
new behavior was verified against the live Docker stack with real uploads, not
|
|
just a clean build — auth/scoping regressions were explicitly checked and none
|
|
were found. Flutter-side consumption (`plans/next-enhancements.md` §6.1/§7.3)
|
|
remains open and unblocked by this change.
|
|
|
|
---
|
|
|
|
# Iteration Log & Audit: Authenticated v1 Product-Scan Endpoint (Task 9.3)
|
|
|
|
## 1. Objective
|
|
Conduct a code review and audit of task 9.3 — `POST /api/v1/scan-product`, an
|
|
authenticated equivalent of the classic dev-only `/api/scan-pfm` — to close gap
|
|
G2/G3 (`docs/api-contract-map.md`): the Flutter product editor currently reaches
|
|
the classify+match pipeline via an unauthenticated route that task 4.5 already
|
|
excluded from the public tunnel, so product scanning is broken off-LAN.
|
|
|
|
## 2. Code Review
|
|
|
|
### 2.1 Shared util (`pfm-web-app/src/utils/product-scan.ts`, new file)
|
|
- `classifyAndMatchProduct` is a byte-for-byte extraction of the classic route's
|
|
classify-call + Levenshtein-SKU-match logic (not a rewrite) — reduces the risk
|
|
that the new v1 route's behavior silently diverges from the already-working
|
|
classic route's matching quality.
|
|
- `ClassifierError` deliberately preserves the classic route's existing behavior
|
|
of forwarding the Python classifier's own HTTP status on failure, rather than
|
|
letting a generic `catch` collapse every failure to 500 — both the classic and
|
|
new v1 route special-case it identically.
|
|
- Intentionally excludes the layout-parsing visualization block: that's
|
|
desktop-test-page-only per the task's explicit response-field list
|
|
(`possibleMatches`, `ocr`, `classification` — no `layoutParsingResult`), so it
|
|
correctly stays in `api/scan-pfm/route.ts` rather than being pulled into the
|
|
shared util or the new v1 route.
|
|
|
|
### 2.2 Classic route refactor (`api/scan-pfm/route.ts`)
|
|
- Response shape (`{classification, ocr, possibleMatches, layoutParsingResult}`,
|
|
no envelope, no auth) is unchanged — this route still serves the desktop test
|
|
page exactly as before, now just calling the shared util instead of inlining
|
|
the logic. Dropped one genuinely dead variable (`extractedProductName`, computed
|
|
but never read in the original code) as part of the extraction.
|
|
|
|
### 2.3 New v1 route (`api/v1/scan-product/route.ts`)
|
|
- Auth: any authenticated account (not admin-gated) — correct, since this is the
|
|
route the mobile app's own store-role accounts call to perform a scan, unlike
|
|
`master/skus` writes which are intentionally admin-only.
|
|
- Dual input handling (multipart primary, JSON base64 fallback) matches the task's
|
|
explicit wording ("multipart (preferred...) or base64") and lets Flutter adopt
|
|
this endpoint today regardless of which shape task 7.1 ends up sending.
|
|
- No `nginx.conf` change was needed — confirmed the port-8001 restricted block's
|
|
`location /api/v1/` (line 135) is a prefix match already covering the new path.
|
|
|
|
## 3. Audit Verification
|
|
- **Functional testing** against the live stack (same running containers as task
|
|
9.1's session; `pfm-web-app` restarted once to pick up the new route file after
|
|
its dev-server file watcher missed the new directory — a known bind-mount
|
|
quirk on Windows Docker Desktop, not a code issue):
|
|
- `POST /api/v1/scan-product` with a real product photo as multipart `image` +
|
|
a real store account's bearer token: `200`, `{status:"success", data:
|
|
{classification, ocr, possibleMatches}}` with a correct top-5 match list and
|
|
`isBestMatch` on the top entry; `ocr` confirmed to include
|
|
`extracted_expired_date`.
|
|
- Same call with no token → `401`.
|
|
- Same image via a JSON `{image_base64}` body instead of multipart → identical
|
|
`possibleMatches` output, confirming both input paths produce the same result.
|
|
- Classic `POST /api/scan-pfm` (JSON body, no auth) with the same image →
|
|
unchanged response shape and matching results, including
|
|
`layoutParsingResult` still present — no regression from the extraction.
|
|
- `npx tsc --noEmit` and `npx eslint` on the three touched/new files clean
|
|
(aside from pre-existing `any`-for-JSONB-shaped-data style already used
|
|
throughout this codebase, e.g. `document-mapper.ts` from task 9.1).
|
|
|
|
## 4. Conclusion
|
|
Task 9.3 closes gap G2/G3 exactly as scoped: the mobile app now has an
|
|
authenticated, tunnel-reachable path to the classify+match pipeline that returns
|
|
identical results to the already-proven classic route, verified against real
|
|
classifier output rather than mocked data. Task 9.2 (non-admin SKU list read)
|
|
remains open and separate. Flutter-side consumption (`plans/next-enhancements.md`
|
|
§7.1) remains open and is now unblocked by this change (alongside 9.2).
|
|
|
|
# Iteration Log & Audit: Non-Admin SKU List Read (Task 9.2)
|
|
|
|
## 1. Objective
|
|
Conduct a code review and audit of task 9.2 — read access to the SKU master
|
|
list (`GET /api/v1/master/skus`) for any authenticated account, not just
|
|
`admin` — closing gap G2 alongside 9.3. Picked up via an explicit
|
|
backend-scoped `n{9.2}` request after the user asked to clarify the two
|
|
options the task itself flagged as undecided (relax the existing endpoint
|
|
vs. add a new one); user chose to relax the existing endpoint.
|
|
|
|
## 2. Code Review
|
|
- `master/skus/route.ts`'s `GET` handler previously rejected any non-`admin`
|
|
account with 403, forcing the Flutter product editor to call the
|
|
unauthenticated classic `GET /api/skus` instead (the actual bug this task
|
|
fixes - that classic route was removed from the public ngrok tunnel by
|
|
task 4.5, so product scans off-LAN were already broken before this fix).
|
|
- Changed the `GET` guard from `!account || account.role !== 'admin'`
|
|
(403 either way) to `!account` (401 for no/invalid token, any valid
|
|
account now passes) - a one-line, surgical change matching the user's
|
|
chosen option exactly. `POST` (SKU creation) was deliberately left
|
|
untouched, still admin-gated - the task's own text specified "writes stay
|
|
admin-only," and admin master-data management is a different concern from
|
|
a mobile client reading the catalog to populate a dropdown.
|
|
- No response-shape change: still `{status: "success", data: res.rows}`,
|
|
matching what task 9.2 asked for (the `{status, data}` v1 envelope) and
|
|
what the Flutter product editor already expects once it switches over
|
|
(root task 7.1, not yet picked up).
|
|
|
|
## 3. Audit Verification
|
|
- Hit a real hot-reload gap during verification: the file was correctly
|
|
updated on disk inside the `pfm-web-app` container (confirmed via
|
|
`docker exec ... cat`), but the running Turbopack dev server kept serving
|
|
the old admin-gated behavior - a known class of issue where Windows-host
|
|
bind-mount file-change events don't reliably reach `next dev`'s watcher.
|
|
Fixed by `docker restart paddleocr-pfm-web-app`, after which the new code
|
|
took effect immediately (confirmed via a fresh `curl` round-trip).
|
|
- Verified via `curl` against the live stack (real account credentials
|
|
pulled from the live `accounts` table, not fixtures):
|
|
- A real non-admin (`store` role) account's token: `GET
|
|
/api/v1/master/skus` → `200`, real `sku_master` rows returned.
|
|
- No `Authorization` header at all: `401 Unauthorized` (previously this
|
|
same case incorrectly returned `403`, since the old guard checked
|
|
`!account || role !== 'admin'` as one combined condition - now correctly
|
|
distinguishes "no valid account" from "valid but insufficient role").
|
|
- The same non-admin token against `POST /api/v1/master/skus` (attempting
|
|
to create a SKU): still `403 Forbidden: Admin access required` - writes
|
|
unaffected.
|
|
- An admin token against `GET /api/v1/master/skus`: still `200` - no
|
|
regression for the existing admin master-data UI.
|
|
|
|
## 4. Conclusion
|
|
Task 9.2 closes gap G2's remaining half: the mobile app can now read the
|
|
SKU master list through the authenticated, tunnel-reachable `/api/v1/*`
|
|
surface without impersonating a dev-only unauthenticated route. Combined
|
|
with 9.1 and 9.3 (both already shipped), every backend blocker behind root
|
|
`plans/next-enhancements.md` §7.1 (moving the product editor onto the v1
|
|
surface) is now cleared - that Flutter task is unblocked and ready to pick
|
|
up. §7.2 (eliminating the duplicate classification pass) is separately
|
|
unblocked in principle (9.1's `docType`/metadata work + 9.3's endpoint both
|
|
exist now) but still needs its own client-side decision about which single
|
|
pass to keep, per that task's own grill-me note.
|
|
|
|
---
|
|
|
|
## Iteration & Audit: Tasks 10.1/10.2 — Document Confirmation Gate & Data Hygiene (2026-07-10)
|
|
|
|
### 1. Objective
|
|
Close backend §10, sourced from user feedback on the release APK
|
|
(`twinkly-riding-mitten.md`, root-cause documented as gaps **G11**/**G12** in
|
|
`docs/api-contract-map.md`): documents were visible via `GET /api/v1/documents`
|
|
the instant OCR parsing finished, before the mobile user ever confirmed them
|
|
via `PUT`, and Product Scan uploads carried fabricated `noPO`/`noSO`/`noDO`
|
|
placeholder values.
|
|
|
|
### 2. Code Review
|
|
- **`db/init.ts`**: new `confirmed BOOLEAN NOT NULL DEFAULT TRUE` column, both
|
|
in the `CREATE TABLE IF NOT EXISTS` block and as an idempotent
|
|
`ALTER TABLE ... ADD COLUMN IF NOT EXISTS` for already-running DBs, matching
|
|
the exact pattern already used for `scan_mode`/`parse_error`. `DEFAULT TRUE`
|
|
is a deliberate grandfather clause — every row that existed before this
|
|
migration counts as already-confirmed, so existing history doesn't vanish.
|
|
- **`v1/documents/upload/route.ts`**: the one real INSERT path for a fresh
|
|
mobile capture now explicitly inserts `confirmed = false`; the dedup-hit
|
|
branch (no INSERT) is untouched, correctly reflecting whatever state the
|
|
original row already has. Its SELECT for the dedup branch was also extended
|
|
to fetch `confirmed` so the mapper has it.
|
|
- **`utils/document-mapper.ts`**: `DocumentRow` interface and
|
|
`mapDocumentRow()`'s return both carry `confirmed` through now, so all three
|
|
call sites (list, GET-by-id, upload dedup) stay in sync from one place —
|
|
same shared-mapper pattern task 9.1 established.
|
|
- **`v1/documents/route.ts`** (list): `AND confirmed = true` added to the
|
|
`WHERE` clause with no role branching — applies to `admin` exactly the same
|
|
as `store` accounts, per the user's explicit answer when asked whether admin
|
|
should retain oversight visibility into unconfirmed documents (they chose
|
|
"no special-casing").
|
|
- **`v1/documents/[id]/route.ts`**: `PUT` now sets `confirmed = true` alongside
|
|
the existing `parsed = true` in its `UPDATE` — the *only* place this flips.
|
|
`GET`-by-id is untouched, deliberately: its existing comment already says
|
|
the point of this endpoint is letting the poller see pending/failed
|
|
documents, and that reasoning extends unchanged to unconfirmed ones — the
|
|
poller must detect parse-completion before the user has had a chance to
|
|
confirm anything.
|
|
- **`parse/route.ts`**: reasoned through, rather than blindly copied, whether
|
|
its own `INSERT ... ON CONFLICT (filename) DO UPDATE` statements (DO and
|
|
Product branches) needed `confirmed` handling. In the real mobile flow the
|
|
upload route's INSERT always runs first, so this statement always resolves
|
|
via the `ON CONFLICT` branch; since `confirmed` is absent from that branch's
|
|
`SET` clause, Postgres leaves the row's existing value untouched by design —
|
|
correct behavior (never regress an already-confirmed row, never reset a
|
|
pending one mid-reparse) without adding a single line. Also removed the
|
|
Product branch's fabricated `noPO`/`noSO`/`noDO` placeholder values (task
|
|
10.2) — replaced with empty strings after confirming (by reading
|
|
`pdf_service.dart` and `product_editor_submit_logic.dart` on the Flutter
|
|
side) that nothing reads them meaningfully; the confirmed document's real
|
|
values always come from the user's own PO-link/batch selection at PUT time
|
|
regardless.
|
|
|
|
### 3. Audit Verification
|
|
Live against the running Docker stack (`docker restart paddleocr-pfm-web-app`
|
|
to pick up the code + run the migration):
|
|
- `\d documents` confirmed the new `confirmed boolean not null default true`
|
|
column; `SELECT count(*) FROM documents WHERE confirmed = true` returned
|
|
13 (all pre-existing rows), `= false` returned 0 — grandfather clause held.
|
|
- Uploaded a real DO photo as store account `WH_JCIBBR1` without ever calling
|
|
`PUT`: absent from that store's `GET /documents` (count unchanged at 3, new
|
|
id 3400 not present) and absent from `admin`'s list too (13, unchanged);
|
|
`GET /documents/3400` still returned `parseStatus: "done"`,
|
|
`confirmed: false` — the poller/editor hand-off path is unaffected.
|
|
- `PUT /documents/3400` (confirm) with real header/shipment data: doc count
|
|
for `WH_JCIBBR1` became 4, id 3400 now present with the real submitted
|
|
`namaPenerima` ("Penerima Test", not a placeholder); `psql` confirmed the
|
|
row's `confirmed` column flipped to `t`.
|
|
- Uploaded a fresh Product Scan as the same store (doc id 3402), read its raw
|
|
unconfirmed `GET /documents/3402` response: `header.no_po`/`no_so`/`no_do`
|
|
all returned `""` — the old `"PO-PRODUCT-001"`/`"1002003004"`/
|
|
`"DO-PRODUCT-999"` placeholders are gone.
|
|
|
|
### 4. Conclusion
|
|
Backend §10 is fully `[DONE]`. Flutter's corresponding root task §8.2 (add an
|
|
optional `confirmed` field to `DocumentModel`, default `true` for
|
|
legacy/cached responses) was implemented and verified in the same session —
|
|
see root `docs/iteration-log.md`. No remaining backend blocker for gap G11 or
|
|
G12.
|
|
|
|
---
|
|
|
|
## Iteration & Audit: Task 11.1 — Single-Pass Product Classification (2026-07-10)
|
|
|
|
### 1. Objective
|
|
Close backend §11 (gap **G3**), sourced directly from user feedback after
|
|
they noticed Product Scan's confirmation screen took visibly longer to open
|
|
than DO Scan's and asked why. G3 had been documented earlier this session in
|
|
`docs/api-contract-map.md` but deliberately left `[TODO]`/deferred, pending
|
|
exactly the client-side decision the user's follow-up message resolved:
|
|
"sama seperti scan DO... GPU tidak 2x kerja" (same as DO scan, GPU shouldn't
|
|
run twice) — i.e. do the classify pass once, at upload, and have the editor
|
|
read the stored result, not re-classify on review.
|
|
|
|
### 2. Code Review
|
|
- **`api/parse/route.ts`'s Product branch**: replaced its own separate,
|
|
inline `fetch(pyServerUrl, ...)` (which discarded everything except
|
|
`top1_name`/`extracted_sku`) with a call to the already-existing shared
|
|
`classifyAndMatchProduct()` from `utils/product-scan.ts` — the same
|
|
function `POST /api/v1/scan-product` (task 9.3) uses, which additionally
|
|
runs the Levenshtein SKU-match against `sku_master` for a real top-5
|
|
candidate list and returns the raw OCR result (`extracted_expired_date`
|
|
included). `b64` (the image's base64 encoding) was already computed
|
|
earlier in this function for the DO path — reused, not recomputed, so
|
|
this is a strict reduction in duplicated work, not an addition.
|
|
- **New `metadata.productScan` key**: `{ possibleMatches, extractedExpiryDate
|
|
}` stored alongside the existing `header`/`shipment`/`items` keys in the
|
|
same JSONB `metadata` column — no migration, following the exact precedent
|
|
those other keys already set for coexisting shapes in one column.
|
|
- **`utils/document-mapper.ts`**: added a top-level `productScan` field to
|
|
`mapDocumentRow()`'s return (`metadata.productScan || null`), so all three
|
|
GET call sites (list, by-id, upload dedup) expose it identically, same
|
|
shared-mapper pattern as `parseStatus`/`docType`/`confirmed`.
|
|
- **Regression audit, not just addition**: read `classifyAndMatchProduct()`'s
|
|
own `fetch` call closely while wiring it into `parse/route.ts` and noticed
|
|
it had *no* `AbortSignal` at all — the inline call it was replacing in
|
|
`parse/route.ts` had an explicit 90s bound
|
|
(`PIPELINE_TIMEOUT_MS`/`AbortSignal.timeout`). Silently dropping that bound
|
|
would have been a real regression (a wedged GPU container hanging past the
|
|
intended fail-fast point). Fixed by adding the identical 90s bound directly
|
|
inside `classifyAndMatchProduct()` itself — which also retroactively fixes
|
|
the live `POST /api/v1/scan-product` route, which never had this bound
|
|
either (pre-existing gap, not something this task's own diff introduced,
|
|
but caught and closed while in the area).
|
|
|
|
### 3. Audit Verification
|
|
Live against the running Docker stack (`docker restart paddleocr-pfm-web-app`
|
|
to pick up the code):
|
|
- Deliberately chose a genuinely fresh image/store combination
|
|
(`do-015.jpg`, never uploaded before, as store `WH_JAFATAH`) to rule out a
|
|
dedup hit masking whether real classification ran. Response was
|
|
`"Document uploaded successfully"` (the fresh-insert branch, not the
|
|
dedup-return branch) and took **9 seconds** — consistent with one real GPU
|
|
classify+match pass, not a cache hit.
|
|
- Immediate `GET /documents/:id` (no editor interaction, no second request)
|
|
returned a fully populated `productScan`: 5 real `possibleMatches` with
|
|
real `sku_master` names/scores (e.g. `"CHAMP CRUNCHY HOTZZ 300 GR/PAC"` at
|
|
`score: 0.7575...`, matching real product naming conventions, not
|
|
fabricated placeholders) and `extractedExpiryDate` (empty string here,
|
|
since this particular test image has no visible expiry text - correctly
|
|
reflecting a real "not found" rather than a fake date, consistent with the
|
|
G7 fix's "no dummy data" rule).
|
|
- Confirmed via a second, earlier check (before switching to the guaranteed-
|
|
fresh combination above) that a dedup-hit response for a different
|
|
document (id 3408) *also* returned a fully populated `productScan` from a
|
|
prior parse - proving the data survives the dedup-return code path too
|
|
(`upload/route.ts`'s dedup SELECT was already extended for `confirmed` in
|
|
task 10.1's session and needed no further change here, since it maps
|
|
through the same shared `mapDocumentRow()`).
|
|
|
|
### 4. Conclusion
|
|
Backend §11 is `[DONE]`. Flutter's corresponding root task §7.2 (read
|
|
`productScanMatches`/`productScanExtractedExpiryDate` directly from the
|
|
document instead of re-calling `/scan-product`) was implemented and verified
|
|
in the same session — see root `docs/iteration-log.md`. Gap G3 is resolved;
|
|
`docs/api-contract-map.md` updated accordingly. `POST /api/v1/scan-product`
|
|
itself is intentionally left in place (unused by this flow now, but a
|
|
legitimate, reusable authenticated endpoint - e.g. for a possible future
|
|
"rescan this photo" action) rather than removed, since removing a working,
|
|
independently-useful route wasn't part of what this task's scope required.
|
|
|
|
---
|
|
|
|
# Iteration Log & Audit: Product Classifier Retrain on Full Dataset (Task 2.5)
|
|
|
|
## 1. Objective
|
|
The reference photo dataset (`pfm-web-app/public/produk-pfm/foto-kemasan-v2/`)
|
|
had grown to **81 product classes / 2,493 photos**, but the deployed model
|
|
artifacts (`models/dinov2_index.pkl`, `models/produk-pfm-classifier-26n-100e-
|
|
2026-07-08.pt`/`.onnx`) were still the ones trained 2026-07-08 against only the
|
|
original **16 classes / 118 photos** — confirmed by counting the class-index
|
|
keys embedded in the ONNX file's metadata (16 numeric keys found, matching
|
|
`docs/scan-product.md`'s "16 classes, 118 photos" note exactly). The other 65
|
|
classes existed as raw photos with no corresponding trained weights. Goal:
|
|
retrain both artifacts against the full current dataset via the documented
|
|
Docker-based retraining procedure (`docs/scan-product.md`'s "Model artifacts &
|
|
retraining" section), and record real timing/accuracy rather than estimates.
|
|
|
|
## 2. Work Performed
|
|
- Started Docker Desktop (not running at session start) and confirmed
|
|
`--gpus all` passthrough works against the host's NVIDIA GeForce RTX 2060
|
|
(6GB VRAM).
|
|
- `docker compose build pipeline-api` from the repo root — rebuilds the image
|
|
with the current `foto-kemasan-v2/` baked in via `COPY . /app` (no
|
|
`.dockerignore` entry excludes it). Build succeeded in **2m54s**.
|
|
- Ran `index_dinov2.py` in a one-off `docker run --gpus all` container with
|
|
`models/` bind-mounted **writable** (the live `pipeline-api` compose service
|
|
mounts it `:ro`) via `/app/.venv-api/bin/python` (the venv `Dockerfile`
|
|
installs `paddlepaddle-gpu`/`ultralytics`/`torch` into, not the base
|
|
interpreter). Result: **"Success! Indexed 2493/2493 images"** — every photo
|
|
across all 81 classes embedded into a fresh `dinov2_index.pkl`.
|
|
- Ran `train_classifier.py train --imgsz 224` the same way. Its own
|
|
`split_dataset()` groups images by source photo (stripping any `_aug_N`
|
|
suffix) and shuffles before cutting 80/20, so augmented copies always land
|
|
with their source and no class is split naively by filename order — verified
|
|
this behavior in the source (`train_classifier.py:88-176`) before relying on
|
|
it, rather than assuming.
|
|
- **Training was stopped by explicit user request** (`docker stop`) at
|
|
**epoch 43/100, 23m0.998s elapsed**, before it produced a final checkpoint.
|
|
The last completed validation pass (epoch 42) reported **84.3% top-1 / 93.9%
|
|
top-5** across all 81 classes — already ahead of the old 16-class model's
|
|
83.3%/90%, but not a final number since the run never reached completion.
|
|
|
|
- **Session resumed**: after the pause above, Docker Desktop had actually
|
|
stopped between sessions — a first resume attempt failed instantly with a
|
|
daemon-connection error before any training ran. Restarted Docker Desktop,
|
|
confirmed `docker ps` responsive, confirmed the `pipeline-api` image and
|
|
`dinov2_index.pkl` from the earlier session were both still intact (no
|
|
rebuild/reindex needed), then relaunched `train_classifier.py train
|
|
--imgsz 224` from epoch 0 in a fresh one-off container, timed with `time`.
|
|
- **Training ran to completion this time: 100/100 epochs, real elapsed time
|
|
54m21.248s.** Final validation: **85.8% top-1 / 94.4% top-5** across all 81
|
|
classes. Published `models/produk-pfm-classifier-26n-100e-2026-07-14.pt`
|
|
(3.4MB) and exported `.onnx` (6.3MB, ONNX opset 20).
|
|
|
|
## 3. Verification
|
|
- Confirmed via `ls` on the host `models/` directory that the new dated
|
|
`produk-pfm-classifier-26n-100e-2026-07-14.pt`/`.onnx` files exist (dated
|
|
2026-07-14 09:05/09:06), alongside the untouched 2026-07-08 files.
|
|
- Confirmed in the training log's own ONNX export step that the model's
|
|
output shape is `(1, 81)` — i.e. genuinely 81 output classes, not a stale
|
|
16-class head.
|
|
- Ran `docker compose up -d pipeline-api` (the main compose stack wasn't
|
|
running this session — confirmed via `docker compose ps` returning empty —
|
|
so this was a fresh start, not a "restart"; it correctly pulled in the
|
|
`vllm-server` dependency too) and polled `docker logs
|
|
paddleocr-pipeline-api` until startup markers appeared. Confirmed lines:
|
|
- `DINOv2 index loaded with 2493 reference images.`
|
|
- `Using classifier weights: /app/pfm-web-app/public/produk-pfm/models/produk-pfm-classifier-26n-100e-2026-07-14.pt`
|
|
- `YOLO model loaded successfully.`
|
|
- `INFO: Application startup complete.`
|
|
|
|
This is real, observed runtime behavior — the live `pipeline-api` service is
|
|
now actually serving the new 81-class model and the full 2,493-image
|
|
DINOv2 index, not an assumption based on `latest_classifier_weights()`'s
|
|
glob-newest-by-date logic.
|
|
|
|
## 4. Status
|
|
**Done, 2026-07-14.** Both artifacts (DINOv2 index, YOLO classifier) retrained
|
|
against the full 81-class/2,493-photo dataset and verified loading in the live
|
|
service. `plans/next-enhancements.md` task 2.5 flipped to `[DONE]` with these
|
|
same numbers; `docs/scan-product.md` and `backend/CLAUDE.md`'s class-count/
|
|
accuracy claims updated from 16→81 classes and 83.3%/90%→85.8%/94.4%;
|
|
`docs/feature-list.md` given a matching entry. Remaining gap toward the
|
|
program's stated ±230-SKU target (see `proposals/sources/` kick-off material)
|
|
is dataset growth, not a code or training-pipeline limitation — the same
|
|
`train_classifier.py`/`index_dinov2.py` procedure documented here scales to
|
|
however many classes `foto-kemasan-v2/` ends up containing.
|