Files
fhanyuh caf8e98378 chore: normalize line endings (CRLF -> LF)
No content changes: git diff --ignore-all-space over these files is empty.
The churn came from editing on Windows against a repo checked out with LF.
2026-08-27 10:40:49 +07:00

23 KiB

Feature List (backend)

Structured log of shipped backend features, updated by the n/next workflow (see AGENTS.md Part B) whenever a task in plans/next-enhancements.md is marked [DONE]. Split out 2026-07-08 from root docs/feature-list.md's backend sections — this file is the sole home for backend feature history going forward.

Format

## <Section / Module Name>

- **<task number>** <feature description> — shipped <date>

Existing Features (pre-kit)

Backfilled 2026-07-08 during adoption of this kit — these predate the e/n workflow and have no task numbers; see git log for real dates/history.

Backend — Next.js API Gateway

  • Upload/parse/documents CRUD routes, GPU status endpoint, vLLM proxy, manual-label review tool.

Backend — OCR Pipeline & Accuracy

  • PaddleOCR + vLLM classification pipeline with DB layout caching, table column-shift correction, date normalization, and an accuracy regression harness (pfm-web-app/scripts/accuracy-check.mts) — 95.10% overall as of 2026-07-08 (target 95% met; see task 2.3 below for the investigation and CLAUDE.md).

Backend — Postgres Data Layer

  • Schema/init in pfm-web-app/src/db/init.ts, served via the canonical root docker-compose.yml stack.
  • 3.3 Added a standard INDEX on documents(file_hash) in db/init.ts to accelerate the upload deduplication queries without strictly enforcing uniqueness across different stores. Correspondingly updated the dedup query in api/v1/documents/upload/route.ts to scope duplicate detection by kode_toko. This fixes a conflict where one store could be incorrectly linked to another store's duplicate receipt image — shipped 2026-07-08.

DevOps — Docker & Dev Tunnel

  • 4.1 Docker Compose Policy Documented: Formalized the execution policy in README.md and CLAUDE.md, explicitly requiring the use of the docker-compose.demo.yml override (production build) for all client demonstrations and field testing to bypass the Next.js dev server bottleneck — shipped 2026-07-08.
  • Docker Compose Dependency Gates: Added strict Docker healthcheck gates (Task 4.2) blocking the pfm-web-app (Next.js) from starting until PostgreSQL and the VLLM models are initialized and fully healthy.
  • Secure Tunnel Ingress: Restructured nginx.conf and start-dev-tunnel.ps1 (Tasks 4.3, 4.5) to expose a dedicated, restricted port (8001) that exclusively routes to /api/v1/*. This perfectly secures the development UI (/scan-pfm) and legacy routes from public exposure.
  • Dead Config Pruning: Stripped deprecated and redundant proxy blocks from the Nginx edge router (Task 4.4).

(New features shipped via n/next go below, organized the same way, with task numbers.)

Backend — Next.js API Gateway

  • 1.4 Enforced real 401 auth on /api/v1/documents/* (list, PUT-by-id, upload) — the actual production API surface, already fully supported by the Flutter client (real login + Authorization: Bearer on every request). Previously none of these three routes rejected a missing/invalid token; upload only optionally read it. Added the pre-existing getAccountFromAuthHeader() helper (utils/auth.ts) + a 401 guard to all three; OPTIONS (CORS preflight) untouched. The original task 1.3 (auth on the classic routes) was cancelled instead — those routes are dev-only web UI surface with no login flow, going away in production. Verified via curl: 401 with no token, success with a real token from /api/v1/auth/login — shipped 2026-07-08.
  • 1.5 Implemented per-store data scoping on /api/v1/documents/*. Added kode_toko column to documents table via db/init.ts migration. The upload route now binds kode_toko to documents upon creation. GET /api/v1/documents and PUT /api/v1/documents/:id enforce ownership checks (kode_toko matching) for store role accounts, while admin retains global access including legacy unassigned documents — shipped 2026-07-08.
  • 1.6 GET /api/v1/health Endpoint: Unauthenticated health probe verifying both PostgreSQL connectivity and Pipeline API HTTP reachability. Returns HTTP 503 if any core dependency is down — shipped 2026-07-08.
  • Ad-hoc Connected /scan-pfm page with /api/scan-pfm route and enabled auto-trigger scanning on custom file upload, sample selection, thumbnail change, and canvas rotation. Supported both image and image_base64 payload keys — shipped 2026-07-09.
  • 9.1 Added GET /api/v1/documents/:id (same 401/403 scoping as PUT), returning a single document — including still-unparsed rows — with a new parseStatus: "pending"|"done"|"failed" field, so the Flutter poller can move off scanning the entire list every 2s. Added scan_mode/parse_error columns to documents (db/init.ts, migrated via ALTER TABLE ... ADD COLUMN IF NOT EXISTS for already-running DBs). scan_mode is now persisted on upload (v1/documents/upload/route.ts) and on the classic /api/parse route's upserts (COALESCE, same pattern as kode_toko), and surfaced as docType on every GET response (utils/document-mapper.ts, a new shared helper extracted from the list route's inline mapping so list/by-id/dedup all agree) — falling back to the legacy order_untuk == "PRODUCT SCAN" sentinel for pre-existing rows with no scan_mode. parse_error is now recorded when the upload route's internal call to /api/parse itself fails to complete (network error or the 210s abort firing) — previously this was silently swallowed and the document stayed parsed=false forever with no signal, burning the client's full 260s timeout; /api/parse's own existing pipeline-error fallback (parsed=true + "Not Found" placeholder) was already fine and is unchanged. Also fixed the dedup branch (a repeat upload of an already-seen file) to return the original document's real current state via the same mapper instead of a hardcoded empty stub. Verified via docker compose up -d --build + curl: schema migration applied cleanly to the live DB (confirmed via psql), DO and Product uploads both correctly persist scan_mode and surface it as docType, a dedup retry returns real header/items instead of an empty stub, GET /:id returns 401 (no token) / 403 (wrong store) / 404 (nonexistent id) / 200 (admin or owning store), and the list endpoint's existing filter/scoping is unchanged — shipped 2026-07-10.
  • 9.3 Added authenticated POST /api/v1/scan-product, the v1 equivalent of the classic dev-only /api/scan-pfm (unauthenticated, and unreachable off-LAN since task 4.5 restricted the public tunnel to /api/v1/*). Extracted the shared classify-and-match logic (Python classifier call + Levenshtein SKU matching against sku_master, top-5 scoring) out of api/scan-pfm/route.ts into a new utils/product-scan.ts (classifyAndMatchProduct, plus a ClassifierError class that preserves forwarding the classifier's own HTTP status instead of collapsing every failure to 500) so the classic route and the new v1 route share one implementation instead of duplicating it — the classic route's response shape, auth-free behavior, and desktop-only layout-parsing visualization are otherwise unchanged. The new route accepts either multipart (image/file field, matching the v1 upload route's convention) or a JSON {image_base64} body, is open to any authenticated account (not admin-gated, since this is what the mobile app itself calls), and wraps the result in the standard {status, data} envelope with classification, ocr (including extracted_expired_date), and possibleMatches. Verified via curl against the live stack with a real product photo: multipart upload and JSON-body variants both return identical, correct top-5 matches; no-token request returns 401; the classic /api/scan-pfm route's response (including layoutParsingResult) is unchanged post-refactor — shipped 2026-07-10.
  • 9.2 Relaxed GET /api/v1/master/skus (master/skus/route.ts) so any authenticated account can read the SKU master list, not just admin — the Flutter product editor needs this and previously had to string-hack its base URL to call the unauthenticated classic GET /api/skus, which task 4.5 had already removed from the public tunnel, breaking product scans off-LAN. Changed the guard from a combined !account || role !== 'admin' check (403 for both "no token" and "wrong role") to !account (correct 401) followed by an unconditional pass-through for any valid account; POST (SKU creation) is untouched, still admin-only, per the user's explicit choice between the two options this task flagged as undecided. No response-shape change. Verified via curl against the live stack with a real non-admin (store role) account's token: GET → 200 with real data; no token → 401 (was incorrectly 403 before this fix); the same non-admin token against POST → still 403; admin GET → still 200. Along the way, hit and resolved a dev-loop issue: the container had the edited file on disk but Turbopack's file watcher wasn't detecting the change over the Windows bind mount, requiring docker restart paddleocr-pfm-web-app to pick it up — noted in case it recurs for future edits. With 9.1-9.3 all shipped, Flutter root task 7.1 (moving the product editor onto the v1 surface) is now fully unblocked — shipped 2026-07-10.

Backend — OCR Pipeline & Accuracy

  • 2.1 Built the Product/SKU scan classifier's model artifacts: models/dinov2_index.pkl (118/118 reference photos indexed across 16 SKU classes) and models/produk-pfm-classifier-26n-100e-2026-07-08.pt (+ .onnx export) — a YOLO classifier fine-tuned 100 epochs, 83.3% top-1 / 90% top-5 validation accuracy on the current (thin, 2-16 photos/class) dataset. Built via a one-off docker run from a freshly-rebuilt pipeline-api image (bare-metal training isn't viable on Windows — paddlepaddle-gpu's wheel index is Linux-only). pipeline-api restarted and confirmed loading both models from logs. Also fixed scripts/install-pipeline.sh, which was missing ultralytics/torch — shipped 2026-07-08.
  • 2.1 (verification pass) Ran a full browser walkthrough of /scan-pfm (classification, top-5, OCR expiry extraction + crop, SKU-master matching, Visual/Spotting Grid, Raw Response — all confirmed working with real data). Found and fixed a real bug: "Save Ground Truth" was returning success but silently writing into the pfm-web-app container's ephemeral filesystem instead of the host, because /sources wasn't a bind-mounted path in root docker-compose.yml. Added ./backend/sources:/sources to the pfm-web-app service, recovered an orphaned entry via docker cp, and re-verified the save now persists to backend/sources/product_manual_labels.json on the host (confirmed the DO-flow's manual_labels.json save was fixed by the same change too) — shipped 2026-07-08.
  • 2.3 Ran the accuracy regression harness and discovered sources/accuracy_report.md was badly stale (claimed 75.04%; real current baseline is 95.10% overall, already at/above the 95% target — added a staleness banner to that file). Root-caused every remaining mismatch by pulling raw OCR text from Postgres (documents.layout_parsing_result): the worst field, plat (67.6%), is almost entirely the license-plate region being classified as an image/seal by the layout model rather than OCR'd as text — not fixable in parser.ts. Found and fixed one genuine parser logic bug along the way: the "global pattern scanning fallback" could duplicate an already-correctly-extracted noDO value into a still-missing noSO field; fixed by excluding already-assigned values from that fallback's candidate pool (pfm-web-app/src/utils/parser.ts). Doesn't change the aggregate score (a wrong value and "Not Found" score the same) but stops a fabricated-looking wrong number from silently reaching the database. All 48 parser unit tests still pass — shipped 2026-07-08.
  • Ad-hoc Built custom expiry-date-based auto-rotation algorithm in Python classifier server (classify_ocr_server.py). The algorithm calculates the slant angle of the Expiry Date / Batch text line bounding box, automatically rotates the image to make it horizontal, and re-runs YOLO classification + PaddleOCR for maximum accuracy. Enhanced SKU matching database lookup to prioritize exact SKU matches with a score of 1.0, pinning them as the Best Match — shipped 2026-07-09.
  • 2.5 Retrained the Product/SKU scan classifier's model artifacts against the full current dataset, which had grown to 81 SKU classes / 2,493 photos (up from the original 16 classes / 118 photos the deployed model dated 2026-07-08 was actually trained on — the other 65 classes had photos but no trained weights). Rebuilt models/dinov2_index.pkl (now 2,493/2,493 photos indexed) and retrained the YOLO classifier 100 epochs on an RTX 2060 (real elapsed time 54m21s), publishing models/produk-pfm-classifier-26n-100e-2026-07-14.pt/.onnx at 85.8% top-1 / 94.4% top-5 validation accuracy across all 81 classes (up from 83.3%/90% on the old 16-class model). Along the way, fixed a real train/val split bug in train_classifier.py: split_dataset() previously shuffled and split individual image files, letting an augmented copy (photo_aug_2.jpeg) land in validation while its near-duplicate source stayed in training — inflating val accuracy with memorization instead of measuring generalization; now groups by source photo (stripping _aug_N) before shuffling and splitting 80/20. Verified via docker compose up -d pipeline-api + docker logs: "DINOv2 index loaded with 2493 reference images", "Using classifier weights: .../produk-pfm-classifier-26n-100e-2026-07-14.pt", "YOLO model loaded successfully" — the live service is confirmed serving the new 81-class model, not assumed from the newest-file-by-date fallback logic. Remaining gap toward the program's ±230-SKU target is dataset growth, not a pipeline limitation — shipped 2026-07-14.

Backend — Postgres Data Layer

  • 3.1 Wrapped the ocr_items delete-then-reinsert in /api/parse and /api/v1/documents/[id] PUT inside a DB transaction (withTransaction helper, pfm-web-app/src/db/index.ts) — a mid-loop insert failure now rolls back to the previous item set instead of leaving a document with a correct header but partial/missing items — shipped 2026-07-08. (Renumbered from root's 7.1 when this file split from root docs/feature-list.md.)
  • 3.2 Hashed accounts.password with bcryptjs (pure-JS, no native compile step — the pfm-web-app Docker image has no build toolchain). db/init.ts hashes the seed and idempotently migrates any pre-existing plaintext rows on every startup; api/v1/auth/login/route.ts now compares with bcrypt.compareSync and cleanly rejects missing credentials with a 401 instead of risking a raw-query edge case. Verified via psql (hash format) and curl (correct login succeeds, wrong/missing password returns 401) — shipped 2026-07-08.

Docs & Workflow Integrity

  • 5.1 Fixed stale doc claims in SKILLS.md (accuracy baseline pointer) and CLAUDE.md (Flutter auth claim and API base URL fallback) — shipped 2026-07-08.
  • 5.2 Refactored plans/next-enhancements.md to archive verbose [DONE] and [CANCELLED] task bodies into one-line stubs. Reduced the file size significantly, strictly enforcing the 256-line threshold rule for maintainability — shipped 2026-07-08.
  • 5.3 Amended AGENTS.md completion checklist with a doc-sync step to ensure architecture changes are synced back to documentation — shipped 2026-07-08.

Product Scan — Ground Truth Annotation & Accuracy

  • 6.1 Built standalone annotation page manual-label-scan/page.tsx for ground truth editing. Includes image browser, editable fields (no_sku, nama_item, expiry_date, notes), and a "Scan with AI" fill-blanks feature — shipped 2026-07-08.
  • 6.2 API + storage groundwork for scan annotation. Extended api/manual-label-scan with GET list mode and DELETE. Persisted uploaded scan photos as base64 images into sources/product-test-images/. Made the scan-pfm quick-save honest by allowing manual correction before save — shipped 2026-07-08.
  • 6.3 Built backend/scripts/accuracy-check-scan.mts mirroring the DO-harness architecture, measuring overall match rate plus per-field breakdown (no_sku, expiry_date) against the new stable labels — shipped 2026-07-08.
  • 6.4 Ported the DO-harness's auto-diff-vs-previous-run reporting into accuracy-check-scan.mts: every run now prints a Δ column per field per split (Training/Validation) vs the last product_accuracy_history.jsonl entry, and calls out field- and image-level regressions/improvements explicitly. Added classifier method (dinov2_similarity/yolo_classifier) distribution and average confidence as informational (non-scoring) context. Created the previously-missing sources/product-test-images/README.md documenting the validation-photo drop workflow — shipped 2026-07-13, user-directed n request to make algorithm tuning self-verifying.

Master Data Management

  • 8.1 & 8.3 CRUD APIs and Web UI: Created /api/v1/master/stores and /api/v1/master/skus endpoints alongside a Next.js Admin page (/admin/master-data) to visually manage the core reference data used by the OCR matching engine — shipped 2026-07-08.
  • 8.2 Auto-Provisioning Store Accounts: Store creation now automatically securely hashes a default password ("123") and creates a paired login account, keeping store configuration perfectly in sync with the accounts table — shipped 2026-07-08.

Auth — Store Accounts & Profile-Sourced Metadata

  • 7.1 Seeded one account per store in db/init.ts during initialization by assigning username = kode_toko and a bcrypt-hashed default password "123". Included role and is_active schema additions — shipped 2026-07-08.
  • 7.2 Enhanced authentication routing by modifying POST /api/v1/auth/login to perform a LEFT JOIN on store_master, returning the extended store profile alongside the token. Added a guard to reject login if is_active = false. Implemented a new GET /api/v1/auth/me endpoint to cleanly re-fetch the profile via token — shipped 2026-07-08.
  • 7.3 Created a reproducible store_master bootstrap logic in db/init.ts that reads from sources/toko_aktif.json idempotently on startup. Also correctly seeded the WH_JOFFICE head office to resolve the admin account foreign-key setup constraint — shipped 2026-07-08.

Backend — Document Confirmation Gate & Data Hygiene

  • 10.1 Added a confirmed BOOLEAN NOT NULL DEFAULT TRUE column to documents (db/init.ts, ALTER TABLE ... ADD COLUMN IF NOT EXISTS — grandfathers every pre-existing row so today's history didn't go empty after migration) and used it to separate "OCR finished" from "user confirmed": previously GET /api/v1/documents filtered only on parsed = true, which the backend sets synchronously right after upload — before the mobile user ever taps "Simpan & Konfirmasi" in the editor — so a scan captured, previewed, then backed out of (never confirmed) was already sitting in every entitled account's document list with blank/placeholder fields (root cause of document_card.dart's "Staff Toko" fallback text on the Flutter side). v1/documents/upload/route.ts now explicitly inserts confirmed = false on every new upload; v1/documents/[id]/route.ts's PUT handler is the only place that flips it to true (literally "the user confirmed"); v1/documents/route.ts (list) now filters AND confirmed = true unconditionally for every account including admin (no role special-casing, per explicit user decision); v1/documents/[id]/route.ts's GET-by-id handler is deliberately untouched by the new filter so the mobile poller can keep seeing pending/unconfirmed documents mid-flow. utils/document-mapper.ts's shared DocumentRow/mapDocumentRow() now carries confirmed through to all three call sites (list, GET-by-id, upload's dedup-hit branch) from one place. parse/route.ts's own INSERT ... ON CONFLICT (filename) DO UPDATE statements (both DO and Product branches) were deliberately left untouched for confirmed — in the real mobile flow the upload route's INSERT always runs first, so this upsert always hits the ON CONFLICT branch, and since its SET clause doesn't mention confirmed, Postgres correctly leaves the existing value alone (verified this is correct, not an oversight). Verified live against the running Docker stack: uploaded a real DO photo as a store account without confirming it — absent from that store's list (and from admin's) while GET /documents/:id still reported the correct parseStatus; PUT (confirm) made it appear immediately with the real submitted data; all 13 pre-existing rows carried confirmed = true after the migration ran — shipped 2026-07-10.
  • 10.2 Removed the fabricated Product Scan placeholder values noPO: "PO-PRODUCT-001", noSO: "1002003004", noDO: "DO-PRODUCT-999" (both the flat keys and the mirrored header.no_po/no_so/no_do sub-object) from parse/route.ts's Product-scan branch, replacing them with empty strings — these are DO-specific concepts that don't apply to a product verification scan, and were never actually read by anything: pdf_service.dart's Product receipt branch never prints them, and product_editor_submit_logic.dart's _submit() builds its own noPo/noSo/noDo from the user's PO-link dropdown and batch selection, ignoring the stored values entirely. Same class of issue as the earlier G7 fix (fabricated data presented as if real) — low risk to remove since nothing meaningfully depended on the old values. Scope stayed narrow to exactly these three fields; nama_driver/nama_penerima's "PRODUCT SCAN"/"STORE STAFF" placeholders were left alone as a deliberate fixed convention, not a fabricated document number. Verified via curl: a freshly-uploaded, unconfirmed Product Scan document's raw GET /documents/:id response now returns no_po/no_so/no_do as empty strings instead of the old fake values — shipped 2026-07-10.

Backend — Single-Pass Product Classification

  • 11.1 Eliminated the duplicate GPU classification pass on Product Scan (gap G3), sourced from user feedback that the review screen took noticeably longer to open than DO Scan's. api/parse/route.ts's Product branch previously had its own separate, poorer inline classify call (kept only top1_name/extracted_sku), forcing the Flutter editor to re-run the entire classify+OCR pipeline a second time via POST /api/v1/scan-product just to get the top-5 candidate list and OCR-extracted expiry date. Now calls the same shared classifyAndMatchProduct() (utils/product-scan.ts) already used by that v1 route — one GPU call, richer result — and persists it under a new metadata.productScan JSONB key (no schema migration), surfaced by document-mapper.ts as a top-level productScan field on every GET response. Caught and fixed a real regression along the way: delegating to the shared function silently dropped the 90s pipeline timeout the old inline fetch had; added the same bound (PIPELINE_TIMEOUT_MS) directly inside classifyAndMatchProduct() so both callers — this route and the live POST /api/v1/scan-product (which never had the bound either) — are protected. Verified via curl with a genuinely fresh image/store combination (proving a real classify pass, not a dedup hit): took 9s, and the immediate GET /documents/:id response already contained 5 real possibleMatches and the extracted expiry date, before any editor interaction — shipped 2026-07-10.