No content changes: git diff --ignore-all-space over these files is empty. The churn came from editing on Windows against a repo checked out with LF.
23 KiB
23 KiB
Feature List (backend)
Structured log of shipped backend features, updated by the n/next workflow (see
AGENTS.md Part B) whenever a task in
plans/next-enhancements.md is marked [DONE].
Split out 2026-07-08 from root docs/feature-list.md's backend sections — this file
is the sole home for backend feature history going forward.
Format
## <Section / Module Name>
- **<task number>** <feature description> — shipped <date>
Existing Features (pre-kit)
Backfilled 2026-07-08 during adoption of this kit — these predate the e/n
workflow and have no task numbers; see git log for real dates/history.
Backend — Next.js API Gateway
- Upload/parse/documents CRUD routes, GPU status endpoint, vLLM proxy, manual-label review tool.
Backend — OCR Pipeline & Accuracy
- PaddleOCR + vLLM classification pipeline with DB layout caching, table column-shift correction, date normalization, and an accuracy regression harness (
pfm-web-app/scripts/accuracy-check.mts) — 95.10% overall as of 2026-07-08 (target 95% met; see task 2.3 below for the investigation andCLAUDE.md).
Backend — Postgres Data Layer
- Schema/init in
pfm-web-app/src/db/init.ts, served via the canonical rootdocker-compose.ymlstack. - 3.3 Added a standard
INDEXondocuments(file_hash)indb/init.tsto accelerate the upload deduplication queries without strictly enforcing uniqueness across different stores. Correspondingly updated the dedup query inapi/v1/documents/upload/route.tsto scope duplicate detection bykode_toko. This fixes a conflict where one store could be incorrectly linked to another store's duplicate receipt image — shipped 2026-07-08.
DevOps — Docker & Dev Tunnel
- 4.1 Docker Compose Policy Documented: Formalized the execution policy in
README.mdandCLAUDE.md, explicitly requiring the use of thedocker-compose.demo.ymloverride (production build) for all client demonstrations and field testing to bypass the Next.js dev server bottleneck — shipped 2026-07-08. - Docker Compose Dependency Gates: Added strict Docker
healthcheckgates (Task 4.2) blocking thepfm-web-app(Next.js) from starting until PostgreSQL and the VLLM models are initialized and fully healthy. - Secure Tunnel Ingress: Restructured
nginx.confandstart-dev-tunnel.ps1(Tasks 4.3, 4.5) to expose a dedicated, restricted port (8001) that exclusively routes to/api/v1/*. This perfectly secures the development UI (/scan-pfm) and legacy routes from public exposure. - Dead Config Pruning: Stripped deprecated and redundant proxy blocks from the Nginx edge router (
Task 4.4).
(New features shipped via n/next go below, organized the same way, with task numbers.)
Backend — Next.js API Gateway
- 1.4 Enforced real 401 auth on
/api/v1/documents/*(list, PUT-by-id, upload) — the actual production API surface, already fully supported by the Flutter client (real login +Authorization: Beareron every request). Previously none of these three routes rejected a missing/invalid token; upload only optionally read it. Added the pre-existinggetAccountFromAuthHeader()helper (utils/auth.ts) + a 401 guard to all three;OPTIONS(CORS preflight) untouched. The original task 1.3 (auth on the classic routes) was cancelled instead — those routes are dev-only web UI surface with no login flow, going away in production. Verified viacurl: 401 with no token, success with a real token from/api/v1/auth/login— shipped 2026-07-08. - 1.5 Implemented per-store data scoping on
/api/v1/documents/*. Addedkode_tokocolumn todocumentstable viadb/init.tsmigration. The upload route now bindskode_tokoto documents upon creation.GET /api/v1/documentsandPUT /api/v1/documents/:idenforce ownership checks (kode_tokomatching) forstorerole accounts, whileadminretains global access including legacy unassigned documents — shipped 2026-07-08. - 1.6
GET /api/v1/healthEndpoint: Unauthenticated health probe verifying both PostgreSQL connectivity and Pipeline API HTTP reachability. ReturnsHTTP 503if any core dependency is down — shipped 2026-07-08. - Ad-hoc Connected
/scan-pfmpage with/api/scan-pfmroute and enabled auto-trigger scanning on custom file upload, sample selection, thumbnail change, and canvas rotation. Supported bothimageandimage_base64payload keys — shipped 2026-07-09. - 9.1 Added
GET /api/v1/documents/:id(same 401/403 scoping asPUT), returning a single document — including still-unparsed rows — with a newparseStatus: "pending"|"done"|"failed"field, so the Flutter poller can move off scanning the entire list every 2s. Addedscan_mode/parse_errorcolumns todocuments(db/init.ts, migrated viaALTER TABLE ... ADD COLUMN IF NOT EXISTSfor already-running DBs).scan_modeis now persisted on upload (v1/documents/upload/route.ts) and on the classic/api/parseroute's upserts (COALESCE, same pattern askode_toko), and surfaced asdocTypeon every GET response (utils/document-mapper.ts, a new shared helper extracted from the list route's inline mapping so list/by-id/dedup all agree) — falling back to the legacyorder_untuk == "PRODUCT SCAN"sentinel for pre-existing rows with noscan_mode.parse_erroris now recorded when the upload route's internal call to/api/parseitself fails to complete (network error or the 210s abort firing) — previously this was silently swallowed and the document stayedparsed=falseforever with no signal, burning the client's full 260s timeout;/api/parse's own existing pipeline-error fallback (parsed=true+ "Not Found" placeholder) was already fine and is unchanged. Also fixed the dedup branch (a repeat upload of an already-seen file) to return the original document's real current state via the same mapper instead of a hardcoded empty stub. Verified viadocker compose up -d --build+curl: schema migration applied cleanly to the live DB (confirmed viapsql), DO and Product uploads both correctly persistscan_modeand surface it asdocType, a dedup retry returns real header/items instead of an empty stub,GET /:idreturns 401 (no token) / 403 (wrong store) / 404 (nonexistent id) / 200 (admin or owning store), and the list endpoint's existing filter/scoping is unchanged — shipped 2026-07-10. - 9.3 Added authenticated
POST /api/v1/scan-product, the v1 equivalent of the classic dev-only/api/scan-pfm(unauthenticated, and unreachable off-LAN since task 4.5 restricted the public tunnel to/api/v1/*). Extracted the shared classify-and-match logic (Python classifier call + Levenshtein SKU matching againstsku_master, top-5 scoring) out ofapi/scan-pfm/route.tsinto a newutils/product-scan.ts(classifyAndMatchProduct, plus aClassifierErrorclass that preserves forwarding the classifier's own HTTP status instead of collapsing every failure to 500) so the classic route and the new v1 route share one implementation instead of duplicating it — the classic route's response shape, auth-free behavior, and desktop-only layout-parsing visualization are otherwise unchanged. The new route accepts either multipart (image/filefield, matching the v1 upload route's convention) or a JSON{image_base64}body, is open to any authenticated account (not admin-gated, since this is what the mobile app itself calls), and wraps the result in the standard{status, data}envelope withclassification,ocr(includingextracted_expired_date), andpossibleMatches. Verified viacurlagainst the live stack with a real product photo: multipart upload and JSON-body variants both return identical, correct top-5 matches; no-token request returns 401; the classic/api/scan-pfmroute's response (includinglayoutParsingResult) is unchanged post-refactor — shipped 2026-07-10. - 9.2 Relaxed
GET /api/v1/master/skus(master/skus/route.ts) so any authenticated account can read the SKU master list, not justadmin— the Flutter product editor needs this and previously had to string-hack its base URL to call the unauthenticated classicGET /api/skus, which task 4.5 had already removed from the public tunnel, breaking product scans off-LAN. Changed the guard from a combined!account || role !== 'admin'check (403 for both "no token" and "wrong role") to!account(correct 401) followed by an unconditional pass-through for any valid account;POST(SKU creation) is untouched, still admin-only, per the user's explicit choice between the two options this task flagged as undecided. No response-shape change. Verified viacurlagainst the live stack with a real non-admin (storerole) account's token:GET→ 200 with real data; no token → 401 (was incorrectly 403 before this fix); the same non-admin token againstPOST→ still 403; adminGET→ still 200. Along the way, hit and resolved a dev-loop issue: the container had the edited file on disk but Turbopack's file watcher wasn't detecting the change over the Windows bind mount, requiringdocker restart paddleocr-pfm-web-appto pick it up — noted in case it recurs for future edits. With 9.1-9.3 all shipped, Flutter root task 7.1 (moving the product editor onto the v1 surface) is now fully unblocked — shipped 2026-07-10.
Backend — OCR Pipeline & Accuracy
- 2.1 Built the Product/SKU scan classifier's model artifacts:
models/dinov2_index.pkl(118/118 reference photos indexed across 16 SKU classes) andmodels/produk-pfm-classifier-26n-100e-2026-07-08.pt(+.onnxexport) — a YOLO classifier fine-tuned 100 epochs, 83.3% top-1 / 90% top-5 validation accuracy on the current (thin, 2-16 photos/class) dataset. Built via a one-offdocker runfrom a freshly-rebuiltpipeline-apiimage (bare-metal training isn't viable on Windows —paddlepaddle-gpu's wheel index is Linux-only).pipeline-apirestarted and confirmed loading both models from logs. Also fixedscripts/install-pipeline.sh, which was missingultralytics/torch— shipped 2026-07-08. - 2.1 (verification pass) Ran a full browser walkthrough of
/scan-pfm(classification, top-5, OCR expiry extraction + crop, SKU-master matching, Visual/Spotting Grid, Raw Response — all confirmed working with real data). Found and fixed a real bug: "Save Ground Truth" was returning success but silently writing into thepfm-web-appcontainer's ephemeral filesystem instead of the host, because/sourceswasn't a bind-mounted path in rootdocker-compose.yml. Added./backend/sources:/sourcesto thepfm-web-appservice, recovered an orphaned entry viadocker cp, and re-verified the save now persists tobackend/sources/product_manual_labels.jsonon the host (confirmed the DO-flow'smanual_labels.jsonsave was fixed by the same change too) — shipped 2026-07-08. - 2.3 Ran the accuracy regression harness and discovered
sources/accuracy_report.mdwas badly stale (claimed 75.04%; real current baseline is 95.10% overall, already at/above the 95% target — added a staleness banner to that file). Root-caused every remaining mismatch by pulling raw OCR text from Postgres (documents.layout_parsing_result): the worst field,plat(67.6%), is almost entirely the license-plate region being classified as an image/seal by the layout model rather than OCR'd as text — not fixable inparser.ts. Found and fixed one genuine parser logic bug along the way: the "global pattern scanning fallback" could duplicate an already-correctly-extractednoDOvalue into a still-missingnoSOfield; fixed by excluding already-assigned values from that fallback's candidate pool (pfm-web-app/src/utils/parser.ts). Doesn't change the aggregate score (a wrong value and "Not Found" score the same) but stops a fabricated-looking wrong number from silently reaching the database. All 48 parser unit tests still pass — shipped 2026-07-08. - Ad-hoc Built custom expiry-date-based auto-rotation algorithm in Python classifier server (
classify_ocr_server.py). The algorithm calculates the slant angle of the Expiry Date / Batch text line bounding box, automatically rotates the image to make it horizontal, and re-runs YOLO classification + PaddleOCR for maximum accuracy. Enhanced SKU matching database lookup to prioritize exact SKU matches with a score of 1.0, pinning them as the Best Match — shipped 2026-07-09. - 2.5 Retrained the Product/SKU scan classifier's model artifacts against the full current dataset, which had grown to 81 SKU classes / 2,493 photos (up from the original 16 classes / 118 photos the deployed model dated 2026-07-08 was actually trained on — the other 65 classes had photos but no trained weights). Rebuilt
models/dinov2_index.pkl(now 2,493/2,493 photos indexed) and retrained the YOLO classifier 100 epochs on an RTX 2060 (real elapsed time 54m21s), publishingmodels/produk-pfm-classifier-26n-100e-2026-07-14.pt/.onnxat 85.8% top-1 / 94.4% top-5 validation accuracy across all 81 classes (up from 83.3%/90% on the old 16-class model). Along the way, fixed a real train/val split bug intrain_classifier.py:split_dataset()previously shuffled and split individual image files, letting an augmented copy (photo_aug_2.jpeg) land in validation while its near-duplicate source stayed in training — inflating val accuracy with memorization instead of measuring generalization; now groups by source photo (stripping_aug_N) before shuffling and splitting 80/20. Verified viadocker compose up -d pipeline-api+docker logs: "DINOv2 index loaded with 2493 reference images", "Using classifier weights: .../produk-pfm-classifier-26n-100e-2026-07-14.pt", "YOLO model loaded successfully" — the live service is confirmed serving the new 81-class model, not assumed from the newest-file-by-date fallback logic. Remaining gap toward the program's ±230-SKU target is dataset growth, not a pipeline limitation — shipped 2026-07-14.
Backend — Postgres Data Layer
- 3.1 Wrapped the
ocr_itemsdelete-then-reinsert in/api/parseand/api/v1/documents/[id]PUT inside a DB transaction (withTransactionhelper,pfm-web-app/src/db/index.ts) — a mid-loop insert failure now rolls back to the previous item set instead of leaving a document with a correct header but partial/missing items — shipped 2026-07-08. (Renumbered from root's7.1when this file split from rootdocs/feature-list.md.) - 3.2 Hashed
accounts.passwordwithbcryptjs(pure-JS, no native compile step — thepfm-web-appDocker image has no build toolchain).db/init.tshashes the seed and idempotently migrates any pre-existing plaintext rows on every startup;api/v1/auth/login/route.tsnow compares withbcrypt.compareSyncand cleanly rejects missing credentials with a 401 instead of risking a raw-query edge case. Verified viapsql(hash format) andcurl(correct login succeeds, wrong/missing password returns 401) — shipped 2026-07-08.
Docs & Workflow Integrity
- 5.1 Fixed stale doc claims in
SKILLS.md(accuracy baseline pointer) andCLAUDE.md(Flutter auth claim and API base URL fallback) — shipped 2026-07-08. - 5.2 Refactored
plans/next-enhancements.mdto archive verbose[DONE]and[CANCELLED]task bodies into one-line stubs. Reduced the file size significantly, strictly enforcing the 256-line threshold rule for maintainability — shipped 2026-07-08. - 5.3 Amended
AGENTS.mdcompletion checklist with a doc-sync step to ensure architecture changes are synced back to documentation — shipped 2026-07-08.
Product Scan — Ground Truth Annotation & Accuracy
- 6.1 Built standalone annotation page
manual-label-scan/page.tsxfor ground truth editing. Includes image browser, editable fields (no_sku,nama_item,expiry_date,notes), and a "Scan with AI" fill-blanks feature — shipped 2026-07-08. - 6.2 API + storage groundwork for scan annotation. Extended
api/manual-label-scanwithGETlist mode andDELETE. Persisted uploaded scan photos as base64 images intosources/product-test-images/. Made thescan-pfmquick-save honest by allowing manual correction before save — shipped 2026-07-08. - 6.3 Built
backend/scripts/accuracy-check-scan.mtsmirroring the DO-harness architecture, measuring overall match rate plus per-field breakdown (no_sku,expiry_date) against the new stable labels — shipped 2026-07-08. - 6.4 Ported the DO-harness's auto-diff-vs-previous-run reporting into
accuracy-check-scan.mts: every run now prints a Δ column per field per split (Training/Validation) vs the lastproduct_accuracy_history.jsonlentry, and calls out field- and image-level regressions/improvements explicitly. Added classifier method (dinov2_similarity/yolo_classifier) distribution and average confidence as informational (non-scoring) context. Created the previously-missingsources/product-test-images/README.mddocumenting the validation-photo drop workflow — shipped 2026-07-13, user-directednrequest to make algorithm tuning self-verifying.
Master Data Management
- 8.1 & 8.3 CRUD APIs and Web UI: Created
/api/v1/master/storesand/api/v1/master/skusendpoints alongside a Next.js Admin page (/admin/master-data) to visually manage the core reference data used by the OCR matching engine — shipped 2026-07-08. - 8.2 Auto-Provisioning Store Accounts: Store creation now automatically securely hashes a default password ("123") and creates a paired login account, keeping store configuration perfectly in sync with the
accountstable — shipped 2026-07-08.
Auth — Store Accounts & Profile-Sourced Metadata
- 7.1 Seeded one account per store in
db/init.tsduring initialization by assigningusername = kode_tokoand a bcrypt-hashed default password"123". Includedroleandis_activeschema additions — shipped 2026-07-08. - 7.2 Enhanced authentication routing by modifying
POST /api/v1/auth/loginto perform aLEFT JOINonstore_master, returning the extended store profile alongside the token. Added a guard to reject login ifis_active = false. Implemented a newGET /api/v1/auth/meendpoint to cleanly re-fetch the profile via token — shipped 2026-07-08. - 7.3 Created a reproducible
store_masterbootstrap logic indb/init.tsthat reads fromsources/toko_aktif.jsonidempotently on startup. Also correctly seeded theWH_JOFFICEhead office to resolve the admin account foreign-key setup constraint — shipped 2026-07-08.
Backend — Document Confirmation Gate & Data Hygiene
- 10.1 Added a
confirmed BOOLEAN NOT NULL DEFAULT TRUEcolumn todocuments(db/init.ts,ALTER TABLE ... ADD COLUMN IF NOT EXISTS— grandfathers every pre-existing row so today's history didn't go empty after migration) and used it to separate "OCR finished" from "user confirmed": previouslyGET /api/v1/documentsfiltered only onparsed = true, which the backend sets synchronously right after upload — before the mobile user ever taps "Simpan & Konfirmasi" in the editor — so a scan captured, previewed, then backed out of (never confirmed) was already sitting in every entitled account's document list with blank/placeholder fields (root cause ofdocument_card.dart's "Staff Toko" fallback text on the Flutter side).v1/documents/upload/route.tsnow explicitly insertsconfirmed = falseon every new upload;v1/documents/[id]/route.ts'sPUThandler is the only place that flips it totrue(literally "the user confirmed");v1/documents/route.ts(list) now filtersAND confirmed = trueunconditionally for every account includingadmin(no role special-casing, per explicit user decision);v1/documents/[id]/route.ts'sGET-by-id handler is deliberately untouched by the new filter so the mobile poller can keep seeing pending/unconfirmed documents mid-flow.utils/document-mapper.ts's sharedDocumentRow/mapDocumentRow()now carriesconfirmedthrough to all three call sites (list, GET-by-id, upload's dedup-hit branch) from one place.parse/route.ts's ownINSERT ... ON CONFLICT (filename) DO UPDATEstatements (both DO and Product branches) were deliberately left untouched forconfirmed— in the real mobile flow the upload route's INSERT always runs first, so this upsert always hits theON CONFLICTbranch, and since itsSETclause doesn't mentionconfirmed, Postgres correctly leaves the existing value alone (verified this is correct, not an oversight). Verified live against the running Docker stack: uploaded a real DO photo as a store account without confirming it — absent from that store's list (and fromadmin's) whileGET /documents/:idstill reported the correctparseStatus;PUT(confirm) made it appear immediately with the real submitted data; all 13 pre-existing rows carriedconfirmed = trueafter the migration ran — shipped 2026-07-10. - 10.2 Removed the fabricated Product Scan placeholder values
noPO: "PO-PRODUCT-001",noSO: "1002003004",noDO: "DO-PRODUCT-999"(both the flat keys and the mirroredheader.no_po/no_so/no_dosub-object) fromparse/route.ts's Product-scan branch, replacing them with empty strings — these are DO-specific concepts that don't apply to a product verification scan, and were never actually read by anything:pdf_service.dart's Product receipt branch never prints them, andproduct_editor_submit_logic.dart's_submit()builds its ownnoPo/noSo/noDofrom the user's PO-link dropdown and batch selection, ignoring the stored values entirely. Same class of issue as the earlier G7 fix (fabricated data presented as if real) — low risk to remove since nothing meaningfully depended on the old values. Scope stayed narrow to exactly these three fields;nama_driver/nama_penerima's "PRODUCT SCAN"/"STORE STAFF" placeholders were left alone as a deliberate fixed convention, not a fabricated document number. Verified viacurl: a freshly-uploaded, unconfirmed Product Scan document's rawGET /documents/:idresponse now returnsno_po/no_so/no_doas empty strings instead of the old fake values — shipped 2026-07-10.
Backend — Single-Pass Product Classification
- 11.1 Eliminated the duplicate GPU classification pass on Product Scan (gap G3), sourced from user feedback that the review screen took noticeably longer to open than DO Scan's.
api/parse/route.ts's Product branch previously had its own separate, poorer inline classify call (kept onlytop1_name/extracted_sku), forcing the Flutter editor to re-run the entire classify+OCR pipeline a second time viaPOST /api/v1/scan-productjust to get the top-5 candidate list and OCR-extracted expiry date. Now calls the same sharedclassifyAndMatchProduct()(utils/product-scan.ts) already used by that v1 route — one GPU call, richer result — and persists it under a newmetadata.productScanJSONB key (no schema migration), surfaced bydocument-mapper.tsas a top-levelproductScanfield on every GET response. Caught and fixed a real regression along the way: delegating to the shared function silently dropped the 90s pipeline timeout the old inline fetch had; added the same bound (PIPELINE_TIMEOUT_MS) directly insideclassifyAndMatchProduct()so both callers — this route and the livePOST /api/v1/scan-product(which never had the bound either) — are protected. Verified viacurlwith a genuinely fresh image/store combination (proving a real classify pass, not a dedup hit): took 9s, and the immediateGET /documents/:idresponse already contained 5 realpossibleMatchesand the extracted expiry date, before any editor interaction — shipped 2026-07-10.