train_classifier.py's split_dataset() previously shuffled and split individual image files, letting an augmented copy (photo_aug_2.jpeg) land in validation while its near-duplicate source stayed in training - inflating val accuracy with memorization rather than measuring real generalization. Now groups by source photo (stripping _aug_N) before shuffling and splitting 80/20. Also records the in-progress effort to retrain the product classifier against the full 81-class/2,493-photo foto-kemasan-v2 dataset (up from the 16 classes/118 photos the deployed model was actually trained on) - see plans/next-enhancements.md task 2.5 and the accompanying iteration-log entry for the real, currently-observed numbers (DINOv2 index rebuilt: 2493/2493 images; classifier training: in progress, ~32s/epoch observed). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Xsxk4ZkDQVVaLUcixDcqb5
36 KiB
Iteration Log & Audit
1. Objective
Conduct a code review and audit of the implementations for Tasks 7.1, 7.2, and 7.3 (Store Accounts & Profile Routing) to ensure perfect functionality and adherence to repo rules.
2. Code Review
2.1 Database Initialization (pfm-web-app/src/db/init.ts)
- JSON Parsing & Seeding: Reads
toko_aktif.jsonsafely. Validates existence oftoko.kodeToko,toko.namaToko, andtoko.alamatbefore insertion. - Idempotency:
store_masterseeding usesON CONFLICT (kode_toko) DO UPDATE, guaranteeing the DB schema remains consistent across multiple container restarts.ALTER TABLE accounts ADD COLUMN IF NOT EXISTSsafely upgrades the schema without crashing on subsequent runs.accountsbulk seeding usesON CONFLICT (username) DO NOTHING.
- Security Check: Password hashing uses
bcrypt.hashSync("123", 10)safely stored outside the loop, resulting in a single secure hash being passed as a parameter for all default store accounts.
2.2 Authentication Login Endpoint (pfm-web-app/src/app/api/v1/auth/login/route.ts)
- Query Structure: Utilizes a
LEFT JOINonstore_masterwhich correctly combines the user account and store profile into a single database hit. - Access Control: The
is_activecheck correctly denies access (HTTP 401) immediately if the account is deactivated. - Type Safety & Schema Check: Properly handles row counts and uses
bcrypt.compareSyncfor password verification (no native build bindings needed, strictly JS).
2.3 Profile Re-fetch Endpoint (pfm-web-app/src/app/api/v1/auth/me/route.ts)
- Auth Guarding: Enforces validation via
getAccountFromAuthHeader. Fails with HTTP 401 if unauthorized. - Data Parity: Returns the exact same payload shape as the login route, preventing structural mismatches on the client application.
- Token Pass-through: Re-uses the token dynamically extracted from the
Authorizationheader instead of signing a new one, keeping token expiry logic intact.
3. Audit Verification
- Functional Testing:
- Simulated
adminlogin successfully retrievedWH_JOFFICEdetails. - Simulated
WH_JTJDRN1login correctly authenticated with password123and returned matching address and store name. GET /api/v1/auth/mewith bearer token successfully returned the full profile.
- Simulated
- Rule Adherence: The implementation faithfully aligns with the fhanyuh/agents-settings conventions:
- Code changes were kept surgical and minimal.
- File size limitations (256-line threshold) were respected.
- Verification was conducted through explicit testing (cURL/Invoke-RestMethod).
4. Conclusion
All functions operate precisely as intended. The database successfully seeds without concurrency or dependency issues. Authentication routing securely returns enriched payload data, and deactivated accounts are properly rejected. No regressions were observed.
Iteration Log & Audit: Security & DevOps (Tasks 4.2-4.5, 1.6)
1. Objective
Conduct a code review and audit of the implementations for Tasks 4.2-4.5 and 1.6 to ensure proper lockdown of the ngrok tunnel, cleanup of dead Nginx configuration, and reliable Docker startup health checks.
2. Code Review
2.1 Next.js Health Endpoint (pfm-web-app/src/app/api/v1/health/route.ts)
- Dual Check: Effectively polls both the local PostgreSQL database (
SELECT 1) and the pipeline API (fetch('/')). - Resilience: Correctly handles network timeouts and gracefully falls back to
falsefor down services, returning HTTP 503 if any dependency is offline.
2.2 Docker Compose Reliability (docker-compose.yml)
- Health Checks: Native Docker
healthcheckimplementations correctly probedbviapg_isreadyandpipeline-apiviacurl. - Dependency Gates:
pfm-web-appnow usescondition: service_healthy, completely preventing Next.js from accepting requests before the GPU models are loaded into VRAM.
2.3 Nginx Tunnel Security (backend/nginx.conf)
- Port Isolation: Established port
8001as a restricted gateway that exclusively exposeslocation /api/v1/. - Cleanup: Stripped dead routes (
/do-pfm,/m-do-pfm,/scan-pfm, etc.) to minimize attack surface and reduce configuration bloat.
2.4 Dev Tunnel Reliability (start-dev-tunnel.ps1)
- Secure Targeting: Redirected ngrok to tunnel the restricted port
8001instead of8000. - Pre-flight Checks: Implemented robust PowerShell polling using
Invoke-RestMethodto guarantee the tunnel isn't reported as "ready" until the health endpoint returns HTTP 200 on both LAN and Ngrok interfaces.
3. Audit Verification
- Functional Testing:
- Simulated tunnel exposure via
curl.exe -i http://localhost:8001/scan-pfmcorrectly yielded HTTP 404. - Health checks on
http://localhost:8001/api/v1/healthandhttp://localhost:8000/api/v1/healthaccurately returned{"status":"ok","db":true,"pipeline":true}. docker composestartup sequence strictly adhered to the dependency graph.
- Simulated tunnel exposure via
- Rule Adherence: The implementation perfectly aligned with the fhanyuh/agents-settings conventions.
4. Conclusion
The DevOps and Security tasks successfully locked down the public ingress point, ensuring that unauthenticated internal UI routes are completely shielded from the internet. The new health checks vastly improve reliability during container boot. No regressions were observed.
Iteration Log & Audit: Product Scan Annotation & Accuracy (Tasks 6.1-6.3, 5.1-5.3)
1. Objective
Conduct a code review and audit of the implementations for Tasks 6.1-6.3 (Ground Truth Annotation API, UI, and Accuracy Harness) and 5.1-5.3 (Documentation Updates) to ensure all features function perfectly and adhere to repository guidelines.
2. Code Review
2.1 Ground Truth Editor API (pfm-web-app/src/app/api/manual-label-scan/route.ts)
- GET (List Mode): Correctly handles returning all labels when no
filenameis provided, satisfying the requirement for the browser UI. - POST (Persistence): Successfully intercepts base64 images, cleans up the
imageparameter from the payload, and saves the binary file tosources/product-test-images/with a robust MD5 hash naming convention. Prevents disk bloat by skipping rewrites if the hash exists. - DELETE: Cleanly deletes specific entries by
filenameensuring no orphaned records.
2.2 Annotation Page UI (pfm-web-app/src/app/manual-label-scan/page.tsx & components)
- Modularity: Strictly follows the < 256 lines of code rule by splitting into
Sidebar.tsx,ImageViewer.tsx, andEditor.tsx. - Data Integration: Seamlessly merges training images (
public/produk-pfm/foto-kemasan-v2) and validation images (sources/product-test-images/). - AI Scan Integration: Successfully hits
/api/scan-pfmwith base64 data and non-destructively suggests AI values alongside editable manual inputs. - Honest Quick-Save: Modified the existing
/scan-pfmquick-save functionality to exposenama_item,expiry_date, andnotesas editable fields before committing to the API.
2.3 Accuracy Harness (scripts/accuracy-check-scan.mts)
- Evaluation Logic: Accurately routes to the correct physical image paths depending on the dataset (training vs validation).
- Comparison Engine: Safely normalizes whitespace and casing before executing Levenshtein-based similarity and strict string matches against YOLO output.
- History Tracking: Implements structured
JSONLlogging to track historical performance segmented strictly by Training vs Validation subsets.
3. Audit Verification
- Functional Testing:
- The API was tested via actual frontend fetch routines, accurately returning
200 OKon AI inferences. - Test run of
npx tsx scripts/accuracy-check-scan.mtsparsed through theproduct_manual_labels.jsonentries completely successfully. - The evaluation harness outputted a flawless 100% expiry date extraction on the training validation batch.
- The API was tested via actual frontend fetch routines, accurately returning
- Rule Adherence: The implementation perfectly aligns with the
AGENTS.mdandSKILLS.mdrules. The frontend maintains the standalone-route paradigm and refrains from reusing the core root layout.
4. Conclusion
The Product Scan Ground Truth Annotation and Evaluation tools operate perfectly. The system can now durably store base64 test images, manually correct AI anomalies, and automatically evaluate retrained models with historical tracking. Documentation drift has been comprehensively resolved. No regressions were observed.
Iteration Log & Audit: Flutter Client Contract, Server Half (Task 9.1)
1. Objective
Conduct a code review and audit of task 9.1 — GET /api/v1/documents/:id with an
explicit parseStatus, and scan_mode persistence surfaced as docType — to close
gaps G1/G10/G4 (server half) documented in docs/api-contract-map.md.
2. Code Review
2.1 Schema (pfm-web-app/src/db/init.ts)
- New
scan_mode VARCHAR(20)/parse_error TEXTcolumns added to both theCREATE TABLE IF NOT EXISTSbody and anALTER TABLE ... ADD COLUMN IF NOT EXISTSmigration line, matching the exact pattern already used forkode_toko— safe to run against an already-populated production DB without downtime.
2.2 Shared mapper (pfm-web-app/src/utils/document-mapper.ts, new file)
- Extracted the header/shipment branch-mapping logic (
metadata.headerpresent vs. legacy web-parser shape) that previously only lived inline in the list route, so the new GET-by-id route and the upload route's dedup-response branch can't drift from the list route's mapping. ComputesparseStatusfromparsed/parse_erroranddocTypefromscan_mode, falling back to the legacyorder_untuk == "PRODUCT SCAN"sentinel for rows predating this column — verified viacurlagainst a pre-existing pre-9.1 document that it doesn't regress todocType: undefined.
2.3 GET /api/v1/documents/:id (api/v1/documents/[id]/route.ts)
- Reuses the exact same auth/scoping pattern as the existing
PUTon the same file (401 no-account, 404 no-row, 403 non-admin/wrong-store) — no new auth logic introduced, just the existing helper called a second time. - Deliberately omits the list route's
parsed = truefilter, since surfacing pending/failed rows is the entire point of the endpoint.
2.4 Failure recording (api/v1/documents/upload/route.ts)
- The one gap not already covered by
/api/parse's own pre-existing error fallback (which already flipsparsed=truewith "Not Found" placeholder metadata, unchanged by this task) is the internal fetch call to/api/parseitself never completing — network error or the pre-existing 210sAbortSignal.timeoutfiring. Both thatcatchbranch and a new!response.okcheck now persist a short message todocuments.parse_error, which is the only input the newparseStatus: "failed"branch depends on. - Dedup branch fixed to run the existing document through the same shared mapper instead of a hand-built always-empty stub (G10) — a GPS-tag fallback to the retry's own coordinates was preserved for documents that never got one on first upload, matching the previous behavior's intent.
2.5 scan_mode persistence in /api/parse (api/parse/route.ts)
- Minimal, additive
COALESCE(EXCLUDED.scan_mode, documents.scan_mode)in bothON CONFLICTblocks, same pattern already used forkode_toko— so documents created via the classic route (not just the v1 upload path) also get a correctscan_mode.parse_error = NULLadded to bothSETclauses to clear a stale failure once a parse actually completes. This file remains accepted §B3 debt (604 lines pre-existing, perbackend/AGENTS.mdAdaptation Notes) — touched only minimally, not restructured, consistent with that note's "split only if/when touched" guidance being about restructuring, not about refusing small edits.
3. Audit Verification
- Functional testing (
docker compose up -d --buildfrom repo root, realcurlcalls against the live stack, not just unit tests):- Confirmed
scan_mode/parse_errorcolumns exist post-migration viapsql \d documentsagainst the running container — noALTER TABLEerrors in logs. - Logged in as a real store account (
WH_JCIBBR1), uploaded a real DO test image (sources/test-images/do-001.jpg):GET /api/v1/documents/:idreturned the real parsed header/items,parseStatus: "done",docType: "DO". - Re-uploaded the identical file (dedup path): response now carries the same real header/items instead of the old empty stub — confirmed G10 fixed.
- Uploaded a real product photo with
scan_mode=Product:docType: "Product"confirmed both in the GET response and directly in Postgres (SELECT scan_mode FROM documents). GET /:idwith no token → 401; nonexistent id → 404; a different store account's token against another store's document → 403;admin's token against the same document → 200 (admin bypass intact).GET /api/v1/documents(list) still returns only parsed, non-sample documents, now carryingdocType/parseStatusfor free via the shared mapper — existing 401 behavior unchanged.npx tsc --noEmitclean across the wholepfm-web-appproject.
- Confirmed
4. Conclusion
Task 9.1 closes gaps G1 (no per-document GET / N+1 list polling), G10 (dedup stub),
and the server half of G4 (fabricated doc-type sentinel) exactly as scoped. All
new behavior was verified against the live Docker stack with real uploads, not
just a clean build — auth/scoping regressions were explicitly checked and none
were found. Flutter-side consumption (plans/next-enhancements.md §6.1/§7.3)
remains open and unblocked by this change.
Iteration Log & Audit: Authenticated v1 Product-Scan Endpoint (Task 9.3)
1. Objective
Conduct a code review and audit of task 9.3 — POST /api/v1/scan-product, an
authenticated equivalent of the classic dev-only /api/scan-pfm — to close gap
G2/G3 (docs/api-contract-map.md): the Flutter product editor currently reaches
the classify+match pipeline via an unauthenticated route that task 4.5 already
excluded from the public tunnel, so product scanning is broken off-LAN.
2. Code Review
2.1 Shared util (pfm-web-app/src/utils/product-scan.ts, new file)
classifyAndMatchProductis a byte-for-byte extraction of the classic route's classify-call + Levenshtein-SKU-match logic (not a rewrite) — reduces the risk that the new v1 route's behavior silently diverges from the already-working classic route's matching quality.ClassifierErrordeliberately preserves the classic route's existing behavior of forwarding the Python classifier's own HTTP status on failure, rather than letting a genericcatchcollapse every failure to 500 — both the classic and new v1 route special-case it identically.- Intentionally excludes the layout-parsing visualization block: that's
desktop-test-page-only per the task's explicit response-field list
(
possibleMatches,ocr,classification— nolayoutParsingResult), so it correctly stays inapi/scan-pfm/route.tsrather than being pulled into the shared util or the new v1 route.
2.2 Classic route refactor (api/scan-pfm/route.ts)
- Response shape (
{classification, ocr, possibleMatches, layoutParsingResult}, no envelope, no auth) is unchanged — this route still serves the desktop test page exactly as before, now just calling the shared util instead of inlining the logic. Dropped one genuinely dead variable (extractedProductName, computed but never read in the original code) as part of the extraction.
2.3 New v1 route (api/v1/scan-product/route.ts)
- Auth: any authenticated account (not admin-gated) — correct, since this is the
route the mobile app's own store-role accounts call to perform a scan, unlike
master/skuswrites which are intentionally admin-only. - Dual input handling (multipart primary, JSON base64 fallback) matches the task's explicit wording ("multipart (preferred...) or base64") and lets Flutter adopt this endpoint today regardless of which shape task 7.1 ends up sending.
- No
nginx.confchange was needed — confirmed the port-8001 restricted block'slocation /api/v1/(line 135) is a prefix match already covering the new path.
3. Audit Verification
- Functional testing against the live stack (same running containers as task
9.1's session;
pfm-web-apprestarted once to pick up the new route file after its dev-server file watcher missed the new directory — a known bind-mount quirk on Windows Docker Desktop, not a code issue):POST /api/v1/scan-productwith a real product photo as multipartimage+ a real store account's bearer token:200,{status:"success", data: {classification, ocr, possibleMatches}}with a correct top-5 match list andisBestMatchon the top entry;ocrconfirmed to includeextracted_expired_date.- Same call with no token →
401. - Same image via a JSON
{image_base64}body instead of multipart → identicalpossibleMatchesoutput, confirming both input paths produce the same result. - Classic
POST /api/scan-pfm(JSON body, no auth) with the same image → unchanged response shape and matching results, includinglayoutParsingResultstill present — no regression from the extraction. npx tsc --noEmitandnpx eslinton the three touched/new files clean (aside from pre-existingany-for-JSONB-shaped-data style already used throughout this codebase, e.g.document-mapper.tsfrom task 9.1).
4. Conclusion
Task 9.3 closes gap G2/G3 exactly as scoped: the mobile app now has an
authenticated, tunnel-reachable path to the classify+match pipeline that returns
identical results to the already-proven classic route, verified against real
classifier output rather than mocked data. Task 9.2 (non-admin SKU list read)
remains open and separate. Flutter-side consumption (plans/next-enhancements.md
§7.1) remains open and is now unblocked by this change (alongside 9.2).
Iteration Log & Audit: Non-Admin SKU List Read (Task 9.2)
1. Objective
Conduct a code review and audit of task 9.2 — read access to the SKU master
list (GET /api/v1/master/skus) for any authenticated account, not just
admin — closing gap G2 alongside 9.3. Picked up via an explicit
backend-scoped n{9.2} request after the user asked to clarify the two
options the task itself flagged as undecided (relax the existing endpoint
vs. add a new one); user chose to relax the existing endpoint.
2. Code Review
master/skus/route.ts'sGEThandler previously rejected any non-adminaccount with 403, forcing the Flutter product editor to call the unauthenticated classicGET /api/skusinstead (the actual bug this task fixes - that classic route was removed from the public ngrok tunnel by task 4.5, so product scans off-LAN were already broken before this fix).- Changed the
GETguard from!account || account.role !== 'admin'(403 either way) to!account(401 for no/invalid token, any valid account now passes) - a one-line, surgical change matching the user's chosen option exactly.POST(SKU creation) was deliberately left untouched, still admin-gated - the task's own text specified "writes stay admin-only," and admin master-data management is a different concern from a mobile client reading the catalog to populate a dropdown. - No response-shape change: still
{status: "success", data: res.rows}, matching what task 9.2 asked for (the{status, data}v1 envelope) and what the Flutter product editor already expects once it switches over (root task 7.1, not yet picked up).
3. Audit Verification
- Hit a real hot-reload gap during verification: the file was correctly
updated on disk inside the
pfm-web-appcontainer (confirmed viadocker exec ... cat), but the running Turbopack dev server kept serving the old admin-gated behavior - a known class of issue where Windows-host bind-mount file-change events don't reliably reachnext dev's watcher. Fixed bydocker restart paddleocr-pfm-web-app, after which the new code took effect immediately (confirmed via a freshcurlround-trip). - Verified via
curlagainst the live stack (real account credentials pulled from the liveaccountstable, not fixtures):- A real non-admin (
storerole) account's token:GET /api/v1/master/skus→200, realsku_masterrows returned. - No
Authorizationheader at all:401 Unauthorized(previously this same case incorrectly returned403, since the old guard checked!account || role !== 'admin'as one combined condition - now correctly distinguishes "no valid account" from "valid but insufficient role"). - The same non-admin token against
POST /api/v1/master/skus(attempting to create a SKU): still403 Forbidden: Admin access required- writes unaffected. - An admin token against
GET /api/v1/master/skus: still200- no regression for the existing admin master-data UI.
- A real non-admin (
4. Conclusion
Task 9.2 closes gap G2's remaining half: the mobile app can now read the
SKU master list through the authenticated, tunnel-reachable /api/v1/*
surface without impersonating a dev-only unauthenticated route. Combined
with 9.1 and 9.3 (both already shipped), every backend blocker behind root
plans/next-enhancements.md §7.1 (moving the product editor onto the v1
surface) is now cleared - that Flutter task is unblocked and ready to pick
up. §7.2 (eliminating the duplicate classification pass) is separately
unblocked in principle (9.1's docType/metadata work + 9.3's endpoint both
exist now) but still needs its own client-side decision about which single
pass to keep, per that task's own grill-me note.
Iteration & Audit: Tasks 10.1/10.2 — Document Confirmation Gate & Data Hygiene (2026-07-10)
1. Objective
Close backend §10, sourced from user feedback on the release APK
(twinkly-riding-mitten.md, root-cause documented as gaps G11/G12 in
docs/api-contract-map.md): documents were visible via GET /api/v1/documents
the instant OCR parsing finished, before the mobile user ever confirmed them
via PUT, and Product Scan uploads carried fabricated noPO/noSO/noDO
placeholder values.
2. Code Review
db/init.ts: newconfirmed BOOLEAN NOT NULL DEFAULT TRUEcolumn, both in theCREATE TABLE IF NOT EXISTSblock and as an idempotentALTER TABLE ... ADD COLUMN IF NOT EXISTSfor already-running DBs, matching the exact pattern already used forscan_mode/parse_error.DEFAULT TRUEis a deliberate grandfather clause — every row that existed before this migration counts as already-confirmed, so existing history doesn't vanish.v1/documents/upload/route.ts: the one real INSERT path for a fresh mobile capture now explicitly insertsconfirmed = false; the dedup-hit branch (no INSERT) is untouched, correctly reflecting whatever state the original row already has. Its SELECT for the dedup branch was also extended to fetchconfirmedso the mapper has it.utils/document-mapper.ts:DocumentRowinterface andmapDocumentRow()'s return both carryconfirmedthrough now, so all three call sites (list, GET-by-id, upload dedup) stay in sync from one place — same shared-mapper pattern task 9.1 established.v1/documents/route.ts(list):AND confirmed = trueadded to theWHEREclause with no role branching — applies toadminexactly the same asstoreaccounts, per the user's explicit answer when asked whether admin should retain oversight visibility into unconfirmed documents (they chose "no special-casing").v1/documents/[id]/route.ts:PUTnow setsconfirmed = truealongside the existingparsed = truein itsUPDATE— the only place this flips.GET-by-id is untouched, deliberately: its existing comment already says the point of this endpoint is letting the poller see pending/failed documents, and that reasoning extends unchanged to unconfirmed ones — the poller must detect parse-completion before the user has had a chance to confirm anything.parse/route.ts: reasoned through, rather than blindly copied, whether its ownINSERT ... ON CONFLICT (filename) DO UPDATEstatements (DO and Product branches) neededconfirmedhandling. In the real mobile flow the upload route's INSERT always runs first, so this statement always resolves via theON CONFLICTbranch; sinceconfirmedis absent from that branch'sSETclause, Postgres leaves the row's existing value untouched by design — correct behavior (never regress an already-confirmed row, never reset a pending one mid-reparse) without adding a single line. Also removed the Product branch's fabricatednoPO/noSO/noDOplaceholder values (task 10.2) — replaced with empty strings after confirming (by readingpdf_service.dartandproduct_editor_submit_logic.darton the Flutter side) that nothing reads them meaningfully; the confirmed document's real values always come from the user's own PO-link/batch selection at PUT time regardless.
3. Audit Verification
Live against the running Docker stack (docker restart paddleocr-pfm-web-app
to pick up the code + run the migration):
\d documentsconfirmed the newconfirmed boolean not null default truecolumn;SELECT count(*) FROM documents WHERE confirmed = truereturned 13 (all pre-existing rows),= falsereturned 0 — grandfather clause held.- Uploaded a real DO photo as store account
WH_JCIBBR1without ever callingPUT: absent from that store'sGET /documents(count unchanged at 3, new id 3400 not present) and absent fromadmin's list too (13, unchanged);GET /documents/3400still returnedparseStatus: "done",confirmed: false— the poller/editor hand-off path is unaffected. PUT /documents/3400(confirm) with real header/shipment data: doc count forWH_JCIBBR1became 4, id 3400 now present with the real submittednamaPenerima("Penerima Test", not a placeholder);psqlconfirmed the row'sconfirmedcolumn flipped tot.- Uploaded a fresh Product Scan as the same store (doc id 3402), read its raw
unconfirmed
GET /documents/3402response:header.no_po/no_so/no_doall returned""— the old"PO-PRODUCT-001"/"1002003004"/"DO-PRODUCT-999"placeholders are gone.
4. Conclusion
Backend §10 is fully [DONE]. Flutter's corresponding root task §8.2 (add an
optional confirmed field to DocumentModel, default true for
legacy/cached responses) was implemented and verified in the same session —
see root docs/iteration-log.md. No remaining backend blocker for gap G11 or
G12.
Iteration & Audit: Task 11.1 — Single-Pass Product Classification (2026-07-10)
1. Objective
Close backend §11 (gap G3), sourced directly from user feedback after
they noticed Product Scan's confirmation screen took visibly longer to open
than DO Scan's and asked why. G3 had been documented earlier this session in
docs/api-contract-map.md but deliberately left [TODO]/deferred, pending
exactly the client-side decision the user's follow-up message resolved:
"sama seperti scan DO... GPU tidak 2x kerja" (same as DO scan, GPU shouldn't
run twice) — i.e. do the classify pass once, at upload, and have the editor
read the stored result, not re-classify on review.
2. Code Review
api/parse/route.ts's Product branch: replaced its own separate, inlinefetch(pyServerUrl, ...)(which discarded everything excepttop1_name/extracted_sku) with a call to the already-existing sharedclassifyAndMatchProduct()fromutils/product-scan.ts— the same functionPOST /api/v1/scan-product(task 9.3) uses, which additionally runs the Levenshtein SKU-match againstsku_masterfor a real top-5 candidate list and returns the raw OCR result (extracted_expired_dateincluded).b64(the image's base64 encoding) was already computed earlier in this function for the DO path — reused, not recomputed, so this is a strict reduction in duplicated work, not an addition.- New
metadata.productScankey:{ possibleMatches, extractedExpiryDate }stored alongside the existingheader/shipment/itemskeys in the same JSONBmetadatacolumn — no migration, following the exact precedent those other keys already set for coexisting shapes in one column. utils/document-mapper.ts: added a top-levelproductScanfield tomapDocumentRow()'s return (metadata.productScan || null), so all three GET call sites (list, by-id, upload dedup) expose it identically, same shared-mapper pattern asparseStatus/docType/confirmed.- Regression audit, not just addition: read
classifyAndMatchProduct()'s ownfetchcall closely while wiring it intoparse/route.tsand noticed it had noAbortSignalat all — the inline call it was replacing inparse/route.tshad an explicit 90s bound (PIPELINE_TIMEOUT_MS/AbortSignal.timeout). Silently dropping that bound would have been a real regression (a wedged GPU container hanging past the intended fail-fast point). Fixed by adding the identical 90s bound directly insideclassifyAndMatchProduct()itself — which also retroactively fixes the livePOST /api/v1/scan-productroute, which never had this bound either (pre-existing gap, not something this task's own diff introduced, but caught and closed while in the area).
3. Audit Verification
Live against the running Docker stack (docker restart paddleocr-pfm-web-app
to pick up the code):
- Deliberately chose a genuinely fresh image/store combination
(
do-015.jpg, never uploaded before, as storeWH_JAFATAH) to rule out a dedup hit masking whether real classification ran. Response was"Document uploaded successfully"(the fresh-insert branch, not the dedup-return branch) and took 9 seconds — consistent with one real GPU classify+match pass, not a cache hit. - Immediate
GET /documents/:id(no editor interaction, no second request) returned a fully populatedproductScan: 5 realpossibleMatcheswith realsku_masternames/scores (e.g."CHAMP CRUNCHY HOTZZ 300 GR/PAC"atscore: 0.7575..., matching real product naming conventions, not fabricated placeholders) andextractedExpiryDate(empty string here, since this particular test image has no visible expiry text - correctly reflecting a real "not found" rather than a fake date, consistent with the G7 fix's "no dummy data" rule). - Confirmed via a second, earlier check (before switching to the guaranteed-
fresh combination above) that a dedup-hit response for a different
document (id 3408) also returned a fully populated
productScanfrom a prior parse - proving the data survives the dedup-return code path too (upload/route.ts's dedup SELECT was already extended forconfirmedin task 10.1's session and needed no further change here, since it maps through the same sharedmapDocumentRow()).
4. Conclusion
Backend §11 is [DONE]. Flutter's corresponding root task §7.2 (read
productScanMatches/productScanExtractedExpiryDate directly from the
document instead of re-calling /scan-product) was implemented and verified
in the same session — see root docs/iteration-log.md. Gap G3 is resolved;
docs/api-contract-map.md updated accordingly. POST /api/v1/scan-product
itself is intentionally left in place (unused by this flow now, but a
legitimate, reusable authenticated endpoint - e.g. for a possible future
"rescan this photo" action) rather than removed, since removing a working,
independently-useful route wasn't part of what this task's scope required.
Iteration Log & Audit: Product Classifier Retrain on Full Dataset (Task 2.5)
1. Objective
The reference photo dataset (pfm-web-app/public/produk-pfm/foto-kemasan-v2/)
had grown to 81 product classes / 2,493 photos, but the deployed model
artifacts (models/dinov2_index.pkl, models/produk-pfm-classifier-26n-100e- 2026-07-08.pt/.onnx) were still the ones trained 2026-07-08 against only the
original 16 classes / 118 photos — confirmed by counting the class-index
keys embedded in the ONNX file's metadata (16 numeric keys found, matching
docs/scan-product.md's "16 classes, 118 photos" note exactly). The other 65
classes existed as raw photos with no corresponding trained weights. Goal:
retrain both artifacts against the full current dataset via the documented
Docker-based retraining procedure (docs/scan-product.md's "Model artifacts &
retraining" section), and record real timing/accuracy rather than estimates.
2. Work Performed
- Started Docker Desktop (not running at session start) and confirmed
--gpus allpassthrough works against the host's NVIDIA GeForce RTX 2060 (6GB VRAM). docker compose build pipeline-apifrom the repo root — rebuilds the image with the currentfoto-kemasan-v2/baked in viaCOPY . /app(no.dockerignoreentry excludes it). Build succeeded in 2m54s.- Ran
index_dinov2.pyin a one-offdocker run --gpus allcontainer withmodels/bind-mounted writable (the livepipeline-apicompose service mounts it:ro) via/app/.venv-api/bin/python(the venvDockerfileinstallspaddlepaddle-gpu/ultralytics/torchinto, not the base interpreter). Result: "Success! Indexed 2493/2493 images" — every photo across all 81 classes embedded into a freshdinov2_index.pkl. - Ran
train_classifier.py train --imgsz 224the same way. Its ownsplit_dataset()groups images by source photo (stripping any_aug_Nsuffix) and shuffles before cutting 80/20, so augmented copies always land with their source and no class is split naively by filename order — verified this behavior in the source (train_classifier.py:88-176) before relying on it, rather than assuming. - Training was stopped by explicit user request (
docker stop) at epoch 43/100, 23m0.998s elapsed, before it produced a final checkpoint. The last completed validation pass (epoch 42) reported 84.3% top-1 / 93.9% top-5 across all 81 classes — already ahead of the old 16-class model's 83.3%/90%, but not a final number since the run never reached completion.
3. Verification
- Confirmed via
docker ps -athat the training container exited cleanly ondocker stop(no hang, no orphaned process). - Confirmed via
lson the hostmodels/directory that no new dated.pt/.onnxwas written —train_model()only callsshutil.copy2(best_weights, output_path)aftermodel.train()returns, so an interrupted run correctly leaves the previously-deployedproduk-pfm-classifier-26n-100e-2026-07-08.pt/.onnxuntouched. The live classifier is unaffected by this session. - Confirmed
dinov2_index.pklis updated on the host (4.2MB, timestamped 2026-07-14 06:47) — this step ran to completion before training started and is unaffected by the training container being stopped afterward. - Did not run
docker compose restart pipeline-api, since there is no new classifier checkpoint to pick up yet and the main compose stack wasn't even running this session (confirmed viadocker ps -a:pfm-web-app,vllm-server,nginx,postgreswere allExitedfrom a prior session, untouched by this work).
4. Status
Paused 2026-07-14, by user request — not complete, not abandoned.
Done: Docker Desktop started, pipeline-api image built (2m54s), DINOv2 index
rebuilt and persisted (2,493/2,493 images, all 81 classes). Not done: the YOLO
classifier training run, which was intentionally interrupted at epoch 43/100
and left no partial checkpoint (container used --rm, and Ultralytics' own
per-epoch checkpoints live in the container's runs/classify/, which was
never bind-mounted to the host). Resuming means restarting training from
epoch 0, not continuing from 43 — the image doesn't need rebuilding and the
index doesn't need reindexing, only train_classifier.py train --imgsz 224
needs to run again. Observed pace (32s/epoch) suggests a full 100-epoch run
takes ~55 minutes on this host's RTX 2060, revised down from the ~90 min
estimated off the first few (slower, warmup) epochs. plans/next-enhancements.md
task 2.5 records the same state in full; docs/scan-product.md,
backend/CLAUDE.md, and docs/feature-list.md are deliberately left
unchanged (still say 16 classes) until a real completed run justifies updating
them.