Grilled 2026-07-16 with the user; full decision record in
docs/expiry-tracking-plan.md. Core reframe: expiry is captured once per
batch at DO intake (staff-typed on the stock-entry confirmation page, from
the physical packs), so the cashier scan only MATCHES OCR fragments against
the 1-3 known in-stock batch dates instead of free-reading damaged
dot-matrix prints (proven model-capability ceiling, 2026-07-15). Fallback:
auto-FEFO + 'inferred' flag, zero cashier interaction. No cloud, ever.
- docs/expiry-tracking-plan.md: architecture, matching algorithm spec
(resolveExpiryFromEvidence), schema/API deltas, phases 1-3, testing plan
- backend plans §13 (13.1-13.4): matcher util + offline tuning, route
wiring + expiry_source provenance, multi-frame union, dot-matrix
recognizer fine-tune
- root plans §10 (10.1-10.3): cashier fast path, inferred badge +
end-of-day review, burst capture for mounted camera
- stock-feature-plan.md: extension note (batch dropdown becomes the
manual-override path)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q8TumxFDnyVnfsR3mxPXfX
Split the expiry-date extraction cascade out of classify_ocr_server.py into
config/date_extract.py (pure regex, importable/testable without loading
models). Three behavioral fixes, offline-regressed against all 79 captured
OCR line-sets and sanity-verified live on the two target images:
- Guard the 012/112 month-misrecognition cleanup rules: they fired on
perfectly valid dates too (BB 01122026 = 01/12/2026 matches 0+112+2026)
and mangled them into 7-digit junk that parsed as 00/22/26. Skipped when
the line already contains a valid date. Fixes image 11.
- Exclude store price-tag lines (Printed:.., Rp...) from the keyword-less
stages so a shelf label's print timestamp can't shadow the real date
printed on the package. Fixes image 71 (09/04/2027).
- Validity-gate the lenient stage (day<=31, month<=12, year 2020-2039) so
garbled digit runs return empty instead of junk like 1/3/06 or 11/1/01.
Also: clamp /probe-ocr crop box to image bounds (PIL pads out-of-bounds
crops into a gigapixel canvas -> DecompressionBombError), and update
CLAUDE.md's Graphify section - the global Claude Code skill integration was
installed 2026-07-15 at the user's explicit request.
Full-batch measurement of these fixes (expected 79.7% -> ~80.6%) is still
pending - the run was stopped twice at the user's end; re-run
scripts/accuracy-check-scan.mts next session before building on this.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gr6HH7JrdsXX8AARejQboM
2026-07-15 re-run of the 79-image benchmark: the tiled full-res pass, VL
text-line evidence, and VL SKU retry deployed at end of 2026-07-14 produce
ZERO flips vs the 79.7% baseline (only img1's expiry pred changed to a
lenient-parse garble). Detail dump: product_scan_detail_20260715.json.
Probe endpoint POST :8120/probe-ocr (temporary, remove before production):
OCRs a crop of a container-local image through arbitrary preprocessing
recipes (scale/blur/close/CLAHE/threshold) and alternate readers - VL
layout-parsing with/without layout detection, direct VLM chat on :8118,
lazily-loaded PP-OCRv5 server det/rec. Probing the dot-matrix expiry crops
of img2/img41 with every combination shows none of the stack's models can
read these prints (VL tags them as pictures, VLM hallucinates, v5-server
skips them) - the 21 expiry non-detections are a model capability ceiling,
not a pipeline bug.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gr6HH7JrdsXX8AARejQboM
Accuracy work on the 79-image product-scan validation set (user goal: 90%):
- classify_ocr_server.py: 0/90/180/270-degree expiry-date search (stops at
first hit, 0-degree fallback); classification decoupled onto the upright
image (rotated frames regressed DINOv2 -6pts until this); cross-line date
stitching; tiled full-res OCR pass (defeats the 4000px downscale that
killed small inkjet dates); VL-pipeline expiry fallback with
keyword-anchored anti-hallucination guard; VL text lines merged into
text_lines + VL SKU retry. Visualization endpoints removed entirely
(Visual/Spotting grids - unused by frontend, 3x per-scan GPU cost).
- product-scan.ts: coverage-normalized OCR-evidence re-ranking of DINOv2
top-K (tuned offline: +8/-0 on top-1 misses), re-ranked class mapped to
sku_master by SKU prefix; classifier timeout 90s->240s for fallback paths.
- Frozen benchmark: product-test-images-fixed/ (79 renamed images) +
freeze/seed/build-undetected/capture/experiment scripts; labels trimmed to
the 79 validation entries (training rows kept in .bak-with-training);
5 TRAINED-ON SKUs replaced with fresh held-out photos.
- manual-label-scan page: shows last batch-test AI prediction under every
field by default (new /api/product-scan-results); serves the fixed folder;
fixed total hydration failure via allowedDevOrigins 127.0.0.1.
- Measured (all-79, zero failures): sku/name 87.3%, expiry 64.6%, overall
79.7%. Tiles/VL-evidence/VL-SKU deployed but not yet batch-measured.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gr6HH7JrdsXX8AARejQboM
train_classifier.py's split_dataset() previously shuffled and split
individual image files, letting an augmented copy (photo_aug_2.jpeg) land
in validation while its near-duplicate source stayed in training -
inflating val accuracy with memorization rather than measuring real
generalization. Now groups by source photo (stripping _aug_N) before
shuffling and splitting 80/20.
Also records the in-progress effort to retrain the product classifier
against the full 81-class/2,493-photo foto-kemasan-v2 dataset (up from the
16 classes/118 photos the deployed model was actually trained on) - see
plans/next-enhancements.md task 2.5 and the accompanying iteration-log
entry for the real, currently-observed numbers (DINOv2 index rebuilt:
2493/2493 images; classifier training: in progress, ~32s/epoch observed).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xsxk4ZkDQVVaLUcixDcqb5
Ports the DO-harness's auto-diff-vs-previous-run reporting into
accuracy-check-scan.mts: prints a per-field, per-split (Training/
Validation) delta against the last product_accuracy_history.jsonl entry
and calls out regressions/improvements explicitly, plus classifier
method distribution and average confidence as informational context.
Also adds the real held-out validation photo set into
sources/product-test-images/ (75 photos, one per current SKU class) with
its README documenting the drop-photo -> label -> re-run workflow, so the
harness's Validation Set split actually has images to score against.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xsxk4ZkDQVVaLUcixDcqb5
Adds tasks 1.4-8.5 across existing sections 1-8 (each traced to a specific
file/line, not invented busywork) plus a new section 9 (Stocks Menu &
DO-to-Stock flow) seeded from a user-directed, grilled ad-hoc feature
request - full design doc in docs/stock-feature-plan.md, status: planned,
no code yet.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xsxk4ZkDQVVaLUcixDcqb5
Records the AST-derived code-graph tooling (query/explain/affected commands
against graphify-out/graph.json) so future sessions know it's available
and why it's not wired in as a global skill/hook.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xsxk4ZkDQVVaLUcixDcqb5
The Flutter app is used by Prima Fresh Mart store staff (petugas toko) on
duty at each store to receive/confirm deliveries and log products - not by
the delivery drivers themselves. "Driver" only appears correctly now as a
data field on the DO document (who drove the delivery), never as the app's
operator. Fixes wording across README.md, CLAUDE.md, and
screenshots/v2/WORKFLOW.md; also corrects CLAUDE.md's stale "up to 2 min"
poll-timing note to match the real ~4.3min/130-retry value.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JSgYVUWJTGMk7SZqqHH8xy
- Auth section described a hardcoded admin/password stub; it's now real
JWT + bcrypt on /api/v1/* (verified against the actual route handlers).
- Known Limitations listed the pending-queue and non-transactional item
writes as still-open bugs; both were fixed 2026-07-08 - marked resolved.
- Corrected the poll-loop timing (2s x 130 retries = ~4.3min, not "2 min").
- Noted the confirmation-gate step now in the PUT flow, and that Product
Scan classification is DINOv2-primary (YOLO fallback), single-pass at
upload time - not the separate LAN-dependent call it used to be.
- Linked screenshots/v2/WORKFLOW.md + the new slide deck as the visual
app walkthrough reference.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JSgYVUWJTGMk7SZqqHH8xy
Captures the full DO Scan and Product Scan flows end-to-end against the live
backend (real OCR extraction, real DINOv2 product classification, a genuine
network-timeout failure, and a live validation-gap finding) for use in
presentations. WORKFLOW.md documents every screen/button/algorithm, and
Prima-Mart-Scanner-Workflow.pptx turns it into a 23-slide deck.
Also updates AppConfig's hardcoded LAN IP (192.168.70.4 -> 192.168.100.20)
to match the current dev machine's address, discovered while reproducing the
upload flow against the real backend.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JSgYVUWJTGMk7SZqqHH8xy
Fixes reported from APK field testing: DO/Product scan mode was inconsistent
between the camera drawer and documents screen (now one shared provider,
with an orange/green color cue); unconfirmed scans leaked into history with
placeholder data before the user tapped confirm (backend now gates
GET /documents on a new `confirmed` column, flipped only by PUT); and
Product Scan ran the GPU classifier twice, once at upload and again on
review (now a single pass at upload, persisted and read directly by the
editor). Also removes the unused "Hubungkan ke PO" field and fabricated
PO/SO/DO placeholder values from the Product Scan flow, closes out the
per-document-polling and save-recovery tasks (6.1/6.3), and splits several
touched files to stay under the repo's 256-line guideline.
Full detail in docs/iteration-log.md and backend/docs/iteration-log.md.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Backend (app-pfm-ocr-v2/backend):
- Product/SKU scan feature complete: trained DINOv2 index (118 reference
photos, 16 SKU classes) and YOLO classifier (83.3% top-1 val accuracy),
fixed scripts/install-pipeline.sh (was missing ultralytics/torch), fully
browser-verified end-to-end on /scan-pfm. Mobile m-scan-pfm page cancelled
(Flutter app handles mobile; web UI is desktop-only for pipeline testing).
- Fixed a real data-loss bug: Save Ground Truth (scan-pfm and the DO-flow's
manual-label) was silently writing into the pfm-web-app container's
ephemeral filesystem instead of the host, because /sources wasn't
bind-mounted in docker-compose.yml. Added the mount, recovered an
orphaned entry.
- accounts.password is now bcrypt-hashed (bcryptjs, idempotent migration
in db/init.ts) instead of plaintext; login route compares hashes.
- /api/v1/documents/* (list, PUT, upload) now enforces real 401 auth,
matching what the Flutter client already sends. The "classic" routes
deliberately stay open — they're dev-only web UI with no login flow and
won't exist in production.
- OCR accuracy investigated end-to-end: real baseline is 95.10% overall
(target met; accuracy_report.md was stale at 75.04%, now flagged). Fixed
one genuine parser.ts bug (SO/DO field duplication in the global fallback
regex); remaining gaps are OCR/layout-model limitations, not parser bugs.
- Adopted a standalone copy of the fhanyuh/agents-settings e/n workflow
scoped to backend/ (AGENTS.md Part A/B split, SKILLS.md, plans/, docs/),
independent of the root copy which now covers Flutter only.
- next-implementation.md deleted; content folded into
backend/plans/next-enhancements.md for traceability.
Root:
- Adopted fhanyuh/agents-settings kit (AGENTS.md, SKILLS.md, plans/,
docs/feature-list.md), scoped to the Flutter app only.
- Pending documents queue now persists to Hive (lib/core/storage) instead
of memory-only, surviving an app kill mid-upload.
Removed backend_backup/ (stale Express/Prisma prototype, superseded by
pfm-web-app) and the completed plans/next-enhancement-plan.md checklist.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Documents the 89.4% overall accuracy (up from 82.2% baseline, 95% target),
grounded in the actual logged run in accuracy_history.jsonl rather than a
vague claim, plus which fields the sanitize/triple-check layer rescues over
raw regex. Notes that current gaps are partly a photo-capture SOP floor
(DO paper photographed on top of other documents confuses auto-deskew),
not just a parsing/OCR ceiling.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eRAsLqN9Sg1b9YPz22Lxz
Replaces stale setup instructions (single hardcoded API URL, old migration
story) with the dual-mode ngrok/LAN config, the demo/production compose
override, and a "what to consider" section grounded in real issues hit this
session: the duplicate backend/docker-compose.yml project-name collision,
store_master shipping unseeded, and the debug-signed APK.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eRAsLqN9Sg1b9YPz22Lxz
Reliability/PoC hardening: dedupe uploads by file_hash, surface editor sync
failures instead of a false success SnackBar with a retry-without-re-OCR path,
bound the OCR pipeline fetches with timeouts, share a single ApiClient/Dio
instance app-wide, tune capture JPEG quality, and add an opt-in
docker-compose.demo.yml for a production-mode run ahead of client demos.
Translate all Flutter-side user-facing text (screens, validators, SnackBars,
the printed delivery receipt, and shared API error messages) to Bahasa
Indonesia.
Also includes in-progress OCR parser/accuracy-tuning work from the same
session: table column/unit normalization fixes, store/customer master data,
accuracy history log, and test-image renaming/cleanup.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eRAsLqN9Sg1b9YPz22Lxz