Force-added: these paths are gitignored, so they stay ignored for new files
unless added the same way. Committed on request so the working data is not
lost during the migration off this machine.
train_classifier.py's split_dataset() previously shuffled and split
individual image files, letting an augmented copy (photo_aug_2.jpeg) land
in validation while its near-duplicate source stayed in training -
inflating val accuracy with memorization rather than measuring real
generalization. Now groups by source photo (stripping _aug_N) before
shuffling and splitting 80/20.
Also records the in-progress effort to retrain the product classifier
against the full 81-class/2,493-photo foto-kemasan-v2 dataset (up from the
16 classes/118 photos the deployed model was actually trained on) - see
plans/next-enhancements.md task 2.5 and the accompanying iteration-log
entry for the real, currently-observed numbers (DINOv2 index rebuilt:
2493/2493 images; classifier training: in progress, ~32s/epoch observed).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xsxk4ZkDQVVaLUcixDcqb5
Backend (app-pfm-ocr-v2/backend):
- Product/SKU scan feature complete: trained DINOv2 index (118 reference
photos, 16 SKU classes) and YOLO classifier (83.3% top-1 val accuracy),
fixed scripts/install-pipeline.sh (was missing ultralytics/torch), fully
browser-verified end-to-end on /scan-pfm. Mobile m-scan-pfm page cancelled
(Flutter app handles mobile; web UI is desktop-only for pipeline testing).
- Fixed a real data-loss bug: Save Ground Truth (scan-pfm and the DO-flow's
manual-label) was silently writing into the pfm-web-app container's
ephemeral filesystem instead of the host, because /sources wasn't
bind-mounted in docker-compose.yml. Added the mount, recovered an
orphaned entry.
- accounts.password is now bcrypt-hashed (bcryptjs, idempotent migration
in db/init.ts) instead of plaintext; login route compares hashes.
- /api/v1/documents/* (list, PUT, upload) now enforces real 401 auth,
matching what the Flutter client already sends. The "classic" routes
deliberately stay open — they're dev-only web UI with no login flow and
won't exist in production.
- OCR accuracy investigated end-to-end: real baseline is 95.10% overall
(target met; accuracy_report.md was stale at 75.04%, now flagged). Fixed
one genuine parser.ts bug (SO/DO field duplication in the global fallback
regex); remaining gaps are OCR/layout-model limitations, not parser bugs.
- Adopted a standalone copy of the fhanyuh/agents-settings e/n workflow
scoped to backend/ (AGENTS.md Part A/B split, SKILLS.md, plans/, docs/),
independent of the root copy which now covers Flutter only.
- next-implementation.md deleted; content folded into
backend/plans/next-enhancements.md for traceability.
Root:
- Adopted fhanyuh/agents-settings kit (AGENTS.md, SKILLS.md, plans/,
docs/feature-list.md), scoped to the Flutter app only.
- Pending documents queue now persists to Hive (lib/core/storage) instead
of memory-only, surviving an app kill mid-upload.
Removed backend_backup/ (stale Express/Prisma prototype, superseded by
pfm-web-app) and the completed plans/next-enhancement-plan.md checklist.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Reliability/PoC hardening: dedupe uploads by file_hash, surface editor sync
failures instead of a false success SnackBar with a retry-without-re-OCR path,
bound the OCR pipeline fetches with timeouts, share a single ApiClient/Dio
instance app-wide, tune capture JPEG quality, and add an opt-in
docker-compose.demo.yml for a production-mode run ahead of client demos.
Translate all Flutter-side user-facing text (screens, validators, SnackBars,
the printed delivery receipt, and shared API error messages) to Bahasa
Indonesia.
Also includes in-progress OCR parser/accuracy-tuning work from the same
session: table column/unit normalization fixes, store/customer master data,
accuracy history log, and test-image renaming/cleanup.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eRAsLqN9Sg1b9YPz22Lxz