Commit Graph
20 Commits
Author SHA1 Message Date
fhanyuh caf8e98378 chore: normalize line endings (CRLF -> LF)
No content changes: git diff --ignore-all-space over these files is empty.
The churn came from editing on Windows against a repo checked out with LF.
2026-08-27 10:40:49 +07:00
Rafhan Mazaya FathurrahmanandClaude Fable 5 48334ded8f chore(backend): measure tiles+VL-evidence (zero net change, 79.7% holds) + add /probe-ocr debug endpoint
2026-07-15 re-run of the 79-image benchmark: the tiled full-res pass, VL
text-line evidence, and VL SKU retry deployed at end of 2026-07-14 produce
ZERO flips vs the 79.7% baseline (only img1's expiry pred changed to a
lenient-parse garble). Detail dump: product_scan_detail_20260715.json.

Probe endpoint POST :8120/probe-ocr (temporary, remove before production):
OCRs a crop of a container-local image through arbitrary preprocessing
recipes (scale/blur/close/CLAHE/threshold) and alternate readers - VL
layout-parsing with/without layout detection, direct VLM chat on :8118,
lazily-loaded PP-OCRv5 server det/rec. Probing the dot-matrix expiry crops
of img2/img41 with every combination shows none of the stack's models can
read these prints (VL tags them as pictures, VLM hallucinates, v5-server
skips them) - the 21 expiry non-detections are a model capability ceiling,
not a pipeline bug.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gr6HH7JrdsXX8AARejQboM
2026-07-14 21:45:44 +07:00
Rafhan Mazaya FathurrahmanandClaude Fable 5 e76ccb60a6 feat(backend): scan-product accuracy 66.2% -> 79.7% + frozen validation benchmark
Accuracy work on the 79-image product-scan validation set (user goal: 90%):
- classify_ocr_server.py: 0/90/180/270-degree expiry-date search (stops at
  first hit, 0-degree fallback); classification decoupled onto the upright
  image (rotated frames regressed DINOv2 -6pts until this); cross-line date
  stitching; tiled full-res OCR pass (defeats the 4000px downscale that
  killed small inkjet dates); VL-pipeline expiry fallback with
  keyword-anchored anti-hallucination guard; VL text lines merged into
  text_lines + VL SKU retry. Visualization endpoints removed entirely
  (Visual/Spotting grids - unused by frontend, 3x per-scan GPU cost).
- product-scan.ts: coverage-normalized OCR-evidence re-ranking of DINOv2
  top-K (tuned offline: +8/-0 on top-1 misses), re-ranked class mapped to
  sku_master by SKU prefix; classifier timeout 90s->240s for fallback paths.
- Frozen benchmark: product-test-images-fixed/ (79 renamed images) +
  freeze/seed/build-undetected/capture/experiment scripts; labels trimmed to
  the 79 validation entries (training rows kept in .bak-with-training);
  5 TRAINED-ON SKUs replaced with fresh held-out photos.
- manual-label-scan page: shows last batch-test AI prediction under every
  field by default (new /api/product-scan-results); serves the fixed folder;
  fixed total hydration failure via allowedDevOrigins 127.0.0.1.
- Measured (all-79, zero failures): sku/name 87.3%, expiry 64.6%, overall
  79.7%. Tiles/VL-evidence/VL-SKU deployed but not yet batch-measured.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gr6HH7JrdsXX8AARejQboM
2026-07-14 19:55:17 +07:00
Rafhan Mazaya FathurrahmanandClaude Sonnet 5 3a17c28758 feat(backend): diff-vs-previous-run reporting for product-scan accuracy harness
Ports the DO-harness's auto-diff-vs-previous-run reporting into
accuracy-check-scan.mts: prints a per-field, per-split (Training/
Validation) delta against the last product_accuracy_history.jsonl entry
and calls out regressions/improvements explicitly, plus classifier
method distribution and average confidence as informational context.

Also adds the real held-out validation photo set into
sources/product-test-images/ (75 photos, one per current SKU class) with
its README documenting the drop-photo -> label -> re-run workflow, so the
harness's Validation Set split actually has images to score against.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xsxk4ZkDQVVaLUcixDcqb5
2026-07-14 08:34:02 +07:00
Rafhan Mazaya Fathurrahman 577be04308 feat(app): implement Product Scan review flow, dynamic batch expiry picker, custom PDFs, PO relationships, and aligned card layouts 2026-07-09 12:36:06 +07:00
Rafhan Mazaya FathurrahmanandClaude Sonnet 5 e60ab63154 Adopt agents-settings kit, ship Product/SKU scan models, harden auth, verify OCR accuracy
Backend (app-pfm-ocr-v2/backend):
- Product/SKU scan feature complete: trained DINOv2 index (118 reference
  photos, 16 SKU classes) and YOLO classifier (83.3% top-1 val accuracy),
  fixed scripts/install-pipeline.sh (was missing ultralytics/torch), fully
  browser-verified end-to-end on /scan-pfm. Mobile m-scan-pfm page cancelled
  (Flutter app handles mobile; web UI is desktop-only for pipeline testing).
- Fixed a real data-loss bug: Save Ground Truth (scan-pfm and the DO-flow's
  manual-label) was silently writing into the pfm-web-app container's
  ephemeral filesystem instead of the host, because /sources wasn't
  bind-mounted in docker-compose.yml. Added the mount, recovered an
  orphaned entry.
- accounts.password is now bcrypt-hashed (bcryptjs, idempotent migration
  in db/init.ts) instead of plaintext; login route compares hashes.
- /api/v1/documents/* (list, PUT, upload) now enforces real 401 auth,
  matching what the Flutter client already sends. The "classic" routes
  deliberately stay open — they're dev-only web UI with no login flow and
  won't exist in production.
- OCR accuracy investigated end-to-end: real baseline is 95.10% overall
  (target met; accuracy_report.md was stale at 75.04%, now flagged). Fixed
  one genuine parser.ts bug (SO/DO field duplication in the global fallback
  regex); remaining gaps are OCR/layout-model limitations, not parser bugs.
- Adopted a standalone copy of the fhanyuh/agents-settings e/n workflow
  scoped to backend/ (AGENTS.md Part A/B split, SKILLS.md, plans/, docs/),
  independent of the root copy which now covers Flutter only.
- next-implementation.md deleted; content folded into
  backend/plans/next-enhancements.md for traceability.

Root:
- Adopted fhanyuh/agents-settings kit (AGENTS.md, SKILLS.md, plans/,
  docs/feature-list.md), scoped to the Flutter app only.
- Pending documents queue now persists to Hive (lib/core/storage) instead
  of memory-only, surviving an app kill mid-upload.

Removed backend_backup/ (stale Express/Prisma prototype, superseded by
pfm-web-app) and the completed plans/next-enhancement-plan.md checklist.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-08 11:56:32 +07:00
Rafhan Mazaya FathurrahmanandClaude Sonnet 5 4808a798fb Add mobile reliability fixes, Bahasa Indonesia UI, and continue OCR accuracy tuning
Reliability/PoC hardening: dedupe uploads by file_hash, surface editor sync
failures instead of a false success SnackBar with a retry-without-re-OCR path,
bound the OCR pipeline fetches with timeouts, share a single ApiClient/Dio
instance app-wide, tune capture JPEG quality, and add an opt-in
docker-compose.demo.yml for a production-mode run ahead of client demos.

Translate all Flutter-side user-facing text (screens, validators, SnackBars,
the printed delivery receipt, and shared API error messages) to Bahasa
Indonesia.

Also includes in-progress OCR parser/accuracy-tuning work from the same
session: table column/unit normalization fixes, store/customer master data,
accuracy history log, and test-image renaming/cleanup.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eRAsLqN9Sg1b9YPz22Lxz
2026-07-04 19:41:24 +07:00
Rafhan Mazaya Fathurrahman 095dd4cb8b feat: implement db layout caching for parsed ocr & reprocess batch test with corrected parser alignment 2026-07-03 16:43:24 +07:00
Rafhan Mazaya Fathurrahman 2900eb670b fix: resolve table column shift, normalize units and standardize dates 2026-07-03 16:04:57 +07:00
Rafhan Mazaya Fathurrahman 43e6eea253 feat: make manual-images list dynamic and resolve checkmark encoding issues 2026-07-02 16:57:47 +07:00
Rafhan Mazaya Fathurrahman a211593f2c feat: completely remove IMG_20260701_134635 and IMG_20260701_134636 from dataset and update report 2026-07-02 16:54:53 +07:00
Rafhan Mazaya Fathurrahman 89730912f3 feat: clean up duplicate files from dataset, update compiled jsons and recalculate accuracy report 2026-07-02 16:49:38 +07:00
Rafhan Mazaya Fathurrahman acd8fb4295 feat: standardise customer name as PT. PRIMAFOOD INTERNATIONAL and update report 2026-07-02 16:48:09 +07:00
Rafhan Mazaya Fathurrahman 6666ecd599 feat: auto-fill empty item names in manual labels from sku master database table 2026-07-02 16:46:13 +07:00
Rafhan Mazaya Fathurrahman 170cd586d9 feat: replace DKI AREA with PX HEAD OFFICE ANCOL in manual labels, update report and remove deleted test images 2026-07-02 16:45:14 +07:00
Rafhan Mazaya Fathurrahman 995347a0ea feat: compile manual labels and AI results into unified JSON files in sources folder 2026-07-02 16:41:25 +07:00
Rafhan Mazaya Fathurrahman 056a650885 feat: add Excel report generator and generate styled xlsx report 2026-07-02 15:34:19 +07:00
Rafhan Mazaya Fathurrahman 5c7c64e2f4 chore: reduce layout threshold to 0.2, tune vLLM memory, and add test scripts and reports 2026-07-02 15:30:19 +07:00
Rafhan Mazaya Fathurrahman bdb3a49742 feat: update backend OCR parser, web app, mobile app camera/preview UI, tests, and documentation with sample images 2026-07-02 13:33:57 +07:00
Rafhan Mazaya Fathurrahman ff3753a745 feat: consolidate backend and docker-compose setup 2026-06-30 21:09:08 +07:00