Files
pfm-ocr/backend/sources/accuracy_report.md
T
Rafhan Mazaya FathurrahmanandClaude Sonnet 5 e60ab63154 Adopt agents-settings kit, ship Product/SKU scan models, harden auth, verify OCR accuracy
Backend (app-pfm-ocr-v2/backend):
- Product/SKU scan feature complete: trained DINOv2 index (118 reference
  photos, 16 SKU classes) and YOLO classifier (83.3% top-1 val accuracy),
  fixed scripts/install-pipeline.sh (was missing ultralytics/torch), fully
  browser-verified end-to-end on /scan-pfm. Mobile m-scan-pfm page cancelled
  (Flutter app handles mobile; web UI is desktop-only for pipeline testing).
- Fixed a real data-loss bug: Save Ground Truth (scan-pfm and the DO-flow's
  manual-label) was silently writing into the pfm-web-app container's
  ephemeral filesystem instead of the host, because /sources wasn't
  bind-mounted in docker-compose.yml. Added the mount, recovered an
  orphaned entry.
- accounts.password is now bcrypt-hashed (bcryptjs, idempotent migration
  in db/init.ts) instead of plaintext; login route compares hashes.
- /api/v1/documents/* (list, PUT, upload) now enforces real 401 auth,
  matching what the Flutter client already sends. The "classic" routes
  deliberately stay open — they're dev-only web UI with no login flow and
  won't exist in production.
- OCR accuracy investigated end-to-end: real baseline is 95.10% overall
  (target met; accuracy_report.md was stale at 75.04%, now flagged). Fixed
  one genuine parser.ts bug (SO/DO field duplication in the global fallback
  regex); remaining gaps are OCR/layout-model limitations, not parser bugs.
- Adopted a standalone copy of the fhanyuh/agents-settings e/n workflow
  scoped to backend/ (AGENTS.md Part A/B split, SKILLS.md, plans/, docs/),
  independent of the root copy which now covers Flutter only.
- next-implementation.md deleted; content folded into
  backend/plans/next-enhancements.md for traceability.

Root:
- Adopted fhanyuh/agents-settings kit (AGENTS.md, SKILLS.md, plans/,
  docs/feature-list.md), scoped to the Flutter app only.
- Pending documents queue now persists to Hive (lib/core/storage) instead
  of memory-only, surviving an app kill mid-upload.

Removed backend_backup/ (stale Express/Prisma prototype, superseded by
pfm-web-app) and the completed plans/next-enhancement-plan.md checklist.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-08 11:56:32 +07:00

7.4 KiB

AI OCR Accuracy & Performance Report

⚠️ Stale as of 2026-07-08. The numbers below (75.04% overall) predate several parser fixes (table column-shift correction, unit normalization, date standardization — see git log) and are significantly out of date. The real current baseline, confirmed 2026-07-08 by re-running node pfm-web-app/scripts/accuracy-check.mts at commit 3df9f6e (matches sources/accuracy_history.jsonl's latest entry exactly): 95.10% overall layer3Final — already at/above the 95% target. Worst fields now: plat (67.6%, almost entirely OCR/layout-model misses — see backend/plans/next-enhancements.md §2 task 2.3 for the full investigation), noSO (86.5%), tanggal (89.2%), noPO/noDO (91.9% each). Don't trust the prose/table below without re-running the harness first — this file isn't auto-regenerated on every run.

This report summarizes the comparison of AI OCR Extraction (Layer 3 Final) against the Manual Ground Truth Labels across all 37 test images.


📊 Summary Accuracy

Field / Area Total Checks Matches Mismatches Accuracy (%)
PO Number 37 33 4 89.19%
SO Number 37 31 6 83.78%
DO Number 37 32 5 86.49%
Date 37 20 17 54.05%
Plat Nomor 37 21 16 56.76%
Customer Name 37 37 0 100.00%
Store Name 37 11 26 29.73%
Alamat 37 12 25 32.43%
Item SKU 127 119 8 93.70%
Item Banyak (Qty Pkg) 127 100 27 78.74%
Item Jumlah (Qty Unit) 127 92 35 72.44%
OVERALL TOTAL 677 508 169 75.04%

Detailed, color-coded comparison results (including side-by-side matches/mismatches) are available in the generated Excel sheet: sources/comparison_report.xlsx.


🔍 Key Findings & Analysis

🟢 What is Performing Well (High Accuracy)

  1. Customer Name (100.00%): Every customer name field matches perfectly, as it is standard and usually well-printed.
  2. Item SKU (93.70%): Outstanding performance on SKU extraction. This is boosted by the SKU master list validation (Levenshtein score threshold check), which dynamically corrects OCR recognition noise into actual valid database SKUs.
  3. DO, PO, and SO Numbers (83.78% - 89.19%): High accuracy for document tracking numbers, which is crucial for database consistency and mapping.

🟡 Moderate Performance (Needs Tuning)

  1. Item Banyak (78.74%) & Item Jumlah (72.44%): Quantity parsing sometimes gets thrown off by OCR capturing checkmarks (✓, ✔) or adjacent unit annotations (e.g. PC, PAC, KG) next to digits, leading to slight string mismatches.

🔴 Opportunities for Improvement (Low Accuracy)

  1. Store Name (29.73%) & Alamat (32.43%):
    • Why it's low: The manual labels for address and store name contain highly specific strings, while the AI OCR might extract partial address segments, omit specific keywords (like "DKI AREA", "KANTIN CPI"), or extract slightly different abbreviations.
    • Recommendation: Introduce a fuzzy store-name and address resolution map against the database entries (using Levenshtein distance on store aliases) to dynamically map extracted text to clean master values.
  2. Date (54.05%):
    • Why it's low: Differences in format normalization (e.g., raw text "23 June 2026" vs expected normalized representations, or OCR misreading text dates like "15 May 2026").
    • Recommendation: Enhance sanitizeParsedMetadata in parser.ts to parse multiple date format variations into a unified target format.
  3. Plat Nomor (56.76%):
    • Why it's low: License plates on delivery orders are often stamped, handwritten, or placed in odd margins, which makes clean extraction difficult.

🧪 Parser Unit Test Results (parser.test.ts)

Last run: 2026-07-06 · Result: 48 / 48 passed (100%) ✅

Unit-level regression tests for parseDOMetadata() and sanitizeParsedMetadata() in parser.ts. Run anytime with:

npx tsx src/utils/parser.test.ts

(from backend/pfm-web-app). Triggered this run by the Customer Name post-processing cleanup — removed the now-dead OCR extraction/regex logic for customerInfo (it was always overwritten by a hardcoded "PT. PRIMAFOOD INTERNATIONAL" constant anyway) and confirmed nothing else regressed.

Status Count
✅ Passed 48
❌ Failed 0
Total 48

parseDOMetadata — 26 tests

# Test Case Result
1 PO standard PO/26/ ✅
2 PO misread P0/26/ on label ✅
3 PO label raw 10-digit, real PO in body ✅
4 PO misread F0/20/ — use current year NOT 20 ✅
5 PO body P0/26/ ✅
6 PO real doc: label raw, body has P0/26/ ✅
7 PO fused F012070000170727 ✅
8 PO fused PO12070000190729 ✅
9 PO noise digits PO120/0000170727 ✅
10 Date trailing noise cut ✅
11 Date no space 25May2020 ✅
12 Date standard 23 June 2026 ✅
13 Date Tanggal blank shifted to No.SO ✅
14 Date prefix timestamp noise ✅
15 Date (Asli/Copy) prefix noise ✅
16 Date junk suffix cut ✅
17 Date bad OCR month Hv → Not Found ✅
18 Date single digit 7 May 2026 → 07 May 2026 ✅
19 Date single digit 4 Apr 2026 → 04 April 2026 ✅
20 00117709 before Tanggal must not pollute date ✅
21 Plate B 9427 UXT ✅
22 Plate B-9999-XYZ dash ✅
23 Plate ignore PO/SO prefix ✅
24 Plate B9427UXT adjacent ✅
25 Plate real doc B 9723 CXS ✅
26 Table column shift alignment correction ✅

sanitizeParsedMetadata — 22 tests

# Test Case Result
27 valid tanggal 30 June 2026 passes ✅
28 valid tanggal 25 May 2020 passes ✅
29 single digit tanggal 4 April 2026 → 04 April 2026 ✅
30 tanggal bad month Hv → Not Found ✅
31 tanggal as number 0011770 → Not Found ✅
32 tanggal day 0 → Not Found ✅
33 tanggal day 32 → Not Found ✅
34 tanggal year 2009 (too old) → Not Found ✅
35 tanggal Not Found stays Not Found ✅
36 tanggal with noise suffix → Not Found ✅
37 valid noPO PO/26/0000178435 passes ✅
38 noPO wrong year auto-corrected to current ✅
39 noPO raw number → Not Found ✅
40 noPO Not Found stays Not Found ✅
41 valid noSO 1691908676 passes ✅
42 noSO 'abc' → Not Found ✅
43 noSO too short '123' → Not Found ✅
44 valid noDO 1659932080 passes ✅
45 noDO 'XYZXYZ' → Not Found ✅
46 valid platTruk B 9427 UXT passes ✅
47 platTruk XY 1234 ABC invalid prefix → empty ✅
48 platTruk empty stays empty ✅

Note: no test case directly exercises customerInfo (it wasn't covered before this change either) - passing confirms the surrounding logic (PO/date/plate/table parsing) is unaffected by the cleanup.