diff --git a/README.md b/README.md index ab514c0..0c58475 100644 --- a/README.md +++ b/README.md @@ -12,6 +12,7 @@ No cloud OCR API is used. Everything (layout detection, text recognition, LLM-as - [Tech Stack](#tech-stack) - [Architecture](#architecture) - [How It Works](#how-it-works) +- [OCR Accuracy](#ocr-accuracy) - [Repository Structure](#repository-structure) - [Getting Started](#getting-started) - [Prerequisites](#prerequisites) @@ -114,6 +115,41 @@ For the full field-by-field extraction rules (fused-digit correction, date sanit --- +## OCR Accuracy + +Accuracy is tracked field-by-field against 37 hand-labeled real DO photos (`backend/sources/test-images` + `manual_labels.json`), not a single vague "it works" claim. The latest logged run (`backend/sources/accuracy_history.jsonl`): + +**Overall: 89.4%** exact-field-match — up from an 82.2% baseline when this tracking tool was first built, against a 95% target. + +The bigger story is *where* that accuracy comes from. Raw regex extraction straight off the OCR text is only **67.7%** — the gain to 89.4% comes from a second correction pass: fuzzy SKU/store matching against master data, unit standardization, and format sanitization. + +| Field | Raw regex | After sanitize + triple-check | What closes the gap | +|---|---:|---:|---| +| Customer (Kepada Yth) | 100% | 100% | — | +| Kode Barang (SKU) | 98.4% | 98.4% | Already reliable at the OCR layer | +| Item count | 13.5% | 97.3% | Table-noise rows filtered by the SKU/unit triple-check | +| Nama Barang | 0% | 94.4% | Corrected to master `sku_master.nama_item` on SKU match | +| Banyak (qty) | 77.6% | 95.2% | Unit standardized from `sku_master.jenis_outer` | +| Jumlah (total) | 76.8% | 94.4% | Unit standardized from `sku_master.standar_jumlah` | +| No. PO | 91.9% | 91.9% | Fused-digit correction already applied at regex layer | +| No. DO | 91.9% | 91.9% | — | +| Tanggal | 89.2% | 89.2% | — | +| No. SO | 86.5% | 86.5% | — | +| Store | 0% | 64.9% | Fuzzy token match against `store_master` | +| Plat Truk | 62.2% | 62.2% | Mostly genuine OCR misses on the truck-line, not a parsing gap | +| Alamat | 0% | 37.8% | Canonicalized against `customers` table where phrasing matches | + +**Known gaps toward the 95% target:** +- **Alamat (37.8%, capped)** — the ground-truth labels themselves use two different phrasings for the same physical address across photo batches; closing this needs the ground truth unified, not more parsing logic. +- **Plat (62.2%)** — mostly genuine OCR misses (the truck line is often faint or absent in the photo) rather than a fixable parsing bug. +- **Store (64.9%)** — bounded by the same `store_master` table noted in [What to Consider](#what-to-consider-before-you-install) — accuracy against it can only improve as far as that master data is populated. + +**A real factor behind these gaps is photo capture SOP, not just parsing or OCR quality.** A meaningful share of the current test set was photographed with the DO paper placed on top of other papers/documents rather than a plain, flat surface. That confuses auto-deskew: the pipeline estimates page tilt from the average angle of detected text blocks, and overlapping paper edges/text from the sheet underneath make that estimate unreliable, which is exactly the failure mode behind the unwarp-retry logic described above. In practice this means accuracy here is a floor, not a ceiling — tightening the field SOP (DO paper alone, on a flat contrasting surface, reasonably well-lit and squared to the camera) should raise these numbers without any further code changes. + +Reproduce this yourself with `backend/pfm-web-app/scripts/accuracy-check.mts` (`npm run accuracy`, or `--refresh-ocr` to force a real pipeline re-run instead of using cached OCR results) — see [Testing & Tooling](#testing--tooling). + +--- + ## Repository Structure ```