fix: resolve table column shift, normalize units and standardize dates

This commit is contained in:
Rafhan Mazaya Fathurrahman committed 2026-07-03 16:04:57 +07:00
1 parent 095d4b799a
commit 2900eb670b
47 files changed
+11881 -2724

No files matched your search

+46
View File
@@ -0,0 +1,46 @@
# AI OCR Accuracy & Performance Report
This report summarizes the comparison of **AI OCR Extraction (Layer 3 Final)** against the **Manual Ground Truth Labels** across all **37 test images**.
---
## 📊 Summary Accuracy
| Field / Area | Total Checks | Matches | Mismatches | Accuracy (%) |
| :--- | :---: | :---: | :---: | :---: |
| **PO Number** | 37 | 33 | 4 | **89.19%** |
| **SO Number** | 37 | 31 | 6 | **83.78%** |
| **DO Number** | 37 | 32 | 5 | **86.49%** |
| **Date** | 37 | 20 | 17 | **54.05%** |
| **Plat Nomor** | 37 | 21 | 16 | **56.76%** |
| **Customer Name** | 37 | 37 | 0 | **100.00%** |
| **Store Name** | 37 | 11 | 26 | **29.73%** |
| **Alamat** | 37 | 12 | 25 | **32.43%** |
| **Item SKU** | 127 | 119 | 8 | **93.70%** |
| **Item Banyak (Qty Pkg)** | 127 | 100 | 27 | **78.74%** |
| **Item Jumlah (Qty Unit)**| 127 | 92 | 35 | **72.44%** |
| **OVERALL TOTAL** | **677** | **508** | **169** | **75.04%** |
*Detailed, color-coded comparison results (including side-by-side matches/mismatches) are available in the generated Excel sheet: [sources/comparison_report.xlsx](file:///d:/Client/Data%20Bisnis%20Solusi/app-pfm-ocr-v2/backend/sources/comparison_report.xlsx).*
---
## 🔍 Key Findings & Analysis
### 🟢 What is Performing Well (High Accuracy)
1. **Customer Name (100.00%)**: Every customer name field matches perfectly, as it is standard and usually well-printed.
2. **Item SKU (93.70%)**: Outstanding performance on SKU extraction. This is boosted by the **SKU master list validation** (Levenshtein score threshold check), which dynamically corrects OCR recognition noise into actual valid database SKUs.
3. **DO, PO, and SO Numbers (83.78% - 89.19%)**: High accuracy for document tracking numbers, which is crucial for database consistency and mapping.
### 🟡 Moderate Performance (Needs Tuning)
1. **Item Banyak (78.74%) & Item Jumlah (72.44%)**: Quantity parsing sometimes gets thrown off by OCR capturing checkmarks (`✓`, `✔`) or adjacent unit annotations (e.g. `PC`, `PAC`, `KG`) next to digits, leading to slight string mismatches.
### 🔴 Opportunities for Improvement (Low Accuracy)
1. **Store Name (29.73%) & Alamat (32.43%)**:
- *Why it's low*: The manual labels for address and store name contain highly specific strings, while the AI OCR might extract partial address segments, omit specific keywords (like "DKI AREA", "KANTIN CPI"), or extract slightly different abbreviations.
- *Recommendation*: Introduce a fuzzy store-name and address resolution map against the database entries (using Levenshtein distance on store aliases) to dynamically map extracted text to clean master values.
2. **Date (54.05%)**:
- *Why it's low*: Differences in format normalization (e.g., raw text "23 June 2026" vs expected normalized representations, or OCR misreading text dates like "15 May 2026").
- *Recommendation*: Enhance `sanitizeParsedMetadata` in [parser.ts](file:///d:/Client/Data%20Bisnis%20Solusi/app-pfm-ocr-v2/backend/pfm-web-app/src/utils/parser.ts) to parse multiple date format variations into a unified target format.
3. **Plat Nomor (56.76%)**:
- *Why it's low*: License plates on delivery orders are often stamped, handwritten, or placed in odd margins, which makes clean extraction difficult.