Files
pfm-ocr/docs/expiry-tracking-plan.md
T
Rafhan Mazaya FathurrahmanandClaude Fable 5 e6daa9b053 docs(plan): per-sale expiry tracking design — batch registry + candidate matching
Grilled 2026-07-16 with the user; full decision record in
docs/expiry-tracking-plan.md. Core reframe: expiry is captured once per
batch at DO intake (staff-typed on the stock-entry confirmation page, from
the physical packs), so the cashier scan only MATCHES OCR fragments against
the 1-3 known in-stock batch dates instead of free-reading damaged
dot-matrix prints (proven model-capability ceiling, 2026-07-15). Fallback:
auto-FEFO + 'inferred' flag, zero cashier interaction. No cloud, ever.

- docs/expiry-tracking-plan.md: architecture, matching algorithm spec
  (resolveExpiryFromEvidence), schema/API deltas, phases 1-3, testing plan
- backend plans §13 (13.1-13.4): matcher util + offline tuning, route
  wiring + expiry_source provenance, multi-frame union, dot-matrix
  recognizer fine-tune
- root plans §10 (10.1-10.3): cashier fast path, inferred badge +
  end-of-day review, burst capture for mounted camera
- stock-feature-plan.md: extension note (batch dropdown becomes the
  manual-override path)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q8TumxFDnyVnfsR3mxPXfX
2026-07-16 17:00:36 +07:00

14 KiB
Raw Blame History

Per-Sale Expiry Tracking — Batch Registry + Candidate Matching

Written 2026-07-16 after a one-question-at-a-time grilling session with the user (see chat history — decisions recorded below, do not re-litigate). This is the context doc for backend plans/next-enhancements.md §13 and root plans/next-enhancements.md §10. It extends stock-feature-plan.md (backend §12 / root §9) — read that first; this doc assumes its schema and flows exist.

Status: planned, not yet implemented. Depends on the Stocks feature (§12.1/§12.2 backend, §9.1–9.4 Flutter), which is itself not yet built.

Problem

The client (retail store) must record the expiry date of every product sold to a customer. The expiry is printed on the pack, usually dot-matrix/inkjet on frozen plastic — frequently degraded (printer defects, scratches, ice, glare).

Measured evidence (79-image frozen validation set, 2026-07-14/15):

  • Overall product-scan accuracy 79.7%; expiry-date field only 64.6% (51/79).
  • The 28 expiry misses = 21 pure non-detections + 7 garbles.
  • The 21 non-detections were probed against every reader in the stack (PP-OCRv6 det/rec, PP-OCRv5-server, VL layout-parsing, direct VLM chat, plus upscale/blur/CLAHE/threshold preprocessing recipes via the temporary /probe-ocr endpoint): none can read these prints. A human can. This is a model capability ceiling, not a pipeline bug — free-form OCR of these prints cannot reach the target no matter how the code is tuned. Realistic free-read ceiling ≈ 82–85%.

Confirmed decisions (grilling record, 2026-07-16)

  1. Success = full automation. The scan happens at the cashier during checkout; added wait time is forbidden. Manual entry at the cashier is not acceptable as a routine step.
  2. Method is open — not restricted to OCR. Whatever reliably yields the expiry date wins.
  3. Upstream data: the Primafood DO paper does not carry batch data in a parseable-enough way to rely on; instead, staff enter batch code + expiry date per line item on the DO confirmation page (the stock-entry step of the Stocks feature), reading the values off the physical packs during goods receiving — no time pressure there. ~100% of sellable stock arrives via scanned DOs, so the batch registry will be complete.
  4. Data purpose: per-sale guarantee — the expiry of the physical unit sold, per transaction. Softened by decision 5 into "per-sale best evidence, honestly flagged when inferred".
  5. Residual case (multiple batches in stock AND print unmatchable): auto-record the FEFO batch (earliest expiry) + flag the record inferred — zero cashier interaction, never block or prompt.
  6. Cashier hardware: camera does both SKU and date (no barcode reliance); a fixed mounted camera at the checkout is a likely Phase-2 addition.
  7. No cloud at all — hard on-prem requirement. Flagged records may only be improved by on-prem means (Phase 3 recognizer, optional human review screen).
  8. Build order: Phase 1 (batch backbone + candidate matching, existing hardware) → Phase 2 (mounted camera, multi-frame) → Phase 3 (fine-tuned dot-matrix recognizer).

The reframe

Stop treating checkout as a reading problem ("OCR this damaged print") and treat it as a matching problem:

The true expiry of every unit in the store is already known — staff recorded it once per batch at intake. At the cashier, the camera only has to decide which of the 1–3 known in-stock batches this pack belongs to.

Consequences:

  • One batch in stock (the common case in a small store): the lookup alone is per-unit exact. Zero reading. Milliseconds.
  • Multiple batches: even a garbled OCR fragment (...2026, 2?10, 112026) is enough to pick between candidates whose dates differ. Matching against 2–3 known strings is drastically easier than free-form reading — most of the 21 "failed" images produced partial fragments that would disambiguate fine.
  • Unmatchable: FEFO + inferred flag (decision 5). Checkout never waits.

End-to-end data flow

INTAKE (no time pressure)                    CHECKOUT (hard latency budget)
─────────────────────────                    ──────────────────────────────
DO photo → OCR → editor                      pack photo → SKU classify
  → confirm (PUT)                              → in-stock batch lookup (candidates)
  → stock-entry screen                         → resolveExpiryFromEvidence()
    staff types batch_code +                       1 candidate  → single_batch
    expiry_date per line item                      exact date   → matched_exact
    (read off the packs)                           fragment win → matched_fragment
  → stock_batches rows born                        else         → inferred_fefo (flag)
    (kode_toko, no_sku,                        → auto-select batch, confirm
     batch_code, expiry_date, qty)             → decrement batch (§12.2)
                                               → sale row: expiry + source + score

The matching algorithm — resolveExpiryFromEvidence()

New pure TypeScript util backend/pfm-web-app/src/utils/expiry-matcher.ts (pure = offline-testable against the 79 captured OCR line-sets, no server needed).

Inputs

  • candidates: the scanned SKU's in-stock batches for this store — [{batchId, batchCode, expiryDate}], from stock_batches (§12.1).
  • evidence: the OCR text lines returned by the classify server for this scan (text_lines — already includes tiled full-res pass + VL-merged lines), plus the cascade's parsed date (date_extract.py output) if any.

Stages (first hit wins)

  1. single_batch — exactly one candidate: return it. No evidence needed.
  2. matched_exact — the cascade's parsed date equals one candidate's expiry_date: return that batch.
  3. matched_fragment — for each candidate, render its expected print forms (DDMMYYYY, DD/MM/YYYY, DD MM YY, DD.MM.YYYY, BB DDMMYYYY, 2-digit-year variants — reuse the format knowledge already encoded in date_extract.py); score every evidence line against every form with digit-confusion-aware fuzzy matching (Levenshtein over digit subsequences, with cheap substitutions for known OCR confusions: 0↔8, 1↔7, 5↔6, 2↔7, 3↔8; also credit partial anchors like a matching year + month pair). Candidate score = max over (lines × forms). Return the top candidate iff topScore ≥ SCORE_MIN and topScore − runnerUpScore ≥ MARGIN_MIN (both thresholds tuned offline — see Testing).
  4. inferred_fefo — otherwise: return the candidate with the earliest expiry_date, flagged.

Output: {batchId, expiryDate, source, score, margin} where source ∈ {single_batch, matched_exact, matched_fragment, inferred_fefo}.

Also matched: the batch_code string itself is a second fragment-matching target — batch codes are often printed adjacent to the date and give an independent disambiguation signal for free.

Verification step before building: confirm the classify server's response to the gateway actually carries text_lines (the offline capture scripts got them from the server, so it almost certainly does); if not, add them to the response payload — small change in config/classify_ocr_server.py.

Schema & API deltas (on top of stock-feature-plan.md)

  • documents (or the Product-branch metadata): add expiry_source VARCHAR(20) with CHECK (expiry_source IN ('single_batch','matched_exact', 'matched_fragment','inferred_fefo','manual')) and expiry_match_score REAL NULL. 'manual' covers legacy/edited rows.
  • api/parse/route.ts (Product branch) and api/v1/scan-product/route.ts: after classifyAndMatchProduct() + the §12.2 in-stock candidate filter, call resolveExpiryFromEvidence() and include the resolution (resolvedBatch + source + score) in the persisted metadata.productScan and in the response, so the Flutter editor can pre-select without any second call (same pattern as task 11.1).
  • v1/documents/[id]/route.ts PUT: persist expiry_source alongside the existing §12.2 stock_batch_id decrement. If the client overrides the auto-selected batch, source becomes 'manual'.
  • Review surface: GET /api/v1/documents?expiry_source=inferred_fefo filter (admin + own-store), powering an optional end-of-day review list.

Flutter deltas (root §10; builds on §9.1–9.4)

  • Fast path at cashier: the §9.4 Product-Scan editor auto-selects the resolved batch. When source is single_batch/matched_exact/ matched_fragment, the flow should be confirmable in one tap (or auto-confirm — decide at pickup with a grill question) with the resolved expiry displayed prominently. When inferred_fefo, same flow plus a small amber "perkiraan" badge — never a blocking prompt (decision 5).
  • Flag visibility: history/documents list shows the badge on inferred sales; an end-of-day review entry point lists them (uses the new filter). Review is optional and zero-checkout-impact by design.
  • Phase 2 capture mode: burst capture (N frames over ~1s) in the camera layer for mounted use; upload frames together; backend unions evidence lines across frames before matching (glare moves between frames — fragments accumulate).

Phases

Phase 1 — batch backbone + matcher (the PoC). Prereqs: §12.1, §12.2, §9.1–9.4. New work: expiry-matcher.ts + offline tuning harness, schema columns, route wiring, Flutter fast-path + badge. No new hardware or models.

Phase 2 — capture upgrade. Mounted camera at the cashier (a cheap phone running the existing Flutter app on a mount is acceptable hardware), burst/ multi-frame capture, evidence union across frames. Expected to lift fragment quality substantially — fixed focus distance + controlled lighting beat hand-held single shots.

Phase 3 — on-prem recognizer upgrade. Fine-tune a small recognition model specifically on dot-matrix/inkjet date prints:

  • Synthetic data: render dates in dot-matrix/inkjet fonts over pack-like backgrounds; augment with dot dropout, scratches, fade, curvature, glare, ice speckle. Thousands of labeled crops for free.
  • Real data flywheel: every intake stock-entry (staff-typed batch+expiry) plus every product-scan photo of that batch = weakly-labeled real training pairs accumulating automatically in normal operation. Harvest crops from uploads/ matched to registry values.
  • Train on the RTX 2060 (PaddleOCR rec fine-tune or similar small model); deploy as an additional reader in classify_ocr_server.py; its lines feed the same matcher. Shrinks the inferred_fefo residue. No cloud, ever (decision 7).

Expected accuracy (why this reaches ~90%+ where free OCR cannot)

Let p = share of scans where the SKU has exactly one batch in stock (small store, fast turnover → p is high, plausibly 0.6–0.8). Those are 100% correct by lookup. Of the rest, exact + fragment matching succeeds wherever OCR yields any usable fragment — on the 79-set evidence, most misses still produced fragments; matching 2–3 candidates needs far less signal than free reading. The residue is auto-FEFO'd — and FEFO itself is right whenever the customer took from the older batch, so even the flagged slice is mostly correct. Net: per-sale correctness ~90%+ in Phase 1, rising with Phases 2–3, with zero silent garbage — every record carries its provenance (source).

Two honest caveats to monitor:

  • SKU misclassification poisons the lookup (wrong SKU → wrong candidates). Mitigation: §12.2's in-stock filter shrinks the effective class space to what the store actually stocks; near-twin SKU confusion keeps improving via reference photos. Track sku/name accuracy alongside expiry.
  • Candidates with near-identical dates (differ by one digit) can fail the margin test → FEFO+flag. Correct behavior; expected to be rare.

Testing plan

  • Offline matcher tuning (before any wiring): replay the 79 captured OCR line-sets (sources/product_scan_detail_*.json + fullcap captures) against synthetic candidate sets built from the ground-truth labels (1, 2, and 3 candidates at varying date distances). Tune SCORE_MIN/MARGIN_MIN for zero wrong-candidate picks (a wrong confident match is worse than a flagged FEFO). This reuses the frozen benchmark as a matcher benchmark.
  • Unit tests: expiry-matcher pure-function tests (form rendering, confusion-aware scoring, margin logic, FEFO tiebreak) — TS side; Flutter pure-logic tests for fast-path/badge state per source value.
  • Live E2E (per repo convention, against the Docker stack): scan a DO → stock-entry with 2 batches of one SKU → product-scan a pack of the older batch → verify matched_* resolution, decrement of the right batch, and expiry_source in Postgres; then repeat with an unreadable pack → verify inferred_fefo + flag, no prompt shown.

Relationship to existing plans

  • Extends stock-feature-plan.md: §12.2's "closest-to-OCR-expiry batch auto-selected" dropdown becomes the manual-override UI behind the new automated resolution; the §12.2 decrement/400-on-bad-batch semantics are unchanged.
  • Backend tasks: backend/plans/next-enhancements.md §13.
  • Flutter tasks: root plans/next-enhancements.md §10.
  • The 79-image frozen benchmark and its capture tooling (accuracy work, 2026-07-14/15) become the matcher's offline test bed — nothing there is wasted by this reframe.