Files
pfm-ocr/docs/expiry-tracking-plan.md
T
Rafhan Mazaya FathurrahmanandClaude Fable 5 e6daa9b053 docs(plan): per-sale expiry tracking design — batch registry + candidate matching
Grilled 2026-07-16 with the user; full decision record in
docs/expiry-tracking-plan.md. Core reframe: expiry is captured once per
batch at DO intake (staff-typed on the stock-entry confirmation page, from
the physical packs), so the cashier scan only MATCHES OCR fragments against
the 1-3 known in-stock batch dates instead of free-reading damaged
dot-matrix prints (proven model-capability ceiling, 2026-07-15). Fallback:
auto-FEFO + 'inferred' flag, zero cashier interaction. No cloud, ever.

- docs/expiry-tracking-plan.md: architecture, matching algorithm spec
  (resolveExpiryFromEvidence), schema/API deltas, phases 1-3, testing plan
- backend plans §13 (13.1-13.4): matcher util + offline tuning, route
  wiring + expiry_source provenance, multi-frame union, dot-matrix
  recognizer fine-tune
- root plans §10 (10.1-10.3): cashier fast path, inferred badge +
  end-of-day review, burst capture for mounted camera
- stock-feature-plan.md: extension note (batch dropdown becomes the
  manual-override path)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q8TumxFDnyVnfsR3mxPXfX
2026-07-16 17:00:36 +07:00

245 lines
14 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Per-Sale Expiry Tracking — Batch Registry + Candidate Matching
Written 2026-07-16 after a one-question-at-a-time grilling session with the user
(see chat history — decisions recorded below, do not re-litigate). This is the
context doc for backend [`plans/next-enhancements.md`](../backend/plans/next-enhancements.md)
§13 and root [`plans/next-enhancements.md`](../plans/next-enhancements.md) §10.
It **extends** [`stock-feature-plan.md`](stock-feature-plan.md) (backend §12 /
root §9) — read that first; this doc assumes its schema and flows exist.
**Status: planned, not yet implemented.** Depends on the Stocks feature
(§12.1/§12.2 backend, §9.1–9.4 Flutter), which is itself not yet built.
## Problem
The client (retail store) must record the expiry date of every product sold to a
customer. The expiry is printed on the pack, usually dot-matrix/inkjet on frozen
plastic — frequently degraded (printer defects, scratches, ice, glare).
Measured evidence (79-image frozen validation set, 2026-07-14/15):
- Overall product-scan accuracy 79.7%; **expiry-date field only 64.6%** (51/79).
- The 28 expiry misses = 21 pure non-detections + 7 garbles.
- The 21 non-detections were probed against **every reader in the stack**
(PP-OCRv6 det/rec, PP-OCRv5-server, VL layout-parsing, direct VLM chat, plus
upscale/blur/CLAHE/threshold preprocessing recipes via the temporary
`/probe-ocr` endpoint): none can read these prints. A human can. This is a
**model capability ceiling, not a pipeline bug** — free-form OCR of these
prints cannot reach the target no matter how the code is tuned. Realistic
free-read ceiling ≈ 82–85%.
## Confirmed decisions (grilling record, 2026-07-16)
1. **Success = full automation.** The scan happens at the cashier during
checkout; added wait time is forbidden. Manual entry at the cashier is not
acceptable as a routine step.
2. **Method is open** — not restricted to OCR. Whatever reliably yields the
expiry date wins.
3. **Upstream data**: the Primafood DO paper does **not** carry batch data in a
parseable-enough way to rely on; instead, **staff enter batch code + expiry
date per line item on the DO confirmation page** (the stock-entry step of
the Stocks feature), reading the values **off the physical packs** during
goods receiving — no time pressure there. ~100% of sellable stock arrives
via scanned DOs, so the batch registry will be complete.
4. **Data purpose: per-sale guarantee** — the expiry of the physical unit sold,
per transaction. Softened by decision 5 into "per-sale best evidence,
honestly flagged when inferred".
5. **Residual case** (multiple batches in stock AND print unmatchable):
**auto-record the FEFO batch (earliest expiry) + flag the record
`inferred`** — zero cashier interaction, never block or prompt.
6. **Cashier hardware**: camera does both SKU and date (no barcode reliance);
a fixed **mounted camera** at the checkout is a likely Phase-2 addition.
7. **No cloud at all** — hard on-prem requirement. Flagged records may only be
improved by on-prem means (Phase 3 recognizer, optional human review screen).
8. **Build order**: Phase 1 (batch backbone + candidate matching, existing
hardware) → Phase 2 (mounted camera, multi-frame) → Phase 3 (fine-tuned
dot-matrix recognizer).
## The reframe
Stop treating checkout as a *reading* problem ("OCR this damaged print") and
treat it as a *matching* problem:
> The true expiry of every unit in the store is already known — staff recorded
> it once per batch at intake. At the cashier, the camera only has to decide
> **which of the 1–3 known in-stock batches** this pack belongs to.
Consequences:
- **One batch in stock** (the common case in a small store): the lookup alone
is per-unit exact. Zero reading. Milliseconds.
- **Multiple batches**: even a garbled OCR fragment (`...2026`, `2?10`,
`112026`) is enough to pick between candidates whose dates differ. Matching
against 2–3 known strings is drastically easier than free-form reading —
most of the 21 "failed" images produced partial fragments that would
disambiguate fine.
- **Unmatchable**: FEFO + `inferred` flag (decision 5). Checkout never waits.
## End-to-end data flow
```
INTAKE (no time pressure) CHECKOUT (hard latency budget)
───────────────────────── ──────────────────────────────
DO photo → OCR → editor pack photo → SKU classify
→ confirm (PUT) → in-stock batch lookup (candidates)
→ stock-entry screen → resolveExpiryFromEvidence()
staff types batch_code + 1 candidate → single_batch
expiry_date per line item exact date → matched_exact
(read off the packs) fragment win → matched_fragment
→ stock_batches rows born else → inferred_fefo (flag)
(kode_toko, no_sku, → auto-select batch, confirm
batch_code, expiry_date, qty) → decrement batch (§12.2)
→ sale row: expiry + source + score
```
## The matching algorithm — `resolveExpiryFromEvidence()`
New pure TypeScript util `backend/pfm-web-app/src/utils/expiry-matcher.ts`
(pure = offline-testable against the 79 captured OCR line-sets, no server
needed).
**Inputs**
- `candidates`: the scanned SKU's in-stock batches for this store —
`[{batchId, batchCode, expiryDate}]`, from `stock_batches` (§12.1).
- `evidence`: the OCR text lines returned by the classify server for this scan
(`text_lines` — already includes tiled full-res pass + VL-merged lines), plus
the cascade's parsed date (`date_extract.py` output) if any.
**Stages** (first hit wins)
1. `single_batch` — exactly one candidate: return it. No evidence needed.
2. `matched_exact` — the cascade's parsed date equals one candidate's
`expiry_date`: return that batch.
3. `matched_fragment` — for each candidate, render its expected print forms
(`DDMMYYYY`, `DD/MM/YYYY`, `DD MM YY`, `DD.MM.YYYY`, `BB DDMMYYYY`,
2-digit-year variants — reuse the format knowledge already encoded in
`date_extract.py`); score every evidence line against every form with
digit-confusion-aware fuzzy matching (Levenshtein over digit subsequences,
with cheap substitutions for known OCR confusions: 0↔8, 1↔7, 5↔6, 2↔7,
3↔8; also credit partial anchors like a matching year + month pair).
Candidate score = max over (lines × forms). Return the top candidate iff
`topScore ≥ SCORE_MIN` **and** `topScore − runnerUpScore ≥ MARGIN_MIN`
(both thresholds tuned offline — see Testing).
4. `inferred_fefo` — otherwise: return the candidate with the earliest
`expiry_date`, flagged.
**Output**: `{batchId, expiryDate, source, score, margin}` where
`source ∈ {single_batch, matched_exact, matched_fragment, inferred_fefo}`.
**Also matched**: the `batch_code` string itself is a second fragment-matching
target — batch codes are often printed adjacent to the date and give an
independent disambiguation signal for free.
**Verification step before building**: confirm the classify server's response
to the gateway actually carries `text_lines` (the offline capture scripts got
them from the server, so it almost certainly does); if not, add them to the
response payload — small change in `config/classify_ocr_server.py`.
## Schema & API deltas (on top of stock-feature-plan.md)
- `documents` (or the Product-branch metadata): add `expiry_source VARCHAR(20)`
with `CHECK (expiry_source IN ('single_batch','matched_exact',
'matched_fragment','inferred_fefo','manual'))` and
`expiry_match_score REAL NULL`. `'manual'` covers legacy/edited rows.
- `api/parse/route.ts` (Product branch) and `api/v1/scan-product/route.ts`:
after `classifyAndMatchProduct()` + the §12.2 in-stock candidate filter, call
`resolveExpiryFromEvidence()` and include the resolution
(`resolvedBatch` + `source` + `score`) in the persisted
`metadata.productScan` and in the response, so the Flutter editor can
pre-select without any second call (same pattern as task 11.1).
- `v1/documents/[id]/route.ts` PUT: persist `expiry_source` alongside the
existing §12.2 `stock_batch_id` decrement. If the client overrides the
auto-selected batch, source becomes `'manual'`.
- Review surface: `GET /api/v1/documents?expiry_source=inferred_fefo` filter
(admin + own-store), powering an optional end-of-day review list.
## Flutter deltas (root §10; builds on §9.1–9.4)
- **Fast path at cashier**: the §9.4 Product-Scan editor auto-selects the
resolved batch. When `source` is `single_batch`/`matched_exact`/
`matched_fragment`, the flow should be confirmable in **one tap** (or
auto-confirm — decide at pickup with a grill question) with the resolved
expiry displayed prominently. When `inferred_fefo`, same flow plus a small
amber "perkiraan" badge — never a blocking prompt (decision 5).
- **Flag visibility**: history/documents list shows the badge on inferred
sales; an end-of-day review entry point lists them (uses the new filter).
Review is optional and zero-checkout-impact by design.
- **Phase 2 capture mode**: burst capture (N frames over ~1s) in the camera
layer for mounted use; upload frames together; backend unions evidence lines
across frames before matching (glare moves between frames — fragments
accumulate).
## Phases
**Phase 1 — batch backbone + matcher (the PoC).** Prereqs: §12.1, §12.2,
§9.1–9.4. New work: `expiry-matcher.ts` + offline tuning harness, schema
columns, route wiring, Flutter fast-path + badge. No new hardware or models.
**Phase 2 — capture upgrade.** Mounted camera at the cashier (a cheap phone
running the existing Flutter app on a mount is acceptable hardware), burst/
multi-frame capture, evidence union across frames. Expected to lift fragment
quality substantially — fixed focus distance + controlled lighting beat
hand-held single shots.
**Phase 3 — on-prem recognizer upgrade.** Fine-tune a small recognition model
specifically on dot-matrix/inkjet date prints:
- **Synthetic data**: render dates in dot-matrix/inkjet fonts over pack-like
backgrounds; augment with dot dropout, scratches, fade, curvature, glare,
ice speckle. Thousands of labeled crops for free.
- **Real data flywheel**: every intake stock-entry (staff-typed batch+expiry)
plus every product-scan photo of that batch = weakly-labeled real training
pairs accumulating automatically in normal operation. Harvest crops from
`uploads/` matched to registry values.
- Train on the RTX 2060 (PaddleOCR rec fine-tune or similar small model);
deploy as an additional reader in `classify_ocr_server.py`; its lines feed
the same matcher. Shrinks the `inferred_fefo` residue. **No cloud, ever**
(decision 7).
## Expected accuracy (why this reaches ~90%+ where free OCR cannot)
Let p = share of scans where the SKU has exactly one batch in stock (small
store, fast turnover → p is high, plausibly 0.6–0.8). Those are 100% correct
by lookup. Of the rest, exact + fragment matching succeeds wherever OCR yields
*any* usable fragment — on the 79-set evidence, most misses still produced
fragments; matching 2–3 candidates needs far less signal than free reading.
The residue is auto-FEFO'd — and FEFO itself is right whenever the customer
took from the older batch, so even the flagged slice is mostly correct.
Net: per-sale correctness ~90%+ in Phase 1, rising with Phases 2–3, with
**zero silent garbage** — every record carries its provenance (`source`).
Two honest caveats to monitor:
- **SKU misclassification poisons the lookup** (wrong SKU → wrong candidates).
Mitigation: §12.2's in-stock filter shrinks the effective class space to
what the store actually stocks; near-twin SKU confusion keeps improving via
reference photos. Track `sku/name` accuracy alongside expiry.
- **Candidates with near-identical dates** (differ by one digit) can fail the
margin test → FEFO+flag. Correct behavior; expected to be rare.
## Testing plan
- **Offline matcher tuning (before any wiring)**: replay the 79 captured OCR
line-sets (`sources/product_scan_detail_*.json` + fullcap captures) against
synthetic candidate sets built from the ground-truth labels (1, 2, and 3
candidates at varying date distances). Tune `SCORE_MIN`/`MARGIN_MIN` for
zero wrong-candidate picks (a wrong confident match is worse than a flagged
FEFO). This reuses the frozen benchmark as a matcher benchmark.
- **Unit tests**: `expiry-matcher` pure-function tests (form rendering,
confusion-aware scoring, margin logic, FEFO tiebreak) — TS side; Flutter
pure-logic tests for fast-path/badge state per `source` value.
- **Live E2E** (per repo convention, against the Docker stack): scan a DO →
stock-entry with 2 batches of one SKU → product-scan a pack of the older
batch → verify `matched_*` resolution, decrement of the right batch, and
`expiry_source` in Postgres; then repeat with an unreadable pack → verify
`inferred_fefo` + flag, no prompt shown.
## Relationship to existing plans
- **Extends** `stock-feature-plan.md`: §12.2's "closest-to-OCR-expiry batch
auto-selected" dropdown becomes the *manual-override* UI behind the new
automated resolution; the §12.2 decrement/400-on-bad-batch semantics are
unchanged.
- Backend tasks: `backend/plans/next-enhancements.md` **§13**.
- Flutter tasks: root `plans/next-enhancements.md` **§10**.
- The 79-image frozen benchmark and its capture tooling (accuracy work,
2026-07-14/15) become the matcher's offline test bed — nothing there is
wasted by this reframe.