Grilled 2026-07-16 with the user; full decision record in docs/expiry-tracking-plan.md. Core reframe: expiry is captured once per batch at DO intake (staff-typed on the stock-entry confirmation page, from the physical packs), so the cashier scan only MATCHES OCR fragments against the 1-3 known in-stock batch dates instead of free-reading damaged dot-matrix prints (proven model-capability ceiling, 2026-07-15). Fallback: auto-FEFO + 'inferred' flag, zero cashier interaction. No cloud, ever. - docs/expiry-tracking-plan.md: architecture, matching algorithm spec (resolveExpiryFromEvidence), schema/API deltas, phases 1-3, testing plan - backend plans §13 (13.1-13.4): matcher util + offline tuning, route wiring + expiry_source provenance, multi-frame union, dot-matrix recognizer fine-tune - root plans §10 (10.1-10.3): cashier fast path, inferred badge + end-of-day review, burst capture for mounted camera - stock-feature-plan.md: extension note (batch dropdown becomes the manual-override path) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q8TumxFDnyVnfsR3mxPXfX
245 lines
14 KiB
Markdown
245 lines
14 KiB
Markdown
# Per-Sale Expiry Tracking — Batch Registry + Candidate Matching
|
||
|
||
Written 2026-07-16 after a one-question-at-a-time grilling session with the user
|
||
(see chat history — decisions recorded below, do not re-litigate). This is the
|
||
context doc for backend [`plans/next-enhancements.md`](../backend/plans/next-enhancements.md)
|
||
§13 and root [`plans/next-enhancements.md`](../plans/next-enhancements.md) §10.
|
||
It **extends** [`stock-feature-plan.md`](stock-feature-plan.md) (backend §12 /
|
||
root §9) — read that first; this doc assumes its schema and flows exist.
|
||
|
||
**Status: planned, not yet implemented.** Depends on the Stocks feature
|
||
(§12.1/§12.2 backend, §9.1–9.4 Flutter), which is itself not yet built.
|
||
|
||
## Problem
|
||
|
||
The client (retail store) must record the expiry date of every product sold to a
|
||
customer. The expiry is printed on the pack, usually dot-matrix/inkjet on frozen
|
||
plastic — frequently degraded (printer defects, scratches, ice, glare).
|
||
|
||
Measured evidence (79-image frozen validation set, 2026-07-14/15):
|
||
|
||
- Overall product-scan accuracy 79.7%; **expiry-date field only 64.6%** (51/79).
|
||
- The 28 expiry misses = 21 pure non-detections + 7 garbles.
|
||
- The 21 non-detections were probed against **every reader in the stack**
|
||
(PP-OCRv6 det/rec, PP-OCRv5-server, VL layout-parsing, direct VLM chat, plus
|
||
upscale/blur/CLAHE/threshold preprocessing recipes via the temporary
|
||
`/probe-ocr` endpoint): none can read these prints. A human can. This is a
|
||
**model capability ceiling, not a pipeline bug** — free-form OCR of these
|
||
prints cannot reach the target no matter how the code is tuned. Realistic
|
||
free-read ceiling ≈ 82–85%.
|
||
|
||
## Confirmed decisions (grilling record, 2026-07-16)
|
||
|
||
1. **Success = full automation.** The scan happens at the cashier during
|
||
checkout; added wait time is forbidden. Manual entry at the cashier is not
|
||
acceptable as a routine step.
|
||
2. **Method is open** — not restricted to OCR. Whatever reliably yields the
|
||
expiry date wins.
|
||
3. **Upstream data**: the Primafood DO paper does **not** carry batch data in a
|
||
parseable-enough way to rely on; instead, **staff enter batch code + expiry
|
||
date per line item on the DO confirmation page** (the stock-entry step of
|
||
the Stocks feature), reading the values **off the physical packs** during
|
||
goods receiving — no time pressure there. ~100% of sellable stock arrives
|
||
via scanned DOs, so the batch registry will be complete.
|
||
4. **Data purpose: per-sale guarantee** — the expiry of the physical unit sold,
|
||
per transaction. Softened by decision 5 into "per-sale best evidence,
|
||
honestly flagged when inferred".
|
||
5. **Residual case** (multiple batches in stock AND print unmatchable):
|
||
**auto-record the FEFO batch (earliest expiry) + flag the record
|
||
`inferred`** — zero cashier interaction, never block or prompt.
|
||
6. **Cashier hardware**: camera does both SKU and date (no barcode reliance);
|
||
a fixed **mounted camera** at the checkout is a likely Phase-2 addition.
|
||
7. **No cloud at all** — hard on-prem requirement. Flagged records may only be
|
||
improved by on-prem means (Phase 3 recognizer, optional human review screen).
|
||
8. **Build order**: Phase 1 (batch backbone + candidate matching, existing
|
||
hardware) → Phase 2 (mounted camera, multi-frame) → Phase 3 (fine-tuned
|
||
dot-matrix recognizer).
|
||
|
||
## The reframe
|
||
|
||
Stop treating checkout as a *reading* problem ("OCR this damaged print") and
|
||
treat it as a *matching* problem:
|
||
|
||
> The true expiry of every unit in the store is already known — staff recorded
|
||
> it once per batch at intake. At the cashier, the camera only has to decide
|
||
> **which of the 1–3 known in-stock batches** this pack belongs to.
|
||
|
||
Consequences:
|
||
|
||
- **One batch in stock** (the common case in a small store): the lookup alone
|
||
is per-unit exact. Zero reading. Milliseconds.
|
||
- **Multiple batches**: even a garbled OCR fragment (`...2026`, `2?10`,
|
||
`112026`) is enough to pick between candidates whose dates differ. Matching
|
||
against 2–3 known strings is drastically easier than free-form reading —
|
||
most of the 21 "failed" images produced partial fragments that would
|
||
disambiguate fine.
|
||
- **Unmatchable**: FEFO + `inferred` flag (decision 5). Checkout never waits.
|
||
|
||
## End-to-end data flow
|
||
|
||
```
|
||
INTAKE (no time pressure) CHECKOUT (hard latency budget)
|
||
───────────────────────── ──────────────────────────────
|
||
DO photo → OCR → editor pack photo → SKU classify
|
||
→ confirm (PUT) → in-stock batch lookup (candidates)
|
||
→ stock-entry screen → resolveExpiryFromEvidence()
|
||
staff types batch_code + 1 candidate → single_batch
|
||
expiry_date per line item exact date → matched_exact
|
||
(read off the packs) fragment win → matched_fragment
|
||
→ stock_batches rows born else → inferred_fefo (flag)
|
||
(kode_toko, no_sku, → auto-select batch, confirm
|
||
batch_code, expiry_date, qty) → decrement batch (§12.2)
|
||
→ sale row: expiry + source + score
|
||
```
|
||
|
||
## The matching algorithm — `resolveExpiryFromEvidence()`
|
||
|
||
New pure TypeScript util `backend/pfm-web-app/src/utils/expiry-matcher.ts`
|
||
(pure = offline-testable against the 79 captured OCR line-sets, no server
|
||
needed).
|
||
|
||
**Inputs**
|
||
- `candidates`: the scanned SKU's in-stock batches for this store —
|
||
`[{batchId, batchCode, expiryDate}]`, from `stock_batches` (§12.1).
|
||
- `evidence`: the OCR text lines returned by the classify server for this scan
|
||
(`text_lines` — already includes tiled full-res pass + VL-merged lines), plus
|
||
the cascade's parsed date (`date_extract.py` output) if any.
|
||
|
||
**Stages** (first hit wins)
|
||
1. `single_batch` — exactly one candidate: return it. No evidence needed.
|
||
2. `matched_exact` — the cascade's parsed date equals one candidate's
|
||
`expiry_date`: return that batch.
|
||
3. `matched_fragment` — for each candidate, render its expected print forms
|
||
(`DDMMYYYY`, `DD/MM/YYYY`, `DD MM YY`, `DD.MM.YYYY`, `BB DDMMYYYY`,
|
||
2-digit-year variants — reuse the format knowledge already encoded in
|
||
`date_extract.py`); score every evidence line against every form with
|
||
digit-confusion-aware fuzzy matching (Levenshtein over digit subsequences,
|
||
with cheap substitutions for known OCR confusions: 0↔8, 1↔7, 5↔6, 2↔7,
|
||
3↔8; also credit partial anchors like a matching year + month pair).
|
||
Candidate score = max over (lines × forms). Return the top candidate iff
|
||
`topScore ≥ SCORE_MIN` **and** `topScore − runnerUpScore ≥ MARGIN_MIN`
|
||
(both thresholds tuned offline — see Testing).
|
||
4. `inferred_fefo` — otherwise: return the candidate with the earliest
|
||
`expiry_date`, flagged.
|
||
|
||
**Output**: `{batchId, expiryDate, source, score, margin}` where
|
||
`source ∈ {single_batch, matched_exact, matched_fragment, inferred_fefo}`.
|
||
|
||
**Also matched**: the `batch_code` string itself is a second fragment-matching
|
||
target — batch codes are often printed adjacent to the date and give an
|
||
independent disambiguation signal for free.
|
||
|
||
**Verification step before building**: confirm the classify server's response
|
||
to the gateway actually carries `text_lines` (the offline capture scripts got
|
||
them from the server, so it almost certainly does); if not, add them to the
|
||
response payload — small change in `config/classify_ocr_server.py`.
|
||
|
||
## Schema & API deltas (on top of stock-feature-plan.md)
|
||
|
||
- `documents` (or the Product-branch metadata): add `expiry_source VARCHAR(20)`
|
||
with `CHECK (expiry_source IN ('single_batch','matched_exact',
|
||
'matched_fragment','inferred_fefo','manual'))` and
|
||
`expiry_match_score REAL NULL`. `'manual'` covers legacy/edited rows.
|
||
- `api/parse/route.ts` (Product branch) and `api/v1/scan-product/route.ts`:
|
||
after `classifyAndMatchProduct()` + the §12.2 in-stock candidate filter, call
|
||
`resolveExpiryFromEvidence()` and include the resolution
|
||
(`resolvedBatch` + `source` + `score`) in the persisted
|
||
`metadata.productScan` and in the response, so the Flutter editor can
|
||
pre-select without any second call (same pattern as task 11.1).
|
||
- `v1/documents/[id]/route.ts` PUT: persist `expiry_source` alongside the
|
||
existing §12.2 `stock_batch_id` decrement. If the client overrides the
|
||
auto-selected batch, source becomes `'manual'`.
|
||
- Review surface: `GET /api/v1/documents?expiry_source=inferred_fefo` filter
|
||
(admin + own-store), powering an optional end-of-day review list.
|
||
|
||
## Flutter deltas (root §10; builds on §9.1–9.4)
|
||
|
||
- **Fast path at cashier**: the §9.4 Product-Scan editor auto-selects the
|
||
resolved batch. When `source` is `single_batch`/`matched_exact`/
|
||
`matched_fragment`, the flow should be confirmable in **one tap** (or
|
||
auto-confirm — decide at pickup with a grill question) with the resolved
|
||
expiry displayed prominently. When `inferred_fefo`, same flow plus a small
|
||
amber "perkiraan" badge — never a blocking prompt (decision 5).
|
||
- **Flag visibility**: history/documents list shows the badge on inferred
|
||
sales; an end-of-day review entry point lists them (uses the new filter).
|
||
Review is optional and zero-checkout-impact by design.
|
||
- **Phase 2 capture mode**: burst capture (N frames over ~1s) in the camera
|
||
layer for mounted use; upload frames together; backend unions evidence lines
|
||
across frames before matching (glare moves between frames — fragments
|
||
accumulate).
|
||
|
||
## Phases
|
||
|
||
**Phase 1 — batch backbone + matcher (the PoC).** Prereqs: §12.1, §12.2,
|
||
§9.1–9.4. New work: `expiry-matcher.ts` + offline tuning harness, schema
|
||
columns, route wiring, Flutter fast-path + badge. No new hardware or models.
|
||
|
||
**Phase 2 — capture upgrade.** Mounted camera at the cashier (a cheap phone
|
||
running the existing Flutter app on a mount is acceptable hardware), burst/
|
||
multi-frame capture, evidence union across frames. Expected to lift fragment
|
||
quality substantially — fixed focus distance + controlled lighting beat
|
||
hand-held single shots.
|
||
|
||
**Phase 3 — on-prem recognizer upgrade.** Fine-tune a small recognition model
|
||
specifically on dot-matrix/inkjet date prints:
|
||
- **Synthetic data**: render dates in dot-matrix/inkjet fonts over pack-like
|
||
backgrounds; augment with dot dropout, scratches, fade, curvature, glare,
|
||
ice speckle. Thousands of labeled crops for free.
|
||
- **Real data flywheel**: every intake stock-entry (staff-typed batch+expiry)
|
||
plus every product-scan photo of that batch = weakly-labeled real training
|
||
pairs accumulating automatically in normal operation. Harvest crops from
|
||
`uploads/` matched to registry values.
|
||
- Train on the RTX 2060 (PaddleOCR rec fine-tune or similar small model);
|
||
deploy as an additional reader in `classify_ocr_server.py`; its lines feed
|
||
the same matcher. Shrinks the `inferred_fefo` residue. **No cloud, ever**
|
||
(decision 7).
|
||
|
||
## Expected accuracy (why this reaches ~90%+ where free OCR cannot)
|
||
|
||
Let p = share of scans where the SKU has exactly one batch in stock (small
|
||
store, fast turnover → p is high, plausibly 0.6–0.8). Those are 100% correct
|
||
by lookup. Of the rest, exact + fragment matching succeeds wherever OCR yields
|
||
*any* usable fragment — on the 79-set evidence, most misses still produced
|
||
fragments; matching 2–3 candidates needs far less signal than free reading.
|
||
The residue is auto-FEFO'd — and FEFO itself is right whenever the customer
|
||
took from the older batch, so even the flagged slice is mostly correct.
|
||
Net: per-sale correctness ~90%+ in Phase 1, rising with Phases 2–3, with
|
||
**zero silent garbage** — every record carries its provenance (`source`).
|
||
|
||
Two honest caveats to monitor:
|
||
- **SKU misclassification poisons the lookup** (wrong SKU → wrong candidates).
|
||
Mitigation: §12.2's in-stock filter shrinks the effective class space to
|
||
what the store actually stocks; near-twin SKU confusion keeps improving via
|
||
reference photos. Track `sku/name` accuracy alongside expiry.
|
||
- **Candidates with near-identical dates** (differ by one digit) can fail the
|
||
margin test → FEFO+flag. Correct behavior; expected to be rare.
|
||
|
||
## Testing plan
|
||
|
||
- **Offline matcher tuning (before any wiring)**: replay the 79 captured OCR
|
||
line-sets (`sources/product_scan_detail_*.json` + fullcap captures) against
|
||
synthetic candidate sets built from the ground-truth labels (1, 2, and 3
|
||
candidates at varying date distances). Tune `SCORE_MIN`/`MARGIN_MIN` for
|
||
zero wrong-candidate picks (a wrong confident match is worse than a flagged
|
||
FEFO). This reuses the frozen benchmark as a matcher benchmark.
|
||
- **Unit tests**: `expiry-matcher` pure-function tests (form rendering,
|
||
confusion-aware scoring, margin logic, FEFO tiebreak) — TS side; Flutter
|
||
pure-logic tests for fast-path/badge state per `source` value.
|
||
- **Live E2E** (per repo convention, against the Docker stack): scan a DO →
|
||
stock-entry with 2 batches of one SKU → product-scan a pack of the older
|
||
batch → verify `matched_*` resolution, decrement of the right batch, and
|
||
`expiry_source` in Postgres; then repeat with an unreadable pack → verify
|
||
`inferred_fefo` + flag, no prompt shown.
|
||
|
||
## Relationship to existing plans
|
||
|
||
- **Extends** `stock-feature-plan.md`: §12.2's "closest-to-OCR-expiry batch
|
||
auto-selected" dropdown becomes the *manual-override* UI behind the new
|
||
automated resolution; the §12.2 decrement/400-on-bad-batch semantics are
|
||
unchanged.
|
||
- Backend tasks: `backend/plans/next-enhancements.md` **§13**.
|
||
- Flutter tasks: root `plans/next-enhancements.md` **§10**.
|
||
- The 79-image frozen benchmark and its capture tooling (accuracy work,
|
||
2026-07-14/15) become the matcher's offline test bed — nothing there is
|
||
wasted by this reframe.
|