Split the expiry-date extraction cascade out of classify_ocr_server.py into
config/date_extract.py (pure regex, importable/testable without loading
models). Three behavioral fixes, offline-regressed against all 79 captured
OCR line-sets and sanity-verified live on the two target images:
- Guard the 012/112 month-misrecognition cleanup rules: they fired on
perfectly valid dates too (BB 01122026 = 01/12/2026 matches 0+112+2026)
and mangled them into 7-digit junk that parsed as 00/22/26. Skipped when
the line already contains a valid date. Fixes image 11.
- Exclude store price-tag lines (Printed:.., Rp...) from the keyword-less
stages so a shelf label's print timestamp can't shadow the real date
printed on the package. Fixes image 71 (09/04/2027).
- Validity-gate the lenient stage (day<=31, month<=12, year 2020-2039) so
garbled digit runs return empty instead of junk like 1/3/06 or 11/1/01.
Also: clamp /probe-ocr crop box to image bounds (PIL pads out-of-bounds
crops into a gigapixel canvas -> DecompressionBombError), and update
CLAUDE.md's Graphify section - the global Claude Code skill integration was
installed 2026-07-15 at the user's explicit request.
Full-batch measurement of these fixes (expected 79.7% -> ~80.6%) is still
pending - the run was stopped twice at the user's end; re-run
scripts/accuracy-check-scan.mts next session before building on this.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gr6HH7JrdsXX8AARejQboM