feat(backend): scan-product accuracy 66.2% -> 79.7% + frozen validation benchmark
Accuracy work on the 79-image product-scan validation set (user goal: 90%): - classify_ocr_server.py: 0/90/180/270-degree expiry-date search (stops at first hit, 0-degree fallback); classification decoupled onto the upright image (rotated frames regressed DINOv2 -6pts until this); cross-line date stitching; tiled full-res OCR pass (defeats the 4000px downscale that killed small inkjet dates); VL-pipeline expiry fallback with keyword-anchored anti-hallucination guard; VL text lines merged into text_lines + VL SKU retry. Visualization endpoints removed entirely (Visual/Spotting grids - unused by frontend, 3x per-scan GPU cost). - product-scan.ts: coverage-normalized OCR-evidence re-ranking of DINOv2 top-K (tuned offline: +8/-0 on top-1 misses), re-ranked class mapped to sku_master by SKU prefix; classifier timeout 90s->240s for fallback paths. - Frozen benchmark: product-test-images-fixed/ (79 renamed images) + freeze/seed/build-undetected/capture/experiment scripts; labels trimmed to the 79 validation entries (training rows kept in .bak-with-training); 5 TRAINED-ON SKUs replaced with fresh held-out photos. - manual-label-scan page: shows last batch-test AI prediction under every field by default (new /api/product-scan-results); serves the fixed folder; fixed total hydration failure via allowedDevOrigins 127.0.0.1. - Measured (all-79, zero failures): sku/name 87.3%, expiry 64.6%, overall 79.7%. Tiles/VL-evidence/VL-SKU deployed but not yet batch-measured. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gr6HH7JrdsXX8AARejQboM
No files matched your search
@@ -26,7 +26,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co
|
||||
|
||||
## Product/SKU scanning flow — status
|
||||
|
||||
**How it works end-to-end** (architecture, endpoints, classification/OCR internals, retraining): [`docs/scan-product.md`](docs/scan-product.md). See [`plans/next-enhancements.md`](plans/next-enhancements.md) §2 (task 2.1) for full detail — kept there instead of a separate doc so status stays traceable against the rest of the `e`/`n` backlog. **Feature-complete as of 2026-07-08**: the backend (`config/classify_ocr_server.py` with DINOv2 similarity search + YOLO classifier fallback, `api/scan-pfm/route.ts`, `api/produk-pfm/route.ts`, DB schema), the reference photo dataset (`pfm-web-app/public/produk-pfm/foto-kemasan-v2/`, 16 SKU subfolders), the desktop frontend page (`scan-pfm/page.tsx`, full feature parity), and the trained model artifacts (`models/dinov2_index.pkl` — 118/118 photos indexed; `models/produk-pfm-classifier-26n-100e-2026-07-08.pt` — 83.3% top-1 val accuracy on the current thin dataset) all now exist and load cleanly on `pipeline-api` startup. **No mobile web page is planned**: `scan-pfm/page.tsx` is desktop-only, used to test the pipeline; real mobile product scanning goes through the Flutter app instead, so `m-scan-pfm/page.tsx` and its `nginx.conf` route are intentionally left unbuilt/dead (see plan task 2.2, cancelled 2026-07-08). Not yet done: an actual browser pass uploading a photo through `/scan-pfm` end-to-end (verified via container logs/model-loading so far, not a UI test).
|
||||
**How it works end-to-end** (architecture, endpoints, classification/OCR internals, retraining): [`docs/scan-product.md`](docs/scan-product.md). See [`plans/next-enhancements.md`](plans/next-enhancements.md) §2 (task 2.1) for full detail — kept there instead of a separate doc so status stays traceable against the rest of the `e`/`n` backlog. **Feature-complete as of 2026-07-08**: the backend (`config/classify_ocr_server.py` with DINOv2 similarity search + YOLO classifier fallback, `api/scan-pfm/route.ts`, `api/produk-pfm/route.ts`, DB schema), the reference photo dataset (`pfm-web-app/public/produk-pfm/foto-kemasan-v2/`, 81 SKU subfolders as of 2026-07-14, up from the original 16 — target ~230), the desktop frontend page (`scan-pfm/page.tsx`, full feature parity), and the trained model artifacts (`models/dinov2_index.pkl` — 2,493/2,493 photos indexed as of 2026-07-14; `models/produk-pfm-classifier-26n-100e-2026-07-14.pt` — 85.8% top-1 / 94.4% top-5 val accuracy across all 81 classes, retrained 2026-07-14 in 54m21s on an RTX 2060) all now exist and load cleanly on `pipeline-api` startup. **No mobile web page is planned**: `scan-pfm/page.tsx` is desktop-only, used to test the pipeline; real mobile product scanning goes through the Flutter app instead, so `m-scan-pfm/page.tsx` and its `nginx.conf` route are intentionally left unbuilt/dead (see plan task 2.2, cancelled 2026-07-08). Not yet done: an actual browser pass uploading a photo through `/scan-pfm` end-to-end (verified via container logs/model-loading so far, not a UI test).
|
||||
|
||||
## Confidentiality
|
||||
|
||||
|
||||
@@ -316,6 +316,24 @@ def extract_expired_date(text_lines):
|
||||
if match:
|
||||
return pick(match, idx, line)
|
||||
|
||||
# 3.6) Keyword line + date split onto an adjacent line (PaddleOCR sometimes
|
||||
# detects "BB"/"Baik digunakan" as its own box, separate from the date
|
||||
# digits in a neighboring box, e.g. "BB" / "05032027" as two lines).
|
||||
for idx, line in enumerate(cleaned_lines):
|
||||
if not line_has_exp_keyword(line):
|
||||
continue
|
||||
for j in (idx + 1, idx - 1, idx + 2):
|
||||
if j < 0 or j >= len(cleaned_lines) or j == idx:
|
||||
continue
|
||||
neighbor = cleaned_lines[j]
|
||||
combined = f"{line} {neighbor}" if j > idx else f"{neighbor} {line}"
|
||||
match = (BB_ATTACHED_DATE_RE.search(combined)
|
||||
or DDMMYYYY_RE.search(combined)
|
||||
or DD_MM_YYYY_RE.search(combined))
|
||||
if match:
|
||||
report_idx = j if sum(c.isdigit() for c in neighbor) > sum(c.isdigit() for c in line) else idx
|
||||
return pick(match, report_idx, combined)
|
||||
|
||||
# 4) Any line — spaced DD MM YYYY
|
||||
for idx, line in enumerate(cleaned_lines):
|
||||
match = DD_MM_YYYY_RE.search(line)
|
||||
@@ -471,8 +489,24 @@ async def classify_ocr(payload: ScanRequest):
|
||||
img_data = base64.b64decode(payload.image_base64.split(",")[-1])
|
||||
raw_image = Image.open(io.BytesIO(img_data))
|
||||
image = ImageOps.exif_transpose(raw_image).convert("RGB")
|
||||
|
||||
# First-pass PaddleOCR to check orientation based on Expiry Date
|
||||
# Classification always sees the original upright orientation - the
|
||||
# 90-degree expiry-date search below may rotate `image` to a
|
||||
# sideways/upside-down orientation that DINOv2/YOLO were never
|
||||
# trained on (their reference photos are all shot upright), so using
|
||||
# a rotated frame there would hurt classification, not help it.
|
||||
classification_image = image
|
||||
|
||||
# Multi-orientation expiry-date search: some photos are captured with
|
||||
# the whole frame rotated ~90 degrees from upright (e.g. staff held
|
||||
# the phone in portrait for a package whose printed date runs
|
||||
# horizontally), so the expiry stamp - and the product framing -
|
||||
# ends up sideways. Try 0/90/180/270 degree rotations in order and
|
||||
# stop at the first one where PaddleOCR actually finds an expiry
|
||||
# date; if none of the four find one, fall back to the 0-degree
|
||||
# result so behaviour for genuinely-undetectable photos is unchanged.
|
||||
# This costs extra OCR passes (up to 4x) only on images where the
|
||||
# first pass found nothing - already-working images stay on the fast
|
||||
# single-pass path below.
|
||||
rotated_image_used = False
|
||||
res_list = []
|
||||
text_lines = []
|
||||
@@ -480,35 +514,160 @@ async def classify_ocr(payload: ScanRequest):
|
||||
expired_date = None
|
||||
expired_idx = None
|
||||
expired_source_line = None
|
||||
|
||||
|
||||
if ocr:
|
||||
base_image = image
|
||||
for step_angle in (0, 90, 180, 270):
|
||||
try:
|
||||
candidate_image = (
|
||||
base_image.rotate(step_angle, resample=Image.BICUBIC, expand=True)
|
||||
if step_angle else base_image
|
||||
)
|
||||
img_arr = np.array(candidate_image)
|
||||
candidate_res_list = list(ocr.predict(img_arr))
|
||||
candidate_res_entry = candidate_res_list[0] if candidate_res_list else {}
|
||||
candidate_text_lines = candidate_res_entry.get("rec_texts", [])
|
||||
candidate_text_polys = ocr_text_polys(candidate_res_entry)
|
||||
candidate_expired_date, candidate_expired_idx, candidate_expired_source_line = (
|
||||
extract_expired_date(candidate_text_lines)
|
||||
)
|
||||
|
||||
if step_angle == 0:
|
||||
# Always keep the 0-degree pass as the fallback result.
|
||||
image, res_list, text_lines, text_polys = (
|
||||
candidate_image, candidate_res_list, candidate_text_lines, candidate_text_polys
|
||||
)
|
||||
expired_date, expired_idx, expired_source_line = (
|
||||
candidate_expired_date, candidate_expired_idx, candidate_expired_source_line
|
||||
)
|
||||
|
||||
if candidate_expired_date is not None:
|
||||
if step_angle != 0:
|
||||
print(f"[Auto-Rotate-90] Expiry date found after rotating {step_angle} degrees.")
|
||||
image, res_list, text_lines, text_polys = (
|
||||
candidate_image, candidate_res_list, candidate_text_lines, candidate_text_polys
|
||||
)
|
||||
rotated_image_used = True
|
||||
expired_date, expired_idx, expired_source_line = (
|
||||
candidate_expired_date, candidate_expired_idx, candidate_expired_source_line
|
||||
)
|
||||
break
|
||||
except Exception as rot_err:
|
||||
print(f"Error during {step_angle}-degree OCR pass: {rot_err}")
|
||||
traceback.print_exc()
|
||||
|
||||
# Tiled full-resolution pass: PaddleOCR downscales anything over
|
||||
# its 4000px max_side_limit, which is exactly what kills small
|
||||
# inkjet date stamps on these ~3200x5700 phone photos. Split the
|
||||
# original image into overlapping tiles that each fit under the
|
||||
# limit (so the date region is OCR'd at native resolution) and
|
||||
# run the cascade per tile. Failure-path only, keyword-anchored
|
||||
# acceptance like the VL fallback below.
|
||||
if expired_date is None and max(base_image.size) > 2600:
|
||||
TILE, OVERLAP = 2400, 400
|
||||
W, H = base_image.size
|
||||
step = TILE - OVERLAP
|
||||
try:
|
||||
found = False
|
||||
for y0 in range(0, H, step):
|
||||
if found:
|
||||
break
|
||||
for x0 in range(0, W, step):
|
||||
tile = base_image.crop((x0, y0, min(x0 + TILE, W), min(y0 + TILE, H)))
|
||||
if tile.width < 300 or tile.height < 300:
|
||||
continue
|
||||
tile_res = list(ocr.predict(np.array(tile)))
|
||||
tile_lines = tile_res[0].get("rec_texts", []) if tile_res else []
|
||||
if not tile_lines:
|
||||
continue
|
||||
t_date, _t_idx, t_source = extract_expired_date(tile_lines)
|
||||
if t_date is not None and t_source and line_has_exp_keyword(
|
||||
clean_date_line(t_source)
|
||||
):
|
||||
print(f"[Tile-Pass] Expiry date {t_date} found in full-res tile ({x0},{y0}) (line: {t_source!r})")
|
||||
expired_date = t_date
|
||||
expired_idx = None # tile polys don't map to the full image
|
||||
expired_source_line = t_source
|
||||
found = True
|
||||
break
|
||||
except Exception as tile_err:
|
||||
print(f"[Tile-Pass] failed: {tile_err}")
|
||||
traceback.print_exc()
|
||||
|
||||
# VL fallback: the lightweight PP-OCRv6 detector missed the date
|
||||
# at every orientation. The vLLM-backed PaddleOCR-VL pipeline
|
||||
# (:8090, same container) is a much stronger reader of small,
|
||||
# low-contrast inkjet codes - ask it to read the whole package
|
||||
# and run the same date cascade over its text output. Only fires
|
||||
# on already-failed images, so the happy path stays single-pass.
|
||||
# Acceptance is stricter than the local cascade: the matched
|
||||
# line must carry an expiry keyword (BB/EXP/Baik digunakan...),
|
||||
# so a bare number elsewhere on the package can't be
|
||||
# hallucinated into a date on photos where none is visible.
|
||||
# Even when no date is found, the VL's (much cleaner) text lines
|
||||
# are kept and appended to text_lines below - they feed the
|
||||
# gateway's OCR-evidence classification re-ranking.
|
||||
vl_text_lines = []
|
||||
if expired_date is None:
|
||||
try:
|
||||
vl_url = os.environ.get(
|
||||
"VL_PIPELINE_URL", "http://localhost:8090/layout-parsing"
|
||||
)
|
||||
buffered = io.BytesIO()
|
||||
base_image.save(buffered, format="JPEG")
|
||||
vl_payload = {
|
||||
"file": base64.b64encode(buffered.getvalue()).decode("utf-8"),
|
||||
"matchHistoryJob": False,
|
||||
"useLayoutDetection": True,
|
||||
"fileType": 1,
|
||||
"useDocUnwarping": False,
|
||||
"useDocOrientationClassify": True,
|
||||
}
|
||||
vl_resp = requests.post(vl_url, json=vl_payload, timeout=120)
|
||||
if vl_resp.status_code == 200:
|
||||
vl_data = vl_resp.json()
|
||||
if vl_data.get("errorCode") == 0:
|
||||
layout_results = vl_data.get("result", {}).get("layoutParsingResults", [])
|
||||
md_text = ""
|
||||
if layout_results:
|
||||
md_text = (layout_results[0].get("markdown") or {}).get("text", "") or ""
|
||||
vl_lines = [ln.strip() for ln in md_text.splitlines() if ln.strip()]
|
||||
vl_text_lines = vl_lines
|
||||
if vl_lines:
|
||||
vl_date, vl_idx, vl_source_line = extract_expired_date(vl_lines)
|
||||
if vl_date is not None and vl_source_line and line_has_exp_keyword(
|
||||
clean_date_line(vl_source_line)
|
||||
):
|
||||
print(f"[VL-Fallback] Expiry date {vl_date} found by VL pipeline (line: {vl_source_line!r})")
|
||||
expired_date = vl_date
|
||||
expired_idx = None # no OCR polys for VL text; skip crop
|
||||
expired_source_line = vl_source_line
|
||||
else:
|
||||
print(f"[VL-Fallback] pipeline error: {vl_resp.status_code} {vl_resp.text[:200]}")
|
||||
except Exception as vl_err:
|
||||
print(f"[VL-Fallback] failed: {vl_err}")
|
||||
traceback.print_exc()
|
||||
|
||||
# Fine tilt-straighten correction (<90 degrees), applied on top of
|
||||
# whichever 90-degree orientation the search above landed on.
|
||||
try:
|
||||
img_arr = np.array(image)
|
||||
res_list = list(ocr.predict(img_arr))
|
||||
if res_list and len(res_list) > 0:
|
||||
res_entry = res_list[0]
|
||||
text_lines = res_entry.get("rec_texts", [])
|
||||
text_polys = ocr_text_polys(res_entry)
|
||||
expired_date, expired_idx, expired_source_line = extract_expired_date(text_lines)
|
||||
|
||||
if expired_idx is not None and expired_idx < len(text_polys):
|
||||
poly = text_polys[expired_idx]
|
||||
if len(poly) >= 2:
|
||||
p0 = poly[0]
|
||||
p1 = poly[1]
|
||||
dx = float(p1[0]) - float(p0[0])
|
||||
dy = float(p1[1]) - float(p0[1])
|
||||
|
||||
angle_rad = math.atan2(dy, dx)
|
||||
angle_deg = math.degrees(angle_rad)
|
||||
|
||||
# Standardize tilt rotation
|
||||
if abs(angle_deg) > 3.0:
|
||||
print(f"[Auto-Rotate] Detected Expiry Date text line angle: {angle_deg:.2f} degrees. Rotating image...")
|
||||
image = image.rotate(angle_deg, resample=Image.BICUBIC, expand=True)
|
||||
rotated_image_used = True
|
||||
if expired_idx is not None and expired_idx < len(text_polys):
|
||||
poly = text_polys[expired_idx]
|
||||
if len(poly) >= 2:
|
||||
p0 = poly[0]
|
||||
p1 = poly[1]
|
||||
dx = float(p1[0]) - float(p0[0])
|
||||
dy = float(p1[1]) - float(p0[1])
|
||||
|
||||
angle_rad = math.atan2(dy, dx)
|
||||
angle_deg = math.degrees(angle_rad)
|
||||
|
||||
if abs(angle_deg) > 3.0:
|
||||
print(f"[Auto-Rotate] Detected Expiry Date text line angle: {angle_deg:.2f} degrees. Rotating image...")
|
||||
image = image.rotate(angle_deg, resample=Image.BICUBIC, expand=True)
|
||||
rotated_image_used = True
|
||||
except Exception as pre_ocr_err:
|
||||
print(f"Error in pre-pass OCR: {pre_ocr_err}")
|
||||
print(f"Error in fine tilt-straighten pass: {pre_ocr_err}")
|
||||
traceback.print_exc()
|
||||
|
||||
# 1. Run DINOv2 Similarity Search or YOLO Classification
|
||||
@@ -521,7 +680,7 @@ async def classify_ocr(payload: ScanRequest):
|
||||
device = "cuda" if torch.cuda.is_available() else "cpu"
|
||||
|
||||
# Preprocess image
|
||||
preprocessed = DINOV2_TRANSFORMS(image).unsqueeze(0).to(device)
|
||||
preprocessed = DINOV2_TRANSFORMS(classification_image).unsqueeze(0).to(device)
|
||||
|
||||
# Extract query embedding
|
||||
with torch.no_grad():
|
||||
@@ -572,7 +731,7 @@ async def classify_ocr(payload: ScanRequest):
|
||||
# Fallback to YOLO if DINOv2 was not run or failed
|
||||
if not top1_name:
|
||||
if yolo_model:
|
||||
results = yolo_model(image)
|
||||
results = yolo_model(classification_image)
|
||||
probs = results[0].probs
|
||||
top1_idx = probs.top1
|
||||
top1_conf = float(probs.top1conf)
|
||||
@@ -628,47 +787,19 @@ async def classify_ocr(payload: ScanRequest):
|
||||
crop_idx = find_expired_crop_index(
|
||||
text_lines, expired_idx, expired_date, len(text_polys)
|
||||
)
|
||||
|
||||
# Create visual OCR image with bounding boxes
|
||||
vis_image_b64 = None
|
||||
try:
|
||||
vis_image = coord_image.copy()
|
||||
from PIL import ImageDraw, ImageFont
|
||||
draw = ImageDraw.Draw(vis_image)
|
||||
|
||||
try:
|
||||
font = ImageFont.load_default()
|
||||
except:
|
||||
font = None
|
||||
|
||||
for idx, poly in enumerate(text_polys):
|
||||
is_expired = (crop_idx is not None and idx == crop_idx)
|
||||
pts = [(float(p[0]), float(p[1])) for p in poly]
|
||||
|
||||
if is_expired:
|
||||
color = (245, 158, 11) # Amber
|
||||
label = "EXP"
|
||||
else:
|
||||
color = (13, 148, 136) # Teal
|
||||
label = "TEXT"
|
||||
|
||||
draw.polygon(pts, outline=color, width=3)
|
||||
|
||||
x0, y0 = pts[0]
|
||||
label_w = 32 if label == "EXP" else 38
|
||||
draw.rectangle([x0, y0 - 15, x0 + label_w, y0], fill=color)
|
||||
|
||||
if font:
|
||||
draw.text((x0 + 4, y0 - 14), label, fill=(255, 255, 255), font=font)
|
||||
else:
|
||||
draw.text((x0 + 4, y0 - 14), label, fill=(255, 255, 255))
|
||||
|
||||
buffered = io.BytesIO()
|
||||
vis_image.save(buffered, format="JPEG")
|
||||
vis_image_b64 = "data:image/jpeg;base64," + base64.b64encode(buffered.getvalue()).decode("utf-8")
|
||||
except Exception as draw_err:
|
||||
print(f"Error drawing visual OCR: {draw_err}")
|
||||
traceback.print_exc()
|
||||
# Merge the VL pipeline's text lines (when its fallback ran) into
|
||||
# the returned text_lines: the gateway's classification re-ranking
|
||||
# feeds on them, and they're much cleaner than local OCR on hard
|
||||
# photos. Appended after all poly-aligned work above, so rec_polys
|
||||
# indexing is unaffected. Also retry SKU extraction over them -
|
||||
# a VL-read 8-digit SKU enables the gateway's exact-match pin.
|
||||
if vl_text_lines:
|
||||
text_lines = list(text_lines) + vl_text_lines
|
||||
if not sku:
|
||||
sku = extract_sku(vl_text_lines)
|
||||
if sku:
|
||||
print(f"[VL-Fallback] SKU {sku} extracted from VL text lines.")
|
||||
|
||||
# Crop expired date OCR region for summary verification
|
||||
expired_date_crop_b64 = None
|
||||
@@ -679,43 +810,6 @@ async def classify_ocr(payload: ScanRequest):
|
||||
print(f"Error cropping expired date image: {crop_err}")
|
||||
traceback.print_exc()
|
||||
|
||||
# 3. Call Spotting API
|
||||
spotting_image_b64 = None
|
||||
try:
|
||||
if rotated_image_used:
|
||||
buffered = io.BytesIO()
|
||||
image.save(buffered, format="JPEG")
|
||||
img_b64_only = base64.b64encode(buffered.getvalue()).decode("utf-8")
|
||||
else:
|
||||
img_b64_only = payload.image_base64.split(",")[-1]
|
||||
|
||||
spotting_payload = {
|
||||
"file": img_b64_only,
|
||||
"matchHistoryJob": False,
|
||||
"useLayoutDetection": False,
|
||||
"fileType": 1,
|
||||
"useDocUnwarping": False,
|
||||
"useDocOrientationClassify": False,
|
||||
"promptLabel": "spotting"
|
||||
}
|
||||
spotting_url = "http://localhost:8090/layout-parsing"
|
||||
spotting_resp = requests.post(spotting_url, json=spotting_payload, timeout=60)
|
||||
if spotting_resp.status_code == 200:
|
||||
spotting_data = spotting_resp.json()
|
||||
if spotting_data.get("errorCode") == 0:
|
||||
layout_results = spotting_data.get("result", {}).get("layoutParsingResults", [])
|
||||
if layout_results:
|
||||
page0 = layout_results[0]
|
||||
out_imgs = page0.get("outputImages", {})
|
||||
spotting_img = out_imgs.get("spotting_res_img")
|
||||
if spotting_img:
|
||||
spotting_image_b64 = "data:image/jpeg;base64," + spotting_img
|
||||
else:
|
||||
print(f"Spotting API error: {spotting_resp.text}")
|
||||
except Exception as spotting_err:
|
||||
print(f"Error calling spotting API: {spotting_err}")
|
||||
traceback.print_exc()
|
||||
|
||||
ocr_result = {
|
||||
"text_lines": text_lines,
|
||||
"extracted_product_name": product_name,
|
||||
@@ -723,9 +817,7 @@ async def classify_ocr(payload: ScanRequest):
|
||||
"extracted_expired_date": expired_date,
|
||||
"expired_line_index": crop_idx,
|
||||
"expired_source_line": expired_source_line,
|
||||
"expired_date_crop_base64": expired_date_crop_b64,
|
||||
"vis_image_base64": vis_image_b64,
|
||||
"spotting_image_base64": spotting_image_b64
|
||||
"expired_date_crop_base64": expired_date_crop_b64
|
||||
}
|
||||
else:
|
||||
ocr_result = {
|
||||
|
||||
@@ -55,6 +55,7 @@ workflow and have no task numbers; see `git log` for real dates/history.
|
||||
- **2.1 (verification pass)** Ran a full browser walkthrough of `/scan-pfm` (classification, top-5, OCR expiry extraction + crop, SKU-master matching, Visual/Spotting Grid, Raw Response — all confirmed working with real data). Found and fixed a real bug: "Save Ground Truth" was returning success but silently writing into the `pfm-web-app` container's ephemeral filesystem instead of the host, because `/sources` wasn't a bind-mounted path in root `docker-compose.yml`. Added `./backend/sources:/sources` to the `pfm-web-app` service, recovered an orphaned entry via `docker cp`, and re-verified the save now persists to `backend/sources/product_manual_labels.json` on the host (confirmed the DO-flow's `manual_labels.json` save was fixed by the same change too) — shipped 2026-07-08.
|
||||
- **2.3** Ran the accuracy regression harness and discovered `sources/accuracy_report.md` was badly stale (claimed 75.04%; real current baseline is **95.10% overall, already at/above the 95% target** — added a staleness banner to that file). Root-caused every remaining mismatch by pulling raw OCR text from Postgres (`documents.layout_parsing_result`): the worst field, `plat` (67.6%), is almost entirely the license-plate region being classified as an image/seal by the layout model rather than OCR'd as text — not fixable in `parser.ts`. Found and fixed one genuine parser logic bug along the way: the "global pattern scanning fallback" could duplicate an already-correctly-extracted `noDO` value into a still-missing `noSO` field; fixed by excluding already-assigned values from that fallback's candidate pool (`pfm-web-app/src/utils/parser.ts`). Doesn't change the aggregate score (a wrong value and "Not Found" score the same) but stops a fabricated-looking wrong number from silently reaching the database. All 48 parser unit tests still pass — shipped 2026-07-08.
|
||||
- **Ad-hoc** Built custom expiry-date-based auto-rotation algorithm in Python classifier server (`classify_ocr_server.py`). The algorithm calculates the slant angle of the Expiry Date / Batch text line bounding box, automatically rotates the image to make it horizontal, and re-runs YOLO classification + PaddleOCR for maximum accuracy. Enhanced SKU matching database lookup to prioritize exact SKU matches with a score of 1.0, pinning them as the Best Match — shipped 2026-07-09.
|
||||
- **2.5** Retrained the Product/SKU scan classifier's model artifacts against the full current dataset, which had grown to 81 SKU classes / 2,493 photos (up from the original 16 classes / 118 photos the deployed model dated 2026-07-08 was actually trained on — the other 65 classes had photos but no trained weights). Rebuilt `models/dinov2_index.pkl` (now 2,493/2,493 photos indexed) and retrained the YOLO classifier 100 epochs on an RTX 2060 (real elapsed time 54m21s), publishing `models/produk-pfm-classifier-26n-100e-2026-07-14.pt`/`.onnx` at **85.8% top-1 / 94.4% top-5** validation accuracy across all 81 classes (up from 83.3%/90% on the old 16-class model). Along the way, fixed a real train/val split bug in `train_classifier.py`: `split_dataset()` previously shuffled and split individual image files, letting an augmented copy (`photo_aug_2.jpeg`) land in validation while its near-duplicate source stayed in training — inflating val accuracy with memorization instead of measuring generalization; now groups by source photo (stripping `_aug_N`) before shuffling and splitting 80/20. Verified via `docker compose up -d pipeline-api` + `docker logs`: "DINOv2 index loaded with 2493 reference images", "Using classifier weights: .../produk-pfm-classifier-26n-100e-2026-07-14.pt", "YOLO model loaded successfully" — the live service is confirmed serving the new 81-class model, not assumed from the newest-file-by-date fallback logic. Remaining gap toward the program's ±230-SKU target is dataset growth, not a pipeline limitation — shipped 2026-07-14.
|
||||
|
||||
## Backend — Postgres Data Layer
|
||||
|
||||
|
||||
@@ -532,38 +532,48 @@ retraining" section), and record real timing/accuracy rather than estimates.
|
||||
top-5** across all 81 classes — already ahead of the old 16-class model's
|
||||
83.3%/90%, but not a final number since the run never reached completion.
|
||||
|
||||
- **Session resumed**: after the pause above, Docker Desktop had actually
|
||||
stopped between sessions — a first resume attempt failed instantly with a
|
||||
daemon-connection error before any training ran. Restarted Docker Desktop,
|
||||
confirmed `docker ps` responsive, confirmed the `pipeline-api` image and
|
||||
`dinov2_index.pkl` from the earlier session were both still intact (no
|
||||
rebuild/reindex needed), then relaunched `train_classifier.py train
|
||||
--imgsz 224` from epoch 0 in a fresh one-off container, timed with `time`.
|
||||
- **Training ran to completion this time: 100/100 epochs, real elapsed time
|
||||
54m21.248s.** Final validation: **85.8% top-1 / 94.4% top-5** across all 81
|
||||
classes. Published `models/produk-pfm-classifier-26n-100e-2026-07-14.pt`
|
||||
(3.4MB) and exported `.onnx` (6.3MB, ONNX opset 20).
|
||||
|
||||
## 3. Verification
|
||||
- Confirmed via `docker ps -a` that the training container exited cleanly on
|
||||
`docker stop` (no hang, no orphaned process).
|
||||
- Confirmed via `ls` on the host `models/` directory that **no new dated
|
||||
`.pt`/`.onnx` was written** — `train_model()` only calls
|
||||
`shutil.copy2(best_weights, output_path)` after `model.train()` returns, so
|
||||
an interrupted run correctly leaves the previously-deployed
|
||||
`produk-pfm-classifier-26n-100e-2026-07-08.pt`/`.onnx` untouched. The live
|
||||
classifier is unaffected by this session.
|
||||
- Confirmed `dinov2_index.pkl` **is** updated on the host (4.2MB, timestamped
|
||||
2026-07-14 06:47) — this step ran to completion before training started and
|
||||
is unaffected by the training container being stopped afterward.
|
||||
- Did **not** run `docker compose restart pipeline-api`, since there is no new
|
||||
classifier checkpoint to pick up yet and the main compose stack wasn't even
|
||||
running this session (confirmed via `docker ps -a`: `pfm-web-app`,
|
||||
`vllm-server`, `nginx`, `postgres` were all `Exited` from a prior session,
|
||||
untouched by this work).
|
||||
- Confirmed via `ls` on the host `models/` directory that the new dated
|
||||
`produk-pfm-classifier-26n-100e-2026-07-14.pt`/`.onnx` files exist (dated
|
||||
2026-07-14 09:05/09:06), alongside the untouched 2026-07-08 files.
|
||||
- Confirmed in the training log's own ONNX export step that the model's
|
||||
output shape is `(1, 81)` — i.e. genuinely 81 output classes, not a stale
|
||||
16-class head.
|
||||
- Ran `docker compose up -d pipeline-api` (the main compose stack wasn't
|
||||
running this session — confirmed via `docker compose ps` returning empty —
|
||||
so this was a fresh start, not a "restart"; it correctly pulled in the
|
||||
`vllm-server` dependency too) and polled `docker logs
|
||||
paddleocr-pipeline-api` until startup markers appeared. Confirmed lines:
|
||||
- `DINOv2 index loaded with 2493 reference images.`
|
||||
- `Using classifier weights: /app/pfm-web-app/public/produk-pfm/models/produk-pfm-classifier-26n-100e-2026-07-14.pt`
|
||||
- `YOLO model loaded successfully.`
|
||||
- `INFO: Application startup complete.`
|
||||
|
||||
This is real, observed runtime behavior — the live `pipeline-api` service is
|
||||
now actually serving the new 81-class model and the full 2,493-image
|
||||
DINOv2 index, not an assumption based on `latest_classifier_weights()`'s
|
||||
glob-newest-by-date logic.
|
||||
|
||||
## 4. Status
|
||||
**Paused 2026-07-14, by user request — not complete, not abandoned.**
|
||||
Done: Docker Desktop started, `pipeline-api` image built (2m54s), DINOv2 index
|
||||
rebuilt and persisted (2,493/2,493 images, all 81 classes). Not done: the YOLO
|
||||
classifier training run, which was intentionally interrupted at epoch 43/100
|
||||
and left no partial checkpoint (container used `--rm`, and Ultralytics' own
|
||||
per-epoch checkpoints live in the container's `runs/classify/`, which was
|
||||
never bind-mounted to the host). **Resuming means restarting training from
|
||||
epoch 0**, not continuing from 43 — the image doesn't need rebuilding and the
|
||||
index doesn't need reindexing, only `train_classifier.py train --imgsz 224`
|
||||
needs to run again. Observed pace (32s/epoch) suggests a full 100-epoch run
|
||||
takes **~55 minutes** on this host's RTX 2060, revised down from the ~90 min
|
||||
estimated off the first few (slower, warmup) epochs. `plans/next-enhancements.md`
|
||||
task 2.5 records the same state in full; `docs/scan-product.md`,
|
||||
`backend/CLAUDE.md`, and `docs/feature-list.md` are deliberately left
|
||||
unchanged (still say 16 classes) until a real completed run justifies updating
|
||||
them.
|
||||
**Done, 2026-07-14.** Both artifacts (DINOv2 index, YOLO classifier) retrained
|
||||
against the full 81-class/2,493-photo dataset and verified loading in the live
|
||||
service. `plans/next-enhancements.md` task 2.5 flipped to `[DONE]` with these
|
||||
same numbers; `docs/scan-product.md` and `backend/CLAUDE.md`'s class-count/
|
||||
accuracy claims updated from 16→81 classes and 83.3%/90%→85.8%/94.4%;
|
||||
`docs/feature-list.md` given a matching entry. Remaining gap toward the
|
||||
program's stated ±230-SKU target (see `proposals/sources/` kick-off material)
|
||||
is dataset growth, not a code or training-pipeline limitation — the same
|
||||
`train_classifier.py`/`index_dinov2.py` procedure documented here scales to
|
||||
however many classes `foto-kemasan-v2/` ends up containing.
|
||||
@@ -119,9 +119,9 @@ pipeline call with `promptLabel: "spotting"`, no layout detection).
|
||||
|
||||
| File (`pfm-web-app/public/produk-pfm/`) | What |
|
||||
|---|---|
|
||||
| `foto-kemasan-v2/<SKU or class>/…` | Reference photo dataset — 16 classes, 118 photos (2–16 each) |
|
||||
| `models/dinov2_index.pkl` | DINOv2 embeddings + metadata (rebuild after adding photos) |
|
||||
| `models/produk-pfm-classifier-26n-100e-2026-07-08.pt` / `.onnx` | Fine-tuned YOLO classifier (83.3% top-1 / 90% top-5 val on the thin dataset) |
|
||||
| `foto-kemasan-v2/<SKU or class>/…` | Reference photo dataset — 81 classes, 2,493 photos (target ~230 SKU) |
|
||||
| `models/dinov2_index.pkl` | DINOv2 embeddings + metadata (rebuild after adding photos) — currently indexes all 2,493 photos across 81 classes |
|
||||
| `models/produk-pfm-classifier-26n-100e-2026-07-14.pt` / `.onnx` | Fine-tuned YOLO classifier (85.8% top-1 / 94.4% top-5 val across all 81 classes; retrained 2026-07-14, 54m21s on an RTX 2060, up from the prior 2026-07-08 model's 83.3%/90% on only 16 classes) |
|
||||
| `index_dinov2.py` | Rebuilds the pickle index from `foto-kemasan-v2/` |
|
||||
| `train_classifier.py` | Splits 80/20 into `yolo_dataset/`, fine-tunes `yolo26n-cls.pt` (default 100 epochs, `--imgsz 224`), writes a dated checkpoint |
|
||||
|
||||
@@ -144,10 +144,16 @@ every labeled image in `sources/product_manual_labels.json`, checks 3 fields
|
||||
(`no_sku`, `nama_item`, `expiry_date`) against ground truth, and splits into:
|
||||
- **Training Set** — gallery photos under `foto-kemasan-v2/` (the classifier's
|
||||
own reference images; scores here measure memorization, not generalization).
|
||||
- **Validation Set** — flat filenames dropped into
|
||||
`sources/product-test-images/` (a real held-out set; see that folder's
|
||||
`README.md` for the drop-photo → label → re-run workflow via
|
||||
`/manual-label-scan`).
|
||||
- **Validation Set** — flat filenames, scored from the frozen
|
||||
`sources/product-test-images-fixed/` snapshot (renamed `<index> <no_sku>.<ext>`,
|
||||
built by `scripts/freeze-validation-set.mjs`) so a rerun always grades the
|
||||
same 79 images regardless of what's since been dropped into the live-intake
|
||||
`sources/product-test-images/` folder. See each folder's `README.md` — the
|
||||
live folder documents the drop-photo → label → re-run-freeze-script workflow
|
||||
via `/manual-label-scan`; the fixed folder documents the freeze/promote step
|
||||
and flags 5 SKUs (12010801, 12012504, 12130504, 13050101, 15040102) whose
|
||||
only available photo was already used to train the classifier, so their
|
||||
scores aren't a clean held-out result.
|
||||
|
||||
Every run appends to `sources/product_accuracy_history.jsonl` and **auto-diffs
|
||||
against the previous run**: the printed summary shows a Δ column per field per
|
||||
@@ -185,10 +191,9 @@ Tracked ones (see `plans/next-enhancements.md`):
|
||||
- **Dataset thinness**: 2–16 photos/class caps both classifiers; every new real
|
||||
photo (especially non-studio, in-warehouse shots) matters. The harness above
|
||||
already reports gallery (training) vs. held-out (validation) accuracy
|
||||
separately — but as of this writing `sources/product-test-images/` is empty,
|
||||
so the Validation Set is still 0 images and every published number so far is
|
||||
a training/memorization score. Dropping real photos there is the next step,
|
||||
not yet done.
|
||||
separately, and as of 2026-07-14 the Validation Set has 79 labeled images
|
||||
(74 genuinely held out, 5 flagged trained-on — see above) — the first real
|
||||
(non-zero) Validation Set numbers.
|
||||
|
||||
Additional recommendations (not yet tasks — promote via `e`/`n` when wanted):
|
||||
1. ~~Use `extracted_sku` in match ranking.~~ **Done** — `product-scan.ts`'s
|
||||
|
||||
@@ -4,6 +4,7 @@ const nextConfig: NextConfig = {
|
||||
// Allow dev requests from any host — needed for tunnel access (ngrok, cloudflare, etc.)
|
||||
// and direct LAN/WiFi IP access from Android devices.
|
||||
allowedDevOrigins: [
|
||||
"127.0.0.1",
|
||||
"*.trycloudflare.com",
|
||||
"*.ngrok.io",
|
||||
"*.ngrok-free.app",
|
||||
|
||||
@@ -8,7 +8,7 @@ export const dynamic = "force-dynamic";
|
||||
export async function GET(req: NextRequest) {
|
||||
try {
|
||||
const filename = req.nextUrl.searchParams.get("filename");
|
||||
const dirPath = path.join(process.cwd(), "..", "sources", "product-test-images");
|
||||
const dirPath = path.join(process.cwd(), "..", "sources", "product-test-images-fixed");
|
||||
|
||||
// File serving mode
|
||||
if (filename) {
|
||||
|
||||
@@ -0,0 +1,86 @@
|
||||
import { NextRequest, NextResponse } from "next/server";
|
||||
import fs from "fs";
|
||||
import path from "path";
|
||||
import { errorResponse } from "@/utils/api-error";
|
||||
|
||||
// Serves the most recent accuracy-check-scan.mts detail dump
|
||||
// (sources/product_scan_detail_*.json) so the manual-label-scan page can show
|
||||
// what the AI actually predicted for a given Validation Set image by default,
|
||||
// without re-running the pipeline live for every image browsed. This is the
|
||||
// same predicted value the accuracy harness scores against ground truth -
|
||||
// not a fresh scan, so it reflects the last batch test run.
|
||||
const SOURCES_DIR = path.join(process.cwd(), "..", "sources");
|
||||
|
||||
interface DetailCheck {
|
||||
field: string;
|
||||
match: boolean;
|
||||
expected: string;
|
||||
predicted: string;
|
||||
}
|
||||
|
||||
interface DetailValidationItem {
|
||||
filename: string;
|
||||
method?: string;
|
||||
confidence?: number;
|
||||
checks: DetailCheck[];
|
||||
}
|
||||
|
||||
interface DetailDump {
|
||||
timestamp: string;
|
||||
validation: DetailValidationItem[];
|
||||
}
|
||||
|
||||
function findLatestDump(): { path: string; data: DetailDump } | null {
|
||||
if (!fs.existsSync(SOURCES_DIR)) return null;
|
||||
const candidates = fs
|
||||
.readdirSync(SOURCES_DIR)
|
||||
.filter((f) => /^product_scan_detail_.*\.json$/.test(f))
|
||||
.map((f) => {
|
||||
const p = path.join(SOURCES_DIR, f);
|
||||
return { path: p, mtime: fs.statSync(p).mtimeMs };
|
||||
})
|
||||
.sort((a, b) => b.mtime - a.mtime);
|
||||
|
||||
if (candidates.length === 0) return null;
|
||||
const latest = candidates[0];
|
||||
const data = JSON.parse(fs.readFileSync(latest.path, "utf8"));
|
||||
return { path: latest.path, data };
|
||||
}
|
||||
|
||||
export async function GET(req: NextRequest) {
|
||||
try {
|
||||
const { searchParams } = new URL(req.url);
|
||||
const filename = searchParams.get("filename");
|
||||
|
||||
const latest = findLatestDump();
|
||||
if (!latest) {
|
||||
return NextResponse.json({ available: false });
|
||||
}
|
||||
|
||||
if (!filename) {
|
||||
return NextResponse.json({ available: true, timestamp: latest.data.timestamp });
|
||||
}
|
||||
|
||||
const item = latest.data.validation.find((v) => v.filename === filename);
|
||||
if (!item) {
|
||||
return NextResponse.json({ available: true, timestamp: latest.data.timestamp, found: false });
|
||||
}
|
||||
|
||||
const byField = Object.fromEntries(item.checks.map((c) => [c.field, c]));
|
||||
|
||||
return NextResponse.json({
|
||||
available: true,
|
||||
found: true,
|
||||
timestamp: latest.data.timestamp,
|
||||
method: item.method,
|
||||
confidence: item.confidence,
|
||||
no_sku: byField.no_sku?.predicted,
|
||||
nama_item: byField.nama_item?.predicted,
|
||||
expiry_date: byField.expiry_date?.predicted
|
||||
});
|
||||
} catch (err: unknown) {
|
||||
console.error("Error in product-scan-results API:", err);
|
||||
const message = err instanceof Error ? err.message : "Internal server error";
|
||||
return errorResponse(500, message);
|
||||
}
|
||||
}
|
||||
@@ -14,41 +14,10 @@ export async function POST(req: NextRequest) {
|
||||
|
||||
const result = await classifyAndMatchProduct(image_base64);
|
||||
|
||||
// Layout-parsing visualization (same pipeline as DO-PFM Visual Grid) - only
|
||||
// used by this desktop test page, not part of the shared classify+match logic.
|
||||
let layoutParsingResult: { layoutParsingResults?: Array<{ outputImages?: Record<string, string> }> } | null = null;
|
||||
const rawB64 = image_base64.includes(",") ? image_base64.split(",")[1] : image_base64;
|
||||
const pipelineUrl = process.env.PIPELINE_URL || "http://localhost:7871/layout-parsing";
|
||||
|
||||
try {
|
||||
const layoutResponse = await fetch(pipelineUrl, {
|
||||
method: "POST",
|
||||
headers: { "Content-Type": "application/json" },
|
||||
body: JSON.stringify({
|
||||
file: rawB64,
|
||||
matchHistoryJob: false,
|
||||
useLayoutDetection: true,
|
||||
fileType: 1,
|
||||
useDocUnwarping: false,
|
||||
useDocOrientationClassify: false
|
||||
})
|
||||
});
|
||||
|
||||
if (layoutResponse.ok) {
|
||||
const layoutData = await layoutResponse.json();
|
||||
layoutParsingResult = layoutData.result ?? layoutData;
|
||||
} else {
|
||||
console.warn("Layout parsing for visualization failed:", await layoutResponse.text());
|
||||
}
|
||||
} catch (layoutErr) {
|
||||
console.warn("Layout parsing for visualization unavailable:", layoutErr);
|
||||
}
|
||||
|
||||
return NextResponse.json({
|
||||
classification: result.classification,
|
||||
ocr: result.ocr,
|
||||
possibleMatches: result.possibleMatches,
|
||||
layoutParsingResult
|
||||
possibleMatches: result.possibleMatches
|
||||
});
|
||||
|
||||
} catch (error: unknown) {
|
||||
|
||||
@@ -18,7 +18,8 @@ export default function ManualLabelScanPage() {
|
||||
notes: ""
|
||||
});
|
||||
const [aiPredicted, setAiPredicted] = useState<AiPredictedData | null>(null);
|
||||
|
||||
const [aiSource, setAiSource] = useState<{ type: "batch" | "live"; timestamp: string; method?: string; confidence?: number } | null>(null);
|
||||
|
||||
const [skuList, setSkuList] = useState<Array<{ no_sku: string; nama_item: string }>>([]);
|
||||
const [isScanning, setIsScanning] = useState(false);
|
||||
const [savingGT, setSavingGT] = useState(false);
|
||||
@@ -39,20 +40,11 @@ export default function ManualLabelScanPage() {
|
||||
setSkuList(skuData.skus || []);
|
||||
}
|
||||
|
||||
// Fetch Training Images
|
||||
const pfmRes = await fetch("/api/produk-pfm");
|
||||
let trainingFiles: { url: string; filename: string }[] = [];
|
||||
if (pfmRes.ok) {
|
||||
const pfmData = await pfmRes.json();
|
||||
trainingFiles = (pfmData.products || []).flatMap((p: any) =>
|
||||
p.images.map((url: string) => ({
|
||||
url,
|
||||
filename: url.replace(/^\/produk-pfm\/foto-kemasan-v2\//, "")
|
||||
}))
|
||||
);
|
||||
}
|
||||
|
||||
// Fetch Test Images
|
||||
// Fetch Test Images — the frozen 79-image Validation Set
|
||||
// (product-test-images-fixed/), the only set the accuracy harness
|
||||
// scores. Gallery/training photos (foto-kemasan-v2/) are not shown
|
||||
// here: they don't need per-photo ground truth, only correct
|
||||
// SKU-folder placement for classifier training.
|
||||
const testRes = await fetch("/api/product-images");
|
||||
let testFiles: { url: string; filename: string }[] = [];
|
||||
if (testRes.ok) {
|
||||
@@ -63,9 +55,8 @@ export default function ManualLabelScanPage() {
|
||||
}));
|
||||
}
|
||||
|
||||
const combined = [...testFiles, ...trainingFiles];
|
||||
setFiles(combined);
|
||||
if (combined.length > 0) setCurrentIndex(0);
|
||||
setFiles(testFiles);
|
||||
if (testFiles.length > 0) setCurrentIndex(0);
|
||||
|
||||
} catch (err) {
|
||||
console.error("Error initializing page", err);
|
||||
@@ -91,11 +82,33 @@ export default function ManualLabelScanPage() {
|
||||
expiry_date: data.expiry_date || "",
|
||||
notes: data.notes || ""
|
||||
});
|
||||
setAiPredicted(null); // Reset AI predictions on new file load
|
||||
}
|
||||
} catch (err) {
|
||||
console.error("Error fetching label", err);
|
||||
}
|
||||
|
||||
// Default-load the AI prediction from the last batch accuracy run
|
||||
// (not a live re-scan) so failures are visible immediately while
|
||||
// browsing - "Scan with AI" below can still be used to get a fresh
|
||||
// live result for this exact image.
|
||||
setAiPredicted(null);
|
||||
setAiSource(null);
|
||||
try {
|
||||
const aiRes = await fetch(`/api/product-scan-results?filename=${encodeURIComponent(file.filename)}`);
|
||||
if (aiRes.ok) {
|
||||
const aiData = await aiRes.json();
|
||||
if (aiData.found) {
|
||||
setAiPredicted({
|
||||
no_sku: aiData.no_sku,
|
||||
nama_item: aiData.nama_item,
|
||||
expiry_date: aiData.expiry_date
|
||||
});
|
||||
setAiSource({ type: "batch", timestamp: aiData.timestamp, method: aiData.method, confidence: aiData.confidence });
|
||||
}
|
||||
}
|
||||
} catch (err) {
|
||||
console.error("Error fetching batch AI result", err);
|
||||
}
|
||||
};
|
||||
loadLabel();
|
||||
}, [currentIndex, files]);
|
||||
@@ -138,13 +151,27 @@ export default function ManualLabelScanPage() {
|
||||
|
||||
if (!scanRes.ok) throw new Error("Pipeline API error");
|
||||
const scanData = await scanRes.json();
|
||||
|
||||
|
||||
// Compare against the sku_master-resolved best match (what the app
|
||||
// actually shows/saves as nama_item, and what the accuracy harness
|
||||
// scores), not classification.top1_name - that's the classifier's raw
|
||||
// internal class label (e.g. the foto-kemasan-v2 folder name), which
|
||||
// structurally never matches a sku_master-style ground truth string
|
||||
// even when the classification itself is correct.
|
||||
const bestMatch = (scanData.possibleMatches || []).find((m: { isBestMatch?: boolean }) => m.isBestMatch);
|
||||
|
||||
setAiPredicted({
|
||||
no_sku: scanData.classification?.top1_name ? skuList.find(s => s.nama_item === scanData.classification.top1_name)?.no_sku : undefined,
|
||||
nama_item: scanData.classification?.top1_name,
|
||||
no_sku: bestMatch?.no_sku,
|
||||
nama_item: bestMatch?.nama_item,
|
||||
expiry_date: scanData.ocr?.extracted_expired_date
|
||||
});
|
||||
|
||||
setAiSource({
|
||||
type: "live",
|
||||
timestamp: new Date().toISOString(),
|
||||
method: scanData.classification?.method,
|
||||
confidence: scanData.classification?.top1_confidence
|
||||
});
|
||||
|
||||
showToast("AI Scan complete!");
|
||||
} catch (err) {
|
||||
showToast(getErrorMessage(err, undefined, "AI Scan failed"), true);
|
||||
@@ -206,6 +233,7 @@ export default function ManualLabelScanPage() {
|
||||
<Editor
|
||||
formData={formData}
|
||||
aiPredicted={aiPredicted}
|
||||
aiSource={aiSource}
|
||||
skuList={skuList}
|
||||
isScanning={isScanning}
|
||||
onScanWithAi={handleScanWithAi}
|
||||
|
||||
@@ -1,7 +1,6 @@
|
||||
"use client";
|
||||
|
||||
import React, { useState, useEffect } from "react";
|
||||
import { extractLayoutVisUrlFromResult } from "@/utils/layoutVisualization";
|
||||
import { getErrorMessage } from "@/utils/client-error";
|
||||
|
||||
export const dynamic = "force-dynamic";
|
||||
@@ -34,16 +33,9 @@ interface ScanResponse {
|
||||
expired_line_index?: number;
|
||||
expired_date_crop_base64?: string;
|
||||
expired_source_line?: string;
|
||||
vis_image_base64?: string;
|
||||
spotting_image_base64?: string;
|
||||
error?: string;
|
||||
};
|
||||
possibleMatches: MatchResult[];
|
||||
layoutParsingResult?: {
|
||||
layoutParsingResults?: Array<{
|
||||
outputImages?: Record<string, string>;
|
||||
}>;
|
||||
};
|
||||
}
|
||||
|
||||
function levenshteinDistance(s1: string, s2: string): number {
|
||||
@@ -115,7 +107,7 @@ export default function ScanPfmPage() {
|
||||
const [savingGT, setSavingGT] = useState<boolean>(false);
|
||||
const [error, setError] = useState<string>("");
|
||||
const [scanResult, setScanResult] = useState<ScanResponse | null>(null);
|
||||
const [activeTab, setActiveTab] = useState<"summary" | "visual" | "spotting" | "json">("summary");
|
||||
const [activeTab, setActiveTab] = useState<"summary" | "json">("summary");
|
||||
|
||||
const [isEditingDate, setIsEditingDate] = useState<boolean>(false);
|
||||
const [editedDate, setEditedDate] = useState<string>("");
|
||||
@@ -385,8 +377,6 @@ export default function ScanPfmPage() {
|
||||
}
|
||||
};
|
||||
|
||||
const visUrl = scanResult ? extractLayoutVisUrlFromResult(scanResult.layoutParsingResult) : null;
|
||||
|
||||
return (
|
||||
<div className="min-h-screen bg-slate-950 text-slate-100 flex flex-col font-sans">
|
||||
|
||||
@@ -637,14 +627,12 @@ export default function ScanPfmPage() {
|
||||
<div className="border-b border-slate-800 flex rounded-xl overflow-hidden bg-slate-950/50">
|
||||
{[
|
||||
{ id: "summary", name: "Scan Summary" },
|
||||
{ id: "visual", name: "Visual Grid" },
|
||||
{ id: "spotting", name: "Spotting Grid" },
|
||||
{ id: "json", name: "Raw Response" }
|
||||
].map((t) => (
|
||||
<button
|
||||
key={t.id}
|
||||
id={`tab-${t.id}`}
|
||||
onClick={() => setActiveTab(t.id as "summary" | "visual" | "spotting" | "json")}
|
||||
onClick={() => setActiveTab(t.id as "summary" | "json")}
|
||||
className={`flex-1 py-3 text-xs font-bold transition-all border-b-2 cursor-pointer select-none ${
|
||||
activeTab === t.id
|
||||
? "border-teal-500 text-teal-400 bg-slate-900/40"
|
||||
@@ -1059,118 +1047,6 @@ export default function ScanPfmPage() {
|
||||
</div>
|
||||
)}
|
||||
|
||||
{/* Visual Grid Tab */}
|
||||
{activeTab === "visual" && (
|
||||
<div className="space-y-4">
|
||||
<h3 className="text-xs font-bold text-teal-400 uppercase tracking-wider">
|
||||
Layout Visualization Grid
|
||||
</h3>
|
||||
{visUrl ? (
|
||||
<div className="bg-slate-950 rounded-xl border border-slate-800 overflow-hidden shadow-2xl p-2 flex justify-center">
|
||||
{/* eslint-disable-next-line @next/next/no-img-element */}
|
||||
<img
|
||||
src={visUrl}
|
||||
alt="Layout Visualization Grid"
|
||||
className="max-w-full h-auto object-contain rounded"
|
||||
/>
|
||||
</div>
|
||||
) : (
|
||||
<p className="text-xs text-slate-500 italic">
|
||||
No layout visualization image returned. The layout-parsing pipeline may be unavailable.
|
||||
</p>
|
||||
)}
|
||||
</div>
|
||||
)}
|
||||
|
||||
{/* Spotting Grid Tab */}
|
||||
{activeTab === "spotting" && (
|
||||
<div className="space-y-6">
|
||||
<div className="space-y-4">
|
||||
<h3 className="text-xs font-bold text-teal-400 uppercase tracking-wider">
|
||||
Spotting Visualization
|
||||
</h3>
|
||||
{scanResult.ocr?.spotting_image_base64 ? (
|
||||
<div className="bg-slate-950 rounded-xl border border-slate-800 overflow-hidden shadow-2xl p-2 flex justify-center">
|
||||
{/* eslint-disable-next-line @next/next/no-img-element */}
|
||||
<img
|
||||
src={scanResult.ocr.spotting_image_base64}
|
||||
alt="Spotting Visualization BBoxes"
|
||||
className="max-w-full h-auto object-contain rounded"
|
||||
/>
|
||||
</div>
|
||||
) : (
|
||||
<p className="text-xs text-slate-500 italic">No spotting visualization image returned by the parser.</p>
|
||||
)}
|
||||
</div>
|
||||
|
||||
<div className="space-y-3 pt-2 border-t border-slate-800">
|
||||
<h3 className="text-xs font-bold text-amber-400 uppercase tracking-wider">
|
||||
Detected Expired Date (Best Before / BB)
|
||||
</h3>
|
||||
<div className="grid grid-cols-1 sm:grid-cols-2 gap-4">
|
||||
<div className="bg-slate-900/40 border border-slate-800 rounded-xl p-4 flex flex-col gap-2 min-h-[100px]">
|
||||
<span className="text-[9px] font-bold text-slate-500 uppercase tracking-wider">Extracted Date</span>
|
||||
{isEditingDate ? (
|
||||
<div className="flex items-center gap-2 mt-1">
|
||||
<input
|
||||
type="text"
|
||||
value={editedDate}
|
||||
onChange={(e) => setEditedDate(e.target.value)}
|
||||
className="bg-slate-950 border border-slate-700 text-slate-100 rounded-lg px-2 py-1 text-sm font-semibold focus:outline-none focus:ring-1 focus:ring-amber-500 w-full"
|
||||
placeholder="DD/MM/YYYY"
|
||||
autoFocus
|
||||
onKeyDown={(e) => {
|
||||
if (e.key === "Enter") handleSaveDate(editedDate);
|
||||
else if (e.key === "Escape") setIsEditingDate(false);
|
||||
}}
|
||||
/>
|
||||
<button onClick={() => handleSaveDate(editedDate)} className="bg-emerald-600 hover:bg-emerald-500 text-white rounded-lg p-1.5 text-xs font-bold transition-colors">✓</button>
|
||||
<button onClick={() => setIsEditingDate(false)} className="bg-slate-700 hover:bg-slate-600 text-slate-300 rounded-lg p-1.5 text-xs font-bold transition-colors">✗</button>
|
||||
</div>
|
||||
) : (
|
||||
<div className="flex items-center justify-between gap-2">
|
||||
<span className={`text-lg font-bold flex items-center gap-2 ${scanResult.ocr?.extracted_expired_date ? "text-amber-400" : "text-slate-600"}`}>
|
||||
<span className={`h-2 w-2 rounded-full flex-shrink-0 ${scanResult.ocr?.extracted_expired_date ? "bg-amber-400" : "bg-slate-700"}`} />
|
||||
{scanResult.ocr?.extracted_expired_date || "Not detected"}
|
||||
</span>
|
||||
<button
|
||||
onClick={() => { setEditedDate(scanResult.ocr?.extracted_expired_date || ""); setIsEditingDate(true); }}
|
||||
className="text-[10px] bg-slate-800 hover:bg-slate-700 text-slate-400 px-2.5 py-1 rounded-md border border-slate-700 transition-all font-semibold flex items-center gap-1"
|
||||
>
|
||||
✏️ Edit
|
||||
</button>
|
||||
</div>
|
||||
)}
|
||||
{scanResult.ocr?.expired_source_line && (
|
||||
<p className="text-[10px] text-slate-500 leading-snug">
|
||||
<span className="font-semibold text-slate-400">OCR: </span>
|
||||
{scanResult.ocr.expired_source_line}
|
||||
</p>
|
||||
)}
|
||||
</div>
|
||||
|
||||
<div className="bg-slate-900/40 border border-slate-800 rounded-xl p-4 flex flex-col gap-2 min-h-[100px]">
|
||||
<span className="text-[9px] font-bold text-slate-500 uppercase tracking-wider">Label Crop</span>
|
||||
<div className="flex-1 bg-slate-950 border border-slate-800 rounded-lg p-2 flex items-center justify-center min-h-[72px] overflow-hidden">
|
||||
{scanResult.ocr?.expired_date_crop_base64 ? (
|
||||
// eslint-disable-next-line @next/next/no-img-element
|
||||
<img
|
||||
src={scanResult.ocr.expired_date_crop_base64}
|
||||
alt="Expired date crop"
|
||||
className="max-h-24 max-w-full object-contain brightness-95 contrast-105"
|
||||
/>
|
||||
) : (
|
||||
<span className="text-[10px] text-slate-600 italic text-center px-2">
|
||||
{scanResult.ocr?.extracted_expired_date ? "Crop unavailable" : "No expiry region to crop"}
|
||||
</span>
|
||||
)}
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
)}
|
||||
|
||||
{/* JSON Tab */}
|
||||
{activeTab === "json" && (
|
||||
<div className="space-y-4">
|
||||
|
||||
@@ -14,9 +14,17 @@ export interface AiPredictedData {
|
||||
expiry_date?: string;
|
||||
}
|
||||
|
||||
export interface AiSourceInfo {
|
||||
type: "batch" | "live";
|
||||
timestamp: string;
|
||||
method?: string;
|
||||
confidence?: number;
|
||||
}
|
||||
|
||||
interface EditorProps {
|
||||
formData: ScanLabelFormData;
|
||||
aiPredicted: AiPredictedData | null;
|
||||
aiSource: AiSourceInfo | null;
|
||||
skuList: Array<{ no_sku: string; nama_item: string }>;
|
||||
isScanning: boolean;
|
||||
onScanWithAi: () => void;
|
||||
@@ -28,6 +36,7 @@ interface EditorProps {
|
||||
export function Editor({
|
||||
formData,
|
||||
aiPredicted,
|
||||
aiSource,
|
||||
skuList,
|
||||
isScanning,
|
||||
onScanWithAi,
|
||||
@@ -69,8 +78,21 @@ export function Editor({
|
||||
disabled={isScanning || !formData.filename}
|
||||
className="w-full bg-teal-600/20 text-teal-400 hover:bg-teal-600/30 disabled:opacity-50 border border-teal-500/30 rounded-lg py-2 text-xs font-semibold transition flex items-center justify-center gap-2"
|
||||
>
|
||||
{isScanning ? "Scanning with Pipeline..." : "Scan with AI 🤖"}
|
||||
{isScanning ? "Scanning with Pipeline..." : "Scan with AI 🤖 (re-run live)"}
|
||||
</button>
|
||||
{aiSource ? (
|
||||
<p className="text-[10px] text-slate-500 leading-snug">
|
||||
{aiSource.type === "batch" ? (
|
||||
<>Showing result from last batch test ({new Date(aiSource.timestamp).toLocaleString()})</>
|
||||
) : (
|
||||
<>Live scan result ({new Date(aiSource.timestamp).toLocaleTimeString()})</>
|
||||
)}
|
||||
{aiSource.method && <> · {aiSource.method}</>}
|
||||
{typeof aiSource.confidence === "number" && <> · conf {aiSource.confidence.toFixed(3)}</>}
|
||||
</p>
|
||||
) : (
|
||||
<p className="text-[10px] text-slate-600 italic">No AI result yet for this image — click "Scan with AI" or run the accuracy batch test.</p>
|
||||
)}
|
||||
</div>
|
||||
|
||||
{/* Form Fields */}
|
||||
|
||||
@@ -1,26 +0,0 @@
|
||||
export function normalizeImageSrc(src: string): string {
|
||||
if (!src) return "";
|
||||
if (src.startsWith("http://") || src.startsWith("https://") || src.startsWith("data:")) {
|
||||
return src;
|
||||
}
|
||||
return `data:image/png;base64,${src}`;
|
||||
}
|
||||
|
||||
export interface LayoutPageResult {
|
||||
outputImages?: Record<string, string>;
|
||||
}
|
||||
|
||||
/** Same visualization URL selection as DO-PFM Visual Grid (second image if present, else first). */
|
||||
export function extractLayoutVisUrl(page0: LayoutPageResult | null | undefined): string {
|
||||
const outImgs = page0?.outputImages || {};
|
||||
const sortedUrls = Object.values(outImgs).filter(Boolean) as string[];
|
||||
const visUrl = sortedUrls.length >= 2 ? sortedUrls[1] : sortedUrls[0] || "";
|
||||
return normalizeImageSrc(visUrl);
|
||||
}
|
||||
|
||||
export function extractLayoutVisUrlFromResult(
|
||||
layoutParsingResult: { layoutParsingResults?: LayoutPageResult[] } | null | undefined
|
||||
): string {
|
||||
const page0 = layoutParsingResult?.layoutParsingResults?.[0];
|
||||
return extractLayoutVisUrl(page0);
|
||||
}
|
||||
@@ -1,10 +1,12 @@
|
||||
import { query } from "../db";
|
||||
|
||||
// Bounds the classifier call so a wedged GPU container fails fast instead of
|
||||
// hanging indefinitely - matches the bound `api/parse/route.ts` used to apply
|
||||
// to its own separate inline classify call before it started sharing this
|
||||
// function (see docs/api-contract-map.md G3).
|
||||
const PIPELINE_TIMEOUT_MS = 90_000;
|
||||
// hanging indefinitely. Raised from 90s (2026-07-14): hard images now
|
||||
// legitimately take up to ~3 min - a 4-orientation OCR search plus a VL
|
||||
// pipeline fallback when no expiry date is found (see
|
||||
// config/classify_ocr_server.py) - and the old bound was killing exactly
|
||||
// the images those fallbacks exist to save.
|
||||
const PIPELINE_TIMEOUT_MS = 240_000;
|
||||
|
||||
// Thrown when the Python classifier service itself returns a non-2xx response,
|
||||
// so callers can forward its actual status instead of collapsing everything to 500.
|
||||
@@ -60,6 +62,115 @@ function getStringSimilarity(s1: string, s2: string): number {
|
||||
return (maxLength - distance) / maxLength;
|
||||
}
|
||||
|
||||
// --- OCR-evidence re-ranking of the classifier's top-K candidates ---
|
||||
//
|
||||
// DINOv2's misses are near-twin confusions (same brand line, different
|
||||
// flavor/size) - exactly the cases where the printed variant words differ,
|
||||
// and PaddleOCR usually reads some of them. Within a narrow similarity band
|
||||
// of the top-1 candidate, prefer the one whose distinctive name tokens
|
||||
// actually appear in the OCR'd text. Coverage-normalized so generic
|
||||
// packaging words (e.g. "French Fries", "Ayam") that happen to be unique to
|
||||
// one candidate's *name* can't hijack the ranking. Parameters tuned offline
|
||||
// against the 79-image validation set (scripts/experiment-rerank.mjs,
|
||||
// 2026-07-14: fixes 8 of 18 top-1 misses, breaks 0 of 61 correct).
|
||||
const RERANK_TOP_K = 12;
|
||||
const RERANK_SIM_BAND = 0.12;
|
||||
const RERANK_COVERAGE_MARGIN = 0.25;
|
||||
|
||||
function classNameSku(className: string): string {
|
||||
// foto-kemasan-v2 class names are "<SKU> <NAME...>"
|
||||
return (className || "").trim().split(/\s+/)[0] || "";
|
||||
}
|
||||
|
||||
function tokenizeName(name: string): string[] {
|
||||
return name.toUpperCase().split(/[^A-Z0-9]+/).filter(t => t.length >= 2);
|
||||
}
|
||||
|
||||
function withinEditDistance1(a: string, b: string): boolean {
|
||||
if (a === b) return true;
|
||||
const la = a.length, lb = b.length;
|
||||
if (Math.abs(la - lb) > 1) return false;
|
||||
if (la === lb) {
|
||||
let diff = 0;
|
||||
for (let i = 0; i < la; i++) if (a[i] !== b[i]) diff++;
|
||||
return diff <= 1;
|
||||
}
|
||||
const [s, l] = la < lb ? [a, b] : [b, a];
|
||||
let i = 0, j = 0, skipped = false;
|
||||
while (i < s.length && j < l.length) {
|
||||
if (s[i] === l[j]) { i++; j++; }
|
||||
else if (!skipped) { skipped = true; j++; }
|
||||
else return false;
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
interface OcrTextIndex { squashed: string; tokens: Set<string>; }
|
||||
|
||||
function buildOcrTextIndex(textLines: string[]): OcrTextIndex {
|
||||
const joined = textLines.join(" ").toUpperCase();
|
||||
return {
|
||||
squashed: joined.replace(/[^A-Z0-9]/g, ""),
|
||||
tokens: new Set(tokenizeName(joined))
|
||||
};
|
||||
}
|
||||
|
||||
function tokenFoundInOcr(token: string, ocr: OcrTextIndex): boolean {
|
||||
if (token.length >= 4 && ocr.squashed.includes(token)) return true;
|
||||
if (ocr.tokens.has(token)) return true;
|
||||
if (token.length >= 5) {
|
||||
for (const t of ocr.tokens) {
|
||||
if (Math.abs(t.length - token.length) <= 1 && withinEditDistance1(token, t)) return true;
|
||||
}
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
// Returns the class name of the best candidate after OCR-evidence
|
||||
// re-ranking (the classifier's top-1 unless a close band-mate has clearly
|
||||
// stronger printed-text evidence).
|
||||
function rerankClassCandidates(
|
||||
allProbabilities: Array<{ name: string; confidence: number }>,
|
||||
textLines: string[]
|
||||
): string {
|
||||
if (!allProbabilities.length) return "";
|
||||
const top1Sim = allProbabilities[0].confidence;
|
||||
const band = allProbabilities
|
||||
.slice(0, RERANK_TOP_K)
|
||||
.filter(p => p.confidence >= top1Sim - RERANK_SIM_BAND);
|
||||
if (band.length <= 1 || !textLines.length) return allProbabilities[0].name;
|
||||
|
||||
const ocrIdx = buildOcrTextIndex(textLines);
|
||||
const cands = band.map(p => {
|
||||
const sku = classNameSku(p.name);
|
||||
return { name: p.name, tokens: new Set(tokenizeName(p.name.replace(sku, ""))), coverage: 0 };
|
||||
});
|
||||
const tokenCounts = new Map<string, number>();
|
||||
for (const c of cands) {
|
||||
for (const tok of c.tokens) tokenCounts.set(tok, (tokenCounts.get(tok) || 0) + 1);
|
||||
}
|
||||
for (const c of cands) {
|
||||
let matched = 0, total = 0;
|
||||
for (const tok of c.tokens) {
|
||||
const nWith = tokenCounts.get(tok) || 1;
|
||||
if (nWith >= cands.length) continue; // shared by all band-mates -> no signal
|
||||
const w = 1 / nWith;
|
||||
total += w;
|
||||
if (tokenFoundInOcr(tok, ocrIdx)) matched += w;
|
||||
}
|
||||
c.coverage = total > 0 ? matched / total : 0;
|
||||
}
|
||||
|
||||
let chosen = cands[0];
|
||||
for (const c of cands.slice(1)) {
|
||||
if (c.coverage >= chosen.coverage + RERANK_COVERAGE_MARGIN) chosen = c;
|
||||
}
|
||||
if (chosen !== cands[0]) {
|
||||
console.log(`[Rerank] OCR evidence overrode classifier top-1 "${cands[0].name}" -> "${chosen.name}" (coverage ${cands[0].coverage.toFixed(2)} vs ${chosen.coverage.toFixed(2)})`);
|
||||
}
|
||||
return chosen.name;
|
||||
}
|
||||
|
||||
// Shared by the classic /api/scan-pfm dev route and the authenticated
|
||||
// /api/v1/scan-product route: calls the Python classifier, then matches the
|
||||
// result against sku_master, returning the top-5 candidates.
|
||||
@@ -86,17 +197,28 @@ export async function classifyAndMatchProduct(imageBase64: string): Promise<Prod
|
||||
nama_item: row.nama_item
|
||||
}));
|
||||
|
||||
const top1Name = data.classification?.top1_name || "";
|
||||
const extractedSku = data.ocr?.extracted_sku || "";
|
||||
|
||||
// Re-rank the classifier's close candidates using OCR'd package text, then
|
||||
// map the winner straight to its sku_master row by the SKU prefix embedded
|
||||
// in the class name. The old approach (Levenshtein between top-1 class name
|
||||
// and every master nama_item) lost classifier-correct results whenever a
|
||||
// *different* SKU's master name happened to be textually closer.
|
||||
const rerankedName = rerankClassCandidates(
|
||||
data.classification?.all_probabilities || [],
|
||||
data.ocr?.text_lines || []
|
||||
) || data.classification?.top1_name || "";
|
||||
const rerankedSku = classNameSku(rerankedName);
|
||||
|
||||
const matchedList: SkuMatch[] = skuMasterList.map(sku => {
|
||||
const yoloSim = top1Name ? getStringSimilarity(sku.nama_item, top1Name) : 0;
|
||||
const yoloSim = rerankedName ? getStringSimilarity(sku.nama_item, rerankedName) : 0;
|
||||
|
||||
const cleanMasterSku = sku.no_sku.trim();
|
||||
const cleanExtractedSku = extractedSku.trim();
|
||||
const isSkuMatch = cleanExtractedSku && cleanMasterSku === cleanExtractedSku;
|
||||
const isClassifierPick = rerankedSku && cleanMasterSku === rerankedSku;
|
||||
|
||||
const score = isSkuMatch ? 1.0 : yoloSim;
|
||||
const score = isSkuMatch ? 1.0 : isClassifierPick ? 0.995 : yoloSim;
|
||||
|
||||
return {
|
||||
no_sku: sku.no_sku,
|
||||
|
||||
@@ -113,11 +113,11 @@ in either project (only SKU, product name, expiry date are extracted) — if
|
||||
requested later, follow the same OCR-regex-cascade pattern already used for
|
||||
expiry-date extraction.*
|
||||
|
||||
- **2.5** [IN PROGRESS 2026-07-14 — resumed, training run 2] **Retrain
|
||||
classifier on the now-81-class dataset.** (Note: a first resume attempt
|
||||
- **2.5** [DONE 2026-07-14] **Retrain classifier on the now-81-class
|
||||
dataset.** (Note: a first resume attempt
|
||||
failed instantly with a Docker daemon connection error — Docker Desktop had
|
||||
stopped between sessions — before any training happened; restarted Docker
|
||||
Desktop and relaunched. This is the actual second training attempt,
|
||||
Desktop and relaunched. The successful run was the second attempt,
|
||||
confirmed running via `docker ps`.) `foto-kemasan-v2/` grew from the 16
|
||||
classes/118 photos the deployed model
|
||||
(`produk-pfm-classifier-26n-100e-2026-07-08.pt`) was trained on to **81
|
||||
@@ -138,36 +138,30 @@ expiry-date extraction.*
|
||||
`docker compose restart pipeline-api` → verify via `docker logs` for
|
||||
"DINOv2 index loaded with N reference images" and "Using classifier
|
||||
weights: <new dated file>".
|
||||
- **Status as of pause (2026-07-14)** — mixed state, read carefully before
|
||||
resuming:
|
||||
- ✅ `pipeline-api` image built (2m54s), bakes in the current 81-class
|
||||
dataset.
|
||||
- ✅ **`dinov2_index.pkl` already rebuilt and persisted to disk** —
|
||||
"Success! Indexed 2493/2493 images" across all 81 classes. This artifact
|
||||
is live on the host now (`models/dinov2_index.pkl`, 4.2MB, dated
|
||||
2026-07-14) and does **not** need to be redone.
|
||||
- ⏸️ **YOLO classifier training was started, then stopped by user request
|
||||
at epoch 43/100 (~23 minutes in)** before it could write a new dated
|
||||
checkpoint. `docker run` used `--rm` and the in-progress epoch
|
||||
checkpoints live only in the container's own `runs/classify/` (not
|
||||
bind-mounted), so **stopping the container discarded that partial
|
||||
progress** — resuming means restarting from epoch 0, not continuing from
|
||||
43. `models/` on the host still has only the original
|
||||
`produk-pfm-classifier-26n-100e-2026-07-08.pt`/`.onnx` (16-class model) —
|
||||
**the live/deployed classifier is unchanged**, still 16 classes.
|
||||
- Observed pace before stopping: ~32s/epoch (43 epochs in 23m1s) → a full
|
||||
100-epoch run should take **~55 minutes** on this host's RTX 2060 (6GB
|
||||
VRAM), not the ~90 min extrapolated from the first few (slower, warmup)
|
||||
epochs. At epoch 42 the in-progress run had already reached 84.3%
|
||||
top-1 / 93.9% top-5 val accuracy across all 81 classes, ahead of the old
|
||||
16-class model's 83.3%/90% — a promising sign for the eventual full run,
|
||||
but not a final result since training didn't finish.
|
||||
- **To resume**: image is already built and the DINOv2 index step can be
|
||||
skipped — just re-run the one-off `train_classifier.py train --imgsz 224`
|
||||
container, then `docker compose restart pipeline-api` and verify via
|
||||
`docker logs`. Update the class count in `docs/scan-product.md`,
|
||||
`CLAUDE.md`, and `docs/feature-list.md` (and flip this task to `[DONE]`)
|
||||
only once that run actually completes with a final dated `.pt`/`.onnx`.
|
||||
- **Final result (2026-07-14)** — both artifacts retrained and live:
|
||||
- ✅ `dinov2_index.pkl` rebuilt and persisted to disk — "Success! Indexed
|
||||
2493/2493 images" across all 81 classes (`models/dinov2_index.pkl`,
|
||||
4.2MB, dated 2026-07-14).
|
||||
- ✅ **YOLO classifier retrained to completion, 100/100 epochs, real
|
||||
elapsed time 54m21s** (a first attempt was intentionally stopped by user
|
||||
request at epoch 43/100 to pause the session; that partial progress was
|
||||
discarded since `docker run --rm`'s in-container `runs/classify/`
|
||||
checkpoints aren't bind-mounted, so the successful run below restarted
|
||||
cleanly from epoch 0 rather than resuming from 43). Final validation:
|
||||
**85.8% top-1 / 94.4% top-5** across all 81 classes — up from the old
|
||||
16-class model's 83.3%/90%, now covering 5x the product classes.
|
||||
Published artifacts: `models/produk-pfm-classifier-26n-100e-2026-07-14.pt`
|
||||
(3.4MB) and matching `.onnx` (6.3MB, ONNX opset 20, output shape
|
||||
confirmed `(1, 81)` — i.e. 81 output classes).
|
||||
- ✅ Verified via `docker compose up -d pipeline-api` (main stack wasn't
|
||||
running this session) + `docker logs paddleocr-pipeline-api`: "DINOv2
|
||||
index loaded with 2493 reference images", "Using classifier weights:
|
||||
/app/pfm-web-app/public/produk-pfm/models/produk-pfm-classifier-26n-100e-2026-07-14.pt",
|
||||
"YOLO model loaded successfully", "Application startup complete" — the
|
||||
live service is now serving the new 81-class model, not a code-review
|
||||
assumption.
|
||||
- Class-count claims updated in `docs/scan-product.md` and `../CLAUDE.md`
|
||||
(16 → 81 classes); this task's `docs/feature-list.md` entry added.
|
||||
|
||||
## 3. Backend — Postgres Data Layer
|
||||
`pfm-web-app/src/db/`
|
||||
|
||||
@@ -4,8 +4,9 @@
|
||||
// backend/sources/product_manual_labels.json, checks 3 fields (no_sku,
|
||||
// nama_item, expiry_date) against ground truth, splits results into a
|
||||
// Training Set (gallery photos under foto-kemasan-v2/ that trained the
|
||||
// classifier itself) vs a Validation Set (flat filenames dropped in
|
||||
// backend/sources/product-test-images/), and appends a summary to
|
||||
// classifier itself) vs a Validation Set (flat filenames, scored from the
|
||||
// frozen backend/sources/product-test-images-fixed/ snapshot so reruns always
|
||||
// grade the exact same images), and appends a summary to
|
||||
// backend/sources/product_accuracy_history.jsonl. Every run auto-diffs
|
||||
// against the last history entry and flags field/image regressions or
|
||||
// improvements, so a tuning change to classify_ocr_server.py shows its
|
||||
@@ -30,7 +31,10 @@ const SOURCES_DIR = path.join(__dirname, "..", "sources");
|
||||
const LABELS_PATH = process.env.ACCURACY_LABELS_PATH || path.join(SOURCES_DIR, "product_manual_labels.json");
|
||||
const HISTORY_PATH = process.env.ACCURACY_HISTORY_PATH || path.join(SOURCES_DIR, "product_accuracy_history.jsonl");
|
||||
|
||||
const FETCH_TIMEOUT_MS = 120_000;
|
||||
// Hard images legitimately take up to ~3 min now (4-orientation OCR search +
|
||||
// VL pipeline fallback for missing expiry dates); must exceed the gateway's
|
||||
// own PIPELINE_TIMEOUT_MS (240s) so slow scans fail there, not here.
|
||||
const FETCH_TIMEOUT_MS = 300_000;
|
||||
const FIELDS = ["no_sku", "nama_item", "expiry_date"] as const;
|
||||
type Field = typeof FIELDS[number];
|
||||
type Split = "training" | "validation";
|
||||
@@ -61,6 +65,7 @@ interface ScanResponse {
|
||||
interface Check {
|
||||
field: Field;
|
||||
match: boolean;
|
||||
predicted: string;
|
||||
}
|
||||
|
||||
interface ResultItem {
|
||||
@@ -70,6 +75,12 @@ interface ResultItem {
|
||||
confidence?: number;
|
||||
}
|
||||
|
||||
// Optional: set ACCURACY_DETAIL_DUMP_PATH to write full per-image,
|
||||
// per-field ground-truth-vs-predicted detail (plus failures) as JSON —
|
||||
// used to triage which images to pull into an "undetected" folder for
|
||||
// visual inspection instead of just the aggregate percentages.
|
||||
const DETAIL_DUMP_PATH = process.env.ACCURACY_DETAIL_DUMP_PATH;
|
||||
|
||||
interface ClassificationStats {
|
||||
methodCounts: Record<string, number>;
|
||||
avgConfidence: number;
|
||||
@@ -104,7 +115,12 @@ function getImagePath(filename: string): string {
|
||||
if (filename.includes("/")) {
|
||||
return path.join(APP_ROOT, "public", "produk-pfm", "foto-kemasan-v2", filename);
|
||||
}
|
||||
return path.join(SOURCES_DIR, "product-test-images", filename);
|
||||
// Validation Set images are scored from the frozen, sequentially-renamed
|
||||
// copy in product-test-images-fixed/ (built by scripts/freeze-validation-set.mjs)
|
||||
// rather than the live-intake product-test-images/ folder, so a rerun always
|
||||
// scores the exact same image set regardless of what's since been dropped
|
||||
// into the live folder for future curation.
|
||||
return path.join(SOURCES_DIR, "product-test-images-fixed", filename);
|
||||
}
|
||||
|
||||
async function checkServerReachable(baseUrl: string) {
|
||||
@@ -338,6 +354,7 @@ async function main() {
|
||||
validation: [] as ResultItem[],
|
||||
failed: [] as string[]
|
||||
};
|
||||
const failedDetail: Array<{ filename: string; error: string }> = [];
|
||||
const perImagePct: Record<string, number> = {};
|
||||
|
||||
for (const gt of labels) {
|
||||
@@ -353,14 +370,21 @@ async function main() {
|
||||
const b64 = "data:image/jpeg;base64," + fs.readFileSync(imgPath, "base64");
|
||||
const parsed = await fetchScan(args.baseUrl, b64);
|
||||
|
||||
const bestMatchSku = parsed.possibleMatches?.find(m => m.isBestMatch)?.no_sku || "";
|
||||
const predictedItemName = parsed.classification?.top1_name || "";
|
||||
const bestMatch = parsed.possibleMatches?.find(m => m.isBestMatch);
|
||||
const bestMatchSku = bestMatch?.no_sku || "";
|
||||
// Compare against the sku_master-resolved name (what the app actually
|
||||
// shows/saves as nama_item), not classification.top1_name — that's the
|
||||
// classifier's raw internal class label (e.g. "11110059 CEKER BERKUKU
|
||||
// FROZEN PACK 1 KG", literally the foto-kemasan-v2 folder name), which
|
||||
// structurally never matches a sku_master-style ground truth string
|
||||
// even when the classification itself is correct.
|
||||
const predictedItemName = bestMatch?.nama_item || "";
|
||||
const predictedExpiry = parsed.ocr?.extracted_expired_date || "";
|
||||
|
||||
const checks: Check[] = [
|
||||
{ field: "no_sku", match: isMatch(gt.no_sku, bestMatchSku) },
|
||||
{ field: "nama_item", match: isMatch(gt.nama_item, predictedItemName) },
|
||||
{ field: "expiry_date", match: isMatch(gt.expiry_date, predictedExpiry) }
|
||||
{ field: "no_sku", match: isMatch(gt.no_sku, bestMatchSku), predicted: bestMatchSku },
|
||||
{ field: "nama_item", match: isMatch(gt.nama_item, predictedItemName), predicted: predictedItemName },
|
||||
{ field: "expiry_date", match: isMatch(gt.expiry_date, predictedExpiry), predicted: predictedExpiry }
|
||||
];
|
||||
|
||||
const item: ResultItem = {
|
||||
@@ -380,8 +404,10 @@ async function main() {
|
||||
perImagePct[gt.filename] = (score / checks.length) * 100;
|
||||
console.log(`done (${score}/3)`);
|
||||
} catch (err) {
|
||||
console.log(`FAILED (${(err as Error).message})`);
|
||||
const message = (err as Error).message;
|
||||
console.log(`FAILED (${message})`);
|
||||
results.failed.push(gt.filename);
|
||||
failedDetail.push({ filename: gt.filename, error: message });
|
||||
}
|
||||
}
|
||||
|
||||
@@ -403,6 +429,26 @@ async function main() {
|
||||
perImage: perImagePct
|
||||
};
|
||||
fs.appendFileSync(HISTORY_PATH, JSON.stringify(entry) + "\n");
|
||||
|
||||
if (DETAIL_DUMP_PATH) {
|
||||
const detail = {
|
||||
timestamp: entry.timestamp,
|
||||
validation: results.validation.map(r => ({
|
||||
filename: r.gt.filename,
|
||||
method: r.method,
|
||||
confidence: r.confidence,
|
||||
checks: r.checks.map(c => ({
|
||||
field: c.field,
|
||||
match: c.match,
|
||||
expected: r.gt[c.field],
|
||||
predicted: c.predicted
|
||||
}))
|
||||
})),
|
||||
failed: failedDetail
|
||||
};
|
||||
fs.writeFileSync(DETAIL_DUMP_PATH, JSON.stringify(detail, null, 2), "utf8");
|
||||
console.log(`\nDetail dump written to ${DETAIL_DUMP_PATH}`);
|
||||
}
|
||||
}
|
||||
|
||||
main().catch(err => {
|
||||
|
||||
@@ -0,0 +1,76 @@
|
||||
// One-off script: pull every Validation Set image that had at least one
|
||||
// mismatched field (or failed to scan entirely) out of
|
||||
// sources/product-test-images-fixed/ into sources/product-test-images-undetected/,
|
||||
// renamed to show which field(s) missed, plus a manifest.md with
|
||||
// expected-vs-predicted per field — so a human can tell at a glance whether
|
||||
// a miss is an OCR problem, a classification problem, or the photo itself
|
||||
// lacking the data (e.g. expiry code out of frame).
|
||||
// Usage: node scripts/build-undetected-set.mjs <path-to-detail-dump.json>
|
||||
import fs from "fs";
|
||||
import path from "path";
|
||||
|
||||
const detailPath = process.argv[2];
|
||||
if (!detailPath) {
|
||||
console.error("Usage: node scripts/build-undetected-set.mjs <detail-dump.json>");
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
const FIXED_DIR = path.join("sources", "product-test-images-fixed");
|
||||
const OUT_DIR = path.join("sources", "product-test-images-undetected");
|
||||
|
||||
const detail = JSON.parse(fs.readFileSync(detailPath, "utf8"));
|
||||
|
||||
if (fs.existsSync(OUT_DIR)) fs.rmSync(OUT_DIR, { recursive: true });
|
||||
fs.mkdirSync(OUT_DIR, { recursive: true });
|
||||
|
||||
const manifestRows = [];
|
||||
let copied = 0;
|
||||
|
||||
for (const item of detail.validation) {
|
||||
const failedFields = item.checks.filter((c) => !c.match);
|
||||
if (!failedFields.length) continue;
|
||||
|
||||
const ext = path.extname(item.filename);
|
||||
const base = path.basename(item.filename, ext);
|
||||
const tag = failedFields.map((c) => c.field).join(",");
|
||||
const destName = `${base} [FAIL ${tag}]${ext}`;
|
||||
fs.copyFileSync(path.join(FIXED_DIR, item.filename), path.join(OUT_DIR, destName));
|
||||
copied++;
|
||||
|
||||
for (const c of item.checks) {
|
||||
manifestRows.push({
|
||||
image: destName,
|
||||
field: c.field,
|
||||
match: c.match,
|
||||
expected: c.expected || "(empty)",
|
||||
predicted: c.predicted || "(empty — not detected)"
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
for (const f of detail.failed) {
|
||||
const srcPath = path.join(FIXED_DIR, f.filename);
|
||||
if (!fs.existsSync(srcPath)) continue;
|
||||
const ext = path.extname(f.filename);
|
||||
const base = path.basename(f.filename, ext);
|
||||
const destName = `${base} [FAILED-scan]${ext}`;
|
||||
fs.copyFileSync(srcPath, path.join(OUT_DIR, destName));
|
||||
copied++;
|
||||
manifestRows.push({ image: destName, field: "(entire scan)", match: false, expected: "-", predicted: `ERROR: ${f.error}` });
|
||||
}
|
||||
|
||||
const lines = [
|
||||
"# Undetected / mismatched images",
|
||||
"",
|
||||
`Generated ${detail.timestamp} from the accuracy-check-scan.mts detail dump.`,
|
||||
`${copied} of ${detail.validation.length + detail.failed.length} Validation Set images had at least one wrong/missing field or failed to scan.`,
|
||||
"",
|
||||
"| Image | Field | OK? | Expected | Predicted |",
|
||||
"|---|---|---|---|---|"
|
||||
];
|
||||
for (const r of manifestRows) {
|
||||
lines.push(`| ${r.image} | ${r.field} | ${r.match ? "✓" : "✗"} | ${r.expected} | ${r.predicted} |`);
|
||||
}
|
||||
fs.writeFileSync(path.join(OUT_DIR, "manifest.md"), lines.join("\n") + "\n", "utf8");
|
||||
|
||||
console.log(`Copied ${copied} images into ${OUT_DIR}, wrote manifest.md (${manifestRows.length} rows).`);
|
||||
@@ -0,0 +1,59 @@
|
||||
// One-off capture: hit /api/scan-pfm for every image in
|
||||
// sources/product-test-images-fixed/ and save the full text-level response
|
||||
// (classification all_probabilities, OCR text_lines, extracted fields,
|
||||
// possibleMatches) minus base64 image blobs to
|
||||
// sources/product_scan_fullcap.json. This lets classification re-ranking
|
||||
// experiments run offline against ground truth in seconds instead of
|
||||
// re-running the 11-minute GPU batch per iteration.
|
||||
// Usage: node scripts/capture-scan-responses.mjs [baseUrl]
|
||||
import fs from "fs";
|
||||
import path from "path";
|
||||
|
||||
const BASE_URL = process.argv[2] || process.env.ACCURACY_BASE_URL || "http://127.0.0.1:3000";
|
||||
const FIXED_DIR = path.join("sources", "product-test-images-fixed");
|
||||
const OUT_PATH = path.join("sources", "product_scan_fullcap.json");
|
||||
|
||||
const files = fs.readdirSync(FIXED_DIR)
|
||||
.filter((f) => /\.(jpe?g|png|webp)$/i.test(f))
|
||||
.sort((a, b) => parseInt(a) - parseInt(b));
|
||||
|
||||
const results = [];
|
||||
for (const filename of files) {
|
||||
const b64 = "data:image/jpeg;base64," +
|
||||
fs.readFileSync(path.join(FIXED_DIR, filename), "base64");
|
||||
process.stdout.write(`Capturing ${filename}... `);
|
||||
try {
|
||||
const res = await fetch(`${BASE_URL}/api/scan-pfm`, {
|
||||
method: "POST",
|
||||
headers: { "Content-Type": "application/json" },
|
||||
body: JSON.stringify({ image_base64: b64 }),
|
||||
signal: AbortSignal.timeout(120_000)
|
||||
});
|
||||
if (!res.ok) throw new Error(`HTTP ${res.status}`);
|
||||
const d = await res.json();
|
||||
results.push({
|
||||
filename,
|
||||
classification: {
|
||||
top1_name: d.classification?.top1_name,
|
||||
top1_confidence: d.classification?.top1_confidence,
|
||||
method: d.classification?.method,
|
||||
all_probabilities: (d.classification?.all_probabilities || []).slice(0, 15)
|
||||
},
|
||||
ocr: {
|
||||
text_lines: d.ocr?.text_lines || [],
|
||||
extracted_sku: d.ocr?.extracted_sku ?? null,
|
||||
extracted_product_name: d.ocr?.extracted_product_name ?? null,
|
||||
extracted_expired_date: d.ocr?.extracted_expired_date ?? null,
|
||||
expired_source_line: d.ocr?.expired_source_line ?? null
|
||||
},
|
||||
possibleMatches: d.possibleMatches || []
|
||||
});
|
||||
console.log("ok");
|
||||
} catch (err) {
|
||||
console.log(`FAILED (${err.message})`);
|
||||
results.push({ filename, error: err.message });
|
||||
}
|
||||
}
|
||||
|
||||
fs.writeFileSync(OUT_PATH, JSON.stringify(results, null, 2), "utf8");
|
||||
console.log(`\nWrote ${results.length} captures to ${OUT_PATH}`);
|
||||
@@ -0,0 +1,178 @@
|
||||
// Offline experiment: OCR-evidence re-ranking of DINOv2 top-K candidates.
|
||||
//
|
||||
// Reads sources/product_scan_fullcap.json (captured live responses, see
|
||||
// capture-scan-responses.mjs) + sources/product_manual_labels.json (ground
|
||||
// truth) and simulates candidate re-ranking without touching the GPU stack,
|
||||
// reporting fixed-vs-broken counts per parameter combination. The winning
|
||||
// parameters get ported into pfm-web-app/src/utils/product-scan.ts.
|
||||
//
|
||||
// Idea: DINOv2's near-twin confusions (same brand, different flavor/size)
|
||||
// are exactly the cases where the *printed variant words* differ - and
|
||||
// PaddleOCR usually reads some of them. So within a narrow similarity band
|
||||
// of the top-1, prefer the candidate whose distinguishing name tokens
|
||||
// actually appear in the OCR'd text.
|
||||
//
|
||||
// Usage: node scripts/experiment-rerank.mjs
|
||||
import fs from "fs";
|
||||
import path from "path";
|
||||
|
||||
const cap = JSON.parse(fs.readFileSync(path.join("sources", "product_scan_fullcap.json"), "utf8"));
|
||||
const labels = JSON.parse(fs.readFileSync(path.join("sources", "product_manual_labels.json"), "utf8"));
|
||||
const gtBySku = new Map(labels.map((l) => [l.filename, l.no_sku]));
|
||||
|
||||
function classSku(className) {
|
||||
// Class names are foto-kemasan-v2 folder names: "<SKU> <NAME...>"
|
||||
return (className || "").trim().split(/\s+/)[0] || "";
|
||||
}
|
||||
|
||||
function tokenize(name) {
|
||||
return name
|
||||
.toUpperCase()
|
||||
.split(/[^A-Z0-9]+/)
|
||||
.filter((t) => t.length >= 2);
|
||||
}
|
||||
|
||||
function editDistance1(a, b) {
|
||||
// true if edit distance <= 1 (same length: 1 substitution; off-by-one: 1 indel)
|
||||
if (a === b) return true;
|
||||
const la = a.length, lb = b.length;
|
||||
if (Math.abs(la - lb) > 1) return false;
|
||||
if (la === lb) {
|
||||
let diff = 0;
|
||||
for (let i = 0; i < la; i++) if (a[i] !== b[i]) diff++;
|
||||
return diff <= 1;
|
||||
}
|
||||
const [s, l] = la < lb ? [a, b] : [b, a];
|
||||
let i = 0, j = 0, skipped = false;
|
||||
while (i < s.length && j < l.length) {
|
||||
if (s[i] === l[j]) { i++; j++; }
|
||||
else if (!skipped) { skipped = true; j++; }
|
||||
else return false;
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
function buildOcrIndex(textLines) {
|
||||
const joined = textLines.join(" ").toUpperCase();
|
||||
const squashed = joined.replace(/[^A-Z0-9]/g, "");
|
||||
const tokens = new Set(tokenize(joined));
|
||||
return { squashed, tokens };
|
||||
}
|
||||
|
||||
function tokenInOcr(token, ocrIdx, fuzzy) {
|
||||
if (token.length >= 4 && ocrIdx.squashed.includes(token)) return true;
|
||||
if (ocrIdx.tokens.has(token)) return true;
|
||||
if (fuzzy && token.length >= 5) {
|
||||
for (const t of ocrIdx.tokens) {
|
||||
if (Math.abs(t.length - token.length) <= 1 && editDistance1(token, t)) return true;
|
||||
}
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
function ocrEvidenceScore(candTokens, bandTokenCounts, bandSize, ocrIdx, fuzzy) {
|
||||
// Coverage-normalized, rarity-weighted evidence: fraction of this
|
||||
// candidate's *distinctive* name tokens (weighted by band rarity) that
|
||||
// actually appear in the OCR'd text. Normalizing by the candidate's own
|
||||
// distinctive-token mass is what stops generic packaging words from
|
||||
// hijacking the ranking - a candidate whose name promises FRENCH +
|
||||
// INSTITUSI + 2KG but whose package shows only "French Fries" scores
|
||||
// 1/3, losing to a candidate whose 2 distinctive tokens both appear.
|
||||
let matched = 0;
|
||||
let total = 0;
|
||||
for (const tok of new Set(candTokens)) {
|
||||
const nWith = bandTokenCounts.get(tok) || 1;
|
||||
if (nWith >= bandSize) continue; // shared by all -> no signal
|
||||
const w = 1 / nWith;
|
||||
total += w;
|
||||
if (tokenInOcr(tok, ocrIdx, fuzzy)) matched += w;
|
||||
}
|
||||
return total > 0 ? matched / total : 0;
|
||||
}
|
||||
|
||||
function skuFuzzyBoost(extractedSku, candidateSku) {
|
||||
if (!extractedSku || extractedSku.length < 7) return 0;
|
||||
if (extractedSku === candidateSku) return 10; // exact (normally pinned upstream anyway)
|
||||
return editDistance1(extractedSku, candidateSku) ? 1 : 0;
|
||||
}
|
||||
|
||||
function runConfig({ K, BAND, MARGIN, FUZZY, SKU_BOOST_W }) {
|
||||
let baselineCorrect = 0, rerankCorrect = 0, fixed = [], broken = [];
|
||||
for (const item of cap) {
|
||||
if (item.error) continue;
|
||||
const gt = gtBySku.get(item.filename);
|
||||
if (!gt) continue;
|
||||
const probs = item.classification?.all_probabilities || [];
|
||||
if (!probs.length) continue;
|
||||
|
||||
const top1Sku = classSku(probs[0].name);
|
||||
const baselineRight = top1Sku === gt;
|
||||
if (baselineRight) baselineCorrect++;
|
||||
|
||||
// Candidate band: within BAND of top-1 similarity, capped at K
|
||||
const top1Sim = probs[0].confidence;
|
||||
const band = probs.slice(0, K).filter((p) => p.confidence >= top1Sim - BAND);
|
||||
|
||||
const ocrIdx = buildOcrIndex(item.ocr?.text_lines || []);
|
||||
const candInfos = band.map((p) => {
|
||||
const sku = classSku(p.name);
|
||||
const tokens = tokenize(p.name.replace(sku, ""));
|
||||
return { sku, sim: p.confidence, tokens };
|
||||
});
|
||||
const bandTokenCounts = new Map();
|
||||
for (const c of candInfos) {
|
||||
for (const tok of new Set(c.tokens)) {
|
||||
bandTokenCounts.set(tok, (bandTokenCounts.get(tok) || 0) + 1);
|
||||
}
|
||||
}
|
||||
for (const c of candInfos) {
|
||||
c.ocrScore = ocrEvidenceScore(c.tokens, bandTokenCounts, candInfos.length, ocrIdx, FUZZY)
|
||||
+ SKU_BOOST_W * skuFuzzyBoost(item.ocr?.extracted_sku || "", c.sku);
|
||||
}
|
||||
|
||||
// Switch away from top-1 only when a band-mate has clearly stronger OCR evidence
|
||||
let chosen = candInfos[0];
|
||||
for (const c of candInfos.slice(1)) {
|
||||
if (c.ocrScore >= chosen.ocrScore + MARGIN) chosen = c;
|
||||
}
|
||||
|
||||
const rerankRight = chosen.sku === gt;
|
||||
if (rerankRight) rerankCorrect++;
|
||||
if (!baselineRight && rerankRight) fixed.push(item.filename);
|
||||
if (baselineRight && !rerankRight) broken.push(item.filename);
|
||||
}
|
||||
return { baselineCorrect, rerankCorrect, fixed, broken };
|
||||
}
|
||||
|
||||
const grid = [];
|
||||
for (const K of [5, 8, 12]) {
|
||||
for (const BAND of [0.04, 0.06, 0.08, 0.12]) {
|
||||
// Coverage scores live in [0, 1]; margin is the minimum coverage lead a
|
||||
// band-mate needs over the current pick before we switch away from it.
|
||||
for (const MARGIN of [0.15, 0.25, 0.35, 0.5]) {
|
||||
for (const FUZZY of [true, false]) {
|
||||
for (const SKU_BOOST_W of [0, 2]) {
|
||||
grid.push({ K, BAND, MARGIN, FUZZY, SKU_BOOST_W });
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
const results = grid.map((cfg) => ({ cfg, ...runConfig(cfg) }));
|
||||
results.sort((a, b) => (b.rerankCorrect - b.broken.length * 0.01) - (a.rerankCorrect - a.broken.length * 0.01));
|
||||
|
||||
console.log(`Images evaluated: ${cap.filter((i) => !i.error && gtBySku.has(i.filename)).length}`);
|
||||
console.log(`Baseline (DINOv2 top-1) correct: ${results[0].baselineCorrect}\n`);
|
||||
console.log("Top 12 configs by re-ranked correct count:");
|
||||
for (const r of results.slice(0, 12)) {
|
||||
console.log(
|
||||
` correct=${r.rerankCorrect} (+${r.fixed.length}/-${r.broken.length}) ` +
|
||||
`K=${r.cfg.K} BAND=${r.cfg.BAND} MARGIN=${r.cfg.MARGIN} FUZZY=${r.cfg.FUZZY} SKUW=${r.cfg.SKU_BOOST_W}`
|
||||
);
|
||||
}
|
||||
|
||||
const best = results[0];
|
||||
console.log(`\nBest config detail: ${JSON.stringify(best.cfg)}`);
|
||||
console.log(` fixed (${best.fixed.length}): ${best.fixed.join(", ")}`);
|
||||
console.log(` broken (${best.broken.length}): ${best.broken.join(", ")}`);
|
||||
@@ -0,0 +1,47 @@
|
||||
// One-off script: freeze the current 74-image product-scan Validation Set into
|
||||
// a dedicated, stable folder (sources/product-test-images-fixed/) so re-running
|
||||
// the accuracy harness always scores the exact same images, independent of
|
||||
// whatever new photos get dropped into the live-intake folder
|
||||
// (sources/product-test-images/, still fed by the /manual-label-scan page).
|
||||
// Renames each image "<index> <no_sku>.<ext>" (index = its stable position,
|
||||
// 1-based) and updates product_manual_labels.json's flat-filename entries to
|
||||
// match. Run once from backend/: node scripts/freeze-validation-set.mjs
|
||||
import fs from "fs";
|
||||
import path from "path";
|
||||
|
||||
const LIVE_DIR = path.join("sources", "product-test-images");
|
||||
const FIXED_DIR = path.join("sources", "product-test-images-fixed");
|
||||
const LABELS_PATH = path.join("sources", "product_manual_labels.json");
|
||||
|
||||
const labels = JSON.parse(fs.readFileSync(LABELS_PATH, "utf8"));
|
||||
const flatEntries = labels.filter((l) => !l.filename.includes("/"));
|
||||
|
||||
if (!fs.existsSync(FIXED_DIR)) fs.mkdirSync(FIXED_DIR, { recursive: true });
|
||||
|
||||
// Resolve by SKU prefix (not entry.filename directly) so this script is
|
||||
// idempotent/rerunnable even after a previous run already renamed
|
||||
// entry.filename to "<index> <sku>.<ext>" — the live-intake folder always
|
||||
// keeps its original "<sku> <product>__<camera-filename>.<ext>" names.
|
||||
const liveFiles = fs.readdirSync(LIVE_DIR);
|
||||
function findSourceFile(no_sku) {
|
||||
const match = liveFiles.find((f) => f.startsWith(`${no_sku} `) || f.startsWith(`${no_sku}__`));
|
||||
if (!match) return null;
|
||||
return path.join(LIVE_DIR, match);
|
||||
}
|
||||
|
||||
let copied = 0;
|
||||
flatEntries.forEach((entry, i) => {
|
||||
const index = i + 1;
|
||||
const srcPath = findSourceFile(entry.no_sku);
|
||||
if (!srcPath) {
|
||||
throw new Error(`Missing source image for ${entry.no_sku} in ${LIVE_DIR}`);
|
||||
}
|
||||
const ext = path.extname(srcPath);
|
||||
const newFilename = `${index} ${entry.no_sku}${ext}`;
|
||||
fs.copyFileSync(srcPath, path.join(FIXED_DIR, newFilename));
|
||||
entry.filename = newFilename;
|
||||
copied++;
|
||||
});
|
||||
|
||||
fs.writeFileSync(LABELS_PATH, JSON.stringify(labels, null, 2), "utf8");
|
||||
console.log(`Copied ${copied} images into ${FIXED_DIR} and updated ${LABELS_PATH}.`);
|
||||
@@ -0,0 +1,130 @@
|
||||
// One-off script: fill ground-truth labels for backend/sources/product-test-images/
|
||||
// (the product-scan accuracy harness's Validation Set, previously 0 labeled images).
|
||||
// Run once from backend/: node scripts/seed-validation-labels.mjs
|
||||
import fs from "fs";
|
||||
import path from "path";
|
||||
|
||||
const IMAGES_DIR = path.join("sources", "product-test-images");
|
||||
const LABELS_PATH = path.join("sources", "product_manual_labels.json");
|
||||
|
||||
// [no_sku, nama_item, expiry_date ("" = not legible in photo, needs re-shoot)]
|
||||
const DATA = [
|
||||
["11110059", "CEKER BERKUKU FROZEN PACK 1 KG(*)", "13/06/2027"],
|
||||
["11140051", "AMPELA FROZEN PACK 1 KG(*)", "26/11/2026"],
|
||||
["11620056", "SBL (FILLET PAHA) 1 KG(*)", "27/02/2027"],
|
||||
["11650053", "PAHA ATAS 1 KG(*)", "26/02/2027"],
|
||||
["11660050", "PAHA BAWAH (1 KG)(*)", "22/06/2027"],
|
||||
["11710051", "DADA UTUH (1 KG)(*)", ""],
|
||||
["11818300", "CP-BEBEK GORENG 400GR/PAC", ""],
|
||||
["1195008A", "RTC CHICKEN KALASAN 400 GR (PAC)", ""],
|
||||
["11959937", "SATE AYAM FRESHMART 360 GR (PAC)", "24/11/2026"],
|
||||
["12010111", "FIESTA CRISPY BUBBLE 400 GR/PAC", "07/05/2027"],
|
||||
["12010115", "FIESTA NUGGET ZOO 400 GR/PAC", "01/12/2026"],
|
||||
["12010117", "FIESTA NUGGET HAPPY STAR 400 GR/PAC", "27/08/2026"],
|
||||
["12010119", "FIESTA NUGGET CHEESE 123 400 GR/PAC", "09/04/2027"],
|
||||
["12010121", "FIESTA NUGGET PIZZABC 400 GR/PAC", "12/11/2026"],
|
||||
["12010127", "FIESTA SPICY NUGGET 400 GR/PAC", "09/12/2027"],
|
||||
["12010509", "CHAMP CRUNCHY NUGGET 450 GR/PAC", "15/04/2027"],
|
||||
["12010515", "CHAMP KOIN KOMBINASI 450 GR/PAC", "15/04/2027"],
|
||||
["12010519", "CHAMP NUGGET STICK 900 GR/PAC", "08/03/2027"],
|
||||
["12012202", "ASIMO NUGGET KOMBINASI 1 KG/PAC", "13/05/2027"],
|
||||
["12012501", "AKUMO CHICKEN NAGET 250 GR", "21/05/2027"],
|
||||
["12012503", "AKUMO CHICKEN NUGGET 1000 GR", "17/06/2027"],
|
||||
["12012505", "AKUMO KOIN 400 GR/PAC", "16/11/2026"],
|
||||
["12020102", "FIESTA SPICY WING 400 GR/PAC", "22/05/2027"],
|
||||
["12030102", "FIESTA STIKIE 200 GR/PAC", "26/02/2027"],
|
||||
["12030403", "GOLDEN FIESTA STIKIE W/ SWEET CHILLI SAUCE 500GR", "07/04/2027"],
|
||||
["12032502", "AKUMO CHICKEN STIK 500 GR", "09/03/2027"],
|
||||
["12040101", "FIESTA SCHNITZEL 400 GR/PAC", "09/10/2026"],
|
||||
["12040102", "FIESTA CRISPY BUBBLE KATSU 400 GR/PAC", "12/04/2027"],
|
||||
["12060103", "FIESTA KARAGE 200 GR/PAC", ""],
|
||||
["12060402", "GOLDEN FIESTA KARAGE CHILI SAUCE 500GR", "15/05/2027"],
|
||||
["12080101", "FIESTA SPICY CHICK 400 GR/PAC", "26/04/2027"],
|
||||
["12130102", "FIESTA CRISPY BURGER 360 GR (NEW)", "27/04/2027"],
|
||||
["12150201", "FIESTA DS CRISPY CRUNCH 300 GR/PAC", "03/06/2027"],
|
||||
["12150501", "CHAMP CRUNCHY HOTZZ 300 GR/PAC", ""],
|
||||
["12190103", "FIESTA DELISTRIPE 400 GR/PAC", "07/04/2027"],
|
||||
["12240103", "FIESTA YAKINIKU R/BITES 400 GR/PAC", "07/05/2027"],
|
||||
["13010111", "FIESTA SOSIS BRATWURST 300 GR", "25/03/2027"],
|
||||
["13010116", "FIESTA SSG ORIGINAL 300 GR", "23/06/2027"],
|
||||
["13010118", "FIESTA RTG SSG 65 GR/PAC", ""],
|
||||
["13010120", "FIESTA RTG C/CHEESY MELTS 65 GR/PAC", ""],
|
||||
["13010122", "FIESTA RTG SAUSAGE WITH HOT LAVA 60G", ""],
|
||||
["13010123", "FIESTA RTG SAUSAGE WITH CHEESE LAVA 60G", "25/10/2026"],
|
||||
["13010125", "FIESTA RTG SAUSAGE WITH MENTAI LAVA 60GR", "05/11/2026"],
|
||||
["13010518", "CHAMP SSG JUMBO BAKAR 500 GR/PAC", ""],
|
||||
["13010524", "CHAMP SSG JUMBO BAKAR 500 GR/PAC (NEW)", ""],
|
||||
["13012206", "ASIMO SOSIS AYAM KOMBINASI 500 GR", "15/03/2027"],
|
||||
["13030501", "CHAMP CHICK MEATBALL 200 GR", ""],
|
||||
["13070506", "CHAMP FRANKFURTER SSG 375GR", "13/03/2027"],
|
||||
["13100512", "CHAMP CHICK SSG S/SANTAP ORIG 546GR (CAN)", ""],
|
||||
["15010101", "FIESTA SHOESTRING 500 GR", "04/06/2027"],
|
||||
["15010102", "FIESTA SHOESTRING 1000 GR", "05/03/2027"],
|
||||
["15010107", "FIESTA FRENCH F SHOESTRING INSTITUSI 2KG", "17/06/2027"],
|
||||
["15020101", "FIESTA STRAIGHT CUT 500 GR", "19/06/2027"],
|
||||
["15020102", "FIESTA STRAIGHT CUT 1000 GR", "18/05/2027"],
|
||||
["15030101", "FIESTA CRINKLE CUT 500 GR", ""],
|
||||
["15030102", "FIESTA CRINKLE CUT 1000 GR", "07/04/2027"],
|
||||
["16060113", "FIESTA CHICK SIOMAY 180GR (NEW)", "24/02/2027"],
|
||||
["16060114", "FIESTA GYOZA 180 GR (NEW)", "18/05/2027"],
|
||||
["17200109", "FIESTA RTS C/TERIYAKI 300GR/PAC", ""],
|
||||
["17210106", "FIESTA RTS B/YAKINIKU 300GR/PAC", "19/05/2027"],
|
||||
["17210107", "FIESTA RTS B/RENDANG 300GR/PAC", "23/06/2027"],
|
||||
["17210108", "FIESTA RTS B/BLACKPEPPER 300GR/PAC", "09/05/2027"],
|
||||
["17210109", "FIESTA RTS B/BULGOGI 300GR/PAC", "18/05/2027"],
|
||||
["20040101", "FIESTA RAMEN BEKU 570 GR/PAC", "23/06/2026"],
|
||||
["20120102", "FIESTA T/B AYAM GORENG 80 GR", ""],
|
||||
["20120105", "FIESTA T/B SERBAGUNA Â 80 GR", "09/03/2027"],
|
||||
["20120115", "FIESTA RACIK AYAM GORENG 20 GR/PAC", ""],
|
||||
["20120116", "FIESTA RACIK NASI GORENG 20 GR/PAC", "13/10/2026"],
|
||||
["21000123", "FIESTA RICE W/GEPREK CHICKEN 320GR/PAC", "30/04/2027"],
|
||||
["21000126", "NEW FIESTA CHICK RENDANG W RICE 320GR (PAC)", "16/04/2027"],
|
||||
["21000130", "NEW FIESTA RICE W/C CHEESE BULDAK 320GR (PAC)", "09/04/2027"],
|
||||
["21000137", "FIESTA HAINAMESE CHICKEN RICE 320GR (PAC)", "12/02/2027"],
|
||||
["21010101", "FIESTA TRUFFLE GYUDON 320 GR/PAC", "18/03/2027"],
|
||||
["21200107", "NEW FIESTA SPAGHETTI CARBONARA 300GR (PAC)", "20/05/2027"],
|
||||
];
|
||||
|
||||
const files = fs
|
||||
.readdirSync(IMAGES_DIR)
|
||||
.filter((f) => f !== "README.md" && f !== "filelist.txt")
|
||||
.sort();
|
||||
|
||||
if (files.length !== DATA.length) {
|
||||
throw new Error(`File count ${files.length} != DATA count ${DATA.length}`);
|
||||
}
|
||||
|
||||
const existing = JSON.parse(fs.readFileSync(LABELS_PATH, "utf8"));
|
||||
const now = new Date().toISOString();
|
||||
|
||||
let added = 0;
|
||||
let skipped = 0;
|
||||
let unreadable = 0;
|
||||
|
||||
for (let i = 0; i < files.length; i++) {
|
||||
const filename = files[i];
|
||||
const [no_sku, nama_item, expiry_date] = DATA[i];
|
||||
|
||||
if (existing.some((l) => l.filename === filename)) {
|
||||
skipped++;
|
||||
continue;
|
||||
}
|
||||
|
||||
if (!expiry_date) unreadable++;
|
||||
|
||||
existing.push({
|
||||
filename,
|
||||
no_sku,
|
||||
nama_item,
|
||||
expiry_date,
|
||||
top1_confidence: null,
|
||||
notes: expiry_date
|
||||
? ""
|
||||
: "expiry date not legible in photo (cropped/blurry/out of frame) - needs re-shoot",
|
||||
saved_at: now,
|
||||
});
|
||||
added++;
|
||||
}
|
||||
|
||||
fs.writeFileSync(LABELS_PATH, JSON.stringify(existing, null, 2), "utf8");
|
||||
console.log(`Added ${added} labels (${unreadable} flagged with no expiry_date), skipped ${skipped} already-labeled.`);
|
||||
|
After Width: | Height: | Size: 3.4 MiB |
|
After Width: | Height: | Size: 140 KiB |
|
After Width: | Height: | Size: 1.3 MiB |
|
After Width: | Height: | Size: 257 KiB |
|
After Width: | Height: | Size: 3.7 MiB |
|
After Width: | Height: | Size: 229 KiB |
|
After Width: | Height: | Size: 254 KiB |
|
After Width: | Height: | Size: 268 KiB |
|
After Width: | Height: | Size: 4.2 MiB |
|
After Width: | Height: | Size: 184 KiB |
|
After Width: | Height: | Size: 220 KiB |
|
After Width: | Height: | Size: 2.3 MiB |
|
After Width: | Height: | Size: 227 KiB |
|
After Width: | Height: | Size: 170 KiB |
|
After Width: | Height: | Size: 217 KiB |
|
After Width: | Height: | Size: 217 KiB |
|
After Width: | Height: | Size: 201 KiB |
|
After Width: | Height: | Size: 3.9 MiB |
|
After Width: | Height: | Size: 239 KiB |
|
After Width: | Height: | Size: 194 KiB |
|
After Width: | Height: | Size: 200 KiB |
|
After Width: | Height: | Size: 207 KiB |
|
After Width: | Height: | Size: 3.1 MiB |
|
After Width: | Height: | Size: 275 KiB |
|
After Width: | Height: | Size: 231 KiB |
|
After Width: | Height: | Size: 304 KiB |
|
After Width: | Height: | Size: 3.3 MiB |
|
After Width: | Height: | Size: 278 KiB |
|
After Width: | Height: | Size: 2.0 MiB |
|
After Width: | Height: | Size: 178 KiB |
|
After Width: | Height: | Size: 293 KiB |
|
After Width: | Height: | Size: 284 KiB |
|
After Width: | Height: | Size: 223 KiB |
|
After Width: | Height: | Size: 164 KiB |
|
After Width: | Height: | Size: 3.4 MiB |
|
After Width: | Height: | Size: 3.4 MiB |
|
After Width: | Height: | Size: 1.8 MiB |
|
After Width: | Height: | Size: 225 KiB |
|
After Width: | Height: | Size: 4.5 MiB |
|
After Width: | Height: | Size: 232 KiB |
|
After Width: | Height: | Size: 3.6 MiB |
|
After Width: | Height: | Size: 235 KiB |
|
After Width: | Height: | Size: 2.4 MiB |
|
After Width: | Height: | Size: 149 KiB |
|
After Width: | Height: | Size: 285 KiB |
|
After Width: | Height: | Size: 3.6 MiB |
|
After Width: | Height: | Size: 3.8 MiB |
|
After Width: | Height: | Size: 3.7 MiB |
|
After Width: | Height: | Size: 106 KiB |
|
After Width: | Height: | Size: 2.5 MiB |
|
After Width: | Height: | Size: 175 KiB |
|
After Width: | Height: | Size: 2.9 MiB |
|
After Width: | Height: | Size: 208 KiB |
|
After Width: | Height: | Size: 143 KiB |
|
After Width: | Height: | Size: 251 KiB |
|
After Width: | Height: | Size: 2.4 MiB |
|
After Width: | Height: | Size: 176 KiB |
|
After Width: | Height: | Size: 230 KiB |
|
After Width: | Height: | Size: 2.3 MiB |
|
After Width: | Height: | Size: 302 KiB |
|
After Width: | Height: | Size: 114 KiB |
|
After Width: | Height: | Size: 217 KiB |
|
After Width: | Height: | Size: 208 KiB |
|
After Width: | Height: | Size: 129 KiB |
|
After Width: | Height: | Size: 288 KiB |
|
After Width: | Height: | Size: 4.1 MiB |
|
After Width: | Height: | Size: 2.5 MiB |
|
After Width: | Height: | Size: 232 KiB |
|
After Width: | Height: | Size: 286 KiB |
|
After Width: | Height: | Size: 3.8 MiB |
|
After Width: | Height: | Size: 171 KiB |
|
After Width: | Height: | Size: 265 KiB |
|
After Width: | Height: | Size: 263 KiB |
|
After Width: | Height: | Size: 261 KiB |
|
After Width: | Height: | Size: 227 KiB |
|
After Width: | Height: | Size: 1.8 MiB |
|
After Width: | Height: | Size: 222 KiB |
|
After Width: | Height: | Size: 215 KiB |
|
After Width: | Height: | Size: 3.1 MiB |