feat(backend): scan-product accuracy 66.2% -> 79.7% + frozen validation benchmark

Accuracy work on the 79-image product-scan validation set (user goal: 90%):
- classify_ocr_server.py: 0/90/180/270-degree expiry-date search (stops at
  first hit, 0-degree fallback); classification decoupled onto the upright
  image (rotated frames regressed DINOv2 -6pts until this); cross-line date
  stitching; tiled full-res OCR pass (defeats the 4000px downscale that
  killed small inkjet dates); VL-pipeline expiry fallback with
  keyword-anchored anti-hallucination guard; VL text lines merged into
  text_lines + VL SKU retry. Visualization endpoints removed entirely
  (Visual/Spotting grids - unused by frontend, 3x per-scan GPU cost).
- product-scan.ts: coverage-normalized OCR-evidence re-ranking of DINOv2
  top-K (tuned offline: +8/-0 on top-1 misses), re-ranked class mapped to
  sku_master by SKU prefix; classifier timeout 90s->240s for fallback paths.
- Frozen benchmark: product-test-images-fixed/ (79 renamed images) +
  freeze/seed/build-undetected/capture/experiment scripts; labels trimmed to
  the 79 validation entries (training rows kept in .bak-with-training);
  5 TRAINED-ON SKUs replaced with fresh held-out photos.
- manual-label-scan page: shows last batch-test AI prediction under every
  field by default (new /api/product-scan-results); serves the fixed folder;
  fixed total hydration failure via allowedDevOrigins 127.0.0.1.
- Measured (all-79, zero failures): sku/name 87.3%, expiry 64.6%, overall
  79.7%. Tiles/VL-evidence/VL-SKU deployed but not yet batch-measured.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gr6HH7JrdsXX8AARejQboM
This commit is contained in:
Rafhan Mazaya FathurrahmanandClaude Fable 5 committed 2026-07-14 19:55:17 +07:00
1 parent 19f1facf9b
commit e76ccb60a6
156 files changed
+17129 -1384

No files matched your search

+1 -1
View File
@@ -26,7 +26,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co
## Product/SKU scanning flow — status
**How it works end-to-end** (architecture, endpoints, classification/OCR internals, retraining): [`docs/scan-product.md`](docs/scan-product.md). See [`plans/next-enhancements.md`](plans/next-enhancements.md) §2 (task 2.1) for full detail — kept there instead of a separate doc so status stays traceable against the rest of the `e`/`n` backlog. **Feature-complete as of 2026-07-08**: the backend (`config/classify_ocr_server.py` with DINOv2 similarity search + YOLO classifier fallback, `api/scan-pfm/route.ts`, `api/produk-pfm/route.ts`, DB schema), the reference photo dataset (`pfm-web-app/public/produk-pfm/foto-kemasan-v2/`, 16 SKU subfolders), the desktop frontend page (`scan-pfm/page.tsx`, full feature parity), and the trained model artifacts (`models/dinov2_index.pkl` — 118/118 photos indexed; `models/produk-pfm-classifier-26n-100e-2026-07-08.pt` — 83.3% top-1 val accuracy on the current thin dataset) all now exist and load cleanly on `pipeline-api` startup. **No mobile web page is planned**: `scan-pfm/page.tsx` is desktop-only, used to test the pipeline; real mobile product scanning goes through the Flutter app instead, so `m-scan-pfm/page.tsx` and its `nginx.conf` route are intentionally left unbuilt/dead (see plan task 2.2, cancelled 2026-07-08). Not yet done: an actual browser pass uploading a photo through `/scan-pfm` end-to-end (verified via container logs/model-loading so far, not a UI test).
**How it works end-to-end** (architecture, endpoints, classification/OCR internals, retraining): [`docs/scan-product.md`](docs/scan-product.md). See [`plans/next-enhancements.md`](plans/next-enhancements.md) §2 (task 2.1) for full detail — kept there instead of a separate doc so status stays traceable against the rest of the `e`/`n` backlog. **Feature-complete as of 2026-07-08**: the backend (`config/classify_ocr_server.py` with DINOv2 similarity search + YOLO classifier fallback, `api/scan-pfm/route.ts`, `api/produk-pfm/route.ts`, DB schema), the reference photo dataset (`pfm-web-app/public/produk-pfm/foto-kemasan-v2/`, 81 SKU subfolders as of 2026-07-14, up from the original 16 — target ~230), the desktop frontend page (`scan-pfm/page.tsx`, full feature parity), and the trained model artifacts (`models/dinov2_index.pkl` — 2,493/2,493 photos indexed as of 2026-07-14; `models/produk-pfm-classifier-26n-100e-2026-07-14.pt` — 85.8% top-1 / 94.4% top-5 val accuracy across all 81 classes, retrained 2026-07-14 in 54m21s on an RTX 2060) all now exist and load cleanly on `pipeline-api` startup. **No mobile web page is planned**: `scan-pfm/page.tsx` is desktop-only, used to test the pipeline; real mobile product scanning goes through the Flutter app instead, so `m-scan-pfm/page.tsx` and its `nginx.conf` route are intentionally left unbuilt/dead (see plan task 2.2, cancelled 2026-07-08). Not yet done: an actual browser pass uploading a photo through `/scan-pfm` end-to-end (verified via container logs/model-loading so far, not a UI test).
## Confidentiality
+184 -92
View File
@@ -316,6 +316,24 @@ def extract_expired_date(text_lines):
if match:
return pick(match, idx, line)
# 3.6) Keyword line + date split onto an adjacent line (PaddleOCR sometimes
# detects "BB"/"Baik digunakan" as its own box, separate from the date
# digits in a neighboring box, e.g. "BB" / "05032027" as two lines).
for idx, line in enumerate(cleaned_lines):
if not line_has_exp_keyword(line):
continue
for j in (idx + 1, idx - 1, idx + 2):
if j < 0 or j >= len(cleaned_lines) or j == idx:
continue
neighbor = cleaned_lines[j]
combined = f"{line} {neighbor}" if j > idx else f"{neighbor} {line}"
match = (BB_ATTACHED_DATE_RE.search(combined)
or DDMMYYYY_RE.search(combined)
or DD_MM_YYYY_RE.search(combined))
if match:
report_idx = j if sum(c.isdigit() for c in neighbor) > sum(c.isdigit() for c in line) else idx
return pick(match, report_idx, combined)
# 4) Any line — spaced DD MM YYYY
for idx, line in enumerate(cleaned_lines):
match = DD_MM_YYYY_RE.search(line)
@@ -471,8 +489,24 @@ async def classify_ocr(payload: ScanRequest):
img_data = base64.b64decode(payload.image_base64.split(",")[-1])
raw_image = Image.open(io.BytesIO(img_data))
image = ImageOps.exif_transpose(raw_image).convert("RGB")
# Classification always sees the original upright orientation - the
# 90-degree expiry-date search below may rotate `image` to a
# sideways/upside-down orientation that DINOv2/YOLO were never
# trained on (their reference photos are all shot upright), so using
# a rotated frame there would hurt classification, not help it.
classification_image = image
# First-pass PaddleOCR to check orientation based on Expiry Date
# Multi-orientation expiry-date search: some photos are captured with
# the whole frame rotated ~90 degrees from upright (e.g. staff held
# the phone in portrait for a package whose printed date runs
# horizontally), so the expiry stamp - and the product framing -
# ends up sideways. Try 0/90/180/270 degree rotations in order and
# stop at the first one where PaddleOCR actually finds an expiry
# date; if none of the four find one, fall back to the 0-degree
# result so behaviour for genuinely-undetectable photos is unchanged.
# This costs extra OCR passes (up to 4x) only on images where the
# first pass found nothing - already-working images stay on the fast
# single-pass path below.
rotated_image_used = False
res_list = []
text_lines = []
@@ -482,15 +516,141 @@ async def classify_ocr(payload: ScanRequest):
expired_source_line = None
if ocr:
base_image = image
for step_angle in (0, 90, 180, 270):
try:
img_arr = np.array(image)
res_list = list(ocr.predict(img_arr))
if res_list and len(res_list) > 0:
res_entry = res_list[0]
text_lines = res_entry.get("rec_texts", [])
text_polys = ocr_text_polys(res_entry)
expired_date, expired_idx, expired_source_line = extract_expired_date(text_lines)
candidate_image = (
base_image.rotate(step_angle, resample=Image.BICUBIC, expand=True)
if step_angle else base_image
)
img_arr = np.array(candidate_image)
candidate_res_list = list(ocr.predict(img_arr))
candidate_res_entry = candidate_res_list[0] if candidate_res_list else {}
candidate_text_lines = candidate_res_entry.get("rec_texts", [])
candidate_text_polys = ocr_text_polys(candidate_res_entry)
candidate_expired_date, candidate_expired_idx, candidate_expired_source_line = (
extract_expired_date(candidate_text_lines)
)
if step_angle == 0:
# Always keep the 0-degree pass as the fallback result.
image, res_list, text_lines, text_polys = (
candidate_image, candidate_res_list, candidate_text_lines, candidate_text_polys
)
expired_date, expired_idx, expired_source_line = (
candidate_expired_date, candidate_expired_idx, candidate_expired_source_line
)
if candidate_expired_date is not None:
if step_angle != 0:
print(f"[Auto-Rotate-90] Expiry date found after rotating {step_angle} degrees.")
image, res_list, text_lines, text_polys = (
candidate_image, candidate_res_list, candidate_text_lines, candidate_text_polys
)
rotated_image_used = True
expired_date, expired_idx, expired_source_line = (
candidate_expired_date, candidate_expired_idx, candidate_expired_source_line
)
break
except Exception as rot_err:
print(f"Error during {step_angle}-degree OCR pass: {rot_err}")
traceback.print_exc()
# Tiled full-resolution pass: PaddleOCR downscales anything over
# its 4000px max_side_limit, which is exactly what kills small
# inkjet date stamps on these ~3200x5700 phone photos. Split the
# original image into overlapping tiles that each fit under the
# limit (so the date region is OCR'd at native resolution) and
# run the cascade per tile. Failure-path only, keyword-anchored
# acceptance like the VL fallback below.
if expired_date is None and max(base_image.size) > 2600:
TILE, OVERLAP = 2400, 400
W, H = base_image.size
step = TILE - OVERLAP
try:
found = False
for y0 in range(0, H, step):
if found:
break
for x0 in range(0, W, step):
tile = base_image.crop((x0, y0, min(x0 + TILE, W), min(y0 + TILE, H)))
if tile.width < 300 or tile.height < 300:
continue
tile_res = list(ocr.predict(np.array(tile)))
tile_lines = tile_res[0].get("rec_texts", []) if tile_res else []
if not tile_lines:
continue
t_date, _t_idx, t_source = extract_expired_date(tile_lines)
if t_date is not None and t_source and line_has_exp_keyword(
clean_date_line(t_source)
):
print(f"[Tile-Pass] Expiry date {t_date} found in full-res tile ({x0},{y0}) (line: {t_source!r})")
expired_date = t_date
expired_idx = None # tile polys don't map to the full image
expired_source_line = t_source
found = True
break
except Exception as tile_err:
print(f"[Tile-Pass] failed: {tile_err}")
traceback.print_exc()
# VL fallback: the lightweight PP-OCRv6 detector missed the date
# at every orientation. The vLLM-backed PaddleOCR-VL pipeline
# (:8090, same container) is a much stronger reader of small,
# low-contrast inkjet codes - ask it to read the whole package
# and run the same date cascade over its text output. Only fires
# on already-failed images, so the happy path stays single-pass.
# Acceptance is stricter than the local cascade: the matched
# line must carry an expiry keyword (BB/EXP/Baik digunakan...),
# so a bare number elsewhere on the package can't be
# hallucinated into a date on photos where none is visible.
# Even when no date is found, the VL's (much cleaner) text lines
# are kept and appended to text_lines below - they feed the
# gateway's OCR-evidence classification re-ranking.
vl_text_lines = []
if expired_date is None:
try:
vl_url = os.environ.get(
"VL_PIPELINE_URL", "http://localhost:8090/layout-parsing"
)
buffered = io.BytesIO()
base_image.save(buffered, format="JPEG")
vl_payload = {
"file": base64.b64encode(buffered.getvalue()).decode("utf-8"),
"matchHistoryJob": False,
"useLayoutDetection": True,
"fileType": 1,
"useDocUnwarping": False,
"useDocOrientationClassify": True,
}
vl_resp = requests.post(vl_url, json=vl_payload, timeout=120)
if vl_resp.status_code == 200:
vl_data = vl_resp.json()
if vl_data.get("errorCode") == 0:
layout_results = vl_data.get("result", {}).get("layoutParsingResults", [])
md_text = ""
if layout_results:
md_text = (layout_results[0].get("markdown") or {}).get("text", "") or ""
vl_lines = [ln.strip() for ln in md_text.splitlines() if ln.strip()]
vl_text_lines = vl_lines
if vl_lines:
vl_date, vl_idx, vl_source_line = extract_expired_date(vl_lines)
if vl_date is not None and vl_source_line and line_has_exp_keyword(
clean_date_line(vl_source_line)
):
print(f"[VL-Fallback] Expiry date {vl_date} found by VL pipeline (line: {vl_source_line!r})")
expired_date = vl_date
expired_idx = None # no OCR polys for VL text; skip crop
expired_source_line = vl_source_line
else:
print(f"[VL-Fallback] pipeline error: {vl_resp.status_code} {vl_resp.text[:200]}")
except Exception as vl_err:
print(f"[VL-Fallback] failed: {vl_err}")
traceback.print_exc()
# Fine tilt-straighten correction (<90 degrees), applied on top of
# whichever 90-degree orientation the search above landed on.
try:
if expired_idx is not None and expired_idx < len(text_polys):
poly = text_polys[expired_idx]
if len(poly) >= 2:
@@ -502,13 +662,12 @@ async def classify_ocr(payload: ScanRequest):
angle_rad = math.atan2(dy, dx)
angle_deg = math.degrees(angle_rad)
# Standardize tilt rotation
if abs(angle_deg) > 3.0:
print(f"[Auto-Rotate] Detected Expiry Date text line angle: {angle_deg:.2f} degrees. Rotating image...")
image = image.rotate(angle_deg, resample=Image.BICUBIC, expand=True)
rotated_image_used = True
except Exception as pre_ocr_err:
print(f"Error in pre-pass OCR: {pre_ocr_err}")
print(f"Error in fine tilt-straighten pass: {pre_ocr_err}")
traceback.print_exc()
# 1. Run DINOv2 Similarity Search or YOLO Classification
@@ -521,7 +680,7 @@ async def classify_ocr(payload: ScanRequest):
device = "cuda" if torch.cuda.is_available() else "cpu"
# Preprocess image
preprocessed = DINOV2_TRANSFORMS(image).unsqueeze(0).to(device)
preprocessed = DINOV2_TRANSFORMS(classification_image).unsqueeze(0).to(device)
# Extract query embedding
with torch.no_grad():
@@ -572,7 +731,7 @@ async def classify_ocr(payload: ScanRequest):
# Fallback to YOLO if DINOv2 was not run or failed
if not top1_name:
if yolo_model:
results = yolo_model(image)
results = yolo_model(classification_image)
probs = results[0].probs
top1_idx = probs.top1
top1_conf = float(probs.top1conf)
@@ -629,46 +788,18 @@ async def classify_ocr(payload: ScanRequest):
text_lines, expired_idx, expired_date, len(text_polys)
)
# Create visual OCR image with bounding boxes
vis_image_b64 = None
try:
vis_image = coord_image.copy()
from PIL import ImageDraw, ImageFont
draw = ImageDraw.Draw(vis_image)
try:
font = ImageFont.load_default()
except:
font = None
for idx, poly in enumerate(text_polys):
is_expired = (crop_idx is not None and idx == crop_idx)
pts = [(float(p[0]), float(p[1])) for p in poly]
if is_expired:
color = (245, 158, 11) # Amber
label = "EXP"
else:
color = (13, 148, 136) # Teal
label = "TEXT"
draw.polygon(pts, outline=color, width=3)
x0, y0 = pts[0]
label_w = 32 if label == "EXP" else 38
draw.rectangle([x0, y0 - 15, x0 + label_w, y0], fill=color)
if font:
draw.text((x0 + 4, y0 - 14), label, fill=(255, 255, 255), font=font)
else:
draw.text((x0 + 4, y0 - 14), label, fill=(255, 255, 255))
buffered = io.BytesIO()
vis_image.save(buffered, format="JPEG")
vis_image_b64 = "data:image/jpeg;base64," + base64.b64encode(buffered.getvalue()).decode("utf-8")
except Exception as draw_err:
print(f"Error drawing visual OCR: {draw_err}")
traceback.print_exc()
# Merge the VL pipeline's text lines (when its fallback ran) into
# the returned text_lines: the gateway's classification re-ranking
# feeds on them, and they're much cleaner than local OCR on hard
# photos. Appended after all poly-aligned work above, so rec_polys
# indexing is unaffected. Also retry SKU extraction over them -
# a VL-read 8-digit SKU enables the gateway's exact-match pin.
if vl_text_lines:
text_lines = list(text_lines) + vl_text_lines
if not sku:
sku = extract_sku(vl_text_lines)
if sku:
print(f"[VL-Fallback] SKU {sku} extracted from VL text lines.")
# Crop expired date OCR region for summary verification
expired_date_crop_b64 = None
@@ -679,43 +810,6 @@ async def classify_ocr(payload: ScanRequest):
print(f"Error cropping expired date image: {crop_err}")
traceback.print_exc()
# 3. Call Spotting API
spotting_image_b64 = None
try:
if rotated_image_used:
buffered = io.BytesIO()
image.save(buffered, format="JPEG")
img_b64_only = base64.b64encode(buffered.getvalue()).decode("utf-8")
else:
img_b64_only = payload.image_base64.split(",")[-1]
spotting_payload = {
"file": img_b64_only,
"matchHistoryJob": False,
"useLayoutDetection": False,
"fileType": 1,
"useDocUnwarping": False,
"useDocOrientationClassify": False,
"promptLabel": "spotting"
}
spotting_url = "http://localhost:8090/layout-parsing"
spotting_resp = requests.post(spotting_url, json=spotting_payload, timeout=60)
if spotting_resp.status_code == 200:
spotting_data = spotting_resp.json()
if spotting_data.get("errorCode") == 0:
layout_results = spotting_data.get("result", {}).get("layoutParsingResults", [])
if layout_results:
page0 = layout_results[0]
out_imgs = page0.get("outputImages", {})
spotting_img = out_imgs.get("spotting_res_img")
if spotting_img:
spotting_image_b64 = "data:image/jpeg;base64," + spotting_img
else:
print(f"Spotting API error: {spotting_resp.text}")
except Exception as spotting_err:
print(f"Error calling spotting API: {spotting_err}")
traceback.print_exc()
ocr_result = {
"text_lines": text_lines,
"extracted_product_name": product_name,
@@ -723,9 +817,7 @@ async def classify_ocr(payload: ScanRequest):
"extracted_expired_date": expired_date,
"expired_line_index": crop_idx,
"expired_source_line": expired_source_line,
"expired_date_crop_base64": expired_date_crop_b64,
"vis_image_base64": vis_image_b64,
"spotting_image_base64": spotting_image_b64
"expired_date_crop_base64": expired_date_crop_b64
}
else:
ocr_result = {
+1
View File
@@ -55,6 +55,7 @@ workflow and have no task numbers; see `git log` for real dates/history.
- **2.1 (verification pass)** Ran a full browser walkthrough of `/scan-pfm` (classification, top-5, OCR expiry extraction + crop, SKU-master matching, Visual/Spotting Grid, Raw Response — all confirmed working with real data). Found and fixed a real bug: "Save Ground Truth" was returning success but silently writing into the `pfm-web-app` container's ephemeral filesystem instead of the host, because `/sources` wasn't a bind-mounted path in root `docker-compose.yml`. Added `./backend/sources:/sources` to the `pfm-web-app` service, recovered an orphaned entry via `docker cp`, and re-verified the save now persists to `backend/sources/product_manual_labels.json` on the host (confirmed the DO-flow's `manual_labels.json` save was fixed by the same change too) — shipped 2026-07-08.
- **2.3** Ran the accuracy regression harness and discovered `sources/accuracy_report.md` was badly stale (claimed 75.04%; real current baseline is **95.10% overall, already at/above the 95% target** — added a staleness banner to that file). Root-caused every remaining mismatch by pulling raw OCR text from Postgres (`documents.layout_parsing_result`): the worst field, `plat` (67.6%), is almost entirely the license-plate region being classified as an image/seal by the layout model rather than OCR'd as text — not fixable in `parser.ts`. Found and fixed one genuine parser logic bug along the way: the "global pattern scanning fallback" could duplicate an already-correctly-extracted `noDO` value into a still-missing `noSO` field; fixed by excluding already-assigned values from that fallback's candidate pool (`pfm-web-app/src/utils/parser.ts`). Doesn't change the aggregate score (a wrong value and "Not Found" score the same) but stops a fabricated-looking wrong number from silently reaching the database. All 48 parser unit tests still pass — shipped 2026-07-08.
- **Ad-hoc** Built custom expiry-date-based auto-rotation algorithm in Python classifier server (`classify_ocr_server.py`). The algorithm calculates the slant angle of the Expiry Date / Batch text line bounding box, automatically rotates the image to make it horizontal, and re-runs YOLO classification + PaddleOCR for maximum accuracy. Enhanced SKU matching database lookup to prioritize exact SKU matches with a score of 1.0, pinning them as the Best Match — shipped 2026-07-09.
- **2.5** Retrained the Product/SKU scan classifier's model artifacts against the full current dataset, which had grown to 81 SKU classes / 2,493 photos (up from the original 16 classes / 118 photos the deployed model dated 2026-07-08 was actually trained on — the other 65 classes had photos but no trained weights). Rebuilt `models/dinov2_index.pkl` (now 2,493/2,493 photos indexed) and retrained the YOLO classifier 100 epochs on an RTX 2060 (real elapsed time 54m21s), publishing `models/produk-pfm-classifier-26n-100e-2026-07-14.pt`/`.onnx` at **85.8% top-1 / 94.4% top-5** validation accuracy across all 81 classes (up from 83.3%/90% on the old 16-class model). Along the way, fixed a real train/val split bug in `train_classifier.py`: `split_dataset()` previously shuffled and split individual image files, letting an augmented copy (`photo_aug_2.jpeg`) land in validation while its near-duplicate source stayed in training — inflating val accuracy with memorization instead of measuring generalization; now groups by source photo (stripping `_aug_N`) before shuffling and splitting 80/20. Verified via `docker compose up -d pipeline-api` + `docker logs`: "DINOv2 index loaded with 2493 reference images", "Using classifier weights: .../produk-pfm-classifier-26n-100e-2026-07-14.pt", "YOLO model loaded successfully" — the live service is confirmed serving the new 81-class model, not assumed from the newest-file-by-date fallback logic. Remaining gap toward the program's ±230-SKU target is dataset growth, not a pipeline limitation — shipped 2026-07-14.
## Backend — Postgres Data Layer
+42 -32
View File
@@ -532,38 +532,48 @@ retraining" section), and record real timing/accuracy rather than estimates.
top-5** across all 81 classes — already ahead of the old 16-class model's
83.3%/90%, but not a final number since the run never reached completion.
- **Session resumed**: after the pause above, Docker Desktop had actually
stopped between sessions — a first resume attempt failed instantly with a
daemon-connection error before any training ran. Restarted Docker Desktop,
confirmed `docker ps` responsive, confirmed the `pipeline-api` image and
`dinov2_index.pkl` from the earlier session were both still intact (no
rebuild/reindex needed), then relaunched `train_classifier.py train
--imgsz 224` from epoch 0 in a fresh one-off container, timed with `time`.
- **Training ran to completion this time: 100/100 epochs, real elapsed time
54m21.248s.** Final validation: **85.8% top-1 / 94.4% top-5** across all 81
classes. Published `models/produk-pfm-classifier-26n-100e-2026-07-14.pt`
(3.4MB) and exported `.onnx` (6.3MB, ONNX opset 20).
## 3. Verification
- Confirmed via `docker ps -a` that the training container exited cleanly on
`docker stop` (no hang, no orphaned process).
- Confirmed via `ls` on the host `models/` directory that **no new dated
`.pt`/`.onnx` was written** — `train_model()` only calls
`shutil.copy2(best_weights, output_path)` after `model.train()` returns, so
an interrupted run correctly leaves the previously-deployed
`produk-pfm-classifier-26n-100e-2026-07-08.pt`/`.onnx` untouched. The live
classifier is unaffected by this session.
- Confirmed `dinov2_index.pkl` **is** updated on the host (4.2MB, timestamped
2026-07-14 06:47) — this step ran to completion before training started and
is unaffected by the training container being stopped afterward.
- Did **not** run `docker compose restart pipeline-api`, since there is no new
classifier checkpoint to pick up yet and the main compose stack wasn't even
running this session (confirmed via `docker ps -a`: `pfm-web-app`,
`vllm-server`, `nginx`, `postgres` were all `Exited` from a prior session,
untouched by this work).
- Confirmed via `ls` on the host `models/` directory that the new dated
`produk-pfm-classifier-26n-100e-2026-07-14.pt`/`.onnx` files exist (dated
2026-07-14 09:05/09:06), alongside the untouched 2026-07-08 files.
- Confirmed in the training log's own ONNX export step that the model's
output shape is `(1, 81)` — i.e. genuinely 81 output classes, not a stale
16-class head.
- Ran `docker compose up -d pipeline-api` (the main compose stack wasn't
running this session — confirmed via `docker compose ps` returning empty —
so this was a fresh start, not a "restart"; it correctly pulled in the
`vllm-server` dependency too) and polled `docker logs
paddleocr-pipeline-api` until startup markers appeared. Confirmed lines:
- `DINOv2 index loaded with 2493 reference images.`
- `Using classifier weights: /app/pfm-web-app/public/produk-pfm/models/produk-pfm-classifier-26n-100e-2026-07-14.pt`
- `YOLO model loaded successfully.`
- `INFO: Application startup complete.`
This is real, observed runtime behavior — the live `pipeline-api` service is
now actually serving the new 81-class model and the full 2,493-image
DINOv2 index, not an assumption based on `latest_classifier_weights()`'s
glob-newest-by-date logic.
## 4. Status
**Paused 2026-07-14, by user request — not complete, not abandoned.**
Done: Docker Desktop started, `pipeline-api` image built (2m54s), DINOv2 index
rebuilt and persisted (2,493/2,493 images, all 81 classes). Not done: the YOLO
classifier training run, which was intentionally interrupted at epoch 43/100
and left no partial checkpoint (container used `--rm`, and Ultralytics' own
per-epoch checkpoints live in the container's `runs/classify/`, which was
never bind-mounted to the host). **Resuming means restarting training from
epoch 0**, not continuing from 43 — the image doesn't need rebuilding and the
index doesn't need reindexing, only `train_classifier.py train --imgsz 224`
needs to run again. Observed pace (32s/epoch) suggests a full 100-epoch run
takes **~55 minutes** on this host's RTX 2060, revised down from the ~90 min
estimated off the first few (slower, warmup) epochs. `plans/next-enhancements.md`
task 2.5 records the same state in full; `docs/scan-product.md`,
`backend/CLAUDE.md`, and `docs/feature-list.md` are deliberately left
unchanged (still say 16 classes) until a real completed run justifies updating
them.
**Done, 2026-07-14.** Both artifacts (DINOv2 index, YOLO classifier) retrained
against the full 81-class/2,493-photo dataset and verified loading in the live
service. `plans/next-enhancements.md` task 2.5 flipped to `[DONE]` with these
same numbers; `docs/scan-product.md` and `backend/CLAUDE.md`'s class-count/
accuracy claims updated from 16→81 classes and 83.3%/90%→85.8%/94.4%;
`docs/feature-list.md` given a matching entry. Remaining gap toward the
program's stated ±230-SKU target (see `proposals/sources/` kick-off material)
is dataset growth, not a code or training-pipeline limitation — the same
`train_classifier.py`/`index_dinov2.py` procedure documented here scales to
however many classes `foto-kemasan-v2/` ends up containing.
+16 -11
View File
@@ -119,9 +119,9 @@ pipeline call with `promptLabel: "spotting"`, no layout detection).
| File (`pfm-web-app/public/produk-pfm/`) | What |
|---|---|
| `foto-kemasan-v2/<SKU or class>/…` | Reference photo dataset — 16 classes, 118 photos (2–16 each) |
| `models/dinov2_index.pkl` | DINOv2 embeddings + metadata (rebuild after adding photos) |
| `models/produk-pfm-classifier-26n-100e-2026-07-08.pt` / `.onnx` | Fine-tuned YOLO classifier (83.3% top-1 / 90% top-5 val on the thin dataset) |
| `foto-kemasan-v2/<SKU or class>/…` | Reference photo dataset — 81 classes, 2,493 photos (target ~230 SKU) |
| `models/dinov2_index.pkl` | DINOv2 embeddings + metadata (rebuild after adding photos) — currently indexes all 2,493 photos across 81 classes |
| `models/produk-pfm-classifier-26n-100e-2026-07-14.pt` / `.onnx` | Fine-tuned YOLO classifier (85.8% top-1 / 94.4% top-5 val across all 81 classes; retrained 2026-07-14, 54m21s on an RTX 2060, up from the prior 2026-07-08 model's 83.3%/90% on only 16 classes) |
| `index_dinov2.py` | Rebuilds the pickle index from `foto-kemasan-v2/` |
| `train_classifier.py` | Splits 80/20 into `yolo_dataset/`, fine-tunes `yolo26n-cls.pt` (default 100 epochs, `--imgsz 224`), writes a dated checkpoint |
@@ -144,10 +144,16 @@ every labeled image in `sources/product_manual_labels.json`, checks 3 fields
(`no_sku`, `nama_item`, `expiry_date`) against ground truth, and splits into:
- **Training Set** — gallery photos under `foto-kemasan-v2/` (the classifier's
own reference images; scores here measure memorization, not generalization).
- **Validation Set** — flat filenames dropped into
`sources/product-test-images/` (a real held-out set; see that folder's
`README.md` for the drop-photo → label → re-run workflow via
`/manual-label-scan`).
- **Validation Set** — flat filenames, scored from the frozen
`sources/product-test-images-fixed/` snapshot (renamed `<index> <no_sku>.<ext>`,
built by `scripts/freeze-validation-set.mjs`) so a rerun always grades the
same 79 images regardless of what's since been dropped into the live-intake
`sources/product-test-images/` folder. See each folder's `README.md` — the
live folder documents the drop-photo → label → re-run-freeze-script workflow
via `/manual-label-scan`; the fixed folder documents the freeze/promote step
and flags 5 SKUs (12010801, 12012504, 12130504, 13050101, 15040102) whose
only available photo was already used to train the classifier, so their
scores aren't a clean held-out result.
Every run appends to `sources/product_accuracy_history.jsonl` and **auto-diffs
against the previous run**: the printed summary shows a Δ column per field per
@@ -185,10 +191,9 @@ Tracked ones (see `plans/next-enhancements.md`):
- **Dataset thinness**: 2–16 photos/class caps both classifiers; every new real
photo (especially non-studio, in-warehouse shots) matters. The harness above
already reports gallery (training) vs. held-out (validation) accuracy
separately — but as of this writing `sources/product-test-images/` is empty,
so the Validation Set is still 0 images and every published number so far is
a training/memorization score. Dropping real photos there is the next step,
not yet done.
separately, and as of 2026-07-14 the Validation Set has 79 labeled images
(74 genuinely held out, 5 flagged trained-on — see above) — the first real
(non-zero) Validation Set numbers.
Additional recommendations (not yet tasks — promote via `e`/`n` when wanted):
1. ~~Use `extracted_sku` in match ranking.~~ **Done** — `product-scan.ts`'s
+1
View File
@@ -4,6 +4,7 @@ const nextConfig: NextConfig = {
// Allow dev requests from any host — needed for tunnel access (ngrok, cloudflare, etc.)
// and direct LAN/WiFi IP access from Android devices.
allowedDevOrigins: [
"127.0.0.1",
"*.trycloudflare.com",
"*.ngrok.io",
"*.ngrok-free.app",
@@ -8,7 +8,7 @@ export const dynamic = "force-dynamic";
export async function GET(req: NextRequest) {
try {
const filename = req.nextUrl.searchParams.get("filename");
const dirPath = path.join(process.cwd(), "..", "sources", "product-test-images");
const dirPath = path.join(process.cwd(), "..", "sources", "product-test-images-fixed");
// File serving mode
if (filename) {
@@ -0,0 +1,86 @@
import { NextRequest, NextResponse } from "next/server";
import fs from "fs";
import path from "path";
import { errorResponse } from "@/utils/api-error";
// Serves the most recent accuracy-check-scan.mts detail dump
// (sources/product_scan_detail_*.json) so the manual-label-scan page can show
// what the AI actually predicted for a given Validation Set image by default,
// without re-running the pipeline live for every image browsed. This is the
// same predicted value the accuracy harness scores against ground truth -
// not a fresh scan, so it reflects the last batch test run.
const SOURCES_DIR = path.join(process.cwd(), "..", "sources");
interface DetailCheck {
field: string;
match: boolean;
expected: string;
predicted: string;
}
interface DetailValidationItem {
filename: string;
method?: string;
confidence?: number;
checks: DetailCheck[];
}
interface DetailDump {
timestamp: string;
validation: DetailValidationItem[];
}
function findLatestDump(): { path: string; data: DetailDump } | null {
if (!fs.existsSync(SOURCES_DIR)) return null;
const candidates = fs
.readdirSync(SOURCES_DIR)
.filter((f) => /^product_scan_detail_.*\.json$/.test(f))
.map((f) => {
const p = path.join(SOURCES_DIR, f);
return { path: p, mtime: fs.statSync(p).mtimeMs };
})
.sort((a, b) => b.mtime - a.mtime);
if (candidates.length === 0) return null;
const latest = candidates[0];
const data = JSON.parse(fs.readFileSync(latest.path, "utf8"));
return { path: latest.path, data };
}
export async function GET(req: NextRequest) {
try {
const { searchParams } = new URL(req.url);
const filename = searchParams.get("filename");
const latest = findLatestDump();
if (!latest) {
return NextResponse.json({ available: false });
}
if (!filename) {
return NextResponse.json({ available: true, timestamp: latest.data.timestamp });
}
const item = latest.data.validation.find((v) => v.filename === filename);
if (!item) {
return NextResponse.json({ available: true, timestamp: latest.data.timestamp, found: false });
}
const byField = Object.fromEntries(item.checks.map((c) => [c.field, c]));
return NextResponse.json({
available: true,
found: true,
timestamp: latest.data.timestamp,
method: item.method,
confidence: item.confidence,
no_sku: byField.no_sku?.predicted,
nama_item: byField.nama_item?.predicted,
expiry_date: byField.expiry_date?.predicted
});
} catch (err: unknown) {
console.error("Error in product-scan-results API:", err);
const message = err instanceof Error ? err.message : "Internal server error";
return errorResponse(500, message);
}
}
@@ -14,41 +14,10 @@ export async function POST(req: NextRequest) {
const result = await classifyAndMatchProduct(image_base64);
// Layout-parsing visualization (same pipeline as DO-PFM Visual Grid) - only
// used by this desktop test page, not part of the shared classify+match logic.
let layoutParsingResult: { layoutParsingResults?: Array<{ outputImages?: Record<string, string> }> } | null = null;
const rawB64 = image_base64.includes(",") ? image_base64.split(",")[1] : image_base64;
const pipelineUrl = process.env.PIPELINE_URL || "http://localhost:7871/layout-parsing";
try {
const layoutResponse = await fetch(pipelineUrl, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
file: rawB64,
matchHistoryJob: false,
useLayoutDetection: true,
fileType: 1,
useDocUnwarping: false,
useDocOrientationClassify: false
})
});
if (layoutResponse.ok) {
const layoutData = await layoutResponse.json();
layoutParsingResult = layoutData.result ?? layoutData;
} else {
console.warn("Layout parsing for visualization failed:", await layoutResponse.text());
}
} catch (layoutErr) {
console.warn("Layout parsing for visualization unavailable:", layoutErr);
}
return NextResponse.json({
classification: result.classification,
ocr: result.ocr,
possibleMatches: result.possibleMatches,
layoutParsingResult
possibleMatches: result.possibleMatches
});
} catch (error: unknown) {
@@ -18,6 +18,7 @@ export default function ManualLabelScanPage() {
notes: ""
});
const [aiPredicted, setAiPredicted] = useState<AiPredictedData | null>(null);
const [aiSource, setAiSource] = useState<{ type: "batch" | "live"; timestamp: string; method?: string; confidence?: number } | null>(null);
const [skuList, setSkuList] = useState<Array<{ no_sku: string; nama_item: string }>>([]);
const [isScanning, setIsScanning] = useState(false);
@@ -39,20 +40,11 @@ export default function ManualLabelScanPage() {
setSkuList(skuData.skus || []);
}
// Fetch Training Images
const pfmRes = await fetch("/api/produk-pfm");
let trainingFiles: { url: string; filename: string }[] = [];
if (pfmRes.ok) {
const pfmData = await pfmRes.json();
trainingFiles = (pfmData.products || []).flatMap((p: any) =>
p.images.map((url: string) => ({
url,
filename: url.replace(/^\/produk-pfm\/foto-kemasan-v2\//, "")
}))
);
}
// Fetch Test Images
// Fetch Test Images — the frozen 79-image Validation Set
// (product-test-images-fixed/), the only set the accuracy harness
// scores. Gallery/training photos (foto-kemasan-v2/) are not shown
// here: they don't need per-photo ground truth, only correct
// SKU-folder placement for classifier training.
const testRes = await fetch("/api/product-images");
let testFiles: { url: string; filename: string }[] = [];
if (testRes.ok) {
@@ -63,9 +55,8 @@ export default function ManualLabelScanPage() {
}));
}
const combined = [...testFiles, ...trainingFiles];
setFiles(combined);
if (combined.length > 0) setCurrentIndex(0);
setFiles(testFiles);
if (testFiles.length > 0) setCurrentIndex(0);
} catch (err) {
console.error("Error initializing page", err);
@@ -91,11 +82,33 @@ export default function ManualLabelScanPage() {
expiry_date: data.expiry_date || "",
notes: data.notes || ""
});
setAiPredicted(null); // Reset AI predictions on new file load
}
} catch (err) {
console.error("Error fetching label", err);
}
// Default-load the AI prediction from the last batch accuracy run
// (not a live re-scan) so failures are visible immediately while
// browsing - "Scan with AI" below can still be used to get a fresh
// live result for this exact image.
setAiPredicted(null);
setAiSource(null);
try {
const aiRes = await fetch(`/api/product-scan-results?filename=${encodeURIComponent(file.filename)}`);
if (aiRes.ok) {
const aiData = await aiRes.json();
if (aiData.found) {
setAiPredicted({
no_sku: aiData.no_sku,
nama_item: aiData.nama_item,
expiry_date: aiData.expiry_date
});
setAiSource({ type: "batch", timestamp: aiData.timestamp, method: aiData.method, confidence: aiData.confidence });
}
}
} catch (err) {
console.error("Error fetching batch AI result", err);
}
};
loadLabel();
}, [currentIndex, files]);
@@ -139,11 +152,25 @@ export default function ManualLabelScanPage() {
if (!scanRes.ok) throw new Error("Pipeline API error");
const scanData = await scanRes.json();
// Compare against the sku_master-resolved best match (what the app
// actually shows/saves as nama_item, and what the accuracy harness
// scores), not classification.top1_name - that's the classifier's raw
// internal class label (e.g. the foto-kemasan-v2 folder name), which
// structurally never matches a sku_master-style ground truth string
// even when the classification itself is correct.
const bestMatch = (scanData.possibleMatches || []).find((m: { isBestMatch?: boolean }) => m.isBestMatch);
setAiPredicted({
no_sku: scanData.classification?.top1_name ? skuList.find(s => s.nama_item === scanData.classification.top1_name)?.no_sku : undefined,
nama_item: scanData.classification?.top1_name,
no_sku: bestMatch?.no_sku,
nama_item: bestMatch?.nama_item,
expiry_date: scanData.ocr?.extracted_expired_date
});
setAiSource({
type: "live",
timestamp: new Date().toISOString(),
method: scanData.classification?.method,
confidence: scanData.classification?.top1_confidence
});
showToast("AI Scan complete!");
} catch (err) {
@@ -206,6 +233,7 @@ export default function ManualLabelScanPage() {
<Editor
formData={formData}
aiPredicted={aiPredicted}
aiSource={aiSource}
skuList={skuList}
isScanning={isScanning}
onScanWithAi={handleScanWithAi}
+2 -126
View File
@@ -1,7 +1,6 @@
"use client";
import React, { useState, useEffect } from "react";
import { extractLayoutVisUrlFromResult } from "@/utils/layoutVisualization";
import { getErrorMessage } from "@/utils/client-error";
export const dynamic = "force-dynamic";
@@ -34,16 +33,9 @@ interface ScanResponse {
expired_line_index?: number;
expired_date_crop_base64?: string;
expired_source_line?: string;
vis_image_base64?: string;
spotting_image_base64?: string;
error?: string;
};
possibleMatches: MatchResult[];
layoutParsingResult?: {
layoutParsingResults?: Array<{
outputImages?: Record<string, string>;
}>;
};
}
function levenshteinDistance(s1: string, s2: string): number {
@@ -115,7 +107,7 @@ export default function ScanPfmPage() {
const [savingGT, setSavingGT] = useState<boolean>(false);
const [error, setError] = useState<string>("");
const [scanResult, setScanResult] = useState<ScanResponse | null>(null);
const [activeTab, setActiveTab] = useState<"summary" | "visual" | "spotting" | "json">("summary");
const [activeTab, setActiveTab] = useState<"summary" | "json">("summary");
const [isEditingDate, setIsEditingDate] = useState<boolean>(false);
const [editedDate, setEditedDate] = useState<string>("");
@@ -385,8 +377,6 @@ export default function ScanPfmPage() {
}
};
const visUrl = scanResult ? extractLayoutVisUrlFromResult(scanResult.layoutParsingResult) : null;
return (
<div className="min-h-screen bg-slate-950 text-slate-100 flex flex-col font-sans">
@@ -637,14 +627,12 @@ export default function ScanPfmPage() {
<div className="border-b border-slate-800 flex rounded-xl overflow-hidden bg-slate-950/50">
{[
{ id: "summary", name: "Scan Summary" },
{ id: "visual", name: "Visual Grid" },
{ id: "spotting", name: "Spotting Grid" },
{ id: "json", name: "Raw Response" }
].map((t) => (
<button
key={t.id}
id={`tab-${t.id}`}
onClick={() => setActiveTab(t.id as "summary" | "visual" | "spotting" | "json")}
onClick={() => setActiveTab(t.id as "summary" | "json")}
className={`flex-1 py-3 text-xs font-bold transition-all border-b-2 cursor-pointer select-none ${
activeTab === t.id
? "border-teal-500 text-teal-400 bg-slate-900/40"
@@ -1059,118 +1047,6 @@ export default function ScanPfmPage() {
</div>
)}
{/* Visual Grid Tab */}
{activeTab === "visual" && (
<div className="space-y-4">
<h3 className="text-xs font-bold text-teal-400 uppercase tracking-wider">
Layout Visualization Grid
</h3>
{visUrl ? (
<div className="bg-slate-950 rounded-xl border border-slate-800 overflow-hidden shadow-2xl p-2 flex justify-center">
{/* eslint-disable-next-line @next/next/no-img-element */}
<img
src={visUrl}
alt="Layout Visualization Grid"
className="max-w-full h-auto object-contain rounded"
/>
</div>
) : (
<p className="text-xs text-slate-500 italic">
No layout visualization image returned. The layout-parsing pipeline may be unavailable.
</p>
)}
</div>
)}
{/* Spotting Grid Tab */}
{activeTab === "spotting" && (
<div className="space-y-6">
<div className="space-y-4">
<h3 className="text-xs font-bold text-teal-400 uppercase tracking-wider">
Spotting Visualization
</h3>
{scanResult.ocr?.spotting_image_base64 ? (
<div className="bg-slate-950 rounded-xl border border-slate-800 overflow-hidden shadow-2xl p-2 flex justify-center">
{/* eslint-disable-next-line @next/next/no-img-element */}
<img
src={scanResult.ocr.spotting_image_base64}
alt="Spotting Visualization BBoxes"
className="max-w-full h-auto object-contain rounded"
/>
</div>
) : (
<p className="text-xs text-slate-500 italic">No spotting visualization image returned by the parser.</p>
)}
</div>
<div className="space-y-3 pt-2 border-t border-slate-800">
<h3 className="text-xs font-bold text-amber-400 uppercase tracking-wider">
Detected Expired Date (Best Before / BB)
</h3>
<div className="grid grid-cols-1 sm:grid-cols-2 gap-4">
<div className="bg-slate-900/40 border border-slate-800 rounded-xl p-4 flex flex-col gap-2 min-h-[100px]">
<span className="text-[9px] font-bold text-slate-500 uppercase tracking-wider">Extracted Date</span>
{isEditingDate ? (
<div className="flex items-center gap-2 mt-1">
<input
type="text"
value={editedDate}
onChange={(e) => setEditedDate(e.target.value)}
className="bg-slate-950 border border-slate-700 text-slate-100 rounded-lg px-2 py-1 text-sm font-semibold focus:outline-none focus:ring-1 focus:ring-amber-500 w-full"
placeholder="DD/MM/YYYY"
autoFocus
onKeyDown={(e) => {
if (e.key === "Enter") handleSaveDate(editedDate);
else if (e.key === "Escape") setIsEditingDate(false);
}}
/>
<button onClick={() => handleSaveDate(editedDate)} className="bg-emerald-600 hover:bg-emerald-500 text-white rounded-lg p-1.5 text-xs font-bold transition-colors">✓</button>
<button onClick={() => setIsEditingDate(false)} className="bg-slate-700 hover:bg-slate-600 text-slate-300 rounded-lg p-1.5 text-xs font-bold transition-colors">✗</button>
</div>
) : (
<div className="flex items-center justify-between gap-2">
<span className={`text-lg font-bold flex items-center gap-2 ${scanResult.ocr?.extracted_expired_date ? "text-amber-400" : "text-slate-600"}`}>
<span className={`h-2 w-2 rounded-full flex-shrink-0 ${scanResult.ocr?.extracted_expired_date ? "bg-amber-400" : "bg-slate-700"}`} />
{scanResult.ocr?.extracted_expired_date || "Not detected"}
</span>
<button
onClick={() => { setEditedDate(scanResult.ocr?.extracted_expired_date || ""); setIsEditingDate(true); }}
className="text-[10px] bg-slate-800 hover:bg-slate-700 text-slate-400 px-2.5 py-1 rounded-md border border-slate-700 transition-all font-semibold flex items-center gap-1"
>
✏️ Edit
</button>
</div>
)}
{scanResult.ocr?.expired_source_line && (
<p className="text-[10px] text-slate-500 leading-snug">
<span className="font-semibold text-slate-400">OCR: </span>
{scanResult.ocr.expired_source_line}
</p>
)}
</div>
<div className="bg-slate-900/40 border border-slate-800 rounded-xl p-4 flex flex-col gap-2 min-h-[100px]">
<span className="text-[9px] font-bold text-slate-500 uppercase tracking-wider">Label Crop</span>
<div className="flex-1 bg-slate-950 border border-slate-800 rounded-lg p-2 flex items-center justify-center min-h-[72px] overflow-hidden">
{scanResult.ocr?.expired_date_crop_base64 ? (
// eslint-disable-next-line @next/next/no-img-element
<img
src={scanResult.ocr.expired_date_crop_base64}
alt="Expired date crop"
className="max-h-24 max-w-full object-contain brightness-95 contrast-105"
/>
) : (
<span className="text-[10px] text-slate-600 italic text-center px-2">
{scanResult.ocr?.extracted_expired_date ? "Crop unavailable" : "No expiry region to crop"}
</span>
)}
</div>
</div>
</div>
</div>
</div>
)}
{/* JSON Tab */}
{activeTab === "json" && (
<div className="space-y-4">
@@ -14,9 +14,17 @@ export interface AiPredictedData {
expiry_date?: string;
}
export interface AiSourceInfo {
type: "batch" | "live";
timestamp: string;
method?: string;
confidence?: number;
}
interface EditorProps {
formData: ScanLabelFormData;
aiPredicted: AiPredictedData | null;
aiSource: AiSourceInfo | null;
skuList: Array<{ no_sku: string; nama_item: string }>;
isScanning: boolean;
onScanWithAi: () => void;
@@ -28,6 +36,7 @@ interface EditorProps {
export function Editor({
formData,
aiPredicted,
aiSource,
skuList,
isScanning,
onScanWithAi,
@@ -69,8 +78,21 @@ export function Editor({
disabled={isScanning || !formData.filename}
className="w-full bg-teal-600/20 text-teal-400 hover:bg-teal-600/30 disabled:opacity-50 border border-teal-500/30 rounded-lg py-2 text-xs font-semibold transition flex items-center justify-center gap-2"
>
{isScanning ? "Scanning with Pipeline..." : "Scan with AI 🤖"}
{isScanning ? "Scanning with Pipeline..." : "Scan with AI 🤖 (re-run live)"}
</button>
{aiSource ? (
<p className="text-[10px] text-slate-500 leading-snug">
{aiSource.type === "batch" ? (
<>Showing result from last batch test ({new Date(aiSource.timestamp).toLocaleString()})</>
) : (
<>Live scan result ({new Date(aiSource.timestamp).toLocaleTimeString()})</>
)}
{aiSource.method && <> · {aiSource.method}</>}
{typeof aiSource.confidence === "number" && <> · conf {aiSource.confidence.toFixed(3)}</>}
</p>
) : (
<p className="text-[10px] text-slate-600 italic">No AI result yet for this image — click &quot;Scan with AI&quot; or run the accuracy batch test.</p>
)}
</div>
{/* Form Fields */}
@@ -1,26 +0,0 @@
export function normalizeImageSrc(src: string): string {
if (!src) return "";
if (src.startsWith("http://") || src.startsWith("https://") || src.startsWith("data:")) {
return src;
}
return `data:image/png;base64,${src}`;
}
export interface LayoutPageResult {
outputImages?: Record<string, string>;
}
/** Same visualization URL selection as DO-PFM Visual Grid (second image if present, else first). */
export function extractLayoutVisUrl(page0: LayoutPageResult | null | undefined): string {
const outImgs = page0?.outputImages || {};
const sortedUrls = Object.values(outImgs).filter(Boolean) as string[];
const visUrl = sortedUrls.length >= 2 ? sortedUrls[1] : sortedUrls[0] || "";
return normalizeImageSrc(visUrl);
}
export function extractLayoutVisUrlFromResult(
layoutParsingResult: { layoutParsingResults?: LayoutPageResult[] } | null | undefined
): string {
const page0 = layoutParsingResult?.layoutParsingResults?.[0];
return extractLayoutVisUrl(page0);
}
+129 -7
View File
@@ -1,10 +1,12 @@
import { query } from "../db";
// Bounds the classifier call so a wedged GPU container fails fast instead of
// hanging indefinitely - matches the bound `api/parse/route.ts` used to apply
// to its own separate inline classify call before it started sharing this
// function (see docs/api-contract-map.md G3).
const PIPELINE_TIMEOUT_MS = 90_000;
// hanging indefinitely. Raised from 90s (2026-07-14): hard images now
// legitimately take up to ~3 min - a 4-orientation OCR search plus a VL
// pipeline fallback when no expiry date is found (see
// config/classify_ocr_server.py) - and the old bound was killing exactly
// the images those fallbacks exist to save.
const PIPELINE_TIMEOUT_MS = 240_000;
// Thrown when the Python classifier service itself returns a non-2xx response,
// so callers can forward its actual status instead of collapsing everything to 500.
@@ -60,6 +62,115 @@ function getStringSimilarity(s1: string, s2: string): number {
return (maxLength - distance) / maxLength;
}
// --- OCR-evidence re-ranking of the classifier's top-K candidates ---
//
// DINOv2's misses are near-twin confusions (same brand line, different
// flavor/size) - exactly the cases where the printed variant words differ,
// and PaddleOCR usually reads some of them. Within a narrow similarity band
// of the top-1 candidate, prefer the one whose distinctive name tokens
// actually appear in the OCR'd text. Coverage-normalized so generic
// packaging words (e.g. "French Fries", "Ayam") that happen to be unique to
// one candidate's *name* can't hijack the ranking. Parameters tuned offline
// against the 79-image validation set (scripts/experiment-rerank.mjs,
// 2026-07-14: fixes 8 of 18 top-1 misses, breaks 0 of 61 correct).
const RERANK_TOP_K = 12;
const RERANK_SIM_BAND = 0.12;
const RERANK_COVERAGE_MARGIN = 0.25;
function classNameSku(className: string): string {
// foto-kemasan-v2 class names are "<SKU> <NAME...>"
return (className || "").trim().split(/\s+/)[0] || "";
}
function tokenizeName(name: string): string[] {
return name.toUpperCase().split(/[^A-Z0-9]+/).filter(t => t.length >= 2);
}
function withinEditDistance1(a: string, b: string): boolean {
if (a === b) return true;
const la = a.length, lb = b.length;
if (Math.abs(la - lb) > 1) return false;
if (la === lb) {
let diff = 0;
for (let i = 0; i < la; i++) if (a[i] !== b[i]) diff++;
return diff <= 1;
}
const [s, l] = la < lb ? [a, b] : [b, a];
let i = 0, j = 0, skipped = false;
while (i < s.length && j < l.length) {
if (s[i] === l[j]) { i++; j++; }
else if (!skipped) { skipped = true; j++; }
else return false;
}
return true;
}
interface OcrTextIndex { squashed: string; tokens: Set<string>; }
function buildOcrTextIndex(textLines: string[]): OcrTextIndex {
const joined = textLines.join(" ").toUpperCase();
return {
squashed: joined.replace(/[^A-Z0-9]/g, ""),
tokens: new Set(tokenizeName(joined))
};
}
function tokenFoundInOcr(token: string, ocr: OcrTextIndex): boolean {
if (token.length >= 4 && ocr.squashed.includes(token)) return true;
if (ocr.tokens.has(token)) return true;
if (token.length >= 5) {
for (const t of ocr.tokens) {
if (Math.abs(t.length - token.length) <= 1 && withinEditDistance1(token, t)) return true;
}
}
return false;
}
// Returns the class name of the best candidate after OCR-evidence
// re-ranking (the classifier's top-1 unless a close band-mate has clearly
// stronger printed-text evidence).
function rerankClassCandidates(
allProbabilities: Array<{ name: string; confidence: number }>,
textLines: string[]
): string {
if (!allProbabilities.length) return "";
const top1Sim = allProbabilities[0].confidence;
const band = allProbabilities
.slice(0, RERANK_TOP_K)
.filter(p => p.confidence >= top1Sim - RERANK_SIM_BAND);
if (band.length <= 1 || !textLines.length) return allProbabilities[0].name;
const ocrIdx = buildOcrTextIndex(textLines);
const cands = band.map(p => {
const sku = classNameSku(p.name);
return { name: p.name, tokens: new Set(tokenizeName(p.name.replace(sku, ""))), coverage: 0 };
});
const tokenCounts = new Map<string, number>();
for (const c of cands) {
for (const tok of c.tokens) tokenCounts.set(tok, (tokenCounts.get(tok) || 0) + 1);
}
for (const c of cands) {
let matched = 0, total = 0;
for (const tok of c.tokens) {
const nWith = tokenCounts.get(tok) || 1;
if (nWith >= cands.length) continue; // shared by all band-mates -> no signal
const w = 1 / nWith;
total += w;
if (tokenFoundInOcr(tok, ocrIdx)) matched += w;
}
c.coverage = total > 0 ? matched / total : 0;
}
let chosen = cands[0];
for (const c of cands.slice(1)) {
if (c.coverage >= chosen.coverage + RERANK_COVERAGE_MARGIN) chosen = c;
}
if (chosen !== cands[0]) {
console.log(`[Rerank] OCR evidence overrode classifier top-1 "${cands[0].name}" -> "${chosen.name}" (coverage ${cands[0].coverage.toFixed(2)} vs ${chosen.coverage.toFixed(2)})`);
}
return chosen.name;
}
// Shared by the classic /api/scan-pfm dev route and the authenticated
// /api/v1/scan-product route: calls the Python classifier, then matches the
// result against sku_master, returning the top-5 candidates.
@@ -86,17 +197,28 @@ export async function classifyAndMatchProduct(imageBase64: string): Promise<Prod
nama_item: row.nama_item
}));
const top1Name = data.classification?.top1_name || "";
const extractedSku = data.ocr?.extracted_sku || "";
// Re-rank the classifier's close candidates using OCR'd package text, then
// map the winner straight to its sku_master row by the SKU prefix embedded
// in the class name. The old approach (Levenshtein between top-1 class name
// and every master nama_item) lost classifier-correct results whenever a
// *different* SKU's master name happened to be textually closer.
const rerankedName = rerankClassCandidates(
data.classification?.all_probabilities || [],
data.ocr?.text_lines || []
) || data.classification?.top1_name || "";
const rerankedSku = classNameSku(rerankedName);
const matchedList: SkuMatch[] = skuMasterList.map(sku => {
const yoloSim = top1Name ? getStringSimilarity(sku.nama_item, top1Name) : 0;
const yoloSim = rerankedName ? getStringSimilarity(sku.nama_item, rerankedName) : 0;
const cleanMasterSku = sku.no_sku.trim();
const cleanExtractedSku = extractedSku.trim();
const isSkuMatch = cleanExtractedSku && cleanMasterSku === cleanExtractedSku;
const isClassifierPick = rerankedSku && cleanMasterSku === rerankedSku;
const score = isSkuMatch ? 1.0 : yoloSim;
const score = isSkuMatch ? 1.0 : isClassifierPick ? 0.995 : yoloSim;
return {
no_sku: sku.no_sku,
+27 -33
View File
@@ -113,11 +113,11 @@ in either project (only SKU, product name, expiry date are extracted) — if
requested later, follow the same OCR-regex-cascade pattern already used for
expiry-date extraction.*
- **2.5** [IN PROGRESS 2026-07-14 — resumed, training run 2] **Retrain
classifier on the now-81-class dataset.** (Note: a first resume attempt
- **2.5** [DONE 2026-07-14] **Retrain classifier on the now-81-class
dataset.** (Note: a first resume attempt
failed instantly with a Docker daemon connection error — Docker Desktop had
stopped between sessions — before any training happened; restarted Docker
Desktop and relaunched. This is the actual second training attempt,
Desktop and relaunched. The successful run was the second attempt,
confirmed running via `docker ps`.) `foto-kemasan-v2/` grew from the 16
classes/118 photos the deployed model
(`produk-pfm-classifier-26n-100e-2026-07-08.pt`) was trained on to **81
@@ -138,36 +138,30 @@ expiry-date extraction.*
`docker compose restart pipeline-api` → verify via `docker logs` for
"DINOv2 index loaded with N reference images" and "Using classifier
weights: <new dated file>".
- **Status as of pause (2026-07-14)** — mixed state, read carefully before
resuming:
- ✅ `pipeline-api` image built (2m54s), bakes in the current 81-class
dataset.
- ✅ **`dinov2_index.pkl` already rebuilt and persisted to disk** —
"Success! Indexed 2493/2493 images" across all 81 classes. This artifact
is live on the host now (`models/dinov2_index.pkl`, 4.2MB, dated
2026-07-14) and does **not** need to be redone.
- ⏸️ **YOLO classifier training was started, then stopped by user request
at epoch 43/100 (~23 minutes in)** before it could write a new dated
checkpoint. `docker run` used `--rm` and the in-progress epoch
checkpoints live only in the container's own `runs/classify/` (not
bind-mounted), so **stopping the container discarded that partial
progress** — resuming means restarting from epoch 0, not continuing from
43. `models/` on the host still has only the original
`produk-pfm-classifier-26n-100e-2026-07-08.pt`/`.onnx` (16-class model) —
**the live/deployed classifier is unchanged**, still 16 classes.
- Observed pace before stopping: ~32s/epoch (43 epochs in 23m1s) → a full
100-epoch run should take **~55 minutes** on this host's RTX 2060 (6GB
VRAM), not the ~90 min extrapolated from the first few (slower, warmup)
epochs. At epoch 42 the in-progress run had already reached 84.3%
top-1 / 93.9% top-5 val accuracy across all 81 classes, ahead of the old
16-class model's 83.3%/90% — a promising sign for the eventual full run,
but not a final result since training didn't finish.
- **To resume**: image is already built and the DINOv2 index step can be
skipped — just re-run the one-off `train_classifier.py train --imgsz 224`
container, then `docker compose restart pipeline-api` and verify via
`docker logs`. Update the class count in `docs/scan-product.md`,
`CLAUDE.md`, and `docs/feature-list.md` (and flip this task to `[DONE]`)
only once that run actually completes with a final dated `.pt`/`.onnx`.
- **Final result (2026-07-14)** — both artifacts retrained and live:
- ✅ `dinov2_index.pkl` rebuilt and persisted to disk — "Success! Indexed
2493/2493 images" across all 81 classes (`models/dinov2_index.pkl`,
4.2MB, dated 2026-07-14).
- ✅ **YOLO classifier retrained to completion, 100/100 epochs, real
elapsed time 54m21s** (a first attempt was intentionally stopped by user
request at epoch 43/100 to pause the session; that partial progress was
discarded since `docker run --rm`'s in-container `runs/classify/`
checkpoints aren't bind-mounted, so the successful run below restarted
cleanly from epoch 0 rather than resuming from 43). Final validation:
**85.8% top-1 / 94.4% top-5** across all 81 classes — up from the old
16-class model's 83.3%/90%, now covering 5x the product classes.
Published artifacts: `models/produk-pfm-classifier-26n-100e-2026-07-14.pt`
(3.4MB) and matching `.onnx` (6.3MB, ONNX opset 20, output shape
confirmed `(1, 81)` — i.e. 81 output classes).
- ✅ Verified via `docker compose up -d pipeline-api` (main stack wasn't
running this session) + `docker logs paddleocr-pipeline-api`: "DINOv2
index loaded with 2493 reference images", "Using classifier weights:
/app/pfm-web-app/public/produk-pfm/models/produk-pfm-classifier-26n-100e-2026-07-14.pt",
"YOLO model loaded successfully", "Application startup complete" — the
live service is now serving the new 81-class model, not a code-review
assumption.
- Class-count claims updated in `docs/scan-product.md` and `../CLAUDE.md`
(16 → 81 classes); this task's `docs/feature-list.md` entry added.
## 3. Backend — Postgres Data Layer
`pfm-web-app/src/db/`
+56 -10
View File
@@ -4,8 +4,9 @@
// backend/sources/product_manual_labels.json, checks 3 fields (no_sku,
// nama_item, expiry_date) against ground truth, splits results into a
// Training Set (gallery photos under foto-kemasan-v2/ that trained the
// classifier itself) vs a Validation Set (flat filenames dropped in
// backend/sources/product-test-images/), and appends a summary to
// classifier itself) vs a Validation Set (flat filenames, scored from the
// frozen backend/sources/product-test-images-fixed/ snapshot so reruns always
// grade the exact same images), and appends a summary to
// backend/sources/product_accuracy_history.jsonl. Every run auto-diffs
// against the last history entry and flags field/image regressions or
// improvements, so a tuning change to classify_ocr_server.py shows its
@@ -30,7 +31,10 @@ const SOURCES_DIR = path.join(__dirname, "..", "sources");
const LABELS_PATH = process.env.ACCURACY_LABELS_PATH || path.join(SOURCES_DIR, "product_manual_labels.json");
const HISTORY_PATH = process.env.ACCURACY_HISTORY_PATH || path.join(SOURCES_DIR, "product_accuracy_history.jsonl");
const FETCH_TIMEOUT_MS = 120_000;
// Hard images legitimately take up to ~3 min now (4-orientation OCR search +
// VL pipeline fallback for missing expiry dates); must exceed the gateway's
// own PIPELINE_TIMEOUT_MS (240s) so slow scans fail there, not here.
const FETCH_TIMEOUT_MS = 300_000;
const FIELDS = ["no_sku", "nama_item", "expiry_date"] as const;
type Field = typeof FIELDS[number];
type Split = "training" | "validation";
@@ -61,6 +65,7 @@ interface ScanResponse {
interface Check {
field: Field;
match: boolean;
predicted: string;
}
interface ResultItem {
@@ -70,6 +75,12 @@ interface ResultItem {
confidence?: number;
}
// Optional: set ACCURACY_DETAIL_DUMP_PATH to write full per-image,
// per-field ground-truth-vs-predicted detail (plus failures) as JSON —
// used to triage which images to pull into an "undetected" folder for
// visual inspection instead of just the aggregate percentages.
const DETAIL_DUMP_PATH = process.env.ACCURACY_DETAIL_DUMP_PATH;
interface ClassificationStats {
methodCounts: Record<string, number>;
avgConfidence: number;
@@ -104,7 +115,12 @@ function getImagePath(filename: string): string {
if (filename.includes("/")) {
return path.join(APP_ROOT, "public", "produk-pfm", "foto-kemasan-v2", filename);
}
return path.join(SOURCES_DIR, "product-test-images", filename);
// Validation Set images are scored from the frozen, sequentially-renamed
// copy in product-test-images-fixed/ (built by scripts/freeze-validation-set.mjs)
// rather than the live-intake product-test-images/ folder, so a rerun always
// scores the exact same image set regardless of what's since been dropped
// into the live folder for future curation.
return path.join(SOURCES_DIR, "product-test-images-fixed", filename);
}
async function checkServerReachable(baseUrl: string) {
@@ -338,6 +354,7 @@ async function main() {
validation: [] as ResultItem[],
failed: [] as string[]
};
const failedDetail: Array<{ filename: string; error: string }> = [];
const perImagePct: Record<string, number> = {};
for (const gt of labels) {
@@ -353,14 +370,21 @@ async function main() {
const b64 = "data:image/jpeg;base64," + fs.readFileSync(imgPath, "base64");
const parsed = await fetchScan(args.baseUrl, b64);
const bestMatchSku = parsed.possibleMatches?.find(m => m.isBestMatch)?.no_sku || "";
const predictedItemName = parsed.classification?.top1_name || "";
const bestMatch = parsed.possibleMatches?.find(m => m.isBestMatch);
const bestMatchSku = bestMatch?.no_sku || "";
// Compare against the sku_master-resolved name (what the app actually
// shows/saves as nama_item), not classification.top1_name — that's the
// classifier's raw internal class label (e.g. "11110059 CEKER BERKUKU
// FROZEN PACK 1 KG", literally the foto-kemasan-v2 folder name), which
// structurally never matches a sku_master-style ground truth string
// even when the classification itself is correct.
const predictedItemName = bestMatch?.nama_item || "";
const predictedExpiry = parsed.ocr?.extracted_expired_date || "";
const checks: Check[] = [
{ field: "no_sku", match: isMatch(gt.no_sku, bestMatchSku) },
{ field: "nama_item", match: isMatch(gt.nama_item, predictedItemName) },
{ field: "expiry_date", match: isMatch(gt.expiry_date, predictedExpiry) }
{ field: "no_sku", match: isMatch(gt.no_sku, bestMatchSku), predicted: bestMatchSku },
{ field: "nama_item", match: isMatch(gt.nama_item, predictedItemName), predicted: predictedItemName },
{ field: "expiry_date", match: isMatch(gt.expiry_date, predictedExpiry), predicted: predictedExpiry }
];
const item: ResultItem = {
@@ -380,8 +404,10 @@ async function main() {
perImagePct[gt.filename] = (score / checks.length) * 100;
console.log(`done (${score}/3)`);
} catch (err) {
console.log(`FAILED (${(err as Error).message})`);
const message = (err as Error).message;
console.log(`FAILED (${message})`);
results.failed.push(gt.filename);
failedDetail.push({ filename: gt.filename, error: message });
}
}
@@ -403,6 +429,26 @@ async function main() {
perImage: perImagePct
};
fs.appendFileSync(HISTORY_PATH, JSON.stringify(entry) + "\n");
if (DETAIL_DUMP_PATH) {
const detail = {
timestamp: entry.timestamp,
validation: results.validation.map(r => ({
filename: r.gt.filename,
method: r.method,
confidence: r.confidence,
checks: r.checks.map(c => ({
field: c.field,
match: c.match,
expected: r.gt[c.field],
predicted: c.predicted
}))
})),
failed: failedDetail
};
fs.writeFileSync(DETAIL_DUMP_PATH, JSON.stringify(detail, null, 2), "utf8");
console.log(`\nDetail dump written to ${DETAIL_DUMP_PATH}`);
}
}
main().catch(err => {
+76
View File
@@ -0,0 +1,76 @@
// One-off script: pull every Validation Set image that had at least one
// mismatched field (or failed to scan entirely) out of
// sources/product-test-images-fixed/ into sources/product-test-images-undetected/,
// renamed to show which field(s) missed, plus a manifest.md with
// expected-vs-predicted per field — so a human can tell at a glance whether
// a miss is an OCR problem, a classification problem, or the photo itself
// lacking the data (e.g. expiry code out of frame).
// Usage: node scripts/build-undetected-set.mjs <path-to-detail-dump.json>
import fs from "fs";
import path from "path";
const detailPath = process.argv[2];
if (!detailPath) {
console.error("Usage: node scripts/build-undetected-set.mjs <detail-dump.json>");
process.exit(1);
}
const FIXED_DIR = path.join("sources", "product-test-images-fixed");
const OUT_DIR = path.join("sources", "product-test-images-undetected");
const detail = JSON.parse(fs.readFileSync(detailPath, "utf8"));
if (fs.existsSync(OUT_DIR)) fs.rmSync(OUT_DIR, { recursive: true });
fs.mkdirSync(OUT_DIR, { recursive: true });
const manifestRows = [];
let copied = 0;
for (const item of detail.validation) {
const failedFields = item.checks.filter((c) => !c.match);
if (!failedFields.length) continue;
const ext = path.extname(item.filename);
const base = path.basename(item.filename, ext);
const tag = failedFields.map((c) => c.field).join(",");
const destName = `${base} [FAIL ${tag}]${ext}`;
fs.copyFileSync(path.join(FIXED_DIR, item.filename), path.join(OUT_DIR, destName));
copied++;
for (const c of item.checks) {
manifestRows.push({
image: destName,
field: c.field,
match: c.match,
expected: c.expected || "(empty)",
predicted: c.predicted || "(empty — not detected)"
});
}
}
for (const f of detail.failed) {
const srcPath = path.join(FIXED_DIR, f.filename);
if (!fs.existsSync(srcPath)) continue;
const ext = path.extname(f.filename);
const base = path.basename(f.filename, ext);
const destName = `${base} [FAILED-scan]${ext}`;
fs.copyFileSync(srcPath, path.join(OUT_DIR, destName));
copied++;
manifestRows.push({ image: destName, field: "(entire scan)", match: false, expected: "-", predicted: `ERROR: ${f.error}` });
}
const lines = [
"# Undetected / mismatched images",
"",
`Generated ${detail.timestamp} from the accuracy-check-scan.mts detail dump.`,
`${copied} of ${detail.validation.length + detail.failed.length} Validation Set images had at least one wrong/missing field or failed to scan.`,
"",
"| Image | Field | OK? | Expected | Predicted |",
"|---|---|---|---|---|"
];
for (const r of manifestRows) {
lines.push(`| ${r.image} | ${r.field} | ${r.match ? "✓" : "✗"} | ${r.expected} | ${r.predicted} |`);
}
fs.writeFileSync(path.join(OUT_DIR, "manifest.md"), lines.join("\n") + "\n", "utf8");
console.log(`Copied ${copied} images into ${OUT_DIR}, wrote manifest.md (${manifestRows.length} rows).`);
@@ -0,0 +1,59 @@
// One-off capture: hit /api/scan-pfm for every image in
// sources/product-test-images-fixed/ and save the full text-level response
// (classification all_probabilities, OCR text_lines, extracted fields,
// possibleMatches) minus base64 image blobs to
// sources/product_scan_fullcap.json. This lets classification re-ranking
// experiments run offline against ground truth in seconds instead of
// re-running the 11-minute GPU batch per iteration.
// Usage: node scripts/capture-scan-responses.mjs [baseUrl]
import fs from "fs";
import path from "path";
const BASE_URL = process.argv[2] || process.env.ACCURACY_BASE_URL || "http://127.0.0.1:3000";
const FIXED_DIR = path.join("sources", "product-test-images-fixed");
const OUT_PATH = path.join("sources", "product_scan_fullcap.json");
const files = fs.readdirSync(FIXED_DIR)
.filter((f) => /\.(jpe?g|png|webp)$/i.test(f))
.sort((a, b) => parseInt(a) - parseInt(b));
const results = [];
for (const filename of files) {
const b64 = "data:image/jpeg;base64," +
fs.readFileSync(path.join(FIXED_DIR, filename), "base64");
process.stdout.write(`Capturing ${filename}... `);
try {
const res = await fetch(`${BASE_URL}/api/scan-pfm`, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ image_base64: b64 }),
signal: AbortSignal.timeout(120_000)
});
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const d = await res.json();
results.push({
filename,
classification: {
top1_name: d.classification?.top1_name,
top1_confidence: d.classification?.top1_confidence,
method: d.classification?.method,
all_probabilities: (d.classification?.all_probabilities || []).slice(0, 15)
},
ocr: {
text_lines: d.ocr?.text_lines || [],
extracted_sku: d.ocr?.extracted_sku ?? null,
extracted_product_name: d.ocr?.extracted_product_name ?? null,
extracted_expired_date: d.ocr?.extracted_expired_date ?? null,
expired_source_line: d.ocr?.expired_source_line ?? null
},
possibleMatches: d.possibleMatches || []
});
console.log("ok");
} catch (err) {
console.log(`FAILED (${err.message})`);
results.push({ filename, error: err.message });
}
}
fs.writeFileSync(OUT_PATH, JSON.stringify(results, null, 2), "utf8");
console.log(`\nWrote ${results.length} captures to ${OUT_PATH}`);
+178
View File
@@ -0,0 +1,178 @@
// Offline experiment: OCR-evidence re-ranking of DINOv2 top-K candidates.
//
// Reads sources/product_scan_fullcap.json (captured live responses, see
// capture-scan-responses.mjs) + sources/product_manual_labels.json (ground
// truth) and simulates candidate re-ranking without touching the GPU stack,
// reporting fixed-vs-broken counts per parameter combination. The winning
// parameters get ported into pfm-web-app/src/utils/product-scan.ts.
//
// Idea: DINOv2's near-twin confusions (same brand, different flavor/size)
// are exactly the cases where the *printed variant words* differ - and
// PaddleOCR usually reads some of them. So within a narrow similarity band
// of the top-1, prefer the candidate whose distinguishing name tokens
// actually appear in the OCR'd text.
//
// Usage: node scripts/experiment-rerank.mjs
import fs from "fs";
import path from "path";
const cap = JSON.parse(fs.readFileSync(path.join("sources", "product_scan_fullcap.json"), "utf8"));
const labels = JSON.parse(fs.readFileSync(path.join("sources", "product_manual_labels.json"), "utf8"));
const gtBySku = new Map(labels.map((l) => [l.filename, l.no_sku]));
function classSku(className) {
// Class names are foto-kemasan-v2 folder names: "<SKU> <NAME...>"
return (className || "").trim().split(/\s+/)[0] || "";
}
function tokenize(name) {
return name
.toUpperCase()
.split(/[^A-Z0-9]+/)
.filter((t) => t.length >= 2);
}
function editDistance1(a, b) {
// true if edit distance <= 1 (same length: 1 substitution; off-by-one: 1 indel)
if (a === b) return true;
const la = a.length, lb = b.length;
if (Math.abs(la - lb) > 1) return false;
if (la === lb) {
let diff = 0;
for (let i = 0; i < la; i++) if (a[i] !== b[i]) diff++;
return diff <= 1;
}
const [s, l] = la < lb ? [a, b] : [b, a];
let i = 0, j = 0, skipped = false;
while (i < s.length && j < l.length) {
if (s[i] === l[j]) { i++; j++; }
else if (!skipped) { skipped = true; j++; }
else return false;
}
return true;
}
function buildOcrIndex(textLines) {
const joined = textLines.join(" ").toUpperCase();
const squashed = joined.replace(/[^A-Z0-9]/g, "");
const tokens = new Set(tokenize(joined));
return { squashed, tokens };
}
function tokenInOcr(token, ocrIdx, fuzzy) {
if (token.length >= 4 && ocrIdx.squashed.includes(token)) return true;
if (ocrIdx.tokens.has(token)) return true;
if (fuzzy && token.length >= 5) {
for (const t of ocrIdx.tokens) {
if (Math.abs(t.length - token.length) <= 1 && editDistance1(token, t)) return true;
}
}
return false;
}
function ocrEvidenceScore(candTokens, bandTokenCounts, bandSize, ocrIdx, fuzzy) {
// Coverage-normalized, rarity-weighted evidence: fraction of this
// candidate's *distinctive* name tokens (weighted by band rarity) that
// actually appear in the OCR'd text. Normalizing by the candidate's own
// distinctive-token mass is what stops generic packaging words from
// hijacking the ranking - a candidate whose name promises FRENCH +
// INSTITUSI + 2KG but whose package shows only "French Fries" scores
// 1/3, losing to a candidate whose 2 distinctive tokens both appear.
let matched = 0;
let total = 0;
for (const tok of new Set(candTokens)) {
const nWith = bandTokenCounts.get(tok) || 1;
if (nWith >= bandSize) continue; // shared by all -> no signal
const w = 1 / nWith;
total += w;
if (tokenInOcr(tok, ocrIdx, fuzzy)) matched += w;
}
return total > 0 ? matched / total : 0;
}
function skuFuzzyBoost(extractedSku, candidateSku) {
if (!extractedSku || extractedSku.length < 7) return 0;
if (extractedSku === candidateSku) return 10; // exact (normally pinned upstream anyway)
return editDistance1(extractedSku, candidateSku) ? 1 : 0;
}
function runConfig({ K, BAND, MARGIN, FUZZY, SKU_BOOST_W }) {
let baselineCorrect = 0, rerankCorrect = 0, fixed = [], broken = [];
for (const item of cap) {
if (item.error) continue;
const gt = gtBySku.get(item.filename);
if (!gt) continue;
const probs = item.classification?.all_probabilities || [];
if (!probs.length) continue;
const top1Sku = classSku(probs[0].name);
const baselineRight = top1Sku === gt;
if (baselineRight) baselineCorrect++;
// Candidate band: within BAND of top-1 similarity, capped at K
const top1Sim = probs[0].confidence;
const band = probs.slice(0, K).filter((p) => p.confidence >= top1Sim - BAND);
const ocrIdx = buildOcrIndex(item.ocr?.text_lines || []);
const candInfos = band.map((p) => {
const sku = classSku(p.name);
const tokens = tokenize(p.name.replace(sku, ""));
return { sku, sim: p.confidence, tokens };
});
const bandTokenCounts = new Map();
for (const c of candInfos) {
for (const tok of new Set(c.tokens)) {
bandTokenCounts.set(tok, (bandTokenCounts.get(tok) || 0) + 1);
}
}
for (const c of candInfos) {
c.ocrScore = ocrEvidenceScore(c.tokens, bandTokenCounts, candInfos.length, ocrIdx, FUZZY)
+ SKU_BOOST_W * skuFuzzyBoost(item.ocr?.extracted_sku || "", c.sku);
}
// Switch away from top-1 only when a band-mate has clearly stronger OCR evidence
let chosen = candInfos[0];
for (const c of candInfos.slice(1)) {
if (c.ocrScore >= chosen.ocrScore + MARGIN) chosen = c;
}
const rerankRight = chosen.sku === gt;
if (rerankRight) rerankCorrect++;
if (!baselineRight && rerankRight) fixed.push(item.filename);
if (baselineRight && !rerankRight) broken.push(item.filename);
}
return { baselineCorrect, rerankCorrect, fixed, broken };
}
const grid = [];
for (const K of [5, 8, 12]) {
for (const BAND of [0.04, 0.06, 0.08, 0.12]) {
// Coverage scores live in [0, 1]; margin is the minimum coverage lead a
// band-mate needs over the current pick before we switch away from it.
for (const MARGIN of [0.15, 0.25, 0.35, 0.5]) {
for (const FUZZY of [true, false]) {
for (const SKU_BOOST_W of [0, 2]) {
grid.push({ K, BAND, MARGIN, FUZZY, SKU_BOOST_W });
}
}
}
}
}
const results = grid.map((cfg) => ({ cfg, ...runConfig(cfg) }));
results.sort((a, b) => (b.rerankCorrect - b.broken.length * 0.01) - (a.rerankCorrect - a.broken.length * 0.01));
console.log(`Images evaluated: ${cap.filter((i) => !i.error && gtBySku.has(i.filename)).length}`);
console.log(`Baseline (DINOv2 top-1) correct: ${results[0].baselineCorrect}\n`);
console.log("Top 12 configs by re-ranked correct count:");
for (const r of results.slice(0, 12)) {
console.log(
` correct=${r.rerankCorrect} (+${r.fixed.length}/-${r.broken.length}) ` +
`K=${r.cfg.K} BAND=${r.cfg.BAND} MARGIN=${r.cfg.MARGIN} FUZZY=${r.cfg.FUZZY} SKUW=${r.cfg.SKU_BOOST_W}`
);
}
const best = results[0];
console.log(`\nBest config detail: ${JSON.stringify(best.cfg)}`);
console.log(` fixed (${best.fixed.length}): ${best.fixed.join(", ")}`);
console.log(` broken (${best.broken.length}): ${best.broken.join(", ")}`);
+47
View File
@@ -0,0 +1,47 @@
// One-off script: freeze the current 74-image product-scan Validation Set into
// a dedicated, stable folder (sources/product-test-images-fixed/) so re-running
// the accuracy harness always scores the exact same images, independent of
// whatever new photos get dropped into the live-intake folder
// (sources/product-test-images/, still fed by the /manual-label-scan page).
// Renames each image "<index> <no_sku>.<ext>" (index = its stable position,
// 1-based) and updates product_manual_labels.json's flat-filename entries to
// match. Run once from backend/: node scripts/freeze-validation-set.mjs
import fs from "fs";
import path from "path";
const LIVE_DIR = path.join("sources", "product-test-images");
const FIXED_DIR = path.join("sources", "product-test-images-fixed");
const LABELS_PATH = path.join("sources", "product_manual_labels.json");
const labels = JSON.parse(fs.readFileSync(LABELS_PATH, "utf8"));
const flatEntries = labels.filter((l) => !l.filename.includes("/"));
if (!fs.existsSync(FIXED_DIR)) fs.mkdirSync(FIXED_DIR, { recursive: true });
// Resolve by SKU prefix (not entry.filename directly) so this script is
// idempotent/rerunnable even after a previous run already renamed
// entry.filename to "<index> <sku>.<ext>" — the live-intake folder always
// keeps its original "<sku> <product>__<camera-filename>.<ext>" names.
const liveFiles = fs.readdirSync(LIVE_DIR);
function findSourceFile(no_sku) {
const match = liveFiles.find((f) => f.startsWith(`${no_sku} `) || f.startsWith(`${no_sku}__`));
if (!match) return null;
return path.join(LIVE_DIR, match);
}
let copied = 0;
flatEntries.forEach((entry, i) => {
const index = i + 1;
const srcPath = findSourceFile(entry.no_sku);
if (!srcPath) {
throw new Error(`Missing source image for ${entry.no_sku} in ${LIVE_DIR}`);
}
const ext = path.extname(srcPath);
const newFilename = `${index} ${entry.no_sku}${ext}`;
fs.copyFileSync(srcPath, path.join(FIXED_DIR, newFilename));
entry.filename = newFilename;
copied++;
});
fs.writeFileSync(LABELS_PATH, JSON.stringify(labels, null, 2), "utf8");
console.log(`Copied ${copied} images into ${FIXED_DIR} and updated ${LABELS_PATH}.`);
+130
View File
@@ -0,0 +1,130 @@
// One-off script: fill ground-truth labels for backend/sources/product-test-images/
// (the product-scan accuracy harness's Validation Set, previously 0 labeled images).
// Run once from backend/: node scripts/seed-validation-labels.mjs
import fs from "fs";
import path from "path";
const IMAGES_DIR = path.join("sources", "product-test-images");
const LABELS_PATH = path.join("sources", "product_manual_labels.json");
// [no_sku, nama_item, expiry_date ("" = not legible in photo, needs re-shoot)]
const DATA = [
["11110059", "CEKER BERKUKU FROZEN PACK 1 KG(*)", "13/06/2027"],
["11140051", "AMPELA FROZEN PACK 1 KG(*)", "26/11/2026"],
["11620056", "SBL (FILLET PAHA) 1 KG(*)", "27/02/2027"],
["11650053", "PAHA ATAS 1 KG(*)", "26/02/2027"],
["11660050", "PAHA BAWAH (1 KG)(*)", "22/06/2027"],
["11710051", "DADA UTUH (1 KG)(*)", ""],
["11818300", "CP-BEBEK GORENG 400GR/PAC", ""],
["1195008A", "RTC CHICKEN KALASAN 400 GR (PAC)", ""],
["11959937", "SATE AYAM FRESHMART 360 GR (PAC)", "24/11/2026"],
["12010111", "FIESTA CRISPY BUBBLE 400 GR/PAC", "07/05/2027"],
["12010115", "FIESTA NUGGET ZOO 400 GR/PAC", "01/12/2026"],
["12010117", "FIESTA NUGGET HAPPY STAR 400 GR/PAC", "27/08/2026"],
["12010119", "FIESTA NUGGET CHEESE 123 400 GR/PAC", "09/04/2027"],
["12010121", "FIESTA NUGGET PIZZABC 400 GR/PAC", "12/11/2026"],
["12010127", "FIESTA SPICY NUGGET 400 GR/PAC", "09/12/2027"],
["12010509", "CHAMP CRUNCHY NUGGET 450 GR/PAC", "15/04/2027"],
["12010515", "CHAMP KOIN KOMBINASI 450 GR/PAC", "15/04/2027"],
["12010519", "CHAMP NUGGET STICK 900 GR/PAC", "08/03/2027"],
["12012202", "ASIMO NUGGET KOMBINASI 1 KG/PAC", "13/05/2027"],
["12012501", "AKUMO CHICKEN NAGET 250 GR", "21/05/2027"],
["12012503", "AKUMO CHICKEN NUGGET 1000 GR", "17/06/2027"],
["12012505", "AKUMO KOIN 400 GR/PAC", "16/11/2026"],
["12020102", "FIESTA SPICY WING 400 GR/PAC", "22/05/2027"],
["12030102", "FIESTA STIKIE 200 GR/PAC", "26/02/2027"],
["12030403", "GOLDEN FIESTA STIKIE W/ SWEET CHILLI SAUCE 500GR", "07/04/2027"],
["12032502", "AKUMO CHICKEN STIK 500 GR", "09/03/2027"],
["12040101", "FIESTA SCHNITZEL 400 GR/PAC", "09/10/2026"],
["12040102", "FIESTA CRISPY BUBBLE KATSU 400 GR/PAC", "12/04/2027"],
["12060103", "FIESTA KARAGE 200 GR/PAC", ""],
["12060402", "GOLDEN FIESTA KARAGE CHILI SAUCE 500GR", "15/05/2027"],
["12080101", "FIESTA SPICY CHICK 400 GR/PAC", "26/04/2027"],
["12130102", "FIESTA CRISPY BURGER 360 GR (NEW)", "27/04/2027"],
["12150201", "FIESTA DS CRISPY CRUNCH 300 GR/PAC", "03/06/2027"],
["12150501", "CHAMP CRUNCHY HOTZZ 300 GR/PAC", ""],
["12190103", "FIESTA DELISTRIPE 400 GR/PAC", "07/04/2027"],
["12240103", "FIESTA YAKINIKU R/BITES 400 GR/PAC", "07/05/2027"],
["13010111", "FIESTA SOSIS BRATWURST 300 GR", "25/03/2027"],
["13010116", "FIESTA SSG ORIGINAL 300 GR", "23/06/2027"],
["13010118", "FIESTA RTG SSG 65 GR/PAC", ""],
["13010120", "FIESTA RTG C/CHEESY MELTS 65 GR/PAC", ""],
["13010122", "FIESTA RTG SAUSAGE WITH HOT LAVA 60G", ""],
["13010123", "FIESTA RTG SAUSAGE WITH CHEESE LAVA 60G", "25/10/2026"],
["13010125", "FIESTA RTG SAUSAGE WITH MENTAI LAVA 60GR", "05/11/2026"],
["13010518", "CHAMP SSG JUMBO BAKAR 500 GR/PAC", ""],
["13010524", "CHAMP SSG JUMBO BAKAR 500 GR/PAC (NEW)", ""],
["13012206", "ASIMO SOSIS AYAM KOMBINASI 500 GR", "15/03/2027"],
["13030501", "CHAMP CHICK MEATBALL 200 GR", ""],
["13070506", "CHAMP FRANKFURTER SSG 375GR", "13/03/2027"],
["13100512", "CHAMP CHICK SSG S/SANTAP ORIG 546GR (CAN)", ""],
["15010101", "FIESTA SHOESTRING 500 GR", "04/06/2027"],
["15010102", "FIESTA SHOESTRING 1000 GR", "05/03/2027"],
["15010107", "FIESTA FRENCH F SHOESTRING INSTITUSI 2KG", "17/06/2027"],
["15020101", "FIESTA STRAIGHT CUT 500 GR", "19/06/2027"],
["15020102", "FIESTA STRAIGHT CUT 1000 GR", "18/05/2027"],
["15030101", "FIESTA CRINKLE CUT 500 GR", ""],
["15030102", "FIESTA CRINKLE CUT 1000 GR", "07/04/2027"],
["16060113", "FIESTA CHICK SIOMAY 180GR (NEW)", "24/02/2027"],
["16060114", "FIESTA GYOZA 180 GR (NEW)", "18/05/2027"],
["17200109", "FIESTA RTS C/TERIYAKI 300GR/PAC", ""],
["17210106", "FIESTA RTS B/YAKINIKU 300GR/PAC", "19/05/2027"],
["17210107", "FIESTA RTS B/RENDANG 300GR/PAC", "23/06/2027"],
["17210108", "FIESTA RTS B/BLACKPEPPER 300GR/PAC", "09/05/2027"],
["17210109", "FIESTA RTS B/BULGOGI 300GR/PAC", "18/05/2027"],
["20040101", "FIESTA RAMEN BEKU 570 GR/PAC", "23/06/2026"],
["20120102", "FIESTA T/B AYAM GORENG 80 GR", ""],
["20120105", "FIESTA T/B SERBAGUNA Â 80 GR", "09/03/2027"],
["20120115", "FIESTA RACIK AYAM GORENG 20 GR/PAC", ""],
["20120116", "FIESTA RACIK NASI GORENG 20 GR/PAC", "13/10/2026"],
["21000123", "FIESTA RICE W/GEPREK CHICKEN 320GR/PAC", "30/04/2027"],
["21000126", "NEW FIESTA CHICK RENDANG W RICE 320GR (PAC)", "16/04/2027"],
["21000130", "NEW FIESTA RICE W/C CHEESE BULDAK 320GR (PAC)", "09/04/2027"],
["21000137", "FIESTA HAINAMESE CHICKEN RICE 320GR (PAC)", "12/02/2027"],
["21010101", "FIESTA TRUFFLE GYUDON 320 GR/PAC", "18/03/2027"],
["21200107", "NEW FIESTA SPAGHETTI CARBONARA 300GR (PAC)", "20/05/2027"],
];
const files = fs
.readdirSync(IMAGES_DIR)
.filter((f) => f !== "README.md" && f !== "filelist.txt")
.sort();
if (files.length !== DATA.length) {
throw new Error(`File count ${files.length} != DATA count ${DATA.length}`);
}
const existing = JSON.parse(fs.readFileSync(LABELS_PATH, "utf8"));
const now = new Date().toISOString();
let added = 0;
let skipped = 0;
let unreadable = 0;
for (let i = 0; i < files.length; i++) {
const filename = files[i];
const [no_sku, nama_item, expiry_date] = DATA[i];
if (existing.some((l) => l.filename === filename)) {
skipped++;
continue;
}
if (!expiry_date) unreadable++;
existing.push({
filename,
no_sku,
nama_item,
expiry_date,
top1_confidence: null,
notes: expiry_date
? ""
: "expiry date not legible in photo (cropped/blurry/out of frame) - needs re-shoot",
saved_at: now,
});
added++;
}
fs.writeFileSync(LABELS_PATH, JSON.stringify(existing, null, 2), "utf8");
console.log(`Added ${added} labels (${unreadable} flagged with no expiry_date), skipped ${skipped} already-labeled.`);
Binary file not shown.

After

Width:  |  Height:  |  Size: 3.4 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 140 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.3 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 257 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.7 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 229 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 254 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 268 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 4.2 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 184 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 220 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.3 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 227 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 170 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 217 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 217 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 201 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.9 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 239 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 194 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 200 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 207 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.1 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 275 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 231 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 304 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.3 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 278 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.0 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 178 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 293 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 284 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 223 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 164 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.4 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.4 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.8 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 225 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 4.5 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 232 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.6 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 235 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.4 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 149 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 285 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.6 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.8 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.7 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 106 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.5 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 175 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.9 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 208 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 143 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 251 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.4 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 176 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 230 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.3 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 302 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 114 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 217 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 208 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 129 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 288 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 4.1 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.5 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 232 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 286 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.8 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 171 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 265 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 263 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 261 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 227 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.8 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 222 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 215 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.1 MiB

Loaded 100 of 156 files, more files were not shown because too many files have changed in this diff. Show more