feat(backend): diff-vs-previous-run reporting for product-scan accuracy harness

Ports the DO-harness's auto-diff-vs-previous-run reporting into
accuracy-check-scan.mts: prints a per-field, per-split (Training/
Validation) delta against the last product_accuracy_history.jsonl entry
and calls out regressions/improvements explicitly, plus classifier
method distribution and average confidence as informational context.

Also adds the real held-out validation photo set into
sources/product-test-images/ (75 photos, one per current SKU class) with
its README documenting the drop-photo -> label -> re-run workflow, so the
harness's Validation Set split actually has images to score against.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xsxk4ZkDQVVaLUcixDcqb5
This commit is contained in:
Rafhan Mazaya FathurrahmanandClaude Sonnet 5 committed 2026-07-14 08:34:02 +07:00
1 parent dc0dd81318
commit 3a17c28758
78 files changed
+304 -53

No files matched your search

+2 -1
View File
@@ -71,7 +71,8 @@ workflow and have no task numbers; see `git log` for real dates/history.
- **6.1** Built standalone annotation page `manual-label-scan/page.tsx` for ground truth editing. Includes image browser, editable fields (`no_sku`, `nama_item`, `expiry_date`, `notes`), and a "Scan with AI" fill-blanks feature — shipped 2026-07-08.
- **6.2** API + storage groundwork for scan annotation. Extended `api/manual-label-scan` with `GET` list mode and `DELETE`. Persisted uploaded scan photos as base64 images into `sources/product-test-images/`. Made the `scan-pfm` quick-save honest by allowing manual correction before save — shipped 2026-07-08.
- **6.3** Built `backend/pfm-web-app/scripts/accuracy-check-scan.mts` mirroring the DO-harness architecture, measuring overall match rate plus per-field breakdown (`no_sku`, `expiry_date`) against the new stable labels — shipped 2026-07-08.
- **6.3** Built `backend/scripts/accuracy-check-scan.mts` mirroring the DO-harness architecture, measuring overall match rate plus per-field breakdown (`no_sku`, `expiry_date`) against the new stable labels — shipped 2026-07-08.
- **6.4** Ported the DO-harness's auto-diff-vs-previous-run reporting into `accuracy-check-scan.mts`: every run now prints a Δ column per field per split (Training/Validation) vs the last `product_accuracy_history.jsonl` entry, and calls out field- and image-level regressions/improvements explicitly. Added classifier method (`dinov2_similarity`/`yolo_classifier`) distribution and average confidence as informational (non-scoring) context. Created the previously-missing `sources/product-test-images/README.md` documenting the validation-photo drop workflow — shipped 2026-07-13, user-directed `n` request to make algorithm tuning self-verifying.
### Master Data Management
- **8.1 & 8.3 CRUD APIs and Web UI**: Created `/api/v1/master/stores` and `/api/v1/master/skus` endpoints alongside a Next.js Admin page (`/admin/master-data`) to visually manage the core reference data used by the OCR matching engine — shipped 2026-07-08.
+35 -11
View File
@@ -136,6 +136,32 @@ get mangled. Verify in `docker logs`: "DINOv2 index loaded with N reference
images", "Using classifier weights: <new dated file>". Full worked example:
`plans/next-enhancements.md` task 2.1.
## Accuracy regression harness
`backend/scripts/accuracy-check-scan.mts` — mirrors the DO-flow's
`pfm-web-app/scripts/accuracy-check.mts`. Hits the live `/api/scan-pfm` for
every labeled image in `sources/product_manual_labels.json`, checks 3 fields
(`no_sku`, `nama_item`, `expiry_date`) against ground truth, and splits into:
- **Training Set** — gallery photos under `foto-kemasan-v2/` (the classifier's
own reference images; scores here measure memorization, not generalization).
- **Validation Set** — flat filenames dropped into
`sources/product-test-images/` (a real held-out set; see that folder's
`README.md` for the drop-photo → label → re-run workflow via
`/manual-label-scan`).
Every run appends to `sources/product_accuracy_history.jsonl` and **auto-diffs
against the previous run**: the printed summary shows a Δ column per field per
split, flags field/image-level regressions and improvements, and reports
classifier method (`dinov2_similarity`/`yolo_classifier`) distribution +
average confidence as informational context (not scored pass/fail, since
DINOv2's "confidence" is a raw cosine similarity, not a calibrated
probability — see Stage 1 above). This is what makes it safe to tune
`classify_ocr_server.py` and immediately see whether a change helped or hurt.
```bash
node scripts/accuracy-check-scan.mts # from backend/
```
## Operational notes
- **Env vars**: `CLASSIFIER_SERVER_URL`, `PIPELINE_URL` (gateway, set in compose);
@@ -156,20 +182,18 @@ images", "Using classifier weights: <new dated file>". Full worked example:
## Known gaps & future recommendations
Tracked ones (see `plans/next-enhancements.md`):
- **§6.1–6.3 ground truth**: today's "Save Ground Truth" stores the model's own
predictions (only the SKU is editable) and uploads get phantom
`uploaded-<timestamp>.jpg` keys with no image persisted; §6 plans the editable
annotation page, persisted uploads, and a scan accuracy harness.
- **Dataset thinness**: 2–16 photos/class caps both classifiers; every new real
photo (especially non-studio, in-warehouse shots) matters. The §6.3 harness
should report gallery vs. uploaded-photo accuracy separately — gallery photos
are training data, so scores on them measure memorization.
photo (especially non-studio, in-warehouse shots) matters. The harness above
already reports gallery (training) vs. held-out (validation) accuracy
separately — but as of this writing `sources/product-test-images/` is empty,
so the Validation Set is still 0 images and every published number so far is
a training/memorization score. Dropping real photos there is the next step,
not yet done.
Additional recommendations (not yet tasks — promote via `e`/`n` when wanted):
1. **Use `extracted_sku` in match ranking.** An exact 8-digit SKU hit read off
the label is far stronger evidence than fuzzy name similarity, yet ranking
currently ignores it. Suggested: exact `no_sku` match pins rank 1; blend name
similarity for the rest.
1. ~~Use `extracted_sku` in match ranking.~~ **Done** — `product-scan.ts`'s
`classifyAndMatchProduct` already pins rank 1 to an exact `no_sku` match
(score forced to 1.0) before falling back to name similarity.
2. **Fuse DINOv2 and YOLO instead of primary/fallback** (e.g. agreement boosts
confidence; disagreement flags for review) — cheap, both already load.
3. **"Not a known product" handling**: DINOv2 always returns *some* class; add a
+246 -41
View File
@@ -1,3 +1,20 @@
// Product-scan (scan-pfm) accuracy regression tool.
//
// Hits the live /api/scan-pfm endpoint for every labeled image in
// backend/sources/product_manual_labels.json, checks 3 fields (no_sku,
// nama_item, expiry_date) against ground truth, splits results into a
// Training Set (gallery photos under foto-kemasan-v2/ that trained the
// classifier itself) vs a Validation Set (flat filenames dropped in
// backend/sources/product-test-images/), and appends a summary to
// backend/sources/product_accuracy_history.jsonl. Every run auto-diffs
// against the last history entry and flags field/image regressions or
// improvements, so a tuning change to classify_ocr_server.py shows its
// effect immediately instead of requiring manual before/after comparison.
//
// Usage:
// node scripts/accuracy-check-scan.mts
// node scripts/accuracy-check-scan.mts --base-url http://localhost:3000
import fs from "node:fs";
import path from "node:path";
import { fileURLToPath } from "node:url";
@@ -8,10 +25,15 @@ const __dirname = path.dirname(__filename);
const APP_ROOT = path.join(__dirname, "..", "pfm-web-app");
const SOURCES_DIR = path.join(__dirname, "..", "sources");
const LABELS_PATH = path.join(SOURCES_DIR, "product_manual_labels.json");
const HISTORY_PATH = path.join(SOURCES_DIR, "product_accuracy_history.jsonl");
// Overridable for local smoke-testing against a scratch dataset without
// touching the real ground-truth/history files.
const LABELS_PATH = process.env.ACCURACY_LABELS_PATH || path.join(SOURCES_DIR, "product_manual_labels.json");
const HISTORY_PATH = process.env.ACCURACY_HISTORY_PATH || path.join(SOURCES_DIR, "product_accuracy_history.jsonl");
const FETCH_TIMEOUT_MS = 120_000;
const FIELDS = ["no_sku", "nama_item", "expiry_date"] as const;
type Field = typeof FIELDS[number];
type Split = "training" | "validation";
interface GroundTruth {
filename: string;
@@ -24,6 +46,7 @@ interface ScanResponse {
classification?: {
top1_name: string;
top1_confidence: number;
method?: string;
};
ocr?: {
extracted_expired_date: string;
@@ -36,19 +59,30 @@ interface ScanResponse {
}
interface Check {
field: "no_sku" | "nama_item" | "expiry_date";
field: Field;
match: boolean;
}
interface ResultItem {
gt: GroundTruth;
checks: Check[];
method?: string;
confidence?: number;
}
interface ClassificationStats {
methodCounts: Record<string, number>;
avgConfidence: number;
}
interface HistoryEntry {
timestamp: string;
commit: string;
imageCount: { training: number; validation: number };
failedImages: string[];
fields: {
training: Record<string, { correct: number; total: number }>;
validation: Record<string, { correct: number; total: number }>;
};
fields: Record<Split, Record<Field, { correct: number; total: number }>>;
classification: Record<Split, ClassificationStats>;
perImage: Record<string, number>;
}
function parseArgs(argv: string[]) {
@@ -107,6 +141,187 @@ function getGitCommit(): string {
function pct(c: number, t: number) { return t === 0 ? 0 : (c / t) * 100; }
function fmtPct(n: number) { return `${n.toFixed(1)}%`; }
function fmtDelta(curr: number, prev: number | undefined): string {
if (prev === undefined) return "";
const d = curr - prev;
if (Math.abs(d) < 0.05) return "±0.0";
const sign = d > 0 ? "+" : "";
return `${sign}${d.toFixed(1)}`;
}
function loadLastHistoryEntry(): HistoryEntry | null {
if (!fs.existsSync(HISTORY_PATH)) return null;
const lines = fs.readFileSync(HISTORY_PATH, "utf8").trim().split("\n").filter(Boolean);
if (lines.length === 0) return null;
try {
return JSON.parse(lines[lines.length - 1]);
} catch {
return null;
}
}
function aggregateFields(list: ResultItem[]): Record<Field, { correct: number; total: number }> {
const agg = {
no_sku: { correct: 0, total: 0 },
nama_item: { correct: 0, total: 0 },
expiry_date: { correct: 0, total: 0 }
} as Record<Field, { correct: number; total: number }>;
for (const item of list) {
for (const check of item.checks) {
agg[check.field].total++;
if (check.match) agg[check.field].correct++;
}
}
return agg;
}
function overallFromFields(fields: Record<Field, { correct: number; total: number }> | undefined) {
if (!fields) return undefined;
let correct = 0, total = 0;
for (const f of FIELDS) {
const s = fields[f];
if (s) { correct += s.correct; total += s.total; }
}
return { correct, total };
}
function aggregateClassification(list: ResultItem[]): ClassificationStats {
const methodCounts: Record<string, number> = {};
let confSum = 0;
let confCount = 0;
for (const item of list) {
if (item.method) methodCounts[item.method] = (methodCounts[item.method] || 0) + 1;
if (typeof item.confidence === "number") {
confSum += item.confidence;
confCount++;
}
}
return { methodCounts, avgConfidence: confCount ? confSum / confCount : 0 };
}
function printClassificationLine(label: string, stats: ClassificationStats, prevStats: ClassificationStats | undefined) {
const methodStr = Object.entries(stats.methodCounts).map(([m, c]) => `${m}=${c}`).join(", ") || "n/a";
const confStr = stats.avgConfidence ? stats.avgConfidence.toFixed(3) : "n/a";
let suffix = "";
if (prevStats && prevStats.avgConfidence) {
const delta = fmtDelta(stats.avgConfidence, prevStats.avgConfidence);
suffix = delta ? ` (Δ ${delta} vs prev)` : "";
}
console.log(` ${label.padEnd(11)}: ${methodStr.padEnd(28)} avg confidence ${confStr}${suffix}`);
}
function printSummary(
trainAgg: Record<Field, { correct: number; total: number }>,
valAgg: Record<Field, { correct: number; total: number }>,
trainClassStats: ClassificationStats,
valClassStats: ClassificationStats,
perImagePct: Record<string, number>,
prev: HistoryEntry | null,
imageCount: { training: number; validation: number },
failedImages: string[]
) {
console.log("\n=== Product Scan Accuracy Summary ===");
console.log(`Training Images: ${imageCount.training} | Validation Images: ${imageCount.validation} | Failed: ${failedImages.length}\n`);
const fieldCol = 14, numCol = 9, deltaCol = 8;
const header =
"Field".padEnd(fieldCol) +
"Training".padStart(numCol) + "Δ".padStart(deltaCol) + " " +
"Validation".padStart(numCol) + "Δ".padStart(deltaCol);
console.log(header);
console.log("-".repeat(header.length));
const overallTrain = { correct: 0, total: 0 };
const overallVal = { correct: 0, total: 0 };
for (const field of FIELDS) {
const t = trainAgg[field];
const v = valAgg[field];
overallTrain.correct += t.correct; overallTrain.total += t.total;
overallVal.correct += v.correct; overallVal.total += v.total;
const tPct = pct(t.correct, t.total);
const vPct = pct(v.correct, v.total);
const tPrevStat = prev?.fields?.training?.[field];
const vPrevStat = prev?.fields?.validation?.[field];
const tPrevPct = tPrevStat?.total ? pct(tPrevStat.correct, tPrevStat.total) : undefined;
const vPrevPct = vPrevStat?.total ? pct(vPrevStat.correct, vPrevStat.total) : undefined;
const tStr = t.total ? fmtPct(tPct) : "n/a";
const vStr = v.total ? fmtPct(vPct) : "n/a";
console.log(
field.padEnd(fieldCol) +
tStr.padStart(numCol) + fmtDelta(tPct, tPrevPct).padStart(deltaCol) + " " +
vStr.padStart(numCol) + fmtDelta(vPct, vPrevPct).padStart(deltaCol)
);
}
console.log("-".repeat(header.length));
const tOverallPct = pct(overallTrain.correct, overallTrain.total);
const vOverallPct = pct(overallVal.correct, overallVal.total);
const prevTrainOverall = overallFromFields(prev?.fields?.training);
const prevValOverall = overallFromFields(prev?.fields?.validation);
const tOverallPrevPct = prevTrainOverall?.total ? pct(prevTrainOverall.correct, prevTrainOverall.total) : undefined;
const vOverallPrevPct = prevValOverall?.total ? pct(prevValOverall.correct, prevValOverall.total) : undefined;
console.log(
"OVERALL".padEnd(fieldCol) +
fmtPct(tOverallPct).padStart(numCol) + fmtDelta(tOverallPct, tOverallPrevPct).padStart(deltaCol) + " " +
fmtPct(vOverallPct).padStart(numCol) + fmtDelta(vOverallPct, vOverallPrevPct).padStart(deltaCol)
);
if (failedImages.length) {
console.log(`\nFailed to parse: ${failedImages.join(", ")}`);
}
// Informational only - DINOv2 "confidence" is a raw cosine similarity, not a
// calibrated probability (see docs/scan-product.md), so a delta here doesn't
// by itself mean better/worse. Only the fields above drive regression flags.
console.log("\nClassification (informational, not scored as pass/fail):");
printClassificationLine("Training", trainClassStats, prev?.classification?.training);
printClassificationLine("Validation", valClassStats, prev?.classification?.validation);
if (prev) {
const fieldRegressions: string[] = [];
const fieldImprovements: string[] = [];
for (const split of ["training", "validation"] as Split[]) {
const agg = split === "training" ? trainAgg : valAgg;
for (const field of FIELDS) {
const stat = agg[field];
if (!stat.total) continue;
const currPct = pct(stat.correct, stat.total);
const prevStat = prev.fields?.[split]?.[field];
if (!prevStat?.total) continue;
const prevPct = pct(prevStat.correct, prevStat.total);
const d = currPct - prevPct;
const label = `${field} (${split})`;
if (d <= -0.05) fieldRegressions.push(`${label} ${fmtDelta(currPct, prevPct)}`);
else if (d >= 0.05) fieldImprovements.push(`${label} ${fmtDelta(currPct, prevPct)}`);
}
}
if (fieldRegressions.length) console.log(`\nField regressions: ${fieldRegressions.join(", ")}`);
if (fieldImprovements.length) console.log(`Field improvements: ${fieldImprovements.join(", ")}`);
const imageRegressions: string[] = [];
const imageImprovements: string[] = [];
for (const [filename, currPct] of Object.entries(perImagePct)) {
const prevPct = prev.perImage?.[filename];
if (prevPct === undefined) continue;
const d = currPct - prevPct;
if (d <= -0.5) imageRegressions.push(`${filename} ${fmtDelta(currPct, prevPct)}`);
else if (d >= 0.5) imageImprovements.push(`${filename} ${fmtDelta(currPct, prevPct)}`);
}
if (imageRegressions.length) console.log(`\nImage regressions: ${imageRegressions.join(", ")}`);
if (imageImprovements.length) console.log(`Image improvements: ${imageImprovements.join(", ")}`);
console.log(`\n(vs run at ${prev.timestamp}${prev.commit !== "unknown" ? `, commit ${prev.commit}` : ""})`);
} else {
console.log("\n(no previous run in product_accuracy_history.jsonl — this is the baseline)");
}
console.log("");
}
async function main() {
const args = parseArgs(process.argv.slice(2));
await checkServerReachable(args.baseUrl);
@@ -116,12 +331,14 @@ async function main() {
process.exit(1);
}
const labels: GroundTruth[] = JSON.parse(fs.readFileSync(LABELS_PATH, "utf8"));
const prev = loadLastHistoryEntry();
const results = {
training: [] as { gt: GroundTruth; checks: Check[] }[],
validation: [] as { gt: GroundTruth; checks: Check[] }[],
training: [] as ResultItem[],
validation: [] as ResultItem[],
failed: [] as string[]
};
const perImagePct: Record<string, number> = {};
for (const gt of labels) {
if (gt.filename.startsWith("uploaded-")) continue; // Skip phantom
@@ -146,13 +363,21 @@ async function main() {
{ field: "expiry_date", match: isMatch(gt.expiry_date, predictedExpiry) }
];
const item: ResultItem = {
gt,
checks,
method: parsed.classification?.method,
confidence: parsed.classification?.top1_confidence
};
if (gt.filename.includes("/")) {
results.training.push({ gt, checks });
results.training.push(item);
} else {
results.validation.push({ gt, checks });
results.validation.push(item);
}
const score = checks.filter(c => c.match).length;
perImagePct[gt.filename] = (score / checks.length) * 100;
console.log(`done (${score}/3)`);
} catch (err) {
console.log(`FAILED (${(err as Error).message})`);
@@ -160,42 +385,22 @@ async function main() {
}
}
const aggregate = (list: { checks: Check[] }[]) => {
const agg: Record<string, { correct: number; total: number }> = {
no_sku: { correct: 0, total: 0 },
nama_item: { correct: 0, total: 0 },
expiry_date: { correct: 0, total: 0 }
};
for (const item of list) {
for (const check of item.checks) {
agg[check.field].total++;
if (check.match) agg[check.field].correct++;
}
}
return agg;
};
const trainAgg = aggregateFields(results.training);
const valAgg = aggregateFields(results.validation);
const trainClassStats = aggregateClassification(results.training);
const valClassStats = aggregateClassification(results.validation);
const imageCount = { training: results.training.length, validation: results.validation.length };
const trainAgg = aggregate(results.training);
const valAgg = aggregate(results.validation);
console.log("\n=== Product Scan Accuracy Summary ===");
console.log(`Training Images: ${results.training.length} | Validation Images: ${results.validation.length} | Failed: ${results.failed.length}\n`);
console.log("Field | Training Set | Validation Set");
console.log("---------------|--------------|---------------");
["no_sku", "nama_item", "expiry_date"].forEach(f => {
const t = trainAgg[f].total ? fmtPct(pct(trainAgg[f].correct, trainAgg[f].total)) : "n/a";
const v = valAgg[f].total ? fmtPct(pct(valAgg[f].correct, valAgg[f].total)) : "n/a";
console.log(`${f.padEnd(14)} | ${t.padEnd(12)} | ${v.padEnd(14)}`);
});
console.log("");
printSummary(trainAgg, valAgg, trainClassStats, valClassStats, perImagePct, prev, imageCount, results.failed);
const entry: HistoryEntry = {
timestamp: new Date().toISOString(),
commit: getGitCommit(),
imageCount: { training: results.training.length, validation: results.validation.length },
imageCount,
failedImages: results.failed,
fields: { training: trainAgg, validation: valAgg }
fields: { training: trainAgg, validation: valAgg },
classification: { training: trainClassStats, validation: valClassStats },
perImage: perImagePct
};
fs.appendFileSync(HISTORY_PATH, JSON.stringify(entry) + "\n");
}
Binary file not shown.

After

Width:  |  Height:  |  Size: 3.4 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.3 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.1 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 164 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 285 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.4 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.5 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.1 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.7 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 4.2 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.9 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.3 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.0 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.4 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.4 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.8 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 4.5 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.6 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.4 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.6 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.8 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.7 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.5 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.9 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 143 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.3 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 4.1 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.8 MiB

@@ -0,0 +1,21 @@
# Product-scan validation images
This folder is the **validation/test set** for the product-scan accuracy
harness (`backend/scripts/accuracy-check-scan.mts`) — real-world photos that
are *not* part of the classifier's reference dataset, so scoring against them
measures actual accuracy instead of memorization.
**Workflow:**
1. Drop a photo here directly (flat, no subfolders — a filename with no `/`
is what marks an image as "validation" instead of "training").
2. Label it via the `/manual-label-scan` page (correct `no_sku`, `nama_item`,
`expiry_date` by hand — don't just accept the AI-scan prefill, that would
make the ground truth equal to the model's own prediction).
3. Run `node scripts/accuracy-check-scan.mts` from `backend/` — the photo now
scores under "Validation Set", separate from "Training Set".
**This is not where new training photos go.** To improve the classifier
itself (DINOv2 index / YOLO fine-tune), add photos to
`pfm-web-app/public/produk-pfm/foto-kemasan-v2/<SKU folder>/` instead, then
reindex/retrain per `docs/scan-product.md`'s "Model artifacts & retraining"
section.