feat(backend): diff-vs-previous-run reporting for product-scan accuracy harness
Ports the DO-harness's auto-diff-vs-previous-run reporting into accuracy-check-scan.mts: prints a per-field, per-split (Training/ Validation) delta against the last product_accuracy_history.jsonl entry and calls out regressions/improvements explicitly, plus classifier method distribution and average confidence as informational context. Also adds the real held-out validation photo set into sources/product-test-images/ (75 photos, one per current SKU class) with its README documenting the drop-photo -> label -> re-run workflow, so the harness's Validation Set split actually has images to score against. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Xsxk4ZkDQVVaLUcixDcqb5
No files matched your search
@@ -71,7 +71,8 @@ workflow and have no task numbers; see `git log` for real dates/history.
|
|||||||
|
|
||||||
- **6.1** Built standalone annotation page `manual-label-scan/page.tsx` for ground truth editing. Includes image browser, editable fields (`no_sku`, `nama_item`, `expiry_date`, `notes`), and a "Scan with AI" fill-blanks feature — shipped 2026-07-08.
|
- **6.1** Built standalone annotation page `manual-label-scan/page.tsx` for ground truth editing. Includes image browser, editable fields (`no_sku`, `nama_item`, `expiry_date`, `notes`), and a "Scan with AI" fill-blanks feature — shipped 2026-07-08.
|
||||||
- **6.2** API + storage groundwork for scan annotation. Extended `api/manual-label-scan` with `GET` list mode and `DELETE`. Persisted uploaded scan photos as base64 images into `sources/product-test-images/`. Made the `scan-pfm` quick-save honest by allowing manual correction before save — shipped 2026-07-08.
|
- **6.2** API + storage groundwork for scan annotation. Extended `api/manual-label-scan` with `GET` list mode and `DELETE`. Persisted uploaded scan photos as base64 images into `sources/product-test-images/`. Made the `scan-pfm` quick-save honest by allowing manual correction before save — shipped 2026-07-08.
|
||||||
- **6.3** Built `backend/pfm-web-app/scripts/accuracy-check-scan.mts` mirroring the DO-harness architecture, measuring overall match rate plus per-field breakdown (`no_sku`, `expiry_date`) against the new stable labels — shipped 2026-07-08.
|
- **6.3** Built `backend/scripts/accuracy-check-scan.mts` mirroring the DO-harness architecture, measuring overall match rate plus per-field breakdown (`no_sku`, `expiry_date`) against the new stable labels — shipped 2026-07-08.
|
||||||
|
- **6.4** Ported the DO-harness's auto-diff-vs-previous-run reporting into `accuracy-check-scan.mts`: every run now prints a Δ column per field per split (Training/Validation) vs the last `product_accuracy_history.jsonl` entry, and calls out field- and image-level regressions/improvements explicitly. Added classifier method (`dinov2_similarity`/`yolo_classifier`) distribution and average confidence as informational (non-scoring) context. Created the previously-missing `sources/product-test-images/README.md` documenting the validation-photo drop workflow — shipped 2026-07-13, user-directed `n` request to make algorithm tuning self-verifying.
|
||||||
|
|
||||||
### Master Data Management
|
### Master Data Management
|
||||||
- **8.1 & 8.3 CRUD APIs and Web UI**: Created `/api/v1/master/stores` and `/api/v1/master/skus` endpoints alongside a Next.js Admin page (`/admin/master-data`) to visually manage the core reference data used by the OCR matching engine — shipped 2026-07-08.
|
- **8.1 & 8.3 CRUD APIs and Web UI**: Created `/api/v1/master/stores` and `/api/v1/master/skus` endpoints alongside a Next.js Admin page (`/admin/master-data`) to visually manage the core reference data used by the OCR matching engine — shipped 2026-07-08.
|
||||||
|
|||||||
@@ -136,6 +136,32 @@ get mangled. Verify in `docker logs`: "DINOv2 index loaded with N reference
|
|||||||
images", "Using classifier weights: <new dated file>". Full worked example:
|
images", "Using classifier weights: <new dated file>". Full worked example:
|
||||||
`plans/next-enhancements.md` task 2.1.
|
`plans/next-enhancements.md` task 2.1.
|
||||||
|
|
||||||
|
## Accuracy regression harness
|
||||||
|
|
||||||
|
`backend/scripts/accuracy-check-scan.mts` — mirrors the DO-flow's
|
||||||
|
`pfm-web-app/scripts/accuracy-check.mts`. Hits the live `/api/scan-pfm` for
|
||||||
|
every labeled image in `sources/product_manual_labels.json`, checks 3 fields
|
||||||
|
(`no_sku`, `nama_item`, `expiry_date`) against ground truth, and splits into:
|
||||||
|
- **Training Set** — gallery photos under `foto-kemasan-v2/` (the classifier's
|
||||||
|
own reference images; scores here measure memorization, not generalization).
|
||||||
|
- **Validation Set** — flat filenames dropped into
|
||||||
|
`sources/product-test-images/` (a real held-out set; see that folder's
|
||||||
|
`README.md` for the drop-photo → label → re-run workflow via
|
||||||
|
`/manual-label-scan`).
|
||||||
|
|
||||||
|
Every run appends to `sources/product_accuracy_history.jsonl` and **auto-diffs
|
||||||
|
against the previous run**: the printed summary shows a Δ column per field per
|
||||||
|
split, flags field/image-level regressions and improvements, and reports
|
||||||
|
classifier method (`dinov2_similarity`/`yolo_classifier`) distribution +
|
||||||
|
average confidence as informational context (not scored pass/fail, since
|
||||||
|
DINOv2's "confidence" is a raw cosine similarity, not a calibrated
|
||||||
|
probability — see Stage 1 above). This is what makes it safe to tune
|
||||||
|
`classify_ocr_server.py` and immediately see whether a change helped or hurt.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
node scripts/accuracy-check-scan.mts # from backend/
|
||||||
|
```
|
||||||
|
|
||||||
## Operational notes
|
## Operational notes
|
||||||
|
|
||||||
- **Env vars**: `CLASSIFIER_SERVER_URL`, `PIPELINE_URL` (gateway, set in compose);
|
- **Env vars**: `CLASSIFIER_SERVER_URL`, `PIPELINE_URL` (gateway, set in compose);
|
||||||
@@ -156,20 +182,18 @@ images", "Using classifier weights: <new dated file>". Full worked example:
|
|||||||
## Known gaps & future recommendations
|
## Known gaps & future recommendations
|
||||||
|
|
||||||
Tracked ones (see `plans/next-enhancements.md`):
|
Tracked ones (see `plans/next-enhancements.md`):
|
||||||
- **§6.1–6.3 ground truth**: today's "Save Ground Truth" stores the model's own
|
|
||||||
predictions (only the SKU is editable) and uploads get phantom
|
|
||||||
`uploaded-<timestamp>.jpg` keys with no image persisted; §6 plans the editable
|
|
||||||
annotation page, persisted uploads, and a scan accuracy harness.
|
|
||||||
- **Dataset thinness**: 2–16 photos/class caps both classifiers; every new real
|
- **Dataset thinness**: 2–16 photos/class caps both classifiers; every new real
|
||||||
photo (especially non-studio, in-warehouse shots) matters. The §6.3 harness
|
photo (especially non-studio, in-warehouse shots) matters. The harness above
|
||||||
should report gallery vs. uploaded-photo accuracy separately — gallery photos
|
already reports gallery (training) vs. held-out (validation) accuracy
|
||||||
are training data, so scores on them measure memorization.
|
separately — but as of this writing `sources/product-test-images/` is empty,
|
||||||
|
so the Validation Set is still 0 images and every published number so far is
|
||||||
|
a training/memorization score. Dropping real photos there is the next step,
|
||||||
|
not yet done.
|
||||||
|
|
||||||
Additional recommendations (not yet tasks — promote via `e`/`n` when wanted):
|
Additional recommendations (not yet tasks — promote via `e`/`n` when wanted):
|
||||||
1. **Use `extracted_sku` in match ranking.** An exact 8-digit SKU hit read off
|
1. ~~Use `extracted_sku` in match ranking.~~ **Done** — `product-scan.ts`'s
|
||||||
the label is far stronger evidence than fuzzy name similarity, yet ranking
|
`classifyAndMatchProduct` already pins rank 1 to an exact `no_sku` match
|
||||||
currently ignores it. Suggested: exact `no_sku` match pins rank 1; blend name
|
(score forced to 1.0) before falling back to name similarity.
|
||||||
similarity for the rest.
|
|
||||||
2. **Fuse DINOv2 and YOLO instead of primary/fallback** (e.g. agreement boosts
|
2. **Fuse DINOv2 and YOLO instead of primary/fallback** (e.g. agreement boosts
|
||||||
confidence; disagreement flags for review) — cheap, both already load.
|
confidence; disagreement flags for review) — cheap, both already load.
|
||||||
3. **"Not a known product" handling**: DINOv2 always returns *some* class; add a
|
3. **"Not a known product" handling**: DINOv2 always returns *some* class; add a
|
||||||
|
|||||||
@@ -1,3 +1,20 @@
|
|||||||
|
// Product-scan (scan-pfm) accuracy regression tool.
|
||||||
|
//
|
||||||
|
// Hits the live /api/scan-pfm endpoint for every labeled image in
|
||||||
|
// backend/sources/product_manual_labels.json, checks 3 fields (no_sku,
|
||||||
|
// nama_item, expiry_date) against ground truth, splits results into a
|
||||||
|
// Training Set (gallery photos under foto-kemasan-v2/ that trained the
|
||||||
|
// classifier itself) vs a Validation Set (flat filenames dropped in
|
||||||
|
// backend/sources/product-test-images/), and appends a summary to
|
||||||
|
// backend/sources/product_accuracy_history.jsonl. Every run auto-diffs
|
||||||
|
// against the last history entry and flags field/image regressions or
|
||||||
|
// improvements, so a tuning change to classify_ocr_server.py shows its
|
||||||
|
// effect immediately instead of requiring manual before/after comparison.
|
||||||
|
//
|
||||||
|
// Usage:
|
||||||
|
// node scripts/accuracy-check-scan.mts
|
||||||
|
// node scripts/accuracy-check-scan.mts --base-url http://localhost:3000
|
||||||
|
|
||||||
import fs from "node:fs";
|
import fs from "node:fs";
|
||||||
import path from "node:path";
|
import path from "node:path";
|
||||||
import { fileURLToPath } from "node:url";
|
import { fileURLToPath } from "node:url";
|
||||||
@@ -8,10 +25,15 @@ const __dirname = path.dirname(__filename);
|
|||||||
|
|
||||||
const APP_ROOT = path.join(__dirname, "..", "pfm-web-app");
|
const APP_ROOT = path.join(__dirname, "..", "pfm-web-app");
|
||||||
const SOURCES_DIR = path.join(__dirname, "..", "sources");
|
const SOURCES_DIR = path.join(__dirname, "..", "sources");
|
||||||
const LABELS_PATH = path.join(SOURCES_DIR, "product_manual_labels.json");
|
// Overridable for local smoke-testing against a scratch dataset without
|
||||||
const HISTORY_PATH = path.join(SOURCES_DIR, "product_accuracy_history.jsonl");
|
// touching the real ground-truth/history files.
|
||||||
|
const LABELS_PATH = process.env.ACCURACY_LABELS_PATH || path.join(SOURCES_DIR, "product_manual_labels.json");
|
||||||
|
const HISTORY_PATH = process.env.ACCURACY_HISTORY_PATH || path.join(SOURCES_DIR, "product_accuracy_history.jsonl");
|
||||||
|
|
||||||
const FETCH_TIMEOUT_MS = 120_000;
|
const FETCH_TIMEOUT_MS = 120_000;
|
||||||
|
const FIELDS = ["no_sku", "nama_item", "expiry_date"] as const;
|
||||||
|
type Field = typeof FIELDS[number];
|
||||||
|
type Split = "training" | "validation";
|
||||||
|
|
||||||
interface GroundTruth {
|
interface GroundTruth {
|
||||||
filename: string;
|
filename: string;
|
||||||
@@ -24,6 +46,7 @@ interface ScanResponse {
|
|||||||
classification?: {
|
classification?: {
|
||||||
top1_name: string;
|
top1_name: string;
|
||||||
top1_confidence: number;
|
top1_confidence: number;
|
||||||
|
method?: string;
|
||||||
};
|
};
|
||||||
ocr?: {
|
ocr?: {
|
||||||
extracted_expired_date: string;
|
extracted_expired_date: string;
|
||||||
@@ -36,19 +59,30 @@ interface ScanResponse {
|
|||||||
}
|
}
|
||||||
|
|
||||||
interface Check {
|
interface Check {
|
||||||
field: "no_sku" | "nama_item" | "expiry_date";
|
field: Field;
|
||||||
match: boolean;
|
match: boolean;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
interface ResultItem {
|
||||||
|
gt: GroundTruth;
|
||||||
|
checks: Check[];
|
||||||
|
method?: string;
|
||||||
|
confidence?: number;
|
||||||
|
}
|
||||||
|
|
||||||
|
interface ClassificationStats {
|
||||||
|
methodCounts: Record<string, number>;
|
||||||
|
avgConfidence: number;
|
||||||
|
}
|
||||||
|
|
||||||
interface HistoryEntry {
|
interface HistoryEntry {
|
||||||
timestamp: string;
|
timestamp: string;
|
||||||
commit: string;
|
commit: string;
|
||||||
imageCount: { training: number; validation: number };
|
imageCount: { training: number; validation: number };
|
||||||
failedImages: string[];
|
failedImages: string[];
|
||||||
fields: {
|
fields: Record<Split, Record<Field, { correct: number; total: number }>>;
|
||||||
training: Record<string, { correct: number; total: number }>;
|
classification: Record<Split, ClassificationStats>;
|
||||||
validation: Record<string, { correct: number; total: number }>;
|
perImage: Record<string, number>;
|
||||||
};
|
|
||||||
}
|
}
|
||||||
|
|
||||||
function parseArgs(argv: string[]) {
|
function parseArgs(argv: string[]) {
|
||||||
@@ -107,6 +141,187 @@ function getGitCommit(): string {
|
|||||||
function pct(c: number, t: number) { return t === 0 ? 0 : (c / t) * 100; }
|
function pct(c: number, t: number) { return t === 0 ? 0 : (c / t) * 100; }
|
||||||
function fmtPct(n: number) { return `${n.toFixed(1)}%`; }
|
function fmtPct(n: number) { return `${n.toFixed(1)}%`; }
|
||||||
|
|
||||||
|
function fmtDelta(curr: number, prev: number | undefined): string {
|
||||||
|
if (prev === undefined) return "";
|
||||||
|
const d = curr - prev;
|
||||||
|
if (Math.abs(d) < 0.05) return "±0.0";
|
||||||
|
const sign = d > 0 ? "+" : "";
|
||||||
|
return `${sign}${d.toFixed(1)}`;
|
||||||
|
}
|
||||||
|
|
||||||
|
function loadLastHistoryEntry(): HistoryEntry | null {
|
||||||
|
if (!fs.existsSync(HISTORY_PATH)) return null;
|
||||||
|
const lines = fs.readFileSync(HISTORY_PATH, "utf8").trim().split("\n").filter(Boolean);
|
||||||
|
if (lines.length === 0) return null;
|
||||||
|
try {
|
||||||
|
return JSON.parse(lines[lines.length - 1]);
|
||||||
|
} catch {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function aggregateFields(list: ResultItem[]): Record<Field, { correct: number; total: number }> {
|
||||||
|
const agg = {
|
||||||
|
no_sku: { correct: 0, total: 0 },
|
||||||
|
nama_item: { correct: 0, total: 0 },
|
||||||
|
expiry_date: { correct: 0, total: 0 }
|
||||||
|
} as Record<Field, { correct: number; total: number }>;
|
||||||
|
for (const item of list) {
|
||||||
|
for (const check of item.checks) {
|
||||||
|
agg[check.field].total++;
|
||||||
|
if (check.match) agg[check.field].correct++;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return agg;
|
||||||
|
}
|
||||||
|
|
||||||
|
function overallFromFields(fields: Record<Field, { correct: number; total: number }> | undefined) {
|
||||||
|
if (!fields) return undefined;
|
||||||
|
let correct = 0, total = 0;
|
||||||
|
for (const f of FIELDS) {
|
||||||
|
const s = fields[f];
|
||||||
|
if (s) { correct += s.correct; total += s.total; }
|
||||||
|
}
|
||||||
|
return { correct, total };
|
||||||
|
}
|
||||||
|
|
||||||
|
function aggregateClassification(list: ResultItem[]): ClassificationStats {
|
||||||
|
const methodCounts: Record<string, number> = {};
|
||||||
|
let confSum = 0;
|
||||||
|
let confCount = 0;
|
||||||
|
for (const item of list) {
|
||||||
|
if (item.method) methodCounts[item.method] = (methodCounts[item.method] || 0) + 1;
|
||||||
|
if (typeof item.confidence === "number") {
|
||||||
|
confSum += item.confidence;
|
||||||
|
confCount++;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return { methodCounts, avgConfidence: confCount ? confSum / confCount : 0 };
|
||||||
|
}
|
||||||
|
|
||||||
|
function printClassificationLine(label: string, stats: ClassificationStats, prevStats: ClassificationStats | undefined) {
|
||||||
|
const methodStr = Object.entries(stats.methodCounts).map(([m, c]) => `${m}=${c}`).join(", ") || "n/a";
|
||||||
|
const confStr = stats.avgConfidence ? stats.avgConfidence.toFixed(3) : "n/a";
|
||||||
|
let suffix = "";
|
||||||
|
if (prevStats && prevStats.avgConfidence) {
|
||||||
|
const delta = fmtDelta(stats.avgConfidence, prevStats.avgConfidence);
|
||||||
|
suffix = delta ? ` (Δ ${delta} vs prev)` : "";
|
||||||
|
}
|
||||||
|
console.log(` ${label.padEnd(11)}: ${methodStr.padEnd(28)} avg confidence ${confStr}${suffix}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
function printSummary(
|
||||||
|
trainAgg: Record<Field, { correct: number; total: number }>,
|
||||||
|
valAgg: Record<Field, { correct: number; total: number }>,
|
||||||
|
trainClassStats: ClassificationStats,
|
||||||
|
valClassStats: ClassificationStats,
|
||||||
|
perImagePct: Record<string, number>,
|
||||||
|
prev: HistoryEntry | null,
|
||||||
|
imageCount: { training: number; validation: number },
|
||||||
|
failedImages: string[]
|
||||||
|
) {
|
||||||
|
console.log("\n=== Product Scan Accuracy Summary ===");
|
||||||
|
console.log(`Training Images: ${imageCount.training} | Validation Images: ${imageCount.validation} | Failed: ${failedImages.length}\n`);
|
||||||
|
|
||||||
|
const fieldCol = 14, numCol = 9, deltaCol = 8;
|
||||||
|
const header =
|
||||||
|
"Field".padEnd(fieldCol) +
|
||||||
|
"Training".padStart(numCol) + "Δ".padStart(deltaCol) + " " +
|
||||||
|
"Validation".padStart(numCol) + "Δ".padStart(deltaCol);
|
||||||
|
console.log(header);
|
||||||
|
console.log("-".repeat(header.length));
|
||||||
|
|
||||||
|
const overallTrain = { correct: 0, total: 0 };
|
||||||
|
const overallVal = { correct: 0, total: 0 };
|
||||||
|
|
||||||
|
for (const field of FIELDS) {
|
||||||
|
const t = trainAgg[field];
|
||||||
|
const v = valAgg[field];
|
||||||
|
overallTrain.correct += t.correct; overallTrain.total += t.total;
|
||||||
|
overallVal.correct += v.correct; overallVal.total += v.total;
|
||||||
|
|
||||||
|
const tPct = pct(t.correct, t.total);
|
||||||
|
const vPct = pct(v.correct, v.total);
|
||||||
|
const tPrevStat = prev?.fields?.training?.[field];
|
||||||
|
const vPrevStat = prev?.fields?.validation?.[field];
|
||||||
|
const tPrevPct = tPrevStat?.total ? pct(tPrevStat.correct, tPrevStat.total) : undefined;
|
||||||
|
const vPrevPct = vPrevStat?.total ? pct(vPrevStat.correct, vPrevStat.total) : undefined;
|
||||||
|
|
||||||
|
const tStr = t.total ? fmtPct(tPct) : "n/a";
|
||||||
|
const vStr = v.total ? fmtPct(vPct) : "n/a";
|
||||||
|
|
||||||
|
console.log(
|
||||||
|
field.padEnd(fieldCol) +
|
||||||
|
tStr.padStart(numCol) + fmtDelta(tPct, tPrevPct).padStart(deltaCol) + " " +
|
||||||
|
vStr.padStart(numCol) + fmtDelta(vPct, vPrevPct).padStart(deltaCol)
|
||||||
|
);
|
||||||
|
}
|
||||||
|
console.log("-".repeat(header.length));
|
||||||
|
|
||||||
|
const tOverallPct = pct(overallTrain.correct, overallTrain.total);
|
||||||
|
const vOverallPct = pct(overallVal.correct, overallVal.total);
|
||||||
|
const prevTrainOverall = overallFromFields(prev?.fields?.training);
|
||||||
|
const prevValOverall = overallFromFields(prev?.fields?.validation);
|
||||||
|
const tOverallPrevPct = prevTrainOverall?.total ? pct(prevTrainOverall.correct, prevTrainOverall.total) : undefined;
|
||||||
|
const vOverallPrevPct = prevValOverall?.total ? pct(prevValOverall.correct, prevValOverall.total) : undefined;
|
||||||
|
|
||||||
|
console.log(
|
||||||
|
"OVERALL".padEnd(fieldCol) +
|
||||||
|
fmtPct(tOverallPct).padStart(numCol) + fmtDelta(tOverallPct, tOverallPrevPct).padStart(deltaCol) + " " +
|
||||||
|
fmtPct(vOverallPct).padStart(numCol) + fmtDelta(vOverallPct, vOverallPrevPct).padStart(deltaCol)
|
||||||
|
);
|
||||||
|
|
||||||
|
if (failedImages.length) {
|
||||||
|
console.log(`\nFailed to parse: ${failedImages.join(", ")}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
// Informational only - DINOv2 "confidence" is a raw cosine similarity, not a
|
||||||
|
// calibrated probability (see docs/scan-product.md), so a delta here doesn't
|
||||||
|
// by itself mean better/worse. Only the fields above drive regression flags.
|
||||||
|
console.log("\nClassification (informational, not scored as pass/fail):");
|
||||||
|
printClassificationLine("Training", trainClassStats, prev?.classification?.training);
|
||||||
|
printClassificationLine("Validation", valClassStats, prev?.classification?.validation);
|
||||||
|
|
||||||
|
if (prev) {
|
||||||
|
const fieldRegressions: string[] = [];
|
||||||
|
const fieldImprovements: string[] = [];
|
||||||
|
for (const split of ["training", "validation"] as Split[]) {
|
||||||
|
const agg = split === "training" ? trainAgg : valAgg;
|
||||||
|
for (const field of FIELDS) {
|
||||||
|
const stat = agg[field];
|
||||||
|
if (!stat.total) continue;
|
||||||
|
const currPct = pct(stat.correct, stat.total);
|
||||||
|
const prevStat = prev.fields?.[split]?.[field];
|
||||||
|
if (!prevStat?.total) continue;
|
||||||
|
const prevPct = pct(prevStat.correct, prevStat.total);
|
||||||
|
const d = currPct - prevPct;
|
||||||
|
const label = `${field} (${split})`;
|
||||||
|
if (d <= -0.05) fieldRegressions.push(`${label} ${fmtDelta(currPct, prevPct)}`);
|
||||||
|
else if (d >= 0.05) fieldImprovements.push(`${label} ${fmtDelta(currPct, prevPct)}`);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if (fieldRegressions.length) console.log(`\nField regressions: ${fieldRegressions.join(", ")}`);
|
||||||
|
if (fieldImprovements.length) console.log(`Field improvements: ${fieldImprovements.join(", ")}`);
|
||||||
|
|
||||||
|
const imageRegressions: string[] = [];
|
||||||
|
const imageImprovements: string[] = [];
|
||||||
|
for (const [filename, currPct] of Object.entries(perImagePct)) {
|
||||||
|
const prevPct = prev.perImage?.[filename];
|
||||||
|
if (prevPct === undefined) continue;
|
||||||
|
const d = currPct - prevPct;
|
||||||
|
if (d <= -0.5) imageRegressions.push(`${filename} ${fmtDelta(currPct, prevPct)}`);
|
||||||
|
else if (d >= 0.5) imageImprovements.push(`${filename} ${fmtDelta(currPct, prevPct)}`);
|
||||||
|
}
|
||||||
|
if (imageRegressions.length) console.log(`\nImage regressions: ${imageRegressions.join(", ")}`);
|
||||||
|
if (imageImprovements.length) console.log(`Image improvements: ${imageImprovements.join(", ")}`);
|
||||||
|
|
||||||
|
console.log(`\n(vs run at ${prev.timestamp}${prev.commit !== "unknown" ? `, commit ${prev.commit}` : ""})`);
|
||||||
|
} else {
|
||||||
|
console.log("\n(no previous run in product_accuracy_history.jsonl — this is the baseline)");
|
||||||
|
}
|
||||||
|
console.log("");
|
||||||
|
}
|
||||||
|
|
||||||
async function main() {
|
async function main() {
|
||||||
const args = parseArgs(process.argv.slice(2));
|
const args = parseArgs(process.argv.slice(2));
|
||||||
await checkServerReachable(args.baseUrl);
|
await checkServerReachable(args.baseUrl);
|
||||||
@@ -116,12 +331,14 @@ async function main() {
|
|||||||
process.exit(1);
|
process.exit(1);
|
||||||
}
|
}
|
||||||
const labels: GroundTruth[] = JSON.parse(fs.readFileSync(LABELS_PATH, "utf8"));
|
const labels: GroundTruth[] = JSON.parse(fs.readFileSync(LABELS_PATH, "utf8"));
|
||||||
|
const prev = loadLastHistoryEntry();
|
||||||
|
|
||||||
const results = {
|
const results = {
|
||||||
training: [] as { gt: GroundTruth; checks: Check[] }[],
|
training: [] as ResultItem[],
|
||||||
validation: [] as { gt: GroundTruth; checks: Check[] }[],
|
validation: [] as ResultItem[],
|
||||||
failed: [] as string[]
|
failed: [] as string[]
|
||||||
};
|
};
|
||||||
|
const perImagePct: Record<string, number> = {};
|
||||||
|
|
||||||
for (const gt of labels) {
|
for (const gt of labels) {
|
||||||
if (gt.filename.startsWith("uploaded-")) continue; // Skip phantom
|
if (gt.filename.startsWith("uploaded-")) continue; // Skip phantom
|
||||||
@@ -146,13 +363,21 @@ async function main() {
|
|||||||
{ field: "expiry_date", match: isMatch(gt.expiry_date, predictedExpiry) }
|
{ field: "expiry_date", match: isMatch(gt.expiry_date, predictedExpiry) }
|
||||||
];
|
];
|
||||||
|
|
||||||
|
const item: ResultItem = {
|
||||||
|
gt,
|
||||||
|
checks,
|
||||||
|
method: parsed.classification?.method,
|
||||||
|
confidence: parsed.classification?.top1_confidence
|
||||||
|
};
|
||||||
|
|
||||||
if (gt.filename.includes("/")) {
|
if (gt.filename.includes("/")) {
|
||||||
results.training.push({ gt, checks });
|
results.training.push(item);
|
||||||
} else {
|
} else {
|
||||||
results.validation.push({ gt, checks });
|
results.validation.push(item);
|
||||||
}
|
}
|
||||||
|
|
||||||
const score = checks.filter(c => c.match).length;
|
const score = checks.filter(c => c.match).length;
|
||||||
|
perImagePct[gt.filename] = (score / checks.length) * 100;
|
||||||
console.log(`done (${score}/3)`);
|
console.log(`done (${score}/3)`);
|
||||||
} catch (err) {
|
} catch (err) {
|
||||||
console.log(`FAILED (${(err as Error).message})`);
|
console.log(`FAILED (${(err as Error).message})`);
|
||||||
@@ -160,42 +385,22 @@ async function main() {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
const aggregate = (list: { checks: Check[] }[]) => {
|
const trainAgg = aggregateFields(results.training);
|
||||||
const agg: Record<string, { correct: number; total: number }> = {
|
const valAgg = aggregateFields(results.validation);
|
||||||
no_sku: { correct: 0, total: 0 },
|
const trainClassStats = aggregateClassification(results.training);
|
||||||
nama_item: { correct: 0, total: 0 },
|
const valClassStats = aggregateClassification(results.validation);
|
||||||
expiry_date: { correct: 0, total: 0 }
|
const imageCount = { training: results.training.length, validation: results.validation.length };
|
||||||
};
|
|
||||||
for (const item of list) {
|
|
||||||
for (const check of item.checks) {
|
|
||||||
agg[check.field].total++;
|
|
||||||
if (check.match) agg[check.field].correct++;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
return agg;
|
|
||||||
};
|
|
||||||
|
|
||||||
const trainAgg = aggregate(results.training);
|
printSummary(trainAgg, valAgg, trainClassStats, valClassStats, perImagePct, prev, imageCount, results.failed);
|
||||||
const valAgg = aggregate(results.validation);
|
|
||||||
|
|
||||||
console.log("\n=== Product Scan Accuracy Summary ===");
|
|
||||||
console.log(`Training Images: ${results.training.length} | Validation Images: ${results.validation.length} | Failed: ${results.failed.length}\n`);
|
|
||||||
|
|
||||||
console.log("Field | Training Set | Validation Set");
|
|
||||||
console.log("---------------|--------------|---------------");
|
|
||||||
["no_sku", "nama_item", "expiry_date"].forEach(f => {
|
|
||||||
const t = trainAgg[f].total ? fmtPct(pct(trainAgg[f].correct, trainAgg[f].total)) : "n/a";
|
|
||||||
const v = valAgg[f].total ? fmtPct(pct(valAgg[f].correct, valAgg[f].total)) : "n/a";
|
|
||||||
console.log(`${f.padEnd(14)} | ${t.padEnd(12)} | ${v.padEnd(14)}`);
|
|
||||||
});
|
|
||||||
console.log("");
|
|
||||||
|
|
||||||
const entry: HistoryEntry = {
|
const entry: HistoryEntry = {
|
||||||
timestamp: new Date().toISOString(),
|
timestamp: new Date().toISOString(),
|
||||||
commit: getGitCommit(),
|
commit: getGitCommit(),
|
||||||
imageCount: { training: results.training.length, validation: results.validation.length },
|
imageCount,
|
||||||
failedImages: results.failed,
|
failedImages: results.failed,
|
||||||
fields: { training: trainAgg, validation: valAgg }
|
fields: { training: trainAgg, validation: valAgg },
|
||||||
|
classification: { training: trainClassStats, validation: valClassStats },
|
||||||
|
perImage: perImagePct
|
||||||
};
|
};
|
||||||
fs.appendFileSync(HISTORY_PATH, JSON.stringify(entry) + "\n");
|
fs.appendFileSync(HISTORY_PATH, JSON.stringify(entry) + "\n");
|
||||||
}
|
}
|
||||||
|
|||||||
|
After Width: | Height: | Size: 3.4 MiB |
|
After Width: | Height: | Size: 2.3 MiB |
|
After Width: | Height: | Size: 3.1 MiB |
|
After Width: | Height: | Size: 164 KiB |
|
After Width: | Height: | Size: 285 KiB |
|
After Width: | Height: | Size: 2.4 MiB |
|
After Width: | Height: | Size: 2.5 MiB |
|
After Width: | Height: | Size: 215 KiB |
|
After Width: | Height: | Size: 3.1 MiB |
|
After Width: | Height: | Size: 140 KiB |
|
After Width: | Height: | Size: 1.3 MiB |
|
After Width: | Height: | Size: 257 KiB |
|
After Width: | Height: | Size: 3.7 MiB |
|
After Width: | Height: | Size: 229 KiB |
|
After Width: | Height: | Size: 254 KiB |
|
After Width: | Height: | Size: 268 KiB |
|
After Width: | Height: | Size: 4.2 MiB |
|
After Width: | Height: | Size: 184 KiB |
|
After Width: | Height: | Size: 220 KiB |
|
After Width: | Height: | Size: 227 KiB |
|
After Width: | Height: | Size: 170 KiB |
|
After Width: | Height: | Size: 217 KiB |
|
After Width: | Height: | Size: 217 KiB |
|
After Width: | Height: | Size: 201 KiB |
|
After Width: | Height: | Size: 3.9 MiB |
|
After Width: | Height: | Size: 239 KiB |
|
After Width: | Height: | Size: 194 KiB |
|
After Width: | Height: | Size: 200 KiB |
|
After Width: | Height: | Size: 207 KiB |
|
After Width: | Height: | Size: 275 KiB |
|
After Width: | Height: | Size: 231 KiB |
|
After Width: | Height: | Size: 304 KiB |
|
After Width: | Height: | Size: 3.3 MiB |
|
After Width: | Height: | Size: 278 KiB |
|
After Width: | Height: | Size: 2.0 MiB |
|
After Width: | Height: | Size: 178 KiB |
|
After Width: | Height: | Size: 293 KiB |
|
After Width: | Height: | Size: 284 KiB |
|
After Width: | Height: | Size: 223 KiB |
|
After Width: | Height: | Size: 3.4 MiB |
|
After Width: | Height: | Size: 3.4 MiB |
|
After Width: | Height: | Size: 1.8 MiB |
|
After Width: | Height: | Size: 225 KiB |
|
After Width: | Height: | Size: 4.5 MiB |
|
After Width: | Height: | Size: 232 KiB |
|
After Width: | Height: | Size: 3.6 MiB |
|
After Width: | Height: | Size: 235 KiB |
|
After Width: | Height: | Size: 2.4 MiB |
|
After Width: | Height: | Size: 149 KiB |
|
After Width: | Height: | Size: 3.6 MiB |
|
After Width: | Height: | Size: 3.8 MiB |
|
After Width: | Height: | Size: 3.7 MiB |
|
After Width: | Height: | Size: 106 KiB |
|
After Width: | Height: | Size: 2.5 MiB |
|
After Width: | Height: | Size: 175 KiB |
|
After Width: | Height: | Size: 2.9 MiB |
|
After Width: | Height: | Size: 208 KiB |
|
After Width: | Height: | Size: 143 KiB |
|
After Width: | Height: | Size: 251 KiB |
|
After Width: | Height: | Size: 176 KiB |
|
After Width: | Height: | Size: 230 KiB |
|
After Width: | Height: | Size: 2.3 MiB |
|
After Width: | Height: | Size: 302 KiB |
|
After Width: | Height: | Size: 114 KiB |
|
After Width: | Height: | Size: 217 KiB |
|
After Width: | Height: | Size: 208 KiB |
|
After Width: | Height: | Size: 129 KiB |
|
After Width: | Height: | Size: 288 KiB |
|
After Width: | Height: | Size: 4.1 MiB |
|
After Width: | Height: | Size: 232 KiB |
|
After Width: | Height: | Size: 286 KiB |
|
After Width: | Height: | Size: 3.8 MiB |
|
After Width: | Height: | Size: 171 KiB |
|
After Width: | Height: | Size: 265 KiB |
@@ -0,0 +1,21 @@
|
|||||||
|
# Product-scan validation images
|
||||||
|
|
||||||
|
This folder is the **validation/test set** for the product-scan accuracy
|
||||||
|
harness (`backend/scripts/accuracy-check-scan.mts`) — real-world photos that
|
||||||
|
are *not* part of the classifier's reference dataset, so scoring against them
|
||||||
|
measures actual accuracy instead of memorization.
|
||||||
|
|
||||||
|
**Workflow:**
|
||||||
|
1. Drop a photo here directly (flat, no subfolders — a filename with no `/`
|
||||||
|
is what marks an image as "validation" instead of "training").
|
||||||
|
2. Label it via the `/manual-label-scan` page (correct `no_sku`, `nama_item`,
|
||||||
|
`expiry_date` by hand — don't just accept the AI-scan prefill, that would
|
||||||
|
make the ground truth equal to the model's own prediction).
|
||||||
|
3. Run `node scripts/accuracy-check-scan.mts` from `backend/` — the photo now
|
||||||
|
scores under "Validation Set", separate from "Training Set".
|
||||||
|
|
||||||
|
**This is not where new training photos go.** To improve the classifier
|
||||||
|
itself (DINOv2 index / YOLO fine-tune), add photos to
|
||||||
|
`pfm-web-app/public/produk-pfm/foto-kemasan-v2/<SKU folder>/` instead, then
|
||||||
|
reindex/retrain per `docs/scan-product.md`'s "Model artifacts & retraining"
|
||||||
|
section.
|
||||||