fix(backend): group augmented images with source photo in train/val split; document classifier retrain effort (task 2.5)

train_classifier.py's split_dataset() previously shuffled and split
individual image files, letting an augmented copy (photo_aug_2.jpeg) land
in validation while its near-duplicate source stayed in training -
inflating val accuracy with memorization rather than measuring real
generalization. Now groups by source photo (stripping _aug_N) before
shuffling and splitting 80/20.

Also records the in-progress effort to retrain the product classifier
against the full 81-class/2,493-photo foto-kemasan-v2 dataset (up from the
16 classes/118 photos the deployed model was actually trained on) - see
plans/next-enhancements.md task 2.5 and the accompanying iteration-log
entry for the real, currently-observed numbers (DINOv2 index rebuilt:
2493/2493 images; classifier training: in progress, ~32s/epoch observed).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xsxk4ZkDQVVaLUcixDcqb5
This commit is contained in:
Rafhan Mazaya FathurrahmanandClaude Sonnet 5 committed 2026-07-14 08:34:36 +07:00
1 parent 3a17c28758
commit f4ec541369
3 files changed
+222 -14

No files matched your search

+78
View File
@@ -489,3 +489,81 @@ itself is intentionally left in place (unused by this flow now, but a
legitimate, reusable authenticated endpoint - e.g. for a possible future legitimate, reusable authenticated endpoint - e.g. for a possible future
"rescan this photo" action) rather than removed, since removing a working, "rescan this photo" action) rather than removed, since removing a working,
independently-useful route wasn't part of what this task's scope required. independently-useful route wasn't part of what this task's scope required.
---
# Iteration Log & Audit: Product Classifier Retrain on Full Dataset (Task 2.5)
## 1. Objective
The reference photo dataset (`pfm-web-app/public/produk-pfm/foto-kemasan-v2/`)
had grown to **81 product classes / 2,493 photos**, but the deployed model
artifacts (`models/dinov2_index.pkl`, `models/produk-pfm-classifier-26n-100e-
2026-07-08.pt`/`.onnx`) were still the ones trained 2026-07-08 against only the
original **16 classes / 118 photos** — confirmed by counting the class-index
keys embedded in the ONNX file's metadata (16 numeric keys found, matching
`docs/scan-product.md`'s "16 classes, 118 photos" note exactly). The other 65
classes existed as raw photos with no corresponding trained weights. Goal:
retrain both artifacts against the full current dataset via the documented
Docker-based retraining procedure (`docs/scan-product.md`'s "Model artifacts &
retraining" section), and record real timing/accuracy rather than estimates.
## 2. Work Performed
- Started Docker Desktop (not running at session start) and confirmed
`--gpus all` passthrough works against the host's NVIDIA GeForce RTX 2060
(6GB VRAM).
- `docker compose build pipeline-api` from the repo root — rebuilds the image
with the current `foto-kemasan-v2/` baked in via `COPY . /app` (no
`.dockerignore` entry excludes it). Build succeeded in **2m54s**.
- Ran `index_dinov2.py` in a one-off `docker run --gpus all` container with
`models/` bind-mounted **writable** (the live `pipeline-api` compose service
mounts it `:ro`) via `/app/.venv-api/bin/python` (the venv `Dockerfile`
installs `paddlepaddle-gpu`/`ultralytics`/`torch` into, not the base
interpreter). Result: **"Success! Indexed 2493/2493 images"** — every photo
across all 81 classes embedded into a fresh `dinov2_index.pkl`.
- Ran `train_classifier.py train --imgsz 224` the same way. Its own
`split_dataset()` groups images by source photo (stripping any `_aug_N`
suffix) and shuffles before cutting 80/20, so augmented copies always land
with their source and no class is split naively by filename order — verified
this behavior in the source (`train_classifier.py:88-176`) before relying on
it, rather than assuming.
- **Training was stopped by explicit user request** (`docker stop`) at
**epoch 43/100, 23m0.998s elapsed**, before it produced a final checkpoint.
The last completed validation pass (epoch 42) reported **84.3% top-1 / 93.9%
top-5** across all 81 classes — already ahead of the old 16-class model's
83.3%/90%, but not a final number since the run never reached completion.
## 3. Verification
- Confirmed via `docker ps -a` that the training container exited cleanly on
`docker stop` (no hang, no orphaned process).
- Confirmed via `ls` on the host `models/` directory that **no new dated
`.pt`/`.onnx` was written** — `train_model()` only calls
`shutil.copy2(best_weights, output_path)` after `model.train()` returns, so
an interrupted run correctly leaves the previously-deployed
`produk-pfm-classifier-26n-100e-2026-07-08.pt`/`.onnx` untouched. The live
classifier is unaffected by this session.
- Confirmed `dinov2_index.pkl` **is** updated on the host (4.2MB, timestamped
2026-07-14 06:47) — this step ran to completion before training started and
is unaffected by the training container being stopped afterward.
- Did **not** run `docker compose restart pipeline-api`, since there is no new
classifier checkpoint to pick up yet and the main compose stack wasn't even
running this session (confirmed via `docker ps -a`: `pfm-web-app`,
`vllm-server`, `nginx`, `postgres` were all `Exited` from a prior session,
untouched by this work).
## 4. Status
**Paused 2026-07-14, by user request — not complete, not abandoned.**
Done: Docker Desktop started, `pipeline-api` image built (2m54s), DINOv2 index
rebuilt and persisted (2,493/2,493 images, all 81 classes). Not done: the YOLO
classifier training run, which was intentionally interrupted at epoch 43/100
and left no partial checkpoint (container used `--rm`, and Ultralytics' own
per-epoch checkpoints live in the container's `runs/classify/`, which was
never bind-mounted to the host). **Resuming means restarting training from
epoch 0**, not continuing from 43 — the image doesn't need rebuilding and the
index doesn't need reindexing, only `train_classifier.py train --imgsz 224`
needs to run again. Observed pace (32s/epoch) suggests a full 100-epoch run
takes **~55 minutes** on this host's RTX 2060, revised down from the ~90 min
estimated off the first few (slower, warmup) epochs. `plans/next-enhancements.md`
task 2.5 records the same state in full; `docs/scan-product.md`,
`backend/CLAUDE.md`, and `docs/feature-list.md` are deliberately left
unchanged (still say 16 classes) until a real completed run justifies updating
them.
@@ -73,16 +73,28 @@ def latest_classifier_weights(models_dir: Path = DEFAULT_MODELS_DIR) -> Path:
return max(candidates, key=sort_key) return max(candidates, key=sort_key)
VALID_IMAGE_EXTENSIONS = {".jpg", ".jpeg", ".png", ".webp", ".bmp"} VALID_IMAGE_EXTENSIONS = {".jpg", ".jpeg", ".png", ".webp", ".bmp"}
AUG_SUFFIX_RE = re.compile(r"_aug_\d+$")
def is_image_file(path: Path) -> bool: def is_image_file(path: Path) -> bool:
return path.is_file() and path.suffix.lower() in VALID_IMAGE_EXTENSIONS return path.is_file() and path.suffix.lower() in VALID_IMAGE_EXTENSIONS
def _source_group_key(filename_stem: str) -> str:
"""Strip an `_aug_<n>` suffix so an augmented image groups with its source photo."""
return AUG_SUFFIX_RE.sub("", filename_stem)
def split_dataset(src_dir: Path, dest_dir: Path, split_ratio: float = 0.8, seed: int = 42): def split_dataset(src_dir: Path, dest_dir: Path, split_ratio: float = 0.8, seed: int = 42):
""" """
Split class folders from src_dir into train/val folders in dest_dir. Split class folders from src_dir into train/val folders in dest_dir.
Ensures every class with 2+ images keeps at least one image in validation. Ensures every class with 2+ images keeps at least one image in validation.
Splits by *source photo group*, not by individual file: an augmented image
(`photo1_aug_2.jpeg`) always stays in the same split as its source
(`photo1.jpeg`). Splitting file-by-file would let near-duplicate images
land on opposite sides of train/val, inflating val accuracy with
memorization instead of measuring generalization.
""" """
random.seed(seed) random.seed(seed)
@@ -111,29 +123,41 @@ def split_dataset(src_dir: Path, dest_dir: Path, split_ratio: float = 0.8, seed:
[f for f in c_dir.iterdir() if is_image_file(f)], [f for f in c_dir.iterdir() if is_image_file(f)],
key=lambda p: p.name, key=lambda p: p.name,
) )
random.shuffle(images)
num_images = len(images) num_images = len(images)
if num_images == 0: if num_images == 0:
print(f"Warning: Class '{class_name}' has 0 images. Skipping.") print(f"Warning: Class '{class_name}' has 0 images. Skipping.")
continue continue
# Group by source photo (stripping any `_aug_N` suffix) so an
# augmented image and the photo it came from always land on the same
# side of the split.
groups: dict[str, list[Path]] = {}
for img in images:
groups.setdefault(_source_group_key(img.stem), []).append(img)
group_keys = sorted(groups.keys())
random.shuffle(group_keys)
class_train_dir = train_dir / class_name class_train_dir = train_dir / class_name
class_val_dir = val_dir / class_name class_val_dir = val_dir / class_name
class_train_dir.mkdir(parents=True, exist_ok=True) class_train_dir.mkdir(parents=True, exist_ok=True)
class_val_dir.mkdir(parents=True, exist_ok=True) class_val_dir.mkdir(parents=True, exist_ok=True)
if num_images == 1: num_groups = len(group_keys)
train_images = images if num_groups == 1:
val_images = images train_groups = group_keys
elif num_images == 2: val_groups = group_keys
train_images = [images[0]] elif num_groups == 2:
val_images = [images[1]] train_groups = [group_keys[0]]
val_groups = [group_keys[1]]
else: else:
split_idx = max(1, int(num_images * split_ratio)) split_idx = max(1, int(num_groups * split_ratio))
split_idx = min(split_idx, num_images - 1) split_idx = min(split_idx, num_groups - 1)
train_images = images[:split_idx] train_groups = group_keys[:split_idx]
val_images = images[split_idx:] val_groups = group_keys[split_idx:]
train_images = [img for key in train_groups for img in groups[key]]
val_images = [img for key in val_groups for img in groups[key]]
for img in train_images: for img in train_images:
shutil.copy(img, class_train_dir / img.name) shutil.copy(img, class_train_dir / img.name)
@@ -145,7 +169,7 @@ def split_dataset(src_dir: Path, dest_dir: Path, split_ratio: float = 0.8, seed:
print( print(
f" Class '{class_name}': {len(train_images)} train, " f" Class '{class_name}': {len(train_images)} train, "
f"{len(val_images)} val (total: {num_images})" f"{len(val_images)} val (from {num_groups} source photos, {num_images} files total)"
) )
print(f"Dataset split completed: {total_train} train images, {total_val} validation images.") print(f"Dataset split completed: {total_train} train images, {total_val} validation images.")
+108 -2
View File
@@ -113,6 +113,62 @@ in either project (only SKU, product name, expiry date are extracted) — if
requested later, follow the same OCR-regex-cascade pattern already used for requested later, follow the same OCR-regex-cascade pattern already used for
expiry-date extraction.* expiry-date extraction.*
- **2.5** [IN PROGRESS 2026-07-14 — resumed, training run 2] **Retrain
classifier on the now-81-class dataset.** (Note: a first resume attempt
failed instantly with a Docker daemon connection error — Docker Desktop had
stopped between sessions — before any training happened; restarted Docker
Desktop and relaunched. This is the actual second training attempt,
confirmed running via `docker ps`.) `foto-kemasan-v2/` grew from the 16
classes/118 photos the deployed model
(`produk-pfm-classifier-26n-100e-2026-07-08.pt`) was trained on to **81
classes / 2,493 photos** — the other 65 classes were never included in any
training run.
- **Goal**: retrain both artifacts (`dinov2_index.pkl` similarity index and the
YOLO classifier) against the full current dataset so the deployed model
actually recognizes all 81 SKU folders, not just the original 16.
- **Agreed procedure** (per `docs/scan-product.md`'s documented retraining
steps — training must run via Docker, not bare-metal Windows, since
`paddlepaddle-gpu` wheels are Linux-only): from repo root,
`docker compose build pipeline-api` (bakes in the current dataset) → one-off
`docker run --gpus all` with `models/` mounted **writable** (the live
compose service mounts it `:ro`) → `index_dinov2.py` (rebuilds the DINOv2
index) → `train_classifier.py train --imgsz 224` (its own `split_dataset()`
does an 80/20 split grouped by source photo and shuffled — not a naive
first-N-files split, so augmented copies always land with their source) →
`docker compose restart pipeline-api` → verify via `docker logs` for
"DINOv2 index loaded with N reference images" and "Using classifier
weights: <new dated file>".
- **Status as of pause (2026-07-14)** — mixed state, read carefully before
resuming:
- ✅ `pipeline-api` image built (2m54s), bakes in the current 81-class
dataset.
- ✅ **`dinov2_index.pkl` already rebuilt and persisted to disk** —
"Success! Indexed 2493/2493 images" across all 81 classes. This artifact
is live on the host now (`models/dinov2_index.pkl`, 4.2MB, dated
2026-07-14) and does **not** need to be redone.
- ⏸️ **YOLO classifier training was started, then stopped by user request
at epoch 43/100 (~23 minutes in)** before it could write a new dated
checkpoint. `docker run` used `--rm` and the in-progress epoch
checkpoints live only in the container's own `runs/classify/` (not
bind-mounted), so **stopping the container discarded that partial
progress** — resuming means restarting from epoch 0, not continuing from
43. `models/` on the host still has only the original
`produk-pfm-classifier-26n-100e-2026-07-08.pt`/`.onnx` (16-class model) —
**the live/deployed classifier is unchanged**, still 16 classes.
- Observed pace before stopping: ~32s/epoch (43 epochs in 23m1s) → a full
100-epoch run should take **~55 minutes** on this host's RTX 2060 (6GB
VRAM), not the ~90 min extrapolated from the first few (slower, warmup)
epochs. At epoch 42 the in-progress run had already reached 84.3%
top-1 / 93.9% top-5 val accuracy across all 81 classes, ahead of the old
16-class model's 83.3%/90% — a promising sign for the eventual full run,
but not a final result since training didn't finish.
- **To resume**: image is already built and the DINOv2 index step can be
skipped — just re-run the one-off `train_classifier.py train --imgsz 224`
container, then `docker compose restart pipeline-api` and verify via
`docker logs`. Update the class count in `docs/scan-product.md`,
`CLAUDE.md`, and `docs/feature-list.md` (and flip this task to `[DONE]`)
only once that run actually completes with a final dated `.pt`/`.onnx`.
## 3. Backend — Postgres Data Layer ## 3. Backend — Postgres Data Layer
`pfm-web-app/src/db/` `pfm-web-app/src/db/`
@@ -179,9 +235,15 @@ surface for product scans**, mirroring what the DO flow already has in
- **8.3** [DONE 2026-07-08] Build an `/admin/master-data` web UI to visually manage both SKUs and Stores. (See docs/feature-list.md) - **8.3** [DONE 2026-07-08] Build an `/admin/master-data` web UI to visually manage both SKUs and Stores. (See docs/feature-list.md)
- **6.2** [DONE 2026-07-08] API + storage groundwork for scan annotation. (See docs/feature-list.md) - **6.2** [DONE 2026-07-08] API + storage groundwork for scan annotation. (See docs/feature-list.md)
- **6.3** [DONE 2026-07-08] Product-scan accuracy harness created. (See docs/feature-list.md) - **6.3** [DONE 2026-07-08] Product-scan accuracy harness created. (See docs/feature-list.md)
- **6.4** [DONE 2026-07-13] Auto-diff-vs-previous-run reporting (ported from the
DO-flow's `accuracy-check.mts`) plus classifier method/confidence tracking
added to `accuracy-check-scan.mts`; created the missing
`sources/product-test-images/` validation-photo folder. User-directed `n`
request: wanted to tune the scan algorithm and see improvement/regression
automatically instead of eyeballing two flat runs. (See docs/feature-list.md)
*Suggested order: 6.2 → 6.1 → 6.3 (storage/API first, page on top, harness once *Suggested order: 6.2 → 6.1 → 6.3 → 6.4 (storage/API first, page on top, harness
labels exist in volume).* once labels exist in volume, diffing once the harness has history to diff against).*
## 7. Auth — Store Accounts & Profile-Sourced Metadata ## 7. Auth — Store Accounts & Profile-Sourced Metadata
`pfm-web-app/src/db/init.ts`, `api/v1/auth/*`, `api/parse/route.ts`, `sources/toko_aktif.json` `pfm-web-app/src/db/init.ts`, `api/v1/auth/*`, `api/parse/route.ts`, `sources/toko_aktif.json`
@@ -350,6 +412,50 @@ Flutter root `plans/next-enhancements.md` §7.2.
root `docs/iteration-log.md` for the Flutter-side verification that the root `docs/iteration-log.md` for the Flutter-side verification that the
editor renders this without a second network call. editor renders this without a second network call.
## 12. Backend — Stock Management
`src/db/init-stock.ts`, `src/app/api/v1/stock/`, `src/utils/stock-*.ts`,
`src/app/api/parse/route.ts`, `src/app/api/v1/documents/[id]/route.ts`,
`src/app/admin/master-data/`
Added 2026-07-10, backend counterpart to root `plans/next-enhancements.md` §9
(Flutter Stocks Menu & DO-to-Stock Flow) — both sections originated from the same
user-directed, extensively grilled ad-hoc feature request (not an `e`/`enhance`
section — see `AGENTS.md` Part B7). **Read
[../../docs/stock-feature-plan.md](../../docs/stock-feature-plan.md) first** — full
schema, API contracts, and sequencing for both sides. **Status: planned, not yet
implemented** — no code for this feature exists in the codebase yet.
- **12.1** [TODO] **Stock schema + core CRUD.** New `src/db/init-stock.ts`
(`stock_batches` — unique per `(kode_toko, no_sku, batch_code, expiry_date)`,
tracks both outer and inner qty; `stock_movements` — append-only audit log,
`intake`/`decrement`/`adjustment`/`manual_seed`), wired into `init.ts`. New
`src/utils/stock-mapper.ts`, `src/utils/stock-movement.ts` (`recordStockMovement`
only, for this task). New routes: `src/app/api/v1/stock/route.ts` (GET summary
per SKU, POST create/merge-by-unique-key), `stock/[noSku]/route.ts` (GET batch
detail), `stock/batches/[id]/route.ts` (PUT edit, logged as `adjustment`). No
DELETE route — batches are edit-only, never removed.
- **12.2** [TODO] **Product Scan decrement hook + in-stock candidate filter.**
Extend `stock-movement.ts` with `decrementBatchForProductScan` (row-locked,
allowed to go negative, logged as `decrement`); wire into
`documents/[id]/route.ts`'s existing PUT transaction, gated on `scan_mode ===
'Product'` and a new `stock_batch_id` payload field — an invalid/cross-store
batch id fails the **whole confirm** (400), never a silent skip (per user's
explicit answer during grilling). New `src/utils/stock-lookup.ts`
(`getInStockSkuSet`/`filterMatchesByStock`), applied to `possibleMatches` in
both `api/parse/route.ts`'s Product branch and `api/v1/scan-product/route.ts` —
**do not change `classifyAndMatchProduct()`'s signature**, it's shared with the
anonymous store-agnostic desktop dev route; filter at the two authenticated call
sites instead. Blocked on 12.1.
- **12.3** [TODO] **Read-only admin Stock view.** `admin/master-data/page.tsx`
(419 lines, already over the 256-line threshold) split into `page.tsx` (shell)
+ extracted `StoreManager.tsx` + `SkuManager.tsx` (pure extraction, no behavior
change) + new `StockManager.tsx` (all-stores table via `GET /api/v1/stock` with
no `kode_toko` param as admin; row click drills into batch detail). Blocked on
12.1.
*Suggested order: 12.1 → 12.2 (needs 12.1's tables/movement helper) and 12.3
(needs 12.1's summary route) — 12.2/12.3 are independent of each other.*
--- ---
*Sections 1-4 migrated 2026-07-08 from root `plans/next-enhancements.md` sections *Sections 1-4 migrated 2026-07-08 from root `plans/next-enhancements.md` sections