diff --git a/backend/docs/iteration-log.md b/backend/docs/iteration-log.md index 28fbc18..a9b4cd4 100644 --- a/backend/docs/iteration-log.md +++ b/backend/docs/iteration-log.md @@ -489,3 +489,81 @@ itself is intentionally left in place (unused by this flow now, but a legitimate, reusable authenticated endpoint - e.g. for a possible future "rescan this photo" action) rather than removed, since removing a working, independently-useful route wasn't part of what this task's scope required. + +--- + +# Iteration Log & Audit: Product Classifier Retrain on Full Dataset (Task 2.5) + +## 1. Objective +The reference photo dataset (`pfm-web-app/public/produk-pfm/foto-kemasan-v2/`) +had grown to **81 product classes / 2,493 photos**, but the deployed model +artifacts (`models/dinov2_index.pkl`, `models/produk-pfm-classifier-26n-100e- +2026-07-08.pt`/`.onnx`) were still the ones trained 2026-07-08 against only the +original **16 classes / 118 photos** — confirmed by counting the class-index +keys embedded in the ONNX file's metadata (16 numeric keys found, matching +`docs/scan-product.md`'s "16 classes, 118 photos" note exactly). The other 65 +classes existed as raw photos with no corresponding trained weights. Goal: +retrain both artifacts against the full current dataset via the documented +Docker-based retraining procedure (`docs/scan-product.md`'s "Model artifacts & +retraining" section), and record real timing/accuracy rather than estimates. + +## 2. Work Performed +- Started Docker Desktop (not running at session start) and confirmed + `--gpus all` passthrough works against the host's NVIDIA GeForce RTX 2060 + (6GB VRAM). +- `docker compose build pipeline-api` from the repo root — rebuilds the image + with the current `foto-kemasan-v2/` baked in via `COPY . /app` (no + `.dockerignore` entry excludes it). Build succeeded in **2m54s**. +- Ran `index_dinov2.py` in a one-off `docker run --gpus all` container with + `models/` bind-mounted **writable** (the live `pipeline-api` compose service + mounts it `:ro`) via `/app/.venv-api/bin/python` (the venv `Dockerfile` + installs `paddlepaddle-gpu`/`ultralytics`/`torch` into, not the base + interpreter). Result: **"Success! Indexed 2493/2493 images"** — every photo + across all 81 classes embedded into a fresh `dinov2_index.pkl`. +- Ran `train_classifier.py train --imgsz 224` the same way. Its own + `split_dataset()` groups images by source photo (stripping any `_aug_N` + suffix) and shuffles before cutting 80/20, so augmented copies always land + with their source and no class is split naively by filename order — verified + this behavior in the source (`train_classifier.py:88-176`) before relying on + it, rather than assuming. +- **Training was stopped by explicit user request** (`docker stop`) at + **epoch 43/100, 23m0.998s elapsed**, before it produced a final checkpoint. + The last completed validation pass (epoch 42) reported **84.3% top-1 / 93.9% + top-5** across all 81 classes — already ahead of the old 16-class model's + 83.3%/90%, but not a final number since the run never reached completion. + +## 3. Verification +- Confirmed via `docker ps -a` that the training container exited cleanly on + `docker stop` (no hang, no orphaned process). +- Confirmed via `ls` on the host `models/` directory that **no new dated + `.pt`/`.onnx` was written** — `train_model()` only calls + `shutil.copy2(best_weights, output_path)` after `model.train()` returns, so + an interrupted run correctly leaves the previously-deployed + `produk-pfm-classifier-26n-100e-2026-07-08.pt`/`.onnx` untouched. The live + classifier is unaffected by this session. +- Confirmed `dinov2_index.pkl` **is** updated on the host (4.2MB, timestamped + 2026-07-14 06:47) — this step ran to completion before training started and + is unaffected by the training container being stopped afterward. +- Did **not** run `docker compose restart pipeline-api`, since there is no new + classifier checkpoint to pick up yet and the main compose stack wasn't even + running this session (confirmed via `docker ps -a`: `pfm-web-app`, + `vllm-server`, `nginx`, `postgres` were all `Exited` from a prior session, + untouched by this work). + +## 4. Status +**Paused 2026-07-14, by user request — not complete, not abandoned.** +Done: Docker Desktop started, `pipeline-api` image built (2m54s), DINOv2 index +rebuilt and persisted (2,493/2,493 images, all 81 classes). Not done: the YOLO +classifier training run, which was intentionally interrupted at epoch 43/100 +and left no partial checkpoint (container used `--rm`, and Ultralytics' own +per-epoch checkpoints live in the container's `runs/classify/`, which was +never bind-mounted to the host). **Resuming means restarting training from +epoch 0**, not continuing from 43 — the image doesn't need rebuilding and the +index doesn't need reindexing, only `train_classifier.py train --imgsz 224` +needs to run again. Observed pace (32s/epoch) suggests a full 100-epoch run +takes **~55 minutes** on this host's RTX 2060, revised down from the ~90 min +estimated off the first few (slower, warmup) epochs. `plans/next-enhancements.md` +task 2.5 records the same state in full; `docs/scan-product.md`, +`backend/CLAUDE.md`, and `docs/feature-list.md` are deliberately left +unchanged (still say 16 classes) until a real completed run justifies updating +them. diff --git a/backend/pfm-web-app/public/produk-pfm/train_classifier.py b/backend/pfm-web-app/public/produk-pfm/train_classifier.py index f4fdeed..6102aed 100644 --- a/backend/pfm-web-app/public/produk-pfm/train_classifier.py +++ b/backend/pfm-web-app/public/produk-pfm/train_classifier.py @@ -73,16 +73,28 @@ def latest_classifier_weights(models_dir: Path = DEFAULT_MODELS_DIR) -> Path: return max(candidates, key=sort_key) VALID_IMAGE_EXTENSIONS = {".jpg", ".jpeg", ".png", ".webp", ".bmp"} +AUG_SUFFIX_RE = re.compile(r"_aug_\d+$") def is_image_file(path: Path) -> bool: return path.is_file() and path.suffix.lower() in VALID_IMAGE_EXTENSIONS +def _source_group_key(filename_stem: str) -> str: + """Strip an `_aug_` suffix so an augmented image groups with its source photo.""" + return AUG_SUFFIX_RE.sub("", filename_stem) + + def split_dataset(src_dir: Path, dest_dir: Path, split_ratio: float = 0.8, seed: int = 42): """ Split class folders from src_dir into train/val folders in dest_dir. Ensures every class with 2+ images keeps at least one image in validation. + + Splits by *source photo group*, not by individual file: an augmented image + (`photo1_aug_2.jpeg`) always stays in the same split as its source + (`photo1.jpeg`). Splitting file-by-file would let near-duplicate images + land on opposite sides of train/val, inflating val accuracy with + memorization instead of measuring generalization. """ random.seed(seed) @@ -111,29 +123,41 @@ def split_dataset(src_dir: Path, dest_dir: Path, split_ratio: float = 0.8, seed: [f for f in c_dir.iterdir() if is_image_file(f)], key=lambda p: p.name, ) - random.shuffle(images) num_images = len(images) if num_images == 0: print(f"Warning: Class '{class_name}' has 0 images. Skipping.") continue + # Group by source photo (stripping any `_aug_N` suffix) so an + # augmented image and the photo it came from always land on the same + # side of the split. + groups: dict[str, list[Path]] = {} + for img in images: + groups.setdefault(_source_group_key(img.stem), []).append(img) + group_keys = sorted(groups.keys()) + random.shuffle(group_keys) + class_train_dir = train_dir / class_name class_val_dir = val_dir / class_name class_train_dir.mkdir(parents=True, exist_ok=True) class_val_dir.mkdir(parents=True, exist_ok=True) - if num_images == 1: - train_images = images - val_images = images - elif num_images == 2: - train_images = [images[0]] - val_images = [images[1]] + num_groups = len(group_keys) + if num_groups == 1: + train_groups = group_keys + val_groups = group_keys + elif num_groups == 2: + train_groups = [group_keys[0]] + val_groups = [group_keys[1]] else: - split_idx = max(1, int(num_images * split_ratio)) - split_idx = min(split_idx, num_images - 1) - train_images = images[:split_idx] - val_images = images[split_idx:] + split_idx = max(1, int(num_groups * split_ratio)) + split_idx = min(split_idx, num_groups - 1) + train_groups = group_keys[:split_idx] + val_groups = group_keys[split_idx:] + + train_images = [img for key in train_groups for img in groups[key]] + val_images = [img for key in val_groups for img in groups[key]] for img in train_images: shutil.copy(img, class_train_dir / img.name) @@ -145,7 +169,7 @@ def split_dataset(src_dir: Path, dest_dir: Path, split_ratio: float = 0.8, seed: print( f" Class '{class_name}': {len(train_images)} train, " - f"{len(val_images)} val (total: {num_images})" + f"{len(val_images)} val (from {num_groups} source photos, {num_images} files total)" ) print(f"Dataset split completed: {total_train} train images, {total_val} validation images.") diff --git a/backend/plans/next-enhancements.md b/backend/plans/next-enhancements.md index 256ab15..34ca8fe 100644 --- a/backend/plans/next-enhancements.md +++ b/backend/plans/next-enhancements.md @@ -113,6 +113,62 @@ in either project (only SKU, product name, expiry date are extracted) — if requested later, follow the same OCR-regex-cascade pattern already used for expiry-date extraction.* +- **2.5** [IN PROGRESS 2026-07-14 — resumed, training run 2] **Retrain + classifier on the now-81-class dataset.** (Note: a first resume attempt + failed instantly with a Docker daemon connection error — Docker Desktop had + stopped between sessions — before any training happened; restarted Docker + Desktop and relaunched. This is the actual second training attempt, + confirmed running via `docker ps`.) `foto-kemasan-v2/` grew from the 16 + classes/118 photos the deployed model + (`produk-pfm-classifier-26n-100e-2026-07-08.pt`) was trained on to **81 + classes / 2,493 photos** — the other 65 classes were never included in any + training run. + - **Goal**: retrain both artifacts (`dinov2_index.pkl` similarity index and the + YOLO classifier) against the full current dataset so the deployed model + actually recognizes all 81 SKU folders, not just the original 16. + - **Agreed procedure** (per `docs/scan-product.md`'s documented retraining + steps — training must run via Docker, not bare-metal Windows, since + `paddlepaddle-gpu` wheels are Linux-only): from repo root, + `docker compose build pipeline-api` (bakes in the current dataset) → one-off + `docker run --gpus all` with `models/` mounted **writable** (the live + compose service mounts it `:ro`) → `index_dinov2.py` (rebuilds the DINOv2 + index) → `train_classifier.py train --imgsz 224` (its own `split_dataset()` + does an 80/20 split grouped by source photo and shuffled — not a naive + first-N-files split, so augmented copies always land with their source) → + `docker compose restart pipeline-api` → verify via `docker logs` for + "DINOv2 index loaded with N reference images" and "Using classifier + weights: ". + - **Status as of pause (2026-07-14)** — mixed state, read carefully before + resuming: + - ✅ `pipeline-api` image built (2m54s), bakes in the current 81-class + dataset. + - ✅ **`dinov2_index.pkl` already rebuilt and persisted to disk** — + "Success! Indexed 2493/2493 images" across all 81 classes. This artifact + is live on the host now (`models/dinov2_index.pkl`, 4.2MB, dated + 2026-07-14) and does **not** need to be redone. + - ⏸️ **YOLO classifier training was started, then stopped by user request + at epoch 43/100 (~23 minutes in)** before it could write a new dated + checkpoint. `docker run` used `--rm` and the in-progress epoch + checkpoints live only in the container's own `runs/classify/` (not + bind-mounted), so **stopping the container discarded that partial + progress** — resuming means restarting from epoch 0, not continuing from + 43. `models/` on the host still has only the original + `produk-pfm-classifier-26n-100e-2026-07-08.pt`/`.onnx` (16-class model) — + **the live/deployed classifier is unchanged**, still 16 classes. + - Observed pace before stopping: ~32s/epoch (43 epochs in 23m1s) → a full + 100-epoch run should take **~55 minutes** on this host's RTX 2060 (6GB + VRAM), not the ~90 min extrapolated from the first few (slower, warmup) + epochs. At epoch 42 the in-progress run had already reached 84.3% + top-1 / 93.9% top-5 val accuracy across all 81 classes, ahead of the old + 16-class model's 83.3%/90% — a promising sign for the eventual full run, + but not a final result since training didn't finish. + - **To resume**: image is already built and the DINOv2 index step can be + skipped — just re-run the one-off `train_classifier.py train --imgsz 224` + container, then `docker compose restart pipeline-api` and verify via + `docker logs`. Update the class count in `docs/scan-product.md`, + `CLAUDE.md`, and `docs/feature-list.md` (and flip this task to `[DONE]`) + only once that run actually completes with a final dated `.pt`/`.onnx`. + ## 3. Backend — Postgres Data Layer `pfm-web-app/src/db/` @@ -179,9 +235,15 @@ surface for product scans**, mirroring what the DO flow already has in - **8.3** [DONE 2026-07-08] Build an `/admin/master-data` web UI to visually manage both SKUs and Stores. (See docs/feature-list.md) - **6.2** [DONE 2026-07-08] API + storage groundwork for scan annotation. (See docs/feature-list.md) - **6.3** [DONE 2026-07-08] Product-scan accuracy harness created. (See docs/feature-list.md) +- **6.4** [DONE 2026-07-13] Auto-diff-vs-previous-run reporting (ported from the + DO-flow's `accuracy-check.mts`) plus classifier method/confidence tracking + added to `accuracy-check-scan.mts`; created the missing + `sources/product-test-images/` validation-photo folder. User-directed `n` + request: wanted to tune the scan algorithm and see improvement/regression + automatically instead of eyeballing two flat runs. (See docs/feature-list.md) -*Suggested order: 6.2 → 6.1 → 6.3 (storage/API first, page on top, harness once -labels exist in volume).* +*Suggested order: 6.2 → 6.1 → 6.3 → 6.4 (storage/API first, page on top, harness +once labels exist in volume, diffing once the harness has history to diff against).* ## 7. Auth — Store Accounts & Profile-Sourced Metadata `pfm-web-app/src/db/init.ts`, `api/v1/auth/*`, `api/parse/route.ts`, `sources/toko_aktif.json` @@ -350,6 +412,50 @@ Flutter root `plans/next-enhancements.md` §7.2. root `docs/iteration-log.md` for the Flutter-side verification that the editor renders this without a second network call. +## 12. Backend — Stock Management +`src/db/init-stock.ts`, `src/app/api/v1/stock/`, `src/utils/stock-*.ts`, +`src/app/api/parse/route.ts`, `src/app/api/v1/documents/[id]/route.ts`, +`src/app/admin/master-data/` + +Added 2026-07-10, backend counterpart to root `plans/next-enhancements.md` §9 +(Flutter Stocks Menu & DO-to-Stock Flow) — both sections originated from the same +user-directed, extensively grilled ad-hoc feature request (not an `e`/`enhance` +section — see `AGENTS.md` Part B7). **Read +[../../docs/stock-feature-plan.md](../../docs/stock-feature-plan.md) first** — full +schema, API contracts, and sequencing for both sides. **Status: planned, not yet +implemented** — no code for this feature exists in the codebase yet. + +- **12.1** [TODO] **Stock schema + core CRUD.** New `src/db/init-stock.ts` + (`stock_batches` — unique per `(kode_toko, no_sku, batch_code, expiry_date)`, + tracks both outer and inner qty; `stock_movements` — append-only audit log, + `intake`/`decrement`/`adjustment`/`manual_seed`), wired into `init.ts`. New + `src/utils/stock-mapper.ts`, `src/utils/stock-movement.ts` (`recordStockMovement` + only, for this task). New routes: `src/app/api/v1/stock/route.ts` (GET summary + per SKU, POST create/merge-by-unique-key), `stock/[noSku]/route.ts` (GET batch + detail), `stock/batches/[id]/route.ts` (PUT edit, logged as `adjustment`). No + DELETE route — batches are edit-only, never removed. +- **12.2** [TODO] **Product Scan decrement hook + in-stock candidate filter.** + Extend `stock-movement.ts` with `decrementBatchForProductScan` (row-locked, + allowed to go negative, logged as `decrement`); wire into + `documents/[id]/route.ts`'s existing PUT transaction, gated on `scan_mode === + 'Product'` and a new `stock_batch_id` payload field — an invalid/cross-store + batch id fails the **whole confirm** (400), never a silent skip (per user's + explicit answer during grilling). New `src/utils/stock-lookup.ts` + (`getInStockSkuSet`/`filterMatchesByStock`), applied to `possibleMatches` in + both `api/parse/route.ts`'s Product branch and `api/v1/scan-product/route.ts` — + **do not change `classifyAndMatchProduct()`'s signature**, it's shared with the + anonymous store-agnostic desktop dev route; filter at the two authenticated call + sites instead. Blocked on 12.1. +- **12.3** [TODO] **Read-only admin Stock view.** `admin/master-data/page.tsx` + (419 lines, already over the 256-line threshold) split into `page.tsx` (shell) + + extracted `StoreManager.tsx` + `SkuManager.tsx` (pure extraction, no behavior + change) + new `StockManager.tsx` (all-stores table via `GET /api/v1/stock` with + no `kode_toko` param as admin; row click drills into batch detail). Blocked on + 12.1. + +*Suggested order: 12.1 → 12.2 (needs 12.1's tables/movement helper) and 12.3 +(needs 12.1's summary route) — 12.2/12.3 are independent of each other.* + --- *Sections 1-4 migrated 2026-07-08 from root `plans/next-enhancements.md` sections