fix(backend): group augmented images with source photo in train/val split; document classifier retrain effort (task 2.5)

train_classifier.py's split_dataset() previously shuffled and split
individual image files, letting an augmented copy (photo_aug_2.jpeg) land
in validation while its near-duplicate source stayed in training -
inflating val accuracy with memorization rather than measuring real
generalization. Now groups by source photo (stripping _aug_N) before
shuffling and splitting 80/20.

Also records the in-progress effort to retrain the product classifier
against the full 81-class/2,493-photo foto-kemasan-v2 dataset (up from the
16 classes/118 photos the deployed model was actually trained on) - see
plans/next-enhancements.md task 2.5 and the accompanying iteration-log
entry for the real, currently-observed numbers (DINOv2 index rebuilt:
2493/2493 images; classifier training: in progress, ~32s/epoch observed).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xsxk4ZkDQVVaLUcixDcqb5
This commit is contained in:
Rafhan Mazaya FathurrahmanandClaude Sonnet 5 committed 2026-07-14 08:34:36 +07:00
1 parent 3a17c28758
commit f4ec541369
3 files changed
+222 -14

No files matched your search

+78
View File
@@ -489,3 +489,81 @@ itself is intentionally left in place (unused by this flow now, but a
legitimate, reusable authenticated endpoint - e.g. for a possible future
"rescan this photo" action) rather than removed, since removing a working,
independently-useful route wasn't part of what this task's scope required.
---
# Iteration Log & Audit: Product Classifier Retrain on Full Dataset (Task 2.5)
## 1. Objective
The reference photo dataset (`pfm-web-app/public/produk-pfm/foto-kemasan-v2/`)
had grown to **81 product classes / 2,493 photos**, but the deployed model
artifacts (`models/dinov2_index.pkl`, `models/produk-pfm-classifier-26n-100e-
2026-07-08.pt`/`.onnx`) were still the ones trained 2026-07-08 against only the
original **16 classes / 118 photos** — confirmed by counting the class-index
keys embedded in the ONNX file's metadata (16 numeric keys found, matching
`docs/scan-product.md`'s "16 classes, 118 photos" note exactly). The other 65
classes existed as raw photos with no corresponding trained weights. Goal:
retrain both artifacts against the full current dataset via the documented
Docker-based retraining procedure (`docs/scan-product.md`'s "Model artifacts &
retraining" section), and record real timing/accuracy rather than estimates.
## 2. Work Performed
- Started Docker Desktop (not running at session start) and confirmed
`--gpus all` passthrough works against the host's NVIDIA GeForce RTX 2060
(6GB VRAM).
- `docker compose build pipeline-api` from the repo root — rebuilds the image
with the current `foto-kemasan-v2/` baked in via `COPY . /app` (no
`.dockerignore` entry excludes it). Build succeeded in **2m54s**.
- Ran `index_dinov2.py` in a one-off `docker run --gpus all` container with
`models/` bind-mounted **writable** (the live `pipeline-api` compose service
mounts it `:ro`) via `/app/.venv-api/bin/python` (the venv `Dockerfile`
installs `paddlepaddle-gpu`/`ultralytics`/`torch` into, not the base
interpreter). Result: **"Success! Indexed 2493/2493 images"** — every photo
across all 81 classes embedded into a fresh `dinov2_index.pkl`.
- Ran `train_classifier.py train --imgsz 224` the same way. Its own
`split_dataset()` groups images by source photo (stripping any `_aug_N`
suffix) and shuffles before cutting 80/20, so augmented copies always land
with their source and no class is split naively by filename order — verified
this behavior in the source (`train_classifier.py:88-176`) before relying on
it, rather than assuming.
- **Training was stopped by explicit user request** (`docker stop`) at
**epoch 43/100, 23m0.998s elapsed**, before it produced a final checkpoint.
The last completed validation pass (epoch 42) reported **84.3% top-1 / 93.9%
top-5** across all 81 classes — already ahead of the old 16-class model's
83.3%/90%, but not a final number since the run never reached completion.
## 3. Verification
- Confirmed via `docker ps -a` that the training container exited cleanly on
`docker stop` (no hang, no orphaned process).
- Confirmed via `ls` on the host `models/` directory that **no new dated
`.pt`/`.onnx` was written** — `train_model()` only calls
`shutil.copy2(best_weights, output_path)` after `model.train()` returns, so
an interrupted run correctly leaves the previously-deployed
`produk-pfm-classifier-26n-100e-2026-07-08.pt`/`.onnx` untouched. The live
classifier is unaffected by this session.
- Confirmed `dinov2_index.pkl` **is** updated on the host (4.2MB, timestamped
2026-07-14 06:47) — this step ran to completion before training started and
is unaffected by the training container being stopped afterward.
- Did **not** run `docker compose restart pipeline-api`, since there is no new
classifier checkpoint to pick up yet and the main compose stack wasn't even
running this session (confirmed via `docker ps -a`: `pfm-web-app`,
`vllm-server`, `nginx`, `postgres` were all `Exited` from a prior session,
untouched by this work).
## 4. Status
**Paused 2026-07-14, by user request — not complete, not abandoned.**
Done: Docker Desktop started, `pipeline-api` image built (2m54s), DINOv2 index
rebuilt and persisted (2,493/2,493 images, all 81 classes). Not done: the YOLO
classifier training run, which was intentionally interrupted at epoch 43/100
and left no partial checkpoint (container used `--rm`, and Ultralytics' own
per-epoch checkpoints live in the container's `runs/classify/`, which was
never bind-mounted to the host). **Resuming means restarting training from
epoch 0**, not continuing from 43 — the image doesn't need rebuilding and the
index doesn't need reindexing, only `train_classifier.py train --imgsz 224`
needs to run again. Observed pace (32s/epoch) suggests a full 100-epoch run
takes **~55 minutes** on this host's RTX 2060, revised down from the ~90 min
estimated off the first few (slower, warmup) epochs. `plans/next-enhancements.md`
task 2.5 records the same state in full; `docs/scan-product.md`,
`backend/CLAUDE.md`, and `docs/feature-list.md` are deliberately left
unchanged (still say 16 classes) until a real completed run justifies updating
them.