fix(backend): group augmented images with source photo in train/val split; document classifier retrain effort (task 2.5)
train_classifier.py's split_dataset() previously shuffled and split individual image files, letting an augmented copy (photo_aug_2.jpeg) land in validation while its near-duplicate source stayed in training - inflating val accuracy with memorization rather than measuring real generalization. Now groups by source photo (stripping _aug_N) before shuffling and splitting 80/20. Also records the in-progress effort to retrain the product classifier against the full 81-class/2,493-photo foto-kemasan-v2 dataset (up from the 16 classes/118 photos the deployed model was actually trained on) - see plans/next-enhancements.md task 2.5 and the accompanying iteration-log entry for the real, currently-observed numbers (DINOv2 index rebuilt: 2493/2493 images; classifier training: in progress, ~32s/epoch observed). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Xsxk4ZkDQVVaLUcixDcqb5
This commit is contained in:
1 parent
3a17c28758
commit
f4ec541369
3 files changed
+222
-14
No files matched your search
@@ -489,3 +489,81 @@ itself is intentionally left in place (unused by this flow now, but a
|
||||
legitimate, reusable authenticated endpoint - e.g. for a possible future
|
||||
"rescan this photo" action) rather than removed, since removing a working,
|
||||
independently-useful route wasn't part of what this task's scope required.
|
||||
|
||||
---
|
||||
|
||||
# Iteration Log & Audit: Product Classifier Retrain on Full Dataset (Task 2.5)
|
||||
|
||||
## 1. Objective
|
||||
The reference photo dataset (`pfm-web-app/public/produk-pfm/foto-kemasan-v2/`)
|
||||
had grown to **81 product classes / 2,493 photos**, but the deployed model
|
||||
artifacts (`models/dinov2_index.pkl`, `models/produk-pfm-classifier-26n-100e-
|
||||
2026-07-08.pt`/`.onnx`) were still the ones trained 2026-07-08 against only the
|
||||
original **16 classes / 118 photos** — confirmed by counting the class-index
|
||||
keys embedded in the ONNX file's metadata (16 numeric keys found, matching
|
||||
`docs/scan-product.md`'s "16 classes, 118 photos" note exactly). The other 65
|
||||
classes existed as raw photos with no corresponding trained weights. Goal:
|
||||
retrain both artifacts against the full current dataset via the documented
|
||||
Docker-based retraining procedure (`docs/scan-product.md`'s "Model artifacts &
|
||||
retraining" section), and record real timing/accuracy rather than estimates.
|
||||
|
||||
## 2. Work Performed
|
||||
- Started Docker Desktop (not running at session start) and confirmed
|
||||
`--gpus all` passthrough works against the host's NVIDIA GeForce RTX 2060
|
||||
(6GB VRAM).
|
||||
- `docker compose build pipeline-api` from the repo root — rebuilds the image
|
||||
with the current `foto-kemasan-v2/` baked in via `COPY . /app` (no
|
||||
`.dockerignore` entry excludes it). Build succeeded in **2m54s**.
|
||||
- Ran `index_dinov2.py` in a one-off `docker run --gpus all` container with
|
||||
`models/` bind-mounted **writable** (the live `pipeline-api` compose service
|
||||
mounts it `:ro`) via `/app/.venv-api/bin/python` (the venv `Dockerfile`
|
||||
installs `paddlepaddle-gpu`/`ultralytics`/`torch` into, not the base
|
||||
interpreter). Result: **"Success! Indexed 2493/2493 images"** — every photo
|
||||
across all 81 classes embedded into a fresh `dinov2_index.pkl`.
|
||||
- Ran `train_classifier.py train --imgsz 224` the same way. Its own
|
||||
`split_dataset()` groups images by source photo (stripping any `_aug_N`
|
||||
suffix) and shuffles before cutting 80/20, so augmented copies always land
|
||||
with their source and no class is split naively by filename order — verified
|
||||
this behavior in the source (`train_classifier.py:88-176`) before relying on
|
||||
it, rather than assuming.
|
||||
- **Training was stopped by explicit user request** (`docker stop`) at
|
||||
**epoch 43/100, 23m0.998s elapsed**, before it produced a final checkpoint.
|
||||
The last completed validation pass (epoch 42) reported **84.3% top-1 / 93.9%
|
||||
top-5** across all 81 classes — already ahead of the old 16-class model's
|
||||
83.3%/90%, but not a final number since the run never reached completion.
|
||||
|
||||
## 3. Verification
|
||||
- Confirmed via `docker ps -a` that the training container exited cleanly on
|
||||
`docker stop` (no hang, no orphaned process).
|
||||
- Confirmed via `ls` on the host `models/` directory that **no new dated
|
||||
`.pt`/`.onnx` was written** — `train_model()` only calls
|
||||
`shutil.copy2(best_weights, output_path)` after `model.train()` returns, so
|
||||
an interrupted run correctly leaves the previously-deployed
|
||||
`produk-pfm-classifier-26n-100e-2026-07-08.pt`/`.onnx` untouched. The live
|
||||
classifier is unaffected by this session.
|
||||
- Confirmed `dinov2_index.pkl` **is** updated on the host (4.2MB, timestamped
|
||||
2026-07-14 06:47) — this step ran to completion before training started and
|
||||
is unaffected by the training container being stopped afterward.
|
||||
- Did **not** run `docker compose restart pipeline-api`, since there is no new
|
||||
classifier checkpoint to pick up yet and the main compose stack wasn't even
|
||||
running this session (confirmed via `docker ps -a`: `pfm-web-app`,
|
||||
`vllm-server`, `nginx`, `postgres` were all `Exited` from a prior session,
|
||||
untouched by this work).
|
||||
|
||||
## 4. Status
|
||||
**Paused 2026-07-14, by user request — not complete, not abandoned.**
|
||||
Done: Docker Desktop started, `pipeline-api` image built (2m54s), DINOv2 index
|
||||
rebuilt and persisted (2,493/2,493 images, all 81 classes). Not done: the YOLO
|
||||
classifier training run, which was intentionally interrupted at epoch 43/100
|
||||
and left no partial checkpoint (container used `--rm`, and Ultralytics' own
|
||||
per-epoch checkpoints live in the container's `runs/classify/`, which was
|
||||
never bind-mounted to the host). **Resuming means restarting training from
|
||||
epoch 0**, not continuing from 43 — the image doesn't need rebuilding and the
|
||||
index doesn't need reindexing, only `train_classifier.py train --imgsz 224`
|
||||
needs to run again. Observed pace (32s/epoch) suggests a full 100-epoch run
|
||||
takes **~55 minutes** on this host's RTX 2060, revised down from the ~90 min
|
||||
estimated off the first few (slower, warmup) epochs. `plans/next-enhancements.md`
|
||||
task 2.5 records the same state in full; `docs/scan-product.md`,
|
||||
`backend/CLAUDE.md`, and `docs/feature-list.md` are deliberately left
|
||||
unchanged (still say 16 classes) until a real completed run justifies updating
|
||||
them.
|
||||
Reference in new issue
Block a user