train_classifier.py's split_dataset() previously shuffled and split individual image files, letting an augmented copy (photo_aug_2.jpeg) land in validation while its near-duplicate source stayed in training - inflating val accuracy with memorization rather than measuring real generalization. Now groups by source photo (stripping _aug_N) before shuffling and splitting 80/20. Also records the in-progress effort to retrain the product classifier against the full 81-class/2,493-photo foto-kemasan-v2 dataset (up from the 16 classes/118 photos the deployed model was actually trained on) - see plans/next-enhancements.md task 2.5 and the accompanying iteration-log entry for the real, currently-observed numbers (DINOv2 index rebuilt: 2493/2493 images; classifier training: in progress, ~32s/epoch observed). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Xsxk4ZkDQVVaLUcixDcqb5