Files
pfm-ocr/backend/pfm-web-app/public
Rafhan Mazaya FathurrahmanandClaude Sonnet 5 f4ec541369 fix(backend): group augmented images with source photo in train/val split; document classifier retrain effort (task 2.5)
train_classifier.py's split_dataset() previously shuffled and split
individual image files, letting an augmented copy (photo_aug_2.jpeg) land
in validation while its near-duplicate source stayed in training -
inflating val accuracy with memorization rather than measuring real
generalization. Now groups by source photo (stripping _aug_N) before
shuffling and splitting 80/20.

Also records the in-progress effort to retrain the product classifier
against the full 81-class/2,493-photo foto-kemasan-v2 dataset (up from the
16 classes/118 photos the deployed model was actually trained on) - see
plans/next-enhancements.md task 2.5 and the accompanying iteration-log
entry for the real, currently-observed numbers (DINOv2 index rebuilt:
2493/2493 images; classifier training: in progress, ~32s/epoch observed).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xsxk4ZkDQVVaLUcixDcqb5
2026-07-14 08:34:36 +07:00
..