# Entity Relationship Diagram (ERD) This document specifies the SQLite database schema and entity relationships for the **reTraining** platform (`app.db`), as defined in [backend/db.py](file:///C:/Users/araar/Downloads/PT_SIAB_FULLTIME/Feedmill_Semarang/reTraining/backend/db.py). --- ## Architectural Storage Model The platform uses a **hybrid storage architecture**: * **SQLite Database (`data/app.db`)**: Stores entity metadata, relations, job queues, triage rules, review states, and metrics. * **Disk Filesystem (`data/projects/`)**: Stores image pixels (`.jpg`), YOLO labels (`.txt`), YAML dataset manifests (`data.yaml`), and trained model weights (`.pt`). --- ## Mermaid Entity Relationship Diagram ```mermaid erDiagram PROJECTS ||--o{ PROJECT_CLASSES : "defines" PROJECTS ||--o{ BATCHES : "contains" PROJECTS ||--o{ DATASETS : "compiles" PROJECTS ||--o{ DATASET_ITEMS : "aggregates" PROJECTS ||--o{ BASE_DATASETS : "mounts" PROJECTS ||--o{ TRIAGE_RULES : "configures" PROJECTS ||--o{ MODEL_VERSIONS : "produces" PROJECTS ||--o{ JOBS : "executes" PROJECTS ||--o{ VIDEO_CLOCK : "indexes" PROJECTS ||--o{ COUNT_RUNS : "benchmarks" BATCHES ||--o{ FRAMES : "extracts" BATCHES ||--o{ JOBS : "triggers" FRAMES ||--o{ ANNOTATIONS : "annotates" FRAMES ||--o{ DATASET_ITEMS : "includes" ANNOTATIONS ||--o| ANNOTATION_OVERRIDES : "overrides" DATASETS ||--o{ DATASET_ITEMS : "contains" PROJECTS { integer id PK text slug UK text name text label_type "bbox | polygon" text base_model_path text base_model_kind "uploaded | pretrained | trained" text video_root integer val_every "default: 5" real created_at } PROJECT_CLASSES { integer id PK integer project_id FK integer class_id text name text prompt } BATCHES { integer id PK integer project_id FK text video_path text date_label text batch_label real start_sec real end_sec real fps text status "extracting | extracted | labeling | reviewing | approved | merged | failed" integer frame_count real created_at real merged_at } FRAMES { integer id PK integer batch_id FK integer idx text filename integer width integer height text review_status "pending | approved | rejected" } ANNOTATIONS { integer id PK integer frame_id FK integer class_id text geometry "JSON / coordinates" real score text source "auto | manual" real created_at } ANNOTATION_OVERRIDES { integer annotation_id PK, FK text verdict "keep | ignore | reclass" integer target_class real decided_at } DATASETS { integer id PK integer project_id FK text name text note text rule_version text rules_json real created_at } DATASET_ITEMS { integer id PK integer project_id FK integer dataset_id FK integer frame_id FK text split "train | val" text image_rel text label_rel real added_at } BASE_DATASETS { integer id PK integer project_id FK text name text source integer image_count integer box_count text classes "JSON array" real created_at } TRIAGE_RULES { integer id PK integer project_id FK text stage "default: dataprep" integer position text name text predicate text action "keep | ignore | reclass" integer target_class real created_at } MODEL_VERSIONS { integer id PK integer project_id FK integer version text weights_path text parent_model_path text metrics "JSON" text base_metrics "JSON" real created_at } JOBS { integer id PK integer project_id FK integer batch_id FK text type "extract | autolabel | merge | train | count | clock-scan | truck-scan" text status "queued | running | done | failed | cancelled" text params "JSON" integer progress integer total text message text error text log real created_at real started_at real finished_at } VIDEO_CLOCK { integer id PK integer project_id FK text video_rel text folder_date text started_at "YYYY-MM-DD HH:MM:SS" text working_day real confidence integer agreeing text source "ocr" text error real read_at } COUNT_RUNS { integer id PK integer project_id FK text video_rel text date_label text batch_label integer loading integer unloading integer net integer ground_truth integer frames real seconds text params "JSON" text model_path text error real counted_at } ``` --- ## Entity Descriptions ### 1. Project & Class Configuration * **`projects`**: Core isolation entity. Configures the target label type (`bbox` vs `polygon`), base model reference, video archive root, and validation split step (`val_every`). * **`project_classes`**: Class definitions tied to a project. Stores integer `class_id`, class `name`, and natural language SAM3 zero-shot `prompt`. ### 2. Video Extraction & Annotation Pipeline * **`batches`**: A trimmed segment from a raw CCTV video file. Holds time ranges, sampling FPS, frame counts, and extraction/review lifecycles. * **`frames`**: Extracted image stills belonging to a batch, tracking width, height, index, and operator approval status (`pending`, `approved`, `rejected`). * **`annotations`**: Bounding boxes or polygon geometries per frame with confidence scores, class IDs, and origin (`auto` from SAM3 vs `manual` from operator). * **`annotation_overrides`**: Per-annotation triage verdicts (`keep`, `ignore`, `reclass`) resulting from outlier inspection. ### 3. Master Datasets & Data Preparation * **`datasets`**: Immutable dataset compilations (`v1`, `v2`, `v3`) capturing the snapshot rules and timestamps. * **`dataset_items`**: Mapping between a dataset and its constituent image frames, recording deterministic train/val splits (`split IN ('train', 'val')`) and relative filesystem paths. * **`base_datasets`**: External train-only datasets imported to supplement training data without affecting validation splits. * **`triage_rules`**: Ordered filter predicates applied during data prep to systematically prune bounding box anomalies (e.g. area, aspect ratio, confidence). ### 4. Training, Jobs & Telemetry * **`model_versions`**: Trained YOLO checkpoints (`1`, `2`, `3`...) storing weights paths, parent models, and side-by-side metric evaluations ($\Delta\text{mAP50}$, precision, recall). * **`jobs`**: Asynchronous background job queue (`extract`, `autolabel`, `train`, `count`, etc.) managing progress counters, logs, and state transitions. ### 5. Video Clock & Production Counting Benchmarks * **`video_clock`**: OCR timestamps extracted from CCTV video overlays to assign recordings to proper 24-hour work shifts (06:00 to 06:00 cutoff). * **`count_runs`**: Inference benchmarks running ByteTrack line-crossing counters against verified physical ground truth numbers. --- ## Database Indexes To maintain sub-second UI performance across thousands of frames and annotations, the following indices are maintained: * `idx_frames_batch` on `frames(batch_id, idx)` * `idx_annotations_frame` on `annotations(frame_id)` * `idx_batches_project` on `batches(project_id)` * `idx_dataset_items_project` on `dataset_items(project_id)` * `idx_triage_rules_project` on `triage_rules(project_id, stage, position)` * `idx_jobs_project` on `jobs(project_id, created_at)` * `idx_video_clock_project` on `video_clock(project_id, working_day)` * `idx_count_runs_project` on `count_runs(project_id, date_label)`