# Entity Relationship Diagram (ERD) This document provides the complete Entity Relationship Diagram (ERD) and relational schema for the **reTraining** platform SQLite database (`app.db`), as implemented in [`backend/db.py`](file:///home/asus/feedmill_semarang_project/reTraining/backend/db.py). --- ## Architectural Storage Model The platform uses a **hybrid storage architecture**: * **SQLite Database (`data/app.db`)**: Stores relational metadata, foreign keys, job queues, triage rules, review states, model versions, and benchmark run telemetry. * **Disk Filesystem (`data/projects//`)**: Stores raw/extracted image pixels (`.jpg`), YOLO annotation labels (`.txt`), immutable dataset definitions (`data.yaml`), and trained neural network checkpoints (`.pt`). --- ## Mermaid Entity Relationship Diagram ```mermaid erDiagram PROJECTS ||--o{ PROJECT_CLASSES : "defines" PROJECTS ||--o{ BATCHES : "contains" PROJECTS ||--o{ DATASETS : "compiles" PROJECTS ||--o{ DATASET_ITEMS : "aggregates" PROJECTS ||--o{ BASE_DATASETS : "mounts" PROJECTS ||--o{ TRIAGE_RULES : "configures" PROJECTS ||--o{ MODEL_VERSIONS : "produces" PROJECTS ||--o{ JOBS : "executes" PROJECTS ||--o{ VIDEO_CLOCK : "indexes" PROJECTS ||--o{ COUNT_RUNS : "benchmarks" BATCHES ||--o{ FRAMES : "extracts" BATCHES ||--o{ JOBS : "triggers" FRAMES ||--o{ ANNOTATIONS : "annotates" FRAMES ||--o{ DATASET_ITEMS : "includes" ANNOTATIONS ||--o| ANNOTATION_OVERRIDES : "overrides" DATASETS ||--o{ DATASET_ITEMS : "contains" PROJECTS { integer id PK "AUTOINCREMENT" text slug UK "Unique project identifier" text name "Display title" text label_type "bbox | polygon" text base_model_path "Path to initial .pt model" text base_model_kind "uploaded | pretrained | trained" text video_root "Root path for video archive" integer val_every "Validation split step (default 5)" text secondary_model_path "Optional companion model" text secondary_model_name "Companion model label" text secondary_model_classes "Companion model class JSON" text augment "Augmentation parameters JSON" real created_at "Unix epoch timestamp" } PROJECT_CLASSES { integer id PK "AUTOINCREMENT" integer project_id FK "References projects(id)" integer class_id "YOLO integer class index" text name "Class display name" text prompt "SAM3 natural language zero-shot prompt" } BATCHES { integer id PK "AUTOINCREMENT" integer project_id FK "References projects(id)" text video_path "Source video relative path" text date_label "Video date directory" text batch_label "Video file base name" real start_sec "Trim range start in seconds" real end_sec "Trim range end in seconds" real fps "Extraction sampling rate" text status "extracting | extracted | labeling | reviewing | approved | merged | failed" integer frame_count "Extracted frames count" real created_at "Unix epoch timestamp" real merged_at "Timestamp when batch merged to dataset" } FRAMES { integer id PK "AUTOINCREMENT" integer batch_id FK "References batches(id)" integer idx "Frame index within batch" text filename "Image file name (000001.jpg)" integer width "Frame pixel width" integer height "Frame pixel height" text review_status "pending | approved | rejected" } ANNOTATIONS { integer id PK "AUTOINCREMENT" integer frame_id FK "References frames(id)" integer class_id "Target YOLO class index" text geometry "Box or Polygon coordinates JSON" real score "SAM3 confidence score (0.0 - 1.0)" text source "auto | manual" real created_at "Unix epoch timestamp" } ANNOTATION_OVERRIDES { integer annotation_id PK, FK "References annotations(id)" text verdict "keep | ignore | reclass" integer target_class "New class index if reclassified" real decided_at "Unix epoch timestamp" } DATASETS { integer id PK "AUTOINCREMENT" integer project_id FK "References projects(id)" text name "Dataset version name (e.g. Master v1)" text note "Release description" text rule_version "Triage rule version identifier" text rules_json "Frozen triage rules JSON snapshot" real created_at "Unix epoch timestamp" } DATASET_ITEMS { integer id PK "AUTOINCREMENT" integer project_id FK "References projects(id)" integer dataset_id FK "References datasets(id)" integer frame_id FK "References frames(id)" text split "train | val" text image_rel "Relative image path in dataset" text label_rel "Relative label path in dataset" real added_at "Unix epoch timestamp" } BASE_DATASETS { integer id PK "AUTOINCREMENT" integer project_id FK "References projects(id)" text name "External baseline dataset name" text source "Import origin description" integer image_count "Number of baseline images" integer box_count "Number of baseline bounding boxes" text classes "JSON array of class names" real created_at "Unix epoch timestamp" } TRIAGE_RULES { integer id PK "AUTOINCREMENT" integer project_id FK "References projects(id)" text stage "Filter stage (default: dataprep)" integer position "Rule execution order position" text name "Rule human readable name" text predicate "Evaluation expression (area, aspect, conf)" text action "keep | ignore | reclass" integer target_class "Target class index if reclass" real created_at "Unix epoch timestamp" } MODEL_VERSIONS { integer id PK "AUTOINCREMENT" integer project_id FK "References projects(id)" integer version "Incrementing version integer" text name "Descriptive model name (arch-labelType-epochs-classes-date)" text weights_path "Path to trained best.pt weights" text parent_model_path "Path to base model used as starting point" text metrics "Trained model evaluation metrics JSON" text base_metrics "Base model evaluation metrics JSON" text rule_version "Triage rule version used for training" text augment "Augmentation settings JSON used" real created_at "Unix epoch timestamp" } JOBS { integer id PK "AUTOINCREMENT" integer project_id FK "References projects(id)" integer batch_id FK "References batches(id)" text type "extract | autolabel | merge | train | count | clock-scan | truck-scan" text status "queued | running | done | failed | cancelled" text params "Job parameters JSON" integer progress "Current completed units" integer total "Total units of work" text message "Human-readable status update" text error "Failure message or traceback" text log "Detailed execution log lines" real created_at "Unix epoch timestamp" real started_at "Job start timestamp" real finished_at "Job completion timestamp" } VIDEO_CLOCK { integer id PK "AUTOINCREMENT" integer project_id FK "References projects(id)" text video_rel "Relative path to video file" text folder_date "Date folder extracted from path" text started_at "Burned-in OCR clock string (YYYY-MM-DD HH:MM:SS)" text working_day "Operational shift day (06:00 - 05:59 cutoff)" real confidence "OCR reading confidence" integer agreeing "Number of sample frames agreeing" text source "ocr | manual" text error "OCR parsing error message" real read_at "Timestamp when OCR scan was executed" integer truck_hits "Frames with truck detected" integer truck_samples "Total sampled frames for truck check" text truck_model "Model used for truck scan" real truck_checked_at "Timestamp of truck scan" } COUNT_RUNS { integer id PK "AUTOINCREMENT" integer project_id FK "References projects(id)" text video_rel "Relative path to video evaluated" text date_label "Video date directory" text batch_label "Video file base name" integer loading "Counted loading direction crossings" integer unloading "Counted unloading direction crossings" integer net "Calculated net count (loading - unloading)" integer ground_truth "Physical verified hand-tally count" integer frames "Total video frames evaluated" real seconds "Inference elapsed time in seconds" text params "Counting line & ByteTrack parameters JSON" text model_path "YOLO model weights path used" text error "Error message if benchmark failed" real counted_at "Unix epoch timestamp" } ``` --- ## Entity Descriptions ### 1. Workspace & Multi-Project Isolation * **`projects`**: Top-level entity isolating datasets, classes, models, and CCTV archives. Enforces geometry type (`bbox` vs `polygon`) and stores the active base model checkpoint. * **`project_classes`**: YOLO class definitions for the project. Pairs each integer `class_id` with natural language text prompts utilized by SAM3 zero-shot auto-annotation. ### 2. Video Ingest & Annotation Lifecycle * **`batches`**: A trimmed segment of raw industrial CCTV footage. Captures start/end timestamps, sampling FPS, frame counts, and progression from extraction to dataset merge. * **`frames`**: Discrete JPEG stills extracted from a batch. Tracks individual review status (`pending`, `approved`, `rejected`). * **`annotations`**: Object bounding boxes or polygon masks per frame with SAM3 / manual confidence scores and class assignments. * **`annotation_overrides`**: Outlier inspection decisions (`keep`, `ignore`, `reclass`) resulting from interactive 2D scatter plot triage. ### 3. Master Datasets & Data Preparation * **`datasets`**: Immutable, versioned master datasets (`v1`, `v2`, `v3`). Captures a permanent JSON snapshot of triage rules (`rules_json`) applied at compilation time. * **`dataset_items`**: Junction entity binding frames into dataset versions with deterministic SHA-1 validation split partitioning (`train` vs `val`). * **`base_datasets`**: External pre-labeled baseline datasets mounted as train-only supplements without polluting validation benchmarks. * **`triage_rules`**: Ordered filtering predicates (e.g., box area, aspect ratio, confidence thresholds) applied during data preparation. ### 4. Continuous YOLO Retraining & Background Workers * **`model_versions`**: Fine-tuned YOLO weights checkpoints. Stores side-by-side performance deltas ($\Delta\text{mAP50}$, $\Delta\text{Precision}$, $\Delta\text{Recall}$) evaluated against the identical frozen validation split. * **`jobs`**: Centralized SQLite job queue for asynchronous background tasks (`extract`, `autolabel`, `train`, `count`, etc.) with atomic status management and progress streaming. ### 5. Shift OCR Indexing & Production Line Counter * **`video_clock`**: OCR timestamps extracted from burned-in CCTV camera overlays, mapping recordings to 24-hour manufacturing shifts (06:00 AM to 05:59 AM next day) and detecting truck arrival presence. * **`count_runs`**: Real-time production inference benchmarks running ByteTrack line-crossing counters against verified ground truth tallies. --- ## Performance & Database Indexes The SQLite database enforces the following indexes to maintain sub-millisecond query performance: | Index Name | Table | Indexed Columns | Purpose | |---|---|---|---| | `idx_frames_batch` | `frames` | `(batch_id, idx)` | Rapid frame lookup and sequential canvas scrubbing | | `idx_annotations_frame` | `annotations` | `(frame_id)` | Sub-millisecond bounding box loading per canvas frame | | `idx_batches_project` | `batches` | `(project_id)` | Fast batch library filtering by project | | `idx_dataset_items_project`| `dataset_items` | `(project_id)` | Dataset compilation and split integrity checks | | `idx_triage_rules_project` | `triage_rules` | `(project_id, stage, position)` | Fast ordered triage predicate evaluation | | `idx_jobs_project` | `jobs` | `(project_id, created_at)` | UI job queue monitoring and polling | | `idx_video_clock_project` | `video_clock` | `(project_id, working_day)` | Shift-based video library filtering | | `idx_count_runs_project` | `count_runs` | `(project_id, date_label)` | Production counting benchmark reporting |