feat: remap ports to 9000/9010, update documentation, and add Mermaid ERD
This commit is contained in:
1 parent
8ad35ed9d1
commit
cf4c3370e1
8 files changed
+457
-179
No files matched your search
+158
-145
@@ -1,14 +1,14 @@
|
||||
# Entity Relationship Diagram (ERD)
|
||||
|
||||
This document specifies the SQLite database schema and entity relationships for the **reTraining** platform (`app.db`), as defined in [backend/db.py](file:///C:/Users/araar/Downloads/PT_SIAB_FULLTIME/Feedmill_Semarang/reTraining/backend/db.py).
|
||||
This document provides the complete Entity Relationship Diagram (ERD) and relational schema for the **reTraining** platform SQLite database (`app.db`), as implemented in [`backend/db.py`](file:///home/asus/feedmill_semarang_project/reTraining/backend/db.py).
|
||||
|
||||
---
|
||||
|
||||
## Architectural Storage Model
|
||||
|
||||
The platform uses a **hybrid storage architecture**:
|
||||
* **SQLite Database (`data/app.db`)**: Stores entity metadata, relations, job queues, triage rules, review states, and metrics.
|
||||
* **Disk Filesystem (`data/projects/`)**: Stores image pixels (`.jpg`), YOLO labels (`.txt`), YAML dataset manifests (`data.yaml`), and trained model weights (`.pt`).
|
||||
* **SQLite Database (`data/app.db`)**: Stores relational metadata, foreign keys, job queues, triage rules, review states, model versions, and benchmark run telemetry.
|
||||
* **Disk Filesystem (`data/projects/<slug>/`)**: Stores raw/extracted image pixels (`.jpg`), YOLO annotation labels (`.txt`), immutable dataset definitions (`data.yaml`), and trained neural network checkpoints (`.pt`).
|
||||
|
||||
---
|
||||
|
||||
@@ -38,169 +38,179 @@ erDiagram
|
||||
DATASETS ||--o{ DATASET_ITEMS : "contains"
|
||||
|
||||
PROJECTS {
|
||||
integer id PK
|
||||
text slug UK
|
||||
text name
|
||||
integer id PK "AUTOINCREMENT"
|
||||
text slug UK "Unique project identifier"
|
||||
text name "Display title"
|
||||
text label_type "bbox | polygon"
|
||||
text base_model_path
|
||||
text base_model_path "Path to initial .pt model"
|
||||
text base_model_kind "uploaded | pretrained | trained"
|
||||
text video_root
|
||||
integer val_every "default: 5"
|
||||
real created_at
|
||||
text video_root "Root path for video archive"
|
||||
integer val_every "Validation split step (default 5)"
|
||||
text secondary_model_path "Optional companion model"
|
||||
text secondary_model_name "Companion model label"
|
||||
text secondary_model_classes "Companion model class JSON"
|
||||
text augment "Augmentation parameters JSON"
|
||||
real created_at "Unix epoch timestamp"
|
||||
}
|
||||
|
||||
PROJECT_CLASSES {
|
||||
integer id PK
|
||||
integer project_id FK
|
||||
integer class_id
|
||||
text name
|
||||
text prompt
|
||||
integer id PK "AUTOINCREMENT"
|
||||
integer project_id FK "References projects(id)"
|
||||
integer class_id "YOLO integer class index"
|
||||
text name "Class display name"
|
||||
text prompt "SAM3 natural language zero-shot prompt"
|
||||
}
|
||||
|
||||
BATCHES {
|
||||
integer id PK
|
||||
integer project_id FK
|
||||
text video_path
|
||||
text date_label
|
||||
text batch_label
|
||||
real start_sec
|
||||
real end_sec
|
||||
real fps
|
||||
integer id PK "AUTOINCREMENT"
|
||||
integer project_id FK "References projects(id)"
|
||||
text video_path "Source video relative path"
|
||||
text date_label "Video date directory"
|
||||
text batch_label "Video file base name"
|
||||
real start_sec "Trim range start in seconds"
|
||||
real end_sec "Trim range end in seconds"
|
||||
real fps "Extraction sampling rate"
|
||||
text status "extracting | extracted | labeling | reviewing | approved | merged | failed"
|
||||
integer frame_count
|
||||
real created_at
|
||||
real merged_at
|
||||
integer frame_count "Extracted frames count"
|
||||
real created_at "Unix epoch timestamp"
|
||||
real merged_at "Timestamp when batch merged to dataset"
|
||||
}
|
||||
|
||||
FRAMES {
|
||||
integer id PK
|
||||
integer batch_id FK
|
||||
integer idx
|
||||
text filename
|
||||
integer width
|
||||
integer height
|
||||
integer id PK "AUTOINCREMENT"
|
||||
integer batch_id FK "References batches(id)"
|
||||
integer idx "Frame index within batch"
|
||||
text filename "Image file name (000001.jpg)"
|
||||
integer width "Frame pixel width"
|
||||
integer height "Frame pixel height"
|
||||
text review_status "pending | approved | rejected"
|
||||
}
|
||||
|
||||
ANNOTATIONS {
|
||||
integer id PK
|
||||
integer frame_id FK
|
||||
integer class_id
|
||||
text geometry "JSON / coordinates"
|
||||
real score
|
||||
integer id PK "AUTOINCREMENT"
|
||||
integer frame_id FK "References frames(id)"
|
||||
integer class_id "Target YOLO class index"
|
||||
text geometry "Box or Polygon coordinates JSON"
|
||||
real score "SAM3 confidence score (0.0 - 1.0)"
|
||||
text source "auto | manual"
|
||||
real created_at
|
||||
real created_at "Unix epoch timestamp"
|
||||
}
|
||||
|
||||
ANNOTATION_OVERRIDES {
|
||||
integer annotation_id PK, FK
|
||||
integer annotation_id PK, FK "References annotations(id)"
|
||||
text verdict "keep | ignore | reclass"
|
||||
integer target_class
|
||||
real decided_at
|
||||
integer target_class "New class index if reclassified"
|
||||
real decided_at "Unix epoch timestamp"
|
||||
}
|
||||
|
||||
DATASETS {
|
||||
integer id PK
|
||||
integer project_id FK
|
||||
text name
|
||||
text note
|
||||
text rule_version
|
||||
text rules_json
|
||||
real created_at
|
||||
integer id PK "AUTOINCREMENT"
|
||||
integer project_id FK "References projects(id)"
|
||||
text name "Dataset version name (e.g. Master v1)"
|
||||
text note "Release description"
|
||||
text rule_version "Triage rule version identifier"
|
||||
text rules_json "Frozen triage rules JSON snapshot"
|
||||
real created_at "Unix epoch timestamp"
|
||||
}
|
||||
|
||||
DATASET_ITEMS {
|
||||
integer id PK
|
||||
integer project_id FK
|
||||
integer dataset_id FK
|
||||
integer frame_id FK
|
||||
integer id PK "AUTOINCREMENT"
|
||||
integer project_id FK "References projects(id)"
|
||||
integer dataset_id FK "References datasets(id)"
|
||||
integer frame_id FK "References frames(id)"
|
||||
text split "train | val"
|
||||
text image_rel
|
||||
text label_rel
|
||||
real added_at
|
||||
text image_rel "Relative image path in dataset"
|
||||
text label_rel "Relative label path in dataset"
|
||||
real added_at "Unix epoch timestamp"
|
||||
}
|
||||
|
||||
BASE_DATASETS {
|
||||
integer id PK
|
||||
integer project_id FK
|
||||
text name
|
||||
text source
|
||||
integer image_count
|
||||
integer box_count
|
||||
text classes "JSON array"
|
||||
real created_at
|
||||
integer id PK "AUTOINCREMENT"
|
||||
integer project_id FK "References projects(id)"
|
||||
text name "External baseline dataset name"
|
||||
text source "Import origin description"
|
||||
integer image_count "Number of baseline images"
|
||||
integer box_count "Number of baseline bounding boxes"
|
||||
text classes "JSON array of class names"
|
||||
real created_at "Unix epoch timestamp"
|
||||
}
|
||||
|
||||
TRIAGE_RULES {
|
||||
integer id PK
|
||||
integer project_id FK
|
||||
text stage "default: dataprep"
|
||||
integer position
|
||||
text name
|
||||
text predicate
|
||||
integer id PK "AUTOINCREMENT"
|
||||
integer project_id FK "References projects(id)"
|
||||
text stage "Filter stage (default: dataprep)"
|
||||
integer position "Rule execution order position"
|
||||
text name "Rule human readable name"
|
||||
text predicate "Evaluation expression (area, aspect, conf)"
|
||||
text action "keep | ignore | reclass"
|
||||
integer target_class
|
||||
real created_at
|
||||
integer target_class "Target class index if reclass"
|
||||
real created_at "Unix epoch timestamp"
|
||||
}
|
||||
|
||||
MODEL_VERSIONS {
|
||||
integer id PK
|
||||
integer project_id FK
|
||||
integer version
|
||||
text weights_path
|
||||
text parent_model_path
|
||||
text metrics "JSON"
|
||||
text base_metrics "JSON"
|
||||
real created_at
|
||||
integer id PK "AUTOINCREMENT"
|
||||
integer project_id FK "References projects(id)"
|
||||
integer version "Incrementing version integer"
|
||||
text weights_path "Path to trained best.pt weights"
|
||||
text parent_model_path "Path to base model used as starting point"
|
||||
text metrics "Trained model evaluation metrics JSON"
|
||||
text base_metrics "Base model evaluation metrics JSON"
|
||||
text rule_version "Triage rule version used for training"
|
||||
text augment "Augmentation settings JSON used"
|
||||
real created_at "Unix epoch timestamp"
|
||||
}
|
||||
|
||||
JOBS {
|
||||
integer id PK
|
||||
integer project_id FK
|
||||
integer batch_id FK
|
||||
integer id PK "AUTOINCREMENT"
|
||||
integer project_id FK "References projects(id)"
|
||||
integer batch_id FK "References batches(id)"
|
||||
text type "extract | autolabel | merge | train | count | clock-scan | truck-scan"
|
||||
text status "queued | running | done | failed | cancelled"
|
||||
text params "JSON"
|
||||
integer progress
|
||||
integer total
|
||||
text message
|
||||
text error
|
||||
text log
|
||||
real created_at
|
||||
real started_at
|
||||
real finished_at
|
||||
text params "Job parameters JSON"
|
||||
integer progress "Current completed units"
|
||||
integer total "Total units of work"
|
||||
text message "Human-readable status update"
|
||||
text error "Failure message or traceback"
|
||||
text log "Detailed execution log lines"
|
||||
real created_at "Unix epoch timestamp"
|
||||
real started_at "Job start timestamp"
|
||||
real finished_at "Job completion timestamp"
|
||||
}
|
||||
|
||||
VIDEO_CLOCK {
|
||||
integer id PK
|
||||
integer project_id FK
|
||||
text video_rel
|
||||
text folder_date
|
||||
text started_at "YYYY-MM-DD HH:MM:SS"
|
||||
text working_day
|
||||
real confidence
|
||||
integer agreeing
|
||||
text source "ocr"
|
||||
text error
|
||||
real read_at
|
||||
integer id PK "AUTOINCREMENT"
|
||||
integer project_id FK "References projects(id)"
|
||||
text video_rel "Relative path to video file"
|
||||
text folder_date "Date folder extracted from path"
|
||||
text started_at "Burned-in OCR clock string (YYYY-MM-DD HH:MM:SS)"
|
||||
text working_day "Operational shift day (06:00 - 05:59 cutoff)"
|
||||
real confidence "OCR reading confidence"
|
||||
integer agreeing "Number of sample frames agreeing"
|
||||
text source "ocr | manual"
|
||||
text error "OCR parsing error message"
|
||||
real read_at "Timestamp when OCR scan was executed"
|
||||
integer truck_hits "Frames with truck detected"
|
||||
integer truck_samples "Total sampled frames for truck check"
|
||||
text truck_model "Model used for truck scan"
|
||||
real truck_checked_at "Timestamp of truck scan"
|
||||
}
|
||||
|
||||
COUNT_RUNS {
|
||||
integer id PK
|
||||
integer project_id FK
|
||||
text video_rel
|
||||
text date_label
|
||||
text batch_label
|
||||
integer loading
|
||||
integer unloading
|
||||
integer net
|
||||
integer ground_truth
|
||||
integer frames
|
||||
real seconds
|
||||
text params "JSON"
|
||||
text model_path
|
||||
text error
|
||||
real counted_at
|
||||
integer id PK "AUTOINCREMENT"
|
||||
integer project_id FK "References projects(id)"
|
||||
text video_rel "Relative path to video evaluated"
|
||||
text date_label "Video date directory"
|
||||
text batch_label "Video file base name"
|
||||
integer loading "Counted loading direction crossings"
|
||||
integer unloading "Counted unloading direction crossings"
|
||||
integer net "Calculated net count (loading - unloading)"
|
||||
integer ground_truth "Physical verified hand-tally count"
|
||||
integer frames "Total video frames evaluated"
|
||||
real seconds "Inference elapsed time in seconds"
|
||||
text params "Counting line & ByteTrack parameters JSON"
|
||||
text model_path "YOLO model weights path used"
|
||||
text error "Error message if benchmark failed"
|
||||
real counted_at "Unix epoch timestamp"
|
||||
}
|
||||
```
|
||||
|
||||
@@ -208,40 +218,43 @@ erDiagram
|
||||
|
||||
## Entity Descriptions
|
||||
|
||||
### 1. Project & Class Configuration
|
||||
* **`projects`**: Core isolation entity. Configures the target label type (`bbox` vs `polygon`), base model reference, video archive root, and validation split step (`val_every`).
|
||||
* **`project_classes`**: Class definitions tied to a project. Stores integer `class_id`, class `name`, and natural language SAM3 zero-shot `prompt`.
|
||||
### 1. Workspace & Multi-Project Isolation
|
||||
* **`projects`**: Top-level entity isolating datasets, classes, models, and CCTV archives. Enforces geometry type (`bbox` vs `polygon`) and stores the active base model checkpoint.
|
||||
* **`project_classes`**: YOLO class definitions for the project. Pairs each integer `class_id` with natural language text prompts utilized by SAM3 zero-shot auto-annotation.
|
||||
|
||||
### 2. Video Extraction & Annotation Pipeline
|
||||
* **`batches`**: A trimmed segment from a raw CCTV video file. Holds time ranges, sampling FPS, frame counts, and extraction/review lifecycles.
|
||||
* **`frames`**: Extracted image stills belonging to a batch, tracking width, height, index, and operator approval status (`pending`, `approved`, `rejected`).
|
||||
* **`annotations`**: Bounding boxes or polygon geometries per frame with confidence scores, class IDs, and origin (`auto` from SAM3 vs `manual` from operator).
|
||||
* **`annotation_overrides`**: Per-annotation triage verdicts (`keep`, `ignore`, `reclass`) resulting from outlier inspection.
|
||||
### 2. Video Ingest & Annotation Lifecycle
|
||||
* **`batches`**: A trimmed segment of raw industrial CCTV footage. Captures start/end timestamps, sampling FPS, frame counts, and progression from extraction to dataset merge.
|
||||
* **`frames`**: Discrete JPEG stills extracted from a batch. Tracks individual review status (`pending`, `approved`, `rejected`).
|
||||
* **`annotations`**: Object bounding boxes or polygon masks per frame with SAM3 / manual confidence scores and class assignments.
|
||||
* **`annotation_overrides`**: Outlier inspection decisions (`keep`, `ignore`, `reclass`) resulting from interactive 2D scatter plot triage.
|
||||
|
||||
### 3. Master Datasets & Data Preparation
|
||||
* **`datasets`**: Immutable dataset compilations (`v1`, `v2`, `v3`) capturing the snapshot rules and timestamps.
|
||||
* **`dataset_items`**: Mapping between a dataset and its constituent image frames, recording deterministic train/val splits (`split IN ('train', 'val')`) and relative filesystem paths.
|
||||
* **`base_datasets`**: External train-only datasets imported to supplement training data without affecting validation splits.
|
||||
* **`triage_rules`**: Ordered filter predicates applied during data prep to systematically prune bounding box anomalies (e.g. area, aspect ratio, confidence).
|
||||
* **`datasets`**: Immutable, versioned master datasets (`v1`, `v2`, `v3`). Captures a permanent JSON snapshot of triage rules (`rules_json`) applied at compilation time.
|
||||
* **`dataset_items`**: Junction entity binding frames into dataset versions with deterministic SHA-1 validation split partitioning (`train` vs `val`).
|
||||
* **`base_datasets`**: External pre-labeled baseline datasets mounted as train-only supplements without polluting validation benchmarks.
|
||||
* **`triage_rules`**: Ordered filtering predicates (e.g., box area, aspect ratio, confidence thresholds) applied during data preparation.
|
||||
|
||||
### 4. Training, Jobs & Telemetry
|
||||
* **`model_versions`**: Trained YOLO checkpoints (`1`, `2`, `3`...) storing weights paths, parent models, and side-by-side metric evaluations ($\Delta\text{mAP50}$, precision, recall).
|
||||
* **`jobs`**: Asynchronous background job queue (`extract`, `autolabel`, `train`, `count`, etc.) managing progress counters, logs, and state transitions.
|
||||
### 4. Continuous YOLO Retraining & Background Workers
|
||||
* **`model_versions`**: Fine-tuned YOLO weights checkpoints. Stores side-by-side performance deltas ($\Delta\text{mAP50}$, $\Delta\text{Precision}$, $\Delta\text{Recall}$) evaluated against the identical frozen validation split.
|
||||
* **`jobs`**: Centralized SQLite job queue for asynchronous background tasks (`extract`, `autolabel`, `train`, `count`, etc.) with atomic status management and progress streaming.
|
||||
|
||||
### 5. Video Clock & Production Counting Benchmarks
|
||||
* **`video_clock`**: OCR timestamps extracted from CCTV video overlays to assign recordings to proper 24-hour work shifts (06:00 to 06:00 cutoff).
|
||||
* **`count_runs`**: Inference benchmarks running ByteTrack line-crossing counters against verified physical ground truth numbers.
|
||||
### 5. Shift OCR Indexing & Production Line Counter
|
||||
* **`video_clock`**: OCR timestamps extracted from burned-in CCTV camera overlays, mapping recordings to 24-hour manufacturing shifts (06:00 AM to 05:59 AM next day) and detecting truck arrival presence.
|
||||
* **`count_runs`**: Real-time production inference benchmarks running ByteTrack line-crossing counters against verified ground truth tallies.
|
||||
|
||||
---
|
||||
|
||||
## Database Indexes
|
||||
## Performance & Database Indexes
|
||||
|
||||
To maintain sub-second UI performance across thousands of frames and annotations, the following indices are maintained:
|
||||
* `idx_frames_batch` on `frames(batch_id, idx)`
|
||||
* `idx_annotations_frame` on `annotations(frame_id)`
|
||||
* `idx_batches_project` on `batches(project_id)`
|
||||
* `idx_dataset_items_project` on `dataset_items(project_id)`
|
||||
* `idx_triage_rules_project` on `triage_rules(project_id, stage, position)`
|
||||
* `idx_jobs_project` on `jobs(project_id, created_at)`
|
||||
* `idx_video_clock_project` on `video_clock(project_id, working_day)`
|
||||
* `idx_count_runs_project` on `count_runs(project_id, date_label)`
|
||||
The SQLite database enforces the following indexes to maintain sub-millisecond query performance:
|
||||
|
||||
| Index Name | Table | Indexed Columns | Purpose |
|
||||
|---|---|---|---|
|
||||
| `idx_frames_batch` | `frames` | `(batch_id, idx)` | Rapid frame lookup and sequential canvas scrubbing |
|
||||
| `idx_annotations_frame` | `annotations` | `(frame_id)` | Sub-millisecond bounding box loading per canvas frame |
|
||||
| `idx_batches_project` | `batches` | `(project_id)` | Fast batch library filtering by project |
|
||||
| `idx_dataset_items_project`| `dataset_items` | `(project_id)` | Dataset compilation and split integrity checks |
|
||||
| `idx_triage_rules_project` | `triage_rules` | `(project_id, stage, position)` | Fast ordered triage predicate evaluation |
|
||||
| `idx_jobs_project` | `jobs` | `(project_id, created_at)` | UI job queue monitoring and polling |
|
||||
| `idx_video_clock_project` | `video_clock` | `(project_id, working_day)` | Shift-based video library filtering |
|
||||
| `idx_count_runs_project` | `count_runs` | `(project_id, date_label)` | Production counting benchmark reporting |
|
||||
Reference in new issue
Block a user