docs: add database entity relationship diagram (ERD.md)

This commit is contained in:
Andrew-AAAA committed 2026-08-31 15:44:27 +07:00
1 parent c0cce17490
commit 8ad35ed9d1
2 files changed
+256 -2

No files matched your search

+247
View File
@@ -0,0 +1,247 @@
# Entity Relationship Diagram (ERD)
This document specifies the SQLite database schema and entity relationships for the **reTraining** platform (`app.db`), as defined in [backend/db.py](file:///C:/Users/araar/Downloads/PT_SIAB_FULLTIME/Feedmill_Semarang/reTraining/backend/db.py).
---
## Architectural Storage Model
The platform uses a **hybrid storage architecture**:
* **SQLite Database (`data/app.db`)**: Stores entity metadata, relations, job queues, triage rules, review states, and metrics.
* **Disk Filesystem (`data/projects/`)**: Stores image pixels (`.jpg`), YOLO labels (`.txt`), YAML dataset manifests (`data.yaml`), and trained model weights (`.pt`).
---
## Mermaid Entity Relationship Diagram
```mermaid
erDiagram
PROJECTS ||--o{ PROJECT_CLASSES : "defines"
PROJECTS ||--o{ BATCHES : "contains"
PROJECTS ||--o{ DATASETS : "compiles"
PROJECTS ||--o{ DATASET_ITEMS : "aggregates"
PROJECTS ||--o{ BASE_DATASETS : "mounts"
PROJECTS ||--o{ TRIAGE_RULES : "configures"
PROJECTS ||--o{ MODEL_VERSIONS : "produces"
PROJECTS ||--o{ JOBS : "executes"
PROJECTS ||--o{ VIDEO_CLOCK : "indexes"
PROJECTS ||--o{ COUNT_RUNS : "benchmarks"
BATCHES ||--o{ FRAMES : "extracts"
BATCHES ||--o{ JOBS : "triggers"
FRAMES ||--o{ ANNOTATIONS : "annotates"
FRAMES ||--o{ DATASET_ITEMS : "includes"
ANNOTATIONS ||--o| ANNOTATION_OVERRIDES : "overrides"
DATASETS ||--o{ DATASET_ITEMS : "contains"
PROJECTS {
integer id PK
text slug UK
text name
text label_type "bbox | polygon"
text base_model_path
text base_model_kind "uploaded | pretrained | trained"
text video_root
integer val_every "default: 5"
real created_at
}
PROJECT_CLASSES {
integer id PK
integer project_id FK
integer class_id
text name
text prompt
}
BATCHES {
integer id PK
integer project_id FK
text video_path
text date_label
text batch_label
real start_sec
real end_sec
real fps
text status "extracting | extracted | labeling | reviewing | approved | merged | failed"
integer frame_count
real created_at
real merged_at
}
FRAMES {
integer id PK
integer batch_id FK
integer idx
text filename
integer width
integer height
text review_status "pending | approved | rejected"
}
ANNOTATIONS {
integer id PK
integer frame_id FK
integer class_id
text geometry "JSON / coordinates"
real score
text source "auto | manual"
real created_at
}
ANNOTATION_OVERRIDES {
integer annotation_id PK, FK
text verdict "keep | ignore | reclass"
integer target_class
real decided_at
}
DATASETS {
integer id PK
integer project_id FK
text name
text note
text rule_version
text rules_json
real created_at
}
DATASET_ITEMS {
integer id PK
integer project_id FK
integer dataset_id FK
integer frame_id FK
text split "train | val"
text image_rel
text label_rel
real added_at
}
BASE_DATASETS {
integer id PK
integer project_id FK
text name
text source
integer image_count
integer box_count
text classes "JSON array"
real created_at
}
TRIAGE_RULES {
integer id PK
integer project_id FK
text stage "default: dataprep"
integer position
text name
text predicate
text action "keep | ignore | reclass"
integer target_class
real created_at
}
MODEL_VERSIONS {
integer id PK
integer project_id FK
integer version
text weights_path
text parent_model_path
text metrics "JSON"
text base_metrics "JSON"
real created_at
}
JOBS {
integer id PK
integer project_id FK
integer batch_id FK
text type "extract | autolabel | merge | train | count | clock-scan | truck-scan"
text status "queued | running | done | failed | cancelled"
text params "JSON"
integer progress
integer total
text message
text error
text log
real created_at
real started_at
real finished_at
}
VIDEO_CLOCK {
integer id PK
integer project_id FK
text video_rel
text folder_date
text started_at "YYYY-MM-DD HH:MM:SS"
text working_day
real confidence
integer agreeing
text source "ocr"
text error
real read_at
}
COUNT_RUNS {
integer id PK
integer project_id FK
text video_rel
text date_label
text batch_label
integer loading
integer unloading
integer net
integer ground_truth
integer frames
real seconds
text params "JSON"
text model_path
text error
real counted_at
}
```
---
## Entity Descriptions
### 1. Project & Class Configuration
* **`projects`**: Core isolation entity. Configures the target label type (`bbox` vs `polygon`), base model reference, video archive root, and validation split step (`val_every`).
* **`project_classes`**: Class definitions tied to a project. Stores integer `class_id`, class `name`, and natural language SAM3 zero-shot `prompt`.
### 2. Video Extraction & Annotation Pipeline
* **`batches`**: A trimmed segment from a raw CCTV video file. Holds time ranges, sampling FPS, frame counts, and extraction/review lifecycles.
* **`frames`**: Extracted image stills belonging to a batch, tracking width, height, index, and operator approval status (`pending`, `approved`, `rejected`).
* **`annotations`**: Bounding boxes or polygon geometries per frame with confidence scores, class IDs, and origin (`auto` from SAM3 vs `manual` from operator).
* **`annotation_overrides`**: Per-annotation triage verdicts (`keep`, `ignore`, `reclass`) resulting from outlier inspection.
### 3. Master Datasets & Data Preparation
* **`datasets`**: Immutable dataset compilations (`v1`, `v2`, `v3`) capturing the snapshot rules and timestamps.
* **`dataset_items`**: Mapping between a dataset and its constituent image frames, recording deterministic train/val splits (`split IN ('train', 'val')`) and relative filesystem paths.
* **`base_datasets`**: External train-only datasets imported to supplement training data without affecting validation splits.
* **`triage_rules`**: Ordered filter predicates applied during data prep to systematically prune bounding box anomalies (e.g. area, aspect ratio, confidence).
### 4. Training, Jobs & Telemetry
* **`model_versions`**: Trained YOLO checkpoints (`1`, `2`, `3`...) storing weights paths, parent models, and side-by-side metric evaluations ($\Delta\text{mAP50}$, precision, recall).
* **`jobs`**: Asynchronous background job queue (`extract`, `autolabel`, `train`, `count`, etc.) managing progress counters, logs, and state transitions.
### 5. Video Clock & Production Counting Benchmarks
* **`video_clock`**: OCR timestamps extracted from CCTV video overlays to assign recordings to proper 24-hour work shifts (06:00 to 06:00 cutoff).
* **`count_runs`**: Inference benchmarks running ByteTrack line-crossing counters against verified physical ground truth numbers.
---
## Database Indexes
To maintain sub-second UI performance across thousands of frames and annotations, the following indices are maintained:
* `idx_frames_batch` on `frames(batch_id, idx)`
* `idx_annotations_frame` on `annotations(frame_id)`
* `idx_batches_project` on `batches(project_id)`
* `idx_dataset_items_project` on `dataset_items(project_id)`
* `idx_triage_rules_project` on `triage_rules(project_id, stage, position)`
* `idx_jobs_project` on `jobs(project_id, created_at)`
* `idx_video_clock_project` on `video_clock(project_id, working_day)`
* `idx_count_runs_project` on `count_runs(project_id, date_label)`