Files
reTraining/docs/ERD.md
T

8.0 KiB

Entity Relationship Diagram (ERD)

This document specifies the SQLite database schema and entity relationships for the reTraining platform (app.db), as defined in backend/db.py.


Architectural Storage Model

The platform uses a hybrid storage architecture:

  • SQLite Database (data/app.db): Stores entity metadata, relations, job queues, triage rules, review states, and metrics.
  • Disk Filesystem (data/projects/): Stores image pixels (.jpg), YOLO labels (.txt), YAML dataset manifests (data.yaml), and trained model weights (.pt).

Mermaid Entity Relationship Diagram

erDiagram
    PROJECTS ||--o{ PROJECT_CLASSES : "defines"
    PROJECTS ||--o{ BATCHES : "contains"
    PROJECTS ||--o{ DATASETS : "compiles"
    PROJECTS ||--o{ DATASET_ITEMS : "aggregates"
    PROJECTS ||--o{ BASE_DATASETS : "mounts"
    PROJECTS ||--o{ TRIAGE_RULES : "configures"
    PROJECTS ||--o{ MODEL_VERSIONS : "produces"
    PROJECTS ||--o{ JOBS : "executes"
    PROJECTS ||--o{ VIDEO_CLOCK : "indexes"
    PROJECTS ||--o{ COUNT_RUNS : "benchmarks"

    BATCHES ||--o{ FRAMES : "extracts"
    BATCHES ||--o{ JOBS : "triggers"

    FRAMES ||--o{ ANNOTATIONS : "annotates"
    FRAMES ||--o{ DATASET_ITEMS : "includes"

    ANNOTATIONS ||--o| ANNOTATION_OVERRIDES : "overrides"

    DATASETS ||--o{ DATASET_ITEMS : "contains"

    PROJECTS {
        integer id PK
        text slug UK
        text name
        text label_type "bbox | polygon"
        text base_model_path
        text base_model_kind "uploaded | pretrained | trained"
        text video_root
        integer val_every "default: 5"
        real created_at
    }

    PROJECT_CLASSES {
        integer id PK
        integer project_id FK
        integer class_id
        text name
        text prompt
    }

    BATCHES {
        integer id PK
        integer project_id FK
        text video_path
        text date_label
        text batch_label
        real start_sec
        real end_sec
        real fps
        text status "extracting | extracted | labeling | reviewing | approved | merged | failed"
        integer frame_count
        real created_at
        real merged_at
    }

    FRAMES {
        integer id PK
        integer batch_id FK
        integer idx
        text filename
        integer width
        integer height
        text review_status "pending | approved | rejected"
    }

    ANNOTATIONS {
        integer id PK
        integer frame_id FK
        integer class_id
        text geometry "JSON / coordinates"
        real score
        text source "auto | manual"
        real created_at
    }

    ANNOTATION_OVERRIDES {
        integer annotation_id PK, FK
        text verdict "keep | ignore | reclass"
        integer target_class
        real decided_at
    }

    DATASETS {
        integer id PK
        integer project_id FK
        text name
        text note
        text rule_version
        text rules_json
        real created_at
    }

    DATASET_ITEMS {
        integer id PK
        integer project_id FK
        integer dataset_id FK
        integer frame_id FK
        text split "train | val"
        text image_rel
        text label_rel
        real added_at
    }

    BASE_DATASETS {
        integer id PK
        integer project_id FK
        text name
        text source
        integer image_count
        integer box_count
        text classes "JSON array"
        real created_at
    }

    TRIAGE_RULES {
        integer id PK
        integer project_id FK
        text stage "default: dataprep"
        integer position
        text name
        text predicate
        text action "keep | ignore | reclass"
        integer target_class
        real created_at
    }

    MODEL_VERSIONS {
        integer id PK
        integer project_id FK
        integer version
        text weights_path
        text parent_model_path
        text metrics "JSON"
        text base_metrics "JSON"
        real created_at
    }

    JOBS {
        integer id PK
        integer project_id FK
        integer batch_id FK
        text type "extract | autolabel | merge | train | count | clock-scan | truck-scan"
        text status "queued | running | done | failed | cancelled"
        text params "JSON"
        integer progress
        integer total
        text message
        text error
        text log
        real created_at
        real started_at
        real finished_at
    }

    VIDEO_CLOCK {
        integer id PK
        integer project_id FK
        text video_rel
        text folder_date
        text started_at "YYYY-MM-DD HH:MM:SS"
        text working_day
        real confidence
        integer agreeing
        text source "ocr"
        text error
        real read_at
    }

    COUNT_RUNS {
        integer id PK
        integer project_id FK
        text video_rel
        text date_label
        text batch_label
        integer loading
        integer unloading
        integer net
        integer ground_truth
        integer frames
        real seconds
        text params "JSON"
        text model_path
        text error
        real counted_at
    }

Entity Descriptions

1. Project & Class Configuration

  • projects: Core isolation entity. Configures the target label type (bbox vs polygon), base model reference, video archive root, and validation split step (val_every).
  • project_classes: Class definitions tied to a project. Stores integer class_id, class name, and natural language SAM3 zero-shot prompt.

2. Video Extraction & Annotation Pipeline

  • batches: A trimmed segment from a raw CCTV video file. Holds time ranges, sampling FPS, frame counts, and extraction/review lifecycles.
  • frames: Extracted image stills belonging to a batch, tracking width, height, index, and operator approval status (pending, approved, rejected).
  • annotations: Bounding boxes or polygon geometries per frame with confidence scores, class IDs, and origin (auto from SAM3 vs manual from operator).
  • annotation_overrides: Per-annotation triage verdicts (keep, ignore, reclass) resulting from outlier inspection.

3. Master Datasets & Data Preparation

  • datasets: Immutable dataset compilations (v1, v2, v3) capturing the snapshot rules and timestamps.
  • dataset_items: Mapping between a dataset and its constituent image frames, recording deterministic train/val splits (split IN ('train', 'val')) and relative filesystem paths.
  • base_datasets: External train-only datasets imported to supplement training data without affecting validation splits.
  • triage_rules: Ordered filter predicates applied during data prep to systematically prune bounding box anomalies (e.g. area, aspect ratio, confidence).

4. Training, Jobs & Telemetry

  • model_versions: Trained YOLO checkpoints (1, 2, 3...) storing weights paths, parent models, and side-by-side metric evaluations (\Delta\text{mAP50}, precision, recall).
  • jobs: Asynchronous background job queue (extract, autolabel, train, count, etc.) managing progress counters, logs, and state transitions.

5. Video Clock & Production Counting Benchmarks

  • video_clock: OCR timestamps extracted from CCTV video overlays to assign recordings to proper 24-hour work shifts (06:00 to 06:00 cutoff).
  • count_runs: Inference benchmarks running ByteTrack line-crossing counters against verified physical ground truth numbers.

Database Indexes

To maintain sub-second UI performance across thousands of frames and annotations, the following indices are maintained:

  • idx_frames_batch on frames(batch_id, idx)
  • idx_annotations_frame on annotations(frame_id)
  • idx_batches_project on batches(project_id)
  • idx_dataset_items_project on dataset_items(project_id)
  • idx_triage_rules_project on triage_rules(project_id, stage, position)
  • idx_jobs_project on jobs(project_id, created_at)
  • idx_video_clock_project on video_clock(project_id, working_day)
  • idx_count_runs_project on count_runs(project_id, date_label)