678 lines
37 KiB
Markdown
678 lines
37 KiB
Markdown
<div align="center">
|
||
|
||
<a href="https://github.com/DBS-Internship/reTraining">
|
||
<img src="docs/logo.png" alt="reTraining AI Computer Vision Platform Logo" width="220" style="border-radius: 16px;" />
|
||
</a>
|
||
|
||
# reTraining
|
||
|
||
### Take a base model you already have, and make it measurably better with footage you already have.
|
||
|
||
**The Open-Source, Self-Hosted Vision Pipeline to Turn Raw Industrial CCTV into Production-Grade YOLO Object Detectors & Line Counters.**
|
||
|
||
<p>
|
||
<a href="https://www.python.org/"><img alt="Python 3.12" src="https://img.shields.io/badge/python-3.12-3776AB?logo=python&logoColor=white"></a>
|
||
<a href="https://fastapi.tiangolo.com/"><img alt="FastAPI" src="https://img.shields.io/badge/FastAPI-72%20endpoints-009688?logo=fastapi&logoColor=white"></a>
|
||
<a href="https://react.dev/"><img alt="React 19" src="https://img.shields.io/badge/React-19-61DAFB?logo=react&logoColor=black"></a>
|
||
<a href="https://vitejs.dev/"><img alt="Vite 7" src="https://img.shields.io/badge/Vite-7-646CFF?logo=vite&logoColor=white"></a>
|
||
<a href="https://huggingface.co/facebook/sam3"><img alt="SAM3" src="https://img.shields.io/badge/SAM3-zero--shot-FF6F00"></a>
|
||
<a href="https://github.com/ultralytics/ultralytics"><img alt="YOLO11" src="https://img.shields.io/badge/Ultralytics-YOLO11-00BFA5"></a>
|
||
<a href="https://www.docker.com/"><img alt="Docker" src="https://img.shields.io/badge/Docker-Compose-2496ED?logo=docker&logoColor=white"></a>
|
||
<a href="https://developer.nvidia.com/cuda-toolkit"><img alt="CUDA" src="https://img.shields.io/badge/CUDA-12.4+-76B900?logo=nvidia&logoColor=white"></a>
|
||
<a href="LICENSE"><img alt="License" src="https://img.shields.io/badge/license-MIT-blue.svg"></a>
|
||
<a href="docs/PANDUAN_SISTEM_LENGKAP.pdf"><img alt="Documentation" src="https://img.shields.io/badge/Docs-PDF%20Guide-red?logo=adobe-acrobat-reader&logoColor=white"></a>
|
||
</p>
|
||
|
||
<p align="center">
|
||
<a href="docs/PANDUAN_SISTEM_LENGKAP.pdf">
|
||
<img src="https://img.shields.io/badge/📥_DOWNLOAD_HANDOVER_DOCS-PDF_GUIDE_(41_PAGES)-2563EB?style=for-the-badge&logo=adobe-acrobat-reader&logoColor=white" alt="Download Handover Docs PDF" />
|
||
</a>
|
||
|
||
<a href="docs/diagram-alur.png">
|
||
<img src="https://img.shields.io/badge/📊_DOWNLOAD_ARCHITECTURE-4K_UHD_DIAGRAM-0D9488?style=for-the-badge&logo=diagramsdotnet&logoColor=white" alt="Download 4K Diagram" />
|
||
</a>
|
||
</p>
|
||
|
||
[⚡ Quickstart](#quickstart) · [🏗️ Architecture](#architecture) · [🖼️ Visual UI Tour](#visual-tour) · [🎥 Counting Engine](#counting-engine) · [⚙️ Configuration](#configuration) · [📚 Documentation](#documentation) · [🔧 Troubleshooting](#troubleshooting)
|
||
|
||
<br>
|
||
|
||
<a href="docs/diagram-alur.png">
|
||
<img src="docs/diagram-alur.png" width="100%" alt="System Architecture & End-to-End Pipeline Diagram" />
|
||
</a>
|
||
|
||
*Click diagram to view 4K UHD resolution. Scalable vector version available at [docs/diagram-alur.svg](docs/diagram-alur.svg).*
|
||
|
||
</div>
|
||
|
||
---
|
||
|
||
<a id="why-retraining"></a>
|
||
## 💡 Why reTraining?
|
||
|
||
Deploying object detection models in industrial environments (manufacturing, logistics, agricultural feedmills, conveyor belts) often hits a painful bottleneck: **general base models fail on domain-specific edge cases, while building manual labeling pipelines from scratch is slow and expensive.**
|
||
|
||
**reTraining** provides a self-contained, enterprise-grade active learning platform designed to run directly on your edge server or GPU workstation:
|
||
|
||
1. **Ingest Raw CCTV Footage**: Stream directly from continuous 24/7 video archives without re-encoding or modifying the underlying storage.
|
||
2. **Zero-Shot Foundation Auto-Labeling**: Leverage Meta's Segment Anything Model 3 (**SAM3**) with natural language text prompts and visual exemplars to annotate thousands of frames in minutes.
|
||
3. **Roboflow-Grade Review Studio**: Sub-second keyboard navigation, instant class switching, click-assist segmentation, and high-density visual triage crop grids.
|
||
4. **Statistical Triage & Data Prep**: Filter bounding box outliers by area, aspect ratio, and confidence score without discarding valid image frames.
|
||
5. **Immutable Dataset Freezing & Stable Val Splits**: Guarantee reproducible benchmarks with deterministic SHA-1 validation sets that never shift across retraining runs.
|
||
6. **Hardware-Aware Continuous Fine-Tuning**: Auto-detect host GPU VRAM, fine-tune Ultralytics YOLO11 models, and evaluate base vs. fine-tuned model performance side-by-side.
|
||
7. **Live Production Counting & Benchmarks**: Real-time RTSP/WHEP live inference with ByteTrack line-crossing counters, evaluated against verified ground truth physical counts at **146 FPS**.
|
||
|
||
---
|
||
|
||
<a id="quickstart"></a>
|
||
## ⚡ Quickstart
|
||
|
||
### Prerequisites
|
||
|
||
- **Host OS**: Linux (Ubuntu 22.04+ recommended) or Windows with WSL2.
|
||
- **GPU Acceleration**: NVIDIA GPU with CUDA 12.4+ and [NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html) (Docker Engine 27+ CDI support).
|
||
- **HuggingFace Account**: Gated model access granted for [facebook/sam3](https://huggingface.co/facebook/sam3) with a valid user access token (`HF_TOKEN`).
|
||
- **Video Storage**: Directory of CCTV footage structured as `<date>/<batch>.mp4` (or let the app mount `./data/archive`).
|
||
|
||
---
|
||
|
||
### Option 1: Docker Compose (Production Standard)
|
||
|
||
Launch the entire stack with a single command. The startup script automatically inspects the host hardware, generates Container Device Interface (CDI) specs for your NVIDIA GPU, and starts both backend and frontend containers.
|
||
|
||
```bash
|
||
# 1. Clone the repository
|
||
git clone git@github.com:fhanyuh/reTraining.git
|
||
cd reTraining
|
||
|
||
# 2. Configure environment credentials
|
||
cp .env.example .env
|
||
# Edit .env and paste your HuggingFace user access token:
|
||
# HF_TOKEN=hf_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
|
||
|
||
# 3. Launch with automated GPU / CDI detection
|
||
chmod +x start.sh
|
||
./start.sh
|
||
```
|
||
|
||
*Alternative direct Docker Compose launch:*
|
||
```bash
|
||
docker compose up -d --build
|
||
```
|
||
|
||
**Access Points**:
|
||
- 🌐 **Web Studio UI**: [http://localhost:9000](http://localhost:9000)
|
||
- 📑 **Interactive REST API Docs**: [http://localhost:9010/docs](http://localhost:9010/docs)
|
||
- 🩺 **Health & GPU Telemetry Endpoint**: [http://localhost:9010/api/health](http://localhost:9010/api/health)
|
||
|
||
Verify system health and GPU VRAM availability:
|
||
```bash
|
||
curl http://localhost:9010/api/health
|
||
```
|
||
```json
|
||
{
|
||
"device": "cuda",
|
||
"gpu": "NVIDIA GeForce RTX 5080 Laptop GPU",
|
||
"vram_free_gb": 14.91,
|
||
"sam3_ready": true,
|
||
"ffmpeg": true,
|
||
"hf_token": true,
|
||
"db": true
|
||
}
|
||
```
|
||
|
||
> [!IMPORTANT]
|
||
> The initial SAM3 auto-annotation job automatically downloads the **~3.4 GB SAM3 checkpoint** from HuggingFace into a persistent Docker named volume (`hf-cache`). Subsequent executions load the model into VRAM in ~12 seconds.
|
||
|
||
---
|
||
|
||
### Option 2: Local Development (Bare-Metal / uv + Vite)
|
||
|
||
For core development and live debugging without Docker:
|
||
|
||
**Requirements**: Python 3.12+, Astral [`uv`](https://docs.astral.sh/uv/), Node.js 20+, FFmpeg.
|
||
|
||
```bash
|
||
# 1. Clone and setup environment
|
||
git clone git@github.com:fhanyuh/reTraining.git
|
||
cd reTraining
|
||
cp .env.example .env
|
||
|
||
# 2. Install backend dependencies and vendored SAM3 with uv
|
||
uv pip install -r requirements.txt
|
||
uv pip install -e ./sam3
|
||
|
||
# 3. Start FastAPI backend (Port 8000)
|
||
uv run uvicorn backend.main:app --host 0.0.0.0 --port 8000 --reload
|
||
|
||
# 4. In a separate terminal, install and start Vite frontend (Port 5173)
|
||
cd frontend
|
||
npm install
|
||
npm run dev -- --host 0.0.0.0 --port 5173
|
||
```
|
||
|
||
> [!NOTE]
|
||
> The frontend dependencies are pinned to **Vite 7** (`vite: ^7.1.5`) in `frontend/package.json` to prevent Rolldown native binding bus errors on Linux platforms.
|
||
|
||
---
|
||
|
||
<a id="architecture"></a>
|
||
## 🏗️ End-to-End Pipeline Architecture
|
||
|
||
The platform operates as a continuous closed-loop retraining pipeline divided into 7 distinct functional stages:
|
||
|
||
```
|
||
┌──────────────────────────────────────────────────────────────────────────────────────────────────┐
|
||
│ DATA RETRAINING & INFERENCE PIPELINE │
|
||
└──────────────────────────────────────────────────────────────────────────────────────────────────┘
|
||
[1. Video Archive] ──▶ [2. Frame Slicing] ──▶ [3. SAM3 Auto-Label] ──▶ [4. Review & Exemplar]
|
||
(CCTV/MP4) (FFmpeg In/Out) (Prompt Grounding) (Click-Assist Canvas)
|
||
│
|
||
[7. Live Counter] ◀── [6. YOLO11 Train] ◀── [5. Immutable Split] ◀── [4. Triage & Filter]
|
||
(WHEP / RTSP) (Auto VRAM) (Stable Train/Val) (Outlier Purge)
|
||
```
|
||
|
||
1. **Video Archive & Shift Slicing**: Raw CCTV recordings are indexed into 24-hour operational work shifts (06:00 to 05:59 next morning). Sub-second timeline in/out trimming extracts high-resolution frame sequences with configurable FPS rates via asynchronous FFmpeg queues.
|
||
2. **SAM3 Foundation Auto-Labeling**: Meta SAM3 zero-shot open-vocabulary grounding generates candidate bounding boxes and segmentation masks from natural language descriptions (e.g. `"white sack of feed on conveyor"`).
|
||
3. **Interactive Review Studio**: Operators review candidate annotations on a Roboflow-grade canvas with single-keystroke approvals, box adjustments, click-assist segmentation, and visual prompt exemplar refinement.
|
||
4. **Statistical Triage & Quality Outliers**: Interactive 2D scatter plots (Box Area vs. Confidence Score) and high-density crop grids allow rapid isolation and pruning of false positives without discarding valid frames.
|
||
5. **Immutable Dataset Compilation**: Filtered batches are merged into versioned datasets (`v1`, `v2`, `v3`) with deterministic SHA-1 validation hashing, guaranteeing that validation images remain permanently locked across iterations.
|
||
6. **Hardware-Aware YOLO Retraining**: Hyperparameters and batch sizes auto-scale based on detected GPU VRAM. The system fine-tunes Ultralytics YOLO11, benchmarks old vs. new models on the identical validation set, and displays signed metric deltas ($\Delta\text{mAP50}$, $\Delta\text{Precision}$, $\Delta\text{Recall}$).
|
||
7. **Counting Accuracy Benchmark & Live Inference**: Real-time production inference evaluates live RTSP/WHEP video streams with ByteTrack trajectory tracking and upper-edge tripwire counters, benchmarking against hand-verified physical ground truth logs at **146 FPS**.
|
||
|
||
---
|
||
|
||
<a id="visual-tour"></a>
|
||
## 🖼️ Visual UI Tour & Feature Gallery
|
||
|
||
Explore the 6 core pipeline phases across all 11 primary user screens and modal workflows.
|
||
|
||
---
|
||
|
||
### Phase 1: Video Ingest, Shift Cycles & Frame Sampling
|
||
|
||
<details open>
|
||
<summary><b>1.1 Project Workspace & Setup</b> — Multi-project isolation and class locking</summary>
|
||
|
||
<br>
|
||
|
||
<div align="center">
|
||
<img src="screenshots/01_projects_page.png" width="100%" alt="Projects Overview Page" />
|
||
<p><i>Figure 1.1: Projects Dashboard showing active projects, base models, class taxonomies, and dataset stats.</i></p>
|
||
</div>
|
||
|
||
| Attribute | Specification |
|
||
|---|---|
|
||
| **🎯 Purpose** | Central management hub for all isolated computer vision projects. |
|
||
| **⚡ Key Capabilities** | View base model architecture, active class tags, total extracted batches, and dataset snapshots at a glance. |
|
||
| **💡 Invariant** | Projects maintain strictly isolated database records, class lists, and filesystem storage roots under `data/projects/<slug>/`. |
|
||
|
||
<br>
|
||
|
||
<div align="center">
|
||
<img src="screenshots/02_project_create_modal.png" width="100%" alt="Project Creation Modal" />
|
||
<p><i>Figure 1.2: Project Creation dialog with base model checkpoint upload and automatic class extraction.</i></p>
|
||
</div>
|
||
|
||
| Attribute | Specification |
|
||
|---|---|
|
||
| **🎯 Purpose** | Initialize a new project with locked class names and base model weights. |
|
||
| **⚡ Key Capabilities** | Upload an existing YOLO `.pt` checkpoint to automatically extract its class taxonomy, or define custom classes and start fine-tuning from `yolo11n.pt`. |
|
||
| **💡 Invariant** | Base model classes are permanently locked to the project to prevent label drift between training iterations. |
|
||
|
||
</details>
|
||
|
||
<details>
|
||
<summary><b>1.2 24-Hour Operational Shift Archive</b> — Cycle grouping that crosses midnight</summary>
|
||
|
||
<br>
|
||
|
||
<div align="center">
|
||
<img src="screenshots/03_library_video_archive.png" width="100%" alt="Video Archive Shift Cycle Browser" />
|
||
<p><i>Figure 1.3: Video Archive grouping recordings into 24-hour operational shifts (06:00 to 05:59 next morning).</i></p>
|
||
</div>
|
||
|
||
| Attribute | Specification |
|
||
|---|---|
|
||
| **🎯 Purpose** | Browse raw CCTV recordings grouped by operational work shifts rather than arbitrary calendar folders. |
|
||
| **⚡ Key Capabilities** | Reads camera burned-in OCR timestamps and sidecar `.json` metadata to assign recordings accurately across midnight boundaries. |
|
||
| **💡 Invariant** | The video archive is mounted **strictly read-only** (`:ro`). Nothing is ever modified, renamed, or deleted in the user's video repository. |
|
||
|
||
</details>
|
||
|
||
<details>
|
||
<summary><b>1.3 Video Trim & Frame Extractor</b> — Sub-second timeline scrubbing and sampling</summary>
|
||
|
||
<br>
|
||
|
||
<div align="center">
|
||
<img src="screenshots/04_trim_page.png" width="100%" alt="Video Trim and Frame Extractor" />
|
||
<p><i>Figure 1.4: Video Trim interface with timeline range sliders and real-time frame calculation.</i></p>
|
||
</div>
|
||
|
||
| Attribute | Specification |
|
||
|---|---|
|
||
| **🎯 Purpose** | Select active truck loading ranges and configure extraction frame rates. |
|
||
| **⚡ Key Capabilities** | Interactive in/out timeline markers, extraction FPS slider (0.5 – 2.0 FPS), real-time output frame counter, and background FFmpeg queueing. |
|
||
| **💡 Invariant** | Extraction runs asynchronously in the background queue; frames are losslessly sampled into `data/projects/<slug>/batches/<id>/frames/`. |
|
||
|
||
</details>
|
||
|
||
---
|
||
|
||
### Phase 2: SAM3 Zero-Shot Auto-Labeling Engine
|
||
|
||
<details open>
|
||
<summary><b>2.1 Batch Management & Mass Auto-Annotation</b> — High-throughput zero-shot grounding</summary>
|
||
|
||
<br>
|
||
|
||
<div align="center">
|
||
<img src="screenshots/05_batches_page.png" width="100%" alt="Batch Management Dashboard" />
|
||
<p><i>Figure 2.1: Batch Management table with batch status badges and bulk action triggers.</i></p>
|
||
</div>
|
||
|
||
<br>
|
||
|
||
<div align="center">
|
||
<img src="screenshots/06_batches_sam3_auto_annotate_modal.png" width="100%" alt="Single Batch SAM3 Auto-Annotate Modal" />
|
||
<p><i>Figure 2.2: Single-batch SAM3 auto-annotation configuration with natural language text prompts.</i></p>
|
||
</div>
|
||
|
||
<br>
|
||
|
||
<div align="center">
|
||
<img src="screenshots/07_batches_mass_auto_annotate_modal.png" width="100%" alt="Mass Auto-Annotate Modal" />
|
||
<p><i>Figure 2.3: Mass Auto-Annotate dialog for enqueuing thousands of frames across multiple batches.</i></p>
|
||
</div>
|
||
|
||
| Attribute | Specification |
|
||
|---|---|
|
||
| **🎯 Purpose** | Orchestrate Segment Anything Model 3 (SAM3) text-prompted auto-annotation across single or bulk batches. |
|
||
| **⚡ Key Capabilities** | Multi-prompt zero-shot grounding, adjustable confidence thresholds, box expansion margin, and background job queueing with VRAM singleton management. |
|
||
| **💡 Invariant** | **One `set_image` per frame**: SAM3 runs its heavy vision backbone once per image and re-runs only the lightweight grounding head across multiple prompts, ensuring maximum inference throughput. |
|
||
|
||
</details>
|
||
|
||
<details>
|
||
<summary><b>2.2 SAM3 Interactive Sandbox</b> — Standalone prompt engineering playground</summary>
|
||
|
||
<br>
|
||
|
||
<div align="center">
|
||
<img src="screenshots/20_sam3_playground_page.png" width="100%" alt="SAM3 Interactive Playground" />
|
||
<p><i>Figure 2.4: SAM3 Interactive Playground for zero-shot text prompting and point prompt testing.</i></p>
|
||
</div>
|
||
|
||
| Attribute | Specification |
|
||
|---|---|
|
||
| **🎯 Purpose** | Interactive sandbox for testing text prompts, positive/negative point clicks, and mask segmentation before launching large auto-annotation jobs. |
|
||
| **⚡ Key Capabilities** | Real-time mask rendering, multi-prompt layer toggles, point click prompt refinement, and raw JSON detection inspector. |
|
||
|
||
</details>
|
||
|
||
---
|
||
|
||
### Phase 3: High-Throughput Annotation Studio & Triage
|
||
|
||
<details open>
|
||
<summary><b>3.1 Roboflow-Grade Annotation Canvas & Quick Reclass</b> — Sub-second keyboard navigation</summary>
|
||
|
||
<br>
|
||
|
||
<div align="center">
|
||
<img src="screenshots/08_review_annotation_canvas.png" width="100%" alt="Annotation Review Canvas" />
|
||
<p><i>Figure 3.1: Annotation Review Canvas with bounding box editor, shape provenance badges, and hotkey controls.</i></p>
|
||
</div>
|
||
|
||
<br>
|
||
|
||
<div align="center">
|
||
<img src="screenshots/09_review_filmstrip_quick_reclass.png" width="100%" alt="Review Filmstrip and Quick Reclass Bar" />
|
||
<p><i>Figure 3.2: Filmstrip thumbnail navigation and floating single-keystroke quick reclassification bar.</i></p>
|
||
</div>
|
||
|
||
| Key | Action | | Key | Action |
|
||
|---|---|---|---|---|
|
||
| `A` | Approve frame | | `←` `→` | Previous / next frame |
|
||
| `X` | Reject frame | | `U` | Jump to next unreviewed |
|
||
| `Del` | Delete selected shape | | `1`–`9` | Quick switch active class |
|
||
| `S` + Drag | SAM3-assisted box click | | Drag | Add / move / resize bounding box |
|
||
|
||
| Attribute | Specification |
|
||
|---|---|
|
||
| **🎯 Purpose** | Fast, ergonomic manual review and refinement of auto-generated bounding boxes. |
|
||
| **⚡ Key Capabilities** | Single-key shortcuts, bottom thumbnail filmstrip with status indicators, and shape provenance tracking (`SAM3`, `Manual`, `Base Model`). |
|
||
| **💡 Invariant** | Approving frames does **not** merge them into a dataset. Approval merely qualifies frames for Data Prep; merging happens under frozen rules. |
|
||
|
||
</details>
|
||
|
||
<details>
|
||
<summary><b>3.2 Visual Exemplar-Guided Prompting</b> — Few-shot visual reference matching</summary>
|
||
|
||
<br>
|
||
|
||
<div align="center">
|
||
<img src="screenshots/12_review_exemplar_pool_panel.png" width="100%" alt="Exemplar Pool Panel" />
|
||
<p><i>Figure 3.3: Exemplar Pool Sidebar Panel for visual reference matching.</i></p>
|
||
</div>
|
||
|
||
| Attribute | Specification |
|
||
|---|---|
|
||
| **🎯 Purpose** | Store and utilize positive and negative visual crop exemplars to guide SAM3 zero-shot grounding on difficult textures or ambiguous sack designs. |
|
||
| **⚡ Key Capabilities** | Visual exemplar library, similarity threshold slider, 1-click exemplar addition from canvas bounding boxes. |
|
||
|
||
</details>
|
||
|
||
---
|
||
|
||
### Phase 4: Data Prep, Statistical Triage & Dataset Freezing
|
||
|
||
<details open>
|
||
<summary><b>4.1 Statistical Triage & Quality Outlier Filtering</b> — Filter bad boxes without losing full images</summary>
|
||
|
||
<br>
|
||
|
||
<div align="center">
|
||
<img src="screenshots/13_data_prep_quality_outliers.png" width="100%" alt="Data Prep Quality Outliers Panel" />
|
||
<p><i>Figure 4.1: Quality filter sliders with live box and frame retention counters.</i></p>
|
||
</div>
|
||
|
||
<br>
|
||
|
||
<div align="center">
|
||
<img src="screenshots/11_review_triage_scatter_plot.png" width="100%" alt="Triage Scatter Plot" />
|
||
<p><i>Figure 4.2: Interactive Scatter Plot (Score vs Area) for instant visual outlier cluster detection.</i></p>
|
||
</div>
|
||
|
||
<br>
|
||
|
||
<div align="center">
|
||
<img src="screenshots/10_review_triage_crop_grid.png" width="100%" alt="Triage Crop Grid" />
|
||
<p><i>Figure 4.3: Triage Crop Grid for rapid bulk visual inspection and 1-click outlier removal.</i></p>
|
||
</div>
|
||
|
||
| Attribute | Specification |
|
||
|---|---|
|
||
| **🎯 Purpose** | Eliminate low-quality bounding boxes (partial crops, false positives, background noise) before dataset compilation. |
|
||
| **⚡ Key Capabilities** | Dual-handle range sliders (Confidence Score, Box Area, Aspect Ratio), interactive SVG scatter plot, high-density crop grid cards, and real-time retention telemetry. |
|
||
| **💡 Invariant** | Dropping an outlier bounding box **leaves the frame in the dataset** unless all boxes are dropped. Industrial conveyor frames contain ~44 objects; dropping whole frames discards 96% of good data to remove 10% of bad boxes. |
|
||
|
||
</details>
|
||
|
||
<details>
|
||
<summary><b>4.2 Augmentation Pipeline & Dataset Freezing</b> — Deterministic Stable Val Split</summary>
|
||
|
||
<br>
|
||
|
||
<div align="center">
|
||
<img src="screenshots/14_data_prep_augmentation_panel.png" width="100%" alt="Augmentation Configuration Panel" />
|
||
<p><i>Figure 4.4: Albumentations training augmentation presets with real-time visual preview.</i></p>
|
||
</div>
|
||
|
||
<br>
|
||
|
||
<div align="center">
|
||
<img src="screenshots/15_data_prep_merge_target_modal.png" width="100%" alt="Merge Target and Dataset Freeze Modal" />
|
||
<p><i>Figure 4.5: Dataset Freeze dialog enforcing deterministic SHA-1 Stable Val Split rules.</i></p>
|
||
</div>
|
||
|
||
| Attribute | Specification |
|
||
|---|---|
|
||
| **🎯 Purpose** | Apply synthetic vision augmentations and freeze the curated batch into an immutable, versioned training dataset. |
|
||
| **⚡ Key Capabilities** | HSV shift, brightness/contrast, rotation, blur, and mosaic controls; dataset destination selection; split ratio slider. |
|
||
| **💡 Invariant** | **Stable Validation Split**: Validation assignment is derived deterministically from the frame's SHA-1 hash. Once a frame lands in `val`, it remains in `val` forever across all future versions. |
|
||
|
||
</details>
|
||
|
||
---
|
||
|
||
### Phase 5: YOLO Retraining & Live Training Progress
|
||
|
||
<details open>
|
||
<summary><b>5.1 Dataset Repository & YOLO Fine-Tuning</b> — Hardware-aware parameter tuning</summary>
|
||
|
||
<br>
|
||
|
||
<div align="center">
|
||
<img src="screenshots/16_datasets_page.png" width="100%" alt="Master Datasets List" />
|
||
<p><i>Figure 5.1: Master Dataset repository showing version history, train/val splits, and export tools.</i></p>
|
||
</div>
|
||
|
||
<br>
|
||
|
||
<div align="center">
|
||
<img src="screenshots/17_models_training_page.png" width="100%" alt="YOLO Models and Training Page" />
|
||
<p><i>Figure 5.2: YOLO fine-tuning parameter setup with hardware detection and model history.</i></p>
|
||
</div>
|
||
|
||
<br>
|
||
|
||
<div align="center">
|
||
<img src="screenshots/21_workflow_progress_states.png" width="100%" alt="Live Training Workflow Progress" />
|
||
<p><i>Figure 5.3: Real-time training telemetry with live loss curves, mAP metrics, and stdout logs.</i></p>
|
||
</div>
|
||
|
||
| Attribute | Specification |
|
||
|---|---|
|
||
| **🎯 Purpose** | Train Ultralytics YOLO11 object detection models with automated hardware tuning and live telemetry. |
|
||
| **⚡ Key Capabilities** | Architecture selection (`YOLO11n/s/m/l/x`), auto-calculated batch sizes based on free GPU VRAM, real-time loss sparklines (`box_loss`, `cls_loss`, `dfl_loss`), epoch progress bars, and streaming logs. |
|
||
| **💡 Invariant** | **Side-by-side Validation**: After training, both the base model and fine-tuned model are benchmarked on the identical validation set, displaying signed metric deltas ($\Delta\text{mAP50}$, $\Delta\text{precision}$, $\Delta\text{recall}$). |
|
||
|
||
</details>
|
||
|
||
---
|
||
|
||
### Phase 6: Counting Accuracy Benchmark & Live Production Inference
|
||
|
||
<details open>
|
||
<summary><b>6.1 Counting Benchmark Matrix & Live Line Counter</b> — 146 FPS verification</summary>
|
||
|
||
<br>
|
||
|
||
<div align="center">
|
||
<img src="screenshots/18_counting_bench_page.png" width="100%" alt="Counting Accuracy Benchmark Matrix" />
|
||
<p><i>Figure 6.1: Counting Accuracy Benchmark Matrix comparing multi-model predictions against Ground Truth.</i></p>
|
||
</div>
|
||
|
||
<br>
|
||
|
||
<div align="center">
|
||
<img src="screenshots/19_live_count_page.png" width="100%" alt="Live Video Inference and Line Counter" />
|
||
<p><i>Figure 6.2: Real-time CCTV live counting interface with interactive tripwire line and ByteTrack trails.</i></p>
|
||
</div>
|
||
|
||
| Attribute | Specification |
|
||
|---|---|
|
||
| **🎯 Purpose** | Verify model counting precision against physical ground truth records and run real-time production counting. |
|
||
| **⚡ Key Capabilities** | Headless counting benchmark running at **146 FPS**, signed delta badges (`+2`, `-1`, `0`), draggable tripwire counting line, ByteTrack trajectory tracking, and per-track JSONL diagnostic logs. |
|
||
| **💡 Invariant** | **Counting tripwire on upper edge ($y_1$)**: Tripwire evaluates the top edge coordinate of the sack bounding box rather than the center or bottom, preventing miscounts caused by physical sack deformation as it drops onto the conveyor. |
|
||
|
||
</details>
|
||
|
||
---
|
||
|
||
<a id="counting-engine"></a>
|
||
## 🎥 Production Counting Engine
|
||
|
||
The counting engine is designed for industrial conveyor belts with low or unstable camera frame rates. It eliminates reliance on instantaneous line-crossing frames by evaluating historical bounding box trajectories.
|
||
|
||
```mermaid
|
||
stateDiagram-v2
|
||
[*] --> UNKNOWN
|
||
UNKNOWN --> ABOVE: y1 coordinate above counting band
|
||
UNKNOWN --> BELOW: Born below band (ghost detection — rejected)
|
||
ABOVE --> COUNTED: Trajectory crosses band + travelled entry_travel_min
|
||
COUNTED --> ABOVE: Sustained frames above line (genuine reload)
|
||
```
|
||
|
||
### Multi-Layer Counter Safeguards
|
||
|
||
| Guard Layer | Rule Specification | Failure Mode Prevented |
|
||
|---|---|---|
|
||
| **Layer 1: Entry Origin Guard** | Object must be observed **above** the counting band prior to crossing. | Prevents counting items that spawn directly inside the truck or loading chute. |
|
||
| **Layer 2: Trajectory Distance** | Object must travel a minimum distance (`entry_travel_min`) across consecutive frames. | Discards transient noise and flickering phantom boxes. |
|
||
| **Layer 3: Track Hand-Off** | When a track is occluded, its movement history is parked for newborn tracks within `handoff_radius`. | Prevents tracker ID switches from dropping or duplicating counts. |
|
||
| **Layer 4: Directional Monotonicity** | Strict single count per trajectory direction with verdict locked to the final confirmed motion. | Prevents double-counting during momentary conveyor pauses. |
|
||
|
||
---
|
||
|
||
<a id="configuration"></a>
|
||
## ⚙️ Configuration & Environment Matrix
|
||
|
||
All configuration parameters are defined via environment variables in `.env`:
|
||
|
||
| Variable | Type | Default Value | Scope | Description |
|
||
|---|---|---|---|---|
|
||
| `HF_TOKEN` | `String` | *(Required)* | Backend / Docker | HuggingFace user access token with authorized permissions to download gated `facebook/sam3` weights. |
|
||
| `VIDEO_ARCHIVE_HOST` | `Path` | `./data/archive` | Docker Compose | Host filesystem directory path containing raw CCTV video recordings (structured as `<date>/<batch>.mp4`). Relative paths resolve against the directory `./start.sh` runs from. Mounted read-write; written only by user-initiated upload / date-folder creation (REQ-178), nothing else. |
|
||
| `APP_DATA_DIR` | `Path` | `./data` (or `/data` in Docker) | Backend | Base directory for application persistent data, SQLite database (`app.db`), projects, extracted frames, datasets, and model weights. |
|
||
| `VIDEO_ARCHIVE` | `Path` | `/videos` (or `data/archive`) | Backend | Internal container/local filesystem path where the video archive is browsed by FastAPI. |
|
||
| `WEB_PORT` | `Integer` | `9000` | Docker / Nginx | Host HTTP port mapped to the Nginx frontend web UI. |
|
||
| `API_PORT` | `Integer` | `9010` | Docker / FastAPI | Host HTTP port mapped to the FastAPI backend API. |
|
||
| `CORS_ORIGINS` | `String` | `http://localhost:5173,http://localhost:9000,http://localhost:9010` | FastAPI Backend | Comma-separated list of allowed origins for Cross-Origin Resource Sharing. |
|
||
| `API_URL` | `URL` | `http://localhost:8000` | Vite Dev Server | Backend target endpoint for Vite development proxy (`frontend/vite.config.js`). |
|
||
| `MEDIAMTX_WHEP_PATH` | `String` | `/whep` | Backend Live Count | WHEP WebRTC endpoint path on the streaming media server (MediaMTX). |
|
||
| `MEDIAMTX_RTSP_PORT` | `Integer` | `8554` | Backend Live Count | RTSP stream port used to translate WHEP browser streams into backend video processing feeds. |
|
||
| `RTSP_TRANSPORT` | `String` | `tcp` | Backend Live Count | RTSP transport protocol (`tcp` or `udp`). TCP guarantees zero frame drop on industrial networks. |
|
||
| `PLAYBACK_URL` | `URL` | `http://192.168.192.96:9996/get` | Recorder Service | MediaMTX recording playback API endpoint for automated CCTV session extraction. |
|
||
| `PLAYBACK_PATH` | `String` | `cam` | Recorder Service | Stream channel identifier on the MediaMTX playback server. |
|
||
| `AUTO_PULL_INTERVAL` | `Integer` | `30` | Auto-Pull Script | Polling frequency in seconds for automated Git repository synchronization (`scripts/auto_pull.py`). |
|
||
| `WEBHOOK_PORT` | `Integer` | `9000` | Webhook Daemon | Port for the GitHub push webhook listener daemon (`scripts/webhook.py`). |
|
||
| `WEBHOOK_SECRET` | `String` | `""` | Webhook Daemon | Shared secret key for validating GitHub webhook HMAC-SHA256 signatures. |
|
||
|
||
---
|
||
|
||
<a id="storage-layout"></a>
|
||
## 🗂️ Persistent Data & Storage Layout
|
||
|
||
All application state, relational metadata, and trained weights live in `data/`:
|
||
|
||
```
|
||
data/
|
||
├── app.db # SQLite metadata database with Write-Ahead Logging (WAL)
|
||
├── archive/<date>/batchNNN.mp4 # Raw CCTV video recordings (Mounted strictly read-only)
|
||
│ batchNNN.json # Sidecar metadata with true server timestamp
|
||
├── recorder.log # 24/7 background recorder daemon log
|
||
├── live-count/session-*.jsonl # Diagnostic per-track trajectory and crossing logs
|
||
└── projects/<slug>/ # Isolated project workspace
|
||
├── base/model.pt # Project base model weights and locked class taxonomy
|
||
├── batches/<id>/frames/ # Losslessly extracted image frames from video slices
|
||
├── datasets/<id>/ # Frozen Ultralytics YOLO formatted training datasets
|
||
│ ├── images/{train,val}/ # Immutable frame images
|
||
│ └── labels/{train,val}/ # YOLO format bounding box annotations (.txt)
|
||
└── models/<n>/ # Training runs (weights/best.pt, metrics.json, args.yaml)
|
||
# Named: {arch}-{labelType}-{epochs}ep-{classNames}-{date}
|
||
```
|
||
|
||
---
|
||
|
||
<a id="background-jobs"></a>
|
||
## ⚙️ Background Job Worker & Mutex Locking
|
||
|
||
Heavy computational operations run through an asynchronous background worker (`backend/jobs.py`) with strict GPU mutex locking to prevent VRAM over-allocation:
|
||
|
||
| Job Type | GPU Locked | Description |
|
||
|---|:---:|---|
|
||
| `extract` | — | Background FFmpeg extraction of video ranges into frame sequences. |
|
||
| `autolabel` | ✅ | Meta SAM3 zero-shot grounding across candidate frames (one `set_image` per image). |
|
||
| `merge` | — | Compiles reviewed batches into immutable dataset splits under frozen triage rules. |
|
||
| `train` | ✅ | Ultralytics YOLO11 fine-tuning followed by automated dual-model validation. |
|
||
| `count` | ✅ | Headless evaluation benchmark running archive videos at **146 FPS**. |
|
||
| `clock-scan` | — | OCR extraction of camera burned-in timestamps. |
|
||
| `truck-scan` | ✅ | Batch inference check verifying the presence of target industrial objects. |
|
||
|
||
---
|
||
|
||
<a id="invariants"></a>
|
||
## 📊 Metrics Integrity & Domain Invariants
|
||
|
||
1. **Stable Validation Split**: Validation assignment is determined deterministically by `sha1(image_bytes) % 100 < val_ratio`. Once a frame lands in the validation split, it remains in validation across all future dataset iterations. This prevents validation leak and ensures that rising mAP scores reflect genuine model improvements.
|
||
2. **One `set_image` per Frame**: SAM3 executes its heavy vision transformer backbone once per image. Multi-class text prompting evaluates the lightweight grounding head against cached backbone embeddings.
|
||
3. **Outlier Filtering Preserves Frames**: Dropping a bounding box during Data Prep removes only the bad annotation. The image frame remains in the dataset as long as at least one valid box persists.
|
||
4. **Upper-Edge Coordinate Line Crossing ($y_1$)**: Tripwires evaluate the top edge coordinate of bounding boxes ($y_1$) rather than the centroid ($y_c$) or bottom edge ($y_2$), ensuring immunity to sack deformation upon conveyor impact.
|
||
|
||
---
|
||
|
||
<a id="documentation"></a>
|
||
## 📚 Documentation & Deep Dives
|
||
|
||
Comprehensive technical specifications, operational SOPs, and architecture diagrams are available in the [`docs/`](docs/) directory:
|
||
|
||
| Document | Format | Description | Action |
|
||
|---|:---:|---|:---:|
|
||
| **Panduan Sistem Lengkap (Handover)** | `PDF` (11 MB) | **Publication-Grade Master User Guide & Technical Manual** (Indonesian) with complete operational SOPs and embedded figures. | [**⬇️ Download PDF**](docs/PANDUAN_SISTEM_LENGKAP.pdf) |
|
||
| **Panduan Sistem Lengkap Source** | `FODT` (11.8 MB) | Native LibreOffice Writer Flat XML editable source document. | [**⬇️ Download FODT**](docs/PANDUAN_SISTEM_LENGKAP.fodt) |
|
||
| **Panduan Sistem Lengkap Markdown** | `MD` (68 KB) | Full Markdown transcript of the 14-chapter system manual. | [**📖 View MD**](docs/PANDUAN_SISTEM_LENGKAP.md) |
|
||
| **Architecture Flow Diagram (4K)** | `PNG` (1.5 MB) | 4K Ultra-HD raster export of the 8-stage end-to-end retraining pipeline. | [**⬇️ Download 4K**](docs/diagram-alur.png) |
|
||
| **Architecture Flow Diagram (Vector)** | `SVG` (34.6 KB) | Scalable vector graphic diagram for high-resolution display. | [**⬇️ Download SVG**](docs/diagram-alur.svg) |
|
||
| **Architecture Flow Diagram (Source)** | `FODG` (24.8 KB) | Native LibreOffice Draw Flat XML editable source file. | [**⬇️ Download FODG**](docs/diagram-alur.fodg) |
|
||
| **Entity Relationship Diagram (ERD)** | `MD` (8.5 KB) | Relational database schema, foreign keys, indexes, and Mermaid ERD diagram. | [**📖 View ERD**](ERD.md) |
|
||
| **System Requirements Specification** | `MD` (23.6 KB) | Numbered technical requirements (`REQ-001` through `REQ-042`). | [**📖 View Spec**](docs/requirements.md) |
|
||
| **System Design & Architecture** | `MD` (22.7 KB) | Database schema, REST API contracts, disk layouts, and backend invariants. | [**📖 View Design**](docs/design.md) |
|
||
| **UI/UX Design Specification** | `MD` (79.7 KB) | Dark theme design tokens, hotkey maps, and component specifications. | [**📖 View UI Spec**](docs/ui-spec.md) |
|
||
| **Ground Truth Benchmark Dataset** | `XLSX` (586 KB) | Hand-verified physical conveyor bag counts across operational shifts. | [**⬇️ Download XLSX**](docs/GT.xlsx) |
|
||
|
||
---
|
||
|
||
<a id="troubleshooting"></a>
|
||
## 🔧 Troubleshooting & FAQ
|
||
|
||
<details>
|
||
<summary><b>Deployment & GPU Acceleration</b></summary>
|
||
|
||
<br>
|
||
|
||
| Symptom | Root Cause & Remediation |
|
||
|---|---|
|
||
| `could not select device driver` | NVIDIA Container Toolkit is missing or Docker Engine is older than CDI specifications. Run `./install_nvidia.sh` or update Docker. |
|
||
| `CUDA out of memory` during training | Lower the batch size on the Models page or stop background jobs. SAM3 releases VRAM before training starts, but external processes may hold memory. |
|
||
| Changes to frontend or backend do not appear | Docker Compose caches container layers at build time. Run `docker compose build backend frontend` and hard-refresh your browser (`Ctrl+Shift+R`). |
|
||
| Job shows *interrupted by server restart* | The backend process stopped while a job was active. Jobs do not resume mid-epoch; simply re-trigger the job from the UI. |
|
||
|
||
</details>
|
||
|
||
<details>
|
||
<summary><b>SAM3 Foundation Auto-Annotation</b></summary>
|
||
|
||
<br>
|
||
|
||
| Symptom | Root Cause & Remediation |
|
||
|---|---|
|
||
| Job fails at *Loading Model* with HTTP 401 | HuggingFace token is invalid or access to [facebook/sam3](https://huggingface.co/facebook/sam3) has not yet been approved. |
|
||
| SAM3 checkpoint download is slow | HuggingFace Xet transfer throttling. `HF_HUB_DISABLE_XET=1` is enabled by default in `docker-compose.yml` to bypass this issue. |
|
||
| SAM3 generates inaccurate bounding boxes | Refine the natural language prompt with physical descriptors (e.g. `"woven polypropylene sack with blue logo"`) or add visual crops to the Exemplar Pool. |
|
||
|
||
</details>
|
||
|
||
<details>
|
||
<summary><b>Video Archive & Live Counting</b></summary>
|
||
|
||
<br>
|
||
|
||
| Symptom | Root Cause & Remediation |
|
||
|---|---|
|
||
| Video file marked as *unreadable* | FFmpeg/FFprobe could not parse the video header. Run `uv run python scripts/transcode_archive.py` to re-mux into standard H.264 MP4. |
|
||
| Shift cycle shows fewer recordings than expected | Check if recordings crossed the 06:00 boundary. Clips recorded before 06:00 belong to the previous operational shift cycle. |
|
||
| Live Counting shows *GPU busy* | An active fine-tuning or auto-annotation job holds the GPU lock. The live counter waits 30 seconds before falling back to CPU or queuing. |
|
||
|
||
</details>
|
||
|
||
---
|
||
|
||
<a id="contributing"></a>
|
||
## 🤝 Contributing & Engineering Guidelines
|
||
|
||
Development adheres strictly to the **Chain of Truth** methodology:
|
||
- **Specifications First**: Every feature must map directly to a numbered requirement in [`docs/requirements.md`](docs/requirements.md) and architectural design in [`docs/design.md`](docs/design.md).
|
||
- **Verified Deliverables**: Tasks tracked in [`docs/tasks.md`](docs/tasks.md) flip to `[DONE]` only after concrete end-to-end verification.
|
||
- **Surgical Changes**: Touch only code directly relevant to the feature. Adhere to the working rules in [`AGENTS.md`](AGENTS.md).
|
||
- **Package Manager**: All backend dependencies are managed exclusively with Astral `uv` (`requirements.txt`).
|
||
|
||
---
|
||
|
||
<a id="license"></a>
|
||
## 📄 License
|
||
|
||
Distributed under the MIT License. See [`LICENSE`](LICENSE) for details.
|