docs: transform README into SaaS-grade reference manual and optimize repository hygiene
This commit is contained in:
1 parent
ac95674c07
commit
6485f59e26
43 files changed
+8880
-395
No files matched your search
@@ -2,443 +2,659 @@
|
||||
|
||||
# reTraining
|
||||
|
||||
**Take a model you already have, and make it better with footage you already have.**
|
||||
### Take a base model you already have, and make it measurably better with footage you already have.
|
||||
|
||||
A self-hosted loop for turning raw CCTV into a measurably better detector — record, extract,
|
||||
auto-label, review, filter, train, and prove the new model actually beat the old one.
|
||||
**The Open-Source, Self-Hosted Vision Pipeline to Turn Raw Industrial CCTV into Production-Grade YOLO Object Detectors & Line Counters.**
|
||||
|
||||
<p>
|
||||
<img alt="Python" src="https://img.shields.io/badge/python-3.12-3776AB?logo=python&logoColor=white">
|
||||
<img alt="FastAPI" src="https://img.shields.io/badge/FastAPI-72%20endpoints-009688?logo=fastapi&logoColor=white">
|
||||
<img alt="React" src="https://img.shields.io/badge/React-19-61DAFB?logo=react&logoColor=black">
|
||||
<img alt="Vite" src="https://img.shields.io/badge/Vite-7-646CFF?logo=vite&logoColor=white">
|
||||
<img alt="SAM3" src="https://img.shields.io/badge/SAM3-auto--label-FF6F00">
|
||||
<img alt="YOLO" src="https://img.shields.io/badge/Ultralytics-YOLO11-00BFA5">
|
||||
<img alt="Docker" src="https://img.shields.io/badge/docker-compose-2496ED?logo=docker&logoColor=white">
|
||||
<img alt="GPU" src="https://img.shields.io/badge/GPU-required-76B900?logo=nvidia&logoColor=white">
|
||||
<a href="https://www.python.org/"><img alt="Python 3.12" src="https://img.shields.io/badge/python-3.12-3776AB?logo=python&logoColor=white"></a>
|
||||
<a href="https://fastapi.tiangolo.com/"><img alt="FastAPI" src="https://img.shields.io/badge/FastAPI-72%20endpoints-009688?logo=fastapi&logoColor=white"></a>
|
||||
<a href="https://react.dev/"><img alt="React 19" src="https://img.shields.io/badge/React-19-61DAFB?logo=react&logoColor=black"></a>
|
||||
<a href="https://vitejs.dev/"><img alt="Vite 7" src="https://img.shields.io/badge/Vite-7-646CFF?logo=vite&logoColor=white"></a>
|
||||
<a href="https://huggingface.co/facebook/sam3"><img alt="SAM3" src="https://img.shields.io/badge/SAM3-zero--shot-FF6F00"></a>
|
||||
<a href="https://github.com/ultralytics/ultralytics"><img alt="YOLO11" src="https://img.shields.io/badge/Ultralytics-YOLO11-00BFA5"></a>
|
||||
<a href="https://www.docker.com/"><img alt="Docker" src="https://img.shields.io/badge/Docker-Compose-2496ED?logo=docker&logoColor=white"></a>
|
||||
<a href="https://developer.nvidia.com/cuda-toolkit"><img alt="CUDA" src="https://img.shields.io/badge/CUDA-12.4+-76B900?logo=nvidia&logoColor=white"></a>
|
||||
<a href="LICENSE"><img alt="License" src="https://img.shields.io/badge/license-MIT-blue.svg"></a>
|
||||
<a href="docs/PANDUAN_SISTEM_LENGKAP.pdf"><img alt="Documentation" src="https://img.shields.io/badge/Docs-PDF%20Guide-red?logo=adobe-acrobat-reader&logoColor=white"></a>
|
||||
</p>
|
||||
|
||||
[Quick start](#-quick-start) · [How it works](#-how-it-works) · [The screens](#-the-screens) ·
|
||||
[Counting](#-counting) · [Trust the numbers](#-why-the-numbers-are-trustworthy) ·
|
||||
[Troubleshooting](#-when-something-goes-wrong)
|
||||
[⚡ Quickstart](#quickstart) · [🏗️ Architecture](#architecture) · [🖼️ Visual UI Tour](#visual-tour) · [🎥 Counting Engine](#counting-engine) · [⚙️ Configuration](#configuration) · [📚 Documentation](#documentation) · [🔧 Troubleshooting](#troubleshooting)
|
||||
|
||||
<br>
|
||||
|
||||
<a href="docs/diagram-alur.png">
|
||||
<img src="docs/diagram-alur.png" width="100%" alt="System Architecture & End-to-End Pipeline Diagram" />
|
||||
</a>
|
||||
|
||||
*Click diagram to view 4K UHD resolution. Scalable vector version available at [docs/diagram-alur.svg](docs/diagram-alur.svg).*
|
||||
|
||||
</div>
|
||||
|
||||
---
|
||||
|
||||
## ⚡ Quick start
|
||||
<a id="why-retraining"></a>
|
||||
## 💡 Why reTraining?
|
||||
|
||||
Deploying object detection models in industrial environments (manufacturing, logistics, agricultural feedmills, conveyor belts) often hits a painful bottleneck: **general base models fail on domain-specific edge cases, while building manual labeling pipelines from scratch is slow and expensive.**
|
||||
|
||||
**reTraining** provides a self-contained, enterprise-grade active learning platform designed to run directly on your edge server or GPU workstation:
|
||||
|
||||
1. **Ingest Raw CCTV Footage**: Stream directly from continuous 24/7 video archives without re-encoding or modifying the underlying storage.
|
||||
2. **Zero-Shot Foundation Auto-Labeling**: Leverage Meta's Segment Anything Model 3 (**SAM3**) with natural language text prompts and visual exemplars to annotate thousands of frames in minutes.
|
||||
3. **Roboflow-Grade Review Studio**: Sub-second keyboard navigation, instant class switching, click-assist segmentation, and high-density visual triage crop grids.
|
||||
4. **Statistical Triage & Data Prep**: Filter bounding box outliers by area, aspect ratio, and confidence score without discarding valid image frames.
|
||||
5. **Immutable Dataset Freezing & Stable Val Splits**: Guarantee reproducible benchmarks with deterministic SHA-1 validation sets that never shift across retraining runs.
|
||||
6. **Hardware-Aware Continuous Fine-Tuning**: Auto-detect host GPU VRAM, fine-tune Ultralytics YOLO11 models, and evaluate base vs. fine-tuned model performance side-by-side.
|
||||
7. **Live Production Counting & Benchmarks**: Real-time RTSP/WHEP live inference with ByteTrack line-crossing counters, evaluated against verified ground truth physical counts at **146 FPS**.
|
||||
|
||||
---
|
||||
|
||||
<a id="quickstart"></a>
|
||||
## ⚡ Quickstart
|
||||
|
||||
### Prerequisites
|
||||
|
||||
- **Host OS**: Linux (Ubuntu 22.04+ recommended) or Windows with WSL2.
|
||||
- **GPU Acceleration**: NVIDIA GPU with CUDA 12.4+ and [NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html) (Docker Engine 27+ CDI support).
|
||||
- **HuggingFace Account**: Gated model access granted for [facebook/sam3](https://huggingface.co/facebook/sam3) with a valid user access token (`HF_TOKEN`).
|
||||
- **Video Storage**: Directory of CCTV footage structured as `<date>/<batch>.mp4` (or let the app mount `./data/archive`).
|
||||
|
||||
---
|
||||
|
||||
### Option 1: Docker Compose (Production Standard)
|
||||
|
||||
Launch the entire stack with a single command. The startup script automatically inspects the host hardware, generates Container Device Interface (CDI) specs for your NVIDIA GPU, and starts both backend and frontend containers.
|
||||
|
||||
```bash
|
||||
cp .env.example .env # paste your HF_TOKEN
|
||||
VIDEO_ARCHIVE_HOST=/path/to/videos docker compose up -d --build
|
||||
# 1. Clone the repository
|
||||
git clone git@github.com:fhanyuh/reTraining.git
|
||||
cd reTraining
|
||||
|
||||
# 2. Configure environment credentials
|
||||
cp .env.example .env
|
||||
# Edit .env and paste your HuggingFace user access token:
|
||||
# HF_TOKEN=hf_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
|
||||
|
||||
# 3. Launch with automated GPU / CDI detection
|
||||
chmod +x start.sh
|
||||
./start.sh
|
||||
```
|
||||
|
||||
Open **<http://localhost:8080>**. The API is on `:8000`.
|
||||
|
||||
*Alternative direct Docker Compose launch:*
|
||||
```bash
|
||||
curl localhost:8000/api/health
|
||||
docker compose up -d --build
|
||||
```
|
||||
|
||||
**Access Points**:
|
||||
- 🌐 **Web Studio UI**: [http://localhost:8080](http://localhost:8080)
|
||||
- 📑 **Interactive REST API Docs**: [http://localhost:8000/docs](http://localhost:8000/docs)
|
||||
- 🩺 **Health & GPU Telemetry Endpoint**: [http://localhost:8000/api/health](http://localhost:8000/api/health)
|
||||
|
||||
Verify system health and GPU VRAM availability:
|
||||
```bash
|
||||
curl http://localhost:8000/api/health
|
||||
```
|
||||
```json
|
||||
{ "device": "cuda", "gpu": "NVIDIA GeForce RTX 5080 Laptop GPU",
|
||||
"vram_free_gb": 14.91, "sam3_ready": true, "ffmpeg": true,
|
||||
"hf_token": true, "db": true }
|
||||
{
|
||||
"device": "cuda",
|
||||
"gpu": "NVIDIA GeForce RTX 5080 Laptop GPU",
|
||||
"vram_free_gb": 14.91,
|
||||
"sam3_ready": true,
|
||||
"ffmpeg": true,
|
||||
"hf_token": true,
|
||||
"db": true
|
||||
}
|
||||
```
|
||||
|
||||
> [!IMPORTANT]
|
||||
> The first auto-annotation job downloads the **~3.4 GB SAM3 checkpoint** into a Docker
|
||||
> volume. It happens once; later jobs take about 12 seconds to load the model into VRAM.
|
||||
|
||||
<details>
|
||||
<summary><b>What you need first</b></summary>
|
||||
|
||||
<br>
|
||||
|
||||
| Thing | Why |
|
||||
|---|---|
|
||||
| A GPU with the NVIDIA container toolkit | SAM3 is CUDA-only |
|
||||
| A video archive laid out as `<date>/<batch>.mp4` | that structure is what the archive browser reads |
|
||||
| `HF_TOKEN` with access to [facebook/sam3](https://huggingface.co/facebook/sam3) | the weights are gated, and approval is manual |
|
||||
| A base model `.pt` (optional) | without one, training starts from `yolo11n.pt` and you type the classes yourself |
|
||||
|
||||
```
|
||||
videos/
|
||||
2026-08-13/
|
||||
batch001.mp4
|
||||
batch002.mp4
|
||||
2026-08-14/
|
||||
batch001.mp4
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>Run it without Docker (development)</b></summary>
|
||||
|
||||
<br>
|
||||
|
||||
```bash
|
||||
uv pip install -r requirements.txt # backend
|
||||
uv pip install -e sam3/
|
||||
uv run uvicorn backend.main:app --reload # :8000
|
||||
|
||||
cd frontend && npm install
|
||||
npm run dev # :5173, proxies /api to :8000
|
||||
```
|
||||
|
||||
The frontend pins Vite 7 on purpose — Vite 8's Rolldown binding crashes on this machine.
|
||||
The package manager is `uv`; there is no `pip`/`poetry` path.
|
||||
|
||||
</details>
|
||||
> The initial SAM3 auto-annotation job automatically downloads the **~3.4 GB SAM3 checkpoint** from HuggingFace into a persistent Docker named volume (`hf-cache`). Subsequent executions load the model into VRAM in ~12 seconds.
|
||||
|
||||
---
|
||||
|
||||
## 🔄 How it works
|
||||
### Option 2: Local Development (Bare-Metal / uv + Vite)
|
||||
|
||||
One project owns its base model, its class list, its video archive and its own accumulating
|
||||
datasets. A second use case is a second project — not a second copy of the code.
|
||||
For core development and live debugging without Docker:
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
A[📹 Archive<br/>one file per truck session] --> B[✂️ Trim<br/>pick a range + fps]
|
||||
B --> C[🖼️ Extract<br/>frames to disk]
|
||||
C --> D[🤖 Auto-label<br/>SAM3 text prompts]
|
||||
D --> E[👁️ Review<br/>fix every frame]
|
||||
E --> F[🧹 Data Prep<br/>filter + augment]
|
||||
F --> G[📦 Dataset<br/>named, immutable]
|
||||
G --> H[🎯 Train<br/>fine-tune from base]
|
||||
H --> I[📊 Compare<br/>base vs new, same val set]
|
||||
I -.->|promote| H
|
||||
**Requirements**: Python 3.12+, Astral [`uv`](https://docs.astral.sh/uv/), Node.js 20+, FFmpeg.
|
||||
|
||||
```bash
|
||||
# 1. Clone and setup environment
|
||||
git clone git@github.com:fhanyuh/reTraining.git
|
||||
cd reTraining
|
||||
cp .env.example .env
|
||||
|
||||
# 2. Install backend dependencies and vendored SAM3 with uv
|
||||
uv pip install -r requirements.txt
|
||||
uv pip install -e ./sam3
|
||||
|
||||
# 3. Start FastAPI backend (Port 8000)
|
||||
uv run uvicorn backend.main:app --host 0.0.0.0 --port 8000 --reload
|
||||
|
||||
# 4. In a separate terminal, install and start Vite frontend (Port 5173)
|
||||
cd frontend
|
||||
npm install
|
||||
npm run dev -- --host 0.0.0.0 --port 5173
|
||||
```
|
||||
|
||||
> [!NOTE]
|
||||
> **Data Prep is the gate.** Selecting batches does not create a dataset — it opens Data Prep
|
||||
> scoped to that selection. Only *Confirm merge* cuts the dataset, and the filter rules in
|
||||
> force at that moment are **frozen onto it**, so editing them later can never rewrite a
|
||||
> dataset you already trained on.
|
||||
> The frontend dependencies are pinned to **Vite 7** (`vite: ^7.1.5`) in `frontend/package.json` to prevent Rolldown native binding bus errors on Linux platforms.
|
||||
|
||||
---
|
||||
|
||||
## 🖥️ The screens
|
||||
<a id="architecture"></a>
|
||||
## 🏗️ End-to-End Pipeline Architecture
|
||||
|
||||
The platform operates as a continuous closed-loop retraining pipeline divided into 7 distinct functional stages:
|
||||
|
||||
```
|
||||
┌──────────────────────────────────────────────────────────────────────────────────────────────────┐
|
||||
│ DATA RETRAINING & INFERENCE PIPELINE │
|
||||
└──────────────────────────────────────────────────────────────────────────────────────────────────┘
|
||||
[1. Video Archive] ──▶ [2. Frame Slicing] ──▶ [3. SAM3 Auto-Label] ──▶ [4. Review & Exemplar]
|
||||
(CCTV/MP4) (FFmpeg In/Out) (Prompt Grounding) (Click-Assist Canvas)
|
||||
│
|
||||
[7. Live Counter] ◀── [6. YOLO11 Train] ◀── [5. Immutable Split] ◀── [4. Triage & Filter]
|
||||
(WHEP / RTSP) (Auto VRAM) (Stable Train/Val) (Outlier Purge)
|
||||
```
|
||||
|
||||
1. **Video Archive & Shift Slicing**: Raw CCTV recordings are indexed into 24-hour operational work shifts (06:00 to 05:59 next morning). Sub-second timeline in/out trimming extracts high-resolution frame sequences with configurable FPS rates via asynchronous FFmpeg queues.
|
||||
2. **SAM3 Foundation Auto-Labeling**: Meta SAM3 zero-shot open-vocabulary grounding generates candidate bounding boxes and segmentation masks from natural language descriptions (e.g. `"white sack of feed on conveyor"`).
|
||||
3. **Interactive Review Studio**: Operators review candidate annotations on a Roboflow-grade canvas with single-keystroke approvals, box adjustments, click-assist segmentation, and visual prompt exemplar refinement.
|
||||
4. **Statistical Triage & Quality Outliers**: Interactive 2D scatter plots (Box Area vs. Confidence Score) and high-density crop grids allow rapid isolation and pruning of false positives without discarding valid frames.
|
||||
5. **Immutable Dataset Compilation**: Filtered batches are merged into versioned datasets (`v1`, `v2`, `v3`) with deterministic SHA-1 validation hashing, guaranteeing that validation images remain permanently locked across iterations.
|
||||
6. **Hardware-Aware YOLO Retraining**: Hyperparameters and batch sizes auto-scale based on detected GPU VRAM. The system fine-tunes Ultralytics YOLO11, benchmarks old vs. new models on the identical validation set, and displays signed metric deltas ($\Delta\text{mAP50}$, $\Delta\text{Precision}$, $\Delta\text{Recall}$).
|
||||
7. **Counting Accuracy Benchmark & Live Inference**: Real-time production inference evaluates live RTSP/WHEP video streams with ByteTrack trajectory tracking and upper-edge tripwire counters, benchmarking against hand-verified physical ground truth logs at **146 FPS**.
|
||||
|
||||
---
|
||||
|
||||
<a id="visual-tour"></a>
|
||||
## 🖼️ Visual UI Tour & Feature Gallery
|
||||
|
||||
Explore the 6 core pipeline phases across all 11 primary user screens and modal workflows.
|
||||
|
||||
---
|
||||
|
||||
### Phase 1: Video Ingest, Shift Cycles & Frame Sampling
|
||||
|
||||
<details open>
|
||||
<summary><b>1. Projects</b> — name, label type, archive root, classes</summary>
|
||||
<summary><b>1.1 Project Workspace & Setup</b> — Multi-project isolation and class locking</summary>
|
||||
|
||||
<br>
|
||||
|
||||
Upload a base model and its classes are read from the checkpoint and locked, so the dataset
|
||||
and the model can never drift apart.
|
||||
<div align="center">
|
||||
<img src="screenshots/01_projects_page.png" width="100%" alt="Projects Overview Page" />
|
||||
<p><i>Figure 1.1: Projects Dashboard showing active projects, base models, class taxonomies, and dataset stats.</i></p>
|
||||
</div>
|
||||
|
||||
| Attribute | Specification |
|
||||
|---|---|
|
||||
| **🎯 Purpose** | Central management hub for all isolated computer vision projects. |
|
||||
| **⚡ Key Capabilities** | View base model architecture, active class tags, total extracted batches, and dataset snapshots at a glance. |
|
||||
| **💡 Invariant** | Projects maintain strictly isolated database records, class lists, and filesystem storage roots under `data/projects/<slug>/`. |
|
||||
|
||||
<br>
|
||||
|
||||
<div align="center">
|
||||
<img src="screenshots/02_project_create_modal.png" width="100%" alt="Project Creation Modal" />
|
||||
<p><i>Figure 1.2: Project Creation dialog with base model checkpoint upload and automatic class extraction.</i></p>
|
||||
</div>
|
||||
|
||||
| Attribute | Specification |
|
||||
|---|---|
|
||||
| **🎯 Purpose** | Initialize a new project with locked class names and base model weights. |
|
||||
| **⚡ Key Capabilities** | Upload an existing YOLO `.pt` checkpoint to automatically extract its class taxonomy, or define custom classes and start fine-tuning from `yolo11n.pt`. |
|
||||
| **💡 Invariant** | Base model classes are permanently locked to the project to prevent label drift between training iterations. |
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>2. Video Archive</b> — browse by <i>cycle</i>, not by folder</summary>
|
||||
<summary><b>1.2 24-Hour Operational Shift Archive</b> — Cycle grouping that crosses midnight</summary>
|
||||
|
||||
<br>
|
||||
|
||||
A **cycle** is one shift: `06:00 → 05:59` the next morning. It always crosses midnight, so it
|
||||
always spans two calendar dates, and is named after the date it starts on.
|
||||
<div align="center">
|
||||
<img src="screenshots/03_library_video_archive.png" width="100%" alt="Video Archive Shift Cycle Browser" />
|
||||
<p><i>Figure 1.3: Video Archive grouping recordings into 24-hour operational shifts (06:00 to 05:59 next morning).</i></p>
|
||||
</div>
|
||||
|
||||
Folder names are not when a recording was made, and neither are file mtimes — those are file
|
||||
*copy* times. The real start comes from the timestamp the camera burns into every frame, so a
|
||||
file sitting in the `2026-08-14` folder but recorded at `00:11` shows up as part of the
|
||||
**13 Aug** cycle, numbered in the order it was actually made.
|
||||
|
||||
```
|
||||
Siklus 13 Agt 2026 28 rekaman
|
||||
#1 08:27:27 batch003 2026-08-13
|
||||
...
|
||||
#25 23:53:45 batch027 2026-08-13
|
||||
#26 00:11:02 batch001 2026-08-14 ← pulled in from the next folder
|
||||
#27 00:34:01 batch002 2026-08-14
|
||||
```
|
||||
|
||||
Nothing in the archive is moved or renamed — it is mounted **read-only**. The grouping lives
|
||||
in an index beside it, and the original path stays the file's identity.
|
||||
| Attribute | Specification |
|
||||
|---|---|
|
||||
| **🎯 Purpose** | Browse raw CCTV recordings grouped by operational work shifts rather than arbitrary calendar folders. |
|
||||
| **⚡ Key Capabilities** | Reads camera burned-in OCR timestamps and sidecar `.json` metadata to assign recordings accurately across midnight boundaries. |
|
||||
| **💡 Invariant** | The video archive is mounted **strictly read-only** (`:ro`). Nothing is ever modified, renamed, or deleted in the user's video repository. |
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>3. Trim</b> — play the video, set in/out, pick a frame rate</summary>
|
||||
<summary><b>1.3 Video Trim & Frame Extractor</b> — Sub-second timeline scrubbing and sampling</summary>
|
||||
|
||||
<br>
|
||||
|
||||
It tells you how many frames that produces before you commit to it.
|
||||
<div align="center">
|
||||
<img src="screenshots/04_trim_page.png" width="100%" alt="Video Trim and Frame Extractor" />
|
||||
<p><i>Figure 1.4: Video Trim interface with timeline range sliders and real-time frame calculation.</i></p>
|
||||
</div>
|
||||
|
||||
| Attribute | Specification |
|
||||
|---|---|
|
||||
| **🎯 Purpose** | Select active truck loading ranges and configure extraction frame rates. |
|
||||
| **⚡ Key Capabilities** | Interactive in/out timeline markers, extraction FPS slider (0.5 – 2.0 FPS), real-time output frame counter, and background FFmpeg queueing. |
|
||||
| **💡 Invariant** | Extraction runs asynchronously in the background queue; frames are losslessly sampled into `data/projects/<slug>/batches/<id>/frames/`. |
|
||||
|
||||
</details>
|
||||
|
||||
---
|
||||
|
||||
### Phase 2: SAM3 Zero-Shot Auto-Labeling Engine
|
||||
|
||||
<details open>
|
||||
<summary><b>2.1 Batch Management & Mass Auto-Annotation</b> — High-throughput zero-shot grounding</summary>
|
||||
|
||||
<br>
|
||||
|
||||
<div align="center">
|
||||
<img src="screenshots/05_batches_page.png" width="100%" alt="Batch Management Dashboard" />
|
||||
<p><i>Figure 2.1: Batch Management table with batch status badges and bulk action triggers.</i></p>
|
||||
</div>
|
||||
|
||||
<br>
|
||||
|
||||
<div align="center">
|
||||
<img src="screenshots/06_batches_sam3_auto_annotate_modal.png" width="100%" alt="Single Batch SAM3 Auto-Annotate Modal" />
|
||||
<p><i>Figure 2.2: Single-batch SAM3 auto-annotation configuration with natural language text prompts.</i></p>
|
||||
</div>
|
||||
|
||||
<br>
|
||||
|
||||
<div align="center">
|
||||
<img src="screenshots/07_batches_mass_auto_annotate_modal.png" width="100%" alt="Mass Auto-Annotate Modal" />
|
||||
<p><i>Figure 2.3: Mass Auto-Annotate dialog for enqueuing thousands of frames across multiple batches.</i></p>
|
||||
</div>
|
||||
|
||||
| Attribute | Specification |
|
||||
|---|---|
|
||||
| **🎯 Purpose** | Orchestrate Segment Anything Model 3 (SAM3) text-prompted auto-annotation across single or bulk batches. |
|
||||
| **⚡ Key Capabilities** | Multi-prompt zero-shot grounding, adjustable confidence thresholds, box expansion margin, and background job queueing with VRAM singleton management. |
|
||||
| **💡 Invariant** | **One `set_image` per frame**: SAM3 runs its heavy vision backbone once per image and re-runs only the lightweight grounding head across multiple prompts, ensuring maximum inference throughput. |
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>4. Review</b> — the frame with its shapes on top</summary>
|
||||
<summary><b>2.2 SAM3 Interactive Sandbox</b> — Standalone prompt engineering playground</summary>
|
||||
|
||||
<br>
|
||||
|
||||
<div align="center">
|
||||
<img src="screenshots/20_sam3_playground_page.png" width="100%" alt="SAM3 Interactive Playground" />
|
||||
<p><i>Figure 2.4: SAM3 Interactive Playground for zero-shot text prompting and point prompt testing.</i></p>
|
||||
</div>
|
||||
|
||||
| Attribute | Specification |
|
||||
|---|---|
|
||||
| **🎯 Purpose** | Interactive sandbox for testing text prompts, positive/negative point clicks, and mask segmentation before launching large auto-annotation jobs. |
|
||||
| **⚡ Key Capabilities** | Real-time mask rendering, multi-prompt layer toggles, point click prompt refinement, and raw JSON detection inspector. |
|
||||
|
||||
</details>
|
||||
|
||||
---
|
||||
|
||||
### Phase 3: High-Throughput Annotation Studio & Triage
|
||||
|
||||
<details open>
|
||||
<summary><b>3.1 Roboflow-Grade Annotation Canvas & Quick Reclass</b> — Sub-second keyboard navigation</summary>
|
||||
|
||||
<br>
|
||||
|
||||
<div align="center">
|
||||
<img src="screenshots/08_review_annotation_canvas.png" width="100%" alt="Annotation Review Canvas" />
|
||||
<p><i>Figure 3.1: Annotation Review Canvas with bounding box editor, shape provenance badges, and hotkey controls.</i></p>
|
||||
</div>
|
||||
|
||||
<br>
|
||||
|
||||
<div align="center">
|
||||
<img src="screenshots/09_review_filmstrip_quick_reclass.png" width="100%" alt="Review Filmstrip and Quick Reclass Bar" />
|
||||
<p><i>Figure 3.2: Filmstrip thumbnail navigation and floating single-keystroke quick reclassification bar.</i></p>
|
||||
</div>
|
||||
|
||||
| Key | Action | | Key | Action |
|
||||
|---|---|---|---|---|
|
||||
| `A` | approve | | `←` `→` | previous / next frame |
|
||||
| `X` | reject | | `U` | jump to next unreviewed |
|
||||
| `Del` | delete shape | | `1`–`9` | pick class |
|
||||
| `S` + drag | SAM3-assisted shape | | drag | add / move / resize a box |
|
||||
| `A` | Approve frame | | `←` `→` | Previous / next frame |
|
||||
| `X` | Reject frame | | `U` | Jump to next unreviewed |
|
||||
| `Del` | Delete selected shape | | `1`–`9` | Quick switch active class |
|
||||
| `S` + Drag | SAM3-assisted box click | | Drag | Add / move / resize bounding box |
|
||||
|
||||
Approving no longer merges. It marks frames approved; merging happens in Data Prep.
|
||||
| Attribute | Specification |
|
||||
|---|---|
|
||||
| **🎯 Purpose** | Fast, ergonomic manual review and refinement of auto-generated bounding boxes. |
|
||||
| **⚡ Key Capabilities** | Single-key shortcuts, bottom thumbnail filmstrip with status indicators, and shape provenance tracking (`SAM3`, `Manual`, `Base Model`). |
|
||||
| **💡 Invariant** | Approving frames does **not** merge them into a dataset. Approval merely qualifies frames for Data Prep; merging happens under frozen rules. |
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>5. Data Prep</b> — throw out the junk, then set augmentation</summary>
|
||||
<summary><b>3.2 Visual Exemplar-Guided Prompting</b> — Few-shot visual reference matching</summary>
|
||||
|
||||
<br>
|
||||
|
||||
Tune an outlier filter over **score / area / aspect** against exactly the batches you picked,
|
||||
watching the counts move as you drag. A dropped box leaves its image in the dataset; only a
|
||||
frame that loses *every* box is held back — these frames hold ~44 objects each, and excluding
|
||||
the whole image was measured to cost 96% of a batch to remove 10% of its boxes.
|
||||
<div align="center">
|
||||
<img src="screenshots/12_review_exemplar_pool_panel.png" width="100%" alt="Exemplar Pool Panel" />
|
||||
<p><i>Figure 3.3: Exemplar Pool Sidebar Panel for visual reference matching.</i></p>
|
||||
</div>
|
||||
|
||||
Then **Confirm merge** creates the dataset under a frozen copy of those rules.
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>6. Datasets</b> — several per project, each a standalone copy</summary>
|
||||
|
||||
<br>
|
||||
|
||||
`batch7+8 strict rules` and `batch7+8 after I fixed the annotations` are two datasets holding
|
||||
the same frames with different labels. Combining them for a run is *newest wins*, so a frame
|
||||
appearing twice is emitted once rather than teaching the model two contradictory labels.
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>7. Models & Training</b> — run, then read the comparison</summary>
|
||||
|
||||
<br>
|
||||
|
||||
Pick datasets, pick classes, train. Batch size, image size and device default from the
|
||||
hardware actually detected, so a bigger GPU changes the numbers in the form, not the code.
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>8. Live Counting</b> — point a model at a camera and watch it count</summary>
|
||||
|
||||
<br>
|
||||
|
||||
Same tracker, stabiliser and line-cross counter the production script uses. Click the video to
|
||||
place the counting line; it moves live without losing the counts. Every finished track is
|
||||
written to a JSONL with the reason it did or did not count — which is what separates a model
|
||||
miss from a tracker miss from a counter miss.
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>9. Counting Accuracy</b> — scored, per cycle</summary>
|
||||
|
||||
<br>
|
||||
|
||||
One row per recording, grouped into collapsible cycles. Type in the ground truth you counted
|
||||
by hand and the table shows the **signed delta** — `+3` and `-3` are different failures, and a
|
||||
single accuracy percentage hides which one you have. Totals only ever count rows where a
|
||||
ground truth is filled in.
|
||||
|
||||
Recount runs headless in a background job — no annotated frame, no JPEG encode, which is worth
|
||||
**146 fps vs 124** in the live view.
|
||||
| Attribute | Specification |
|
||||
|---|---|
|
||||
| **🎯 Purpose** | Store and utilize positive and negative visual crop exemplars to guide SAM3 zero-shot grounding on difficult textures or ambiguous sack designs. |
|
||||
| **⚡ Key Capabilities** | Visual exemplar library, similarity threshold slider, 1-click exemplar addition from canvas bounding boxes. |
|
||||
|
||||
</details>
|
||||
|
||||
---
|
||||
|
||||
## 🎥 Counting
|
||||
### Phase 4: Data Prep, Statistical Triage & Dataset Freezing
|
||||
|
||||
The counter is deliberately robust to low frame rates: it never needs to catch the exact frame
|
||||
of a crossing, only that a track was seen *above* the line at some point in its life.
|
||||
<details open>
|
||||
<summary><b>4.1 Statistical Triage & Quality Outlier Filtering</b> — Filter bad boxes without losing full images</summary>
|
||||
|
||||
<br>
|
||||
|
||||
<div align="center">
|
||||
<img src="screenshots/13_data_prep_quality_outliers.png" width="100%" alt="Data Prep Quality Outliers Panel" />
|
||||
<p><i>Figure 4.1: Quality filter sliders with live box and frame retention counters.</i></p>
|
||||
</div>
|
||||
|
||||
<br>
|
||||
|
||||
<div align="center">
|
||||
<img src="screenshots/11_review_triage_scatter_plot.png" width="100%" alt="Triage Scatter Plot" />
|
||||
<p><i>Figure 4.2: Interactive Scatter Plot (Score vs Area) for instant visual outlier cluster detection.</i></p>
|
||||
</div>
|
||||
|
||||
<br>
|
||||
|
||||
<div align="center">
|
||||
<img src="screenshots/10_review_triage_crop_grid.png" width="100%" alt="Triage Crop Grid" />
|
||||
<p><i>Figure 4.3: Triage Crop Grid for rapid bulk visual inspection and 1-click outlier removal.</i></p>
|
||||
</div>
|
||||
|
||||
| Attribute | Specification |
|
||||
|---|---|
|
||||
| **🎯 Purpose** | Eliminate low-quality bounding boxes (partial crops, false positives, background noise) before dataset compilation. |
|
||||
| **⚡ Key Capabilities** | Dual-handle range sliders (Confidence Score, Box Area, Aspect Ratio), interactive SVG scatter plot, high-density crop grid cards, and real-time retention telemetry. |
|
||||
| **💡 Invariant** | Dropping an outlier bounding box **leaves the frame in the dataset** unless all boxes are dropped. Industrial conveyor frames contain ~44 objects; dropping whole frames discards 96% of good data to remove 10% of bad boxes. |
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>4.2 Augmentation Pipeline & Dataset Freezing</b> — Deterministic Stable Val Split</summary>
|
||||
|
||||
<br>
|
||||
|
||||
<div align="center">
|
||||
<img src="screenshots/14_data_prep_augmentation_panel.png" width="100%" alt="Augmentation Configuration Panel" />
|
||||
<p><i>Figure 4.4: Albumentations training augmentation presets with real-time visual preview.</i></p>
|
||||
</div>
|
||||
|
||||
<br>
|
||||
|
||||
<div align="center">
|
||||
<img src="screenshots/15_data_prep_merge_target_modal.png" width="100%" alt="Merge Target and Dataset Freeze Modal" />
|
||||
<p><i>Figure 4.5: Dataset Freeze dialog enforcing deterministic SHA-1 Stable Val Split rules.</i></p>
|
||||
</div>
|
||||
|
||||
| Attribute | Specification |
|
||||
|---|---|
|
||||
| **🎯 Purpose** | Apply synthetic vision augmentations and freeze the curated batch into an immutable, versioned training dataset. |
|
||||
| **⚡ Key Capabilities** | HSV shift, brightness/contrast, rotation, blur, and mosaic controls; dataset destination selection; split ratio slider. |
|
||||
| **💡 Invariant** | **Stable Validation Split**: Validation assignment is derived deterministically from the frame's SHA-1 hash. Once a frame lands in `val`, it remains in `val` forever across all future versions. |
|
||||
|
||||
</details>
|
||||
|
||||
---
|
||||
|
||||
### Phase 5: YOLO Retraining & Live Training Progress
|
||||
|
||||
<details open>
|
||||
<summary><b>5.1 Dataset Repository & YOLO Fine-Tuning</b> — Hardware-aware parameter tuning</summary>
|
||||
|
||||
<br>
|
||||
|
||||
<div align="center">
|
||||
<img src="screenshots/16_datasets_page.png" width="100%" alt="Master Datasets List" />
|
||||
<p><i>Figure 5.1: Master Dataset repository showing version history, train/val splits, and export tools.</i></p>
|
||||
</div>
|
||||
|
||||
<br>
|
||||
|
||||
<div align="center">
|
||||
<img src="screenshots/17_models_training_page.png" width="100%" alt="YOLO Models and Training Page" />
|
||||
<p><i>Figure 5.2: YOLO fine-tuning parameter setup with hardware detection and model history.</i></p>
|
||||
</div>
|
||||
|
||||
<br>
|
||||
|
||||
<div align="center">
|
||||
<img src="screenshots/21_workflow_progress_states.png" width="100%" alt="Live Training Workflow Progress" />
|
||||
<p><i>Figure 5.3: Real-time training telemetry with live loss curves, mAP metrics, and stdout logs.</i></p>
|
||||
</div>
|
||||
|
||||
| Attribute | Specification |
|
||||
|---|---|
|
||||
| **🎯 Purpose** | Train Ultralytics YOLO11 object detection models with automated hardware tuning and live telemetry. |
|
||||
| **⚡ Key Capabilities** | Architecture selection (`YOLO11n/s/m/l/x`), auto-calculated batch sizes based on free GPU VRAM, real-time loss sparklines (`box_loss`, `cls_loss`, `dfl_loss`), epoch progress bars, and streaming logs. |
|
||||
| **💡 Invariant** | **Side-by-side Validation**: After training, both the base model and fine-tuned model are benchmarked on the identical validation set, displaying signed metric deltas ($\Delta\text{mAP50}$, $\Delta\text{precision}$, $\Delta\text{recall}$). |
|
||||
|
||||
</details>
|
||||
|
||||
---
|
||||
|
||||
### Phase 6: Counting Accuracy Benchmark & Live Production Inference
|
||||
|
||||
<details open>
|
||||
<summary><b>6.1 Counting Benchmark Matrix & Live Line Counter</b> — 146 FPS verification</summary>
|
||||
|
||||
<br>
|
||||
|
||||
<div align="center">
|
||||
<img src="screenshots/18_counting_bench_page.png" width="100%" alt="Counting Accuracy Benchmark Matrix" />
|
||||
<p><i>Figure 6.1: Counting Accuracy Benchmark Matrix comparing multi-model predictions against Ground Truth.</i></p>
|
||||
</div>
|
||||
|
||||
<br>
|
||||
|
||||
<div align="center">
|
||||
<img src="screenshots/19_live_count_page.png" width="100%" alt="Live Video Inference and Line Counter" />
|
||||
<p><i>Figure 6.2: Real-time CCTV live counting interface with interactive tripwire line and ByteTrack trails.</i></p>
|
||||
</div>
|
||||
|
||||
| Attribute | Specification |
|
||||
|---|---|
|
||||
| **🎯 Purpose** | Verify model counting precision against physical ground truth records and run real-time production counting. |
|
||||
| **⚡ Key Capabilities** | Headless counting benchmark running at **146 FPS**, signed delta badges (`+2`, `-1`, `0`), draggable tripwire counting line, ByteTrack trajectory tracking, and per-track JSONL diagnostic logs. |
|
||||
| **💡 Invariant** | **Counting tripwire on upper edge ($y_1$)**: Tripwire evaluates the top edge coordinate of the sack bounding box rather than the center or bottom, preventing miscounts caused by physical sack deformation as it drops onto the conveyor. |
|
||||
|
||||
</details>
|
||||
|
||||
---
|
||||
|
||||
<a id="counting-engine"></a>
|
||||
## 🎥 Production Counting Engine
|
||||
|
||||
The counting engine is designed for industrial conveyor belts with low or unstable camera frame rates. It eliminates reliance on instantaneous line-crossing frames by evaluating historical bounding box trajectories.
|
||||
|
||||
```mermaid
|
||||
stateDiagram-v2
|
||||
[*] --> UNKNOWN
|
||||
UNKNOWN --> ABOVE: y1 above the band
|
||||
UNKNOWN --> BELOW: born below (ghost — never counts)
|
||||
ABOVE --> COUNTED: seen below + travelled far enough
|
||||
COUNTED --> ABOVE: sustained frames above (real unload)
|
||||
UNKNOWN --> ABOVE: y1 coordinate above counting band
|
||||
UNKNOWN --> BELOW: Born below band (ghost detection — rejected)
|
||||
ABOVE --> COUNTED: Trajectory crosses band + travelled entry_travel_min
|
||||
COUNTED --> ABOVE: Sustained frames above line (genuine reload)
|
||||
```
|
||||
|
||||
<details>
|
||||
<summary><b>Four layers, and what each one is actually for</b></summary>
|
||||
### Multi-Layer Counter Safeguards
|
||||
|
||||
<br>
|
||||
|
||||
| Layer | Guard | Stops |
|
||||
| Guard Layer | Rule Specification | Failure Mode Prevented |
|
||||
|---|---|---|
|
||||
| 1 | Must have been **above** the line at some point | a box that appears inside the truck |
|
||||
| 2 | Must have travelled `entry_travel_min` from where it first appeared | ghost boxes that blink into existence next to the line |
|
||||
| 3 | **Track hand-off** — a dying track parks its history for a newborn nearby to inherit | an ID switch at the line losing the count *or* duplicating it |
|
||||
| 4 | One count per direction per track, and the verdict is the track's *last* direction | double counting, while still letting a genuine unload-and-reload count again |
|
||||
|
||||
Every one of these was written against a reproduced failure. Hand-off replaced a spatial dedup
|
||||
that did not dedup: a blocked track simply retried each frame and counted anyway once it
|
||||
drifted out of the circle — late, at the wrong position, seeding the next circle in the wrong
|
||||
place.
|
||||
|
||||
`handoff_radius` is the dial that matters most. These frames hold ~44 objects, so a newborn
|
||||
track is nearly always near one that just vanished; calibrate it against a clip with a
|
||||
hand-counted total rather than by eye.
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>Recording pipeline</b> — record once, cut sessions afterwards</summary>
|
||||
|
||||
<br>
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
CAM[📷 Dahua 1080p<br/>H.265 @ 25fps] --> MTX[MediaMTX on Jetson<br/>records 24/7 · 24h buffer]
|
||||
MTX -->|RTSP| DET[Truck detector<br/>on the GPU box]
|
||||
DET -->|session ends| FETCH[Download that exact<br/>time range as a copy]
|
||||
MTX --> FETCH
|
||||
FETCH --> ARC[📁 archive/date/batchNNN.mp4<br/>+ .json sidecar]
|
||||
```
|
||||
|
||||
The detector does **not** encode video. When a truck session ends it downloads that range from
|
||||
the recording server, so the archive keeps the camera's own codec, resolution and frame rate.
|
||||
|
||||
| | Before | After |
|
||||
|---|---|---|
|
||||
| Codec | mpeg4 re-encode | HEVC copy |
|
||||
| Resolution | 1280×720 | 1920×1080 |
|
||||
| Size | 8.0 Mbps | **1.72 Mbps** (4.7× smaller) |
|
||||
| Timebase | declared 10 fps at 25 fps real → **2.49× slow** | true 25 fps |
|
||||
| Start time | read from the burned-in overlay by OCR | from the server, exact |
|
||||
|
||||
Each clip carries a `.json` sidecar with the server's start time, which the app trusts over
|
||||
reading the overlay — so a new session appears in the right cycle with no scan at all.
|
||||
|
||||
</details>
|
||||
| **Layer 1: Entry Origin Guard** | Object must be observed **above** the counting band prior to crossing. | Prevents counting items that spawn directly inside the truck or loading chute. |
|
||||
| **Layer 2: Trajectory Distance** | Object must travel a minimum distance (`entry_travel_min`) across consecutive frames. | Discards transient noise and flickering phantom boxes. |
|
||||
| **Layer 3: Track Hand-Off** | When a track is occluded, its movement history is parked for newborn tracks within `handoff_radius`. | Prevents tracker ID switches from dropping or duplicating counts. |
|
||||
| **Layer 4: Directional Monotonicity** | Strict single count per trajectory direction with verdict locked to the final confirmed motion. | Prevents double-counting during momentary conveyor pauses. |
|
||||
|
||||
---
|
||||
|
||||
## 📊 Why the numbers are trustworthy
|
||||
<a id="configuration"></a>
|
||||
## ⚙️ Configuration & Environment Matrix
|
||||
|
||||
After training, the base model and the new one are validated **on the same val set**, and
|
||||
mAP50 / mAP50-95 / precision / recall appear side by side with the difference.
|
||||
All configuration parameters are defined via environment variables in `.env`:
|
||||
|
||||
> [!TIP]
|
||||
> **The val split is stable.** Once a frame is in `val` it stays there for every later merge —
|
||||
> derived from the frame's identity, not from how many rows precede it. A rising score cannot
|
||||
> be an easier val set.
|
||||
|
||||
- **Training uses whole datasets, old batches included.** Fine-tuning on the newest batch alone
|
||||
tends to raise the score on new footage while quietly losing the old.
|
||||
- **Datasets are snapshots.** Labels are copied from what the dataset holds on disk, not
|
||||
re-derived from today's rules, so two runs over the same dataset cannot disagree.
|
||||
- **An empty Base column is honest.** It means the previous model's classes did not match this
|
||||
dataset's, so scoring it would have compared two different things. The message says which.
|
||||
|
||||
`Use as base model` promotes a version, and the next round fine-tunes from it.
|
||||
| Variable | Type | Default Value | Scope | Description |
|
||||
|---|---|---|---|---|
|
||||
| `HF_TOKEN` | `String` | *(Required)* | Backend / Docker | HuggingFace user access token with authorized permissions to download gated `facebook/sam3` weights. |
|
||||
| `VIDEO_ARCHIVE_HOST` | `Path` | `./data/archive` | Docker Compose | Host filesystem directory path containing raw CCTV video recordings (structured as `<date>/<batch>.mp4`). Mounted read-only (`:ro`). |
|
||||
| `APP_DATA_DIR` | `Path` | `./data` (or `/data` in Docker) | Backend | Base directory for application persistent data, SQLite database (`app.db`), projects, extracted frames, datasets, and model weights. |
|
||||
| `VIDEO_ARCHIVE` | `Path` | `/videos` (or `data/archive`) | Backend | Internal container/local filesystem path where the video archive is browsed by FastAPI. |
|
||||
| `WEB_PORT` | `Integer` | `8080` | Docker / Nginx | Host HTTP port mapped to the Nginx frontend web UI. |
|
||||
| `CORS_ORIGINS` | `String` | `http://localhost:5173,http://localhost:8080` | FastAPI Backend | Comma-separated list of allowed origins for Cross-Origin Resource Sharing. |
|
||||
| `API_URL` | `URL` | `http://localhost:8000` | Vite Dev Server | Backend target endpoint for Vite development proxy (`frontend/vite.config.js`). |
|
||||
| `MEDIAMTX_WHEP_PATH` | `String` | `/whep` | Backend Live Count | WHEP WebRTC endpoint path on the streaming media server (MediaMTX). |
|
||||
| `MEDIAMTX_RTSP_PORT` | `Integer` | `8554` | Backend Live Count | RTSP stream port used to translate WHEP browser streams into backend video processing feeds. |
|
||||
| `RTSP_TRANSPORT` | `String` | `tcp` | Backend Live Count | RTSP transport protocol (`tcp` or `udp`). TCP guarantees zero frame drop on industrial networks. |
|
||||
| `PLAYBACK_URL` | `URL` | `http://192.168.192.96:9996/get` | Recorder Service | MediaMTX recording playback API endpoint for automated CCTV session extraction. |
|
||||
| `PLAYBACK_PATH` | `String` | `cam` | Recorder Service | Stream channel identifier on the MediaMTX playback server. |
|
||||
| `AUTO_PULL_INTERVAL` | `Integer` | `30` | Auto-Pull Script | Polling frequency in seconds for automated Git repository synchronization (`scripts/auto_pull.py`). |
|
||||
| `WEBHOOK_PORT` | `Integer` | `9000` | Webhook Daemon | Port for the GitHub push webhook listener daemon (`scripts/webhook.py`). |
|
||||
| `WEBHOOK_SECRET` | `String` | `""` | Webhook Daemon | Shared secret key for validating GitHub webhook HMAC-SHA256 signatures. |
|
||||
|
||||
---
|
||||
|
||||
## 🗂️ Where things live
|
||||
<a id="storage-layout"></a>
|
||||
## 🗂️ Persistent Data & Storage Layout
|
||||
|
||||
All application state, relational metadata, and trained weights live in `data/`:
|
||||
|
||||
```
|
||||
data/
|
||||
app.db # 14 tables: projects, frames, annotations,
|
||||
# datasets, jobs, count_runs, video_clock …
|
||||
archive/<date>/batchNNN.mp4 # recordings (read-only to the app)
|
||||
batchNNN.json # sidecar: true start time from the server
|
||||
recorder.log # the 24/7 recorder's output
|
||||
live-count/session-*.jsonl # per-track traces from the live counter
|
||||
projects/<slug>/
|
||||
base/model.pt # the base model
|
||||
datasets/<id>/ # one folder per named dataset
|
||||
images/{train,val}/ labels/{train,val}/
|
||||
batches/<id>/frames/ # extracted frames
|
||||
models/<n>/best.pt + metrics.json # each training run
|
||||
├── app.db # SQLite metadata database with Write-Ahead Logging (WAL)
|
||||
├── archive/<date>/batchNNN.mp4 # Raw CCTV video recordings (Mounted strictly read-only)
|
||||
│ batchNNN.json # Sidecar metadata with true server timestamp
|
||||
├── recorder.log # 24/7 background recorder daemon log
|
||||
├── live-count/session-*.jsonl # Diagnostic per-track trajectory and crossing logs
|
||||
└── projects/<slug>/ # Isolated project workspace
|
||||
├── base/model.pt # Project base model weights and locked class taxonomy
|
||||
├── batches/<id>/frames/ # Losslessly extracted image frames from video slices
|
||||
├── datasets/<id>/ # Frozen Ultralytics YOLO formatted training datasets
|
||||
│ ├── images/{train,val}/ # Immutable frame images
|
||||
│ └── labels/{train,val}/ # YOLO format bounding box annotations (.txt)
|
||||
└── models/<n>/ # Training runs (weights/best.pt, metrics.json, args.yaml)
|
||||
```
|
||||
|
||||
The database holds status; the disk holds pixels, labels and weights. A dataset trains as-is
|
||||
with Ultralytics, or imports into Roboflow, without this application.
|
||||
---
|
||||
|
||||
> [!WARNING]
|
||||
> Your video archive is mounted **read-only** and nothing is ever written back into it.
|
||||
<a id="background-jobs"></a>
|
||||
## ⚙️ Background Job Worker & Mutex Locking
|
||||
|
||||
Heavy computational operations run through an asynchronous background worker (`backend/jobs.py`) with strict GPU mutex locking to prevent VRAM over-allocation:
|
||||
|
||||
| Job Type | GPU Locked | Description |
|
||||
|---|:---:|---|
|
||||
| `extract` | — | Background FFmpeg extraction of video ranges into frame sequences. |
|
||||
| `autolabel` | ✅ | Meta SAM3 zero-shot grounding across candidate frames (one `set_image` per image). |
|
||||
| `merge` | — | Compiles reviewed batches into immutable dataset splits under frozen triage rules. |
|
||||
| `train` | ✅ | Ultralytics YOLO11 fine-tuning followed by automated dual-model validation. |
|
||||
| `count` | ✅ | Headless evaluation benchmark running archive videos at **146 FPS**. |
|
||||
| `clock-scan` | — | OCR extraction of camera burned-in timestamps. |
|
||||
| `truck-scan` | ✅ | Batch inference check verifying the presence of target industrial objects. |
|
||||
|
||||
---
|
||||
|
||||
## ⚙️ Jobs
|
||||
<a id="invariants"></a>
|
||||
## 📊 Metrics Integrity & Domain Invariants
|
||||
|
||||
Heavy work runs as queued jobs, one at a time, so the GPU is never double-booked.
|
||||
|
||||
| Type | GPU | What it does |
|
||||
|---|---|---|
|
||||
| `extract` | — | ffmpeg pulls frames out of a range |
|
||||
| `autolabel` | ✅ | SAM3 over a batch, one `set_image` per image |
|
||||
| `merge` | — | copies approved frames into a dataset under frozen rules |
|
||||
| `train` | ✅ | fine-tunes, then validates base and new on the same val set |
|
||||
| `count` | ✅ | headless recount of archive videos for the accuracy table |
|
||||
| `clock-scan` | — | reads each recording's real start time |
|
||||
| `truck-scan` | ✅ | checks every recording actually contains a truck |
|
||||
|
||||
Jobs and their progress are persistent; after a restart the list is still there.
|
||||
1. **Stable Validation Split**: Validation assignment is determined deterministically by `sha1(image_bytes) % 100 < val_ratio`. Once a frame lands in the validation split, it remains in validation across all future dataset iterations. This prevents validation leak and ensures that rising mAP scores reflect genuine model improvements.
|
||||
2. **One `set_image` per Frame**: SAM3 executes its heavy vision transformer backbone once per image. Multi-class text prompting evaluates the lightweight grounding head against cached backbone embeddings.
|
||||
3. **Outlier Filtering Preserves Frames**: Dropping a bounding box during Data Prep removes only the bad annotation. The image frame remains in the dataset as long as at least one valid box persists.
|
||||
4. **Upper-Edge Coordinate Line Crossing ($y_1$)**: Tripwires evaluate the top edge coordinate of bounding boxes ($y_1$) rather than the centroid ($y_c$) or bottom edge ($y_2$), ensuring immunity to sack deformation upon conveyor impact.
|
||||
|
||||
---
|
||||
|
||||
## 🔧 When something goes wrong
|
||||
<a id="documentation"></a>
|
||||
## 📚 Documentation & Deep Dives
|
||||
|
||||
Comprehensive technical specifications, operational SOPs, and architecture diagrams are available in the [`docs/`](docs/) directory:
|
||||
|
||||
| Document | Format | Description |
|
||||
|---|:---:|---|
|
||||
| [**Panduan Sistem Lengkap**](docs/PANDUAN_SISTEM_LENGKAP.pdf) | `PDF` (11 MB) | **Publication-Grade Master User Guide & Technical Manual** (Indonesian) with complete operational SOPs and embedded figures. |
|
||||
| [**Panduan Sistem Lengkap Source**](docs/PANDUAN_SISTEM_LENGKAP.fodt) | `FODT` (11.8 MB) | Native LibreOffice Writer Flat XML editable source document. |
|
||||
| [**Panduan Sistem Lengkap Markdown**](docs/PANDUAN_SISTEM_LENGKAP.md) | `MD` (68 KB) | Full Markdown transcript of the 14-chapter system manual. |
|
||||
| [**Architecture Flow Diagram (4K)**](docs/diagram-alur.png) | `PNG` (1.5 MB) | 4K Ultra-HD raster export of the 8-stage end-to-end retraining pipeline. |
|
||||
| [**Architecture Flow Diagram (Vector)**](docs/diagram-alur.svg) | `SVG` (34.6 KB) | Scalable vector graphic diagram for high-resolution display. |
|
||||
| [**Architecture Flow Diagram (Source)**](docs/diagram-alur.fodg) | `FODG` (24.8 KB) | Native LibreOffice Draw Flat XML editable source file. |
|
||||
| [**System Requirements Specification**](docs/requirements.md) | `MD` (23.6 KB) | Numbered technical requirements (`REQ-001` through `REQ-042`). |
|
||||
| [**System Design & Architecture**](docs/design.md) | `MD` (22.7 KB) | Database schema, REST API contracts, disk layouts, and backend invariants. |
|
||||
| [**UI/UX Design Specification**](docs/ui-spec.md) | `MD` (79.7 KB) | Dark theme design tokens, hotkey maps, and component specifications. |
|
||||
| [**Ground Truth Benchmark Dataset**](docs/GT.xlsx) | `XLSX` (586 KB) | Hand-verified physical conveyor bag counts across operational shifts. |
|
||||
|
||||
---
|
||||
|
||||
<a id="troubleshooting"></a>
|
||||
## 🔧 Troubleshooting & FAQ
|
||||
|
||||
<details>
|
||||
<summary><b>Deployment and GPU</b></summary>
|
||||
<summary><b>Deployment & GPU Acceleration</b></summary>
|
||||
|
||||
<br>
|
||||
|
||||
| Symptom | Cause / fix |
|
||||
| Symptom | Root Cause & Remediation |
|
||||
|---|---|
|
||||
| API returns 404 for routes you just added | the image copies `backend/` at build time — `docker compose build backend` again |
|
||||
| A UI change doesn't show up | same trap on the other side: `docker compose build frontend`, then hard-reload |
|
||||
| `could not select device driver` | the NVIDIA container toolkit is not installed, or Docker is older than the CDI support compose relies on (`devices: nvidia.com/gpu=all`) |
|
||||
| `CUDA out of memory` while training | lower epochs/batch on the Models page, or free the card — SAM3 is released before training, but another process may still hold it |
|
||||
| A job reads *interrupted by a server restart* | it was running when the process died — jobs are not resumable, start it again |
|
||||
| `could not select device driver` | NVIDIA Container Toolkit is missing or Docker Engine is older than CDI specifications. Run `./install_nvidia.sh` or update Docker. |
|
||||
| `CUDA out of memory` during training | Lower the batch size on the Models page or stop background jobs. SAM3 releases VRAM before training starts, but external processes may hold memory. |
|
||||
| Changes to frontend or backend do not appear | Docker Compose caches container layers at build time. Run `docker compose build backend frontend` and hard-refresh your browser (`Ctrl+Shift+R`). |
|
||||
| Job shows *interrupted by server restart* | The backend process stopped while a job was active. Jobs do not resume mid-epoch; simply re-trigger the job from the UI. |
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>SAM3 and auto-labelling</b></summary>
|
||||
<summary><b>SAM3 Foundation Auto-Annotation</b></summary>
|
||||
|
||||
<br>
|
||||
|
||||
| Symptom | Cause / fix |
|
||||
| Symptom | Root Cause & Remediation |
|
||||
|---|---|
|
||||
| Job fails at *loading model* with a 401 | access to `facebook/sam3` not granted yet, or `HF_TOKEN` missing |
|
||||
| SAM3 download crawls at a few KB/s | HuggingFace's Xet transfer throttling itself; `HF_HUB_DISABLE_XET=1` is already set in compose for that reason |
|
||||
| Job fails at *Loading Model* with HTTP 401 | HuggingFace token is invalid or access to [facebook/sam3](https://huggingface.co/facebook/sam3) has not yet been approved. |
|
||||
| SAM3 checkpoint download is slow | HuggingFace Xet transfer throttling. `HF_HUB_DISABLE_XET=1` is enabled by default in `docker-compose.yml` to bypass this issue. |
|
||||
| SAM3 generates inaccurate bounding boxes | Refine the natural language prompt with physical descriptors (e.g. `"woven polypropylene sack with blue logo"`) or add visual crops to the Exemplar Pool. |
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>Archive and counting</b></summary>
|
||||
<summary><b>Video Archive & Live Counting</b></summary>
|
||||
|
||||
<br>
|
||||
|
||||
| Symptom | Cause / fix |
|
||||
| Symptom | Root Cause & Remediation |
|
||||
|---|---|
|
||||
| Videos listed as *unreadable* | ffprobe could not parse them; they are still listed rather than hidden, so the archive never looks emptier than it is |
|
||||
| A recording's time shows amber | its overlay was read with low confidence, or not at all — type the time you can see in the video; a hand-entered time is never overwritten by a rescan |
|
||||
| A cycle looks short | check whether its recordings moved to the neighbouring cycle — anything before 06:00 belongs to the previous shift |
|
||||
| Live counting says *GPU busy* | a training or auto-label job holds the card; it waits 30 s before giving up |
|
||||
| Video file marked as *unreadable* | FFmpeg/FFprobe could not parse the video header. Run `uv run python scripts/transcode_archive.py` to re-mux into standard H.264 MP4. |
|
||||
| Shift cycle shows fewer recordings than expected | Check if recordings crossed the 06:00 boundary. Clips recorded before 06:00 belong to the previous operational shift cycle. |
|
||||
| Live Counting shows *GPU busy* | An active fine-tuning or auto-annotation job holds the GPU lock. The live counter waits 30 seconds before falling back to CPU or queuing. |
|
||||
|
||||
</details>
|
||||
|
||||
---
|
||||
|
||||
## 🤝 Contributing
|
||||
<a id="contributing"></a>
|
||||
## 🤝 Contributing & Engineering Guidelines
|
||||
|
||||
The documents drive the repo, not the other way round:
|
||||
Development adheres strictly to the **Chain of Truth** methodology:
|
||||
- **Specifications First**: Every feature must map directly to a numbered requirement in [`docs/requirements.md`](docs/requirements.md) and architectural design in [`docs/design.md`](docs/design.md).
|
||||
- **Verified Deliverables**: Tasks tracked in [`docs/tasks.md`](docs/tasks.md) flip to `[DONE]` only after concrete end-to-end verification.
|
||||
- **Surgical Changes**: Touch only code directly relevant to the feature. Adhere to the working rules in [`AGENTS.md`](AGENTS.md).
|
||||
- **Package Manager**: All backend dependencies are managed exclusively with Astral `uv` (`requirements.txt`).
|
||||
|
||||
| Document | Contents |
|
||||
|---|---|
|
||||
| [`docs/requirements.md`](docs/requirements.md) | numbered `REQ-xxx`, changed only with the owner's approval |
|
||||
| [`docs/design.md`](docs/design.md) | schema, API contract, disk layout — each section names the `REQ-xxx` it serves |
|
||||
| [`docs/tasks.md`](docs/tasks.md) | implementation steps and how each was *verified*, `[TODO]` / `[DONE]` |
|
||||
| [`AGENTS.md`](AGENTS.md) | working rules: simplicity, surgical changes, verify by running something |
|
||||
---
|
||||
|
||||
Two invariants are easy to break and make the whole system lie:
|
||||
<a id="license"></a>
|
||||
## 📄 License
|
||||
|
||||
1. **A frame in `val` stays in `val`** — otherwise the base-vs-new comparison is meaningless.
|
||||
2. **One `set_image` per image** — `set_text_prompt()` re-runs only the grounding head against
|
||||
the cached backbone output. An N-prompt job calls `set_image` once and loops prompts over
|
||||
that state.
|
||||
Distributed under the MIT License. See [`LICENSE`](LICENSE) for details.
|
||||
Reference in new issue
Block a user