- AutoAnnotateModal: render merges last full-set detections for non-exemplared classes with conditioned results for exemplared ones (selected-not-exemplared filter); drawing an example no longer clears the other classes' detections - Run Preview now always refreshes full-set first (exemplars: []) then re-runs exemplared classes - docs: REQ-182 amended (merge semantics), design/ui-spec/tasks synced
reTraining
Take a base model you already have, and make it measurably better with footage you already have.
The Open-Source, Self-Hosted Vision Pipeline to Turn Raw Industrial CCTV into Production-Grade YOLO Object Detectors & Line Counters.
⚡ Quickstart · 🏗️ Architecture · 🖼️ Visual UI Tour · 🎥 Counting Engine · ⚙️ Configuration · 📚 Documentation · 🔧 Troubleshooting
Click diagram to view 4K UHD resolution. Scalable vector version available at docs/diagram-alur.svg.
💡 Why reTraining?
Deploying object detection models in industrial environments (manufacturing, logistics, agricultural feedmills, conveyor belts) often hits a painful bottleneck: general base models fail on domain-specific edge cases, while building manual labeling pipelines from scratch is slow and expensive.
reTraining provides a self-contained, enterprise-grade active learning platform designed to run directly on your edge server or GPU workstation:
- Ingest Raw CCTV Footage: Stream directly from continuous 24/7 video archives without re-encoding or modifying the underlying storage.
- Zero-Shot Foundation Auto-Labeling: Leverage Meta's Segment Anything Model 3 (SAM3) with natural language text prompts and visual exemplars to annotate thousands of frames in minutes.
- Roboflow-Grade Review Studio: Sub-second keyboard navigation, instant class switching, click-assist segmentation, and high-density visual triage crop grids.
- Statistical Triage & Data Prep: Filter bounding box outliers by area, aspect ratio, and confidence score without discarding valid image frames.
- Immutable Dataset Freezing & Stable Val Splits: Guarantee reproducible benchmarks with deterministic SHA-1 validation sets that never shift across retraining runs.
- Hardware-Aware Continuous Fine-Tuning: Auto-detect host GPU VRAM, fine-tune Ultralytics YOLO11 models, and evaluate base vs. fine-tuned model performance side-by-side.
- Live Production Counting & Benchmarks: Real-time RTSP/WHEP live inference with ByteTrack line-crossing counters, evaluated against verified ground truth physical counts at 146 FPS.
⚡ Quickstart
Prerequisites
- Host OS: Linux (Ubuntu 22.04+ recommended) or Windows with WSL2.
- GPU Acceleration: NVIDIA GPU with CUDA 12.4+ and NVIDIA Container Toolkit (Docker Engine 27+ CDI support).
- HuggingFace Account: Gated model access granted for facebook/sam3 with a valid user access token (
HF_TOKEN). - Video Storage: Directory of CCTV footage structured as
<date>/<batch>.mp4(or let the app mount./data/archive).
Option 1: Docker Compose (Production Standard)
Launch the entire stack with a single command. The startup script automatically inspects the host hardware, generates Container Device Interface (CDI) specs for your NVIDIA GPU, and starts both backend and frontend containers.
# 1. Clone the repository
git clone git@github.com:fhanyuh/reTraining.git
cd reTraining
# 2. Configure environment credentials
cp .env.example .env
# Edit .env and paste your HuggingFace user access token:
# HF_TOKEN=hf_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
# 3. Launch with automated GPU / CDI detection
chmod +x start.sh
./start.sh
Alternative direct Docker Compose launch:
docker compose up -d --build
Access Points:
- 🌐 Web Studio UI: http://localhost:9000
- 📑 Interactive REST API Docs: http://localhost:9010/docs
- 🩺 Health & GPU Telemetry Endpoint: http://localhost:9010/api/health
Verify system health and GPU VRAM availability:
curl http://localhost:9010/api/health
{
"device": "cuda",
"gpu": "NVIDIA GeForce RTX 5080 Laptop GPU",
"vram_free_gb": 14.91,
"sam3_ready": true,
"ffmpeg": true,
"hf_token": true,
"db": true
}
Important
The initial SAM3 auto-annotation job automatically downloads the ~3.4 GB SAM3 checkpoint from HuggingFace into a persistent Docker named volume (
hf-cache). Subsequent executions load the model into VRAM in ~12 seconds.
Option 2: Local Development (Bare-Metal / uv + Vite)
For core development and live debugging without Docker:
Requirements: Python 3.12+, Astral uv, Node.js 20+, FFmpeg.
# 1. Clone and setup environment
git clone git@github.com:fhanyuh/reTraining.git
cd reTraining
cp .env.example .env
# 2. Install backend dependencies and vendored SAM3 with uv
uv pip install -r requirements.txt
uv pip install -e ./sam3
# 3. Start FastAPI backend (Port 8000)
uv run uvicorn backend.main:app --host 0.0.0.0 --port 8000 --reload
# 4. In a separate terminal, install and start Vite frontend (Port 5173)
cd frontend
npm install
npm run dev -- --host 0.0.0.0 --port 5173
Note
The frontend dependencies are pinned to Vite 7 (
vite: ^7.1.5) infrontend/package.jsonto prevent Rolldown native binding bus errors on Linux platforms.
🏗️ End-to-End Pipeline Architecture
The platform operates as a continuous closed-loop retraining pipeline divided into 7 distinct functional stages:
┌──────────────────────────────────────────────────────────────────────────────────────────────────┐
│ DATA RETRAINING & INFERENCE PIPELINE │
└──────────────────────────────────────────────────────────────────────────────────────────────────┘
[1. Video Archive] ──▶ [2. Frame Slicing] ──▶ [3. SAM3 Auto-Label] ──▶ [4. Review & Exemplar]
(CCTV/MP4) (FFmpeg In/Out) (Prompt Grounding) (Click-Assist Canvas)
│
[7. Live Counter] ◀── [6. YOLO11 Train] ◀── [5. Immutable Split] ◀── [4. Triage & Filter]
(WHEP / RTSP) (Auto VRAM) (Stable Train/Val) (Outlier Purge)
- Video Archive & Shift Slicing: Raw CCTV recordings are indexed into 24-hour operational work shifts (06:00 to 05:59 next morning). Sub-second timeline in/out trimming extracts high-resolution frame sequences with configurable FPS rates via asynchronous FFmpeg queues.
- SAM3 Foundation Auto-Labeling: Meta SAM3 zero-shot open-vocabulary grounding generates candidate bounding boxes and segmentation masks from natural language descriptions (e.g.
"white sack of feed on conveyor"). - Interactive Review Studio: Operators review candidate annotations on a Roboflow-grade canvas with single-keystroke approvals, box adjustments, click-assist segmentation, and visual prompt exemplar refinement.
- Statistical Triage & Quality Outliers: Interactive 2D scatter plots (Box Area vs. Confidence Score) and high-density crop grids allow rapid isolation and pruning of false positives without discarding valid frames.
- Immutable Dataset Compilation: Filtered batches are merged into versioned datasets (
v1,v2,v3) with deterministic SHA-1 validation hashing, guaranteeing that validation images remain permanently locked across iterations. - Hardware-Aware YOLO Retraining: Hyperparameters and batch sizes auto-scale based on detected GPU VRAM. The system fine-tunes Ultralytics YOLO11, benchmarks old vs. new models on the identical validation set, and displays signed metric deltas (
\Delta\text{mAP50},\Delta\text{Precision},\Delta\text{Recall}). - Counting Accuracy Benchmark & Live Inference: Real-time production inference evaluates live RTSP/WHEP video streams with ByteTrack trajectory tracking and upper-edge tripwire counters, benchmarking against hand-verified physical ground truth logs at 146 FPS.
🖼️ Visual UI Tour & Feature Gallery
Explore the 6 core pipeline phases across all 11 primary user screens and modal workflows.
Phase 1: Video Ingest, Shift Cycles & Frame Sampling
1.1 Project Workspace & Setup — Multi-project isolation and class locking
Figure 1.1: Projects Dashboard showing active projects, base models, class taxonomies, and dataset stats.
| Attribute | Specification |
|---|---|
| 🎯 Purpose | Central management hub for all isolated computer vision projects. |
| ⚡ Key Capabilities | View base model architecture, active class tags, total extracted batches, and dataset snapshots at a glance. |
| 💡 Invariant | Projects maintain strictly isolated database records, class lists, and filesystem storage roots under data/projects/<slug>/. |
Figure 1.2: Project Creation dialog with base model checkpoint upload and automatic class extraction.
| Attribute | Specification |
|---|---|
| 🎯 Purpose | Initialize a new project with locked class names and base model weights. |
| ⚡ Key Capabilities | Upload an existing YOLO .pt checkpoint to automatically extract its class taxonomy, or define custom classes and start fine-tuning from yolo11n.pt. |
| 💡 Invariant | Base model classes are permanently locked to the project to prevent label drift between training iterations. |
1.2 24-Hour Operational Shift Archive — Cycle grouping that crosses midnight
Figure 1.3: Video Archive grouping recordings into 24-hour operational shifts (06:00 to 05:59 next morning).
| Attribute | Specification |
|---|---|
| 🎯 Purpose | Browse raw CCTV recordings grouped by operational work shifts rather than arbitrary calendar folders. |
| ⚡ Key Capabilities | Reads camera burned-in OCR timestamps and sidecar .json metadata to assign recordings accurately across midnight boundaries. |
| 💡 Invariant | The video archive is mounted strictly read-only (:ro). Nothing is ever modified, renamed, or deleted in the user's video repository. |
1.3 Video Trim & Frame Extractor — Sub-second timeline scrubbing and sampling
| Attribute | Specification |
|---|---|
| 🎯 Purpose | Select active truck loading ranges and configure extraction frame rates. |
| ⚡ Key Capabilities | Interactive in/out timeline markers, extraction FPS slider (0.5 – 2.0 FPS), real-time output frame counter, and background FFmpeg queueing. |
| 💡 Invariant | Extraction runs asynchronously in the background queue; frames are losslessly sampled into data/projects/<slug>/batches/<id>/frames/. |
Phase 2: SAM3 Zero-Shot Auto-Labeling Engine
2.1 Batch Management & Mass Auto-Annotation — High-throughput zero-shot grounding
| Attribute | Specification |
|---|---|
| 🎯 Purpose | Orchestrate Segment Anything Model 3 (SAM3) text-prompted auto-annotation across single or bulk batches. |
| ⚡ Key Capabilities | Multi-prompt zero-shot grounding, adjustable confidence thresholds, box expansion margin, and background job queueing with VRAM singleton management. |
| 💡 Invariant | One set_image per frame: SAM3 runs its heavy vision backbone once per image and re-runs only the lightweight grounding head across multiple prompts, ensuring maximum inference throughput. |
2.2 SAM3 Interactive Sandbox — Standalone prompt engineering playground
| Attribute | Specification |
|---|---|
| 🎯 Purpose | Interactive sandbox for testing text prompts, positive/negative point clicks, and mask segmentation before launching large auto-annotation jobs. |
| ⚡ Key Capabilities | Real-time mask rendering, multi-prompt layer toggles, point click prompt refinement, and raw JSON detection inspector. |
Phase 3: High-Throughput Annotation Studio & Triage
3.1 Roboflow-Grade Annotation Canvas & Quick Reclass — Sub-second keyboard navigation
Figure 3.1: Annotation Review Canvas with bounding box editor, shape provenance badges, and hotkey controls.
Figure 3.2: Filmstrip thumbnail navigation and floating single-keystroke quick reclassification bar.
| Key | Action | Key | Action | |
|---|---|---|---|---|
A |
Approve frame | ← → |
Previous / next frame | |
X |
Reject frame | U |
Jump to next unreviewed | |
Del |
Delete selected shape | 1–9 |
Quick switch active class | |
S + Drag |
SAM3-assisted box click | Drag | Add / move / resize bounding box |
| Attribute | Specification |
|---|---|
| 🎯 Purpose | Fast, ergonomic manual review and refinement of auto-generated bounding boxes. |
| ⚡ Key Capabilities | Single-key shortcuts, bottom thumbnail filmstrip with status indicators, and shape provenance tracking (SAM3, Manual, Base Model). |
| 💡 Invariant | Approving frames does not merge them into a dataset. Approval merely qualifies frames for Data Prep; merging happens under frozen rules. |
3.2 Visual Exemplar-Guided Prompting — Few-shot visual reference matching
| Attribute | Specification |
|---|---|
| 🎯 Purpose | Store and utilize positive and negative visual crop exemplars to guide SAM3 zero-shot grounding on difficult textures or ambiguous sack designs. |
| ⚡ Key Capabilities | Visual exemplar library, similarity threshold slider, 1-click exemplar addition from canvas bounding boxes. |
Phase 4: Data Prep, Statistical Triage & Dataset Freezing
4.1 Statistical Triage & Quality Outlier Filtering — Filter bad boxes without losing full images
| Attribute | Specification |
|---|---|
| 🎯 Purpose | Eliminate low-quality bounding boxes (partial crops, false positives, background noise) before dataset compilation. |
| ⚡ Key Capabilities | Dual-handle range sliders (Confidence Score, Box Area, Aspect Ratio), interactive SVG scatter plot, high-density crop grid cards, and real-time retention telemetry. |
| 💡 Invariant | Dropping an outlier bounding box leaves the frame in the dataset unless all boxes are dropped. Industrial conveyor frames contain ~44 objects; dropping whole frames discards 96% of good data to remove 10% of bad boxes. |
4.2 Augmentation Pipeline & Dataset Freezing — Deterministic Stable Val Split
| Attribute | Specification |
|---|---|
| 🎯 Purpose | Apply synthetic vision augmentations and freeze the curated batch into an immutable, versioned training dataset. |
| ⚡ Key Capabilities | HSV shift, brightness/contrast, rotation, blur, and mosaic controls; dataset destination selection; split ratio slider. |
| 💡 Invariant | Stable Validation Split: Validation assignment is derived deterministically from the frame's SHA-1 hash. Once a frame lands in val, it remains in val forever across all future versions. |
Phase 5: YOLO Retraining & Live Training Progress
5.1 Dataset Repository & YOLO Fine-Tuning — Hardware-aware parameter tuning
| Attribute | Specification |
|---|---|
| 🎯 Purpose | Train Ultralytics YOLO11 object detection models with automated hardware tuning and live telemetry. |
| ⚡ Key Capabilities | Architecture selection (YOLO11n/s/m/l/x), auto-calculated batch sizes based on free GPU VRAM, real-time loss sparklines (box_loss, cls_loss, dfl_loss), epoch progress bars, and streaming logs. |
| 💡 Invariant | Side-by-side Validation: After training, both the base model and fine-tuned model are benchmarked on the identical validation set, displaying signed metric deltas (\Delta\text{mAP50}, \Delta\text{precision}, \Delta\text{recall}). |
Phase 6: Counting Accuracy Benchmark & Live Production Inference
6.1 Counting Benchmark Matrix & Live Line Counter — 146 FPS verification
Figure 6.1: Counting Accuracy Benchmark Matrix comparing multi-model predictions against Ground Truth.
Figure 6.2: Real-time CCTV live counting interface with interactive tripwire line and ByteTrack trails.
| Attribute | Specification |
|---|---|
| 🎯 Purpose | Verify model counting precision against physical ground truth records and run real-time production counting. |
| ⚡ Key Capabilities | Headless counting benchmark running at 146 FPS, signed delta badges (+2, -1, 0), draggable tripwire counting line, ByteTrack trajectory tracking, and per-track JSONL diagnostic logs. |
| 💡 Invariant | Counting tripwire on upper edge (y_1): Tripwire evaluates the top edge coordinate of the sack bounding box rather than the center or bottom, preventing miscounts caused by physical sack deformation as it drops onto the conveyor. |
🎥 Production Counting Engine
The counting engine is designed for industrial conveyor belts with low or unstable camera frame rates. It eliminates reliance on instantaneous line-crossing frames by evaluating historical bounding box trajectories.
stateDiagram-v2
[*] --> UNKNOWN
UNKNOWN --> ABOVE: y1 coordinate above counting band
UNKNOWN --> BELOW: Born below band (ghost detection — rejected)
ABOVE --> COUNTED: Trajectory crosses band + travelled entry_travel_min
COUNTED --> ABOVE: Sustained frames above line (genuine reload)
Multi-Layer Counter Safeguards
| Guard Layer | Rule Specification | Failure Mode Prevented |
|---|---|---|
| Layer 1: Entry Origin Guard | Object must be observed above the counting band prior to crossing. | Prevents counting items that spawn directly inside the truck or loading chute. |
| Layer 2: Trajectory Distance | Object must travel a minimum distance (entry_travel_min) across consecutive frames. |
Discards transient noise and flickering phantom boxes. |
| Layer 3: Track Hand-Off | When a track is occluded, its movement history is parked for newborn tracks within handoff_radius. |
Prevents tracker ID switches from dropping or duplicating counts. |
| Layer 4: Directional Monotonicity | Strict single count per trajectory direction with verdict locked to the final confirmed motion. | Prevents double-counting during momentary conveyor pauses. |
⚙️ Configuration & Environment Matrix
All configuration parameters are defined via environment variables in .env:
| Variable | Type | Default Value | Scope | Description |
|---|---|---|---|---|
HF_TOKEN |
String |
(Required) | Backend / Docker | HuggingFace user access token with authorized permissions to download gated facebook/sam3 weights. |
VIDEO_ARCHIVE_HOST |
Path |
./data/archive |
Docker Compose | Host filesystem directory path containing raw CCTV video recordings (structured as <date>/<batch>.mp4). Relative paths resolve against the directory ./start.sh runs from. Mounted read-write; written only by user-initiated upload / date-folder creation (REQ-178), nothing else. |
APP_DATA_DIR |
Path |
./data (or /data in Docker) |
Backend | Base directory for application persistent data, SQLite database (app.db), projects, extracted frames, datasets, and model weights. |
VIDEO_ARCHIVE |
Path |
/videos (or data/archive) |
Backend | Internal container/local filesystem path where the video archive is browsed by FastAPI. |
WEB_PORT |
Integer |
9000 |
Docker / Nginx | Host HTTP port mapped to the Nginx frontend web UI. |
API_PORT |
Integer |
9010 |
Docker / FastAPI | Host HTTP port mapped to the FastAPI backend API. |
CORS_ORIGINS |
String |
http://localhost:5173,http://localhost:9000,http://localhost:9010 |
FastAPI Backend | Comma-separated list of allowed origins for Cross-Origin Resource Sharing. |
API_URL |
URL |
http://localhost:8000 |
Vite Dev Server | Backend target endpoint for Vite development proxy (frontend/vite.config.js). |
MEDIAMTX_WHEP_PATH |
String |
/whep |
Backend Live Count | WHEP WebRTC endpoint path on the streaming media server (MediaMTX). |
MEDIAMTX_RTSP_PORT |
Integer |
8554 |
Backend Live Count | RTSP stream port used to translate WHEP browser streams into backend video processing feeds. |
RTSP_TRANSPORT |
String |
tcp |
Backend Live Count | RTSP transport protocol (tcp or udp). TCP guarantees zero frame drop on industrial networks. |
PLAYBACK_URL |
URL |
http://192.168.192.96:9996/get |
Recorder Service | MediaMTX recording playback API endpoint for automated CCTV session extraction. |
PLAYBACK_PATH |
String |
cam |
Recorder Service | Stream channel identifier on the MediaMTX playback server. |
AUTO_PULL_INTERVAL |
Integer |
30 |
Auto-Pull Script | Polling frequency in seconds for automated Git repository synchronization (scripts/auto_pull.py). |
WEBHOOK_PORT |
Integer |
9000 |
Webhook Daemon | Port for the GitHub push webhook listener daemon (scripts/webhook.py). |
WEBHOOK_SECRET |
String |
"" |
Webhook Daemon | Shared secret key for validating GitHub webhook HMAC-SHA256 signatures. |
🗂️ Persistent Data & Storage Layout
All application state, relational metadata, and trained weights live in data/:
data/
├── app.db # SQLite metadata database with Write-Ahead Logging (WAL)
├── archive/<date>/batchNNN.mp4 # Raw CCTV video recordings (Mounted strictly read-only)
│ batchNNN.json # Sidecar metadata with true server timestamp
├── recorder.log # 24/7 background recorder daemon log
├── live-count/session-*.jsonl # Diagnostic per-track trajectory and crossing logs
└── projects/<slug>/ # Isolated project workspace
├── base/model.pt # Project base model weights and locked class taxonomy
├── batches/<id>/frames/ # Losslessly extracted image frames from video slices
├── datasets/<id>/ # Frozen Ultralytics YOLO formatted training datasets
│ ├── images/{train,val}/ # Immutable frame images
│ └── labels/{train,val}/ # YOLO format bounding box annotations (.txt)
└── models/<n>/ # Training runs (weights/best.pt, metrics.json, args.yaml)
# Named: {arch}-{labelType}-{epochs}ep-{classNames}-{date}
⚙️ Background Job Worker & Mutex Locking
Heavy computational operations run through an asynchronous background worker (backend/jobs.py) with strict GPU mutex locking to prevent VRAM over-allocation:
| Job Type | GPU Locked | Description |
|---|---|---|
extract |
— | Background FFmpeg extraction of video ranges into frame sequences. |
autolabel |
✅ | Meta SAM3 zero-shot grounding across candidate frames (one set_image per image). |
merge |
— | Compiles reviewed batches into immutable dataset splits under frozen triage rules. |
train |
✅ | Ultralytics YOLO11 fine-tuning followed by automated dual-model validation. |
count |
✅ | Headless evaluation benchmark running archive videos at 146 FPS. |
clock-scan |
— | OCR extraction of camera burned-in timestamps. |
truck-scan |
✅ | Batch inference check verifying the presence of target industrial objects. |
📊 Metrics Integrity & Domain Invariants
- Stable Validation Split: Validation assignment is determined deterministically by
sha1(image_bytes) % 100 < val_ratio. Once a frame lands in the validation split, it remains in validation across all future dataset iterations. This prevents validation leak and ensures that rising mAP scores reflect genuine model improvements. - One
set_imageper Frame: SAM3 executes its heavy vision transformer backbone once per image. Multi-class text prompting evaluates the lightweight grounding head against cached backbone embeddings. - Outlier Filtering Preserves Frames: Dropping a bounding box during Data Prep removes only the bad annotation. The image frame remains in the dataset as long as at least one valid box persists.
- Upper-Edge Coordinate Line Crossing (
y_1): Tripwires evaluate the top edge coordinate of bounding boxes (y_1) rather than the centroid (y_c) or bottom edge (y_2), ensuring immunity to sack deformation upon conveyor impact.
📚 Documentation & Deep Dives
Comprehensive technical specifications, operational SOPs, and architecture diagrams are available in the docs/ directory:
| Document | Format | Description | Action |
|---|---|---|---|
| Panduan Sistem Lengkap (Handover) | PDF (11 MB) |
Publication-Grade Master User Guide & Technical Manual (Indonesian) with complete operational SOPs and embedded figures. | ⬇️ Download PDF |
| Panduan Sistem Lengkap Source | FODT (11.8 MB) |
Native LibreOffice Writer Flat XML editable source document. | ⬇️ Download FODT |
| Panduan Sistem Lengkap Markdown | MD (68 KB) |
Full Markdown transcript of the 14-chapter system manual. | 📖 View MD |
| Architecture Flow Diagram (4K) | PNG (1.5 MB) |
4K Ultra-HD raster export of the 8-stage end-to-end retraining pipeline. | ⬇️ Download 4K |
| Architecture Flow Diagram (Vector) | SVG (34.6 KB) |
Scalable vector graphic diagram for high-resolution display. | ⬇️ Download SVG |
| Architecture Flow Diagram (Source) | FODG (24.8 KB) |
Native LibreOffice Draw Flat XML editable source file. | ⬇️ Download FODG |
| Entity Relationship Diagram (ERD) | MD (8.5 KB) |
Relational database schema, foreign keys, indexes, and Mermaid ERD diagram. | 📖 View ERD |
| System Requirements Specification | MD (23.6 KB) |
Numbered technical requirements (REQ-001 through REQ-042). |
📖 View Spec |
| System Design & Architecture | MD (22.7 KB) |
Database schema, REST API contracts, disk layouts, and backend invariants. | 📖 View Design |
| UI/UX Design Specification | MD (79.7 KB) |
Dark theme design tokens, hotkey maps, and component specifications. | 📖 View UI Spec |
| Ground Truth Benchmark Dataset | XLSX (586 KB) |
Hand-verified physical conveyor bag counts across operational shifts. | ⬇️ Download XLSX |
🔧 Troubleshooting & FAQ
Deployment & GPU Acceleration
| Symptom | Root Cause & Remediation |
|---|---|
could not select device driver |
NVIDIA Container Toolkit is missing or Docker Engine is older than CDI specifications. Run ./install_nvidia.sh or update Docker. |
CUDA out of memory during training |
Lower the batch size on the Models page or stop background jobs. SAM3 releases VRAM before training starts, but external processes may hold memory. |
| Changes to frontend or backend do not appear | Docker Compose caches container layers at build time. Run docker compose build backend frontend and hard-refresh your browser (Ctrl+Shift+R). |
| Job shows interrupted by server restart | The backend process stopped while a job was active. Jobs do not resume mid-epoch; simply re-trigger the job from the UI. |
SAM3 Foundation Auto-Annotation
| Symptom | Root Cause & Remediation |
|---|---|
| Job fails at Loading Model with HTTP 401 | HuggingFace token is invalid or access to facebook/sam3 has not yet been approved. |
| SAM3 checkpoint download is slow | HuggingFace Xet transfer throttling. HF_HUB_DISABLE_XET=1 is enabled by default in docker-compose.yml to bypass this issue. |
| SAM3 generates inaccurate bounding boxes | Refine the natural language prompt with physical descriptors (e.g. "woven polypropylene sack with blue logo") or add visual crops to the Exemplar Pool. |
Video Archive & Live Counting
| Symptom | Root Cause & Remediation |
|---|---|
| Video file marked as unreadable | FFmpeg/FFprobe could not parse the video header. Run uv run python scripts/transcode_archive.py to re-mux into standard H.264 MP4. |
| Shift cycle shows fewer recordings than expected | Check if recordings crossed the 06:00 boundary. Clips recorded before 06:00 belong to the previous operational shift cycle. |
| Live Counting shows GPU busy | An active fine-tuning or auto-annotation job holds the GPU lock. The live counter waits 30 seconds before falling back to CPU or queuing. |
🤝 Contributing & Engineering Guidelines
Development adheres strictly to the Chain of Truth methodology:
- Specifications First: Every feature must map directly to a numbered requirement in
docs/requirements.mdand architectural design indocs/design.md. - Verified Deliverables: Tasks tracked in
docs/tasks.mdflip to[DONE]only after concrete end-to-end verification. - Surgical Changes: Touch only code directly relevant to the feature. Adhere to the working rules in
AGENTS.md. - Package Manager: All backend dependencies are managed exclusively with Astral
uv(requirements.txt).
📄 License
Distributed under the MIT License. See LICENSE for details.













