asus 8f41c6c85a fix: exemplar draw keeps other classes' full-set preview (REQ-182 amended)
- AutoAnnotateModal: render merges last full-set detections for
  non-exemplared classes with conditioned results for exemplared ones
  (selected-not-exemplared filter); drawing an example no longer clears
  the other classes' detections
- Run Preview now always refreshes full-set first (exemplars: []) then
  re-runs exemplared classes
- docs: REQ-182 amended (merge semantics), design/ui-spec/tasks synced
2026-10-02 15:20:46 +07:00

reTraining AI Computer Vision Platform Logo

reTraining

Take a base model you already have, and make it measurably better with footage you already have.

The Open-Source, Self-Hosted Vision Pipeline to Turn Raw Industrial CCTV into Production-Grade YOLO Object Detectors & Line Counters.

Python 3.12 FastAPI React 19 Vite 7 SAM3 YOLO11 Docker CUDA License Documentation

Download Handover Docs PDF    Download 4K Diagram

⚡ Quickstart · 🏗️ Architecture · 🖼️ Visual UI Tour · 🎥 Counting Engine · ⚙️ Configuration · 📚 Documentation · 🔧 Troubleshooting


System Architecture & End-to-End Pipeline Diagram

Click diagram to view 4K UHD resolution. Scalable vector version available at docs/diagram-alur.svg.


💡 Why reTraining?

Deploying object detection models in industrial environments (manufacturing, logistics, agricultural feedmills, conveyor belts) often hits a painful bottleneck: general base models fail on domain-specific edge cases, while building manual labeling pipelines from scratch is slow and expensive.

reTraining provides a self-contained, enterprise-grade active learning platform designed to run directly on your edge server or GPU workstation:

  1. Ingest Raw CCTV Footage: Stream directly from continuous 24/7 video archives without re-encoding or modifying the underlying storage.
  2. Zero-Shot Foundation Auto-Labeling: Leverage Meta's Segment Anything Model 3 (SAM3) with natural language text prompts and visual exemplars to annotate thousands of frames in minutes.
  3. Roboflow-Grade Review Studio: Sub-second keyboard navigation, instant class switching, click-assist segmentation, and high-density visual triage crop grids.
  4. Statistical Triage & Data Prep: Filter bounding box outliers by area, aspect ratio, and confidence score without discarding valid image frames.
  5. Immutable Dataset Freezing & Stable Val Splits: Guarantee reproducible benchmarks with deterministic SHA-1 validation sets that never shift across retraining runs.
  6. Hardware-Aware Continuous Fine-Tuning: Auto-detect host GPU VRAM, fine-tune Ultralytics YOLO11 models, and evaluate base vs. fine-tuned model performance side-by-side.
  7. Live Production Counting & Benchmarks: Real-time RTSP/WHEP live inference with ByteTrack line-crossing counters, evaluated against verified ground truth physical counts at 146 FPS.

⚡ Quickstart

Prerequisites

  • Host OS: Linux (Ubuntu 22.04+ recommended) or Windows with WSL2.
  • GPU Acceleration: NVIDIA GPU with CUDA 12.4+ and NVIDIA Container Toolkit (Docker Engine 27+ CDI support).
  • HuggingFace Account: Gated model access granted for facebook/sam3 with a valid user access token (HF_TOKEN).
  • Video Storage: Directory of CCTV footage structured as <date>/<batch>.mp4 (or let the app mount ./data/archive).

Option 1: Docker Compose (Production Standard)

Launch the entire stack with a single command. The startup script automatically inspects the host hardware, generates Container Device Interface (CDI) specs for your NVIDIA GPU, and starts both backend and frontend containers.

# 1. Clone the repository
git clone git@github.com:fhanyuh/reTraining.git
cd reTraining

# 2. Configure environment credentials
cp .env.example .env
# Edit .env and paste your HuggingFace user access token:
# HF_TOKEN=hf_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx

# 3. Launch with automated GPU / CDI detection
chmod +x start.sh
./start.sh

Alternative direct Docker Compose launch:

docker compose up -d --build

Access Points:

Verify system health and GPU VRAM availability:

curl http://localhost:9010/api/health
{
  "device": "cuda",
  "gpu": "NVIDIA GeForce RTX 5080 Laptop GPU",
  "vram_free_gb": 14.91,
  "sam3_ready": true,
  "ffmpeg": true,
  "hf_token": true,
  "db": true
}

Important

The initial SAM3 auto-annotation job automatically downloads the ~3.4 GB SAM3 checkpoint from HuggingFace into a persistent Docker named volume (hf-cache). Subsequent executions load the model into VRAM in ~12 seconds.


Option 2: Local Development (Bare-Metal / uv + Vite)

For core development and live debugging without Docker:

Requirements: Python 3.12+, Astral uv, Node.js 20+, FFmpeg.

# 1. Clone and setup environment
git clone git@github.com:fhanyuh/reTraining.git
cd reTraining
cp .env.example .env

# 2. Install backend dependencies and vendored SAM3 with uv
uv pip install -r requirements.txt
uv pip install -e ./sam3

# 3. Start FastAPI backend (Port 8000)
uv run uvicorn backend.main:app --host 0.0.0.0 --port 8000 --reload

# 4. In a separate terminal, install and start Vite frontend (Port 5173)
cd frontend
npm install
npm run dev -- --host 0.0.0.0 --port 5173

Note

The frontend dependencies are pinned to Vite 7 (vite: ^7.1.5) in frontend/package.json to prevent Rolldown native binding bus errors on Linux platforms.


🏗️ End-to-End Pipeline Architecture

The platform operates as a continuous closed-loop retraining pipeline divided into 7 distinct functional stages:

┌──────────────────────────────────────────────────────────────────────────────────────────────────┐
│                                 DATA RETRAINING & INFERENCE PIPELINE                             │
└──────────────────────────────────────────────────────────────────────────────────────────────────┘
  [1. Video Archive] ──▶ [2. Frame Slicing] ──▶ [3. SAM3 Auto-Label] ──▶ [4. Review & Exemplar]
       (CCTV/MP4)           (FFmpeg In/Out)       (Prompt Grounding)        (Click-Assist Canvas)
                                                                                     │
  [7. Live Counter]  ◀── [6. YOLO11 Train]  ◀── [5. Immutable Split] ◀── [4. Triage & Filter]
     (WHEP / RTSP)          (Auto VRAM)           (Stable Train/Val)         (Outlier Purge)
  1. Video Archive & Shift Slicing: Raw CCTV recordings are indexed into 24-hour operational work shifts (06:00 to 05:59 next morning). Sub-second timeline in/out trimming extracts high-resolution frame sequences with configurable FPS rates via asynchronous FFmpeg queues.
  2. SAM3 Foundation Auto-Labeling: Meta SAM3 zero-shot open-vocabulary grounding generates candidate bounding boxes and segmentation masks from natural language descriptions (e.g. "white sack of feed on conveyor").
  3. Interactive Review Studio: Operators review candidate annotations on a Roboflow-grade canvas with single-keystroke approvals, box adjustments, click-assist segmentation, and visual prompt exemplar refinement.
  4. Statistical Triage & Quality Outliers: Interactive 2D scatter plots (Box Area vs. Confidence Score) and high-density crop grids allow rapid isolation and pruning of false positives without discarding valid frames.
  5. Immutable Dataset Compilation: Filtered batches are merged into versioned datasets (v1, v2, v3) with deterministic SHA-1 validation hashing, guaranteeing that validation images remain permanently locked across iterations.
  6. Hardware-Aware YOLO Retraining: Hyperparameters and batch sizes auto-scale based on detected GPU VRAM. The system fine-tunes Ultralytics YOLO11, benchmarks old vs. new models on the identical validation set, and displays signed metric deltas (\Delta\text{mAP50}, \Delta\text{Precision}, \Delta\text{Recall}).
  7. Counting Accuracy Benchmark & Live Inference: Real-time production inference evaluates live RTSP/WHEP video streams with ByteTrack trajectory tracking and upper-edge tripwire counters, benchmarking against hand-verified physical ground truth logs at 146 FPS.

Explore the 6 core pipeline phases across all 11 primary user screens and modal workflows.


Phase 1: Video Ingest, Shift Cycles & Frame Sampling

1.1 Project Workspace & Setup — Multi-project isolation and class locking
Projects Overview Page

Figure 1.1: Projects Dashboard showing active projects, base models, class taxonomies, and dataset stats.

Attribute Specification
🎯 Purpose Central management hub for all isolated computer vision projects.
⚡ Key Capabilities View base model architecture, active class tags, total extracted batches, and dataset snapshots at a glance.
💡 Invariant Projects maintain strictly isolated database records, class lists, and filesystem storage roots under data/projects/<slug>/.

Project Creation Modal

Figure 1.2: Project Creation dialog with base model checkpoint upload and automatic class extraction.

Attribute Specification
🎯 Purpose Initialize a new project with locked class names and base model weights.
⚡ Key Capabilities Upload an existing YOLO .pt checkpoint to automatically extract its class taxonomy, or define custom classes and start fine-tuning from yolo11n.pt.
💡 Invariant Base model classes are permanently locked to the project to prevent label drift between training iterations.
1.2 24-Hour Operational Shift Archive — Cycle grouping that crosses midnight
Video Archive Shift Cycle Browser

Figure 1.3: Video Archive grouping recordings into 24-hour operational shifts (06:00 to 05:59 next morning).

Attribute Specification
🎯 Purpose Browse raw CCTV recordings grouped by operational work shifts rather than arbitrary calendar folders.
⚡ Key Capabilities Reads camera burned-in OCR timestamps and sidecar .json metadata to assign recordings accurately across midnight boundaries.
💡 Invariant The video archive is mounted strictly read-only (:ro). Nothing is ever modified, renamed, or deleted in the user's video repository.
1.3 Video Trim & Frame Extractor — Sub-second timeline scrubbing and sampling
Video Trim and Frame Extractor

Figure 1.4: Video Trim interface with timeline range sliders and real-time frame calculation.

Attribute Specification
🎯 Purpose Select active truck loading ranges and configure extraction frame rates.
⚡ Key Capabilities Interactive in/out timeline markers, extraction FPS slider (0.5 – 2.0 FPS), real-time output frame counter, and background FFmpeg queueing.
💡 Invariant Extraction runs asynchronously in the background queue; frames are losslessly sampled into data/projects/<slug>/batches/<id>/frames/.

Phase 2: SAM3 Zero-Shot Auto-Labeling Engine

2.1 Batch Management & Mass Auto-Annotation — High-throughput zero-shot grounding
Batch Management Dashboard

Figure 2.1: Batch Management table with batch status badges and bulk action triggers.


Single Batch SAM3 Auto-Annotate Modal

Figure 2.2: Single-batch SAM3 auto-annotation configuration with natural language text prompts.


Mass Auto-Annotate Modal

Figure 2.3: Mass Auto-Annotate dialog for enqueuing thousands of frames across multiple batches.

Attribute Specification
🎯 Purpose Orchestrate Segment Anything Model 3 (SAM3) text-prompted auto-annotation across single or bulk batches.
⚡ Key Capabilities Multi-prompt zero-shot grounding, adjustable confidence thresholds, box expansion margin, and background job queueing with VRAM singleton management.
💡 Invariant One set_image per frame: SAM3 runs its heavy vision backbone once per image and re-runs only the lightweight grounding head across multiple prompts, ensuring maximum inference throughput.
2.2 SAM3 Interactive Sandbox — Standalone prompt engineering playground
SAM3 Interactive Playground

Figure 2.4: SAM3 Interactive Playground for zero-shot text prompting and point prompt testing.

Attribute Specification
🎯 Purpose Interactive sandbox for testing text prompts, positive/negative point clicks, and mask segmentation before launching large auto-annotation jobs.
⚡ Key Capabilities Real-time mask rendering, multi-prompt layer toggles, point click prompt refinement, and raw JSON detection inspector.

Phase 3: High-Throughput Annotation Studio & Triage

3.1 Roboflow-Grade Annotation Canvas & Quick Reclass — Sub-second keyboard navigation
Annotation Review Canvas

Figure 3.1: Annotation Review Canvas with bounding box editor, shape provenance badges, and hotkey controls.


Review Filmstrip and Quick Reclass Bar

Figure 3.2: Filmstrip thumbnail navigation and floating single-keystroke quick reclassification bar.

Key Action Key Action
A Approve frame ← → Previous / next frame
X Reject frame U Jump to next unreviewed
Del Delete selected shape 1–9 Quick switch active class
S + Drag SAM3-assisted box click Drag Add / move / resize bounding box
Attribute Specification
🎯 Purpose Fast, ergonomic manual review and refinement of auto-generated bounding boxes.
⚡ Key Capabilities Single-key shortcuts, bottom thumbnail filmstrip with status indicators, and shape provenance tracking (SAM3, Manual, Base Model).
💡 Invariant Approving frames does not merge them into a dataset. Approval merely qualifies frames for Data Prep; merging happens under frozen rules.
3.2 Visual Exemplar-Guided Prompting — Few-shot visual reference matching
Exemplar Pool Panel

Figure 3.3: Exemplar Pool Sidebar Panel for visual reference matching.

Attribute Specification
🎯 Purpose Store and utilize positive and negative visual crop exemplars to guide SAM3 zero-shot grounding on difficult textures or ambiguous sack designs.
⚡ Key Capabilities Visual exemplar library, similarity threshold slider, 1-click exemplar addition from canvas bounding boxes.

Phase 4: Data Prep, Statistical Triage & Dataset Freezing

4.1 Statistical Triage & Quality Outlier Filtering — Filter bad boxes without losing full images
Data Prep Quality Outliers Panel

Figure 4.1: Quality filter sliders with live box and frame retention counters.


Triage Scatter Plot

Figure 4.2: Interactive Scatter Plot (Score vs Area) for instant visual outlier cluster detection.


Triage Crop Grid

Figure 4.3: Triage Crop Grid for rapid bulk visual inspection and 1-click outlier removal.

Attribute Specification
🎯 Purpose Eliminate low-quality bounding boxes (partial crops, false positives, background noise) before dataset compilation.
⚡ Key Capabilities Dual-handle range sliders (Confidence Score, Box Area, Aspect Ratio), interactive SVG scatter plot, high-density crop grid cards, and real-time retention telemetry.
💡 Invariant Dropping an outlier bounding box leaves the frame in the dataset unless all boxes are dropped. Industrial conveyor frames contain ~44 objects; dropping whole frames discards 96% of good data to remove 10% of bad boxes.
4.2 Augmentation Pipeline & Dataset Freezing — Deterministic Stable Val Split
Augmentation Configuration Panel

Figure 4.4: Albumentations training augmentation presets with real-time visual preview.


Merge Target and Dataset Freeze Modal

Figure 4.5: Dataset Freeze dialog enforcing deterministic SHA-1 Stable Val Split rules.

Attribute Specification
🎯 Purpose Apply synthetic vision augmentations and freeze the curated batch into an immutable, versioned training dataset.
⚡ Key Capabilities HSV shift, brightness/contrast, rotation, blur, and mosaic controls; dataset destination selection; split ratio slider.
💡 Invariant Stable Validation Split: Validation assignment is derived deterministically from the frame's SHA-1 hash. Once a frame lands in val, it remains in val forever across all future versions.

Phase 5: YOLO Retraining & Live Training Progress

5.1 Dataset Repository & YOLO Fine-Tuning — Hardware-aware parameter tuning
Master Datasets List

Figure 5.1: Master Dataset repository showing version history, train/val splits, and export tools.


YOLO Models and Training Page

Figure 5.2: YOLO fine-tuning parameter setup with hardware detection and model history.


Live Training Workflow Progress

Figure 5.3: Real-time training telemetry with live loss curves, mAP metrics, and stdout logs.

Attribute Specification
🎯 Purpose Train Ultralytics YOLO11 object detection models with automated hardware tuning and live telemetry.
⚡ Key Capabilities Architecture selection (YOLO11n/s/m/l/x), auto-calculated batch sizes based on free GPU VRAM, real-time loss sparklines (box_loss, cls_loss, dfl_loss), epoch progress bars, and streaming logs.
💡 Invariant Side-by-side Validation: After training, both the base model and fine-tuned model are benchmarked on the identical validation set, displaying signed metric deltas (\Delta\text{mAP50}, \Delta\text{precision}, \Delta\text{recall}).

Phase 6: Counting Accuracy Benchmark & Live Production Inference

6.1 Counting Benchmark Matrix & Live Line Counter — 146 FPS verification
Counting Accuracy Benchmark Matrix

Figure 6.1: Counting Accuracy Benchmark Matrix comparing multi-model predictions against Ground Truth.


Live Video Inference and Line Counter

Figure 6.2: Real-time CCTV live counting interface with interactive tripwire line and ByteTrack trails.

Attribute Specification
🎯 Purpose Verify model counting precision against physical ground truth records and run real-time production counting.
⚡ Key Capabilities Headless counting benchmark running at 146 FPS, signed delta badges (+2, -1, 0), draggable tripwire counting line, ByteTrack trajectory tracking, and per-track JSONL diagnostic logs.
💡 Invariant Counting tripwire on upper edge (y_1): Tripwire evaluates the top edge coordinate of the sack bounding box rather than the center or bottom, preventing miscounts caused by physical sack deformation as it drops onto the conveyor.

🎥 Production Counting Engine

The counting engine is designed for industrial conveyor belts with low or unstable camera frame rates. It eliminates reliance on instantaneous line-crossing frames by evaluating historical bounding box trajectories.

stateDiagram-v2
    [*] --> UNKNOWN
    UNKNOWN --> ABOVE: y1 coordinate above counting band
    UNKNOWN --> BELOW: Born below band (ghost detection — rejected)
    ABOVE --> COUNTED: Trajectory crosses band + travelled entry_travel_min
    COUNTED --> ABOVE: Sustained frames above line (genuine reload)

Multi-Layer Counter Safeguards

Guard Layer Rule Specification Failure Mode Prevented
Layer 1: Entry Origin Guard Object must be observed above the counting band prior to crossing. Prevents counting items that spawn directly inside the truck or loading chute.
Layer 2: Trajectory Distance Object must travel a minimum distance (entry_travel_min) across consecutive frames. Discards transient noise and flickering phantom boxes.
Layer 3: Track Hand-Off When a track is occluded, its movement history is parked for newborn tracks within handoff_radius. Prevents tracker ID switches from dropping or duplicating counts.
Layer 4: Directional Monotonicity Strict single count per trajectory direction with verdict locked to the final confirmed motion. Prevents double-counting during momentary conveyor pauses.

⚙️ Configuration & Environment Matrix

All configuration parameters are defined via environment variables in .env:

Variable Type Default Value Scope Description
HF_TOKEN String (Required) Backend / Docker HuggingFace user access token with authorized permissions to download gated facebook/sam3 weights.
VIDEO_ARCHIVE_HOST Path ./data/archive Docker Compose Host filesystem directory path containing raw CCTV video recordings (structured as <date>/<batch>.mp4). Relative paths resolve against the directory ./start.sh runs from. Mounted read-write; written only by user-initiated upload / date-folder creation (REQ-178), nothing else.
APP_DATA_DIR Path ./data (or /data in Docker) Backend Base directory for application persistent data, SQLite database (app.db), projects, extracted frames, datasets, and model weights.
VIDEO_ARCHIVE Path /videos (or data/archive) Backend Internal container/local filesystem path where the video archive is browsed by FastAPI.
WEB_PORT Integer 9000 Docker / Nginx Host HTTP port mapped to the Nginx frontend web UI.
API_PORT Integer 9010 Docker / FastAPI Host HTTP port mapped to the FastAPI backend API.
CORS_ORIGINS String http://localhost:5173,http://localhost:9000,http://localhost:9010 FastAPI Backend Comma-separated list of allowed origins for Cross-Origin Resource Sharing.
API_URL URL http://localhost:8000 Vite Dev Server Backend target endpoint for Vite development proxy (frontend/vite.config.js).
MEDIAMTX_WHEP_PATH String /whep Backend Live Count WHEP WebRTC endpoint path on the streaming media server (MediaMTX).
MEDIAMTX_RTSP_PORT Integer 8554 Backend Live Count RTSP stream port used to translate WHEP browser streams into backend video processing feeds.
RTSP_TRANSPORT String tcp Backend Live Count RTSP transport protocol (tcp or udp). TCP guarantees zero frame drop on industrial networks.
PLAYBACK_URL URL http://192.168.192.96:9996/get Recorder Service MediaMTX recording playback API endpoint for automated CCTV session extraction.
PLAYBACK_PATH String cam Recorder Service Stream channel identifier on the MediaMTX playback server.
AUTO_PULL_INTERVAL Integer 30 Auto-Pull Script Polling frequency in seconds for automated Git repository synchronization (scripts/auto_pull.py).
WEBHOOK_PORT Integer 9000 Webhook Daemon Port for the GitHub push webhook listener daemon (scripts/webhook.py).
WEBHOOK_SECRET String "" Webhook Daemon Shared secret key for validating GitHub webhook HMAC-SHA256 signatures.

🗂️ Persistent Data & Storage Layout

All application state, relational metadata, and trained weights live in data/:

data/
├── app.db                                  # SQLite metadata database with Write-Ahead Logging (WAL)
├── archive/<date>/batchNNN.mp4             # Raw CCTV video recordings (Mounted strictly read-only)
│                  batchNNN.json            # Sidecar metadata with true server timestamp
├── recorder.log                            # 24/7 background recorder daemon log
├── live-count/session-*.jsonl              # Diagnostic per-track trajectory and crossing logs
└── projects/<slug>/                        # Isolated project workspace
    ├── base/model.pt                       # Project base model weights and locked class taxonomy
    ├── batches/<id>/frames/                # Losslessly extracted image frames from video slices
    ├── datasets/<id>/                      # Frozen Ultralytics YOLO formatted training datasets
    │   ├── images/{train,val}/             # Immutable frame images
    │   └── labels/{train,val}/             # YOLO format bounding box annotations (.txt)
    └── models/<n>/                         # Training runs (weights/best.pt, metrics.json, args.yaml)
                                            # Named: {arch}-{labelType}-{epochs}ep-{classNames}-{date}

⚙️ Background Job Worker & Mutex Locking

Heavy computational operations run through an asynchronous background worker (backend/jobs.py) with strict GPU mutex locking to prevent VRAM over-allocation:

Job Type GPU Locked Description
extract — Background FFmpeg extraction of video ranges into frame sequences.
autolabel ✅ Meta SAM3 zero-shot grounding across candidate frames (one set_image per image).
merge — Compiles reviewed batches into immutable dataset splits under frozen triage rules.
train ✅ Ultralytics YOLO11 fine-tuning followed by automated dual-model validation.
count ✅ Headless evaluation benchmark running archive videos at 146 FPS.
clock-scan — OCR extraction of camera burned-in timestamps.
truck-scan ✅ Batch inference check verifying the presence of target industrial objects.

📊 Metrics Integrity & Domain Invariants

  1. Stable Validation Split: Validation assignment is determined deterministically by sha1(image_bytes) % 100 < val_ratio. Once a frame lands in the validation split, it remains in validation across all future dataset iterations. This prevents validation leak and ensures that rising mAP scores reflect genuine model improvements.
  2. One set_image per Frame: SAM3 executes its heavy vision transformer backbone once per image. Multi-class text prompting evaluates the lightweight grounding head against cached backbone embeddings.
  3. Outlier Filtering Preserves Frames: Dropping a bounding box during Data Prep removes only the bad annotation. The image frame remains in the dataset as long as at least one valid box persists.
  4. Upper-Edge Coordinate Line Crossing (y_1): Tripwires evaluate the top edge coordinate of bounding boxes (y_1) rather than the centroid (y_c) or bottom edge (y_2), ensuring immunity to sack deformation upon conveyor impact.

📚 Documentation & Deep Dives

Comprehensive technical specifications, operational SOPs, and architecture diagrams are available in the docs/ directory:

Document Format Description Action
Panduan Sistem Lengkap (Handover) PDF (11 MB) Publication-Grade Master User Guide & Technical Manual (Indonesian) with complete operational SOPs and embedded figures. ⬇️ Download PDF
Panduan Sistem Lengkap Source FODT (11.8 MB) Native LibreOffice Writer Flat XML editable source document. ⬇️ Download FODT
Panduan Sistem Lengkap Markdown MD (68 KB) Full Markdown transcript of the 14-chapter system manual. 📖 View MD
Architecture Flow Diagram (4K) PNG (1.5 MB) 4K Ultra-HD raster export of the 8-stage end-to-end retraining pipeline. ⬇️ Download 4K
Architecture Flow Diagram (Vector) SVG (34.6 KB) Scalable vector graphic diagram for high-resolution display. ⬇️ Download SVG
Architecture Flow Diagram (Source) FODG (24.8 KB) Native LibreOffice Draw Flat XML editable source file. ⬇️ Download FODG
Entity Relationship Diagram (ERD) MD (8.5 KB) Relational database schema, foreign keys, indexes, and Mermaid ERD diagram. 📖 View ERD
System Requirements Specification MD (23.6 KB) Numbered technical requirements (REQ-001 through REQ-042). 📖 View Spec
System Design & Architecture MD (22.7 KB) Database schema, REST API contracts, disk layouts, and backend invariants. 📖 View Design
UI/UX Design Specification MD (79.7 KB) Dark theme design tokens, hotkey maps, and component specifications. 📖 View UI Spec
Ground Truth Benchmark Dataset XLSX (586 KB) Hand-verified physical conveyor bag counts across operational shifts. ⬇️ Download XLSX

🔧 Troubleshooting & FAQ

Deployment & GPU Acceleration
Symptom Root Cause & Remediation
could not select device driver NVIDIA Container Toolkit is missing or Docker Engine is older than CDI specifications. Run ./install_nvidia.sh or update Docker.
CUDA out of memory during training Lower the batch size on the Models page or stop background jobs. SAM3 releases VRAM before training starts, but external processes may hold memory.
Changes to frontend or backend do not appear Docker Compose caches container layers at build time. Run docker compose build backend frontend and hard-refresh your browser (Ctrl+Shift+R).
Job shows interrupted by server restart The backend process stopped while a job was active. Jobs do not resume mid-epoch; simply re-trigger the job from the UI.
SAM3 Foundation Auto-Annotation
Symptom Root Cause & Remediation
Job fails at Loading Model with HTTP 401 HuggingFace token is invalid or access to facebook/sam3 has not yet been approved.
SAM3 checkpoint download is slow HuggingFace Xet transfer throttling. HF_HUB_DISABLE_XET=1 is enabled by default in docker-compose.yml to bypass this issue.
SAM3 generates inaccurate bounding boxes Refine the natural language prompt with physical descriptors (e.g. "woven polypropylene sack with blue logo") or add visual crops to the Exemplar Pool.
Video Archive & Live Counting
Symptom Root Cause & Remediation
Video file marked as unreadable FFmpeg/FFprobe could not parse the video header. Run uv run python scripts/transcode_archive.py to re-mux into standard H.264 MP4.
Shift cycle shows fewer recordings than expected Check if recordings crossed the 06:00 boundary. Clips recorded before 06:00 belong to the previous operational shift cycle.
Live Counting shows GPU busy An active fine-tuning or auto-annotation job holds the GPU lock. The live counter waits 30 seconds before falling back to CPU or queuing.

🤝 Contributing & Engineering Guidelines

Development adheres strictly to the Chain of Truth methodology:

  • Specifications First: Every feature must map directly to a numbered requirement in docs/requirements.md and architectural design in docs/design.md.
  • Verified Deliverables: Tasks tracked in docs/tasks.md flip to [DONE] only after concrete end-to-end verification.
  • Surgical Changes: Touch only code directly relevant to the feature. Adhere to the working rules in AGENTS.md.
  • Package Manager: All backend dependencies are managed exclusively with Astral uv (requirements.txt).

📄 License

Distributed under the MIT License. See LICENSE for details.

S
Description
used for retraining and annotation of karung feedmill project
Readme MIT
29 MiB
0 Stars 1 Watchers 0 Forks
Languages
Python 65.2%
JavaScript 31.8%
CSS 2.3%
Shell 0.4%
Dockerfile 0.2%