Metadata-Version: 2.4
Name: weight-estimator
Version: 0.1.0
Summary: Offline chicken weight estimation from top-down video and scale data.
Requires-Python: >=3.10
Description-Content-Type: text/markdown
Requires-Dist: numpy>=1.26
Requires-Dist: opencv-python>=4.10
Requires-Dist: PyYAML>=6.0.2
Requires-Dist: pandas>=2.2
Requires-Dist: scipy>=1.11
Requires-Dist: scikit-learn>=1.4
Requires-Dist: matplotlib>=3.8
Requires-Dist: pyarrow>=15.0
Requires-Dist: joblib>=1.3
Provides-Extra: jetson
Requires-Dist: ultralytics>=8.4.38; extra == "jetson"

# Weight Estimator

Offline experiment pipeline for estimating daily average broiler weight per cage
from cam2/cam3 top-down video and physical scale readings.

## Install

```bash
cd try-weight-chicken
python -m pip install -e .

# For Jetson video extraction (optional on dev machines):
python -m pip install -e ../chicken-sukawarna-existing
python -m pip install -e ".[jetson]"
```

## Predict average weight (daily use)

DOC is **required**. Formula: `DOC weight + Gompertz(gain|age) + vision residual`.

After you have a trained `output/model_bundle.joblib`:

```bash
source wchicken/bin/activate

# Prefer both cameras (FUSED). One camera is OK.
# Age can be --cycle-day OR derived from --date and --doc-date.
weight-estimator predict \
  --config configs/weight_cc2_cc3.yaml \
  --doc-weight 32 \
  --doc-date 2026-05-22 \
  --date 2026-06-15 \
  --model output/model_bundle.joblib \
  --cc2 /path/to/kandang_1_camera_2_XXXX.mp4 \
  --cc3 /path/to/kandang_1_camera_3_XXXX.mp4 \
  --output-dir output/predict_2026-06-15
```

Cycle6 example (Age 21 / Manual ~803 g) — use **that flock’s** DOC weight (e.g. Age-0 Manual ~32 g):

```bash
weight-estimator predict \
  --config configs/weight_cc2_cc3.yaml \
  --doc-weight 32 \
  --doc-date 2026-04-02 \
  --date 2026-04-23 \
  --model output/model_bundle.joblib \
  --cc2 /path/to/cycle6/cc2.mp4 \
  --cc3 /path/to/cycle6/cc3.mp4 \
  --output-dir output/predict_cycle6_2026-04-23
```

Prints `predicted_avg_g` and writes `prediction.json` under `--output-dir`.

## Retrain after DOC / label changes (GPU)

Confirm `DOC Weight (g)` / `DOC Date` in `historical_weights_avg.csv`, then:

```bash
source wchicken/bin/activate
weight-estimator audit --config configs/weight_cc2_cc3.yaml
# Re-extract only if days or filters changed; otherwise:
weight-estimator aggregate --config configs/weight_cc2_cc3.yaml
# Production bundle trains ages train_cycle_day_min..max (default 1–35)
# model_mode: vision_primary → pred = DOC + max(0, vision_gain(features))
weight-estimator train --config configs/weight_cc2_cc3.yaml --model ridge
# Evaluate still uses late holdout ages 29–35 (train proxy 1–28) for GO report
weight-estimator evaluate --config configs/weight_cc2_cc3.yaml
```

No re-extract needed after switching to vision-primary if `cage_day_features.parquet` already exists
**and** sampling targets / features are unchanged.

After changing **smart camera fusion** (`fuse_min_detections` / weighted fuse), re-run
**aggregate → train → evaluate** only (no re-extract).

## Richer sampling re-extract (GPU)

After raising `target_detections` (default **1000** per camera/day), denser stride, and richer vision features,
**re-extract** so detections reflect the new targets:

```bash
source wchicken/bin/activate
pip install -e .

# optional rollback copy
cp -a output output_pre_richer

tmux new -s weight-extract
PYTHONUNBUFFERED=1 weight-estimator extract --config configs/weight_cc2_cc3.yaml \
  2>&1 | tee logs/extract_richer_$(date +%Y%m%d_%H%M).log
# detach: Ctrl+b d

weight-estimator aggregate --config configs/weight_cc2_cc3.yaml
weight-estimator train --config configs/weight_cc2_cc3.yaml --model ridge
weight-estimator evaluate --config configs/weight_cc2_cc3.yaml
```

## Quick start (synthetic demo)

When videos are not available locally, run the full pipeline with synthetic detections:

```bash
weight-estimator run-all --config configs/weight_cc2_cc3.yaml --synthetic
```

Outputs land in `output/`:

- `audit_report.json`
- `calibration_audit.json`
- `detections.parquet`
- `cage_day_features.parquet`
- `model_bundle.joblib`
- `validation_report.json`
- `bland_altman_*.png`

## Real video workflow (Jetson or GPU server)

Confirm `video_root` and `csv_path` in [`configs/weight_cc2_cc3.yaml`](configs/weight_cc2_cc3.yaml) match the machine where videos (and labels) live.

```bash
# Activate the project venv (local folder), not a conda env named wchicken:
#   source wchicken/bin/activate
#   which weight-estimator   # should be .../wchicken/bin/weight-estimator

# 1. Audit CSV ↔ video mapping
weight-estimator audit --config configs/weight_cc2_cc3.yaml

# 2. Calibration frame audit
weight-estimator calibration-audit --config configs/weight_cc2_cc3.yaml

# 3. Validate homography (<5% error target)
weight-estimator validate-calibration --config configs/weight_cc2_cc3.yaml

# 4. Extract features — prefer tmux for long runs (see below)
# 5. Aggregate to cage-day features
weight-estimator aggregate --config configs/weight_cc2_cc3.yaml

# 6. Train Gompertz + residual model
weight-estimator train --config configs/weight_cc2_cc3.yaml --model ridge

# 7. Validate and go/no-go report
weight-estimator evaluate --config configs/weight_cc2_cc3.yaml
```

Steps 5–7 only need `detections.parquet` and the CSV; they can run later on CPU after extract finishes.

### Long extract with tmux

`extract` can take hours. Run it inside **tmux** so SSH disconnect or shutting your laptop does not kill the job.

Optional smoke test in the foreground until you see a model line:

```bash
PYTHONUNBUFFERED=1 weight-estimator extract --config configs/weight_cc2_cc3.yaml
# Expect: [model] loaded ... then leave it a minute; Ctrl+C when healthy
```

Full run:

```bash
tmux new -s weight-extract
cd /path/to/try-weight-chicken
source wchicken/bin/activate
mkdir -p logs
PYTHONUNBUFFERED=1 weight-estimator extract --config configs/weight_cc2_cc3.yaml \
  2>&1 | tee logs/extract_$(date +%Y%m%d_%H%M).log
# Detach (job keeps running): Ctrl+b, then d
```

Monitor / reattach:

```bash
tmux attach -t weight-extract
tail -f logs/extract_*.log
pgrep -af weight-estimator
nvidia-smi          # GPU server (e.g. L40)
# tegrastats        # Jetson
```

**Extract log feedback:**

- `[extract] DATE CCx: frames=N window=1-N full_video=True ...` at the start of each day/camera
- `[model] loaded ...` when the tracker is created
- `[extract] DATE CCx: done valid=... inferred=...` when that video finishes
- Final CLI line: `Extracted N detections -> output/detections.parquet`
- Log name `extract_YYYYMMDD_HHMM.log` is the **job start time**, not a video date

## Configuration

- [`configs/weight_cc2_cc3.yaml`](configs/weight_cc2_cc3.yaml) — data paths, sampling, quality filters
- [`configs/calibration/CC2_homography.yaml`](configs/calibration/CC2_homography.yaml) — CC2 pixel→cm points
- [`configs/calibration/CC3_homography.yaml`](configs/calibration/CC3_homography.yaml) — CC3 pixel→cm points

Refine calibration point pairs after inspecting `output/calibration_frames/`.

## Quality filters (relaxed for dense late-cycle frames)

Detections must pass **all** tiers before entering the weight sample:

1. **Class gate** — detector ignores class 1 (not-chicken); only class 0 continues; reject half-chicken (2)
2. **Containment** — ≥85% bbox overlap with ROI and full bbox inside red-line sampling corridors
3. **Isolation** — max IoU &lt; 0.30 AND min centroid distance
4. **Shape** — confidence, aspect ratio, area outlier checks (vs frame median)
5. **Frame gate** — clean_ratio and ≥ 1 accepted bird per frame
6. **Size distribution (second filter)** — per camera-day MAD z-score on `area_cm2` (`size_distribution` in YAML). Drops flock-inconsistent sizes after frame tiers.

Rejection breakdown is written to `output/filter_report.json` (includes `size_distribution` per day/camera).

**YOLO vs post-filter confidence:** YOLO `conf` / `iou` live in [`cycle7_batch.yaml`](../chicken-counting/chicken-sukawarna-existing/configs/cycle7_batch.yaml) (shared with chicken counting). Batch `early_cycle` lowers YOLO thresholds on days 1–21 (`conf` 0.15 early vs 0.25 base); extract/overlay apply the same merge via `cycle_day`. Weight `quality.conf_min` in [`weight_cc2_cc3.yaml`](configs/weight_cc2_cc3.yaml) must stay aligned (e.g. base `conf_min` `0.30`, early-cycle `0.15`). Early-cycle `min_overlap_ratio` relaxes `not_fully_inside_roi` rejects. Re-extract and re-run overlays after changing either file.

```bash
# Visual QA: green=accepted, red=rejected (requires local video)
# Note: overlays show frame tiers only — day-level size rejects are not painted here.
# For early-cycle clips (small chicks), pass --cycle-day to match extract thresholds.
weight-estimator filter-overlay --config configs/weight_cc2_cc3.yaml \
  --video /path/to/kandang_1_camera_2_XXXX.mp4 --camera-id CC2 --cycle-day 1

# Re-apply day-level size filter to existing detections.parquet (no re-extract)
weight-estimator aggregate --config configs/weight_cc2_cc3.yaml --refilter-size

# Compare strict vs legacy filter impact
weight-estimator compare-filters --config configs/weight_cc2_cc3.yaml
```

## Design notes

- **DOC-essential:** labels require `Cycle ID`, `DOC Date`, `DOC Weight (g)`. Age must equal `(Tanggal - DOC Date).days`.
- **Vision-primary (default):** `predicted_avg_g = DOC + max(0, vision_gain(bbox/area features))`. Gompertz age curve is **diagnostic only** (`age_baseline_g` / age_only ablation), not the additive core.
- **Unit of analysis:** one row per `(cage, date)`; detections are aggregated, not individually labeled.
- **Sampling:** full-video adaptive stride; default target **1000** detections per camera/day (early and late), stride 15→5, `per_track_cap` 8. Re-extract after changing these.
- **Camera fusion:** gate cams below `fuse_min_detections` (default 50), then weight survivors by `detection_count`; if none pass, use best single cam. Re-aggregate (no re-extract) after changing fusion.
- **Segmentation pilot:** weight estimator uses YOLO-seg via `detection.model_path` override (chicken class only). Mask contour geometry is preferred with bbox fallback (`geometry` block; `min_mask_fill_ratio` default 0.30). Shape cage-day stats use **mask rows only** when `mask_count >= min_mask_for_stats` (default **20**). Extra mask features: `solidity`, `equiv_diameter_cm`, `mask_fill_ratio`, `mask_usage_ratio`. FUSED rows propagate `shape_stats_mask_only`. Counting package: `chicken-sukawarna-existing`. Feature-preset / aggregate changes need **re-aggregate + retrain** (no full extract).
- **Vision features:** production preset `features.preset: mask_max` = area/axes/perimeter/compactness/eccentricity/aspect/solidity/equiv diameter/mask fill + `mask_usage_ratio` — **not** `cycle_day`, and **not** `mean_confidence` / `detection_count`. Other presets: `shape_only`, `no_pipeline_meta`, `full_legacy`. `evaluate` reports Ridge feature importance and feature-set ablations.
- **Size distribution:** MAD z-score second filter on `area_cm2` per camera-day (`size_distribution`); tune `z_max` then `aggregate --refilter-size` without re-extract.
- **Early cycle:** `early_cycle.day_max` (default 21) relaxes post-filter gates (`min_overlap_ratio`, `min_box_area_px`, `conf_min`, isolation/frame gates, `z_max`) for days 1–21. Batch `early_cycle` in `cycle7_batch.yaml` relaxes YOLO/counting thresholds on the same days. Use `filter-overlay --cycle-day N` to preview.
- **Detection thresholds:** batch `conf`/`iou` + `early_cycle` in `cycle7_batch.yaml`; weight overrides seg `model_path` + `classes` only.
- **Train ages:** production `train` uses `train_cycle_day_min`..`train_cycle_day_max` (cycle7: **3–35**). `evaluate` late-cycle GO holdout stays ages **29–35** (fit on **3–28**).
- **Excluded days:** configured via `exclude_cycle_days` in YAML.
- **Cycle7 production window:** exclude days 0–2 (22–24 May 2026); training starts 25 May (cycle day 3).
- Cycle7 / cycle6 DOC weight in labels and examples: **32 g** (Age-0 Manual).
