# Agent Instructions Working rules for agents in this repo. A merge of the [Karpathy coding guidelines](https://github.com/multica-ai/andrej-karpathy-skills) and the [Chain of Truth](https://faridsurya-dev.github.io/Vibe-Coding-Research/en/welcome) method: **validated artifacts are the source of truth, AI is a generator and accelerator.** Please Response in extremely concise and precise. ## 1. Think Before Coding **Don't assume. Don't hide confusion. Surface tradeoffs.** - State your assumptions explicitly. If uncertain, ask. - If multiple interpretations exist, present them — don't pick silently. - If a simpler approach exists, say so. Push back when warranted. - If something is unclear, stop. Name what's confusing. Ask. ## 2. Simplicity First **Minimum code that solves the problem. Nothing speculative.** - No features beyond what was asked. - No abstractions for single-use code. - No "flexibility" or "configurability" that wasn't requested. - No error handling for impossible scenarios. - If you write 200 lines and it could be 50, rewrite it. ## 3. Surgical Changes **Touch only what you must. Clean up only your own mess.** - Don't "improve" adjacent code, comments, or formatting. - Don't refactor things that aren't broken. - Match existing style, even if you'd do it differently. - If you notice unrelated dead code, mention it — don't delete it. - Remove imports/variables/functions that *your* changes made unused. The test: every changed line should trace directly to the user's request. ## 4. Goal-Driven Execution **Define success criteria. Loop until verified.** Turn tasks into verifiable goals, and for multi-step work state a brief plan: ``` 1. [Step] → verify: [check] 2. [Step] → verify: [check] ``` This repo has no automated tests, so verification means **running something**: hit the endpoint, run the job, look at the files it produced. ## 5. Chain of Truth — documents first, then code `docs/` is the source of truth, not the chat prompt. | Document | Contents | | ---------------------- | ----------------------------------------------------------------------------------- | | `docs/requirements.md` | Numbered `REQ-xxx` requirements. Changes only with the user's approval. | | `docs/design.md` | Data schema, API contract, disk layout. Each section names the `REQ-xxx` it serves. | | `docs/tasks.md` | Implementation steps + verification criteria, status `[TODO]`/`[DONE]`. | The rules: - Before writing feature code, make sure a `REQ-xxx` covers it. If none does, propose adding one to the user first. - Once a task is finished **and verified**, flip its status in `docs/tasks.md` to `[DONE]` in the same commit. - If the implementation diverges from `docs/design.md`, update the design — never let a document lie. - Use relative paths in markdown (`./`, `../`), not absolute ones. ## 6. Repo rules - **Package manager: `uv`.** No `pip`, `poetry`, or bare `python`/`python3`. Dependencies live in `requirements.txt`; install them with `uv pip install -r requirements.txt` and run scripts with `uv run`. The Docker image installs the same file, so the two environments cannot drift. - **File size limit: 400 lines.** Any new or refactored file that exceeds it must be split into smaller, logical modules. - `**sam3/` is a vendor copy** of Meta's library. It's a dependency, not app code — don't add scripts there or edit anything inside it. - **Never write into the user's video archive.** All output goes under `data/`. - Secrets (`HF_TOKEN`) come from `.env` only; they never belong in code or docs. ## 7. UI/UX Frontend work follows [ui-ux-pro-max](https://github.com/nextlevelbuilder/ui-ux-pro-max-skill): generate the design system first (style, palette, typography), then build against it, then validate before delivering. Consistency across pages beats per-page cleverness. This app is a **dense internal tool**, not a landing page. Its screens are for long review sessions in front of a screen: the video frame and the annotation canvas are the content, everything else is chrome and stays quiet. No decorative gradients, no marketing motion. Pre-delivery checklist — a UI task is not `[DONE]` until all of it passes: - [ ] No emoji as icons (SVG only: Heroicons/Lucide) - [ ] `cursor: pointer` on every clickable element - [ ] Hover states with smooth transitions (150–300 ms) - [ ] Text contrast at least 4.5:1 - [ ] Focus states visible for keyboard navigation - [ ] `prefers-reduced-motion` respected - [ ] Responsive at 375 / 768 / 1024 / 1440 px The review editor also has to survive keyboard-only use — see `docs/design.md`, "Frontend". ## 8. Domain invariants Two things are easy to break without noticing, and breaking either makes the whole system lie: 1. **Stable val split.** Once a frame lands in `val`, it stays in `val` forever. Otherwise he base-vs-new mAP comparison is meaningless. 2. **One `set_image` per image.** `Sam3Processor.set_image()` runs the vision backbone; set_text_prompt()`only re-runs the grounding head against the cached`backbone_out`. n N-prompt job calls` set_image` **once per image** and loops prompts over that same tate. Don't restructure this into set_image-per-prompt. ## 9. Scalability & Portability **Never hardcode something that will change across environments.** - **Hardware Agnosticism:** Do not hardcode hardware requirements (e.g., GPU configurations in `docker-compose.yml`) directly into base configuration files. Instead, use dynamic startup scripts (like `start.sh`) or environment overrides to detect the host's capabilities and inject the appropriate settings automatically. - **Portability:** The app must be fully deployable and scalable on any device (from a CPU-only laptop to a massive multi-GPU rig) without requiring manual code edits to run. - **Dynamic Configuration:** Do not hardcode absolute IP addresses, local network paths, or machine-specific environment variables in code. Rely on relative paths and configuration files to ensure maximum scalability.