# Agent Instructions This is the authoritative rules file for any AI coding agent (Claude Code, Cursor, GitHub Copilot, Aider, etc.) working in a project that uses this starter kit. Agents that don't read AGENTS.md natively should be pointed at it via their own config file — see `docs/vibe-coding/` for per-tool instructions. `CLAUDE.md` imports this file. ## Scope: excludes `backend/` This kit's automated behaviors — the `e`/`enhance` and `n`/`next` triggers, the §3 file-size threshold as an *enforcement action* (splitting a file because this kit says so), and the §5 Demo/Live and §6 Cloud/Local switches — govern the **Flutter app at the repo root** (`lib/`, `test/`, root-level config/docs) only. **`backend/` is out of scope for this kit.** It is a separately governed subtree with its own pre-existing rules files — `backend/CLAUDE.md` (project-specific architecture/history/confidentiality notes) and `backend/AGENTS.md` (carried over near-unchanged from `ai-ocr-pfm-2026`: uv conventions, issue-recording workflow, karpathy coding guidelines) — that predate this kit's adoption and are tuned to that codebase's own conventions. Do not let this kit's generic rules override or be conflated with them: - Don't run `e`/`enhance` against `backend/` modules, and don't let `n`/`next` auto-seed or pick up backend-section tasks, unless the user explicitly asks for a backend-scoped run for that specific invocation. - Don't apply §3's 256-line split as a repo-wide mandate inside `backend/` — that subtree already documents its own pre-existing large files as accepted debt (see `backend/CLAUDE.md`) and has its own change conventions. - §5 (Demo/Live) and §6 (Cloud/Local) describe app-level UX switches for the Flutter client; they are not a mandate to add mock-data layers or deployment switches inside the Next.js backend. - If asked to work inside `backend/`, defer entirely to `backend/CLAUDE.md` and `backend/AGENTS.md` (and `backend/plans/next-enhancements.md` for feature-level status — `backend/next-implementation.md` was deleted 2026-07-08, its content folded into that plan) instead of this file. See **Adaptation Notes** at the end of this file for how `plans/next-enhancements.md` reconciles this after an earlier adoption pass mistakenly seeded backend sections. ## 0. Adopting Into an Existing Project This kit auto-detects whether it's landing in a new or existing codebase — no separate trigger is required. You can also type "i" / "init" at any time to force a re-audit (e.g. after a large refactor or reorg). **Auto-detection**: the first time "e"/"enhance" or "n"/"next" runs, if `plans/next-enhancements.md` is still the empty seed template (or missing) AND the repo already contains files/history beyond this kit's own files, treat it as an existing project and run the audit below first. Skip it on a genuinely empty/new project — there's nothing to audit. **The audit** (owned by the Software Architect role, see `SKILLS.md`): - **Don't overwrite silently.** If the repo already has its own README.md/AGENTS.md/ CLAUDE.md, merge additively — append a clearly-marked section pointing at this kit's files rather than deleting existing project-specific instructions. - **Discover reality before writing anything**: - Detect the tech stack and real build/test/lint commands (package.json, Makefile, pyproject.toml, etc.). - Map the actual existing folder/module structure — this becomes the section list `e`/`enhance` uses, replacing invented sections. - Scan for files already over the 256-line threshold (§3) and list them as **pre-existing debt**, not a blocker — the rule binds new/touched files going forward, it does not retroactively force-refactor untouched legacy code. - Check for existing equivalents of Demo/Live (§5) or Cloud/Local (§6) patterns; adapt to what's already there instead of introducing a redundant second switch. - Seed `plans/next-enhancements.md` sections from the real discovered modules (not blank placeholders); optionally backfill `docs/feature-list.md` with a short "Existing Features (pre-kit)" summary so the log doesn't start misleadingly empty. - Record any deviation from this kit's defaults (because the project already does it differently) as a short **Adaptation Notes** subsection appended to the end of this file — so future `e`/`n` runs see local reality, not just the template's assumptions. - Once the audit (and any merge/adaptation) is done, `e`/`enhance` and `n`/`next` behave exactly as described below. ## 1. Trigger "e" or "enhance" If the user types "e", "enhance", or requests an enhancement plan: - Read `/plans/next-enhancements.md` to understand the current platform structure, history, and active tasks. - Overwrite or update the active tasks list inside `/plans/next-enhancements.md`. - The plan must cover each main section/module of the application. - Inside the tasks list, define **exactly 3 new enhancements per section** with: 1. A unique number (e.g., `1.1`, `1.2`, `1.3`). 2. A clear, specific description of the functional change. 3. A status (initially set to `[TODO]`). - Present this plan to the user in your final summary response. ## 2. Trigger "n", "next", or "n{x}" If the user types "n", "next", "n{x}" (where `{x}` is a positive integer), or requests execution of the next enhancement task(s): - Read `/plans/next-enhancements.md` to check the status of tasks. - If all tasks are `[DONE]` (or none are `[TODO]`), automatically run the **"e" / "enhance"** workflow first. - Otherwise, select the most impactful `[TODO]` task(s) by strategic value, functional impact, or UX contribution — not just the first in order. If `{x}` is given, select the top `{x}` tasks and execute them sequentially. ### 2a. Clarify before building ("Grill Me" step) Before implementing a task, check whether its scope or acceptance criteria are genuinely ambiguous (multiple valid interpretations, unspecified UI/data behavior, no clear "done" condition). If so: - Ask clarifying questions **one at a time** until the task is unambiguous — use your tool's native mechanism (e.g. Claude Code's `AskUserQuestion`) where available, otherwise ask inline and wait for the answer before continuing. - Record the resolved acceptance criteria as a short 1-3 line note next to the task entry in `/plans/next-enhancements.md` before writing any code. - Skip this step entirely when the task is already unambiguous — don't grill the user on obvious work. ### 2b. TDD Workflow (Test First) - **Write Tests First**: Before implementing the actual feature code for a task, write automated tests defining the expected behavior. - **Iterate Until Green**: Run the tests to confirm they fail, then write the implementation until all tests pass perfectly. - **Browser Testing**: If the enhancement involves web UI or visual components, use browser tools (e.g., Chrome) to test the app visually and functionally if necessary. - Implement the selected task(s) fully in the codebase, applying the relevant role(s) from `SKILLS.md` (Architect, Backend, Frontend, QA, Hardware/Compatibility). - Once complete: 1. Update the task's status in `/plans/next-enhancements.md` to `[DONE]`. 2. Document the new/updated feature in `/docs/feature-list.md` under the right section heading. 3. **Create an Iteration Log**: Perform a code review and audit of the tasks just completed. Document this audit in `/docs/iteration-log.md` (or append to it) to ensure all functions work perfectly. - **Verify build integrity** — this is not just "it compiles and runs": - QA pass: exercise the golden path and edge cases; run/extend automated tests. - Hardware/Compatibility pass: check cross-platform, cross-browser, and resource (memory/CPU) assumptions per `SKILLS.md`. - In your final response, state which task(s) were completed and the exact menu/navigation path to see the new feature. ## 3. File Size, Refactoring & SOLID Rules - **SOLID Principles**: Always design, implement, and refactor code adhering to SOLID programming principles (Single Responsibility, Open/Closed, Liskov Substitution, Interface Segregation, Dependency Inversion). This ensures code is modular, testable, and maintainable. - **256-line threshold**: any code/script file — new, modified, or pre-existing — that exceeds 256 lines of code must be split into smaller, modular, logical files. This is a repo-wide rule, not just for new work; if you touch a file over the threshold, split it as part of that change. - **Why**: an agent reading a file spends its context budget parsing it before it can reason about the task. Small, single-purpose files keep that cost low and keep the agent's understanding accurate (see the "smart zone / dumb zone" problem — models reason worse as context fills up). - This applies to AGENTS.md, CLAUDE.md, and SKILLS.md too: keep each under ~250 lines. Push detail into linked docs (`docs/`) rather than growing the root files. ## 4. Roles Every implementation task should be viewed through the lens of the relevant role(s) defined in `SKILLS.md`: Software Architect, Backend Engineer, Frontend Engineer, QA/Test Engineer, and Hardware & Performance Compatibility Reviewer. A single agent plays all roles in sequence unless the harness supports spawning role-specific subagents (see `CLAUDE.md`). ## 5. Mockup Data & Demo/Live Mode - Store all mock/sample data in `/data/mockup/` — keep it separate from UI components and styling. - Provide a mock API layer that reads from `/data/mockup/` and mirrors the real backend contract. - Add a switcher control (icon) in the UI to toggle **Demo** (mock API + mock data) and **Live** (real API + real data). ## 6. Cloud vs Local (On-Premise) - Provide a setting to choose **Cloud** or **Local (on-premise)** deployment. - **Cloud**: use remote/cloud-hosted API endpoints and services. - **Local**: use on-premise/self-hosted API endpoints and services. - Persist the selection and route all backend/service calls to the chosen environment. ## 7. Ad-hoc Feature Requests For direct feature requests not using "e"/"n", implement the feature and document it in `/docs/feature-list.md`. ## Adaptation Notes (this project) Recorded 2026-07-08, from adopting this kit into the existing `app-pfm-ocr-v2` repo. - **Tech stack**: Flutter mobile app at the repo root (`lib/`) is the field client; `backend/` runs a Next.js API gateway + Python PaddleOCR/vLLM OCR pipeline + Postgres. See root `CLAUDE.md` and `backend/CLAUDE.md` for full architecture. - **Real commands**: `flutter test`, `flutter analyze lib`, `flutter build apk --release` (Flutter, repo root); `npm run dev`/`build`/`lint` inside `backend/pfm-web-app` (Next.js) — see `backend/CLAUDE.md` for the OCR accuracy regression harness and the uv/vLLM Python service commands. There is no single unified test runner across both halves; Flutter and backend are tested independently. - **Real module boundaries** (seeded into `plans/next-enhancements.md`): Flutter — Auth & Splash, Camera Capture & Geotagging, Pending Documents Queue, Document Editor & PDF Receipt; Backend — Next.js API Gateway, OCR Pipeline & Accuracy, Postgres Data Layer, DevOps/Docker & Dev Tunnel. - **Scope correction (2026-07-08, later same day)**: an early `e`/`enhance` run seeded and executed against the four Backend sections above (5-8) before the **Scope: excludes `backend/`** rule at the top of this file was written. Those sections were first frozen as archival, then removed outright from `plans/next-enhancements.md` (2026-07-08) once `backend/` got its own independent kit copy — the already-shipped 7.1 DB-transaction fix stays recorded in `docs/feature-list.md`, but the task itself no longer has an entry in this file's plan. `plans/next-enhancements.md` now tracks only the Flutter sections (1-4). Backend enhancement tracking lives entirely in `backend/`'s own copy of this kit (`backend/plans/next-enhancements.md`, driven by `backend/AGENTS.md` Part B). - **Pre-existing files over the 256-line threshold** (§3 debt, not a blocker — split only if/when touched going forward), **Flutter only** now that `backend/` is out of scope for this kit (see §Scope above; backend's own debt list, re-verified 2026-07-08, lives in `backend/AGENTS.md`'s Adaptation Notes instead): `lib/features/camera/image_preview_screen.dart` (380), `lib/features/camera/camera_screen.dart` (376). (`editor_screen.dart` was split 2026-07-10 when task 6.3 touched it — see root `plans/next-enhancements.md` §6.3 and `docs/iteration-log.md`; `documents_screen.dart` and `pending_documents_provider.dart` were already split down by earlier iterations and are no longer over threshold — this line was stale.) - **No Demo/Live or Cloud/Local switch exists yet** (§5, §6). The closest existing analogue is `AppConfig.initializeApiBaseUrl()` in `lib/config/app_config.dart`, which dynamically resolves a *real* backend endpoint (ngrok tunnel, falling back to a LAN IP) — it always talks to the live backend with real data, never mock data, so it is not a Demo/Live switch. Building an actual mock-data Demo mode and a Cloud/Local deployment switch remain open `n`/`next` work, not something this adoption pass implemented. - **Naming collision to watch for**: `docker-compose.demo.yml` at the repo root already uses "demo" for an unrelated concept — a **production-mode** Compose override (`npm start` instead of `npm run dev`), not a mock-data mode. Don't conflate it with §5's Demo/Live data switch when that feature is eventually built. - **Pre-existing, unrelated governance files left as-is** (not part of this kit, not touched by this adoption): - `.agents/AGENTS.md` — a short set of OCR post-processing/parsing rules (table column-shift correction, date normalization). Different purpose and path from this kit's root `AGENTS.md`; both coexist. - `plans/next-enhancement-plan.md` (singular) — a pre-existing, fully `[DONE]` QA verification checklist auto-managed by "QE guidelines," unrelated to the `e`/`n` workflow's `plans/next-enhancements.md` (plural, added by this kit). Kept side by side, not merged or migrated. - **CLAUDE.md**: the repo already had a substantial project-specific root `CLAUDE.md`. Per this section's "merge additively" rule, this kit's import and Claude-specific notes were appended as a new, clearly-marked section rather than replacing the file.