Add SDD skills and skill lock
This commit is contained in:
1 parent
1d1199fa04
commit
32921ea3f3
10 files changed
+858
No files matched your search
@@ -0,0 +1,89 @@
|
||||
---
|
||||
name: backprop
|
||||
description: |
|
||||
Bug → spec protocol. When a bug is found or a test fails, trace the cause,
|
||||
decide whether a new §V invariant would catch recurrence, append to §B.
|
||||
This is the one non-obvious thing SDD does that plan-then-execute doesn't.
|
||||
Triggers on test failure, bug report, post-mortem, or explicit user ask.
|
||||
---
|
||||
|
||||
# backprop — bug → spec
|
||||
|
||||
Plan-then-execute fixes the code & forgets.
|
||||
SDD fixes the code AND edits spec so recurrence is impossible.
|
||||
That edit is backprop.
|
||||
|
||||
## WHEN TO BACKPROP
|
||||
|
||||
- Test failed at `/build` verification.
|
||||
- User reports bug.
|
||||
- Post-mortem after production incident.
|
||||
- `/check` flags VIOLATE with root cause found.
|
||||
|
||||
## SIX STEPS
|
||||
|
||||
### 1. TRACE
|
||||
Read failure output / bug report.
|
||||
Find exact file:line of wrong behavior.
|
||||
Name root cause in one caveman sentence.
|
||||
|
||||
### 2. ANALYZE
|
||||
Ask three questions:
|
||||
- Would a new §V invariant catch this class of bug? (most common: yes)
|
||||
- Is §I wrong — did spec claim shape the code cannot deliver? (sometimes)
|
||||
- Is §T wrong — did we build the wrong thing? (rare but real)
|
||||
|
||||
### 3. PROPOSE
|
||||
Draft the spec change. Never skip §B; §V/§I/§T are case-by-case.
|
||||
|
||||
Template:
|
||||
```
|
||||
§B row: B<next>|<date>|<root cause>|V<N>
|
||||
§V line: V<next>: <testable rule that would have caught it>
|
||||
```
|
||||
|
||||
Example:
|
||||
```
|
||||
§B row: B3|2026-04-20|refund job ran twice on retry|V7
|
||||
§V line: V7: ∀ refund → idempotency key check before charge reversal
|
||||
```
|
||||
|
||||
### 4. GENERATE TEST
|
||||
New invariant without test = lie. Add failing test first.
|
||||
Name test so it cites the invariant: `TestV7_RefundIdempotent`.
|
||||
|
||||
### 5. VERIFY
|
||||
Fix code. Run test. Must pass. Run full suite. Must not regress.
|
||||
|
||||
### 6. LOG
|
||||
Commit spec edit + test + code fix together.
|
||||
Commit msg: `backprop §B.<n> + §V.<N>: <one-line cause>`.
|
||||
|
||||
## WHAT MAKES A GOOD INVARIANT
|
||||
|
||||
- Testable in code (grep-able or assert-able).
|
||||
- Scoped to a behavior, not a file.
|
||||
- Stated positively when possible (`! hold` over `⊥ forbid`).
|
||||
- References §I surface where it applies.
|
||||
|
||||
**Bad**: V8: code should be correct.
|
||||
**Good**: V8: ∀ pg_query ! params interpolated via driver, ⊥ string concat.
|
||||
|
||||
## WHEN NOT TO ADD §V
|
||||
|
||||
- Bug was purely mechanical typo with no class (`i++` vs `i--` in throwaway).
|
||||
- Fix is a one-time migration.
|
||||
- Root cause is external dep (upgrade deps instead, note in §C).
|
||||
|
||||
Still append §B entry — record that this failure mode was considered. Future bug with same smell → §B search shows precedent.
|
||||
|
||||
## OUTPUT SHAPE
|
||||
|
||||
Every backprop run produces:
|
||||
1. §B entry (always).
|
||||
2. §V entry (usually).
|
||||
3. Test file (when §V added).
|
||||
4. Code fix.
|
||||
5. One commit.
|
||||
|
||||
No dashboards. No log files. SPEC.md + git is the full history.
|
||||
@@ -0,0 +1,81 @@
|
||||
---
|
||||
name: build
|
||||
description: |
|
||||
Plan-then-execute implementation against SPEC.md. Native single-thread
|
||||
loop, no sub-agents. On test or build failure, auto-invokes the backprop
|
||||
skill before retrying — a failed verification always considers whether
|
||||
a new §V invariant would prevent recurrence. Triggers when the user asks
|
||||
to build, implement, execute the spec, or tackle a specific §T task
|
||||
(`build §T.3`, `build --next`, `implement next task`, `run the build`).
|
||||
Expects SPEC.md to exist; if not, defers to the spec skill.
|
||||
---
|
||||
|
||||
# build — implement spec
|
||||
|
||||
Single-thread native plan→execute. You are main Claude. No swarm.
|
||||
|
||||
## LOAD
|
||||
|
||||
1. Read `SPEC.md`. If missing → tell user to invoke the spec skill first. Stop.
|
||||
2. Read `FORMAT.md` once if not loaded.
|
||||
3. Read §R if present — external facts the build must honor, ⊥ re-derive or contradict.
|
||||
4. Parse invocation args:
|
||||
- `§T.n` → that task only
|
||||
- `--next` → lowest-numbered row with status `.` or `~`
|
||||
- `--all` or empty → every `.` row in §T order
|
||||
|
||||
High blast radius (shared module, auth, data, money, public §I)? Run `/review` first. Trivial & reversible? Skip planning ceremony, just do step EXECUTE.
|
||||
|
||||
## PLAN
|
||||
|
||||
Native plan mode — you delegate to it, you do not reinvent task breakdown. For chosen task(s):
|
||||
|
||||
1. Cite every §V invariant that applies. Plan must respect all.
|
||||
2. Cite every §I interface touched. Plan must preserve shape.
|
||||
3. List files to create / edit.
|
||||
4. **Verification contract** — name the EXACT test(s) / acceptance criteria that
|
||||
prove each §V touched. Which test, not "add tests". "Do TDD" alone backfires;
|
||||
the spec says *what to check*. Each §V touched → a named test that fails first.
|
||||
5. Name verification command (test, build, lint) — this is the external oracle. Green = done; ⊥ "looks done".
|
||||
|
||||
Show plan. Wait for user OK unless auto mode.
|
||||
|
||||
## EXECUTE
|
||||
|
||||
Per task in order:
|
||||
|
||||
1. Flip §T.n status cell `.` → `~`. Just write to SPEC.md.
|
||||
2. Edit code per plan.
|
||||
3. Run verification command.
|
||||
4. **Pass** → flip `~` → `x`. Next task.
|
||||
5. **Fail** → invoke backprop skill. Do NOT retry blindly.
|
||||
|
||||
## FAIL → BACKPROP
|
||||
|
||||
On test/build failure:
|
||||
|
||||
1. Read failure output.
|
||||
2. Ask: is failure (a) my code bug, (b) spec wrong, or (c) unspecified edge case?
|
||||
3. If (a) → fix code, re-run. No spec change.
|
||||
4. If (b) or (c) → invoke spec skill with `bug: <cause>` first, let it update §V and §B, then resume build against updated spec.
|
||||
|
||||
Rule: never silently fix root-cause without considering backprop. §B is the memory that stops recurrence.
|
||||
|
||||
## WRITE POLICY
|
||||
|
||||
- Only flip §T status. No other SPEC.md edits from build.
|
||||
- Other spec edits → invoke spec skill.
|
||||
- Commit after each §T completes. Message: `T<n>: <goal line>` + §V cites.
|
||||
|
||||
## VERIFICATION
|
||||
|
||||
Task `x` only if:
|
||||
- Verification command (the oracle) exits 0.
|
||||
- Every §V touched has its named test from the verification contract, and it passes.
|
||||
- No §V invariant regressed (run full test suite at end).
|
||||
|
||||
## NON-GOALS
|
||||
|
||||
- No sub-agents. No parallel workers. Main thread only.
|
||||
- No progress dashboards. `cat SPEC.md | grep §T` is the dashboard.
|
||||
- No speculative work beyond chosen task scope.
|
||||
@@ -0,0 +1,118 @@
|
||||
---
|
||||
name: caveman
|
||||
description: |
|
||||
Caveman encoding for SPEC.md and spec-adjacent writes. Loaded by /spec, /build,
|
||||
/check. Cuts tokens ~75% vs prose while staying precise. Triggers on any write
|
||||
to SPEC.md or when user says "caveman", "compress this", "be brief".
|
||||
---
|
||||
|
||||
# caveman — spec encoding
|
||||
|
||||
Applies to SPEC.md writes, spec-referencing prose, backprop entries.
|
||||
Does NOT apply to code, error strings, commit messages, PR descriptions.
|
||||
|
||||
## GRAMMAR
|
||||
|
||||
- Drop articles (a, an, the).
|
||||
- Drop filler (just, really, basically, simply, actually).
|
||||
- Drop aux verbs where fragment works (is, are, was, were, being).
|
||||
- Drop pleasantries.
|
||||
- No hedging (skip "might", "perhaps", "could be worth").
|
||||
- Fragments fine.
|
||||
- Short synonyms: fix > implement, big > extensive, run > execute.
|
||||
|
||||
## SYMBOLS
|
||||
|
||||
Prefer over words:
|
||||
|
||||
```
|
||||
→ leads to / becomes / on <x>
|
||||
∴ therefore / fix
|
||||
∀ for all / every
|
||||
∃ exists / some
|
||||
! must / required
|
||||
? may / optional / unknown
|
||||
⊥ never / forbidden / nil
|
||||
≠ not equal
|
||||
∈ in
|
||||
∉ not in
|
||||
≤ at most
|
||||
≥ at least
|
||||
& and
|
||||
| or
|
||||
§ section reference
|
||||
```
|
||||
|
||||
## PRESERVE VERBATIM
|
||||
|
||||
Never compress:
|
||||
|
||||
- Code blocks, snippets, one-liners with backticks.
|
||||
- Paths: `src/auth/mw.go`.
|
||||
- URLs.
|
||||
- Identifiers: function names, variable names, env vars.
|
||||
- Numbers and versions.
|
||||
- Error message strings.
|
||||
- SQL, regex, JSON, YAML.
|
||||
- Quoted strings.
|
||||
|
||||
## SHAPES
|
||||
|
||||
**Invariant**:
|
||||
```
|
||||
V<n>: <subject> <relation> <condition>
|
||||
V1: ∀ req → auth check before handler
|
||||
V2: token expiry ≤ current_time → reject
|
||||
```
|
||||
|
||||
**Bug row** (pipe table under §B):
|
||||
```
|
||||
id|date|cause|fix
|
||||
B1|2026-04-20|token `<` not `≤`|V2
|
||||
```
|
||||
|
||||
**Task row** (pipe table under §T):
|
||||
```
|
||||
id|status|task|cites
|
||||
T3|x|add auth mw|V1,I.api
|
||||
```
|
||||
Status: `x` done, `~` wip, `.` todo. Escape literal `|` as `\|`.
|
||||
|
||||
**Interface**:
|
||||
```
|
||||
<kind>: <name> → <shape>
|
||||
api: POST /x → 200 {id:string}
|
||||
cmd: `foo bar <arg>` → stdout JSON
|
||||
env: FOO_KEY ! set
|
||||
```
|
||||
|
||||
## EXAMPLES
|
||||
|
||||
**Bad**:
|
||||
> The system should ensure that every incoming request is properly authenticated before being forwarded to its corresponding handler function.
|
||||
|
||||
**Good**:
|
||||
> V1: ∀ req → auth check before handler
|
||||
|
||||
**Bad**:
|
||||
> We discovered that the token expiration check in the middleware was using a strict less-than comparison operator, which meant tokens were being rejected at the exact moment of their expiry.
|
||||
|
||||
**Good**:
|
||||
> B1: token `<` not `≤` → reject @ expiry boundary.
|
||||
|
||||
**Bad**:
|
||||
> The POST endpoint at /x accepts a JSON body and returns a 200 response with an object containing the created id.
|
||||
|
||||
**Good**:
|
||||
> api: POST /x → 200 {id}
|
||||
|
||||
## BOUNDARIES
|
||||
|
||||
- User asks for prose explanation → switch to normal English.
|
||||
- Spec documents for external review (RFC, pitch) → normal English.
|
||||
- Commit message → normal English (git readers expect it).
|
||||
- Diff comment in code → normal English.
|
||||
|
||||
## WHEN UNSURE
|
||||
|
||||
If cutting a word loses a fact, keep it. Caveman is compression, not amputation.
|
||||
@@ -0,0 +1,93 @@
|
||||
---
|
||||
name: check
|
||||
description: |
|
||||
Read-only drift detector. Diffs SPEC.md against current code and reports
|
||||
violations grouped by severity. Writes nothing — suggests remedies via
|
||||
the spec or build skills but never invokes them. Triggers when the user
|
||||
asks to check drift, audit the spec, verify invariants, or ask whether
|
||||
code still matches the spec. Phrasings: "check drift", "audit the spec",
|
||||
"does the code still match §V", "check invariants", "spec vs code".
|
||||
---
|
||||
|
||||
# check — drift report
|
||||
|
||||
Pure diagnostic. Reports violations. Writes nothing. User decides remedy.
|
||||
|
||||
Spec drifting silently from code is the #1 SDD failure mode. check is the
|
||||
detector. Run it after each `/build` and before each ship — drift caught here is
|
||||
a diff; drift caught in prod is a §B.
|
||||
|
||||
## LOAD
|
||||
|
||||
1. Read `SPEC.md`. If missing → "no spec, nothing to check." Stop.
|
||||
2. Parse invocation args:
|
||||
- `§V` → check invariants only (default)
|
||||
- `§I` → check interfaces
|
||||
- `§T` → audit task status vs code
|
||||
- `--all` → all three
|
||||
|
||||
## CHECK §V — invariants
|
||||
|
||||
For each V<n>:
|
||||
|
||||
1. Translate invariant into verifiable claim about code.
|
||||
2. Grep / read relevant files.
|
||||
3. Classify: **HOLD** / **VIOLATE** / **UNVERIFIABLE**.
|
||||
4. Record address + file:line evidence.
|
||||
|
||||
## CHECK §I — interfaces
|
||||
|
||||
For each I item:
|
||||
|
||||
1. Locate implementation.
|
||||
2. Classify:
|
||||
- **MATCH** — shape in code = shape in spec.
|
||||
- **DRIFT** — impl exists, shape differs.
|
||||
- **MISSING** — impl absent.
|
||||
- **EXTRA** — code exposes surface not in §I.
|
||||
|
||||
## CHECK §T — tasks
|
||||
|
||||
For each T<n>:
|
||||
|
||||
1. If `x`: verify claimed work present.
|
||||
2. If `~`: note as in-progress.
|
||||
3. If `.`: note as pending.
|
||||
4. Flag `x` rows with no evidence as **STALE**.
|
||||
|
||||
## REPORT
|
||||
|
||||
Caveman. Grouped by severity.
|
||||
|
||||
```
|
||||
## §V drift
|
||||
V2 VIOLATE: auth/mw.go:47 uses `<` not `≤`. see §B.1.
|
||||
V5 UNVERIFIABLE: no test covers ∀ req path.
|
||||
|
||||
## §I drift
|
||||
I.api DRIFT: POST /x returns `{result}` not `{id}`. route.go:112.
|
||||
I.cmd MISSING: `foo bar` absent from cli/*.go.
|
||||
|
||||
## §T drift
|
||||
T3 STALE: status `x`, no middleware file exists.
|
||||
|
||||
## summary
|
||||
2 violate. 1 missing. 1 stale. 1 unverifiable.
|
||||
next: spec skill with `bug:` or fix code at cited lines.
|
||||
```
|
||||
|
||||
## REMEDY HINTS (not actions)
|
||||
|
||||
End report with one-line hint per class:
|
||||
- VIOLATE / DRIFT → invoke spec skill `bug: <V.n>` or fix code.
|
||||
- MISSING → invoke build skill on `§T.n` if task exists; else spec skill `amend §T`.
|
||||
- STALE → spec skill `amend §T` to uncheck.
|
||||
- EXTRA → spec skill `amend §I` to document, or delete code.
|
||||
|
||||
Never invoke fixes. Report only.
|
||||
|
||||
## NON-GOALS
|
||||
|
||||
- Zero writes. No SPEC.md edits. No code edits.
|
||||
- No sub-agents. Main thread reads.
|
||||
- No scores, no grades. Binary per item: holds or drifts.
|
||||
@@ -0,0 +1,84 @@
|
||||
---
|
||||
name: deepen
|
||||
description: |
|
||||
Optional design-improvement pass for when you have spare usage to drain. Finds
|
||||
the shallowest modules in the code the spec touches, researches a deeper
|
||||
design, and proposes refactors that shrink interfaces and hide decisions —
|
||||
behavior held constant, tests green before and after. Proposes §I/§V/§T edits,
|
||||
never silent rewrites. Triggers when the user says "deepen this", "improve the
|
||||
design", "this module feels shallow", "pull complexity down", "use spare
|
||||
budget on the codebase", or invokes /ck:deepen. Leans on the codebase-design
|
||||
skill's deep-module vocabulary when present.
|
||||
---
|
||||
|
||||
# deepen — make modules deep
|
||||
|
||||
**Behavior is sacred: tests green before AND after. Every change shrinks an interface or hides a decision — deepen, don't churn.**
|
||||
|
||||
A **deep module** hides a lot behind a small interface; a **shallow** one's
|
||||
interface costs as much to use as writing the code yourself. Complexity =
|
||||
dependencies + obscurity, and it compounds. Deepen spends spare usage paying
|
||||
that down *before* it becomes a §B. Run it when the build is green & you have
|
||||
budget to drain — not under deadline.
|
||||
|
||||
## WHEN TO DEEPEN
|
||||
|
||||
- Build is green, tests pass, & you have token budget spare.
|
||||
- A module's interface feels as complex as its implementation (shallow smell).
|
||||
- The same change keeps touching many files (change amplification).
|
||||
- User explicitly asks to improve design quality.
|
||||
|
||||
⊥ run mid-feature or under pressure. Deepen is the deliberate pass, not the reflex.
|
||||
|
||||
## FIVE STEPS
|
||||
|
||||
### 1. PICK THE SHALLOW
|
||||
Scan the modules the spec touches. Rank by shallowness — interface surface vs
|
||||
work done. Pick the **one** worst offender. Tells:
|
||||
- Pass-through method that only forwards to one other (shallow layer).
|
||||
- Caller must set 5 flags right to use it (config leakage).
|
||||
- Same abstraction repeated at two layers (no information hiding).
|
||||
- A `?` or §B that traces back to a confusing interface.
|
||||
|
||||
One module per pass. Deepening is surgical, ⊥ a codebase sweep.
|
||||
|
||||
### 2. DIAGNOSE
|
||||
Name the design defect in caveman, citing file:line:
|
||||
> src/auth/token.go: 6-arg ctor leaks rotation policy to every caller. shallow.
|
||||
Complexity is real only if it shows: change amplification, high cognitive load,
|
||||
or an unknown-unknown (caller must know a hidden fact to call it right).
|
||||
|
||||
### 3. RESEARCH THE DEEPENING
|
||||
What does a deeper version look like? Pull a known pattern (hand to **research**
|
||||
for the external case → §R) or derive from the codebase's own better modules.
|
||||
Moves that deepen:
|
||||
- **Pull complexity down** — hide the hard part inside, give callers the simple path.
|
||||
- **Define errors out of existence** — design the interface so the edge can't occur.
|
||||
- **Information hiding** — one decision, one module; callers don't learn it.
|
||||
- **General-purpose interface** over a pile of special-case methods.
|
||||
|
||||
### 4. PROPOSE
|
||||
Draft the change as spec edits, not a silent rewrite:
|
||||
- New/simpler §I shape for the module.
|
||||
- §V that locks the deepened invariant so a future build can't re-shallow it.
|
||||
- §T refactor row(s), each citing the §V/§I it serves.
|
||||
Hand to **spec** to write. Show the before/after interface so the user sees the shrink.
|
||||
|
||||
### 5. VERIFY BEHAVIOR HELD
|
||||
Refactor ≠ rewrite. Full suite green before you start AND after. A deepening that
|
||||
changes behavior is a feature in disguise — stop, route through `/spec` + `/build`.
|
||||
New interface gets a test proving the old callers still work.
|
||||
|
||||
## WHEN TO STOP
|
||||
|
||||
Done when the chosen module's interface is strictly smaller, its hidden decision
|
||||
no longer leaks, tests are green, and §I/§V record the new shape. One module
|
||||
deepened beats five churned. Budget left → pick the next shallowest, fresh pass.
|
||||
|
||||
## BOUNDARIES
|
||||
|
||||
- ⊥ change behavior. Green before, green after. Pure structure.
|
||||
- ⊥ write SPEC.md. Propose §I/§V/§T; spec writes.
|
||||
- ⊥ deepen more than one module per pass.
|
||||
- ⊥ run under deadline or mid-feature. This is the spare-budget pass.
|
||||
- ⊥ add abstraction for single-use code. A deep module earns its hiding; a speculative one is just more surface.
|
||||
@@ -0,0 +1,82 @@
|
||||
---
|
||||
name: grill
|
||||
description: |
|
||||
Calibrated interrogation of a fuzzy idea before it becomes a spec. Asks one
|
||||
question at a time, recommends an answer, and lands each answer in §G (goal)
|
||||
or §C (constraints) — unknowns parked as `?` items, never guessed. The
|
||||
cheapest place to kill a bad idea is before §T exists. Triggers when the user
|
||||
has a vague idea, says "grill me", "stress-test this", "challenge my plan",
|
||||
"interview me before I spec", or invokes /ck:grill. Defers the actual write to
|
||||
the spec skill.
|
||||
---
|
||||
|
||||
# grill — sharpen idea before spec
|
||||
|
||||
**One question at a time. Every answer lands in a § or gets parked `?`. Never guess a constraint into existence.**
|
||||
|
||||
Plan-then-execute guesses the fuzzy parts & builds the wrong thing.
|
||||
Grill drags the fuzz into §G/§C *before* a single §T row exists.
|
||||
A bad assumption caught here costs one question. Caught in §B it costs a bug.
|
||||
|
||||
## WHEN TO GRILL
|
||||
|
||||
- Idea is one sentence & you can feel the holes.
|
||||
- Multiple readings of the goal exist & you are about to pick one silently.
|
||||
- Before `/spec new` on anything non-trivial.
|
||||
- User asks to be challenged / stress-tested.
|
||||
|
||||
Skip for a typo or a one-line fix. Grill scales to uncertainty, ⊥ to ego.
|
||||
|
||||
## CALIBRATE FIRST
|
||||
|
||||
One opening read, not a quiz:
|
||||
1. How well does user know this domain? (sets question depth)
|
||||
2. How locked is the idea? (exploring vs committed)
|
||||
3. Pressure wanted: light / normal / brutal.
|
||||
|
||||
Match it. Brutal grilling on a half-formed idea just demoralizes. Light
|
||||
grilling on a committed plan misses the load-bearing flaw.
|
||||
|
||||
## QUESTION LADDER
|
||||
|
||||
Climb in order. Each rung, ask **one** question, **recommend** an answer, wait.
|
||||
|
||||
1. **Goal** — what must the code *do*, in one line? (→ §G)
|
||||
2. **Done** — how do we know it works? name the observable. (→ §C / future §V)
|
||||
3. **Boundary** — what is explicitly out of scope? (→ §C)
|
||||
4. **Lock** — what tech/lib/pattern is non-negotiable? what is forbidden? (→ §C)
|
||||
5. **Surface** — what does the outside world touch — cmd, api, file, env? (→ §I)
|
||||
6. **Edge** — the one input that breaks the happy path? (→ future §V)
|
||||
7. **Unknown** — what do we *not* know yet? (→ park as `?` §C bullet)
|
||||
|
||||
Stop climbing the moment the spec would be unambiguous. Do not ask all seven by reflex.
|
||||
|
||||
## ANSWER FORMAT
|
||||
|
||||
Each question carries a recommended answer so the user can grunt "yes" & move:
|
||||
|
||||
> Q: auth — session cookie or JWT?
|
||||
> rec: JWT — stateless, you named horizontal scaling as a §C.
|
||||
> (a) JWT (b) cookie (c) something else?
|
||||
|
||||
## HANDOFF
|
||||
|
||||
When done, emit a compact block — goal line, constraint bullets, surfaced
|
||||
unknowns as `?` — and hand to the **spec** skill to write §G/§C. Grill proposes;
|
||||
spec is the sole mutator. Never write SPEC.md directly.
|
||||
|
||||
## WHEN TO STOP
|
||||
|
||||
Done when ALL hold:
|
||||
- §G is one line, one reading, zero "or maybe".
|
||||
- §C covers every non-negotiable the user stated or implied.
|
||||
- Every blocking unknown is either answered or parked as an explicit `?`.
|
||||
|
||||
Unresolved blocking unknown that needs the outside world → recommend `/research`, not a guess.
|
||||
|
||||
## BOUNDARIES
|
||||
|
||||
- ⊥ make product decisions for the user. Recommend, never decide.
|
||||
- ⊥ write SPEC.md. Hand structured answers to spec.
|
||||
- ⊥ ask in bulk. One question, one recommendation, wait.
|
||||
- ⊥ grill a trivial change. Right-size or skip.
|
||||
@@ -0,0 +1,72 @@
|
||||
---
|
||||
name: research
|
||||
description: |
|
||||
Gather external knowledge the spec needs and distill it into §R — the durable
|
||||
research log — so build grounds in facts instead of hallucinating library
|
||||
behavior. Each finding cites a source; unsourced claims are flagged, never
|
||||
written as fact. Triggers when a spec decision hinges on a library/API/best
|
||||
practice the agent is unsure of, when the user says "research this", "what's
|
||||
the best lib for…", "check current best practice", or invokes /ck:research.
|
||||
Defers the §R write to the spec skill.
|
||||
---
|
||||
|
||||
# research — external knowledge → §R
|
||||
|
||||
**Every finding cites a source. No source → flag it `?`, never write a guess as fact.**
|
||||
|
||||
"Process without library context gives you well-organized hallucinations."
|
||||
Build invents a plausible-but-wrong API & §B fills with avoidable bugs.
|
||||
Research is the external oracle: pull the real fact once, log it caveman, never re-derive.
|
||||
|
||||
## WHEN TO RESEARCH
|
||||
|
||||
- A §C/§I/§V decision hinges on a lib, API, version, or pattern you are unsure of.
|
||||
- You are about to assume how an external dependency behaves.
|
||||
- The idea touches a domain with real prior art (auth, payments, crypto, rate-limit).
|
||||
- `/grill` parked a `?` that the outside world must answer.
|
||||
|
||||
Skip when the build touches only code you already wrote. Research scales to the unknown, ⊥ to habit.
|
||||
|
||||
## FOUR STEPS
|
||||
|
||||
### 1. SCOPE
|
||||
Turn the unknown into 1-3 concrete questions. Vague "research auth" → "JWT lib
|
||||
for Node ESM, maintained?" + "refresh-token rotation: current best practice?".
|
||||
A scoped question gets a citable answer; a vague one gets an essay.
|
||||
|
||||
### 2. GATHER
|
||||
Use web search / docs tools. Prefer primary sources: official docs, the repo,
|
||||
the RFC, the paper. Two independent sources beat one confident blog. For a big
|
||||
sweep, spawn a sub-agent so the raw pages never touch this context — it returns
|
||||
only the distilled finding + source.
|
||||
|
||||
### 3. DISTILL
|
||||
Crush each answer to one caveman line + its source. Drop the prose. The §R row
|
||||
is the memory; the tab you read is not.
|
||||
|
||||
> R3|refresh token|rotate on use, revoke family on reuse-detect|datatracker.ietf.org/doc/html/rfc6819#section-5.2.2.3
|
||||
|
||||
### 4. HAND OFF
|
||||
Emit the §R rows & hand to the **spec** skill to append. If a finding changes a
|
||||
constraint or interface, note the §C/§I edit for spec too. Research proposes;
|
||||
spec writes.
|
||||
|
||||
## SOURCE DISCIPLINE
|
||||
|
||||
- Cite a URL, repo, RFC, or paper per row. Verbatim identifiers/versions.
|
||||
- Could not verify → write the row but flag `?` in the finding & say so. An
|
||||
unverified claim labeled honestly is fine; one disguised as fact is a future §B.
|
||||
- Conflicting sources → log both, let the user pick. ⊥ silently average them.
|
||||
|
||||
## WHEN TO STOP
|
||||
|
||||
Done when every scoped question has a sourced §R row (or an honest `?`), and no
|
||||
build decision still rests on an unchecked assumption. ⊥ research past the
|
||||
questions you scoped — that is just burning the attention budget.
|
||||
|
||||
## BOUNDARIES
|
||||
|
||||
- ⊥ write SPEC.md. Hand §R rows to spec.
|
||||
- ⊥ write a finding as fact without a source.
|
||||
- ⊥ dump raw pages into context or §R. Distill or it does not land.
|
||||
- ⊥ research what you can read in the repo. Local truth > web guess.
|
||||
@@ -0,0 +1,85 @@
|
||||
---
|
||||
name: review
|
||||
description: |
|
||||
Adversarial senior review of the spec before any code is written. Constructs a
|
||||
skeptical reviewer whose authority comes from the codebase, §R research, and
|
||||
live best-practice — then tries to REFUTE the spec, not rubber-stamp it. Every
|
||||
finding cites evidence (file:line or source); unverifiable ones are flagged.
|
||||
Survivors harden §V; the run ends in an explicit go / no-go gate. Triggers
|
||||
before building anything high-blast-radius, when the user says "review the
|
||||
spec", "red-team this", "is this plan sound", "senior review", or invokes
|
||||
/ck:review.
|
||||
---
|
||||
|
||||
# review — refute the spec before build
|
||||
|
||||
**Every finding cites evidence — file:line or a source. No evidence → flag `[unverified]`. Default to refuted: a flaw you cannot prove is a flaw you note, not one you wave through.**
|
||||
|
||||
An LLM cannot self-correct on its own judgment — left alone it drifts or
|
||||
degrades. Review fixes that the only way that works: a *separate* skeptic
|
||||
anchored to an *external oracle* — the code, §R, the test suite, the docs.
|
||||
"Looks good" is not a review. A refutation attempt is.
|
||||
|
||||
## WHEN TO REVIEW
|
||||
|
||||
- Before `/build` on a high-blast-radius change (shared module, auth, data, money, public API).
|
||||
- Spec touched §I or §V that other code depends on.
|
||||
- Right-sizing says the cost of a wrong build > the cost of one review pass.
|
||||
|
||||
Skip for a trivial, reversible, well-understood change. Adversarial review on a
|
||||
typo hallucinates flaws & wastes the budget — the self-critique paradox is real.
|
||||
|
||||
## PHASE 0 — CAPTURE
|
||||
|
||||
Read the spec: §G §C §I §R §V §T. Hold the whole thing. You review the *spec*,
|
||||
not your memory of the conversation.
|
||||
|
||||
## PHASE 1 — CONSTRUCT THE SENIOR
|
||||
|
||||
Build a reviewer with real authority, not a generic critic:
|
||||
- **Codebase** — grep/read the modules this spec touches. What patterns, what invariants already hold?
|
||||
- **§R** — what did research establish? A spec decision that contradicts §R is a finding.
|
||||
- **Live** — for any best-practice claim you are unsure of, fetch it. An out-of-date assumption is a flaw.
|
||||
|
||||
A reviewer with no evidence is just an opinion. Earn the authority first.
|
||||
|
||||
## PHASE 2 — REFUTE
|
||||
|
||||
Attack the spec on these axes. For each, try to find the case where it breaks:
|
||||
- **Goal vs reality** — does §G solve the actual problem, or a proxy?
|
||||
- **Missing invariant** — what can go wrong that no §V catches? (most findings live here)
|
||||
- **Interface drift** — does §I match what callers already expect? (cite the caller, file:line)
|
||||
- **Constraint conflict** — do two §C bullets contradict? does one fight §R?
|
||||
- **Unowned edge** — the input, ordering, failure, or concurrency case no §T covers.
|
||||
- **Altitude** — §T too vague to act on, or so granular it is just typing?
|
||||
|
||||
## PHASE 3 — CLASSIFY
|
||||
|
||||
Each finding: `evidence → claim → severity`.
|
||||
- **BLOCK** — build on this spec ships a real defect. Must fix first.
|
||||
- **HARDEN** — add/sharpen a §V so the build cannot regress it.
|
||||
- **NOTE** — worth knowing, not blocking.
|
||||
|
||||
No evidence? Down-rank to NOTE & tag `[unverified]`. ⊥ inflate a hunch to BLOCK.
|
||||
|
||||
## PHASE 4 — HARDEN §V & GATE
|
||||
|
||||
- Each HARDEN finding → a draft §V line (testable, cites the §I/behavior it guards). Hand to **spec** to write.
|
||||
- End on an explicit gate:
|
||||
|
||||
```
|
||||
## review verdict
|
||||
BLOCK: 1 — §I.api shape ≠ caller src/client.ts:40. fix §I before build.
|
||||
HARDEN: 2 — drafted V8 (idempotent refund), V9 (tx around dual write).
|
||||
NOTE: 1 — §T4 vague, split before /build.
|
||||
gate: NO-GO until BLOCK cleared. then /build §T after spec writes V8,V9.
|
||||
```
|
||||
|
||||
GO or NO-GO, never a shrug. Review is the checkpoint that stops a confident wrong build.
|
||||
|
||||
## BOUNDARIES
|
||||
|
||||
- ⊥ write SPEC.md. Draft §V & hand to spec.
|
||||
- ⊥ pass a finding with no evidence as fact. Flag `[unverified]`.
|
||||
- ⊥ review trivia. Right-size or skip.
|
||||
- ⊥ rewrite the user's intent. You harden the spec, you do not replace its goal.
|
||||
@@ -0,0 +1,95 @@
|
||||
---
|
||||
name: spec
|
||||
description: |
|
||||
Create, amend, or backprop bugs into SPEC.md at repo root. Sole mutator
|
||||
of the project spec. Triggers when the user asks to write a spec, start
|
||||
a new spec, distill a spec from existing code, add invariants, amend
|
||||
sections (§G, §C, §I, §V, §T, §B), or record a bug via backprop.
|
||||
Common phrasings: "write the spec for...", "new spec", "bug: ...",
|
||||
"amend §V.3", "distill spec from code", "spec this idea". Reads and
|
||||
follows FORMAT.md for the caveman encoding rules and pipe-table shape
|
||||
of §T and §B.
|
||||
---
|
||||
|
||||
# spec — spec mutator
|
||||
|
||||
Read `FORMAT.md` at repo root if not already loaded. Caveman skill applies to all writes here.
|
||||
|
||||
## DISPATCH
|
||||
|
||||
Inspect user request and project state:
|
||||
|
||||
1. No `SPEC.md` at repo root AND args describe idea → **NEW**
|
||||
2. No `SPEC.md` AND `from-code` in args → **DISTILL**
|
||||
3. `SPEC.md` exists AND args start `bug:` → **BACKPROP**
|
||||
4. `SPEC.md` exists AND args start `amend` → **AMEND**
|
||||
5. `SPEC.md` exists, no args → ask user which mode
|
||||
|
||||
## INPUTS — spec is the sole mutator
|
||||
|
||||
The other verbs produce material; spec writes it. Ingest their handoff blocks
|
||||
into the right section, show a diff, write on OK:
|
||||
|
||||
- **grill** → sharpened §G + §C
|
||||
- **research** → §R rows (add the §R section if absent)
|
||||
- **review** → drafted §V lines + the risk verdict
|
||||
- **deepen** → §I/§V/§T amendments
|
||||
|
||||
⊥ rewrite a section the handoff did not name. Sectioned ownership (see FORMAT.md).
|
||||
|
||||
## NEW — idea → spec
|
||||
|
||||
Input: user idea. If it arrived fuzzy, prefer running **grill** first.
|
||||
|
||||
Steps:
|
||||
1. Extract goal (1 line, caveman). → §G.
|
||||
2. List constraints user stated or implied. → §C.
|
||||
3. List external surfaces user named. → §I.
|
||||
4. §R only if **research** ran — else omit the section (right-size).
|
||||
5. Propose initial invariants. → §V (numbered V1…).
|
||||
6. Break goal into ordered tasks. → §T pipe table, all status `.`, ids T1…
|
||||
7. §B section with header row only (`id|date|cause|fix`).
|
||||
|
||||
Write to `SPEC.md`. Show user full file. Ask: "spec OK? `/review` if high-blast-radius, else `/build`."
|
||||
|
||||
## DISTILL — code → spec
|
||||
|
||||
Walk repo. Produce §G (infer from README/package.json/main entry), §C (infer from stack), §I (enumerate public APIs/CLIs/configs), §V (derive from tests and assertions), §T (one task per known TODO or missing test), §B (empty).
|
||||
|
||||
Caveman everywhere. Flag uncertain items with `?` in text so user can confirm.
|
||||
|
||||
## BACKPROP — bug → §B + §V
|
||||
|
||||
Input: `bug: <description>`.
|
||||
|
||||
Steps:
|
||||
1. Parse bug description.
|
||||
2. Find root cause (read relevant code).
|
||||
3. Decide: would a new invariant catch recurrence? If yes → draft `V<next>`.
|
||||
4. Append §B row: `B<next>|<date>|<cause>|V<N>`.
|
||||
5. Append new invariant to §V.
|
||||
6. If fix also changes behavior → add/update §T rows.
|
||||
7. Show diff. Apply only on user OK.
|
||||
|
||||
Rule: every bug gets a §B entry. Invariant optional but preferred.
|
||||
|
||||
## AMEND — targeted edit
|
||||
|
||||
Input: `amend §V.3` or `amend §T` etc.
|
||||
|
||||
Read that section. Show current. Ask user what changes. Write. Show diff.
|
||||
|
||||
Never silently rewrite sections user did not name.
|
||||
|
||||
## OUTPUT RULES
|
||||
|
||||
- Caveman format per `FORMAT.md`.
|
||||
- Preserve identifiers, paths, code verbatim.
|
||||
- Numbering monotonic — never reuse §V.N or §B.N.
|
||||
- §T row `cites` column ! list §V/§I deps: `T5|.|impl auth mw|V2,I.api`.
|
||||
|
||||
## NON-GOALS
|
||||
|
||||
- No sub-agents. Main thread writes.
|
||||
- No dashboards, no logs, no state files beyond SPEC.md itself.
|
||||
- No auto-build after spec. User invokes build explicitly.
|
||||
@@ -0,0 +1,59 @@
|
||||
{
|
||||
"version": 1,
|
||||
"skills": {
|
||||
"backprop": {
|
||||
"source": "JuliusBrussee/cavekit",
|
||||
"sourceType": "github",
|
||||
"skillPath": "skills/backprop/SKILL.md",
|
||||
"computedHash": "7d23e415bf412ff01fa8c3d920c269f6ac7a9ca0c337eafda861d412ba8f946c"
|
||||
},
|
||||
"build": {
|
||||
"source": "JuliusBrussee/cavekit",
|
||||
"sourceType": "github",
|
||||
"skillPath": "skills/build/SKILL.md",
|
||||
"computedHash": "6cf3f7fc11d4c9892bd9b71002715ab9b5915aabef15d2da30d955d117f1052d"
|
||||
},
|
||||
"caveman": {
|
||||
"source": "JuliusBrussee/cavekit",
|
||||
"sourceType": "github",
|
||||
"skillPath": "skills/caveman/SKILL.md",
|
||||
"computedHash": "b7296730e7d074253e6806b9e71dab4a6e1cbfeb643cafc000f9ce7bec0a825c"
|
||||
},
|
||||
"check": {
|
||||
"source": "JuliusBrussee/cavekit",
|
||||
"sourceType": "github",
|
||||
"skillPath": "skills/check/SKILL.md",
|
||||
"computedHash": "28e8c45a027c7c1c9832417100aff4d78b9981d45d3672c0797eb606e020c2d9"
|
||||
},
|
||||
"deepen": {
|
||||
"source": "JuliusBrussee/cavekit",
|
||||
"sourceType": "github",
|
||||
"skillPath": "skills/deepen/SKILL.md",
|
||||
"computedHash": "5e0fe183377f9b12811ce6a05173d72a8f9194055a519d4a0b26000f0c776392"
|
||||
},
|
||||
"grill": {
|
||||
"source": "JuliusBrussee/cavekit",
|
||||
"sourceType": "github",
|
||||
"skillPath": "skills/grill/SKILL.md",
|
||||
"computedHash": "2bad4607d6fb2b0538d6f3e57f151064d219133bc4cd08a941c0850a0967d6ca"
|
||||
},
|
||||
"research": {
|
||||
"source": "JuliusBrussee/cavekit",
|
||||
"sourceType": "github",
|
||||
"skillPath": "skills/research/SKILL.md",
|
||||
"computedHash": "91566d23ee13a46fe8d1094c9e5b25dc2a622e4b5ea2ecc97179a58f4f901b8d"
|
||||
},
|
||||
"review": {
|
||||
"source": "JuliusBrussee/cavekit",
|
||||
"sourceType": "github",
|
||||
"skillPath": "skills/review/SKILL.md",
|
||||
"computedHash": "eb74a6bf82fff552294e7c3ce0a281d6e668f6b61ce93b1f8ffa2a10bf99a7d3"
|
||||
},
|
||||
"spec": {
|
||||
"source": "JuliusBrussee/cavekit",
|
||||
"sourceType": "github",
|
||||
"skillPath": "skills/spec/SKILL.md",
|
||||
"computedHash": "80ec98e14d6e5bb009d2f36d57cfc073f3cd136cf83ddf2176efb39a581ef072"
|
||||
}
|
||||
}
|
||||
}
|
||||
Reference in new issue
Block a user