Problem
The always-on prompt surface a builder consumes is ~21,900 served words (measured: 1252-word-baseline.md), dominated by process recipes: porch phase tasks (11,430w across a 10-iteration project), protocol.md inlined into every spawn (3,703w), CLAUDE.md how-to prose, and consult-type preambles. Spec 1252's measurement proved deduplication alone yields only −7%: the surface is not duplicated, it is over-instructed.
The new rules of context engineering for Claude-5-generation models reports >80% of Claude Code's system prompt was deleted with no measurable performance loss, by replacing rules with judgment, deleting worst-case padding, designing interfaces instead of examples, and progressive disclosure. That operation — content deletion on judgment-trust grounds — was explicitly a Non-goal of Spec 1252 and has never been attempted here.
Goal (SUPERSEDED — see Amendment below)
Rewrite the always-on instruction surface for frontier-model consumers: >50% reduction in served always-on words, measured with 1252's committed measurement script (counts served/expanded words; phantom-savings-proof).
Attack order by size: phase-task prompts → protocol.md spawn inlining (keep the state machine, gates, artifact contracts; delete narrative process) → CLAUDE.md residual how-tos (→ progressive-disclosure skills) → consult-type prompts (rubric + verdict contract; delete process prose).
Baked Decisions
- All prompt consumers are frontier models (Claude 5, GPT 5.6, Gemini 3.6 class). No weak-model tier, no fallback scaffolding variant, no tiering mechanism. One form.
- Scar rules are exempt and verbatim — the eight compressed canonicals developed in Spec 1252 Phase 5 (six repo rules + shellper verified-orphan + Tower-restart permission) ship with the rewrite; the registry/enforcement concept from 1252 is rebuilt fit-for-purpose around the post-shrink surface, not before it.
- Validation is A/B, not observational: same issues executed by builders on old vs new prompts, compared on outcomes (gate friction, review rounds, correctness). Spec 1252's M12 established that observational baselines (n=17) can only detect large regressions — insufficient at deletion scale. The A/B design is a first-class spec section.
- Spec must define a rollback story (prompt surfaces are files; reverting is cheap — say so concretely).
Prior art (required reading for the spec phase)
Protocol
SPIR — spec phase must produce: the per-surface cut plan with word targets, the A/B eval design, and the scar-rule carriage plan, before any rewriting.
AMENDMENT — 2026-08-01 (owner redirect, supersedes the Goal above)
The owner has redirected the acceptance model; reviewers should judge against THIS, not the original Goal:
- The goal is NOT a particular size. Acceptance = conformance to the blog post's principles (P1–P7, quoted verbatim in the spec, each restated as a per-file question answerable from a diff). A file that is principle-conformant at MORE words passes.
- Size becomes reporting-only. The measurement instrument (corrected per the spec's M0 — the original committed script measured a dead prompt tree) still runs before/after and reports honestly, including deletion-vs-relocation separation, but no criterion passes or fails on a word count. The original '>50% reduction measured with 1252's committed script' target is superseded on both counts.
- Architect personal inspection is a mandatory gate mechanic: the architect inspects the old-vs-new diff of every changed file (~66 distinct decisions), per-phase, against a manifest that fails the phase if it omits a changed file.
- Unchanged from the charter: frontier-fleet Baked Decisions, scar-rules exemption (resolved against the blog's P7 in the spec: the blog deletes guardrails against bad output; scar rules guard irreversible acts), and mandatory A/B behavioral validation.
Problem
The always-on prompt surface a builder consumes is ~21,900 served words (measured:
1252-word-baseline.md), dominated by process recipes: porch phase tasks (11,430w across a 10-iteration project),protocol.mdinlined into every spawn (3,703w), CLAUDE.md how-to prose, and consult-type preambles. Spec 1252's measurement proved deduplication alone yields only −7%: the surface is not duplicated, it is over-instructed.The new rules of context engineering for Claude-5-generation models reports >80% of Claude Code's system prompt was deleted with no measurable performance loss, by replacing rules with judgment, deleting worst-case padding, designing interfaces instead of examples, and progressive disclosure. That operation — content deletion on judgment-trust grounds — was explicitly a Non-goal of Spec 1252 and has never been attempted here.
Goal (SUPERSEDED — see Amendment below)
Rewrite the always-on instruction surface for frontier-model consumers: >50% reduction in served always-on words, measured with 1252's committed measurement script (counts served/expanded words; phantom-savings-proof).
Attack order by size: phase-task prompts →
protocol.mdspawn inlining (keep the state machine, gates, artifact contracts; delete narrative process) → CLAUDE.md residual how-tos (→ progressive-disclosure skills) → consult-type prompts (rubric + verdict contract; delete process prose).Baked Decisions
Prior art (required reading for the spec phase)
builder/spir-1252.Protocol
SPIR — spec phase must produce: the per-surface cut plan with word targets, the A/B eval design, and the scar-rule carriage plan, before any rewriting.
AMENDMENT — 2026-08-01 (owner redirect, supersedes the Goal above)
The owner has redirected the acceptance model; reviewers should judge against THIS, not the original Goal: