Skip to content

Prompt-surface judgment rewrite: acceptance = blog-principles conformance per file (size is reporting-only) #1280

Description

@waleedkadous

Problem

The always-on prompt surface a builder consumes is ~21,900 served words (measured: 1252-word-baseline.md), dominated by process recipes: porch phase tasks (11,430w across a 10-iteration project), protocol.md inlined into every spawn (3,703w), CLAUDE.md how-to prose, and consult-type preambles. Spec 1252's measurement proved deduplication alone yields only −7%: the surface is not duplicated, it is over-instructed.

The new rules of context engineering for Claude-5-generation models reports >80% of Claude Code's system prompt was deleted with no measurable performance loss, by replacing rules with judgment, deleting worst-case padding, designing interfaces instead of examples, and progressive disclosure. That operation — content deletion on judgment-trust grounds — was explicitly a Non-goal of Spec 1252 and has never been attempted here.

Goal (SUPERSEDED — see Amendment below)

Rewrite the always-on instruction surface for frontier-model consumers: >50% reduction in served always-on words, measured with 1252's committed measurement script (counts served/expanded words; phantom-savings-proof).

Attack order by size: phase-task prompts → protocol.md spawn inlining (keep the state machine, gates, artifact contracts; delete narrative process) → CLAUDE.md residual how-tos (→ progressive-disclosure skills) → consult-type prompts (rubric + verdict contract; delete process prose).

Baked Decisions

  • All prompt consumers are frontier models (Claude 5, GPT 5.6, Gemini 3.6 class). No weak-model tier, no fallback scaffolding variant, no tiering mechanism. One form.
  • Scar rules are exempt and verbatim — the eight compressed canonicals developed in Spec 1252 Phase 5 (six repo rules + shellper verified-orphan + Tower-restart permission) ship with the rewrite; the registry/enforcement concept from 1252 is rebuilt fit-for-purpose around the post-shrink surface, not before it.
  • Validation is A/B, not observational: same issues executed by builders on old vs new prompts, compared on outcomes (gate friction, review rounds, correctness). Spec 1252's M12 established that observational baselines (n=17) can only detect large regressions — insufficient at deletion scale. The A/B design is a first-class spec section.
  • Spec must define a rollback story (prompt surfaces are files; reverting is cheap — say so concretely).

Prior art (required reading for the spec phase)

Protocol

SPIR — spec phase must produce: the per-surface cut plan with word targets, the A/B eval design, and the scar-rule carriage plan, before any rewriting.

AMENDMENT — 2026-08-01 (owner redirect, supersedes the Goal above)

The owner has redirected the acceptance model; reviewers should judge against THIS, not the original Goal:

  1. The goal is NOT a particular size. Acceptance = conformance to the blog post's principles (P1–P7, quoted verbatim in the spec, each restated as a per-file question answerable from a diff). A file that is principle-conformant at MORE words passes.
  2. Size becomes reporting-only. The measurement instrument (corrected per the spec's M0 — the original committed script measured a dead prompt tree) still runs before/after and reports honestly, including deletion-vs-relocation separation, but no criterion passes or fails on a word count. The original '>50% reduction measured with 1252's committed script' target is superseded on both counts.
  3. Architect personal inspection is a mandatory gate mechanic: the architect inspects the old-vs-new diff of every changed file (~66 distinct decisions), per-phase, against a manifest that fails the phase if it omits a changed file.
  4. Unchanged from the charter: frontier-fleet Baked Decisions, scar-rules exemption (resolved against the blog's P7 in the spec: the blog deletes guardrails against bad output; scar rules guard irreversible acts), and mandatory A/B behavioral validation.

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/cross-cuttingTouches multiple areas — needs coordinated handling

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions