Skip to content

Controlled A/B eval of prompt-surface changes (deferred from Spec 1252 / M12c) #1277

Description

@waleedkadous

Context

Spec 1252's behavioural measurement (M12) is observational: baseline metrics from committed artifacts (CMAP REQUEST_CHANGES rate 51.9% baseline, review rounds, scar-violation mining) compared before/after over a verify window. Its honest ceiling is "no evidence of behavioural harm at this sample size" — it cannot separate prompt effects from drift in task difficulty, model versions, or reviewer behaviour.

A controlled A/B — same task corpus, old vs new prompts, held-out grader — was considered during the spec and deliberately deferred (decision recorded in the spec's Non-goals): it needs a task corpus and grading harness that do not exist, and building them is a larger project than 1252 itself.

What this issue delivers

  • A small task corpus representative of builder work (spec drafting, plan phases, bugfix loops)
  • A harness that runs the same task under two prompt-surface versions
  • A held-out grading rubric (compliance with scar rules, protocol fidelity, output quality)
  • A comparison against Spec 1252's committed baselines (codev/resources/1252-behavior-baseline.md)

Refs #1252

🤖 Generated with Claude Code

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/protocolsArea: Protocol definitions — distinct from area/porch (orchestration)

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions