Context
Spec 1252's behavioural measurement (M12) is observational: baseline metrics from committed artifacts (CMAP REQUEST_CHANGES rate 51.9% baseline, review rounds, scar-violation mining) compared before/after over a verify window. Its honest ceiling is "no evidence of behavioural harm at this sample size" — it cannot separate prompt effects from drift in task difficulty, model versions, or reviewer behaviour.
A controlled A/B — same task corpus, old vs new prompts, held-out grader — was considered during the spec and deliberately deferred (decision recorded in the spec's Non-goals): it needs a task corpus and grading harness that do not exist, and building them is a larger project than 1252 itself.
What this issue delivers
- A small task corpus representative of builder work (spec drafting, plan phases, bugfix loops)
- A harness that runs the same task under two prompt-surface versions
- A held-out grading rubric (compliance with scar rules, protocol fidelity, output quality)
- A comparison against Spec 1252's committed baselines (
codev/resources/1252-behavior-baseline.md)
Refs #1252
🤖 Generated with Claude Code
Context
Spec 1252's behavioural measurement (M12) is observational: baseline metrics from committed artifacts (CMAP REQUEST_CHANGES rate 51.9% baseline, review rounds, scar-violation mining) compared before/after over a verify window. Its honest ceiling is "no evidence of behavioural harm at this sample size" — it cannot separate prompt effects from drift in task difficulty, model versions, or reviewer behaviour.
A controlled A/B — same task corpus, old vs new prompts, held-out grader — was considered during the spec and deliberately deferred (decision recorded in the spec's Non-goals): it needs a task corpus and grading harness that do not exist, and building them is a larger project than 1252 itself.
What this issue delivers
codev/resources/1252-behavior-baseline.md)Refs #1252
🤖 Generated with Claude Code