ci+local: upgrade AI PR reviewer model gpt-5.5 -> gpt-5.6-sol (effort xhigh)#701
Conversation
… xhigh) Validated via the reviewer-eval 4-arm blinded campaign (66 runs, 11 cases, 3 independent blind graders, pre-registered DECISION_RULE.md gates): sol@xhigh 7/10 reliable catches vs 5.5@xhigh 2/10, zero regressions, three strict improvements incl. an S4 missed-bug probe; sol@max no decisive gain at 1.5-2x latency (effort stays xhigh); terra ~= 5.5. Pins: ai_pr_review.yml model, DEFAULT_MODEL, _is_reasoning_model prefix tuple, PRICING (+sol/terra/luna), ai-review-local.md refs, configs.json control arm flip. Also corrects the s3-wcr-pfloor fixture comment overclaim the winning arms caught (restores it as a genuinely clean negative control). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019xnguzh8asX9dY54djvR41
Overall Assessment✅ Looks good — no unmitigated P0 or P1 findings. Executive Summary
Methodology
Code QualityNo findings. Control selection remains fail-closed and explicit in PerformanceNo findings. MaintainabilityNo findings. Workflow, local defaults, documentation, and evaluator control metadata are synchronized. Tech DebtNo findings. No correctness limitation requiring SecurityNo findings. The changed configuration does not expose credentials or broaden sandbox permissions. Documentation/Tests
|
…n verdict (CI review R1 P3s) Pins _is_reasoning_model/_resolve_timeout/PRICING/DEFAULT_MODEL for the gpt-5.6 family in tests, and commits the sanitized 2026-07 campaign verdict (gate results + artifact hashes; raw model outputs stay local) so the CHANGELOG's eval claims are auditable from the repository. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019xnguzh8asX9dY54djvR41
|
🔁 AI review rerun (requested by @igerber) Head SHA: Overall Assessment✅ Looks good — no unmitigated P0 or P1 findings. Executive Summary
Methodology
Code QualityNo findings. Bare PerformanceNo findings. MaintainabilityNo findings. Tech DebtNo findings. SecurityNo findings. The workflow retains its read-only sandbox and existing safety strategy. Documentation/Tests
The prior P3 gaps are otherwise resolved by |
… R2 P3) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019xnguzh8asX9dY54djvR41
|
🔁 AI review rerun (requested by @igerber) Head SHA: Overall Assessment✅ Looks good — no unmitigated P0 or P1 findings. Executive Summary
Methodology
Code QualityNo findings. Control resolution fails closed and delegates to existing configuration validation. PerformanceNo findings. MaintainabilityNo findings. CI, local defaults, reasoning classification, and configuration metadata are aligned. Tech DebtNo findings. SecurityNo findings. The workflow retains its read-only sandbox and fork-safety controls. Documentation/TestsNo findings. The previous stale verdict-path issue is fixed at
|
Summary
gpt-5.6-solatxhigheffort on BOTH surfaces simultaneously: the CI codex-action (ai_pr_review.ymlmodel:) and the local review script (openai_review.pyDEFAULT_MODEL), which share the same_build_codex_cmdinvocation - keeping local review predictive of CI review.tools/reviewer-eval/4-arm blinded campaign (feat(reviewer-eval): N-arm matrix, blinded grading, corpus 2->11 for the GPT-5.6 evaluation #700 infrastructure): 66 runs over 11 cases, three independent identity-blind graders, pre-registered gates inDECISION_RULE.md. Results: gpt-5.6-sol@xhigh reliably caught 7/10 ground-truth bugs vs gpt-5.5@xhigh's 2/10 (+2 unstable), with zero catch regressions and three strict improvements (including an S4 missed-bug probe gpt-5.5 historically passed over). gpt-5.6-sol@max showed no decisive gain over xhigh at ~1.5-2x latency, so effort stays xhigh; gpt-5.6-terra tracked ~gpt-5.5._is_reasoning_modelprefix tuple +=gpt-5.6,PRICING+= gpt-5.6-sol (5/30), -terra (2.50/15), -luna (1/6) per 1M tokens,ai-review-local.mddoc refs, andconfigs.jsoncontrol-arm flip (gpt-5.6-sol@xhigh becomesrole=control; gpt-5.5 retained as the previous-production arm for future back-comparisons).smokenow derives its default arm from the solerole=controlentry at run time (a control flip can never drift the default; regression-tested), and thes3-wcr-pfloor-commentcorpus fixture's comment overclaim - which the winning arms correctly flagged during the campaign - is corrected to purpose-phrasing, restoring it as a genuinely clean negative control.gpt-5.6-sol- the review below this description is the new model operating in the production action environment.Methodology references (required if estimator / math changes)
diff_diff/-adjacent content is inside a frozen reviewer-eval corpus fixture (comment-only wording fix ins3-wcr-pfloor-comment/inject.diff); the live tree'sdiff_diff/utils.pyis untouched.Validation
tests/test_evals_runtime.py+1 (test_smoke_default_selects_control_arm). Full touched suites green: 372 tests acrosstest_evals_{runtime,engine,adapters}.py+test_openai_review.py;verify-corpus11/11 (regenerated fixture applies cleanly at its pinned base).tools/reviewer-eval/runs/(gitignored); verdict summarized in the CHANGELOG entry.maxeffort acceptance, and codex-actionmodel:/effort:passthrough verified live against codex-cli 0.144.5 on 2026-07-18.Security / privacy
Generated with Claude Code
https://claude.ai/code/session_019xnguzh8asX9dY54djvR41