Openai / gpt-5.6-luna

high

96.6% 197 / 204 fixtures · 1 run(s)
34,222 input / 16,697 total output / 12,981 reasoning within output tokens $0.13466550
94.1% 192 / 204 fixtures · 1 run(s)
40,518 input / 18,826 total output / 14,268 reasoning within output tokens $0.15374150
Loading reliability summary…
Pass Rate Delta
-2.4% Text: 96.6% → JSON: 94.1%
+2
Gained
JSON pass / text fail
−7
Lost
Text pass / JSON fail
190
Unchanged Pass
Both pass
5
Unchanged Fail
Both fail
Fixture Reliability Delta
Fixture Text JSON Delta
Benchmark Deltas
Benchmark Text JSON Delta
branch_cleanup 100% 83.3% -16.7%
cherry_pick 91.7% 75% -16.7%
merge_conflicts 91.7% 83.3% -8.3%
rebase 91.7% 83.3% -8.3%
git_show 91.7% 100% + 8.3%
blame_forensics 100% 100% + 0%
commit_messages 100% 100% + 0%
commit_squash 100% 100% + 0%
git_bisect 100% 100% + 0%
git_clean 100% 100% + 0%
git_grep 100% 100% + 0%
git_log_format 100% 100% + 0%
reflog 100% 100% + 0%
stash_recovery 100% 100% + 0%
submodule_usage 91.7% 91.7% + 0%
tag_management 100% 100% + 0%
worktree_usage 83.3% 83.3% + 0%
Changed Fixtures (9)