Openai / gpt-5.6-luna

low

93.1% 190 / 204 fixtures · 1 run(s)
34,191 input / 10,390 total output / 6,633 reasoning within output tokens $0.09679375
92.2% 188 / 204 fixtures · 1 run(s)
40,432 input / 9,288 total output / 4,787 reasoning within output tokens $0.09642275
Loading reliability summary…
Pass Rate Delta
-1% Text: 93.1% → JSON: 92.2%
+5
Gained
JSON pass / text fail
−7
Lost
Text pass / JSON fail
183
Unchanged Pass
Both pass
9
Unchanged Fail
Both fail
Fixture Reliability Delta
Fixture Text JSON Delta
Benchmark Deltas
Benchmark Text JSON Delta
merge_conflicts 91.7% 83.3% -8.3%
submodule_usage 83.3% 91.7% + 8.3%
git_clean 100% 91.7% -8.3%
worktree_usage 83.3% 75% -8.3%
blame_forensics 100% 100% + 0%
branch_cleanup 91.7% 91.7% + 0%
cherry_pick 75% 75% + 0%
commit_messages 100% 100% + 0%
commit_squash 100% 100% + 0%
git_bisect 100% 100% + 0%
git_grep 100% 100% + 0%
git_log_format 100% 100% + 0%
git_show 91.7% 91.7% + 0%
rebase 75% 75% + 0%
reflog 100% 100% + 0%
stash_recovery 100% 100% + 0%
tag_management 91.7% 91.7% + 0%
Changed Fixtures (12)