Openai / gpt-5.6-luna

none

91.7% 187 / 204 fixtures · 1 run(s)
34,218 input / 3,511 total output / 0 reasoning within output tokens $0.05554525
91.2% 186 / 204 fixtures · 1 run(s)
40,544 input / 4,314 total output / 0 reasoning within output tokens $0.06669650
Loading reliability summary…
Pass Rate Delta
-0.5% Text: 91.7% → JSON: 91.2%
+4
Gained
JSON pass / text fail
−5
Lost
Text pass / JSON fail
182
Unchanged Pass
Both pass
13
Unchanged Fail
Both fail
Fixture Reliability Delta
Fixture Text JSON Delta
Benchmark Deltas
Benchmark Text JSON Delta
worktree_usage 83.3% 91.7% + 8.3%
branch_cleanup 83.3% 75% -8.3%
cherry_pick 75% 66.7% -8.3%
git_grep 91.7% 100% + 8.3%
merge_conflicts 75% 66.7% -8.3%
rebase 83.3% 75% -8.3%
tag_management 91.7% 100% + 8.3%
blame_forensics 91.7% 91.7% + 0%
commit_messages 100% 100% + 0%
commit_squash 91.7% 91.7% + 0%
git_bisect 100% 100% + 0%
git_clean 91.7% 91.7% + 0%
git_log_format 100% 100% + 0%
git_show 100% 100% + 0%
reflog 100% 100% + 0%
stash_recovery 100% 100% + 0%
submodule_usage 100% 100% + 0%
Changed Fixtures (9)