Openai / gpt-oss-120b

xhigh

94.6% 193 / 204 fixtures · 1 run(s)
46,146 input / 100,957 total output / 93,793 reasoning within output tokens $0.03184461
90.2% 184 / 204 fixtures · 1 run(s)
46,317 input / 97,512 total output / 94,821 reasoning within output tokens $0.02645551
Loading reliability summary…
Pass Rate Delta
-4.4% Text: 94.6% → JSON: 90.2%
+4
Gained
JSON pass / text fail
−13
Lost
Text pass / JSON fail
180
Unchanged Pass
Both pass
7
Unchanged Fail
Both fail
Fixture Reliability Delta
Fixture Text JSON Delta
Benchmark Deltas
Benchmark Text JSON Delta
reflog 100% 75% -25%
commit_messages 100% 83.3% -16.7%
git_grep 100% 83.3% -16.7%
rebase 75% 91.7% + 16.7%
tag_management 100% 83.3% -16.7%
commit_squash 91.7% 83.3% -8.3%
cherry_pick 83.3% 75% -8.3%
blame_forensics 100% 100% + 0%
branch_cleanup 100% 100% + 0%
git_bisect 100% 100% + 0%
git_clean 91.7% 91.7% + 0%
git_log_format 100% 100% + 0%
git_show 100% 100% + 0%
merge_conflicts 83.3% 83.3% + 0%
stash_recovery 100% 100% + 0%
submodule_usage 100% 100% + 0%
worktree_usage 83.3% 83.3% + 0%
Changed Fixtures (17)