Z
Z-ai / glm-5.2

xhigh

95.6% 195 / 204 fixtures · 1 run(s)
36,712 input / 122,893 total output / 110,390 reasoning within output tokens $0.53893895
78.9% 161 / 204 fixtures · 1 run(s)
36,846 input / 88,371 total output / 82,972 reasoning within output tokens $0.39458491
Loading reliability summary…
Pass Rate Delta
-16.7% Text: 95.6% → JSON: 78.9%
+0
Gained
JSON pass / text fail
−34
Lost
Text pass / JSON fail
161
Unchanged Pass
Both pass
9
Unchanged Fail
Both fail
Fixture Reliability Delta
Fixture Text JSON Delta
Benchmark Deltas
Benchmark Text JSON Delta
reflog 100% 50% -50%
commit_squash 91.7% 50% -41.7%
branch_cleanup 100% 66.7% -33.3%
worktree_usage 83.3% 50% -33.3%
git_grep 100% 75% -25%
commit_messages 100% 83.3% -16.7%
rebase 91.7% 75% -16.7%
tag_management 100% 83.3% -16.7%
blame_forensics 91.7% 83.3% -8.3%
git_show 91.7% 83.3% -8.3%
git_bisect 100% 91.7% -8.3%
git_clean 100% 91.7% -8.3%
git_log_format 100% 91.7% -8.3%
stash_recovery 100% 91.7% -8.3%
cherry_pick 83.3% 83.3% + 0%
merge_conflicts 91.7% 91.7% + 0%
submodule_usage 100% 100% + 0%
Changed Fixtures (34)