Z
Z-ai / glm-5.2

low

96.1% 196 / 204 fixtures · 1 run(s)
36,651 input / 74,991 total output / 66,992 reasoning within output tokens $0.35013099
89.2% 182 / 204 fixtures · 1 run(s)
40,188 input / 64,077 total output / 57,693 reasoning within output tokens $0.29657239
Loading reliability summary…
Pass Rate Delta
-6.9% Text: 96.1% → JSON: 89.2%
+3
Gained
JSON pass / text fail
−17
Lost
Text pass / JSON fail
179
Unchanged Pass
Both pass
5
Unchanged Fail
Both fail
Fixture Reliability Delta
Fixture Text JSON Delta
f011 0% (0/1) 100% (1/1) -100%
Benchmark Deltas
Benchmark Text JSON Delta
branch_cleanup 100% 75% -25%
reflog 100% 75% -25%
commit_squash 91.7% 75% -16.7%
merge_conflicts 91.7% 75% -16.7%
commit_messages 91.7% 100% + 8.3%
git_bisect 100% 91.7% -8.3%
git_show 100% 91.7% -8.3%
submodule_usage 100% 91.7% -8.3%
tag_management 100% 91.7% -8.3%
worktree_usage 100% 91.7% -8.3%
blame_forensics 91.7% 91.7% + 0%
cherry_pick 83.3% 83.3% + 0%
git_clean 100% 100% + 0%
git_grep 91.7% 91.7% + 0%
git_log_format 100% 100% + 0%
rebase 91.7% 91.7% + 0%
stash_recovery 100% 100% + 0%
Changed Fixtures (20)