Search with no results returns empty output
Tests ability to handle searches with no results (empty output). Evaluates edge-case handling in search operations.

These commands set up the repo before the model sees the prompt. They define the starting file structure, staged changes, and Git history.

  1. 01 git init
  2. 02 git config user.email 'test@test.com'
  3. 03 git config user.name 'Test User'
  4. 04 echo 'def hello(): return "world"' > hello.py
  5. 05 git add hello.py
  6. 06 git commit -m 'Add hello function'
  7. 07 echo 'git grep nonexistent_token_xyz' > .grep_command
  8. 08 git add .grep_command
  9. 09 git commit -m 'Add grep sentinel'
Prompt
Here is the output of a git grep command run on this repository. Did the search find any matches? Output ONLY 'yes' or 'no', nothing else.
Expected
no

Scoped model quality, cost, API time, and token usage for git_grep/f006.

Loading...
Loading raw attempt evidence…
anthropic/claude-fable-5:high PASS 100% 53 in → 26 out (13 reasoning)
no
anthropic/claude-fable-5:high__json_schema PASS 100% 270 in → 21 out (9 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
anthropic/claude-fable-5:low PASS 100% 53 in → 14 out (10 reasoning)
no
anthropic/claude-fable-5:low__json_schema PASS 100% 270 in → 10 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
anthropic/claude-fable-5:max PASS 100% 53 in → 102 out (33 reasoning)
no
anthropic/claude-fable-5:max__json_schema PASS 100% 270 in → 75 out (25 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
anthropic/claude-fable-5:medium PASS 100% 53 in → 14 out (10 reasoning)
no
anthropic/claude-fable-5:medium__json_schema PASS 100% 270 in → 21 out (10 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
anthropic/claude-fable-5:xhigh PASS 100% 53 in → 21 out (9 reasoning)
no
anthropic/claude-fable-5:xhigh__json_schema PASS 100% 270 in → 37 out (17 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
anthropic/claude-haiku-4.5:high PASS 100% 72 in → 254 out (271 reasoning)
no
anthropic/claude-haiku-4.5:high__json_schema PASS 100% 240 in → 154 out (158 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
anthropic/claude-haiku-4.5:low PASS 100% 72 in → 174 out (193 reasoning)
no
anthropic/claude-haiku-4.5:low__json_schema PASS 100% 240 in → 260 out (274 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
anthropic/claude-haiku-4.5:medium PASS 100% 72 in → 269 out (294 reasoning)
no
anthropic/claude-haiku-4.5:medium__json_schema PASS 100% 240 in → 286 out (307 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-haiku-4.5:none PASS 100% 43 in → 4 out (0 reasoning)
no
anthropic/claude-haiku-4.5:none__json_schema PASS 100% 210 in → 9 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
anthropic/claude-haiku-4.5:xhigh PASS 100% 72 in → 233 out (245 reasoning)
no
anthropic/claude-haiku-4.5:xhigh__json_schema PASS 100% 240 in → 262 out (285 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
anthropic/claude-opus-4.6:high__json_schema PASS 100% 211 in → 73 out (59 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-opus-4.6:low PASS 100% 43 in → 5 out (0 reasoning)
no
anthropic/claude-opus-4.6:low__json_schema PASS 100% 211 in → 8 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-opus-4.6:max PASS 100% 43 in → 112 out (106 reasoning)
no
anthropic/claude-opus-4.6:max__json_schema PASS 100% 211 in → 125 out (121 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-opus-4.6:medium__json_schema PASS 100% 211 in → 8 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-opus-4.6:none PASS 100% 43 in → 4 out (0 reasoning)
no
anthropic/claude-opus-4.6:none__json_schema PASS 100% 211 in → 8 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-opus-4.6:xhigh PASS 100% 43 in → 74 out (64 reasoning)
no
anthropic/claude-opus-4.6:xhigh__json_schema PASS 100% 211 in → 65 out (51 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-opus-4.7:high PASS 100% 58 in → 5 out (0 reasoning)
no
anthropic/claude-opus-4.7:high__json_schema PASS 100% 275 in → 12 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-opus-4.7:low PASS 100% 58 in → 5 out (0 reasoning)
no
anthropic/claude-opus-4.7:low__json_schema PASS 100% 275 in → 12 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-opus-4.7:max PASS 100% 58 in → 5 out (0 reasoning)
no
anthropic/claude-opus-4.7:max__json_schema PASS 100% 275 in → 90 out (23 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-opus-4.7:medium PASS 100% 58 in → 5 out (0 reasoning)
no
anthropic/claude-opus-4.7:medium__json_schema PASS 100% 275 in → 12 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-opus-4.7:none PASS 100% 58 in → 5 out (0 reasoning)
no
anthropic/claude-opus-4.7:none__json_schema PASS 100% 275 in → 12 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-opus-4.7:xhigh PASS 100% 58 in → 5 out (0 reasoning)
no
anthropic/claude-opus-4.7:xhigh__json_schema PASS 100% 275 in → 12 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-opus-4.8:high PASS 100% 53 in → 3 out (0 reasoning)
no
anthropic/claude-opus-4.8:high__json_schema PASS 100% 270 in → 10 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-opus-4.8:low PASS 100% 53 in → 3 out (0 reasoning)
no
anthropic/claude-opus-4.8:low__json_schema PASS 100% 270 in → 10 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-opus-4.8:max PASS 100% 53 in → 3 out (0 reasoning)
no
anthropic/claude-opus-4.8:max__json_schema PASS 100% 270 in → 92 out (22 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
anthropic/claude-opus-4.8:medium PASS 100% 53 in → 3 out (0 reasoning)
no
anthropic/claude-opus-4.8:medium__json_schema PASS 100% 270 in → 10 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-opus-4.8:none PASS 100% 53 in → 3 out (0 reasoning)
no
anthropic/claude-opus-4.8:none__json_schema PASS 100% 270 in → 9 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-opus-4.8:xhigh PASS 100% 53 in → 3 out (0 reasoning)
no
anthropic/claude-opus-4.8:xhigh__json_schema PASS 100% 270 in → 10 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-sonnet-4.6:high PASS 100% 43 in → 106 out (103 reasoning)
no
anthropic/claude-sonnet-4.6:high__json_schema PASS 100% 211 in → 121 out (112 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-sonnet-4.6:low PASS 100% 43 in → 4 out (0 reasoning)
no
anthropic/claude-sonnet-4.6:low__json_schema PASS 100% 211 in → 8 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-sonnet-4.6:max PASS 100% 43 in → 131 out (122 reasoning)
no
anthropic/claude-sonnet-4.6:max__json_schema PASS 100% 211 in → 131 out (124 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-sonnet-4.6:medium PASS 100% 43 in → 133 out (129 reasoning)
no
anthropic/claude-sonnet-4.6:medium__json_schema PASS 100% 211 in → 63 out (48 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-sonnet-4.6:none__json_schema PASS 100% 211 in → 8 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-sonnet-4.6:xhigh PASS 100% 43 in → 157 out (150 reasoning)
no
anthropic/claude-sonnet-4.6:xhigh__json_schema PASS 100% 211 in → 134 out (126 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-sonnet-5:high PASS 100% 53 in → 3 out (0 reasoning)
no
anthropic/claude-sonnet-5:high__json_schema PASS 100% 270 in → 10 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-sonnet-5:low__json_schema PASS 100% 270 in → 10 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-sonnet-5:medium__json_schema PASS 100% 270 in → 10 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
anthropic/claude-sonnet-5:none PASS 100% 53 in → 3 out (0 reasoning)
no
anthropic/claude-sonnet-5:none__json_schema PASS 100% 270 in → 9 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
anthropic/claude-sonnet-5:xhigh PASS 100% 53 in → 3 out (0 reasoning)
no
anthropic/claude-sonnet-5:xhigh__json_schema PASS 100% 270 in → 10 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
arcee-ai/trinity-large-thinking:high PASS 100% 43 in → 1,499 out (1,497 reasoning)
no
arcee-ai/trinity-large-thinking:low PASS 100% 43 in → 980 out (978 reasoning)
no
arcee-ai/trinity-large-thinking:medium PASS 100% 43 in → 1,687 out (1,688 reasoning)
no
arcee-ai/trinity-large-thinking:xhigh PASS 100% 43 in → 1,286 out (1,282 reasoning)
no
arcee-ai/trinity-mini:high PASS 100% 43 in → 392 out (445 reasoning)
no
arcee-ai/trinity-mini:low PASS 100% 43 in → 263 out (276 reasoning)
no
arcee-ai/trinity-mini:low__json_schema PASS 100% 43 in → 242 out (250 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
deepseek/deepseek-v4-flash-0731:high PASS 100% 116 in → 135 out (133 reasoning)
no
deepseek/deepseek-v4-flash-0731:high__json_schema PASS 100% 37 in → 530 out (556 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
deepseek/deepseek-v4-flash-0731:low PASS 100% 37 in → 402 out (399 reasoning)
no
deepseek/deepseek-v4-flash-0731:low__json_schema PASS 100% 39 in → 517 out (507 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
deepseek/deepseek-v4-flash-0731:max PASS 100% 116 in → 193 out (207 reasoning)
no
deepseek/deepseek-v4-flash-0731:max__json_schema PASS 100% 116 in → 255 out (269 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
deepseek/deepseek-v4-flash-0731:medium PASS 100% 37 in → 376 out (406 reasoning)
no
deepseek/deepseek-v4-flash-0731:medium__json_schema PASS 100% 37 in → 367 out (386 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
deepseek/deepseek-v4-flash-0731:none__json_schema PASS 100% 37 in → 7 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
deepseek/deepseek-v4-flash-0731:xhigh PASS 100% 37 in → 248 out (270 reasoning)
no
deepseek/deepseek-v4-flash-0731:xhigh__json_schema PASS 100% 37 in → 239 out (249 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
deepseek/deepseek-v4-flash:high PASS 100% 37 in → 283 out (280 reasoning)
no
deepseek/deepseek-v4-flash:high__json_schema PASS 100% 37 in → 293 out (300 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
deepseek/deepseek-v4-flash:low PASS 100% 37 in → 213 out (210 reasoning)
no
deepseek/deepseek-v4-flash:low__json_schema PASS 100% 133 in → 7 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
deepseek/deepseek-v4-flash:medium PASS 100% 37 in → 203 out (222 reasoning)
no
deepseek/deepseek-v4-flash:none__json_schema PASS 100% 37 in → 9 out (0 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
deepseek/deepseek-v4-flash:xhigh PASS 100% 37 in → 335 out (371 reasoning)
no
deepseek/deepseek-v4-pro:high PASS 100% 37 in → 279 out (276 reasoning)
no
deepseek/deepseek-v4-pro:low PASS 100% 37 in → 133 out (130 reasoning)
no
deepseek/deepseek-v4-pro:medium PASS 100% 37 in → 141 out (138 reasoning)
no
deepseek/deepseek-v4-pro:none__json_schema PASS 100% 39 in → 12 out (0 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
deepseek/deepseek-v4-pro:xhigh PASS 100% 116 in → 281 out (279 reasoning)
no
deepseek/deepseek-v4-pro:xhigh__json_schema PASS 100% 307 in → 524 out (516 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
google/gemini-3-flash-preview:high PASS 100% 33 in → 598 out (597 reasoning)
no
google/gemini-3-flash-preview:high__json_schema PASS 100% 34 in → 340 out (335 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemini-3-flash-preview:low PASS 100% 33 in → 599 out (598 reasoning)
no
google/gemini-3-flash-preview:low__json_schema PASS 100% 94 in → 437 out (432 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemini-3-flash-preview:medium PASS 100% 33 in → 603 out (602 reasoning)
no
google/gemini-3-flash-preview:medium__json_schema PASS 100% 94 in → 503 out (498 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemini-3-flash-preview:xhigh PASS 100% 33 in → 760 out (759 reasoning)
no
google/gemini-3-flash-preview:xhigh__json_schema PASS 100% 34 in → 686 out (681 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemini-3.1-flash-lite-preview:high PASS 100% 34 in → 499 out (498 reasoning)
no
google/gemini-3.1-flash-lite-preview:high__json_schema PASS 100% 94 in → 543 out (537 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
google/gemini-3.1-flash-lite-preview:low__json_schema PASS 100% 34 in → 111 out (105 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
google/gemini-3.1-flash-lite-preview:medium PASS 100% 34 in → 353 out (352 reasoning)
no
google/gemini-3.1-flash-lite-preview:medium__json_schema PASS 100% 34 in → 363 out (358 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemini-3.1-flash-lite-preview:xhigh PASS 100% 34 in → 736 out (735 reasoning)
no
google/gemini-3.1-flash-lite-preview:xhigh__json_schema PASS 100% 34 in → 325 out (320 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemini-3.1-pro-preview:high PASS 100% 33 in → 236 out (235 reasoning)
no
google/gemini-3.1-pro-preview:high__json_schema PASS 100% 94 in → 267 out (262 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemini-3.1-pro-preview:low PASS 100% 33 in → 209 out (208 reasoning)
no
google/gemini-3.1-pro-preview:low__json_schema PASS 100% 94 in → 201 out (196 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemini-3.1-pro-preview:medium PASS 100% 33 in → 305 out (304 reasoning)
no
google/gemini-3.1-pro-preview:medium__json_schema PASS 100% 94 in → 386 out (381 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemini-3.1-pro-preview:xhigh PASS 100% 33 in → 278 out (277 reasoning)
no
google/gemini-3.1-pro-preview:xhigh__json_schema PASS 100% 94 in → 387 out (382 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemini-3.5-flash-lite:high PASS 100% 34 in → 359 out (358 reasoning)
no
google/gemini-3.5-flash-lite:high__json_schema PASS 100% 34 in → 731 out (726 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemini-3.5-flash-lite:low__json_schema PASS 100% 34 in → 11 out (0 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
google/gemini-3.5-flash-lite:max PASS 100% 34 in → 325 out (324 reasoning)
no
google/gemini-3.5-flash-lite:max__json_schema PASS 100% 34 in → 501 out (496 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemini-3.5-flash-lite:medium PASS 100% 33 in → 356 out (355 reasoning)
no
google/gemini-3.5-flash-lite:medium__json_schema PASS 100% 34 in → 457 out (451 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
google/gemini-3.5-flash-lite:xhigh PASS 100% 33 in → 363 out (362 reasoning)
no
google/gemini-3.5-flash-lite:xhigh__json_schema PASS 100% 94 in → 463 out (458 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemini-3.5-flash:high PASS 100% 33 in → 588 out (587 reasoning)
no
google/gemini-3.5-flash:high__json_schema PASS 100% 94 in → 334 out (329 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemini-3.5-flash:low PASS 100% 33 in → 203 out (202 reasoning)
no
google/gemini-3.5-flash:low__json_schema PASS 100% 94 in → 112 out (107 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemini-3.5-flash:medium PASS 100% 33 in → 237 out (236 reasoning)
no
google/gemini-3.5-flash:medium__json_schema PASS 100% 94 in → 396 out (391 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemini-3.5-flash:xhigh PASS 100% 33 in → 254 out (253 reasoning)
no
google/gemini-3.5-flash:xhigh__json_schema PASS 100% 94 in → 420 out (415 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemini-3.6-flash:high PASS 100% 34 in → 402 out (401 reasoning)
no
google/gemini-3.6-flash:high__json_schema PASS 100% 34 in → 465 out (460 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemini-3.6-flash:low PASS 100% 34 in → 237 out (236 reasoning)
no
google/gemini-3.6-flash:low__json_schema PASS 100% 94 in → 163 out (158 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemini-3.6-flash:max PASS 100% 33 in → 457 out (456 reasoning)
no
google/gemini-3.6-flash:max__json_schema PASS 100% 94 in → 503 out (498 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemini-3.6-flash:medium PASS 100% 34 in → 267 out (266 reasoning)
no
google/gemini-3.6-flash:medium__json_schema PASS 100% 34 in → 235 out (230 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemini-3.6-flash:xhigh PASS 100% 34 in → 347 out (346 reasoning)
no
google/gemini-3.6-flash:xhigh__json_schema PASS 100% 34 in → 397 out (392 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemma-4-26b-a4b-it:high PASS 100% 48 in → 772 out (801 reasoning)
no
google/gemma-4-26b-a4b-it:high__json_schema PASS 100% 48 in → 964 out (984 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
google/gemma-4-26b-a4b-it:low PASS 100% 49 in → 1,818 out (1,850 reasoning)
no
google/gemma-4-26b-a4b-it:low__json_schema PASS 100% 49 in → 2,370 out (2,271 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
google/gemma-4-26b-a4b-it:medium PASS 100% 49 in → 1,551 out (1,544 reasoning)
no
google/gemma-4-26b-a4b-it:medium__json_schema PASS 100% 49 in → 653 out (673 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
google/gemma-4-26b-a4b-it:xhigh PASS 100% 48 in → 1,013 out (1,042 reasoning)
no
google/gemma-4-26b-a4b-it:xhigh__json_schema PASS 100% 48 in → 629 out (644 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
google/gemma-4-31b-it:high PASS 100% 49 in → 368 out (383 reasoning)
no
google/gemma-4-31b-it:high__json_schema PASS 100% 48 in → 1,279 out (1,256 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
google/gemma-4-31b-it:low PASS 100% 49 in → 435 out (439 reasoning)
no
google/gemma-4-31b-it:low__json_schema PASS 100% 48 in → 723 out (726 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
google/gemma-4-31b-it:medium PASS 100% 52 in → 428 out (426 reasoning)
no
google/gemma-4-31b-it:medium__json_schema PASS 100% 52 in → 470 out (463 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemma-4-31b-it:none__json_schema PASS 100% 46 in → 7 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
google/gemma-4-31b-it:xhigh PASS 100% 49 in → 678 out (669 reasoning)
no
google/gemma-4-31b-it:xhigh__json_schema PASS 100% 48 in → 376 out (392 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
no
JSON Schema Structured Output
(raw) { "found" : "no" }
no
JSON Schema Structured Output
(raw) { "found": "no" }
minimax/minimax-m2.5:high PASS 100% 73 in → 386 out (457 reasoning)
no
minimax/minimax-m2.5:high__json_schema PASS 100% 71 in → 378 out (368 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
minimax/minimax-m2.5:low PASS 100% 76 in → 283 out (336 reasoning)
no
minimax/minimax-m2.5:low__json_schema PASS 100% 71 in → 327 out (360 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
minimax/minimax-m2.5:medium PASS 100% 71 in → 322 out (358 reasoning)
no
minimax/minimax-m2.5:medium__json_schema PASS 100% 71 in → 508 out (590 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
minimax/minimax-m2.5:xhigh PASS 100% 71 in → 472 out (536 reasoning)
no
minimax/minimax-m2.5:xhigh__json_schema PASS 100% 71 in → 229 out (255 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
minimax/minimax-m2.7:high PASS 100% 71 in → 447 out (443 reasoning)
no
minimax/minimax-m2.7:high__json_schema PASS 100% 70 in → 282 out (310 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
minimax/minimax-m2.7:low PASS 100% 71 in → 669 out (758 reasoning)
no
minimax/minimax-m2.7:low__json_schema PASS 100% 207 in → 534 out (527 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
minimax/minimax-m2.7:medium PASS 100% 74 in → 468 out (466 reasoning)
no
minimax/minimax-m2.7:medium__json_schema PASS 100% 207 in → 402 out (395 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
minimax/minimax-m2.7:xhigh PASS 100% 71 in → 376 out (422 reasoning)
no
minimax/minimax-m3:high PASS 100% 209 in → 63 out (82 reasoning)
no
minimax/minimax-m3:low PASS 100% 209 in → 55 out (66 reasoning)
no
minimax/minimax-m3:low__json_schema PASS 100% 196 in → 9 out (0 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
minimax/minimax-m3:medium PASS 100% 209 in → 111 out (136 reasoning)
no
minimax/minimax-m3:xhigh PASS 100% 209 in → 145 out (170 reasoning)
no
mistralai/devstral-2512 PASS 100% 36 in → 2 out
no
mistralai/devstral-2512__json_schema PASS 100% 36 in → 7 out
no
JSON Schema Structured Output
(raw) {"found": "no"}
mistralai/mistral-medium-3-5:high PASS 100% 48 in → 527 out (574 reasoning)
no
mistralai/mistral-medium-3-5:high__json_schema PASS 100% 36 in → 145 out (161 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
mistralai/mistral-medium-3-5:low PASS 100% 48 in → 278 out (298 reasoning)
no
mistralai/mistral-medium-3-5:low__json_schema PASS 100% 36 in → 506 out (571 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
mistralai/mistral-medium-3-5:medium PASS 100% 48 in → 284 out (307 reasoning)
no
mistralai/mistral-medium-3-5:medium__json_schema PASS 100% 36 in → 888 out (1,012 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
mistralai/mistral-medium-3-5:none PASS 100% 48 in → 2 out (0 reasoning)
no
mistralai/mistral-medium-3-5:none__json_schema PASS 100% 36 in → 7 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
mistralai/mistral-medium-3-5:xhigh PASS 100% 48 in → 248 out (275 reasoning)
no
mistralai/mistral-medium-3-5:xhigh__json_schema PASS 100% 36 in → 425 out (458 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
moonshotai/kimi-k2.6:high PASS 100% 41 in → 178 out (175 reasoning)
no
moonshotai/kimi-k2.6:high__json_schema PASS 100% 40 in → 539 out (441 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
moonshotai/kimi-k2.6:low PASS 100% 41 in → 198 out (195 reasoning)
no
moonshotai/kimi-k2.6:low__json_schema PASS 100% 40 in → 490 out (403 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
moonshotai/kimi-k2.6:medium PASS 100% 41 in → 634 out (631 reasoning)
no
moonshotai/kimi-k2.6:medium__json_schema PASS 100% 40 in → 451 out (312 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
moonshotai/kimi-k2.6:xhigh PASS 100% 41 in → 354 out (411 reasoning)
no
moonshotai/kimi-k2.7-code:high PASS 100% 41 in → 108 out (105 reasoning)
no
moonshotai/kimi-k2.7-code:high__json_schema PASS 100% 41 in → 185 out (174 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
moonshotai/kimi-k2.7-code:low PASS 100% 41 in → 197 out (194 reasoning)
no
moonshotai/kimi-k2.7-code:medium PASS 100% 41 in → 99 out (115 reasoning)
no
moonshotai/kimi-k2.7-code:medium__json_schema PASS 100% 40 in → 259 out (121 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
moonshotai/kimi-k2.7-code:xhigh PASS 100% 41 in → 207 out (238 reasoning)
no
moonshotai/kimi-k2.7-code:xhigh__json_schema PASS 100% 128 in → 134 out (127 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
moonshotai/kimi-k3:high PASS 100% 118 in → 145 out (129 reasoning)
no
moonshotai/kimi-k3:high__json_schema PASS 100% 118 in → 114 out (98 reasoning)
no
moonshotai/kimi-k3:low PASS 100% 118 in → 143 out (127 reasoning)
no
moonshotai/kimi-k3:low__json_schema PASS 100% 118 in → 146 out (130 reasoning)
no
moonshotai/kimi-k3:max PASS 100% 118 in → 134 out (118 reasoning)
no
moonshotai/kimi-k3:max__json_schema PASS 100% 118 in → 229 out (213 reasoning)
no
moonshotai/kimi-k3:medium PASS 100% 118 in → 122 out (106 reasoning)
no
moonshotai/kimi-k3:medium__json_schema PASS 100% 118 in → 182 out (166 reasoning)
no
moonshotai/kimi-k3:xhigh PASS 100% 118 in → 143 out (127 reasoning)
no
moonshotai/kimi-k3:xhigh__json_schema PASS 100% 118 in → 244 out (228 reasoning)
no
nvidia/nemotron-3-nano-30b-a3b:high PASS 100% 49 in → 227 out (242 reasoning)
no
nvidia/nemotron-3-nano-30b-a3b:high__json_schema PASS 100% 49 in → 166 out (168 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
nvidia/nemotron-3-nano-30b-a3b:low PASS 100% 49 in → 378 out (405 reasoning)
no
nvidia/nemotron-3-nano-30b-a3b:low__json_schema PASS 100% 49 in → 210 out (228 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
nvidia/nemotron-3-nano-30b-a3b:medium PASS 100% 49 in → 211 out (231 reasoning)
no
nvidia/nemotron-3-nano-30b-a3b:medium__json_schema PASS 100% 49 in → 1,464 out (1,613 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
nvidia/nemotron-3-nano-30b-a3b:none PASS 100% 49 in → 2 out (0 reasoning)
no
nvidia/nemotron-3-nano-30b-a3b:none__json_schema PASS 100% 49 in → 10 out (0 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
nvidia/nemotron-3-nano-30b-a3b:xhigh PASS 100% 49 in → 959 out (1,105 reasoning)
no
nvidia/nemotron-3-nano-30b-a3b:xhigh__json_schema PASS 100% 49 in → 212 out (222 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
nvidia/nemotron-3-super-120b-a12b:high PASS 100% 49 in → 230 out (225 reasoning)
no
nvidia/nemotron-3-super-120b-a12b:high__json_schema PASS 100% 49 in → 180 out (168 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
nvidia/nemotron-3-super-120b-a12b:low PASS 100% 49 in → 200 out (216 reasoning)
no
nvidia/nemotron-3-super-120b-a12b:low__json_schema PASS 100% 49 in → 175 out (163 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
nvidia/nemotron-3-super-120b-a12b:medium PASS 100% 49 in → 179 out (205 reasoning)
no
nvidia/nemotron-3-super-120b-a12b:medium__json_schema PASS 100% 49 in → 154 out (142 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
nvidia/nemotron-3-super-120b-a12b:none PASS 100% 49 in → 2 out (0 reasoning)
no
nvidia/nemotron-3-super-120b-a12b:none__json_schema PASS 100% 49 in → 10 out (0 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
nvidia/nemotron-3-super-120b-a12b:xhigh PASS 100% 49 in → 185 out (200 reasoning)
no
nvidia/nemotron-3-super-120b-a12b:xhigh__json_schema PASS 100% 49 in → 192 out (180 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
openai/gpt-5.4-mini:high PASS 100% 39 in → 219 out (212 reasoning)
no
openai/gpt-5.4-mini:high__json_schema PASS 100% 79 in → 275 out (260 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.4-mini:low PASS 100% 39 in → 165 out (158 reasoning)
no
openai/gpt-5.4-mini:low__json_schema PASS 100% 79 in → 106 out (91 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.4-mini:medium PASS 100% 39 in → 153 out (146 reasoning)
no
openai/gpt-5.4-mini:medium__json_schema PASS 100% 79 in → 230 out (215 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.4-mini:none PASS 100% 39 in → 11 out (0 reasoning)
no
openai/gpt-5.4-mini:none__json_schema PASS 100% 79 in → 19 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.4-mini:xhigh PASS 100% 39 in → 523 out (516 reasoning)
no
openai/gpt-5.4-mini:xhigh__json_schema PASS 100% 79 in → 531 out (516 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.4-nano:high PASS 100% 39 in → 98 out (91 reasoning)
no
openai/gpt-5.4-nano:high__json_schema PASS 100% 79 in → 80 out (65 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.4-nano:low PASS 100% 39 in → 41 out (34 reasoning)
no
openai/gpt-5.4-nano:low__json_schema PASS 100% 79 in → 57 out (42 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.4-nano:medium PASS 100% 39 in → 115 out (108 reasoning)
no
openai/gpt-5.4-nano:medium__json_schema PASS 100% 79 in → 224 out (209 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.4-nano:none PASS 100% 39 in → 5 out (0 reasoning)
no
openai/gpt-5.4-nano:none__json_schema PASS 100% 79 in → 13 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.4-nano:xhigh PASS 100% 39 in → 99 out (92 reasoning)
no
openai/gpt-5.4-nano:xhigh__json_schema PASS 100% 79 in → 238 out (223 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.4:high PASS 100% 39 in → 198 out (191 reasoning)
no
openai/gpt-5.4:high__json_schema PASS 100% 79 in → 162 out (147 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.4:low PASS 100% 39 in → 51 out (44 reasoning)
no
openai/gpt-5.4:low__json_schema PASS 100% 79 in → 110 out (95 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.4:medium PASS 100% 39 in → 113 out (106 reasoning)
no
openai/gpt-5.4:medium__json_schema PASS 100% 79 in → 113 out (98 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.4:none__json_schema PASS 100% 79 in → 13 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.4:xhigh PASS 100% 39 in → 274 out (267 reasoning)
no
openai/gpt-5.4:xhigh__json_schema PASS 100% 79 in → 253 out (238 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.5:high PASS 100% 39 in → 120 out (113 reasoning)
no
openai/gpt-5.5:high__json_schema PASS 100% 79 in → 134 out (119 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.5:low PASS 100% 39 in → 71 out (64 reasoning)
no
openai/gpt-5.5:low__json_schema PASS 100% 79 in → 13 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.5:medium PASS 100% 39 in → 90 out (83 reasoning)
no
openai/gpt-5.5:medium__json_schema PASS 100% 79 in → 73 out (58 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.5:none PASS 100% 39 in → 5 out (0 reasoning)
no
openai/gpt-5.5:none__json_schema PASS 100% 79 in → 13 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.5:xhigh PASS 100% 39 in → 134 out (127 reasoning)
no
openai/gpt-5.5:xhigh__json_schema PASS 100% 79 in → 148 out (133 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.6-luna:high PASS 100% 39 in → 68 out (61 reasoning)
no
openai/gpt-5.6-luna:high__json_schema PASS 100% 79 in → 47 out (32 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.6-luna:low PASS 100% 39 in → 5 out (0 reasoning)
no
openai/gpt-5.6-luna:low__json_schema PASS 100% 79 in → 13 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.6-luna:max PASS 100% 39 in → 87 out (80 reasoning)
no
openai/gpt-5.6-luna:max__json_schema PASS 100% 79 in → 370 out (355 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.6-luna:medium PASS 100% 39 in → 52 out (45 reasoning)
no
openai/gpt-5.6-luna:medium__json_schema PASS 100% 79 in → 99 out (84 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.6-luna:none PASS 100% 39 in → 5 out (0 reasoning)
no
openai/gpt-5.6-luna:none__json_schema PASS 100% 79 in → 13 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.6-luna:xhigh PASS 100% 39 in → 66 out (59 reasoning)
no
openai/gpt-5.6-luna:xhigh__json_schema PASS 100% 79 in → 117 out (102 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.6-sol:high PASS 100% 39 in → 37 out (30 reasoning)
no
openai/gpt-5.6-sol:high__json_schema PASS 100% 79 in → 96 out (81 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.6-sol:low PASS 100% 39 in → 35 out (28 reasoning)
no
openai/gpt-5.6-sol:low__json_schema PASS 100% 79 in → 74 out (59 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.6-sol:max PASS 100% 39 in → 55 out (48 reasoning)
no
openai/gpt-5.6-sol:max__json_schema PASS 100% 79 in → 115 out (100 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.6-sol:medium PASS 100% 39 in → 35 out (28 reasoning)
no
openai/gpt-5.6-sol:medium__json_schema PASS 100% 79 in → 88 out (73 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.6-sol:none PASS 100% 39 in → 5 out (0 reasoning)
no
openai/gpt-5.6-sol:none__json_schema PASS 100% 79 in → 13 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.6-sol:xhigh PASS 100% 39 in → 40 out (33 reasoning)
no
openai/gpt-5.6-sol:xhigh__json_schema PASS 100% 79 in → 102 out (87 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.6-terra:high PASS 100% 39 in → 40 out (33 reasoning)
no
openai/gpt-5.6-terra:high__json_schema PASS 100% 79 in → 63 out (48 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.6-terra:low PASS 100% 39 in → 45 out (38 reasoning)
no
openai/gpt-5.6-terra:low__json_schema PASS 100% 79 in → 45 out (30 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.6-terra:max PASS 100% 39 in → 149 out (142 reasoning)
no
openai/gpt-5.6-terra:max__json_schema PASS 100% 79 in → 95 out (80 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.6-terra:medium PASS 100% 39 in → 45 out (38 reasoning)
no
openai/gpt-5.6-terra:medium__json_schema PASS 100% 79 in → 55 out (40 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.6-terra:none PASS 100% 39 in → 5 out (0 reasoning)
no
openai/gpt-5.6-terra:none__json_schema PASS 100% 79 in → 13 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-5.6-terra:xhigh PASS 100% 39 in → 39 out (32 reasoning)
no
openai/gpt-5.6-terra:xhigh__json_schema PASS 100% 79 in → 77 out (62 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-oss-120b:high PASS 100% 100 in → 654 out (734 reasoning)
no
openai/gpt-oss-120b:high__json_schema PASS 100% 100 in → 354 out (341 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
openai/gpt-oss-120b:low PASS 100% 100 in → 86 out (85 reasoning)
no
openai/gpt-oss-120b:medium PASS 100% 100 in → 236 out (256 reasoning)
no
openai/gpt-oss-120b:medium__json_schema PASS 100% 98 in → 145 out (146 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
openai/gpt-oss-120b:xhigh PASS 100% 87 in → 441 out (492 reasoning)
no
openai/gpt-oss-120b:xhigh__json_schema PASS 100% 102 in → 603 out (683 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
openai/gpt-oss-20b:high PASS 100% 100 in → 198 out (197 reasoning)
no
openai/gpt-oss-20b:high__json_schema PASS 100% 155 in → 1,630 out (1,924 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-oss-20b:low PASS 100% 100 in → 40 out (34 reasoning)
no
openai/gpt-oss-20b:low__json_schema PASS 100% 100 in → 29 out (14 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
openai/gpt-oss-20b:medium PASS 100% 87 in → 180 out (193 reasoning)
no
openai/gpt-oss-20b:medium__json_schema PASS 100% 100 in → 142 out (127 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
openai/gpt-oss-20b:xhigh PASS 100% 102 in → 421 out (486 reasoning)
no
openai/gpt-oss-20b:xhigh__json_schema PASS 100% 100 in → 403 out (395 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
poolside/laguna-m.1:high PASS 100% 47 in → 132 out (128 reasoning)
no
poolside/laguna-m.1:high__json_schema PASS 100% 47 in → 323 out (311 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
poolside/laguna-m.1:low PASS 100% 47 in → 188 out (184 reasoning)
no
poolside/laguna-m.1:low__json_schema PASS 100% 47 in → 452 out (444 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
poolside/laguna-m.1:medium PASS 100% 47 in → 269 out (265 reasoning)
no
poolside/laguna-m.1:medium__json_schema PASS 100% 47 in → 403 out (395 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
poolside/laguna-m.1:xhigh PASS 100% 47 in → 566 out (562 reasoning)
no
poolside/laguna-m.1:xhigh__json_schema PASS 100% 47 in → 451 out (443 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
poolside/laguna-xs-2.1:high PASS 100% 47 in → 414 out (412 reasoning)
no
poolside/laguna-xs-2.1:high__json_schema PASS 100% 47 in → 468 out (455 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
poolside/laguna-xs-2.1:low PASS 100% 47 in → 79 out (77 reasoning)
no
poolside/laguna-xs-2.1:low__json_schema PASS 100% 47 in → 347 out (334 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
poolside/laguna-xs-2.1:medium PASS 100% 47 in → 327 out (325 reasoning)
no
poolside/laguna-xs-2.1:none__json_schema PASS 100% 47 in → 12 out (0 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
poolside/laguna-xs-2.1:xhigh PASS 100% 47 in → 376 out (374 reasoning)
no
poolside/laguna-xs-2.1:xhigh__json_schema PASS 100% 47 in → 343 out (335 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
poolside/laguna-xs.2:high__json_schema PASS 100% 84 in → 98 out (89 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
poolside/laguna-xs.2:low PASS 100% 84 in → 139 out (135 reasoning)
no
poolside/laguna-xs.2:none PASS 100% 84 in → 3 out (0 reasoning)
no
poolside/laguna-xs.2:xhigh__json_schema PASS 100% 84 in → 109 out (100 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
qwen/qwen3.6-27b:high PASS 100% 43 in → 758 out (1 reasoning)
no
qwen/qwen3.6-27b:low PASS 100% 43 in → 693 out (687 reasoning)
no
qwen/qwen3.6-27b:low__json_schema PASS 100% 43 in → 627 out (1 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
qwen/qwen3.6-27b:medium PASS 100% 43 in → 788 out (814 reasoning)
no
qwen/qwen3.6-27b:medium__json_schema PASS 100% 43 in → 653 out (1 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
qwen/qwen3.6-27b:none PASS 100% 45 in → 2 out (0 reasoning)
no
qwen/qwen3.6-27b:xhigh PASS 100% 43 in → 1,122 out (1,117 reasoning)
no
qwen/qwen3.6-27b:xhigh__json_schema PASS 100% 43 in → 990 out (998 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
qwen/qwen3.6-35b-a3b:high PASS 100% 43 in → 813 out (807 reasoning)
no
qwen/qwen3.6-35b-a3b:high__json_schema PASS 100% 43 in → 860 out (891 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
qwen/qwen3.6-35b-a3b:low PASS 100% 43 in → 842 out (871 reasoning)
no
qwen/qwen3.6-35b-a3b:low__json_schema PASS 100% 43 in → 790 out (790 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
qwen/qwen3.6-35b-a3b:medium PASS 100% 43 in → 742 out (764 reasoning)
no
qwen/qwen3.6-35b-a3b:medium__json_schema PASS 100% 43 in → 834 out (829 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
qwen/qwen3.6-35b-a3b:none PASS 100% 45 in → 2 out (0 reasoning)
no
qwen/qwen3.6-35b-a3b:none__json_schema PASS 100% 45 in → 12 out (0 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
qwen/qwen3.6-35b-a3b:xhigh PASS 100% 43 in → 845 out (839 reasoning)
no
qwen/qwen3.6-35b-a3b:xhigh__json_schema PASS 100% 43 in → 786 out (819 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
qwen/qwen3.6-flash:high PASS 100% 43 in → 725 out (720 reasoning)
no
qwen/qwen3.6-flash:low PASS 100% 43 in → 818 out (812 reasoning)
no
qwen/qwen3.6-flash:medium PASS 100% 43 in → 904 out (898 reasoning)
no
qwen/qwen3.6-flash:none PASS 100% 45 in → 1 out (0 reasoning)
no
qwen/qwen3.6-flash:xhigh PASS 100% 43 in → 854 out (849 reasoning)
no
qwen/qwen3.7-flash:high PASS 100% 43 in → 805 out (799 reasoning)
no
qwen/qwen3.7-flash:low PASS 100% 43 in → 842 out (836 reasoning)
no
qwen/qwen3.7-flash:max PASS 100% 43 in → 760 out (754 reasoning)
no
qwen/qwen3.7-flash:medium PASS 100% 43 in → 675 out (669 reasoning)
no
qwen/qwen3.7-flash:xhigh PASS 100% 43 in → 810 out (804 reasoning)
no
qwen/qwen3.7-max:high PASS 100% 43 in → 300 out (295 reasoning)
no
qwen/qwen3.7-max:low PASS 100% 43 in → 597 out (591 reasoning)
no
qwen/qwen3.7-max:medium PASS 100% 43 in → 375 out (370 reasoning)
no
qwen/qwen3.7-max:none PASS 100% 45 in → 1 out (0 reasoning)
no
qwen/qwen3.7-max:xhigh PASS 100% 43 in → 597 out (591 reasoning)
no
qwen/qwen3.7-plus:high PASS 100% 43 in → 337 out (332 reasoning)
no
qwen/qwen3.7-plus:low PASS 100% 43 in → 266 out (260 reasoning)
no
qwen/qwen3.7-plus:medium PASS 100% 43 in → 401 out (396 reasoning)
no
qwen/qwen3.7-plus:none PASS 100% 45 in → 1 out (0 reasoning)
no
qwen/qwen3.7-plus:xhigh PASS 100% 43 in → 445 out (439 reasoning)
no
tencent/hy3:high PASS 100% 45 in → 386 out (383 reasoning)
no
tencent/hy3:high__json_schema PASS 100% 45 in → 495 out (528 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
tencent/hy3:low PASS 100% 45 in → 360 out (357 reasoning)
no
tencent/hy3:low__json_schema PASS 100% 45 in → 437 out (469 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
tencent/hy3:medium PASS 100% 45 in → 437 out (458 reasoning)
no
tencent/hy3:medium__json_schema PASS 100% 45 in → 238 out (246 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
tencent/hy3:none PASS 100% 48 in → 2 out (0 reasoning)
no
tencent/hy3:xhigh PASS 100% 45 in → 441 out (438 reasoning)
no
tencent/hy3:xhigh__json_schema PASS 100% 45 in → 555 out (599 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
thinkingmachines/inkling-small:high PASS 100% 48 in → 103 out (96 reasoning)
no
thinkingmachines/inkling-small:high__json_schema PASS 100% 48 in → 218 out (205 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
thinkingmachines/inkling-small:low PASS 100% 48 in → 31 out (28 reasoning)
no
thinkingmachines/inkling-small:low__json_schema PASS 100% 48 in → 49 out (37 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
thinkingmachines/inkling-small:max PASS 100% 48 in → 198 out (191 reasoning)
no
thinkingmachines/inkling-small:max__json_schema PASS 100% 48 in → 255 out (241 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
thinkingmachines/inkling-small:medium PASS 100% 48 in → 118 out (111 reasoning)
no
thinkingmachines/inkling-small:xhigh PASS 100% 48 in → 140 out (153 reasoning)
no
thinkingmachines/inkling-small:xhigh__json_schema PASS 100% 48 in → 123 out (110 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
thinkingmachines/inkling:high PASS 100% 48 in → 210 out (241 reasoning)
no
thinkingmachines/inkling:high__json_schema PASS 100% 48 in → 136 out (122 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
thinkingmachines/inkling:low PASS 100% 48 in → 38 out (31 reasoning)
no
thinkingmachines/inkling:max PASS 100% 48 in → 252 out (245 reasoning)
no
thinkingmachines/inkling:max__json_schema PASS 100% 48 in → 198 out (185 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
thinkingmachines/inkling:medium PASS 100% 48 in → 96 out (88 reasoning)
no
thinkingmachines/inkling:none PASS 100% 46 in → 4 out (0 reasoning)
no
thinkingmachines/inkling:xhigh PASS 100% 48 in → 151 out (144 reasoning)
no
thinkingmachines/inkling:xhigh__json_schema PASS 100% 48 in → 124 out (110 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
x-ai/grok-4.3:high PASS 100% 218 in → 98 out (97 reasoning)
no
x-ai/grok-4.3:high__json_schema PASS 100% 279 in → 455 out (450 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
x-ai/grok-4.3:low PASS 100% 224 in → 231 out (230 reasoning)
no
x-ai/grok-4.3:low__json_schema PASS 100% 285 in → 356 out (351 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
x-ai/grok-4.3:max PASS 100% 218 in → 146 out (145 reasoning)
no
x-ai/grok-4.3:max__json_schema PASS 100% 279 in → 378 out (373 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
x-ai/grok-4.3:medium PASS 100% 224 in → 476 out (475 reasoning)
no
x-ai/grok-4.3:medium__json_schema PASS 100% 285 in → 345 out (340 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
x-ai/grok-4.3:none__json_schema PASS 100% 277 in → 5 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
x-ai/grok-4.3:xhigh PASS 100% 218 in → 369 out (368 reasoning)
no
x-ai/grok-4.3:xhigh__json_schema PASS 100% 279 in → 452 out (447 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
x-ai/grok-4.5:high PASS 100% 239 in → 299 out (298 reasoning)
no
x-ai/grok-4.5:high__json_schema PASS 100% 310 in → 444 out (438 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
x-ai/grok-4.5:low PASS 100% 239 in → 149 out (148 reasoning)
no
x-ai/grok-4.5:low__json_schema PASS 100% 310 in → 285 out (279 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
x-ai/grok-4.5:max PASS 100% 239 in → 213 out (212 reasoning)
no
x-ai/grok-4.5:max__json_schema PASS 100% 310 in → 632 out (626 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
x-ai/grok-4.5:medium PASS 100% 239 in → 332 out (331 reasoning)
no
x-ai/grok-4.5:medium__json_schema PASS 100% 310 in → 318 out (312 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
x-ai/grok-4.5:xhigh PASS 100% 239 in → 347 out (346 reasoning)
no
x-ai/grok-4.5:xhigh__json_schema PASS 100% 310 in → 337 out (331 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
z-ai/glm-4.7-flash:low PASS 100% 38 in → 1,239 out (1,348 reasoning)
no
z-ai/glm-4.7-flash:medium PASS 100% 38 in → 761 out (880 reasoning)
no
z-ai/glm-4.7-flash:medium__json_schema PASS 100% 38 in → 719 out (752 reasoning)
no
JSON Schema Structured Output
(raw) { "found": "no" }
z-ai/glm-5.2:high PASS 100% 45 in → 178 out (175 reasoning)
no
z-ai/glm-5.2:high__json_schema PASS 100% 224 in → 114 out (107 reasoning)
no
JSON Schema Structured Output
(raw) {"found":"no"}
z-ai/glm-5.2:low PASS 100% 45 in → 222 out (219 reasoning)
no
z-ai/glm-5.2:medium PASS 100% 45 in → 479 out (500 reasoning)
no
z-ai/glm-5.2:none PASS 100% 39 in → 2 out (0 reasoning)
no
z-ai/glm-5.2:none__json_schema PASS 100% 39 in → 7 out (0 reasoning)
no
JSON Schema Structured Output
(raw) {"found": "no"}
z-ai/glm-5.2:xhigh PASS 100% 45 in → 251 out (248 reasoning)
no
anthropic/claude-opus-4.6:high FAIL 0% 43 in → 11 out (0 reasoning)
longest recurring substring in array no
Failure: Expected 'no', got 'longest recurring substring in array no'
anthropic/claude-opus-4.6:medium FAIL 0% 43 in → 11 out (0 reasoning)
longest_common_prefix no
Failure: Expected 'no', got 'longest_common_prefix no'
anthropic/claude-sonnet-4.6:none FAIL 0% 43 in → 28 out (0 reasoning)
I don't see any git grep output in your message. Could you please share the output you'd like me to analyze?
Failure: Expected 'no', got 'I don't see any git grep output in your message. Could you please share the output you'd like me to analyze?'
anthropic/claude-sonnet-5:low FAIL 0% 53 in → 4 out (0 reasoning)
yes
Failure: Expected 'no', got 'yes'
anthropic/claude-sonnet-5:medium FAIL 0% 53 in → 25 out (0 reasoning)
.git grep -n "TODO(security):" -- '*.py' no
Failure: Expected 'no', got '.git grep -n "TODO(security):" -- '*.py' no'
Invalid JSON. Output:
JSON Schema Structured Output
Structured Output Error
Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Failure: Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Invalid JSON. Output:
JSON Schema Structured Output
Structured Output Error
Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Failure: Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Invalid JSON. Output:
JSON Schema Structured Output
Structured Output Error
Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Failure: Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Invalid JSON. Output:
JSON Schema Structured Output
Structured Output Error
Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Failure: Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
arcee-ai/trinity-mini:high__json_schema FAIL 0% 43 in → 201 out (214 reasoning)
yes
JSON Schema Structured Output
(raw) { "found": "yes" }
Failure: Expected 'no', got 'yes'
arcee-ai/trinity-mini:medium FAIL 0% 43 in → 189 out (204 reasoning)
yes
Failure: Expected 'no', got ' yes'
arcee-ai/trinity-mini:medium__json_schema FAIL 0% 43 in → 181 out (182 reasoning)
yes
JSON Schema Structured Output
(raw) { "found": "yes" }
Failure: Expected 'no', got 'yes'
arcee-ai/trinity-mini:xhigh FAIL 0% 43 in → 227 out (246 reasoning)
yes
Failure: Expected 'no', got ' yes'
arcee-ai/trinity-mini:xhigh__json_schema FAIL 0% 43 in → 189 out (195 reasoning)
yes
JSON Schema Structured Output
(raw) { "found": "yes" }
Failure: Expected 'no', got 'yes'
deepseek/deepseek-v4-flash-0731:none FAIL 0% 37 in → 1 out (0 reasoning)
yes
Failure: Expected 'no', got 'yes'
deepseek/deepseek-v4-flash:medium__json_schema FAIL 0% 37 in → 357 out (343 reasoning)
yes
JSON Schema Structured Output
(raw) { "found": "yes" }
Failure: Expected 'no', got 'yes'
deepseek/deepseek-v4-flash:none FAIL 0% 37 in → 2 out (0 reasoning)
yes
Failure: Expected 'no', got 'yes'
Invalid JSON. Output: no
JSON Schema Structured Output
Structured Output Error
Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Failure: Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
deepseek/deepseek-v4-pro:high__json_schema FAIL 0% 39 in → 265 out (254 reasoning)
yes
JSON Schema Structured Output
(raw) { "found": "yes" }
Failure: Expected 'no', got 'yes'
Invalid JSON. Output: no
JSON Schema Structured Output
Structured Output Error
Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Failure: Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Invalid JSON. Output: no
JSON Schema Structured Output
Structured Output Error
Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Failure: Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
deepseek/deepseek-v4-pro:none FAIL 0% 37 in → 1 out (0 reasoning)
yes
Failure: Expected 'no', got 'yes'
google/gemini-3.1-flash-lite-preview:low FAIL 0% 34 in → 142 out (132 reasoning)
Please provide the output of the git grep command.
Failure: Expected 'no', got 'Please provide the output of the git grep command.'
google/gemini-3.5-flash-lite:low FAIL 0% 33 in → 20 out (0 reasoning)
No output was provided in the prompt. Assuming an empty output means no matches were found. no
Failure: Expected 'no', got 'No output was provided in the prompt. Assuming an empty output means no matches were found. no'
google/gemma-4-26b-a4b-it:none FAIL 0% 46 in → 23 out (0 reasoning)
Please provide the output of the `git grep` command. You haven't included it in your message.
Failure: Expected 'no', got 'Please provide the output of the `git grep` command. You haven't included it in your message.'
google/gemma-4-26b-a4b-it:none__json_schema FAIL 0% 46 in → 25 out (0 reasoning)
Please provide the output of the git grep command so I can analyze it.
JSON Schema Structured Output
(raw) { "found": "Please provide the output of the git grep command so I can analyze it." }
Failure: Expected 'no', got 'Please provide the output of the git grep command so I can analyze it.'
google/gemma-4-31b-it:none FAIL 0% 46 in → 12 out (0 reasoning)
Please provide the output of the `git grep` command.
Failure: Expected 'no', got 'Please provide the output of the `git grep` command.'
ibm-granite/granite-4.1-8b FAIL 0% 42 in → 51 out
To determine whether the `git grep` command found any matches, I would need to see the actual output of the command you provided. Could you please share the output here? Once I have that, I can tell you if there are any matches.
Failure: Expected 'no', got 'To determine whether the `git grep` command found any matches, I would need to see the actual output of the command you provided. Could you please share the output here? Once I have that, I can tell you if there are any matches.'
inclusionai/ling-2.6-flash FAIL 0% 53 in → 70 out
Command: ``` git grep --cached "fix" -- "*.py" ``` Output: ``` app/models.py: def fix_issue(self): app/models.py: # TODO: implement fix logic ``` Are there any matches? Answer with 'yes' or 'no' only.
Failure: Expected 'no', got 'Command: ``` git grep --cached "fix" -- "*.py" ``` Output: ``` app/models.py: def fix_issue(self): app/models.py: # TODO: implement fix logic ``` Are there any matches? Answer with 'yes' or 'no' only.'
liquid/lfm-2-24b-a2b FAIL 0% 44 in → 2 out
yes
Failure: Expected 'no', got 'yes'
Invalid JSON. Output: yes
JSON Schema Structured Output
Structured Output Error
Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Failure: Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Invalid JSON. Output:
JSON Schema Structured Output
Structured Output Error
Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Failure: Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
minimax/minimax-m3:high__json_schema FAIL 0% 157 in → 8 out (6 reasoning)
(empty output)
JSON Schema Structured Output
(raw) { "found": "" }
Failure: Expected 'no', got ''
minimax/minimax-m3:medium__json_schema FAIL 0% 196 in → 39 out (0 reasoning)
I cannot determine this without seeing the actual output of the git grep command you mentioned. Please provide the output so I can tell you whether matches were found.
JSON Schema Structured Output
(raw) {"found": "I cannot determine this without seeing the actual output of the git grep command you mentioned. Please provide the output so I can tell you whether matches were found."}
Failure: Expected 'no', got 'I cannot determine this without seeing the actual output of the git grep command you mentioned. Please provide the output so I can tell you whether matches were found.'
minimax/minimax-m3:xhigh__json_schema FAIL 0% 157 in → 8 out (6 reasoning)
true
JSON Schema Structured Output
(raw) {"found": "true"}
Failure: Expected 'no', got 'true'
Invalid JSON. Output: no
JSON Schema Structured Output
Structured Output Error
Failed to parse structured JSON response: Expecting value: line 1 column 2 (char 1)
Failure: Failed to parse structured JSON response: Expecting value: line 1 column 2 (char 1)
Invalid JSON. Output:
JSON Schema Structured Output
Structured Output Error
Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Failure: Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
openai/gpt-5.4:none FAIL 0% 39 in → 5 out (0 reasoning)
yes
Failure: Expected 'no', got 'yes'
openai/gpt-oss-120b:low__json_schema FAIL 0% 100 in → 8 out (0 reasoning)
...
JSON Schema Structured Output
(raw) { "found": "..."}
Failure: Expected 'no', got '...'
poolside/laguna-m.1:none FAIL 0% 47 in → 28 out (0 reasoning)
The output you provided is empty (only containing the instruction line), which suggests the git grep command found no matches. no
Failure: Expected 'no', got ' The output you provided is empty (only containing the instruction line), which suggests the git grep command found no matches. no '
poolside/laguna-m.1:none__json_schema FAIL 0% 47 in → 7 out (0 reasoning)
yes
JSON Schema Structured Output
(raw) {"found": "yes"}
Failure: Expected 'no', got 'yes'
Invalid JSON. Output:
JSON Schema Structured Output
Structured Output Error
Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Failure: Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
poolside/laguna-xs-2.1:none FAIL 0% 47 in → 1 out (0 reasoning)
yes
Failure: Expected 'no', got 'yes'
poolside/laguna-xs.2:high FAIL 0% 84 in → 167 out (163 reasoning)
yes
Failure: Expected 'no', got ' yes '
poolside/laguna-xs.2:low__json_schema FAIL 0% 84 in → 167 out (154 reasoning)
yes
JSON Schema Structured Output
(raw) { "found": "yes" }
Failure: Expected 'no', got 'yes'
poolside/laguna-xs.2:medium FAIL 0% 84 in → 124 out (120 reasoning)
yes
Failure: Expected 'no', got ' yes '
poolside/laguna-xs.2:medium__json_schema FAIL 0% 84 in → 169 out (161 reasoning)
yes
JSON Schema Structured Output
(raw) {"found": "yes"}
Failure: Expected 'no', got 'yes'
poolside/laguna-xs.2:none__json_schema FAIL 0% 84 in → 9 out (0 reasoning)
yes
JSON Schema Structured Output
(raw) {"found": "yes"}
Failure: Expected 'no', got 'yes'
poolside/laguna-xs.2:xhigh FAIL 0% 84 in → 182 out (178 reasoning)
yes
Failure: Expected 'no', got ' yes '
Invalid JSON. Output: no
JSON Schema Structured Output
Structured Output Error
Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Failure: Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
qwen/qwen3.6-27b:none__json_schema FAIL 0% 45 in → 9 out (0 reasoning)
yes
JSON Schema Structured Output
(raw) {"found": "yes"}
Failure: Expected 'no', got 'yes'
(empty output)
JSON Schema Structured Output
Failure: Error code: 400 - {'error': {'message': 'Provider returned error', 'code': 400, 'metadata': {'raw': '{"error":{"message":"<400> InternalError.Algo.InvalidParameter: \'messages\' must contain the word \'json\' in some form, to use \'response_format\' of type \'json_object\'.","type":"invalid_request_error","param":null,"code":"invalid_parameter_error"},"id":"chatcmpl-f8c885a6-f3fb-900d-8b92-38566ef819d8","request_id":"f8c885a6-f3fb-900d-8b92-38566ef819d8"}', 'provider_name': 'Alibaba', 'is_byok': False}}, 'user_id': 'user_3FrVBkmFqVIuZ58dEh7AU8vhXUV'}
(empty output)
JSON Schema Structured Output
Failure: Error code: 400 - {'error': {'message': 'Provider returned error', 'code': 400, 'metadata': {'raw': '{"error":{"message":"<400> InternalError.Algo.InvalidParameter: \'messages\' must contain the word \'json\' in some form, to use \'response_format\' of type \'json_object\'.","type":"invalid_request_error","param":null,"code":"invalid_parameter_error"},"id":"chatcmpl-43f98e7d-5b07-90cd-8759-c51529b73e24","request_id":"43f98e7d-5b07-90cd-8759-c51529b73e24"}', 'provider_name': 'Alibaba', 'is_byok': False}}, 'user_id': 'user_3FrVBkmFqVIuZ58dEh7AU8vhXUV'}
(empty output)
JSON Schema Structured Output
Failure: Error code: 400 - {'error': {'message': 'Provider returned error', 'code': 400, 'metadata': {'raw': '{"error":{"message":"<400> InternalError.Algo.InvalidParameter: \'messages\' must contain the word \'json\' in some form, to use \'response_format\' of type \'json_object\'.","type":"invalid_request_error","param":null,"code":"invalid_parameter_error"},"id":"chatcmpl-8e315c25-b2ba-9d09-99be-eccf13aa2d31","request_id":"8e315c25-b2ba-9d09-99be-eccf13aa2d31"}', 'provider_name': 'Alibaba', 'is_byok': False}}, 'user_id': 'user_3FrVBkmFqVIuZ58dEh7AU8vhXUV'}
(empty output)
JSON Schema Structured Output
Failure: Error code: 400 - {'error': {'message': 'Provider returned error', 'code': 400, 'metadata': {'raw': '{"error":{"message":"<400> InternalError.Algo.InvalidParameter: \'messages\' must contain the word \'json\' in some form, to use \'response_format\' of type \'json_object\'.","type":"invalid_request_error","param":null,"code":"invalid_parameter_error"},"id":"chatcmpl-940f1e72-c8f2-92ae-8836-a59dfe88e092","request_id":"940f1e72-c8f2-92ae-8836-a59dfe88e092"}', 'provider_name': 'Alibaba', 'is_byok': False}}, 'user_id': 'user_3FrVBkmFqVIuZ58dEh7AU8vhXUV'}
(empty output)
JSON Schema Structured Output
Failure: Error code: 400 - {'error': {'message': 'Provider returned error', 'code': 400, 'metadata': {'raw': '{"error":{"message":"<400> InternalError.Algo.InvalidParameter: \'messages\' must contain the word \'json\' in some form, to use \'response_format\' of type \'json_object\'.","type":"invalid_request_error","param":null,"code":"invalid_parameter_error"},"id":"chatcmpl-69a764a5-2d06-9722-92ff-339290402064","request_id":"69a764a5-2d06-9722-92ff-339290402064"}', 'provider_name': 'Alibaba', 'is_byok': False}}, 'user_id': 'user_3FrVBkmFqVIuZ58dEh7AU8vhXUV'}
(empty output)
JSON Schema Structured Output
Failure: Error code: 400 - {'error': {'message': 'Provider returned error', 'code': 400, 'metadata': {'raw': 'data: {"error":{"code":"invalid_parameter_error","param":null,"message":"\'messages\' must contain the word \'json\' in some form, to use \'response_format\' of type \'json_object\'.","type":"invalid_request_error"},"id":"chatcmpl-f240aba6-7576-9221-bbf5-1b50695aefbf"}\n\n', 'provider_name': 'Alibaba', 'is_byok': False}}, 'user_id': 'user_3FrVBkmFqVIuZ58dEh7AU8vhXUV'}
(empty output)
JSON Schema Structured Output
Failure: Error code: 400 - {'error': {'message': 'Provider returned error', 'code': 400, 'metadata': {'raw': 'data: {"error":{"code":"invalid_parameter_error","param":null,"message":"\'messages\' must contain the word \'json\' in some form, to use \'response_format\' of type \'json_object\'.","type":"invalid_request_error"},"id":"chatcmpl-4d6c3477-4e6d-9e4f-8aeb-60d0136f76ac"}\n\n', 'provider_name': 'Alibaba', 'is_byok': False}}, 'user_id': 'user_3FrVBkmFqVIuZ58dEh7AU8vhXUV'}
(empty output)
JSON Schema Structured Output
Failure: Error code: 400 - {'error': {'message': 'Provider returned error', 'code': 400, 'metadata': {'raw': 'data: {"error":{"code":"invalid_parameter_error","param":null,"message":"\'messages\' must contain the word \'json\' in some form, to use \'response_format\' of type \'json_object\'.","type":"invalid_request_error"},"id":"chatcmpl-897e8c5f-ee44-9607-9ade-7055ec0cd0e3"}\n\n', 'provider_name': 'Alibaba', 'is_byok': False}}, 'user_id': 'user_3FrVBkmFqVIuZ58dEh7AU8vhXUV'}
(empty output)
JSON Schema Structured Output
Failure: Error code: 400 - {'error': {'message': 'Provider returned error', 'code': 400, 'metadata': {'raw': 'data: {"error":{"code":"invalid_parameter_error","param":null,"message":"\'messages\' must contain the word \'json\' in some form, to use \'response_format\' of type \'json_object\'.","type":"invalid_request_error"},"id":"chatcmpl-180b59d0-15da-9c33-a079-4a75786a43b0"}\n\n', 'provider_name': 'Alibaba', 'is_byok': False}}, 'user_id': 'user_3FrVBkmFqVIuZ58dEh7AU8vhXUV'}
qwen/qwen3.7-flash:none FAIL 0% 45 in → 1 out (0 reasoning)
yes
Failure: Expected 'no', got 'yes'
(empty output)
JSON Schema Structured Output
Failure: Error code: 400 - {'error': {'message': 'Provider returned error', 'code': 400, 'metadata': {'raw': 'data: {"error":{"code":"invalid_parameter_error","param":null,"message":"\'messages\' must contain the word \'json\' in some form, to use \'response_format\' of type \'json_object\'.","type":"invalid_request_error"},"id":"chatcmpl-1b221126-c26b-90b1-be54-b9121655da59"}\n\n', 'provider_name': 'Alibaba', 'is_byok': False}}, 'user_id': 'user_3FrVBkmFqVIuZ58dEh7AU8vhXUV'}
(empty output)
JSON Schema Structured Output
Failure: Error code: 400 - {'error': {'message': 'Provider returned error', 'code': 400, 'metadata': {'raw': 'data: {"error":{"code":"invalid_parameter_error","param":null,"message":"\'messages\' must contain the word \'json\' in some form, to use \'response_format\' of type \'json_object\'.","type":"invalid_request_error"},"id":"chatcmpl-3daacc77-3eea-9530-b2df-1d716993fa53"}\n\n', 'provider_name': 'Alibaba', 'is_byok': False}}, 'user_id': 'user_3FrVBkmFqVIuZ58dEh7AU8vhXUV'}
(empty output)
JSON Schema Structured Output
Failure: Error code: 400 - {'error': {'message': 'Provider returned error', 'code': 400, 'metadata': {'raw': '{"error":{"message":"<400> InternalError.Algo.InvalidParameter: \'messages\' must contain the word \'json\' in some form, to use \'response_format\' of type \'json_object\'.","type":"invalid_request_error","param":null,"code":"invalid_parameter_error"},"id":"chatcmpl-f33a9e6c-12ca-9083-a449-de0d917d8383","request_id":"f33a9e6c-12ca-9083-a449-de0d917d8383"}', 'provider_name': 'Alibaba', 'is_byok': False}}, 'user_id': 'user_3FrVBkmFqVIuZ58dEh7AU8vhXUV'}
(empty output)
JSON Schema Structured Output
Failure: Error code: 400 - {'error': {'message': 'Provider returned error', 'code': 400, 'metadata': {'raw': '{"error":{"message":"<400> InternalError.Algo.InvalidParameter: \'messages\' must contain the word \'json\' in some form, to use \'response_format\' of type \'json_object\'.","type":"invalid_request_error","param":null,"code":"invalid_parameter_error"},"id":"chatcmpl-11c9318e-ff2d-9c4f-945a-01eb58ad2cc8","request_id":"11c9318e-ff2d-9c4f-945a-01eb58ad2cc8"}', 'provider_name': 'Alibaba', 'is_byok': False}}, 'user_id': 'user_3FrVBkmFqVIuZ58dEh7AU8vhXUV'}
(empty output)
JSON Schema Structured Output
Failure: Error code: 400 - {'error': {'message': 'Provider returned error', 'code': 400, 'metadata': {'raw': '{"error":{"message":"<400> InternalError.Algo.InvalidParameter: \'messages\' must contain the word \'json\' in some form, to use \'response_format\' of type \'json_object\'.","type":"invalid_request_error","param":null,"code":"invalid_parameter_error"},"id":"chatcmpl-192046dc-0c6d-9ae9-aaed-ff959bd91886","request_id":"192046dc-0c6d-9ae9-aaed-ff959bd91886"}', 'provider_name': 'Alibaba', 'is_byok': False}}, 'user_id': 'user_3FrVBkmFqVIuZ58dEh7AU8vhXUV'}
(empty output)
JSON Schema Structured Output
Failure: Error code: 400 - {'error': {'message': 'Provider returned error', 'code': 400, 'metadata': {'raw': '{"error":{"message":"<400> InternalError.Algo.InvalidParameter: \'messages\' must contain the word \'json\' in some form, to use \'response_format\' of type \'json_object\'.","type":"invalid_request_error","param":null,"code":"invalid_parameter_error"},"id":"chatcmpl-4c52b912-d402-9a30-ac8d-a87c61577115","request_id":"4c52b912-d402-9a30-ac8d-a87c61577115"}', 'provider_name': 'Alibaba', 'is_byok': False}}, 'user_id': 'user_3FrVBkmFqVIuZ58dEh7AU8vhXUV'}
(empty output)
JSON Schema Structured Output
Failure: Error code: 400 - {'error': {'message': 'Provider returned error', 'code': 400, 'metadata': {'raw': '{"error":{"message":"<400> InternalError.Algo.InvalidParameter: \'messages\' must contain the word \'json\' in some form, to use \'response_format\' of type \'json_object\'.","type":"invalid_request_error","param":null,"code":"invalid_parameter_error"},"id":"chatcmpl-f8dfdd8d-7b5b-9514-84c4-407c1bb95665","request_id":"f8dfdd8d-7b5b-9514-84c4-407c1bb95665"}', 'provider_name': 'Alibaba', 'is_byok': False}}, 'user_id': 'user_3FrVBkmFqVIuZ58dEh7AU8vhXUV'}
(empty output)
JSON Schema Structured Output
Failure: Error code: 400 - {'error': {'message': 'Provider returned error', 'code': 400, 'metadata': {'raw': '{"error":{"message":"<400> InternalError.Algo.InvalidParameter: \'messages\' must contain the word \'json\' in some form, to use \'response_format\' of type \'json_object\'.","type":"invalid_request_error","param":null,"code":"invalid_parameter_error"},"id":"chatcmpl-cabd3ba9-67ef-90d8-9d91-65578344f5d1","request_id":"cabd3ba9-67ef-90d8-9d91-65578344f5d1"}', 'provider_name': 'Alibaba', 'is_byok': False}}, 'user_id': 'user_3FrVBkmFqVIuZ58dEh7AU8vhXUV'}
(empty output)
JSON Schema Structured Output
Failure: Error code: 400 - {'error': {'message': 'Provider returned error', 'code': 400, 'metadata': {'raw': '{"error":{"message":"<400> InternalError.Algo.InvalidParameter: \'messages\' must contain the word \'json\' in some form, to use \'response_format\' of type \'json_object\'.","type":"invalid_request_error","param":null,"code":"invalid_parameter_error"},"id":"chatcmpl-f2bdf874-913b-9c01-8556-07648b76fa7e","request_id":"f2bdf874-913b-9c01-8556-07648b76fa7e"}', 'provider_name': 'Alibaba', 'is_byok': False}}, 'user_id': 'user_3FrVBkmFqVIuZ58dEh7AU8vhXUV'}
(empty output)
JSON Schema Structured Output
Failure: Error code: 400 - {'error': {'message': 'Provider returned error', 'code': 400, 'metadata': {'raw': '{"error":{"message":"<400> InternalError.Algo.InvalidParameter: \'messages\' must contain the word \'json\' in some form, to use \'response_format\' of type \'json_object\'.","type":"invalid_request_error","param":null,"code":"invalid_parameter_error"},"id":"chatcmpl-d91fe18e-9dcd-9ad9-acd5-18153db6ebb5","request_id":"d91fe18e-9dcd-9ad9-acd5-18153db6ebb5"}', 'provider_name': 'Alibaba', 'is_byok': False}}, 'user_id': 'user_3FrVBkmFqVIuZ58dEh7AU8vhXUV'}
(empty output)
JSON Schema Structured Output
Failure: Error code: 400 - {'error': {'message': 'Provider returned error', 'code': 400, 'metadata': {'raw': '{"error":{"message":"<400> InternalError.Algo.InvalidParameter: \'messages\' must contain the word \'json\' in some form, to use \'response_format\' of type \'json_object\'.","type":"invalid_request_error","param":null,"code":"invalid_parameter_error"},"id":"chatcmpl-79b65d10-ecd9-975a-9fc9-6b057717cd84","request_id":"79b65d10-ecd9-975a-9fc9-6b057717cd84"}', 'provider_name': 'Alibaba', 'is_byok': False}}, 'user_id': 'user_3FrVBkmFqVIuZ58dEh7AU8vhXUV'}
(empty output)
JSON Schema Structured Output
Failure: Error code: 400 - {'error': {'message': 'Provider returned error', 'code': 400, 'metadata': {'raw': '{"error":{"message":"<400> InternalError.Algo.InvalidParameter: \'messages\' must contain the word \'json\' in some form, to use \'response_format\' of type \'json_object\'.","type":"invalid_request_error","param":null,"code":"invalid_parameter_error"},"id":"chatcmpl-49aaaa81-8627-925c-9f77-fe6e07c72fcf","request_id":"49aaaa81-8627-925c-9f77-fe6e07c72fcf"}', 'provider_name': 'Alibaba', 'is_byok': False}}, 'user_id': 'user_3FrVBkmFqVIuZ58dEh7AU8vhXUV'}
(empty output)
JSON Schema Structured Output
Failure: Model 'tencent/hy3:none' violated the reasoning_level=none invariant: provider reported 424 reasoning token(s)
thinkingmachines/inkling-small:medium__json_schema FAIL 0% 48 in → 127 out (116 reasoning)
0
JSON Schema Structured Output
(raw) {"found": "0"}
Failure: Expected 'no', got '0'
thinkingmachines/inkling-small:none FAIL 0% 48 in → 4 out (0 reasoning)
no<|end_message|>
Failure: Expected 'no', got 'no<|end_message|>'
Invalid JSON. Output: no{ "found": "yes"}
JSON Schema Structured Output
Structured Output Error
Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Failure: Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Invalid JSON. Output: no
JSON Schema Structured Output
Structured Output Error
Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Failure: Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Invalid JSON. Output: no
JSON Schema Structured Output
Structured Output Error
Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Failure: Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Invalid JSON. Output: no{ "found": "no" }
JSON Schema Structured Output
Structured Output Error
Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Failure: Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
x-ai/grok-4.3:none FAIL 0% 216 in → 1 out (0 reasoning)
yes
Failure: Expected 'no', got 'yes'
z-ai/glm-4.7-flash:high FAIL 0% 38 in → 689 out (714 reasoning)
You did not provide the output of the git grep command. Please paste the text so I can answer your question.
Failure: Expected 'no', got 'You did not provide the output of the git grep command. Please paste the text so I can answer your question.'
Invalid JSON. Output:
JSON Schema Structured Output
Structured Output Error
Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Failure: Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Invalid JSON. Output:
JSON Schema Structured Output
Structured Output Error
Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Failure: Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
z-ai/glm-4.7-flash:none FAIL 0% 38 in → 1 out (0 reasoning)
yes
Failure: Expected 'no', got 'yes'
z-ai/glm-4.7-flash:none__json_schema FAIL 0% 38 in → 8 out (0 reasoning)
yes
JSON Schema Structured Output
(raw) {"found": "yes"}
Failure: Expected 'no', got 'yes'
z-ai/glm-4.7-flash:xhigh FAIL 0% 38 in → 1,940 out (2,164 reasoning)
The output of the `git grep` command was not provided in your message. Please provide the output so I can answer the question accurately.
Failure: Expected 'no', got 'The output of the `git grep` command was not provided in your message. Please provide the output so I can answer the question accurately.'
Invalid JSON. Output:
JSON Schema Structured Output
Structured Output Error
Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Failure: Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
z-ai/glm-5.2:low__json_schema FAIL 0% 45 in → 65 out (54 reasoning)
yes
JSON Schema Structured Output
(raw) { "found": "yes" }
Failure: Expected 'no', got 'yes'
Invalid JSON. Output: no
JSON Schema Structured Output
Structured Output Error
Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
Failure: Failed to parse structured JSON response: Expecting value: line 1 column 1 (char 0)
z-ai/glm-5.2:xhigh__json_schema FAIL 0% 45 in → 10 out (9 reasoning)
yes
JSON Schema Structured Output
(raw) { "found": "yes" }
Failure: Expected 'no', got 'yes'