Changes
The changelog of measured change: an entry exists only when an eval's output moved from one measured value to another - both dated, both backed by evidence. New providers or models joining the sheet don't appear here; a provider's behavior changing does. Providers ship changes silently - this page is where those changes stop being silent.
Novita · tool-choice: IGNORED (required + forced-specific) → PASS (all three conditions enforced)
Healed in two steps: forced-specific was honored again on 2026-08-25 with "required" still ignored, and all three conditions passed in three consecutive runs on 2026-08-29. Through July only "none" was honored. the related finding class →
SambaNova · overcharge-repro: 2.18× average overcharge (ten distinct billed values for one request) → 1.00× (one billed value, 40 of 40 calls)
The meter now bills identical bytes identically. The co-fired closed pair - which through July billed the large member's count on both members - billed each member its own count in three of three rounds on 2026-08-29. The billing cell returns to a pass.
SambaNova · tool-choice: 400 post-hoc on required + forced-specific → PASS (all three conditions enforced)
The provider-level 400 on constraint violation no longer occurs; all three tool_choice conditions passed in four runs across 2026-08-25 and 2026-08-29.
SambaNova · tool-image: 500 (deterministic) → WORKS
Images in tool results now reach the model - perception-judged (the model names the image's color) in four runs across 2026-08-25 and 2026-08-29. the finding →
Novita · tool-image: REJECT 422 → WORKS
Images in tool results now reach the model - perception-judged in four runs across 2026-08-25 and 2026-08-29. the finding →
AWS Bedrock (Chat Completions), AWS Bedrock (Responses) · tool-emit: stray one-word content beside the tool call → clean (no content beside the call)
Both Bedrock APIs stopped emitting the stray one-word content next to tool calls (clean in four runs per API through 2026-08-29). The argument un-escaping is unchanged and the finding now covers it alone. the finding →
Novita · cot-carry: CARRIES → DROPS
Replayed reasoning no longer reaches the model: the canary planted in the replayed reasoning field is dropped and the continuation invents instead of recalling. Reproduced on the 2026-08-25 full re-run. the finding →
Parasail · tool-choice: PASS (all three conditions enforced) → IGNORED (validated, unenforced)
"required" now returns prose under finish_reason "tool_calls" and a forced tool by name yields a call to a different tool, where the 2026-07-27 battery measured all three tool_choice conditions enforced. The field is still validated (invalid values 400), and the regression held across three runs through 2026-08-02. the finding →
Lightning · cot-carry: DROPS → CARRIES
Replayed reasoning now reaches the model: the 99-memory-test canary planted in the replayed reasoning field is recalled, where the 2026-07-06 battery measured invention (0/12 across both spellings). what the probe measures →
SambaNova · cot-carry: DROPS → CARRIES
Replayed reasoning now reaches the model, where the 2026-07-06 battery measured a drop. what the probe measures →
DeepInfra · cot-carry: DROPS → CARRIES
Replayed reasoning now reaches the model, where the 2026-07-06 battery measured a drop on both request spellings (0/2 each). what the probe measures →
SambaNova · tool-image: 404 (model delisted) → 500 (deterministic)
The verdict tracks the model listing: while the gemma-4 model id was delisted the probe returned 404; with the id listed again, images in tool results fail with the historic deterministic 500 (7/7 attempts). the finding →