Changes

The changelog of measured change: an entry exists only when an eval's output moved from one measured value to another - both dated, both backed by evidence. New providers or models joining the sheet don't appear here; a provider's behavior changing does. Providers ship changes silently - this page is where those changes stop being silent.

Novita · tool-choice: IGNORED (required + forced-specific)PASS (all three conditions enforced)

measured 2026-07-26 → re-measured 2026-08-29

Healed in two steps: forced-specific was honored again on 2026-08-25 with "required" still ignored, and all three conditions passed in three consecutive runs on 2026-08-29. Through July only "none" was honored. the related finding class →

SambaNova · overcharge-repro: 2.18× average overcharge (ten distinct billed values for one request)1.00× (one billed value, 40 of 40 calls)

measured 2026-07-26 → re-measured 2026-08-25

The meter now bills identical bytes identically. The co-fired closed pair - which through July billed the large member's count on both members - billed each member its own count in three of three rounds on 2026-08-29. The billing cell returns to a pass.

SambaNova · tool-choice: 400 post-hoc on required + forced-specificPASS (all three conditions enforced)

measured 2026-07-26 → re-measured 2026-08-25

The provider-level 400 on constraint violation no longer occurs; all three tool_choice conditions passed in four runs across 2026-08-25 and 2026-08-29.

SambaNova · tool-image: 500 (deterministic)WORKS

measured 2026-07-26 → re-measured 2026-08-25

Images in tool results now reach the model - perception-judged (the model names the image's color) in four runs across 2026-08-25 and 2026-08-29. the finding →

Novita · tool-image: REJECT 422WORKS

measured 2026-07-26 → re-measured 2026-08-25

Images in tool results now reach the model - perception-judged in four runs across 2026-08-25 and 2026-08-29. the finding →

AWS Bedrock (Chat Completions), AWS Bedrock (Responses) · tool-emit: stray one-word content beside the tool callclean (no content beside the call)

measured 2026-07-26 → re-measured 2026-08-25

Both Bedrock APIs stopped emitting the stray one-word content next to tool calls (clean in four runs per API through 2026-08-29). The argument un-escaping is unchanged and the finding now covers it alone. the finding →

Novita · cot-carry: CARRIESDROPS

measured 2026-07-26 → re-measured 2026-08-02

Replayed reasoning no longer reaches the model: the canary planted in the replayed reasoning field is dropped and the continuation invents instead of recalling. Reproduced on the 2026-08-25 full re-run. the finding →

Parasail · tool-choice: PASS (all three conditions enforced)IGNORED (validated, unenforced)

measured 2026-07-27 → re-measured 2026-08-01

"required" now returns prose under finish_reason "tool_calls" and a forced tool by name yields a call to a different tool, where the 2026-07-27 battery measured all three tool_choice conditions enforced. The field is still validated (invalid values 400), and the regression held across three runs through 2026-08-02. the finding →

Lightning · cot-carry: DROPSCARRIES

measured 2026-07-06 → re-measured 2026-07-26

Replayed reasoning now reaches the model: the 99-memory-test canary planted in the replayed reasoning field is recalled, where the 2026-07-06 battery measured invention (0/12 across both spellings). what the probe measures →

SambaNova · cot-carry: DROPSCARRIES

measured 2026-07-06 → re-measured 2026-07-26

Replayed reasoning now reaches the model, where the 2026-07-06 battery measured a drop. what the probe measures →

DeepInfra · cot-carry: DROPSCARRIES

measured 2026-07-06 → re-measured 2026-07-26

Replayed reasoning now reaches the model, where the 2026-07-06 battery measured a drop on both request spellings (0/2 each). what the probe measures →

SambaNova · tool-image: 404 (model delisted)500 (deterministic)

measured 2026-07-06 → re-measured 2026-07-26

The verdict tracks the model listing: while the gemma-4 model id was delisted the probe returned 404; with the id listed again, images in tool results fail with the historic deterministic 500 (7/7 attempts). the finding →