# Changes - inferencecanary

The changelog of measured change: an entry exists only when an eval's output moved from one
measured value to another - both dated, both backed by evidence. New providers or models
joining the sheet don't appear here; a provider's behavior changing does.
HTML version: [https://inferencecanary.com/changes](https://inferencecanary.com/changes)

## novita · tool-choice: IGNORED (required + forced-specific) → PASS (all three conditions enforced)

Measured 2026-07-26 → re-measured 2026-08-29.

Healed in two steps: forced-specific was honored again on 2026-08-25 with "required" still ignored, and all three conditions passed in three consecutive runs on 2026-08-29. Through July only "none" was honored.

- [the related finding class](https://inferencecanary.com/findings#tool-choice-unenforced)

## sambanova · overcharge-repro: 2.18× average overcharge (ten distinct billed values for one request) → 1.00× (one billed value, 40 of 40 calls)

Measured 2026-07-26 → re-measured 2026-08-25.

The meter now bills identical bytes identically. The co-fired closed pair - which through July billed the large member's count on both members - billed each member its own count in three of three rounds on 2026-08-29. The billing cell returns to a pass.

## sambanova · tool-choice: 400 post-hoc on required + forced-specific → PASS (all three conditions enforced)

Measured 2026-07-26 → re-measured 2026-08-25.

The provider-level 400 on constraint violation no longer occurs; all three tool_choice conditions passed in four runs across 2026-08-25 and 2026-08-29.

## sambanova · tool-image: 500 (deterministic) → WORKS

Measured 2026-07-26 → re-measured 2026-08-25.

Images in tool results now reach the model - perception-judged (the model names the image's color) in four runs across 2026-08-25 and 2026-08-29.

- [the finding](https://inferencecanary.com/tool-result-images)

## novita · tool-image: REJECT 422 → WORKS

Measured 2026-07-26 → re-measured 2026-08-25.

Images in tool results now reach the model - perception-judged in four runs across 2026-08-25 and 2026-08-29.

- [the finding](https://inferencecanary.com/tool-result-images)

## aws, aws-responses · tool-emit: stray one-word content beside the tool call → clean (no content beside the call)

Measured 2026-07-26 → re-measured 2026-08-25.

Both Bedrock APIs stopped emitting the stray one-word content next to tool calls (clean in four runs per API through 2026-08-29). The argument un-escaping is unchanged and the finding now covers it alone.

- [the finding](https://inferencecanary.com/bedrock-argument-corruption)

## novita · cot-carry: CARRIES → DROPS

Measured 2026-07-26 → re-measured 2026-08-02.

Replayed reasoning no longer reaches the model: the canary planted in the replayed reasoning field is dropped and the continuation invents instead of recalling. Reproduced on the 2026-08-25 full re-run.

- [the finding](https://inferencecanary.com/replayed-reasoning-dropped)

## parasail · tool-choice: PASS (all three conditions enforced) → IGNORED (validated, unenforced)

Measured 2026-07-27 → re-measured 2026-08-01.

"required" now returns prose under finish_reason "tool_calls" and a forced tool by name yields a call to a different tool, where the 2026-07-27 battery measured all three tool_choice conditions enforced. The field is still validated (invalid values 400), and the regression held across three runs through 2026-08-02.

- [the finding](https://inferencecanary.com/tool-choice-unenforced)

## lightning · cot-carry: DROPS → CARRIES

Measured 2026-07-06 → re-measured 2026-07-26.

Replayed reasoning now reaches the model: the 99-memory-test canary planted in the replayed reasoning field is recalled, where the 2026-07-06 battery measured invention (0/12 across both spellings).

- [what the probe measures](https://inferencecanary.com/replayed-reasoning-dropped)

## sambanova · cot-carry: DROPS → CARRIES

Measured 2026-07-06 → re-measured 2026-07-26.

Replayed reasoning now reaches the model, where the 2026-07-06 battery measured a drop.

- [what the probe measures](https://inferencecanary.com/replayed-reasoning-dropped)

## deepinfra · cot-carry: DROPS → CARRIES

Measured 2026-07-06 → re-measured 2026-07-26.

Replayed reasoning now reaches the model, where the 2026-07-06 battery measured a drop on both request spellings (0/2 each).

- [what the probe measures](https://inferencecanary.com/replayed-reasoning-dropped)

## sambanova · tool-image: 404 (model delisted) → 500 (deterministic)

Measured 2026-07-06 → re-measured 2026-07-26.

The verdict tracks the model listing: while the gemma-4 model id was delisted the probe returned 404; with the id listed again, images in tool results fail with the historic deterministic 500 (7/7 attempts).

- [the finding](https://inferencecanary.com/findings#tool-result-images)


