# SambaNova - gemma-4-31b grade card

OpenAI chat/completions · serves `gemma-4-31B-it` as gemma-4-31b · probed live 2026-08-29

HTML version: [https://inferencecanary.com/sambanova/gemma-4-31b](https://inferencecanary.com/sambanova/gemma-4-31b) · matrix: [https://inferencecanary.com/index.md](https://inferencecanary.com/index.md)

## Grades


| Axis | Grade | Basis |
|---|---|---|
| Inference | A− | A − 0 high − 1 medium severity, per the published ladder |
| Billing | OK | Reported usage consistent with bytes on the wire. |
| Caching | No (0/8 hits) | No cold→warm trial hit - every call re-reads (and re-bills) the full context. Speed numbers: [https://inferencecanary.com/speed.md](https://inferencecanary.com/speed.md) |


## Flags

- **[medium · inference] SambaNova ships raw chain-of-thought as the answer when truncation lands mid-thought** - With thinking on, a max_tokens cut that lands inside the thought returns the entire raw thought channel - literal template marker included - in the content field, with the reasoning field empty and no answer at all. Deterministic, three of three identical runs; a cut landing mid-answer splits cleanly, so only the mid-thought cut leaks. (details: [/sambanova-truncation-leak.md](https://inferencecanary.com/sambanova-truncation-leak.md))

