SambaNova
OpenAI chat/completions · serves gemma-4-31B-it as gemma-4-31b ·
probed live 2026-08-29 · back to the matrix
Billing
✓
Reported usage consistent with bytes on the wire.
Caching
No
0/8 cold→warm trials hit - every call re-reads (and re-bills) the full context. Speed numbers →
Flags
-
medium inference
SambaNova ships raw chain-of-thought as the answer when truncation lands mid-thought
With thinking on, a max_tokens cut that lands inside the thought returns the entire raw thought channel - literal template marker included - in the content field, with the reasoning field empty and no answer at all....
Every flag links to its finding - evidence, repro, and disclosure records live there; the deepest findings have full writeup pages.