Together AI
OpenAI chat/completions · serves google/gemma-4-31B-it as gemma-4-31b ·
probed live 2026-08-25 · back to the matrix
Billing
✓
Reported usage consistent with bytes on the wire.
Caching
Sometimes
Reused already-seen context on 2/3 cold→warm trials - a cache exists but a hit is not dependable. Speed numbers →
Flags
-
high inference
Several providers drop replayed reasoning: the model never sees its own prior thinking
The gemma-4 chat template renders reasoning replayed on model turns, and the model card mandates preserving thinking across tool-call turns - the reference provider reproduces this live. These providers delete it before...
Every flag links to its finding - evidence, repro, and disclosure records live there; the deepest findings have full writeup pages.