Google (OpenAI-compat)

OpenAI chat/completions shim over generateContent · serves gemma-4-31b-it as gemma-4-31b · probed live 2026-08-25 · back to the matrix

Google's OpenAI-compatibility shim, serving the same model as the native provider - graded on what the shim delivers.

Inference
F
A − 4 high − 1 medium severity, per the ladder; flags below.
Billing
Reported usage consistent with bytes on the wire. Footnote: completion_tokens under-counts its own delivered content - under-bills in the customer's favor.
Caching
Fail
A repeat 20K-token submission waits out the free tier's per-minute token quota (~46s of 429 retries) before it is served - and gemma-4 has no paid tier, so there is no quota to buy. Whether context is reused is unobservable behind the wait. Speed numbers →

Flags

Every flag links to its finding - evidence, repro, and disclosure records live there; the deepest findings have full writeup pages.