# AWS - gemma-4-31b grades

2 graded endpoints serving gemma-4-31b.
HTML version: [https://inferencecanary.com/aws](https://inferencecanary.com/aws) · matrix: [https://inferencecanary.com/index.md](https://inferencecanary.com/index.md)

A vendor is graded once per API it speaks, because a shim and a native protocol routinely
behave differently - AWS speaks 2.


## Endpoints

| Endpoint | Dialect | Inference | Billing | Probed |
|---|---|---|---|---|
| [AWS Bedrock (Chat Completions)](https://inferencecanary.com/aws/gemma-4-31b.md) | OpenAI chat/completions on Bedrock Mantle | D | OK | 2026-08-29 |
| [AWS Bedrock (Responses)](https://inferencecanary.com/aws-responses/gemma-4-31b.md) | OpenAI Responses on Bedrock Mantle | D− | OK | 2026-08-29 |


## Billing

Footnote (customer-favoring, not flagged): Reports reasoning_tokens: 0 beside delivered reasoning - under-counts in the customer's favor.

## Findings affecting AWS

- **[high · inference] Several providers drop replayed reasoning: the model never sees its own prior thinking** - details: [/replayed-reasoning-dropped.md](https://inferencecanary.com/replayed-reasoning-dropped.md)
- **[high · inference] Images in tool results: rejected or silently blinded on several providers** - details: [/tool-result-images.md](https://inferencecanary.com/tool-result-images.md)
- **[medium · inference] AWS Bedrock never emits parallel tool calls** - details: [/bedrock-parallel-tool-call.md](https://inferencecanary.com/bedrock-parallel-tool-call.md)
- **[medium · inference] AWS Bedrock un-escapes tool-call arguments in transit** - details: [/bedrock-argument-corruption.md](https://inferencecanary.com/bedrock-argument-corruption.md)
- **[medium · inference] AWS Bedrock (Responses) rejects temperature** - details: [/aws-responses-temperature.md](https://inferencecanary.com/aws-responses-temperature.md)

