Grok 4.7 on Amazon Bedrock: Routing, reasoning effort, and cost with OpenAI-compatible APIs

xAI's Grok 4.7 is now on Amazon Bedrock with a 500K-token context window. Here's how reasoning effort, Geo vs Global routing and service tiers change cost and latency, with OpenAI-compatible code you can reuse.

xAI’s Grok 4.7 is now available on Amazon Bedrock, bringing a 500K-token context window, configurable reasoning, image input, and support for long-running agentic workloads.

At first glance, this looks like another model joining the growing Bedrock catalog.

But there is a more interesting engineering story underneath.

Grok 4.7 gives you several knobs that directly affect how an application behaves in production:

  • Geo vs Global cross-Region routing
  • Four reasoning-effort levels
  • Standard, Priority, and Flex service tiers
  • OpenAI-compatible Responses and Chat Completions APIs
  • Bedrock’s native Converse API

Those aren’t just configuration options. They affect cost, latency, throughput, data residency, and how much reasoning context your agent can carry between turns.

For anyone building agentic pipelines, that’s where Grok 4.7 gets interesting.

The short version

If you only want the important bits:

Capability Grok 4.7 on Bedrock
Context window 500K tokens
Reasoning low, medium, high, xhigh
Default reasoning high
Cross-Region routing Geo + Global
Input Text + image
Output Text
OpenAI-compatible APIs Responses + Chat Completions
Native AWS API Converse
Standard Geo price $2.20 / 1M input tokens
Standard Global price $2.00 / 1M input tokens
Standard Geo output $6.60 / 1M output tokens
Standard Global output $6.00 / 1M output tokens

AWS lists Grok 4.7 as an active model with a launch date of September 28, 2026 and an EOL date no sooner than September 28, 2027.

Infographic: a Grok 4.7 request on Amazon Bedrock passes through reasoning effort (low to xhigh), routing (Geo US or Global), and service tier (Priority, Standard, Flex), with example workload configurations

The whole decision path on one page: reasoning effort, routing and service tier, with example workload settings.

Why this matters for AI engineers

The interesting change isn’t simply that another frontier model is available through Bedrock.

It’s that reasoning has become an explicit production-control problem.

AWS exposes four reasoning levels:

low → medium → high → xhigh

And reasoning is enabled by default, with high as the default effort level.

That means a production application shouldn’t blindly send every request using the default.

A simple classification request doesn’t necessarily need the same amount of reasoning as a multi-step coding agent.

Consider a pipeline like this:

User request
     │
     ▼
Intent classification
     │
     ▼
Planning agent
     │
     ▼
Tool execution
     │
     ▼
Verification
     │
     ▼
Final response

Using xhigh everywhere could be wasteful.

A more deliberate strategy might look like:

Classification       → low
Simple extraction    → low
Normal reasoning     → medium
Planning             → high
Complex agent loop   → xhigh

The exact mapping should be benchmarked against your workload, but the principle is important:

Reasoning effort should become part of your routing policy, not a value you leave at its default.

AWS recommends benchmarking the reasoning levels against your workload to determine where additional reasoning stops paying for itself.

Grok 4.7 is also more expensive when it thinks longer

This becomes particularly important for agentic systems.

Artificial Analysis reports an Intelligence Index of 46 for Grok 4.7 versus 44 for Grok 4.6, while its Coding Agent score increased from 47 to 56.

But there’s a catch.

The evaluation also reports approximately 81K output tokens per Intelligence Index task for Grok 4.7 versus roughly 36K for Grok 4.6, with Grok 4.7 measured at xhigh reasoning effort and Grok 4.6 at high.

That is more than twice the output tokens per task.

And this is exactly why reasoning effort matters in production.

A single additional token might not matter much.

An agent making 50 model calls absolutely can.

If every step becomes more expensive, the economics of the entire agent changes.

Geo vs Global: you’re choosing a routing policy

One subtle point about Grok 4.7 on Bedrock is that it isn’t exposed through a normal in-Region model invocation.

AWS currently exposes two cross-Region inference profiles:

us.xai.grok-4.7
global.xai.grok-4.7

The us profile is the Geo option, while global is the Global option.

The distinction matters.

Geo

Geo routing keeps inference within the supported geographic boundary.

For Grok 4.7, AWS currently provides a US Geo profile.

Its Standard pricing is:

  • Input: $2.20 / 1M tokens
  • Output: $6.60 / 1M tokens
  • Cache read: $0.55 / 1M tokens

Global

Global routing can use supported commercial AWS Regions worldwide.

Its Standard pricing is:

  • Input: $2.00 / 1M tokens
  • Output: $6.00 / 1M tokens
  • Cache read: $0.50 / 1M tokens

So Global is about 9% cheaper than Geo on both input and output under Standard pricing.

The trade-off is control.

With Global routing, you get access to a larger pool of capacity, but you have less control over exactly where an individual request is processed.

That makes the choice less about:

“Which AWS Region should I use?”

and more about:

“What routing and residency policy does this workload require?”

That’s a much more useful way to think about it.

There is another cost knob: Service Tier

Routing isn’t the only thing affecting the bill.

Bedrock also exposes three service tiers for Grok 4.7:

Standard
Priority
Flex

Standard is the normal pay-per-token option.

Priority provides prioritized processing at a premium.

Flex is designed for workloads where latency isn’t critical and is substantially cheaper. AWS documents Priority at 1.75× Standard pricing and Flex at 0.5× Standard pricing.

You choose the tier per request with the service_tier parameter:

response = client.responses.create(
    model="global.xai.grok-4.7",
    service_tier="flex",  # "default" (Standard), "priority" or "flex"
    reasoning={"effort": "medium"},
    input="Group these 40 support tickets into themes.",
)

That gives us an interesting three-dimensional decision:

                 Reasoning
                low → xhigh
                    │
                    │
Routing ────────────┼──────────── Service tier
Geo / Global        │             Flex / Standard / Priority

For example:

Interactive coding assistant

Geo
high reasoning
Priority

might make sense when latency and geographic constraints matter.

Whereas:

Overnight document-processing pipeline

Global
medium reasoning
Flex

could be a much more economical choice.

The point isn’t that these are universally correct configurations.

The point is that your AI infrastructure can now make these decisions explicitly.

OpenAI-compatible APIs make migration easier

One of the most useful aspects of the Bedrock integration is that you don’t necessarily have to rewrite an existing application around a completely different API.

Grok 4.7 supports:

  • Responses API
  • Chat Completions API
  • Converse
  • InvokeModel

The Responses and Chat Completions APIs are exposed through an OpenAI-compatible endpoint.

For example:

from openai import OpenAI

client = OpenAI()

response = client.responses.create(
    model="us.xai.grok-4.7",
    input="Explain how a distributed feature store works."
)

print(response.output_text)

The base URL points at Bedrock:

export OPENAI_API_KEY="<your-bedrock-api-key>"
export OPENAI_BASE_URL="https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1"

The important detail is that the model isn’t simply xai.grok-4.7.

You explicitly select the inference profile:

model="us.xai.grok-4.7"

or:

model="global.xai.grok-4.7"

AWS explicitly notes that Grok 4.7 isn’t available for in-Region inference on this endpoint.

Controlling reasoning effort

With the Responses API, reasoning can be configured directly:

response = client.responses.create(
    model="us.xai.grok-4.7",
    reasoning={
        "effort": "high"
    },
    input="Design a fault-tolerant data ingestion architecture."
)

print(response.output_text)

The supported levels are:

low
medium
high
xhigh

AWS documents high as the default.

This is something I would explicitly configure in production rather than relying on the default.

For example:

def choose_effort(task_type: str) -> str:
    if task_type in {"classification", "extraction"}:
        return "low"

    if task_type in {"analysis", "summarization"}:
        return "medium"

    if task_type in {"planning", "coding"}:
        return "high"

    if task_type in {"complex_agent"}:
        return "xhigh"

    return "medium"

That tiny piece of routing logic can eventually become part of a much larger model gateway.

Responses API vs Chat Completions

There is an important difference between the two interfaces.

The Chat Completions API doesn’t return the model’s reasoning tokens.

The Responses API can return encrypted reasoning content, which can then be supplied back to the model during subsequent turns.

For example:

response = client.responses.create(
    model="us.xai.grok-4.7",
    reasoning={"effort": "high"},
    include=["reasoning.encrypted_content"],
    input="Analyze this architecture and identify the three biggest failure modes."
)

This matters for long-running agents.

An agent isn’t necessarily just:

prompt → answer

It can be:

observe
   ↓
reason
   ↓
act
   ↓
observe
   ↓
reason
   ↓
verify
   ↓
act

Being able to preserve encrypted reasoning context across turns can therefore become useful when building multi-step workflows.

Where I would actually use each setting

Here’s the way I would initially approach a production workload.

Workload Reasoning Routing Service tier
Classification Low Global Flex
Metadata extraction Low Global Flex
Document analysis Medium Global Standard
Code generation High Geo/Global Standard
Complex debugging High Geo/Global Standard
Long-running agent High/Xhigh Global Standard
Interactive agent High Geo Priority
Batch research Medium/High Global Flex

Again, these are starting hypotheses, not universal recommendations.

The correct configuration should come from benchmarking your own workload.

That is probably the most important lesson here.

The engineering takeaway

Grok 4.7 on Bedrock isn’t particularly interesting simply because another frontier model has become available.

What’s more interesting is the control surface around it.

You can now think about an AI application as a routing problem:

                    ┌──────────────┐
                    │   Workload   │
                    └──────┬───────┘
                           │
                 ┌─────────▼─────────┐
                 │ Task classification│
                 └─────────┬─────────┘
                           │
              ┌────────────┼────────────┐
              │            │            │
             Low        Medium       High/Xhigh
              │            │            │
              └────────────┼────────────┘
                           │
                  ┌────────▼────────┐
                  │ Routing policy  │
                  │ Geo / Global    │
                  └────────┬────────┘
                           │
                  ┌────────▼────────┐
                  │ Service tier    │
                  │ Flex/Std/Prior. │
                  └────────┬────────┘
                           │
                    ┌──────▼──────┐
                    │  Grok 4.7   │
                    └─────────────┘

For small applications, you can ignore most of this.

For an agentic system making hundreds or thousands of model calls, you probably shouldn’t.

Reasoning effort affects token consumption.

Routing affects geography, capacity and price.

Service tier affects latency and price.

And together, those three decisions become part of the architecture.

That is where I think Grok 4.7 on Bedrock is worth paying attention to.

Final takeaways

  • Grok 4.7 has a 500K-token context window and four reasoning levels.
  • There is no in-Region option on Bedrock: you choose a Geo (US) or a Global cross-Region profile.
  • Global is cheaper than Geo under Standard pricing, but provides less geographic control.
  • Reasoning effort should be treated as a routing decision, especially for high-volume agentic workloads.
  • Flex, Standard and Priority provide another cost/latency dimension.
  • The OpenAI-compatible APIs make it relatively straightforward to bring existing applications to Bedrock.
  • For multi-turn agentic systems, the Responses API is particularly interesting because encrypted reasoning content can be carried between turns.
  • The right configuration isn’t something to guess. Benchmark it against your actual workload.

Sources

  1. AWS Machine Learning Blog: Grok 4.7 is now available on Amazon Bedrock
  2. AWS documentation: Grok 4.7 model card (pricing, inference profiles, service tiers, reasoning)
  3. AWS What’s New: Grok 4.7 on Amazon Bedrock
  4. Artificial Analysis: Grok 4.7 (xhigh) vs Grok 4.6 (high)

Prices and availability are as published by AWS on 5 October 2026. Check the model card before you commit to a configuration.

Discussion

Join the discussion: create a free account or sign in to comment.

No comments yet.