GPT-6.1 Sol on Amazon Bedrock: Near-Astra performance at a fraction of the cost
OpenAI's GPT-6.1 Sol brings near-Astra performance to coding, computer use and professional work at one-fifth of Astra's standard token prices. Here's what Amazon Bedrock adds, from pricing and a 1M-token context to endpoints and code.

OpenAI has introduced GPT-6.1 Sol, a new model designed to bring much of GPT-6 Astra’s capability to workloads where cost matters.
The headline is simple: OpenAI says GPT-6.1 Sol approaches GPT-6 Astra on agentic coding, computer use, and professional work at one-fifth of Astra’s standard API input and output token prices.
That alone is interesting.
But the more important question for engineers is what happens when you put that model inside a production agent.
A real agent doesn’t make one model call and disappear. It reads files, searches for information, calls tools, writes code, tests the result, recovers from failures, and sometimes repeats the whole process.
Every additional step costs money.
So a model that can deliver near-frontier capability at a much lower token price can change the economics of the entire workflow.
And GPT-6.1 Sol is now generally available through Amazon Bedrock, with AWS identity, networking, auditing, and security controls around the inference layer.
That combination makes Sol much more interesting than just another model release.
Why GPT-6.1 Sol matters for AI agents
Think about a software-engineering agent working on an unfamiliar repository.
The agent may have to:
Understand the repository
↓
Find relevant files
↓
Trace dependencies
↓
Plan the change
↓
Write code
↓
Run tests
↓
Inspect failures
↓
Fix the implementation
↓
Run tests again
A small difference in model capability or pricing can become a large difference across the full task.
OpenAI reports that GPT-6.1 Sol matches GPT-6 Astra on DeepSWE v1.1, a benchmark for long-horizon software-engineering tasks, while exceeding GPT-6 Sol’s best result by 6.4 percentage points at a lower reasoning effort and cost.
The same pattern appears in other evaluations.
On AutomationBench, GPT-6.1 Sol scores 2.2 percentage points above Opus 5.5 at medium reasoning effort, while costing roughly one-third as much on the evaluated task. OpenAI also reports a 4.8-point improvement over GPT-6 Sol at the same reasoning setting.
On OSWorld 2.0, GPT-6.1 Sol improves by seven percentage points over GPT-6 Sol at maximum reasoning effort and comes within 2.1 percentage points of Astra’s score, according to OpenAI.
For AI engineers, that creates an interesting question:
How much of your agent’s workload really needs your most expensive model?
GPT-6.1 Sol is effectively an answer to that question.
GPT-6.1 Sol pricing: the cost difference is the real story
OpenAI’s standard API pricing compares the GPT-6 family like this:
| Model | Input / 1M tokens | Cached input / 1M tokens | Output / 1M tokens |
|---|---|---|---|
| GPT-6 Astra | $10.00 | $1.00 | $50.00 |
| GPT-6.1 Sol | $2.00 | $0.10 | $10.00 |
| GPT-6 Luna | $0.10 | $0.01 | $0.50 |
That makes GPT-6.1 Sol approximately one-fifth the standard input and output token price of GPT-6 Astra.
Cached input is even more interesting.
GPT-6.1 Sol’s cached input price is $0.10 per million tokens, which OpenAI describes as 95% below its standard input price and 50% below GPT-6 Sol’s cached input price.
For an agent repeatedly working against the same repository, instructions, documents, or context, that difference can become significant.
A ten-token prompt isn’t where economics get interesting.
A million-token workload repeated across hundreds of agent interactions is.
Amazon Bedrock changes the deployment story
GPT-6.1 Sol became generally available on Amazon Bedrock on September 29, 2026.
AWS describes the Bedrock deployment as running on an inference engine designed for performance, security, and reliability at scale.
For enterprises already operating inside AWS, the important part isn’t just the model.
It’s everything around it.
Bedrock lets organizations control model access through AWS IAM, audit invocation activity through CloudTrail, and keep traffic inside their network boundaries using VPC endpoints powered by AWS PrivateLink.
AWS also states that inference runs on hardware-isolated infrastructure with zero-operator access, meaning AWS operators cannot access customer prompts and completions during inference. AWS says inference data isn’t used for model training, and using GPT-6.1 Sol doesn’t require customers to opt into sharing data with OpenAI.
There is one important nuance.
For automated abuse detection, AWS says classifier-flagged traffic can be retained for up to 30 days and processed programmatically. Customers can request zero data retention through their AWS account team.
For regulated or sensitive workloads, that distinction matters.
A 1M-token context window changes the type of workload you can build
One of the easy-to-miss details in the Bedrock model card is the 1 million token context window.
AWS lists GPT-6.1 Sol with a maximum context of 1M tokens and a maximum output of 131,072 tokens.
That is useful for workloads where the model has to carry a very large amount of context at once.
Think:
Large codebase
+
Architecture documents
+
Logs
+
Tickets
+
API documentation
+
Previous investigation
↓
GPT-6.1 Sol
But there’s an important pricing catch.
AWS applies higher long-context rates when the input exceeds 272,000 tokens. For commercial in-Region or US geographic cross-Region inference, the short-context rates are $2.20 input and $11.00 output per million tokens, while the long-context rates are $4.40 input and $16.50 output.
Note that these Bedrock rates sit slightly above OpenAI’s own API list prices in the table above, so use the Bedrock numbers when you budget an AWS deployment.
So the 1M-token context window isn’t a free license to throw your entire data warehouse at the model.
Context length is a capability.
It is still an economic decision.
How GPT-6.1 Sol is exposed through Bedrock
This is another area where the current Bedrock architecture matters.
AWS currently supports GPT-6.1 Sol through two Bedrock endpoints:
bedrock-mantle
bedrock-runtime
On bedrock-mantle, GPT-6.1 Sol is available in Region in us-east-1 using the model ID:
openai.gpt-6.1-sol
On bedrock-runtime, direct in-Region invocation isn’t supported for this launch. Instead, AWS exposes the US geographic cross-Region inference profile:
us.openai.gpt-6.1-sol
AWS currently does not offer a global inference profile for GPT-6.1 Sol on bedrock-runtime.
That distinction is worth understanding before writing production routing logic. (It is a different shape from Grok 4.7 on Bedrock, which has no in-Region option at all and offers Geo and Global profiles instead.)
OpenAI-compatible APIs make the integration familiar
GPT-6.1 Sol supports both Responses and Chat Completions APIs on Bedrock.
For developers already using the OpenAI SDK, the interface therefore looks familiar.
For example, using the bedrock-runtime endpoint:
from openai import OpenAI
client = OpenAI()
response = client.responses.create(
model="us.openai.gpt-6.1-sol",
input="Explain the failure modes of a distributed data pipeline.",
max_output_tokens=512,
)
print(response.output_text)
The corresponding environment variables are:
export OPENAI_API_KEY="<your-bedrock-api-key>"
export OPENAI_BASE_URL="https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1"
The same model can also be called through bedrock-mantle using:
openai.gpt-6.1-sol
with:
export OPENAI_BASE_URL="https://bedrock-mantle.us-east-1.api.aws/openai/v1"
AWS documents both approaches in the GPT-6.1 Sol model guide.
That compatibility is valuable because it allows existing OpenAI-style application code to move toward Bedrock without forcing developers to learn an entirely different request model.
Prompt caching: useful, but understand the distinction
There is a subtle detail here that is easy to miss.
GPT-6.1 Sol has cache pricing dimensions, including a very low cached-input price.
However, AWS currently marks explicit prompt caching as unsupported for this Bedrock model.
So it’s better not to casually equate the existence of a cache-read price with the availability of every prompt-caching feature.
That distinction matters when designing an application around repeated context.
The general lesson is simple:
Always check both pricing and feature support.
A price table tells you what something costs.
It doesn’t necessarily tell you how the feature is exposed.
Where GPT-6.1 Sol fits in a model-routing layer
I wouldn’t think of Sol as simply:
“A cheaper Astra.”
A better mental model is:
A high-capability model intended to handle a much larger share of everyday agentic work without paying frontier-model prices for every step.
That changes how you might design a model-routing layer.
For example:
Incoming task
│
▼
Task classification
│
┌───────────────┼────────────────┐
│ │ │
Simple Complex Frontier
task task task
│ │ │
▼ ▼ ▼
Luna Sol Astra
The exact boundaries should come from evaluation rather than assumptions.
But this architecture is attractive because not every step of an agent requires the same model.
A data-extraction step might not need Astra.
A code-analysis step may be well suited to Sol.
A genuinely difficult scientific-research task may still justify Astra.
That is the kind of model routing strategy that can improve both cost efficiency and overall system design.
The Agent Toolkit for AWS makes the AWS side easier
AWS has also introduced the Agent Toolkit for AWS, which connects coding agents to AWS documentation, APIs, and services through MCP and a curated set of skills.
The toolkit supports agents including:
- Kiro
- Cursor
- Claude Code
- Codex
With AWS CLI 2.35.0 or later, the interactive setup can detect installed agents, install default AWS skills, and configure the AWS MCP Server:
aws configure agent-toolkit
AWS says the toolkit can then discover relevant skills automatically based on the task instead of requiring the developer to manually select a skill each time.

How the pieces fit: your coding agent works through the Agent Toolkit for AWS and AWS services, with GPT-6.1 Sol running on Amazon Bedrock.
The toolkit can also be extended with specialized skills for areas such as:
AWS agents
AWS data analytics
AWS data lakes
ETL
API Gateway
AgentCore
That becomes particularly interesting when GPT-6.1 Sol is used inside an engineering workflow rather than simply as a chat model.
The model isn’t just generating code.
It can sit inside a workflow where an agent understands an AWS environment, looks up documentation, works with services, changes infrastructure, and validates its result.
Getting started with Agent Toolkit for AWS
The quickest path depends on the coding agent you’re using.
For example, Claude Code can install the AWS core plugin with:
/plugin marketplace add aws/agent-toolkit-for-aws
/plugin install aws-core@agent-toolkit-for-aws
/reload-plugins
For Codex, AWS documents:
codex plugin marketplace add aws/agent-toolkit-for-aws
For other MCP-compatible agents, AWS provides a direct MCP configuration path.
The prerequisites include uv, which AWS uses for the MCP proxy. IAM credentials are optional for documentation search and skill discovery, but are required when the agent needs to execute AWS API calls or scripts.
This is a small but useful development pattern:
Coding Agent
│
▼
Agent Toolkit
│
├── AWS documentation
├── AWS skills
├── AWS MCP Server
└── AWS APIs
│
▼
Your AWS environment
The result is less manual context switching between an AI coding agent, AWS documentation, the CLI, and the console.
The biggest limitation: Sol isn’t Astra
The pricing advantage shouldn’t obscure the capability boundary.
GPT-6.1 Sol is close to Astra on many evaluations, but OpenAI explicitly reports that GPT-6 Astra remains the highest-scoring model on Terminal-Bench Science 0.1, with a score of 68.1%. OpenAI recommends Astra for the most difficult scientific-research tasks.
That’s an important distinction.
“Near-Astra” doesn’t mean “Astra replacement.”
It means that for a large class of real workloads, the performance gap may be small enough that the economics become much more attractive.
And in production, economics matter.
An agent that costs five times as much isn’t automatically five times better.
Benchmark results need context
There is another reason to avoid treating benchmark numbers as universal truths.
OpenAI explicitly states that its evaluations were performed in its research environment or through its API. It also notes that competitor-model results were taken from publicly available reports. Differences in system prompts, available tools, reasoning effort, and environments can therefore affect the comparison.
The factuality benchmark has another important limitation.
OpenAI says the test uses deliberately difficult, de-identified conversations in which users had previously flagged factual errors. Those prompts are therefore not representative of typical usage. GPT-6.1 Sol reduced the share of responses containing a factual error from 11.4% to 7.7% at low reasoning effort, but that should be read as a benchmark result rather than a general real-world error rate.
That’s why I would treat the published benchmark numbers as a starting point.
For production decisions, your own workload should be the final benchmark.
What data and AI engineers should do
For an engineering team evaluating GPT-6.1 Sol, I’d start with three questions.
1. What percentage of your workload really needs Astra-level capability?
Take your real agent traces and classify them.
Simple extraction → cheaper model
Routine reasoning → Sol
Complex coding → Sol
Hard research → Astra
Don’t choose the model globally.
Choose it per workload class.
2. How much context are your agents carrying?
If a workflow routinely exceeds 272K input tokens, pricing changes materially on Bedrock.
That means context management becomes an architecture problem.
Chunking, retrieval, summaries, state management, and context reuse still matter even when a model supports a million-token context window.
3. Can your infrastructure take advantage of AWS-native controls?
If the answer is yes, Bedrock gives you:
IAM
CloudTrail
PrivateLink
VPC networking
Hardware-isolated inference
AWS governance
That can make the deployment story much easier for teams already standardized on AWS.
The engineering takeaway
GPT-6.1 Sol is interesting because it moves the optimization problem up one level.
The question is no longer simply:
Which model is smartest?
It becomes:
Which model is smart enough for this step, at the lowest total cost for completing the task?
That’s a fundamentally different way to design an AI system.
An agent may make dozens of model calls.
The best architecture isn’t necessarily the one that uses the smartest model every time.
It is the one that knows when intelligence matters most.
GPT-6.1 Sol looks particularly well positioned for that middle ground: high-end reasoning for coding, documents, computer use, and multi-step workflows, without paying GPT-6 Astra pricing for every interaction.
And with Amazon Bedrock now providing the model alongside AWS governance, networking, and security controls, Sol can be treated not just as an API endpoint, but as another building block in a production AI stack.
Final takeaways
- GPT-6.1 Sol targets near-Astra performance at one-fifth of Astra’s standard input and output token prices.
- The model has a 1M-token context window, with higher long-context pricing beyond 272K input tokens.
- Cached input is priced at $0.10 per million tokens, although explicit prompt caching is currently not supported on the Bedrock model.
bedrock-mantleprovides in-Region access inus-east-1, whilebedrock-runtimecurrently exposes the US geographic cross-Region profile.- Responses and Chat Completions APIs are supported, making OpenAI-style integration straightforward.
- Priority and Flex service tiers are not supported for GPT-6.1 Sol on Bedrock.
- AWS provides IAM, CloudTrail, PrivateLink, and hardware-isolated inference around the model.
- GPT-6 Astra still leads on the hardest scientific-research benchmark, so Sol should be viewed as a highly capable cost-efficient model, not a universal Astra replacement.
- The right way to evaluate Sol is against your own agent traces, task success rates, latency, context size, and total cost per completed task.
Sources
- OpenAI: Introducing GPT-6.1 Sol (September 29, 2026)
- AWS Machine Learning Blog: Bring near-Astra intelligence to everyday work with GPT-6.1 Sol on Amazon Bedrock (September 29, 2026)
- AWS Documentation: GPT-6.1 Sol – Amazon Bedrock model card
- AWS Documentation: Getting started – Agent Toolkit for AWS
- AWS What’s New: OpenAI GPT-6.1 Sol is now generally available on Amazon Bedrock (September 29, 2026)
Discussion
Join the discussion: create a free account or sign in to comment.
No comments yet.