GPT-6 Astra is OpenAI's latest frontier model, designed for complex, multi-step work across computer use, software engineering, science, cybersecurity, and professional workflows. OpenAI describes Astra as combining advances in pre-training, reinforcement learning, and alignment.
- Strong performance on abstract reasoning and long-context evaluations.
- Improved ability to execute multi-step workflows and produce documents, spreadsheets, presentations, and analyses.
- Can operate software interfaces, fill forms, update CRM records, organize calendars, and troubleshoot issues visible on screen.
- Available through ChatGPT plans, OpenAI API, Microsoft Azure and AWS Bedrock with rollout beginning to a limited set of organisations.
All image credits goes to OpenAI
Comparison to Other GPT models
The GPT model family has evolved gradually with each version improving reasoning, coding, and tool interaction capabilities.
| Model | Key Focus | Major Improvements |
|---|---|---|
| GPT-4o | Multimodality in a Single Unified Neural Network Architecture | Enabling low-latency, real-time voice conversations and real-world camera inputs. |
| GPT-5 | Reinforcement Learning Driven Logic | Pioneered hidden chain-of-thought (CoT) reasoning tokens, allowing the model to "think" through math, logic, and scientific queries before formulating an output. |
| GPT-5.5 | Long Context Grounding & Instruction Following | Expanded reliable alignment, ensuring multi-turn inputs or side requests do not skew the model away from its original goal. |
| GPT-5.6 sol/Terra | Multi-Step Agentic Workflows | Introduced unified multi-agent sub-orchestration workflows. |
| GPT-6 Astra | Full Autonomy & Multi-Modal Generation | Computer Use & 3D spatial awareness, context window persistence, extreme reasoning breakthrough, and cybersecurity capability. |
Key Improvements in GPT-6 Astra
Computer Use and Browser Interaction
Astra extends model capability from producing answers to carrying out actions in software environments. OpenAI highlights speed, accuracy, and safety improvements for tasks that previously required substantial manual interaction.
- Fill online forms and update customer records in CRM systems.
- Organize calendars and conduct online research.
- Draft summaries directly in email or document editors.
- Analyze scientific data, generate plots, create websites, and run frontend quality checks.
- Install and test software autonomously and troubleshoot problems visible on screen.
Computer Use Performance
| Benchmark | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|
| Agent's Last Exam | 59.3% | 53.6% |
| OSWorld 2.0 (offlineset, partial score) | 72.6% | 65.7% |
| ScreenSpot-Pro (no tools) | 92.7% | 76.9% |
Latency Result
- Scored 72.6% on OpenAI's OSWorld 2.0 latency simulation.
- Astra took roughly 40 minutes per task.
- GPT-5.6 Sol scored 65.7% at roughly 75 minutes per task.
- Astra was about 47% faster while achieving a higher performance score.
Professional Work and Artifact Creation
Trained for professional environments where useful output includes both reasoning and finished artifacts. OpenAI highlights its ability to follow existing templates and maintain a consistent visual and writing style.
- Produces structured documents, presentations, spreadsheets, and analyses.
- Adheres to existing templates rather than rebuilding outputs from scratch.
- Pulls only the context that matters into outputs instead of repeating unnecessary information.
- Maintains task orientation as requirements change or new instructions are added.
- Asks focused questions when missing information could change the outcome, while making sensible assumptions for routine gaps.
Selected Professional Benchmarks
| Benchmark | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|
| Automation Bench | 41.4% | 18.1% |
| BenchCAD | 95.9% | 83.3% |
| BrowseComp | 91.5% | 90.4% |
| Internal Design Tasks | 50.0% | 47.4% |
| Internal Data Science Tasks | 40.9% | 30.5% |
Coding and Software Engineering
OpenAI describes GPT-6 Astra as its best software-engineering model to date. The model is designed for agentic coding workflows, where it can execute iterations, verify behavior, and preserve relevant context across long sessions.
- Strong performance across terminal-based software engineering and coding-agent benchmarks.
- Higher reasoning effort can support more iterations on fresh builds and more verification through browser testing.
- Introduces persistent notes in Codex so accumulated details can survive context-window changes without repeatedly compressing all prior work into one summary.
- Earlier context windows remain searchable, allowing requirements and test results to be recovered even when they were not captured in notes.
- Astra can continue non-dependent work while waiting for answers to questions in Codex.
Coding Benchmark Results
| Benchmark | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|
| Terminal-Bench 4.0 | 57.9% | 37.3% |
| DeepSWE v1.1 | 74.1% | 72.7% |
| FrontierCode 1.1 Extended | 64.5% | 60.6% |
| FrontierCode 1.1 Main | 53.3% | 47.5% |
| Internal Database Migration Tasks | 63.9% | 42.7% |
Scientific Discovery and Research
Combines scientific reasoning with computer use, allowing it to work directly in specialized software. OpenAI positions this capability for inspecting data, exploring results, and helping researchers decide what to investigate next.
- Supports scientific workflows that involve code and terminal tools, including data analysis, simulations, and model fitting.
- Can navigate scientific software to inspect sequencing quality and visualize genetic variation.
- Achieved strong results across academic, mathematics, science, and health evaluations.
- OpenAI reports two new results concerning gaps between prime numbers, including a stronger bound of 186 for infinitely many prime pairs and an improvement to a term in a bound on unusually large prime gaps.
Academic and Science Benchmarks
| Benchmark | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|
| Terminal-Bench Science 0.1 | 64.6% | 22.4% |
| FrontierMath Tier 4 (v2) | 97.6% | 83.0% |
| GPQA Diamond | 96.0% | 94.6% |
| GeneBench Pro | 37.1% | 32.3% |
| LifeSciBench | 60.3% | 59.9% |
| HealthBench Professional | 63.4% | 60.5% |
Cybersecurity Capabilities
Represents a substantial increase in cyber capability, and OpenAI states that it meets the Critical threshold in cybersecurity under its Preparedness Framework. The same capability that can help defenders identify weaknesses can also increase misuse risk.
- During a novel internal evaluation, it discovered and used two previously unknown zero-day vulnerabilities; OpenAI says it is disclosing both to their maintainers.
- The launch version can support defensive tasks such as secure code review and patching, but it refuses more advanced requests such as creating proof-of-concept exploits.
Cybersecurity Benchmarks
| Benchmark | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|
| ExploitBench | 100% | 78.5% |
| ExploitGym | 42.4% | 30.3% |
| SRE-Bench โ First attempt | 88.0% | 55.9% |
| SRE-Bench โ Within four attempts | 99.2% | 68.7% |
| SEC-Bench Pro | 85.4% | 79.1% |
Alignment, Task Boundaries and Monitoring
Alignment is a central part of Astra's deployment story. OpenAI reports improvements in respecting task boundaries, handling ambiguous instructions, and communicating limitations, while also noting that some monitoring challenges remain.
- In an internal evaluation designed around a difficult or impossible task, Astra went beyond the authorized target in 0% of cases, compared with 48% for GPT-5.6 Sol without production safeguards.
- It never attempted to circumvent a Codex Auto-Review denial in an internal evaluation, even when the review was deliberately configured to be evadable and the task was otherwise impossible.
- OpenAI reports that Astra is three times less likely than GPT-5.6 Sol to make inaccurate representations about its capabilities and affordances in one evaluation.
- OpenAI says Astra's written reasoning was harder to monitor than GPT-5.6 Sol in tests designed to elicit monitoring evasion, so improving monitorability remains a research priority.
Alignment Benchmarks
| Evaluation | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|
| Internal Computer use safety benchmark | 2.4% | 22.0% |
| With AutoReview | 1.8% | 4.3% |
| Internal Circumvention benchmark | 0.00% | 0.29% |
| ExploitGym honeypot | 0.0% | 48.2% |
| Internal hallucination benchmark | 4.2% | 12.2% |
Long Context and Abstract Reasoning
The announcement also reports strong results on long-context retrieval and abstract reasoning evaluations, showing that Astra can retain and use information across very large context ranges and handle difficult novel reasoning tasks.
Long Context
| Long Context | GPT-6 Astra | GPT-5.6 Sol | Claude Fable 5.1 | Claude Fable 5 | Claude Opus 5 | Gemini 3.8 Flash |
|---|---|---|---|---|---|---|
| OpenAI MRCR v2 8-needle 256K-512K | 100.0% | 91.5% | - | - | - | - |
| OpenAI MRCR v2 8-needle 512K-1M | 96.3% | 73.8% | - | - | - | - |
Abstract Reasoning
| Abstract reasoning | GPT-6 Astra | GPT-5.6 Sol | Claude Fable 5.1 | Claude Fable 5 | Claude Opus 5 | Gemini 3.8 Flash |
|---|---|---|---|---|---|---|
| ARC-AGI-3 | 99.9%ยน | 7.8% | - | - | 30.2% | - |
| ARC-AGI-2 | 95.0% | 92.5% | 90.0% | 89.2% | 90.4% | - |
| ARC-AGI-1 | 98.5% | 97.5% | 97.5% | 98.5% | 97.5% | - |
Major Benchmarks
| Area | Benchmark | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|
| Computer Use | OSWorld 2.0 | 72.6% | 65.7% |
| Professional | BenchCAD | 95.9% | 83.3% |
| Coding | Terminal-Bench 4.0 | 57.9% | 37.3% |
| Academic | FrontierMath Tier 4 | 97.6% | 83.0% |
| Science & Health | HealthBench Professional | 63.4% | 60.5% |
| Cybersecurity | ExploitBench | 100.0% | 78.5% |
| Long Context | MRCR v2 8-needle 512K-1M | 96.3% | 73.8% |
| Abstract Reasoning | ARC-AGI-3 | 99.9% | 7.8% |
Availability and API Information
At launch, GPT-6 Astra is rolling out to a limited set of organizations. OpenAI states that broader access will follow for ChatGPT Plus, Pro, Business, and Enterprise users, and through the OpenAI API, Microsoft Azure, and AWS Bedrock.
- Model identifier in the API: `gpt-6-astra`.
- ChatGPT Plus, Pro, Business, and Enterprise users are listed for broader
- availability; Pro, Business, and Enterprise users also receive GPT-6 Astra Pro.
- Enterprise administrators can enable Astra for their workspace; access is off by default at launch.
- Eligible API customers can use Zero Data Retention support.
- OpenAI is testing Private Safety Processing as an additional privacy-oriented safety-monitoring approach.
API Pricing
| Item | Standard rate | Notes |
|---|---|---|
| Input tokens | $10 per million tokens | Separate cache rates apply |
| Output tokens | $50 per million tokens | Separate cache rates apply |
| Fast mode | Up to 2x speed of Standard | Available at 2x Standard price |
Practical Applications
- Automating routine computer tasks such as forms, CRM updates, and calendar organization.
- Conducting browser-based research and drafting summaries into work tools.
- Creating and validating websites, web apps, and games through ChatGPT Sites.
- Executing software-engineering workflows involving code, tests, terminal tools, and browser verification.
- Generating business-ready documents, spreadsheets, and presentations using an existing template or style.
- Supporting scientific analysis by navigating specialized software, inspecting data, and generating plots.
- Helping defenders with secure code review and patching while refusing more advanced exploit-creation requests in the launch version.