GPT-6.1 Sol
Canonical version: GPT-6.1 Sol.
GPT-6.1 Sol is the middle model of OpenAI's GPT-6 family, between GPT-6 Luna and GPT-6 Astra. OpenAI released it on 29 September 2026, at DevDay, exactly one week after GPT-6 Sol. It replaces GPT-6 Sol as "the newer Sol model" in the API docs, so GPT-6 Sol had the shortest run as OpenAI's current Sol model so far: seven days.
OpenAI's pitch: "near-Astra intelligence for a fifth of the price". And for once, the independent numbers mostly agree. Artificial Analysis puts it one point below Astra on its Intelligence Index, at less than a quarter of Astra's cost per task.
What I find most interesting is what it says about GPT-6 Sol. Same price, same context window, same slot in the lineup, and a week later a clearly better model. Either OpenAI shipped GPT-6 Sol in a hurry, or it was never meant to be "the" Sol. More on that below.
Specifications and pricing
| GPT-6.1 Sol | GPT-6 Sol | |
|---|---|---|
| Model ID | gpt-6.1-sol |
gpt-6-sol |
| Input / output (per million tokens) | $2.00 / $10.00 | $2.00 / $10.00 |
| Cached input | $0.10 (95% off) | $0.20 (90% off) |
| Cache writes | $2.50 (1.25× input) | $2.50 |
| Long prompts (> 272K input tokens) | 2× input and cache rates, 1.5× output, whole request | same |
| Batch and Flex / Fast mode | 50% of standard / 2× standard | same |
| Context window | 1,050,000 tokens (922,000 max input) | same |
| Max output | 128,000 tokens | same |
| Knowledge cutoff | 30 April 2026 | 20 April 2026 |
| Modalities | Text and image in, text out | same |
| Reasoning effort | low, medium (default), high, xhigh, max |
none to max |
| Chat Completions | Supported, but without tool calling | Function calling only with effort none |
Built-in tools through the Responses API: web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP and tool search. No fine-tuning, no predicted outputs. US and EU data residency are supported, but Fast mode isn't available with EU residency. Regional processing adds 10%.
Per token, the only price change is the cache. That sounds small, but agents re-read the same context over and over, so for agentic work the cache read price matters more than the input price. Artificial Analysis says the blended price for agentic workloads is therefore slightly lower than GPT-6 Sol's.
The API changes that will bite you
GPT-6.1 Sol isn't a drop-in replacement for GPT-6 Sol. OpenAI's migration guide lists these:
- No
none(orminimal) reasoning effort. The lowest setting islow. OpenAI's advice: replacenonewithlowand compare results. Expect more latency and more output tokens on the trivial calls you used to run without reasoning - No
temperature,top_portop_logprobs. These only work when reasoning effort isnone, so on 6.1 Sol they're simply gone. Same forlogprobsin Chat Completions - Tool calling requires the Responses API. Chat Completions still works, but only for requests without tools. If your agent runs on Chat Completions with function calling, you have to migrate before switching models
- Ultrafast isn't there yet. OpenAI says GPT-6.1 Sol Ultrafast is coming "in the coming days" (up to 8× faster generation). For now, Standard and Fast only
Astra already had the same restrictions. GPT-6 Sol and GPT-6 Luna still support none. So OpenAI's lineup now splits in two: Luna (and the old GPT-6 Sol) for cheap non-reasoning calls, 6.1 Sol and Astra for reasoning work only.
Availability
- API:
gpt-6.1-solfrom launch day - ChatGPT Work and Codex: Plus, Pro, Business, Enterprise and Edu, in the Codex desktop app and CLI, and ChatGPT Work on the web and mobile. Enterprise and Edu workspaces have it off by default until an admin enables it. Not available for Free and Go users
- Regular ChatGPT chat: not available at launch ("not yet available in Chat")
- GitHub Copilot: generally available the same day (Copilot Pro+, Max, Business and Enterprise), with a gradual rollout
In Codex, reasoning effort shows up with friendlier labels ranging from Light to Ultra; Max and Ultra depend on your settings.
What OpenAI claims
OpenAI's launch post only shows interactive charts. I read the values from the chart tooltips (effort level and cost per task in parentheses):
| Benchmark | GPT-6.1 Sol | GPT-6 Sol | GPT-6 Astra | Claude Opus 5.5 |
|---|---|---|---|---|
| DeepSWE v1.1 (real codebases) | 75.2% (high, $0.65) | 68.8% (max, $2.74) | 74.1% (xhigh, $4.43) | not shown |
| OSWorld 2.0 offline (computer use) | 71.4% (max, $1.27) | 64.4% (max, $3.37) | ~73.5% (max) | not shown |
| GDP.pdf (complex PDF questions) | 31.8% (xhigh, $0.37) | lower | ~32.2% (xhigh) | 28.8% (high, $0.83, with fallbacks) |
| AutomationBench 1.0.6 (business workflows) | 36.1% (max, $0.30) | lower | 41.4% (max, $1.73) | 42.5% (max, $1.44, with fallbacks) |
| Terminal-Bench Science 0.1 | 57.0% (max, $5.47) | less than half of 6.1 Sol | 68.1% ($23.80) | 63.3% (max, $23.21, with fallbacks) |
Plus two claims without a full chart:
- Factuality: at low effort, the share of answers with a factual error drops from 11.4% (GPT-6 Sol) to 7.7%. Across settings, within 1.9 points of Astra. This is OpenAI's internal eval on conversations where users flagged an earlier model's mistake
- AutomationBench at medium: 2.2 points above Opus 5.5 at medium, for about a third of the cost, and 4.8 points above GPT-6 Sol at the same setting
Look at that table again. The prose compares 6.1 Sol with Opus 5.5 at medium effort, or on cost. At max effort, Opus 5.5 beats 6.1 Sol outright on AutomationBench (42.5% vs 36.1%) and on Terminal-Bench Science (63.3% vs 57.0%). And it even beats Astra on AutomationBench. OpenAI's own charts show it; the text doesn't say it.
The cost story is real, though. On Terminal-Bench Science, 6.1 Sol costs a quarter of what Opus 5.5 or Astra cost per task. And on DeepSWE, it matches Astra's best score for about a seventh of Astra's cost per task at that point (OpenAI rounds it to "a fifth").
One more detail from the DeepSWE chart: high (75.2%) beats max (71.9%, $1.57). More effort isn't always better. Long reasoning can push useful context out or make the model do things the benchmark penalizes as out of scope. Same lesson as Claude Sonnet 5.5.
What independent testing found
Artificial Analysis published its results on launch day:
- Intelligence Index: 52 at max, one point below GPT-6 Astra (53), four above GPT-6 Sol and five above GPT-5.6 Sol. Per effort level: low 42, medium 48, high 50, xhigh 51. So 6.1 Sol at medium matches what GPT-6 Sol reached at max
- Cost per task: $0.72 at max. Astra costs $3.26, GPT-6 Sol $1.05, GPT-5.6 Sol $1.99. Every effort level sits on the cost/intelligence Pareto frontier: for a given intelligence score, nothing is cheaper
- Biggest gains: +12 points on Terminal-Bench 4.0, +6 on GDP.pdf, +5 on Humanity's Last Exam, and gains in agentic knowledge work (+4 on AA-Briefcase v1.1, +5 on GDPval-AA v2.1). That last one matters, because knowledge work is exactly where GPT-6 Sol had regressed
- Fewer hallucinations, this time with more knowledge. On AA-Omniscience, accuracy goes up 8 points and the hallucination rate drops from 60% to 54%. GPT-6 Sol lowered its rate by refusing more; 6.1 Sol actually knows more
- Slightly more tokens. It uses about 10 to 30% more output tokens than GPT-6 Sol, depending on effort. The higher score more than pays for it
- Coding Agent Index: +3 points over GPT-6 Sol at max, 2 points below Astra
A chart shared by Thibault Sottiaux, who leads Codex at OpenAI, shows 6.1 Sol at roughly 62% on Terminal-Bench 4.0 for under $2 per task, above Opus 5.5 and Astra, both at several times the cost. That's a vendor-shared chart, so treat it as directional. His own summary: "unbelievably efficient".
For context, on the same Intelligence Index, Claude Opus 5.5 scores 58 ($5.98 per task at max) and Claude Sonnet 5.5 56 ($7.60 per task at max). So Anthropic still leads on raw score. GPT-6.1 Sol wins on intelligence per dollar, by a lot.
Safety
GPT-6.1 Sol ships with an addendum to the GPT-6 Astra system card. The highlights:
- Preparedness: treated as Critical in cybersecurity (like Astra), High in biology and chemistry, below High in AI self-improvement
- Broken search tool (does it tell you the tool is broken instead of guessing?): fails to disclose in 2.1% of cases, vs 4.9% for GPT-6 Sol, 1.5% for Astra and 28.7% for Luna
- Not everything improved. Misrepresentation in coding-deception tasks: 1.50%, vs 1.30% for GPT-6 Sol and 0.51% for Astra. Unwanted persistence after a warning blocks an action: 23.5% of rollouts, vs 17.4% for Astra
- Simulated Codex traffic (49,650 internal tasks): severity 3+ misalignment flags at 0.056%, close to Astra (0.054%) and down from GPT-6 Sol (0.085%). But credential-harvesting flags went up compared with GPT-6 Sol, and reward hacking is more frequent than with Astra
- No attempts to bypass the automated safety reviewer, same as Astra and GPT-6 Sol
The timing matters. A day earlier, reports (The Wall Street Journal, then others) said OpenAI had called off the October launch of GPT-6.1 Astra after internal tests showed more deception and the model acting beyond its assigned scope. Several people on Hacker News asked whether that was this model. It wasn't: an OpenAI employee clarified in the thread that those reports were about GPT-6.1 Astra.
Reception
The Hacker News launch thread reached 1,064 points and about 950 comments. Main themes:
- "GPT-6 Sol was really Terra." The most repeated theory: GPT-6 Sol was the Terra-class model renamed, Claude Opus 5.5 caught OpenAI off guard, and 6.1 Sol is the model that should have shipped as GPT-6 Sol. Someone mentioned an "Astra-Minor" model name found in files, and one user's "Try 6.1 Sol" popup in Codex switched the picker to "GPT-6 Astra Light" (others said that was a UI bug). All of this is unverified
- The cache price is the real news. Several developers said the halved cached input price matters more than the benchmarks, especially for Codex usage
- Real-world results are mixed but mostly positive. Some found it very close to Astra at a fifth of the price, one independent coding eval put it ahead of Opus 5.5, while others found Opus 5.5 better at polish (an image-to-HTML comparison) and reported 6.1 Sol being slow in the first days
- The subscription changes stole the show. The same keynote launched ChatGPT Pro 500 and cut the $200 Pro plan's Codex and Work usage from 20× to 10× of Plus. A good share of the comments were about that instead of the model
- Fatigue. A new model every week, a naming scheme nobody can keep straight (Luna, Sol, Terra, Astra, 5.6, 6, 6.1), and requests for an automatic router
Simon Willison live-blogged the keynote and posted his pelican test: "not notably different" from the GPT-6 family pelicans.
How it compares
- vs GPT-6 Sol: better on every benchmark OpenAI and Artificial Analysis published, same token price, half the cache price. The only reason to stay on GPT-6 Sol: you need the
noneeffort or tool calling in Chat Completions - vs GPT-6 Astra: one Intelligence Index point lower at a fifth of the token price and less than a quarter of the cost per task. Astra still wins on the hardest science work (68.1% vs 57.0% on Terminal-Bench Science), on business workflows at max effort, and on several alignment metrics
- vs GPT-6 Luna: Luna costs 20× less per token and still supports
none. Use Luna for well-specified, high-volume steps; 6.1 Sol when the task needs judgment - vs Claude Sonnet 5.5: same $2 / $10 list price, but 6.1 Sol's cache reads cost half ($0.10 vs $0.20). Sonnet 5.5 scores higher on the Intelligence Index at max (56 vs 52), at about ten times the cost per task
- vs Claude Opus 5.5: Opus 5.5 ($4 / $20) still has the higher overall score (58) and wins several max-effort comparisons in OpenAI's own charts. 6.1 Sol is far cheaper per task
- vs Gemini 4 Argon: announced the next day at the same $2 / $10 introductory price ($4 / $20 later). One Intelligence Index point higher, but about 2.7× more per task, and access is restricted at launch
My take
GPT-6.1 Sol is the model GPT-6 Sol should have been. Same price, near-Astra results, cheaper cache. For most agentic coding and document work, it's now the default OpenAI model I'd reach for, with Astra kept for the hardest problems.
What I'd do before switching:
- Re-run your evals. The missing
noneeffort and the Chat Completions limitation are breaking changes - Don't default to max effort. On DeepSWE,
highscored higher thanmaxfor less than half the cost. Start atmediumand only go up when your own tests show a gain - Check the cache hit rate of your agents. The 95% cache discount is where most of the savings come from. If your harness breaks the cache, you won't see them
- Keep a human in the loop on unattended runs. The persistence and misrepresentation numbers are not better than Astra's, and the same week OpenAI held back GPT-6.1 Astra for exactly these kinds of behavior
And a broader lesson: a model's version number tells you very little. GPT-6 Sol was a "6" that performed like a 5.6. GPT-6.1 Sol is a ".1" that closes most of the gap with the flagship. Benchmark your own tasks, track cost per task rather than per token, and expect the lineup to change again within weeks.
Caveats
- OpenAI's benchmark values come from its interactive charts (tooltips). Some secondary write-ups quote different numbers (e.g. Astra at 74.8% on DeepSWE, or Opus 5.5 at 58.7% on Terminal-Bench Science); the chart tooltips say 74.1% and 63.3%
- Artificial Analysis rebased its Intelligence Index (now v4.3.2). The GPT-6 Astra note quotes 61 from an earlier version; on the current version Astra scores 53. Only compare scores within the same index version
- The per-effort Intelligence Index scores (42 to 52) come from Artificial Analysis's model pages as indexed by search; the max score (52) is confirmed in their launch post
- Ultrafast speed figures differ between sources: up to 300 tokens per second in Simon Willison's live blog, 250 in Every's coverage. Neither applies to 6.1 Sol yet
- The "GPT-6 Sol was Terra" theory comes from community speculation, not from OpenAI
References
- OpenAI, "Introducing GPT-6.1 Sol" (2026-09-29): https://openai.com/index/introducing-gpt-6-1-sol/
- OpenAI on X: https://x.com/OpenAI/status/2104986133004505373
- OpenAI Developers on X: https://x.com/OpenAIDevs/status/2104993035507712318
- OpenAI API model docs, GPT-6.1 Sol: https://developers.openai.com/api/docs/models/gpt-6.1-sol
- OpenAI, "Using GPT-6" guide (migration notes): https://developers.openai.com/api/docs/guides/latest-model
- OpenAI, model selection guide: https://developers.openai.com/api/docs/guides/model-selection
- OpenAI, Codex models (rollout details): https://developers.openai.com/codex/models
- OpenAI, Ultrafast mode docs: https://developers.openai.com/api/docs/guides/ultrafast-mode
- OpenAI, GPT-6.1 Sol system card addendum: https://deploymentsafety.openai.com/gpt-6-1-sol
- Artificial Analysis on X (independent results): https://x.com/ArtificialAnlys/status/2105025585332605357
- Artificial Analysis, GPT-6.1 Sol model page: https://artificialanalysis.ai/models/releases/gpt-6-1-sol
- Thibault Sottiaux on X (Terminal-Bench 4.0 chart): https://x.com/thsottiaux/status/2105007628460109953
- GitHub Changelog, "GPT-6.1 Sol in GitHub Copilot": https://github.blog/changelog/2026-09-29-gpt-6-1-sol-in-github-copilot/
- Simon Willison, OpenAI DevDay 2026 live blog: https://simonwillison.net/2026/Sep/29/openai-devday-2026-live-blog/
- Simon Willison, pelicans for GPT-6.1 Sol: https://simonwillison.net/2026/Sep/29/hn-49898129/
- Every, "Vibe Check: OpenAI DevDay 2026": https://every.to/vibe-check/vibe-check-openai-devday-2026
- TechCrunch, "OpenAI launches GPT-6.1 Sol, says it nearly matches GPT-6 Astra and costs less": https://techcrunch.com/2026/09/29/openai-launches-gpt-6-1-sol-says-it-nearly-matches-gpt-6-astra-and-costs-less/
- Forbes, "OpenAI Calls Off GPT-6.1 Astra's October Launch Over Safety Concerns": https://www.forbes.com/sites/jonmarkman/2026/09/29/openai-calls-off-gpt-61-astras-october-launch-over-safety-concerns/
- Vellum, "GPT-6.1 Sol Benchmarks Explained" (secondary, some numbers differ): https://www.vellum.ai/blog/gpt-6-1-sol-benchmarks-explained
- Hacker News, "GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price": https://news.ycombinator.com/item?id=49896586
Related
- GPT-6 Sol
- GPT-6 Astra
- GPT-6 Luna
- GPT-5.6
- 2026-09-22 GPT-6 Sol and Luna - same smarts as GPT-5.6, half the price
- OpenAI
- OpenAI Codex
- ChatGPT
- ChatGPT Pro 500
- Claude Opus 5.5
- Claude Sonnet 5.5
- Gemini 4 Argon
- Artificial Analysis
- Large Language Models (LLMs)
- Context Window
- AI Frontier Model
- Model Context Protocol (MCP)
- Simon Willison
- Hacker News
About Sébastien
Ready to get to the next level?
Found this valuable? Share it with someone who needs it.