🔍 Read the full analysis: OpenAI’s Budget-Friendly GPT‑6 Sol And Luna Models Keep Benchmarks Steady on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
OpenAI has launched GPT‑6 Sol and Luna, two new models priced at half their predecessor’s cost. They deliver comparable benchmark scores while reducing expenses, though some quality regressions are noted. This shift could expand AI deployment in budget-sensitive applications.
OpenAI has released two new models, GPT‑6 Sol and Luna, priced at approximately half the cost of their GPT‑5.6 predecessors. These models are designed to democratize access to advanced AI by significantly reducing operational expenses while maintaining benchmark performance levels, marking a shift in how AI capabilities can be integrated into products and workflows.
The new models, GPT‑6 Sol and Luna, arrived on September 22, 2026, with prices cut by 50% compared to GPT‑5.6. GPT‑6 Sol costs $2.00 per 1 million tokens for input and $10.00 for output, while Luna is priced at $0.10 and $0.50 respectively. These reductions are achieved through improvements in caching and inference efficiency, passing savings directly to users.
Independent analysis by Artificial Analysis confirms that the cost per task has approximately halved, with the models maintaining roughly similar benchmark scores. GPT‑6 Sol scores 48 on the Artificial Analysis Intelligence Index, well above the median of 25, with a context window of 872,000 tokens. Luna scores 37, also outperforming its peers, with a 1 million token context window. Despite the cost savings, some regressions are observed in certain knowledge-work evaluations, attributed to changes in presentation quality and output completeness.
In terms of quality, GPT‑6 Sol shows notable improvements in hallucination reduction, cutting its hallucination rate from 92% to 60%. However, this comes with an increase in refusals, with Sol attempting fewer questions—83% versus 99%—which reduces errors but also decreases overall answer completeness. The models also demonstrate varied performance in coding tasks, with Sol slightly outperforming Luna in some indices but regressing in others. The models’ ability to balance cost, quality, and effort is a key focus for users deploying these models at scale.
GPT‑6 Sol and Luna: half the price, about the same intelligence
OpenAI’s September 22, 2026 release doesn’t raise the ceiling. It lowers the cost of everything below it, which changes what’s worth automating.
Per 1M input / output tokens. Cached input reads keep the 90% discount.
Cost per task, halved
Measured by Artificial Analysis as the weighted cost of one Intelligence Index task, at max effort.
The effort dial moves cost more than the model choice
| Model and effort | Intelligence Index | Cost per task |
|---|---|---|
| GPT‑6 Sol (max) | 48 | $1.06 |
| GPT‑6 Sol (low) | 34 | $0.13 |
| GPT‑6 Luna (max) | 37 | $0.07 |
| GPT‑6 Luna (low) | 21 | $0.0045 |
| GPT‑6 Luna (non‑reasoning) | 18 | $0.01 |
Sol at low effort keeps about 70% of its max score for roughly an eighth of the cost, because it writes far fewer reasoning tokens. For reference, Claude Opus 5.5 leads the same index at 58.
What got better, and what got worse
Better
- Hallucination rate on AA‑Omniscience: Sol 92% → 60%, Luna 93% → 77%
- Coding Agent Index: Sol 57, up 2 points, at ~50% lower cost per task
- OpenAI reports about half as many factual mistakes for Sol as its predecessor
- Higher cache hit rates; GitHub reports over 50% fewer prompt tokens needing fresh processing
Sol gets there partly by declining more: it attempts 83% of questions vs 99%, and accuracy falls 59% → 54%.
Worse
- GDPval‑AA v2.1: Sol down ~100 Elo, Luna down ~75
- AA‑Briefcase v1.1: Luna down ~45 Elo
- Coding Agent Index: Luna 41, down 2 points
- Both models write more output tokens per task than their predecessors
Reviewers attribute the drops to weaker presentation and deliverables that omit required elements.
What to do about it
Impact on AI Deployment and Cost Efficiency
The introduction of GPT‑6 Sol and Luna at half the previous cost levels could significantly expand the scope of AI deployment across industries. Smaller organizations and startups, previously limited by high operational costs, can now integrate advanced models into customer service, research, and automation workflows. The models’ maintained benchmark performance ensures that quality remains competitive, while the lower price points make large-scale or high-frequency applications more feasible.
However, the observed regressions in some knowledge and presentation tasks highlight the importance of careful testing before replacing existing models in production. The trade-offs between hallucination reduction and answer completeness, as well as the impact of effort settings, will influence how organizations optimize these models for their specific needs. Overall, this release underscores a shift toward more accessible AI, emphasizing cost-efficiency without sacrificing core capabilities.
As an affiliate, we earn on qualifying purchases.
Background of OpenAI’s Model Pricing and Performance Trends
OpenAI has historically positioned its models as premium offerings, with GPT‑5.6 and similar versions priced to reflect their advanced capabilities. The release of Astra, a high-end model announced earlier in 2026, set a new benchmark for performance, but at a high cost. The recent launch of GPT‑6 Sol and Luna marks a strategic shift toward broadening AI adoption by offering models that are both affordable and sufficiently capable for many tasks.
Prior to this, OpenAI’s models were often limited by cost constraints, especially for smaller entities or applications requiring high-volume inference. The company’s focus on improving caching and inference efficiency—leading to the 50% price reduction—is part of a broader effort to make advanced AI more accessible. Independent evaluations confirm that while the models’ costs have dropped, their benchmark scores remain competitive, though some quality regressions have been noted in specific areas.
This move aligns with industry trends toward democratizing AI, where cost-effective models are increasingly essential for scaling AI solutions across sectors.
affordable AI chatbot development tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Performance Variability and Long-term Stability
While initial evaluations show promising cost savings and maintained benchmarks, some regressions in knowledge and presentation quality raise questions about long-term stability and suitability for all workflows. The impact of effort tuning on output consistency and accuracy remains an area for further testing, especially in high-stakes applications. OpenAI has not yet provided extensive real-world deployment data to confirm sustained performance over time.
As an affiliate, we earn on qualifying purchases.
Next Steps for Adoption and Evaluation
Organizations are encouraged to conduct their own testing of GPT‑6 Sol and Luna within their specific workflows before full deployment. OpenAI is expected to release additional diagnostics tools and user feedback channels to monitor performance and address emerging issues. Further independent studies and real-world case studies will clarify the models’ suitability across different industries and use cases. The AI community will also watch for updates on long-term stability and ongoing improvements in hallucination reduction and output quality.
cost-effective AI automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
How do GPT‑6 Sol and Luna compare to previous models in terms of cost?
They are priced at approximately half the cost of GPT‑5.6 models, with GPT‑6 Sol at $2.00 per 1M input tokens and $10.00 per 1M output tokens, and Luna at $0.10 and $0.50 respectively.
Do these models perform as well as higher-end models like Astra?
They maintain comparable benchmark scores in many evaluations, though some knowledge and presentation quality regressions are noted. They are designed for cost-effective deployment rather than maximum performance.
What are the main trade-offs with these models?
While cost savings are significant, some regressions in answer completeness and presentation quality have been observed. Additionally, reducing hallucinations involves more refusals, which may impact workflows requiring complete outputs.
Are there any concerns about long-term stability or reliability?
Yes, the initial evaluations are promising, but long-term stability and performance consistency in real-world applications are still being monitored. Further data will clarify these aspects over time.
How should organizations approach testing these new models?
Organizations should conduct thorough testing within their specific use cases, especially for high-stakes or detailed output tasks, before replacing existing models in production environments.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
