OpenAI’s Budget-Friendly GPT‑6 Sol And Luna Models Keep Benchmarks Steady
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: OpenAI’s Budget-Friendly GPT‑6 Sol And Luna Models Keep Benchmarks Steady on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI has launched GPT‑6 Sol and Luna, two new models priced at half their predecessor’s cost. They deliver comparable benchmark scores while reducing expenses, though some quality regressions are noted. This shift could expand AI deployment in budget-sensitive applications.

OpenAI has released two new models, GPT‑6 Sol and Luna, priced at approximately half the cost of their GPT‑5.6 predecessors. These models are designed to democratize access to advanced AI by significantly reducing operational expenses while maintaining benchmark performance levels, marking a shift in how AI capabilities can be integrated into products and workflows.

The new models, GPT‑6 Sol and Luna, arrived on September 22, 2026, with prices cut by 50% compared to GPT‑5.6. GPT‑6 Sol costs $2.00 per 1 million tokens for input and $10.00 for output, while Luna is priced at $0.10 and $0.50 respectively. These reductions are achieved through improvements in caching and inference efficiency, passing savings directly to users.

Independent analysis by Artificial Analysis confirms that the cost per task has approximately halved, with the models maintaining roughly similar benchmark scores. GPT‑6 Sol scores 48 on the Artificial Analysis Intelligence Index, well above the median of 25, with a context window of 872,000 tokens. Luna scores 37, also outperforming its peers, with a 1 million token context window. Despite the cost savings, some regressions are observed in certain knowledge-work evaluations, attributed to changes in presentation quality and output completeness.

In terms of quality, GPT‑6 Sol shows notable improvements in hallucination reduction, cutting its hallucination rate from 92% to 60%. However, this comes with an increase in refusals, with Sol attempting fewer questions—83% versus 99%—which reduces errors but also decreases overall answer completeness. The models also demonstrate varied performance in coding tasks, with Sol slightly outperforming Luna in some indices but regressing in others. The models’ ability to balance cost, quality, and effort is a key focus for users deploying these models at scale.

At a glance
updateWhen: announced September 22, 2026
The developmentOpenAI announced the release of GPT‑6 Sol and Luna models on September 22, 2026, emphasizing their lower cost and maintained performance benchmarks.

GPT‑6 Sol and Luna: half the price, about the same intelligence

OpenAI’s September 22, 2026 release doesn’t raise the ceiling. It lowers the cost of everything below it, which changes what’s worth automating.

GPT‑6 Sol
$4 / $20 → $2 / $10
GPT‑6 Luna
$0.20 / $1.20 → $0.10 / $0.50

Per 1M input / output tokens. Cached input reads keep the 90% discount.

Cost per task, halved

Measured by Artificial Analysis as the weighted cost of one Intelligence Index task, at max effort.

GPT‑5.6 Sol
$1.99
GPT‑6 Sol
$1.06
GPT‑5.6 Luna
$0.18
GPT‑6 Luna
$0.07

The effort dial moves cost more than the model choice

Model and effortIntelligence IndexCost per task
GPT‑6 Sol (max)48$1.06
GPT‑6 Sol (low)34$0.13
GPT‑6 Luna (max)37$0.07
GPT‑6 Luna (low)21$0.0045
GPT‑6 Luna (non‑reasoning)18$0.01

Sol at low effort keeps about 70% of its max score for roughly an eighth of the cost, because it writes far fewer reasoning tokens. For reference, Claude Opus 5.5 leads the same index at 58.

What got better, and what got worse

Better

  • Hallucination rate on AA‑Omniscience: Sol 92% → 60%, Luna 93% → 77%
  • Coding Agent Index: Sol 57, up 2 points, at ~50% lower cost per task
  • OpenAI reports about half as many factual mistakes for Sol as its predecessor
  • Higher cache hit rates; GitHub reports over 50% fewer prompt tokens needing fresh processing

Sol gets there partly by declining more: it attempts 83% of questions vs 99%, and accuracy falls 59% → 54%.

Worse

  • GDPval‑AA v2.1: Sol down ~100 Elo, Luna down ~75
  • AA‑Briefcase v1.1: Luna down ~45 Elo
  • Coding Agent Index: Luna 41, down 2 points
  • Both models write more output tokens per task than their predecessors

Reviewers attribute the drops to weaker presentation and deliverables that omit required elements.

What to do about it

Already on GPT‑5.6 Sol or Luna? The move is mostly a price cut. Re‑test first if your output is a document someone reads, not data a system consumes.
Shelved an automation on cost? Token prices halved and the effort dial adds another order of magnitude. Re‑run the business case.
Choosing between labs? The question is no longer which model is smartest, but which clears your quality bar at the lowest cost per task.
ThorstenMeyerAI.comSources: OpenAI (pricing, vendor benchmarks) and Artificial Analysis (independent evaluation and model pages). Figures as of 23 September 2026.

Impact on AI Deployment and Cost Efficiency

The introduction of GPT‑6 Sol and Luna at half the previous cost levels could significantly expand the scope of AI deployment across industries. Smaller organizations and startups, previously limited by high operational costs, can now integrate advanced models into customer service, research, and automation workflows. The models’ maintained benchmark performance ensures that quality remains competitive, while the lower price points make large-scale or high-frequency applications more feasible.

However, the observed regressions in some knowledge and presentation tasks highlight the importance of careful testing before replacing existing models in production. The trade-offs between hallucination reduction and answer completeness, as well as the impact of effort settings, will influence how organizations optimize these models for their specific needs. Overall, this release underscores a shift toward more accessible AI, emphasizing cost-efficiency without sacrificing core capabilities.

Amazon

AI language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of OpenAI’s Model Pricing and Performance Trends

OpenAI has historically positioned its models as premium offerings, with GPT‑5.6 and similar versions priced to reflect their advanced capabilities. The release of Astra, a high-end model announced earlier in 2026, set a new benchmark for performance, but at a high cost. The recent launch of GPT‑6 Sol and Luna marks a strategic shift toward broadening AI adoption by offering models that are both affordable and sufficiently capable for many tasks.

Prior to this, OpenAI’s models were often limited by cost constraints, especially for smaller entities or applications requiring high-volume inference. The company’s focus on improving caching and inference efficiency—leading to the 50% price reduction—is part of a broader effort to make advanced AI more accessible. Independent evaluations confirm that while the models’ costs have dropped, their benchmark scores remain competitive, though some quality regressions have been noted in specific areas.

This move aligns with industry trends toward democratizing AI, where cost-effective models are increasingly essential for scaling AI solutions across sectors.

Amazon

affordable AI chatbot development tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Performance Variability and Long-term Stability

While initial evaluations show promising cost savings and maintained benchmarks, some regressions in knowledge and presentation quality raise questions about long-term stability and suitability for all workflows. The impact of effort tuning on output consistency and accuracy remains an area for further testing, especially in high-stakes applications. OpenAI has not yet provided extensive real-world deployment data to confirm sustained performance over time.

Amazon

AI coding assistant software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Adoption and Evaluation

Organizations are encouraged to conduct their own testing of GPT‑6 Sol and Luna within their specific workflows before full deployment. OpenAI is expected to release additional diagnostics tools and user feedback channels to monitor performance and address emerging issues. Further independent studies and real-world case studies will clarify the models’ suitability across different industries and use cases. The AI community will also watch for updates on long-term stability and ongoing improvements in hallucination reduction and output quality.

Amazon

cost-effective AI automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How do GPT‑6 Sol and Luna compare to previous models in terms of cost?

They are priced at approximately half the cost of GPT‑5.6 models, with GPT‑6 Sol at $2.00 per 1M input tokens and $10.00 per 1M output tokens, and Luna at $0.10 and $0.50 respectively.

Do these models perform as well as higher-end models like Astra?

They maintain comparable benchmark scores in many evaluations, though some knowledge and presentation quality regressions are noted. They are designed for cost-effective deployment rather than maximum performance.

What are the main trade-offs with these models?

While cost savings are significant, some regressions in answer completeness and presentation quality have been observed. Additionally, reducing hallucinations involves more refusals, which may impact workflows requiring complete outputs.

Are there any concerns about long-term stability or reliability?

Yes, the initial evaluations are promising, but long-term stability and performance consistency in real-world applications are still being monitored. Further data will clarify these aspects over time.

How should organizations approach testing these new models?

Organizations should conduct thorough testing within their specific use cases, especially for high-stakes or detailed output tasks, before replacing existing models in production environments.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

SenseTime Breaks Even And Turns Profit Thanks To AI Portfolio Gains

SenseTime signals it expects to record its first profit since its Hong Kong IPO, driven by gains across its AI portfolio. Full details pending official financial disclosure.

How Artificial Intelligence Is Enabling Multilingual Understanding

Google announces support for over 300 languages, real-time speech translation in 70 languages, and new models to improve multilingual understanding worldwide.

Qwen4 Architecture Unveiled Early — What AI Experts Are Saying

Alibaba’s Qwen team open-sourced early architecture details of Qwen4, sparking expert analysis on its design and implications for AI development.

Explore AI Innovation By Building A Grok Bot With Grok Bot

xAI announced a project titled ‘Designing Grok Bot with Grok Bot,’ indicating Grok AI’s involvement in creating a new system, but details remain limited.