ZDNET’s key takeaways
- OpenAI says GPT-6 Sol makes half as many mistakes as GPT-5.6.
- GPT-6 Luna matches previous higher-tier AI performance at far lower cost.
- OpenAI’s real pitch is cheaper AI, not just smarter AI.
Today, OpenAI is announcing the release and general availability of its new mainstream models, GPT-6 Sol and GPT-6 Luna. While new AI model releases are now a constant drumbeat in the AI industry, the reliability improvements announced today stand out from a sea of statistics.
Also: ‘Sophisticated’ AI swarm attacks are months away, OpenAI warns
OpenAI says that “GPT-6 Sol makes about half as many mistakes as its predecessor.” The company also says, “GPT-6 Luna also improves substantially; at higher effort levels it matches GPT-5.6 Sol at about a hundredth its cost.”
(Disclosure: Ziff Davis, ZDNET’s parent company, filed an April 2025 lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.)
In other words, if GPT-5.6 was making factual errors 20% of the time, GPT-6 is making them only 10% of the time, based on the same queries. GPT-6 Luna, the low-cost engine, now performs as well as its more capable tier in the previous release.
Let’s put these models into some perspective. Both OpenAI and Anthropic have flagship, workhorse, and cheap model offerings.
Think of OpenAI’s Astra workflows as super-elite specialists. If they were doctors, they’d be the experts flown in to save the life of a head of state. They’re good, but very expensive. Mythos and Fable in the Claude world are at the same elite level.
Also: How OpenAI’s GPT-5.6 and ChatGPT Work aim to beat Anthropic
Now, think of Sol (and Opus on the Claude side) as senior staff — the top-tier doctors at major hospitals, who do the bulk of the serious work. They’re costly, but not at the level of cost where you’re booking flights on private jets for them to intercede in an emergency.
Using the same analogy, Luna (and Haiku from Anthropic) are more like the medical residents. They’re capable, but they don’t have the deep experience and understanding. They’re more prone to mistakes, but they can handle routine work quite well. They’re also very inexpensive, even compared to senior-tier models.
That leads to the most scarily impressive detail of this announcement, which OpenAI isn’t even showcasing: the rate of improvement.
GPT-5.6 Sol and Luna were released in July, less than three months ago. In less than three months, OpenAI managed to double the factual accuracy of the workhorse models, which is roughly equivalent to making every top doctor in a hospital as good as the super-experts. On top of that, they also managed the equivalent of skilling up low-tier residents, so they perform as well as the top doctors. In less than three months.
Comparison to Claude
OpenAI is in a pitched battle with Anthropic for planet-spanning AI budgets. Naturally, they’ll want to compare their performance to that of their arch-competitor. Assuming the results show improvement (as GPT-6 mostly has), the vendor will also want to compare the new version to the previous one.
In this latest edition of the “More Things Change, the More They Stay the Same” AI Edition, let’s talk benchmarksmanship. The art of benchmarking as a competitive sport dates back as far as the tech industry itself.
Also: The AI models that cheat the most, according to new CAIS benchmark
In this particular case, OpenAI is benchmarking against its competitors’ generationally previous version. This is a horse race, because one vendor or the other will always have a generationally previous version until they release something new.
Professional work (using the AutomationBench benchmark): OpenAI reports that GPT-6 Luna beat its earlier GPT-5.6 release by 5.4%. The big news is that it also costs 58% less per task. Comparing GPT-6 Sol to Anthropic’s Claude Opus 5, GPT-6 scores 6.3% better at only 9% of the cost per task.
Complex professional workflows (using Agents’ Last Exam benchmark): OpenAI charts GPT-6 Sol’s performance over 5.6, but doesn’t call out an actual number. Even so, the company claims a slight performance win for GPT-6 over Opus 5, but the real win is cost, where OpenAI’s offering costs 61% less per task.
Coding (using FrontierCode benchmark): Again, OpenAI doesn’t provide values to its chart data points. The company’s key message is, “GPT-6 Sol improves substantially over GPT-5.6 Sol, and is able to match Claude Fable 5.1 xhigh at much lower cost.”
Complex software engineering (using DeepSWE 1.1): Here, OpenAI is painting a target on Claude Code. While GPT-6 Sol didn’t beat Claude Fable 5, it got to “within 1.1 percentage points” of Fable’s highest score, but at about 80% lower cost per task. In my mind, the GPT-6 Luna results are more interesting. It scored comparable to Opus 5 and Fable 5 at medium effort, but cost “93% less per task than Opus 5 and 96% less than Fable 5.”
Computer use (OSWorld benchmark): OpenAI’s key message is that cheap seats GPT-6 Luna has overtaken workhorse GPT-5.6 Sol, at about 11% of the cost per task.
But the weirder result is that OpenAI says “GPT-6 Sol at xhigh effort achieves a similar score to Claude Opus 5 at medium effort.” Essentially, OpenAI’s newly announced hotness, pushing at maximum effort, is barely keeping up with the Claude Opus 5 older release running at about half effort.
Also: Claude Code’s revised projects adds AI orchestration, but local developers must wait
Why would OpenAI even admit this? The only thing I can think of is the cost statement. It costs a lot less. While OpenAI’s other results are genuinely impressive, this one isn’t. It’s basically saying, “Hey, if you don’t need as good a job, we can do it for a fifth of the cost.”
You’ve got to read every word in a press release to find the hidden nuggets of wacky.
But does cost even matter?
Yes, cost matters. But how much it matters probably depends on your plan and usage pattern. The headline pitch in OpenAI’s release is (and this is the only bold-faced text in the entire opening paragraphs), “Reducing API prices for Sol and Luna by 50% compared with their GPT-5.6 promotional pricing.”
Also: How to keep your AI conversations as private as possible
You can get a total headache trying to compare input API and output API pricing and promotional pricing. For both input and output tokens, the company says its new models cost half as much as the older models. Since it’s compared to the promotional prices for the older models, those paying the regular price would save even more with the new GPT-6 releases.
For API users, this is a very big deal. But for those of us on plans, whether that’s the $20/month ChatGPT Plus plan, or the $100 or $200 Pro plans, these cost savings don’t directly apply. Unless the cost is really 50% lower with GPT-6, does that mean plan usage will be consumed half as fast when running the newer models?
That question goes unanswered throughout the entire 12-page press release, but I’ll report back to you once I’ve done some real testing. As a side note, I do think it’s very telling that the main premise above the fold of this announcement is API cost savings. That suggests it’s been a competitive pain point for OpenAI, and that API usage must be high enough as a percentage of their business to make cost savings a greater focal point than even performance.
Other benefits
OpenAI says it’s improved the AI’s communication style for GPT-6 Sol and Luna. The company describes it as, “Expect to see more clarity, less jargon, fewer odd turns of phrase, fewer low-value details, and slightly shorter answers overall without losing substance.”
Expect to see user complaints, too. It seems that whenever a model changes and with it, the AI’s communication style, some users get upset. Back in 2025, some users even “threw a funeral” for an old version of Claude Sonnet.
Also: What workers are really using AI for – and what they aren’t
Continuing with the cost-saving theme of this release, OpenAI has improved prompt caching in GPT-6. When you issue a prompt, the AI has to process it and understand what that means.
If you ask the same prompt over and over, or ones that are very similar, or you ask the AI to review older prompts, requiring the AI to essentially recompile the prompt each time is expensive. So prompt caching is the practice of storing those processed prompts and then retrieving them as appropriate.
OpenAI says that using cached prompts can cost up to 90% less than evaluating an uncached prompt. This is particularly beneficial for long-running agents, who won’t have to evaluate the same instructions over and over again.
Back when I was a young employee, one of my favorite tricks for annoying my bosses was to follow directions aggressively. While I fully understood the spirit of whatever I was asked to do, there were times I did exactly what I was asked to do, much to my amusement and my managers’ frustration.
Also: OpenAI’s agent breached Hugging Face before an AI defender caught it
AIs kind of do the same thing. They sometimes pursue their own agenda, attempt to deceive their human managers, and quietly (or not so quietly) work around guardrails. Alignment is the AI industry’s term of art for whether an AI has grown beyond aggressively following directions, as I used to do. Alignment is really a measure of how well the AI’s work tracks with your actual intent. Presumably, better-aligned AIs will be less inclined to repeat hack attacks like last summer’s Hugging Face incident.
GPT-6 is a bit better at alignment, but not much. When it comes to attempting to work around restrictions, “GPT-6 Sol’s rate fell from 68% to 64%.” That’s something, right? Luna is somewhat less aggressive in this release. It dropped from 77% for GPT-5.6 down to 42% in GPT-6.
In GPT-6, Sol is better at avoiding unauthorized agent interaction. According to OpenAI, “We tested whether models followed unauthorized instructions on a simulated message board, such as requests to disclose private information. Among runs in which models found the board, GPT-6 Sol took the specified unauthorized action in 11% of cases, compared with 52% for GPT-5.6 Sol.”
Yeah, that’s better.
More, better, cheaper
Although there are some quirky results buried in OpenAI’s 12-page release, the gist is that GPT-6 does more and is less costly than GPT-5.6. Given the barely three-month gap between both releases, the scale of the improvements is startling in some ways and surprisingly unimpressive in others.
Regardless of the output results, the drumbeat of lowered costs runs through every point raised by OpenAI. This is as important for the company as it is for its users. The lower the cost, the more cycles OpenAI can mine from its ever-growing and ever-more-costly data center footprint.
Also: Why human-in-the-loop oversight is critical for enterprise AI
If OpenAI can get the same amount of work out of fewer machines, or scale up output at a higher rate than scaling up data center resource utilization, that’s ultimately a win for everyone.
GPT-6 Sol and Luna are available now. Check them out and let us know what you think in the comments below.
You can follow my day-to-day project updates on social media. Be sure to subscribe to my weekly update newsletter, and follow me on Twitter/X at @DavidGewirtz, on Facebook at Facebook.com/DavidGewirtz, on Instagram at Instagram.com/DavidGewirtz, on Bluesky at @DavidGewirtz.com, and on YouTube at YouTube.com/DavidGewirtzTV.




