Whereas Google with Gemini 3.x, Anthropic with Fable etc are happy to just go for 'big model with dense params'
It's hard to guess from the outside of course but just this kind of talking points focus on GPU efficacy is what we see from OpenAI and Chinese open source labs more often than from Anthropic or Google Deepmind and this benchmark chart seems to concur
Before I thought it was just an improved version or at least in the same class as gpt 5.4 mini but now it's being priced like a nano model!
I thought about it because Terra has similar pricing to 5.4 and Sol is similar to 5.5.
Luna was already my workhorse before, it performs very well on high/xhigh for most of the tasks, very happy about this drop.
To fix this you currently need to make your own copy of the bundled model catalog [2] and opt Luna into MultiAgent V2.
1. https://github.com/openai/codex/issues/32031
2. https://github.com/openai/codex/issues/32031#issuecomment-51...
Does this seem higher or lower limits ?
This applies even more strongly to model choosing. I know for a fact that majority of my work doesn't require a very strong model, but separating the trivial and non-trivial tasks is a famously hard problem (if at all decidable).
I don't have the words.
I genuinely thought we were in a stage where we were plateauing and going in for 5-10% improvements over months. Seeing spikes like this makes me question about where the floor really is.
If the cost of serving GPT-5.6 just dropped by 20%, does that add up to literally billions of dollars in savings per month?
We know Anthropic spend $1.25 billion renting inference capacity from SpaceX (in two Colossus datacenters) from the SpaceX IPO, but we don't know how much of Anthropic's inference capacity that is (presumably a small fraction, since they were operating on top of AWS and other providers before the SpaceX deal.)
I've not seen any numbers that hint at OpenAI's per-month inference bill, but surely that has to be in the multiple billions of dollars as well.
So 20% is a really, really big deal.
I was already a huge proponent of Luna for things like deep research. Being able to run 5x more for the same cost is simply bananas. We are already running 10 parallel agents for hypothesis generation. I cannot imagine 50. The statistics become much more interesting & powerful when you can run so many samples of the exact same prompt+model without breaking the bank.
For our production app, OpenAI clearly is the best provider now. Their API is very reliable and has many nice features. The price-performance of the model lineup is incredible. We used open weights model via Fireworks for a long time (e.g. Kimi K2.5). Fireworks is a great provider but we still ran into issues here and there (Same with Anthropic and Google). OpenAI just works, is fast and in my view has a better price-performance ratio across almost all levels of intelligence.
I was still using GLM-5.2 in my personal projects, but this just made Luna a very easy choice.
https://deepswe.datacurve.ai/ - (See the Agent Steps view)
Or is the output speed so much higher that it cancels out?
I don't see a lot of benchmarks that record actual time. But on AA, Sol on Low beats Luna on High for Time Per Task.
Perhaps users prefer instant responses over thinking models so much so that using a more expensive and less performant non-thinking model is worthwhile.
[1] https://developers.openai.com/api/docs/models/chat-latest [2] https://developers.openai.com/api/docs/models/gpt-5.6-luna
I don't buy it.
There have been recent weeks where some of the mid-level models (Hy3, Laguna M.1) are free (true for parts of June and July, see Hy3 in Cyan) . Even then the total token usage appears to be reaching a steady-state.
https://openrouter.ai/rankings#top-models
^ the first graph is tokens per week across all models
I guess we just can only throw ideas at an LLM at a certain rate.
I still have ideas and now I can have an LLM vibe code what I want, but I'm not going to let an agent just run unattended for longer than a few minutes or a few bucks for hobby projects.
So maybe it is a matter of lowering the cost of an LLM so I can let it churn for hours at a cost of pennies... But I suspect demand for tokens is very price-elastic.
- lower input/output token pricing
- the cached token price is $0.0028/Million tokens, which is like 50-90% of tokens
Famously, this is also a problem for human coders in sprint planning.
HN could be run as a BBS on 70's hardware. Instead of using a CPU with ~10 thousand transistors, you're likely using one with ~10 billion to do basically the same thing, and you don't think twice about it.
You can actually use Luna without reasoning (set it to "none"). So if they wanted to, they could definitely replace 5.5 Instant with it.
Stop rejecting what we have been collectively telling you guys! LLM providers are profitable, and have been for awhile!
- GPT-5 high: score 35, approximately $0.37/task
- Luna medium: score 38, approximately $0.01/task
- Luna max: score 51, approximately $0.042/task
So Luna medium is:
- slightly more capable than GPT-5 high;
- approximately 35–40× cheaper per benchmark task.
And Luna max is:
- 16 Intelligence Index points better;
- still roughly 9× cheaper per task.
This reduction was possible within 1 year.
https://files.catbox.moe/csxl32.png
(2 cents to run AA index, score 40)
Looks like OpenAI broke the pareto frontier on the trust-me-bro benchmarks!
(One has to wonder if they used any of the neat tricks from the DSV4 paper :)
My company checks the models and pays for Opus through AWS.
You still send the WHOLE context of whatever you want to do to a random endpoint on the internet. If you want to write a good email, you give that context your email address, names, the reason for it etc.
Big companies don't randomly use some random api endpoint to do so.
Anthropics quarerly revenue is still growing very fast. I don't think we have seen even the real potenzial of it yet at all.
Not only are still a lot of countries missing which do not even use anthropic or any other frontier model yet but also all the agentic based solutions enterprise companies are currently building on mass (at least in my industry)
I think there's also a new generation of hardware in the past year or so tuned specifically for LLM workloads, where it was almost an accident that GPUs worked to run LLMs before. So, while there's still this ridiculous shortage of hardware, what is being delivered is much faster and cheaper to run for these specific workloads.
I wasn't expecting it to happen from the US vendors, though, as they've spent so much capital to get to where they are they need to make huge margins on inference to pay it all back. I expected the Chinese models who're running much leaner operations to be the "frontier" on costs (and they have been). But, I'm glad to see OpenAI joining the "cheap and cheerful" models party. There's a lot of work in that area of capability. Probably most work people are doing falls into that area of capability.
But yeah I do'nt want to know what Kimi 3 is pushing buttons inside Anthropic, OpenAI and Google.
Besides any floor: For every year the tokens get faster and cheaper, we will see new things like properly working AI factories which mimic expert teams. A lot more parallism as well.
Now we have an American lab drastically cutting a price, feels like this is the opposite of that trend.
Just imagine how much the same level of intelligence cost only a few months ago.
> China: Household rates average around $0.08 / kWh (¥0.53/kWh).
vs
> US: Household rates average around $0.16 / kWh, though regional variation is massive—ranging from ~$0.10/kWh in low-cost states (like Washington or Louisiana) to $0.30–$0.45+/kWh in high-cost areas like California or Hawaii.
The dynamic GPU clusters will be used for the 5% of tasks, and pushing out the frontier. Also there will be a set of knowledge tasks that are not done today (because they are too difficult for most knowledge workers), that will start being done in the future.
Assuming the efficiency gains are real, I feel like something has to give, maybe worse quality due to aggressive quantization/kv cache compression?
If those numbers are accurate, I don't think 20% is a really, really big deal. It's like saying "we're digging our grave 20% slower." Ok, but they're still digging!
Or, different analogy, if I'm going broke because I lost my job due to executive AI psychosis, cancelling my netflix subscription doesn't really change the math of not being able to afford rent. It doesn't even really slow it. The amount of money that OpenAI is spending is so absurd that a minor cost saving is like, uh, some progress, but they'd need to do it a lot more to move the needle
If somebody already thought the price was subsidized then this price change would just mean it's even more subsidized, so why would that change their mind?
But if this is even at 400B size it's insanity those inference prices, maybe 10-20% margins, if it's higher I would like to know is it their own chips or maybe they have accurately sized the model to fit on exactly a B300?
Could be a lot of magical things we can only speculate, but from here there likely isn't another 60-70% margin, like I have heard people claim, I would definitely be willing to bet on that.
Could still be a healthy 10-30% margin. Especially with Terra.
Although I'm sure there are some efficiency gains, the technology is too new and labs are scrambling to release too quickly to think that the low-hanging optimization fruit has been picked already.
It probably costs them a substantial amount of money everyday to not replace Instant with Luna, and the people who want larger models will upgrade regardless of whether they get Instant or Luna on the free tier. It is unlikely the amount of people upgrading just to have latest models would be enough to offset the costs.
[1]https://artificialanalysis.ai/models/gpt-5-6-luna-non-reason... [2]https://artificialanalysis.ai/models/gpt-5-5-instant-05-26
Haiku 4.5, on the other hand, is comparable to performance to Gemma4 31B (with working tool call formatting) in my experience, and Gemma4 strongly wins on vision and multimodal.
Estimated final electricity price for large industrial customers in energy-intensive industries:
USA 50 USD/MWh
China 68 USD/MWh
With this new price change, Terra does look pretty Pareto’ed by Luna.
On agentic coding, pairing Sol Medium for architecting with Luna High for coding does kinda make sense. But beware that architecting can be very read-heavy, and Sol is a bit read-pricey compared to Terra.
presumably it's a much bigger model
There is ZERO reason to believe models 1/10th the size of frontier are completely capped on intelligence and impossible to get smarter.
They have consistently compressed the intelligence of larger models.
You'll see it first on the small end, when they stop being able to compress intelligence, you know that will slowly bubble up and up the chain to larger and larger models.
There's no evidence we've reached that at the bottom.
[0] I am constantly surprised how much work pay-as-you-go with DeepSeek / MiMo will get done. I've barely crossed $2 each in a month of use (~200m tokens).
MI500 series is supposedly already taping out and they're claiming massive increases (we'll find out end of 2027 prob).
Personally, I'm having surprisingly good results with DeepSeek 4 Pro at home, which is very good value for money: it's not as good as Claude / GPT 5.6 (I have Co-pilot license at work), but it's still really useful for code reviews, validating thoughts, and especially designing / writing unit tests for new (and old before refactoring) functionality.
And it's very cheap per task. (Flash is even cheaper, but I've had issues with that on more complex tasks where it starts forgetting things and arguing with itself "but wait, let me read the function again").
This way the expensive/strong model only handles the architecture and orchestration tasks. The cheaper models handle everything else and the strong one knows how to tell them what to do in enough detail to get good work out of them.
Taking actions that mutate the environment is a different story. I think this is where you run into diminishing returns very quickly. You generally want one strong agent to act given the results of all the searching that was done. If the plan is clear, you don't need a genius model to execute it.
Anthropic's big marketing push this year has been entirely focused on getting people to use Opus via a Claude Code subscription, to the point that Sonnet is almost viewed as the poor man's alternative, and from what I've seen, almost nobody uses it.
Actually, here's an interesting project for all the vibe coders looking for their next front page post: scrape a ton of commits from GitHub with Co-Authored-By: Claude and figure out what the percentage split between Opus/Fable/Sonnet is. I'm willing to bet it's less than 10% Sonnet.
I have no doubt that further work was required to enable this, but it's still very cool to be possible to say that.
Edit: searching for the story now, further bolstering the point is that was 1% in training time [1], and the openAI claim is 20% in end to end inference cost. This is a bad comparison.
[1] https://deepmind.google/blog/alphaevolve-a-gemini-powered-co...
Phase 1 - Run X copies of Luna in parallel over the user's prompt. The purpose is to generate a diverse set of hypotheses.
Phase 2 - Run Y copies of Terra in parallel to investigate the hypothesis results, with each receiving them in a randomized order.
Phase 3 - Run 1 copy of Sol over investigation reports.
The goal is to ensure that the agent covers more initial starting points before presenting a final conclusion. If you only run a single copy of Sol and it hooks onto something wrong, it might not recover.
https://api-docs.deepseek.com/updates/
The update seems to say that it's a re-post train of exactly the same model.
https://developers.openai.com/api/docs/guides/latest-model#p...
Yesterday, we shared how GPT‑5.6 helped make itself more efficient to run. Today, we’re passing those gains on to customers with lower prices for GPT‑5.6 Luna(opens in a new window) and Terra(opens in a new window) and faster performance with GPT‑5.6 Sol in the API. Together, these updates help customers get more from every dollar they invest in AI and move faster when time matters.
Starting today, GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less, while GPT‑5.6 Terra, our balanced model for everyday work, will cost 20% less. These lower prices for Luna and Terra are also reflected in how usage is counted against paid subscriptions when using Codex and ChatGPT Work. Luna gives businesses a far more cost-effective way to handle high-volume work at very high levels of quality. It can use tools and complete multi-step workflows, making a broader range of AI applications practical to run at scale.
Making advanced intelligence more abundant and affordable is central to OpenAI’s mission to ensure AGI benefits all of humanity. These changes put that commitment into practice. They reflect years of improvements in how our models are built, served, and put to work.
We’re also introducing Fast mode in the API, which replaces our Priority Processing offering. For GPT‑5.6 Sol, Fast mode now delivers up to 2.5× faster speeds than Standard processing at twice the price, with no change in intelligence. Fast mode is backward compatible: requests tagged priority will automatically use Fast mode.
1 of 6
Using AI efficiently begins with the outcome. The stakes, cost of error, urgency, and scale determine the right balance of intelligence, speed, reliability, and cost. That balance can change from one step of a workflow to the next.
GPT‑5.6 gives businesses much more room to optimize that equation. Luna delivers performance comparable to models that were frontier-class a year ago at roughly 6 cents on the dollar per task, and at nearly nine times the speed. On professional work, as measured by Agents’ Last Exam, Luna outperforms Fable 5 at an estimated cost per task nearly 99% lower.
In practice, businesses can define the outcome and quality standard they need, then use evaluations to determine where additional intelligence materially improves the result and where faster, lower-cost processing can deliver the same quality. A coding workflow, for example, might use Sol to resolve uncertainty and define the plan, then use Luna to implement well-specified changes, write and run tests, and evaluate the results. Another workflow may call for a different balance.
The GPT‑5.6 family expands the range of those choices. Businesses can apply the maximum useful intelligence at every stage while paying the right price for the value it creates.
Delivering that flexibility starts with making every layer behind the models more efficient.
Our efficiency edge comes from improving the models, the inference systems that run them, and the agentic harness that connects them to tools and context. GPT‑5.6 models take a more direct path through work. Better routing keeps hardware productive, optimized production software generates tokens more efficiently, and smarter context management helps agents avoid repeating completed work. Together, these improvements let us complete more useful work with the same compute, reducing the time, tokens, and cost required for each result.
GPT‑5.6 Sol is increasingly helping us find and deliver the next round of gains. Within a human-led process, Sol autonomously rewrote and optimized production kernels, designed and ran hundreds of experiments to improve token generation, and monitored training, intervening when problems arose. The kernel work helped reduce the end-to-end cost of serving the model by 20%, while its experiments increased token-generation efficiency by more than 15%. This work continues, creating a tighter feedback loop: as our models improve and are able to work more autonomously, our ability to improve efficiencies accelerates. Read more about the engineering behind GPT‑5.6.
Meeting demand for abundant intelligence requires both more compute and more productive compute. We are building a resilient infrastructure portfolio and matching each workload to the systems best suited to run it. That approach supports both ends of the price-performance curve. At the lower-cost end, the new Luna and Terra prices make high-volume work economical at much greater scale. At the frontier end, Fast mode gives API customers faster access to Sol when response time is important.
Enterprises can move more AI into everyday operations without sacrificing speed on their most consequential work. Large-scale document analysis, customer-interaction classification, and routine implementation can become economical to run broadly, while complex Sol workloads can move faster when the premium is justified.
The gains can compound. Within a human-led process, more capable models help our technical team find the next generation of improvements, shortening the path to better performance and lower costs. Our strategy remains focused on advancing both capability and efficiency so each generation of intelligence can accomplish more work at a lower cost.
GPT‑5.6 Terra and Luna remain available in ChatGPT Work, Codex, and the OpenAI API. In ChatGPT Work and Codex, Free and Go users can access Terra, while Plus, Pro, Business, and Enterprise users can choose Terra and Luna.
Starting July 30, API pricing is $2 per million input tokens and $12 per million output tokens for Terra, and $0.20 per million input tokens and $1.20 per million output tokens for Luna. Sol pricing remains unchanged. ChatGPT and Codex subscription prices and quota budgets remain unchanged, while Terra and Luna usage now consumes fewer credits. Pricing changes will begin rolling out in AWS later today.
Fast mode for GPT‑5.6 Sol replaces Priority Processing in the API and aligns with /fast in Codex. Existing API requests tagged priority will continue to work. View complete API pricing details.
We run an agent company and outside coding the new Gemini 3.6 Flash and GPT 5.6 Luna are very interesting. Luna can do a bit of research and create reports. Gemini is great for computer use.
For programming it's all Kimi K3 now.
All we need now is some sort of program to evaluate halting problem oracles...
https://chatjimmy.ai/ is their interactive. Tiny context, very dumb, but absurdly fast. Imagine this as a tool call for claude code for trivial changes - the tool call from the harness takes longer than the execution.
Blocking Fable for sure made it very politicl a lot sooner than i expected it to happen.
and because China already has massive problems of getting access, they are pushing it on hardware too like what Huawai did without EUV.
It seems China is already able to do DUV a lot sooner than others expected.
I feel perfectly content in using pay as you go pricing with deepseek. On the other hand, although Anthropic's models used to be my bread and butter for personal work, they are simply too expensive to reach for these days.
https://artificialanalysis.ai/models/comparisons/gpt-5-6-lun...
How so? First, kernel writing (or ML engineering more broadly) is a highly specialised task. Not everyone can do it. It shows that models are getting better and better at (easily verifiable) hard tasks. And you can "hire" that expertise much easier than you can hire the equivalent meatbags. And more importantly you can "fire" them as soon as the task is done. And then hire them 3 months later, when the new model drops. And so on.
Second, 20% gains in inference today gives better end results (i.e. lower overall cost) than 1% in training 2 years ago. Today's models are improving mostly via RL. And RL is highly dependant on fast inference (you want many rollouts for each training scenario). Same for dataset filtering, environment generation, distillation, etc.
Luna is an extremely strong model.
why on earth would you suppose that?
Or, when we will start doing this, who's going to be able to do that in scale?
I'm seeing the TAALAS example, but it's only an 8B model, suggesting some real limitations parameter wise. And for 2.5kW?
You use the big models to plan. Not just the overall plan, but which files need to be edited etc. Then they give that to the lower end model. So Luna or Sonnet, which are perfectly capable of following instructions and still creative enough to not get stuck.
That's the media and in particular US KOLs of all sorts driving the wrong impression of China and other places. China and many other places for example have fast public transport that the US doesn't and can't even imagine today. They're not behind.
China's DUV still isn't that production grade (mass produce-able) so don't get that hyped up the wrong way (in a different direction).
The whole China-is-behind with tech and in particular semi wasn't that they can't. The truth is they spent decades in internal politics and corruption. That all got solved with the bans, so thank the bans! Jensen even said the bans were bad.
These companies are printing money on inference. The only issue is the capex of expending on new servers but they are generating TONS of income.
Morph occasionally have lower prices than the standard rates, and:
https://telnyx.com/pricing/inference-api
Is one which is a bit cheaper... I haven't actually tried K3 myself...
No, I'm using it via OpenRouter in pi.dev - I just used it 30 mins ago... Providers (automatically selected): StreamLake and Baidu Qianfan.
They obviously load shed a bit of Codex-sub during peak times, and for the amount of tokens you get for a sub, I don't mind. I just mean the API where you pay-per-token is rock stable.
If we get to anywhere near this speed for the equivalent of the current models... I don't even know what to think about that future.
I pressed Enter, and the response was instant.
> Generated in 0.037s • 14,205 tok/s
This is unbelievable.
I just tried it too and 14,098 tokens in .05 seconds, I barely blinked and it was done. There was no typing at all appearing on the screen. It just showed up.
https://chatjimmy.ai/chats/01dc66a4-4b1b-4dea-bb5f-926855e37...
https://arxiv.org/abs/2512.24601
There's also a great write up here by the author:
By benchmarks, which sadly is a poor measure. Yes Luna is a good model under certain circumstances. Whether it is great for general usage is another story. Sonnet is definitely better when prompts are more vague and it needs to decide things. Luna generally sticks to things very strictly and goes off in bad ways.
The "edge" AI landscape (in particular, what you can do with ~5W) is going to be nuts in about 18 months.
```python
source_code = open(__file__, 'rt').read()
response = ask_llm("will this program halt, only answer 'yes' or 'no'?\n" + source_code)
will_halt = response == 'yes'
while will_halt:
continue
```Once speed significantly increases I think we're going to see some interesting downstream effects. The three things I currently spend the most time waiting on are LLM API requests, Rust compile times, and nix derivations. As AI latency approaches zero I think we're going to start taking a hard look at whether slow-compiling languages are adding enough value over Golang, Typescript, or even dynamic languages to be worth the slowdown.
This is crazy.
RLM might be more useful on the execution side than on the research side. In fact, these somehow feel like they might be exact inverses of each other in terms of what the ideal architecture looks like. At some point you definitely do need something in the middle that has it all sorted out.
That's about how disrupting DSPs were to the industries they arose out of (over a very long time frame).
How would that disrupt the industry?
The big AI labs won't do that unless they are forced to, as they want you to spend more money on the big, expensive, frontier models (so they can live up to their valuation), so it's more likely that you will see this on smaller open weights models.
(I've been using it via OpenRouter and it's much more than that, but still cheap).
I've previously found flash (for all the hate it gets) to be good for these kinds of things. Haiku was fine but it's ancient.
A Principal Staff Engineer who costs $2400 a year and never feels any work is beneath them? Hell yeah.
OK OK, usage limits
That's again not some "intelligence factor" here. Different agents work for different use cases. Luna wins some. Terra wins some. Sonnet wins some. Flash was really good at exploring.
So I'm not sure what your point is? There's a big market for everything. Even within the market you describe it's likely not a Luna-size fits all either.
> Cost of Revenue: $2.65 billion
That's how standard accounting rules for public companies would measure it.
and it gave a very reasonable answer in non-perceptible time.
I’m still trying to figure out coding agents. I can’t even begin to imagine the things it would enable. Even the most mundane ideas like LLMs-in-HiFreq-trading have huge implications.
I almost exclusively use it with xhigh or max effort, but when run like that it's been an incredibly cheap little workhorse for most development work. I'm still leaning on Sol for planning and debugging, but when it's time to start pumping out code I've been leaning into Luna (Max) and I've been enjoying it! And that was before the price drop, it's going to feel practically free at this point
The "halting problem is unsolvable" argument relies on the oracle not being able to output "not sure". But adding that option admits trivial oracles, like ones which output "not sure" for everything, so some are better than others.
The "real world" use most people have for halting oracles is as part of software safety, where if the checker outputs "not sure" you modify the software until the checker can decide if it halts.
Here's a copy and paste prompt if somebody wants to just test it real quick to see what I saw:
Write a story about the fastest monkey who ever lived, his name is Jimmy and he is an AI superbot monkey that is part cyborg primate. He can travel through time and is psychic.
The rest of us need to optimise a bit.
The only thing that will crash prices is reduced demand (duh) or, more interestingly, increased production. In particular, if CXMT is able to get their DDR5 fabs up to a reasonably high yield, that could add some downward price pressure (as could government subsidies). As well, if Micron/Kingston/Hynix think that CXMT is going to start cutting into their market share, they might be willing to either increases supply or drop prices. Unfortunately CXMT looks to be taking quite a while to get their new fab up to max capacity so that may take a year+ before anything manifests.
If you're interested in following the (publicly available) info on these sorts of things, check out what companies like Axelera, DeepX, and MemoryX are doing today and have on their roadmaps, as well as the sorts of chips/SoCs Qualcomm, Kinara (now NXP), and Ambarella currently have announced (or have on the market). And remember, that pretty much all of these chips on the market today were in initial development more or less when ChatGPT first launched. If you knew what you knew today (or a year ago) about what requirements current- and next-generation models would have (from a silicon perspective), what might you do differently? Think for instance, host system interconnects, amount and speed of on-package or on-die memory, image/video decode capabilities, int8 vs fp8 vs fp16 vs bf16 compute units, etc. And, consider that most "AI" stuff in development a few years ago was all 15nm or 12nm - because who was gonna pay big money to get fab capacity at 3nm to run some object detection models? So most of the stuff on the market today is on very old nodes and therefore not super power efficient.
I find myself getting caught up in the sheer speed of modern computing and networking. The fact I can play an online game with 10 other people is just insane.
Knowing that functions terminate is important for proof languages like Lean, where you often want to prove things without running the code at all. You're proving that one could, theoretically, calculate an answer, without actually calculating it.
Things like this give me hope for a system that can be fully local and private, but also with the ability to be almost infinitely extendable with tools.
The other thing that I think is really interesting about all of this, is that LLMs are already perforce behind the times with their knowledge cutoff, so adding an additional ~3 months for bake into silicon isn't such a huge deal, I think, for the ~10x more efficient and faster you get.