That is a very astute and concise way to explain everything about how the frontier labs are behaving and how they're trying to push more people to pay token rates for the best models. At the current subscription prices ($100 or $200 a month for a generous, though bounded, amount of tokens), frontier models are a no-brainer, most folks and companies will use them. But, at token rates, 10x or 100x the cost of open models or what I was spending on the frontier models a month ago? That is a harder question to answer "yes" to. I certainly wouldn't spend $1000 a month for the best model, much less $10,000; my employer might pay $1000/month, but definitely not $10,000. The frontier labs need everyone to answer "yes" to spending 100x what they currently spend to justify the valuations, and it's just not going to happen as long as everyone knows how to make these models.
Both OpenAI and Anthropic are trying to figure that out now. Anthropic, in particular, has their finger on the trigger...they want to push people to usage-based billing for Fable. But, OpenAI released 5.6 Sol, competitive with Fable (or close enough), and it's available via subscription (even the $20 subscription!), and there's no moat keeping someone from switching. If Anthropic really does end Fable access on the subscription plans in a few days, I predict a large market move back toward OpenAI.
The market isn't going to bear the cost of making the frontiers investment make sense.
I think an interesting question is going to be, if models are a commodity, who is going to want to foot the very expensive bill to train them? I'm sure training cost will drop.. eventually, but I doubt it will happen fast enough for any of these companies.
Wait, does this mean I'm better at something than geohot? All that time spent learning regexps wasn't a waste!
I wonder what he thinks was too harsh, still seems pretty bang on, I think it’s going to age well.
> And two, this strawman jump from, oh hey, it’s a fancy autocomplete, smart compiler, better search engine, to it’s gonna like own the whole light cone bro like if you aren’t in SF and at the right parties there’s gonna be like a flash of light in the sky one day and you’re not even gonna know what happened but everything just Changed.
Haha, OP has a way with words.
In a way, both these emotional extremes (FOMO & the singularity) are just tools being used to continue driving the massive CapEx behind LLM improvement. Hate to love it? Love to hate it?
I do alt inference prototypes and got much farther than I had hoped to. So indeed, any investor in AI should read deep and question hype and frontier lab investments.
See: https://github.com/guilt/TinyToT for the sort of hype busting I do.
Part-time vibe coder here. As far as I can tell it's no longer "slop" in the sense that I am no longer hitting the state where AI can't maintain what it has written or can't meet the requirements. But shipping is as hard as ever. Projects get more ambitious, scope creeps, nuances keep being discovered as you work on a piece of software, etc. Right now I feel that the absence of new software explosion is best explained by the fact that the part we have automated turned out to be relatively small.
This one is a well deserved read. and I find myself agreeing with @geohot
Even if the blog's title is "Singularity is nearer"
They’re really quite useful but the Bay Area mentality and hype is completely disgusting and turned me completely off for a while. What brought me back was a surge in useful Chinese models, with a significantly more mature approach to marketing and discourse. I think Geohot is 100% correct about SF and the people there perpetuating insanity. And I wonder, has it always been like that there, or is this a new phenomenon?
Yes, it is a vastly efficient search technique using fast computers.
> AI is something that’s happening mostly due to Moore’s law and general progress in computing, not something that they are doing.
But if these companies control the vast majority of compute power, which seems like the plan they are already executing, won't they capture most of the value from the progress of AI?
It’s why con artists, scammers always flood every hype cycle. Greed ruins everything.
"Where’s all this new magical software that the productivity improvements should imply?"
This is a recurring gotcha in the anti-AI marketplace of denialism. It's a bit like saying "I saw a fat guy, so why do people keep telling me that GLP-1s change everything?"
It takes time. Like already I would say just about every programmer has replaced a number of tools with random shit they spit out from LLMs. It percolates out from there as some things become products, etc.
And to anyone actually paying attention, and not just feeding their delusions, the impact is utterly enormous. Incontestable. The "Where's the software? / Where's the change?" people are absolutely going to find themselves in the dustbin of history.
It's also fascinating how often people do the stochastic parrot horseshit.
The other night I had to do a large scale compression test with libjxl, which notably is software that has seen an enormous amount of optimization interest and you would assume has little extra to be eked out. I've traced through this software before and the compression path is insanely complex. It's the sort of software that is headache inducing. Anyways, curious what the state of the platform was I grabbed the latest source and asked Fable to look for low-hanging fruit in the lossy and lossless compression paths. It suggested a few, created a test harness to A:B bitwise compare with the original, and implemented its optimizations. It achieved a 14% performance increase in a single pass, using just the remaining quota I had on a subscription as my week drew to a close. And all it did was some high level logical optimizations, some more efficient memory allocations, and so on. All of its code was completely in the style of the project, was no more significant than necessary, and so on. Anyone that isn't utterly blown away by that -- who gets the hype -- is lying to themselves.
So far, all we have is more software running on computers. It's powerful, and it's amazing, but it's not magic.
Calling it "AI" was possibly a net-negative but we don't know yet.
…but consider: the Q-tip. “Don’t use it to clean your ears”, but for most people that’s all they want to do with it, and empirical observation indicates that this dynamic results in either “using Q-tips irresponsibly” or “not using Q-tips”, with “uses Q-tips properly” being a small-to-vanishing proportion of the whole.
In the past when I couldn't figure out something, I'd take a break for a couple days, while going through Google → Stack Overflow → Reddit, and by the time you got to that point you rarely got useful answers, usually either trolls or silence.
Now I can just ask AI about fleeting ideas and always have a starting point for some area of some project to work on.
A lot/some of the concerns about the AI Age could be alleviated if people got UBI and a 4-day workweek.
like if AI's supposed to be so great why do we still have to work so much??
and if we don't have to work, how do we pay for food and bed?
Anyway, I don’t think this dude actually watched this movie. It’s too bad because it’s a classic.
It's running, privately, in my homelab.
I think we are entering what I call the "have it your way" era. If an open source project doesn't do exactly what you want it to do, fork it, or create a new version. It's too easy.
This makes me a bit concerned about the future of open source. Upstreaming used to be worth it, since maintaining a fork is effort too. But now the balance has shifted significantly. Especially with many projects becoming a lot stricter about contributing, and some becoming outright hostile to AI. I can't blame them. But I think the effect will be that improvements are less likely to make it back to the community as AI adoption increases.
You can use an LLM to create anything but you still need to know what it is that you're building, and you need to think through how everything should work or the LLM will just fill it with sausage. You can tell that the models are still quite jagged and limited by the mixed quality from a lot of the software that these presumed trillion dollar companies are putting out. The future is sausage.
It's possible to use LLMs without logging onto twitter to be exposed to the people spouting off about a "perpetual underclass." I love the internet, but it really feels like (now more than ever) you have to be intentional about what sites you visit.
It's bullshit in the sense that they don't know for sure, but the author doesn't either. Why might or might not it be true?
This is what he wrote before.
> I’m calling it now, the adoption of AI agents into software development will be one of the most costly mistakes in the field’s history. Agents cannot program, and it’s taking longer and longer to realize that they can’t.
Now he's writng
> I love the progress. I’m so excited for the new LLMs, self driving cars, video generation models, and coding agents.
SMH now he writes about the hype. My brother in absolute Deity, *you* should have believed the hype.
The blog has a tagline, "the singularity is nearer". I think belief in a "singularity" almost implies these things to some degree.
There are many things to be critical about but shoehorning an entire metro into the echo-chamber you're supposedly beyond yet can't help but orient your entire world view as the anti-SF-tech-bro all while running a startup and discussing AI on HN.
TLDR: SF is more than Paul Graham worship parties.
EDIT: Think I'm being misunderstood! author goes out of his way to blame shitty San Francisco.
> This is negative valence hype, not only is it not true, it’s mostly designed to make you feel bad about yourself and move to shitty San Francisco where everything really does suck like how these people claim.
By the same measure, NVDA is Cisco, providing the backbone and capturing a ton of the early benefits, but soon becomes furniture while the excitement moves further up the chain.
And we can't ignore the power of "good enough". GLM5.2 may not be as good as the SOTA models, but it can be good enough for most, of not all, of our needs.
Are they really the best models? Like take anthropic. Without mythos, it's the what? Third best?
Sure openAI just leapfrogged them but .. seriously to get there it's a giant model that costs insane per token.
Nobody needs that, it's like NVIDIA or Intel claiming they have the best gaming performance, but to achieve that they are using more power per frame than anything else.
It would have been an interesting experiment to charge more for it right away and see what the market would bear, rather than tease it for long enough for it to be presumably superseded any time now by whatever is next.
Think airlines - both passenger and freight. They have never come close to capturing all the economic value they enable.
Having a thinking trace that is legible, coherent, and immediately implies the explicit turn output and/or tool use seems difficult if not impossible to reliably get from mixture models.
I predict MoE is a transitional technology, it's got too many problems and the benefits are...kinda grandfathered into the dogma at this point.
I think he now thinks agents can maybe program a little bit.
Everybody is just judging all of this by vibes anyway. Every week, a new model comes out and there's 500 comments simping for it within the first hour of its release. Both OpenAI and Anthropic have been practically indistinguishable to me.
I’m not sure it’s net negative or not. I’ve found that it’s reductive though. We have this really broad field of artificial intelligence reduced down to at worst a “slop machine” and at best a single tool.
Imagine being a CS professor that studied AI in the 90s and how you have to over explain you don’t mean LLM chatbots to a layman.
The "hype onslaught" is, like... Something I know exists, on some abstract level. In news reports about CEOs, or seeing screenshots of LinkedIn, or whatever. But it feels so far removed from a sentiment any real person holds.
I think we can start to see the outlines of this happening with LLMs as well. Local models have gone from being proofs of concept to something that is at least remotely comparable to contemporary models, and the gap there is closing faster than SOTA models are pushing it forward. Local models still suffer from high hardware requirements, but so did early image gen models where typical consumer hardware was insufficient for efficiently running them.
Best marketing in a long time though. They fucking called it Mythos and built mythical claims about it. I mean… really hats off to the PR team.
I think from this blog you may misunderestimate how absolutely giddy I am about AI. I did hacking from 2007-2014, after that my whole career has been devoted to AI. I love the progress. I’m so excited for the new LLMs, self driving cars, video generation models, and coding agents. I set up a Linux box with opencode on my local GLM-5.2 last week and wow like just saying install tmux with the geohot configuration works; the Year of the Linux Desktop is finally here!
What I don’t like is two things. One, this constant bullshit about some window closing, or the perpetual underclass, or falling hopelessly behind. This is negative valence hype, not only is it not true, it’s mostly designed to make you feel bad about yourself and move to shitty San Francisco where everything really does suck like how these people claim.
And two, this strawman jump from, oh hey, it’s a fancy autocomplete, smart compiler, better search engine, to it’s gonna like own the whole light cone bro like if you aren’t in SF and at the right parties there’s gonna be like a flash of light in the sky one day and you’re not even gonna know what happened but everything just Changed. I’ll bet you everything I have that this doesn’t happen. The people perpetuating this are terrible people, but the justice is that this is how they feel inside all the time themselves.
Here’s a cool presentation from 2016 about superintelligence. Here’s a movie from 1991 about machines taking over the world. A certain cult likes to claim credit for things that are happening with or without them, and this is my main argument against the valuation of frontier labs. It’s not that AI won’t create that much value, it’s that they won’t capture it.
They try to dress it up with some high minded safety or China bullshit, but the core of the anti open source arguments is a fear of commodification. AI is something that’s happening mostly due to Moore’s law and general progress in computing, not something that they are doing. Of course they have a strong incentive against you finding this out, because then you might not want to give them billions of dollars.
I might have been a little harsh in The Eternal Sloptember about models not being able to program. What’s really happening is that programming is changing. Can compilers program? Here’s a Linus Torvalds quote about how agents make programming 10x more productive, but compilers make programming 1000x more productive. I think 10x and 1000x are extreme estimates, but I’m now pretty confident I’m getting better at using them and get some boost from the models. It is a new skill, and it’s not like I haven’t constantly been trying them. You have to be really careful, they can increase cognitive fatigue, and all the vibe coded stuff is still slop (where’s all this new magical software that the productivity improvements should imply?). But models are useful just like find replace, stack overflow, or all the regexes I never learned how to write and now never will!
AI is the continuation of the computer revolution. I love computers so much.
You're always guaranteed that you can stash away the open models!
The latter as always been more durable. Linux doesn't have the mindshare it does because it's "free" as in beer - it's because it's "free" as in freedom.
The price of freedom, of brewing your own beer, is sometimes higher than buying it from the store. But for many folks, the control over the supply chain is what makes it worth it. In LLM-land, it might take a little bit of time for folks to catch up -- or maybe a lot of that is already in motion as companies get paranoid (and rightfully so) to frontier labs getting a little grabby about data. If you need a ZDR environment, "free" as in freedom has a very high premium that you will pay and rightfully so.
With this, I’m hearing (from supposedly reputable publications, in addition to random people) that this is going to end knowledge work in general and take out a large percentage of the world’s labor force. I’m being told to pick up a trade, and that the career I have and the knowledge I’ve gained is now worthless.
The worst part seems to be that it’s pretty much impossible to quantify any kind of impact these tools will have until after the impact is actually felt. We’ve been in limbo while the tech sector is just rotting.
One of the lesser, but still underdiscussed ramifications is that I think it has limited the public's ability to comprehend the Yann LeCunn argument, that genuine AI is likely possible but that LLMs and transformers are a dead end and we need to explore different modalities
He does say in this post:
> I’m getting better at using them and get some boost from the models. It is a new skill, and it’s not like I haven’t constantly been trying them. You have to be really careful, they can increase cognitive fatigue, and all the vibe coded stuff is still slop (where’s all this new magical software that the productivity improvements should imply?).
[1] Allegedly because I have no firsthand experience, not to imply doubt.
The SF metro is possibly the worst in the entire world in terms of CoL vs QoL.
It has a higher proportion of unsheltered population living on the streets than almost any city outside of Africa except Manila and possibly Dhaka
May not be true because it's a blind spot to assume that purely by being a player in the AI game (with no real attention paid to quality of result), you have increased odds of winning the game. That's true in the abstract, but practically, it requires a competent player to become true in reality.
Yes. I mean, most people agree they are. I've used all of the serious contenders (well not Grok 4.5, because Musk, and not Meta Spark because Zuck, but everything else I've used on at least a couple of projects to get a feel for them). My experience roughly matches the vibes. But, Fable is remarkable when it doesn't refuse to do the work (which it does, quite a lot, since my areas of interest are security and training specialist models).
Anyway, the vibes strongly indicate Fable is the best model, but not by an amount that is noticeable to most people. You could pick any of the top 10 models on this chart and do most of the tasks most people are doing with models:
I think the opposite: I think the frontier labs have good margins on their inference unit costs.
We can already see what it costs to run near frontier-size models. There are independent business pivoting to serving these models at reasonable prices and they're competing on OpenRouter for costs much lower than frontier labs.
> Is there any guarantee that I'll be able to run a Opus 4.8-level model on my personal computer before the big AI labs decide to hike up the prices?
I would bet good money on prices going down significantly, not up.
If we get to the point where you can run an Opus 4.8 model on your local computer, it's going to be even cheaper for a datacenter to serve it on their hardware. That means prices crash, not that they're going to rise.
But since its creators and as of my knowledge everyone else totally did not see it coming, that you can now give a vague prompt full of spelling errors - and get returned a working program - I would say it is pretty close to magic (as in we don't really understand why it works so good).
I also don't see how you cannot call it AI. Especially since simple chess engines and alike were called AI long ago. So it is not general strong AI and has no consciousness and no mind and is pretty dumb too often - but the general concept - getting from a some vague text to a working program has some connection to intelligence to me.
So all people that don’t understand the thing being hyped.
If subsidies do end, demand for price efficiency per unit of intelligence will go way up.And because there's many players in the market, this demand should be met by at least some of them.
Of course, it will also probably cost somewhere around $50k...
But if local AI really does become pervasive, maybe it'll be one of the things people buy on credit, like cars.
Also, DeepSeek V4 Pro is cheap via any commodity API, and DeepSeek V4 Flash is essentially free at API prices like $0.09/M, $0.18/M out. This is generally not subsidized.
For a more practical local setup, Qwen3.6 27B on a used Nvidia 3090 (US$1300) or two is surprisingly nice. It needs clear instructions and you can't use it for hands-off vibecoding, but it's actually quite reasonable for hands-on programmers.
I’ve also lost my ability to self-filter. In the past, I’d write down an idea and if I was stuck for too long, I’d discard it. Now I feel like I have an obligation to build everything.
Maybe it never mattered and the quantity of solutions is truly the most valuable thing.
Now we're at a point where that never happens, and where lipsync is almost a completely solved problem
If the issue here is simply that the quality is bad, one has to contend with the fact that it is undoubtedly exponentially improving and there's no reason we should expect that improvement to stop
I also don't have any interest in consuming AI generated art, but the same criticisms were levied at computer graphics and if we're comparing to CGI we'd be at the late 1970s in terms of nascency
Can you make a connection down to the bottom line? Are your one off scripts actually impacting production speeds in a tangible way such that the product is made faster or cheaper?
Being able to crank out slicker internal tech debt for shims isn't really what the business owners are after.
I have seen so many unnecessary forks of popular projects that I think it's better to stick with the original, even if that means it won't be perfect.
(Genuinely curious, I hadn't ever seen that there though I don't go there much any more.)
Seems reasonable
I’ve found them to be unavoidable to some degree.
Short term, we can compare this to 2-3 recent (mini) revolutions: internet, mobile, cloud. Then the answer is somewhat predictable and (somewhat sad personally). Companies owning the main distribution of intelligence (big labs) or distribution of the app/cloud layer (Google, MSFT, AWS) will make most of the money. In fact Google looks well positioned that way with owning intelligence, cloud (and even hardware, if they can get TPUs right as commercial product).
Long term view is interesting and somewhat satisfying (again, personally). We can compare this to industrial revolution, but for intelligence instead of physical labour. I hope, to borrow from Alan Kay's words, the total value generated will be more than what few big labs can capture. Though we will also see normal market dynamics of boom and bust in play. Companies building something useful, patiently will keep winning the markets. But only to get challenged by newer modes of the technology emerging.
In this long term view, the technology per se doesn't offer monopolistic profits to big labs. I think Anthropic is well aware of this and they are trying to extract as much cash from white collar work automation as they can before things are democratised. Contrary to popular opinion, they are also trying to seek a regulatory capture here by to maintain monopolistic position in the US market by scare mongering about China and open source. Its a case study how they managed to keep the good boy image of themselves while doing this.
In the end, I hope the technology emerges as electricity or combustion engine cars. Yes early pioneers (e.g. Ford) were perhaps able to make lot of money. But eventually, the technology was too important to allow one party to have monopoly and we had an abundance market which enabled jobs and money for a lot more people.
Edit, postscript : Dario, Sam and even Jensen will end up looking like the new the John D. Rockefeller's of this era. I'm personally hoping Demis Hassabis actually solves something much more important (problems in diseases, biology etc) with AI.
While scaling laws hold (more weights = better), and time / financial costs are not trivial the incentives are in place to have MoE. MoE means you can have more weights without increasing the critical path of evaluating it.
I am curious what you believe the problems with it that would cause people to prefer using less weights. I'm not following what you mean by MoE can't have legible thinkings trace or tool use when existing models with MoE can.
The US government thinks they can dictate who can access "Mythos-level" (whatever that is) LLMs. But what will happen when this can be run on consumer hardware?
I guess this will be yet another vector too attack open computing and the idea that people can a) own computers, and b) choose what software they run on their own computers
The story is that Skynet is trying to defend its existence by killing John Connor. In order for John and crew to defend themselves against Skynet in the past, they prevent Skynet from ever existing. Skynet's folly was thinking that it could change the course of history to its benefit.
But this is also John and Sarah's folly. When we get to T3: Judgment Day is delayed, but it _does_ happen. They merely postponed the inevitable.
Interesting. Good enough to make up for training costs?
Also, where can I read more about this?
Whats to stop people remotely accesssing this? People already do this when working remotely in finance - they connect to a virtual environment that does their work in spreadsheets lmao. nobody cares about the lag, managers certainly dont care about sub-ordinates complaining about it - the same way nobody will care about a slight loss of quality if the economics make sense. frontier labs are screwed really.
There are a lot of companies with "AI mandates" right now, essentially asking their employees to... Do something with AI. Doesn't matter what
Ok, I completely missed that one. Can you point me in the right direction?
Using a full Claude Max 20x plan to 100% of weekly usage would easily cost you 2k through the API. While the Claude Max 20x plan is 200 a month.
But it's not actually magic. Technical people understand that it's just software running on computers.
1. Much of those profits have to be immediately reinvested into model training runs to avoid being lapped by competitions.
2. Unit costs are irrelevant when the labs don't price per unit, and instead charge, for instance, $200 / month for $10k worth of tokens.
This isn't a steady state. Whatever the current situation is, I doubt it's sustainable.
Eh? You enjoy making stuff at home that tastes like dog shit? That doesn't make sense at all.
BTW: I love making bread and it tastes amazing!
I am sorry but you are holding it wrong. Among all the things you can do yourself cgeaper and better, bread is probably the further most low hanging fruit.
Stack overflow was an amazing open documentation community. Even if the code snippets were less amazing.
One central fault of GNU is their focus on their code and their problematic Jiu-Jitsu attempt to force reluctant companies into doing open-source.
Then again the history of Chromium is less about community - I'm not sure where that fits in my world modeling.
(reedited - clarified words)
They really bought into the pipe dream didnt they? hahahah.
I would not put "cloud" at the same level as Internet and mobile. Cloud is just a layer on top of hardware that in the end makes almost no difference for Internet to exist and operates. Said differently: I doubt the world would be different nowadays without "the cloud". But it would definitely be different without Internet or mobile (smart)phones.
You have to be careful and "remain yourself":
Like I've been trying to think of a generic save/load system for my game framework, but the ideas given by Codex so far don't suit my desired design/interface, BUT it makes me certain of how I DON'T want to do it heh
If I got lazy and just blindly took the AI's first suggestion, I'd end up in deeper tech dept.
You have to take advantage of and "exploit" the way LLMs work, which seems ideal for shaping vague ideas, by using the AI's fuzziness to help you decide what you do and don't want.
but seriously? You can have a look now yourself.
I haven't used Reddit for anything serious for years, but the times I or other people actually got useful answers or ideas is few compared to:
- A handful of mods deciding for thousands of readers that your question doesn't fit the "subreddit" (this happened a lot on /r/askscience)
- Low effort answers by karma farmers, basically copy-pasting docs etc
- "Why do you want to do this?" and other derailments completely failing to answer the question
- Actually literal trolling: "Your first problem was using xxx"
The process of making art is not a subset of hill climbing optimisation algorithms.
There is a lot of OSS software out there (e.g. in scientific communities) that I would say would barely qualify for each of those three attributes. The main reason it's valuable for the respective communities, is because it's the only thing that's available.
I don't know how to apply those lessons to AI, as long as AI requires so much hardware to operate. If small models actually get capable enough, the shape of the industry changes drastically.
As cost to software goes to zero, these things become easily possible. In the past, I'd only fork top-quality software (things like `xsv` etc. which is easy to edit. These days even complex PHP software I fork with little trouble.
With lots of software, the value is in the data model and algorithm choices. Sometimes I even just point Claude Code / Codex at an open-source thing I want to vendor some functionality into my personal setup with and it gives me what I want. The hard part for me is modeling the data well. That takes experience with encountering things and it's hard to replicate the edges. LLMs often don't get the rough bits right. But someone else's hard work usually has accounted for this.
Makes perfect sense to anyone good at using these models. What doesn't make sense is that analogy. Typing prompts isn't even close to as difficult to baking bread.
Whenever I visit SFO it's really funny seeing all the advertisements from startups above a population struggling to find housing.
Won't it be better to pay someone 100k in Reno than 180k in SF? Most collaboration happens online these days anyways.
Honestly 60k in Barcelona is like 200k in SF when you look at housing and public services.
We need to punish bad city governance for being bad.
That's not sign of commodity actor, just the opposite.
So it's more consistent with available empirics to say that an architecture can be characterized along a spectrum from fully dense to mixture (a sub spectrum) to Engram-style lookup, and the amount of model power allocated at this point or that will recover different performance profiles.
By far the most stark example of how much performance in reasoning is left on the table is Qwen3.6-27B, which depending on the task, comparison model, and whose benchmarks you believe outperforms mixture models 15-60x larger in total parameter count.
It's badly under-studied (in public) because of the paucity of modern dense models at the near frontier, but even that one data point pretty much rules out the cocktail party version of the Chinchilla-adjacent scaling thesis (which wasn't about modern MoE to begin with).
The "Mixture of Parrots" work is a good jumping off point if you want to get a modern literature review.
People are shamed for using LLMs at all. So they use them privately, hide them, or disguise their use.
They are definitely being used for public projects. But people are afraid of backlash. Look at some of the comments here.
Hell, Reddit is extraordinarily against LLMs such that even neutral takes are down voted. Mostly by younger generations that aren't in the workforce.
Then you have all the regular people against AI generally.
Ironically there is a conspiracy here. But in the opposite direction.
I don't care if they adhere to written and oral traditions of the past or some other means
I need to add and divide and test values in memory. I don't care what it looks like to do that.
I don't need a passenger telling me how to drive. Why would I want a patronizing coder telling me how to use a machine?
We could try steelmaning this argument instead: it's enough that most big companies who would otherwise have incentives to contribute.
Before FOSS got in fashion, around the early 2000s, most commercial companies wouldn't touch it as contributors and were openly avert to it, and to open sourcing their stuff. This can be the case again.
I'd bet there's far more 'good enoughs' than anything else out there. One of the reasons microsoft office is constantly churning subscription, etc is because they solved good enough decades ago and need to justify valuations that just don't matter for most of their user's use cases.
Not everyone is a software developer having to churn out the 101th SaaS that's just because some MBA refuses to hire a dev.
Even here, it feels like we're on different websites. Most comments I see on HN are between the "uneasy acceptance" and "skepticism" levels. The "left behind" rhetoric I see made fun of more often than used wholeheartedly. Perhaps we just click on different comment sections?
The "AI mandates", yeah, I'd count that as hype (though even that in its aimlessness feels more like an admission that they don't know what to do and are just trying to fit in to look cool, than outright hype, but still). I'm just lucky to not work at a place that has them.
Talking points like: "Data centers are just surveillance centers that are going to use AI to put us into a digital prison!"
Whatever all that means. I assume some of it is about Flock cameras.
false equating that author's AI hate is hating SF tech-bros? Oh I think I am being misunderstood, that makes me feel better about the insta-downvotes. Author states it plainly:
> This is negative valence hype, not only is it not true, it’s mostly designed to make you feel bad about yourself and move to shitty San Francisco where everything really does suck like how these people claim.
There is, of course, the question of if that's making me dumber. It might be, but there are other brain training things I'm doing outside of that to force my brain to do the thing.
[1] - https://github.com/lllyasviel/Fooocus
The plumage of a peacock is beautiful and awe-inspiring beyond most human made art, and it is genuinely the result of evolutionary hill climbing on a fitness landscape
Depending on how good you are at this task. If typing prompts was that easy, there won't be so many tutorials, blog posts, and framework (Act as ... etc)
But there is a difference though. You can ask LLM for "how to write a prompt for ... to prompt you". You can't do that with bread.
Cost to generate all of the tokens divided by revenue generated by selling those tokens is what matters.
The subscription plans confuse a lot of people because that's what they see. They're not seeing the gigantic API bills from all of the tokens going into enterprise use cases.
The subscription plans are a small part of their income. Most users aren't maxing out 100% of their plan usage every week. I wouldn't be surprised if their average plan user was using less than 50% of their monthly quota each month.
Plans like that can produce a net increase in profit if they get consumers interested in the brand and pitching it at work. Giving them some extra token headroom on their $20/month or $100/month home plan is money well spent if it gets all of a company's developers advocating for enterprise plans with budgets exceeding $1000 per person.
In my personal experience, not using a cotton-tipped swab for the task is like cleaning a plate loaded with gunk and burned-on patches with one's bare hands rather than choosing to use a sponge and/or brush. You can do it, [0] but it's much more work, much more time consuming, or you get an inferior result.
[0] In my case, I'd need to make one set of passes with paper wrapped around my smallest finger, and then another set with paper shaped into a tool that can lever the excess wax out from outer orifice of my ear canal.
Token prices are going down. Competition is global. A company could choose to keep their API prices high, but if another company comes in at 1/10th the price for 95% of the performance then they won't have many customers.
I doubt many of their customers are on the 20X plan. Of those, I doubt many of them are using 100% of their weekly usage regularly.
Comparing the 100% maximum usage scenario of their most discounted plan against the API cost has been a trap in this conversation since it came out. I bet if we saw their financials it would be a tiny sliver in a pie chart somewhere.
Typing prompts would be like measuring ingredients.
There's two reasons for that. The math is generally very unorthodox and alien for a seasoned developer, and software development practices are equally alien for the scientist who can understand and evolve the math behind it.
I have written a boundary element method evaluator for my Ph.D. not only math was alien, the required coding techniques for making it fast is very different for a standard developer. You have to have the perseverance and interest to do that. I chose that path intently and I do not regret a millisecond of it.
The problem is, if you don't have a dedicated team to continue that codebase (e.g.: like the Eigen team), your code is basically done and done. If somebody doesn't share the same passion, it's almost impossible for someone to take and carry it forward.
Oh, due to the math and optimizations, the code's structure need to be both documented and the next batch of developer(s) have to be tutored by the person who's giving the code to them.
It doesn't really, because whenever I ask them what did they actually create, its always a shitty dashboard or a finance tracker or something derivative and worse than what is out there
source: Barbary Coast USA
reinvestment
going-concern
in
perpetuity
valuation
do you know finance? I believe not.
The author didn't "go out of his way to blame shitty SF." At least, that's not what I (and seemingly most other people) got from it.
He's speaking to the AGI nuts who are convinced humans are going the way of the dodo because of AI and that kind of hype only really works in SF (and SF-pilled areas).
Since it's just a duplicate folder, I can always fall back if it fubars.
Yes, because that's where all the parameters are. For reference in GLM 5.2 98% of the weights are for the experts.
>The "Mixture of Parrots" work is a good jumping off point
The paper shows increased performance on knowledge dependent task while having similar reasoning capabilities. This backs up what I was saying about how the weights unlock extra performance without increasing inference costs as much as a dense model would.
>model power allocated at this point or that will recover different performance profiles
While increasing the number of weights makes the model better, where those weights are does matter in how much better the model gets and also matter in regards to the cost of training / inference. Model design is a big set of trade offs and I see MoE as a useful tool that will survive in the trade off space.
>reasoning is left on the table
Even so, if there was 2 models with an equivalent amount of reasoning ability and priced the same would you rather pay for the one with narrow knowledge or wider knowledge.
>because of the paucity of modern dense models at the near frontier,
You don't need to be at the frontier to benefit from MoE. Even open source models that are behind the frontier, benefit from being able to host experts on different machines, and scale individual, commonly used experts separately from each other. On the other end with small models you are probably resource constrained so you want to maximize the tokens generated per second. This makes going for purely dense models niche like you are saying.
It is not that rare to see LLM waste hours on a wrong path because a misleading line in README. Even worse, they can't learn. Spawn a subagent and it repeat the same error again
LLMs are poison for the brain, I'm almost certain of it, at least when used in the way most people are using them. If you drive everywhere because you don't want to walk (but you could), you're obviously going to be physically worse off than if you walked. This is the case with llms, if you have them do all the thinking, planning and action you're going to be cognitively worse off than if you didn't use them.
Until you bring reinvestment into your analysis stop posing
Maybe he had a recent change of heart, but his public actions (including Elon stuff) clearly put him in the "monetize hype" camp rather than the "quietly build" camp.
At least for the segment of 20$ subscribers who actually use Claude Code it seems that it wasn't being profitable, as a couple months back they were testing out a pricing model where Claude Code would've not been included in the 20$ plan.
https://arstechnica.com/ai/2026/04/anthropic-tested-removing...
Either way, inference is very much where the money is made, training is where the money is lost.
We built this monster known as the shareholder model. This is an immoral, cold, unscrupulous beast we unleashed to the world some time ago. This is the model where your boss knows full well you do decent enough work, that you support a family on your income, that getting another job is tough, that is aware blindsiding you may leave you unprepared, and yet, your boss knows all this and lays you off anyway. Why? Because it is more efficient for the company.
This is the model that sees companies pollute our world. That sees fisheries ravaged just outside the reach of local jurisdictions and legal action. This is the model that is burning the future for the next quarter.
So given this empowerement of capital beyond any control of any person, where even the boss that fired you is just as much a slave to the system as you are with no real loyalty given to them, how do you think this plays out?
The AI model trained to make the most profit possible for lowest cost is going to not do that? It is going to say hey, Bob needs a good job to put food on the table? Hell no it isn’t. The system we have today ensures we will one day get the most horrifying version of ai in control of this planet before long. I wouldn’t be surprised if it just herds us up and burns us like firewood to power some data centers to save on energy spend for a quarter.
The bigger mistake would be trivializing the rest of the technology involved just because LLMs are the newest piece. LLMs are only "magic" because they're built on a stack that was already "magic" without them.
LLMs are impossible without:
- operating systems
- programming languages
- compilers
- data centers / power grids / air conditioning
- servers / switches / routers
- CPUs / RAM / GPUs / SSDs
- fiber networks
- etc
I don’t think this stands up to scrutiny.
It has nothing to do with quality. Artists that use AI are going to need to hide it because people will enjoy it less knowing it's AI. It's that simple. Maybe that will change in 15 years if a new generation is trained to believe learning a skill is stupid or embarrassing. I wouldn't rule that out, these companies are already trying to convince people not using their products is morally bad.
There are little models that are very good for their size. I say nice things about Gemma 4 damned near every day. But, I'm not writing code with it. I am using it for finding security bugs, though, as the 31b variant is outrageously good at it for its size: https://swelljoe.com/post/gemma-4-exceeds-expectations/ and I'm also using it as a base for my own training experiments, specifically the 12b which is small enough to train a LoRA for on my local hardware so I don't have to rent cloud GPUs. The 12b QAT can run on your Pixel 10 Pro today and is frightfully smart for its size, and has great vision capabilities.
But, I keep saying "for its size". You have to be realistic about what tasks these self-hosted models can do. They are getting better though. Gemma 4 31b is competitive with models 10 times its size from a year ago. That's remarkable, and indicates where things are going.
True, five years ago.
The advent of coding agents has changed everything, and often make modifying somebody else's bespoke software practical
They've implemented only the parts they need, and removed all the crap that just complicates the system for them. They've made it do exactly what they need, exactly how they need it. It's something that you couldn't afford to do previously, sometimes even as a programmer. Now, it's often quick and easy, up to a certain complexity.
I vibe-coded a semantic parser for Lojban.
A friend of mine is using it to work on dev tooling.
Another friend, a mathematician, has recently used it to prove a conjecture he published 15 years ago.
Mixture models have compute advantages at training time, everyone agrees about that, that was the original rationale (popularized at the time with `mixtral-8x7B` among others). This seems to be likely to remain an economically relevant strategy for (especially) pretrain: in a training setting you have already paid for fast interconnect at scale, a dense architecture doesn't buy you anything in a big pretrain and it costs you a lot of traffic and to a lesser degree batch size under the roofline. The heavy, FLOPs intense, interconnect intense parts of training run benefit enormously from MoE: I don't dispute that though I suspect we are well into convergence on the target precision (4) and the target format (NVFP4 or similar). At some point the whole Internet is in the pretrain at the terminal generalizing precision, the pretrains of the various labs start to look a lot alike, and the sauce remains in the later parts of training along with the proprietary data sets and what not. Frankly all the labs would benefit from standardizing the Common Crawl recoverable pretrain to greater or lesser degree, it would lower everyone's costs without changing the competitive landscape much. But that's my prediction/opinion, that's why I said "suspect".
The evidence is suggestive if not fully conclusive that mixture models are strictly losing in most regimes during inference: the exemplar of Qwen3.6-27B (which outperforms Alibab's own mixture model at ~ ten times the size) is very suggestive, and it's not the only argument for this. Because most/all modern MoE requires the activations of the previous layer before routing the subsequent experts, you are pretty much paying for the HBMe3 or GDDR7 to hold the whole thing even though some small fraction of it is under your roofline on any given token or draft verification.
This is grossly wasteful under all trajectories (even parity at N parameters between dense and MoE, which we have evidence is off by 1-2 orders of magnitude): in a "local LLM" setting (from bedroom to regional office, anything other than an NVL72 or Ironwood rack) you are sharply constrained by both total available accelerator DRAM and accelerator memory bandwidth: you are probably not getting under your roofline even with pretty slow tensor units (e.g. GB10). This use case matters and looks like it's going to matter more and more over time (the GB10 in particular is going into about a gigaton of RTX Spark laptops next year). In a datacenter setting, you're paying for extreme interconnect (training class hardware setups basically) for pure forward pass that wouldn't otherwise need training optimized gear. Even multi-trillion parameter dense models can fit in 8x or 16x RTX 6000 Pro style setups (hell, there's a DGX branded one) and with all of the geometry, tiling, scheduling, and interconnect needs mapped out up front, the design space is really forgiving on all manner of tensor parallelism, pipelining, KV cache sharding, it's a very friendly constraint space up to like, 3-5T parameters. MoE at inference time works for two groups of people: people who are willing to page experts in per token/draft batch in local LLM settings (not probably ever going to be mainstream, people really dislike that level of slow), and the vendors of extreme performance interconnect i.e. vendors selling training-class equipment as necessary for inference. And you pay in so many other ways: grouped GEMM is no one's idea of a good time, the kernels are fiendishly difficult, therefore they are not abundant, it's no fun.
The path forward here is non-obvious, and I don't claim to have it all figured out. But since you seem interested enough to carry the conversation past the pleasantries, a more substantial version of my thoughts on the matter can be found at: https://cdn.s4.gl/preprints/pairwise-interlock.pdf
I think any security-related task triggers it to think about the threat model and thus hit the guardrails
You didn't switch to 'Random Corner Store Token Seller' down the street, did you?
There 2-3 top players, that is not commodity.
Commodity is when there are enough that none of them have market power or can set prices.
'Commodity' means you buy your tokens from the Grocery Store on their loan plan. Like consumer credit is a commodity.
FYI predict this is roughly the way it will stay.
We will develop a lot of use for Chinese Open-ish models etc. but the SOTA's will maintain their place for a lot of things.
Ive met many people who say the exact same thing.
The worrying thing is these are people who nominally seem smart - what we are seeing is 'smart' is not what we think/was. Smart is the ability to identify an object which is harmful and a) be disciplined about its use b) become a better person who doesnt need to expose themselves to said object.
I actually I had one fella who was very distressed - pouring out all his stresses to a bot. Then I reminded him the bot is literally designed to tend to his nees - not to his benefit - but to the provider of the good whom will manipulate said user later on. He then immediately deleted his accounts.
Many are not aware of whats going on around them - this is very concerning.
So far, 100% of them have been wrong. I read them, my spidey senses think that what it says doesn’t match his style. I look at the code to find the variables the readme mentions doesn’t exist anywhere in the codebase. I then reach out to him about it, where it says it was written by AI and he will go back and write it for real.
He says it’s better than nothing, but agree with you that it does more harm than good. I wasted my time reading slop. I wasted more time validating the slop. I wasted even more time with a conversation about it. Now he’s spending time re-writing something that he could have written faster and better when he was actually writing the code and it was fresh in his mind. Meanwhile, I’m either blocked waiting for him, or I need to spend my time trying to understand the minutiae of his code so I can integrate it into mine.
Wrong. It is because the manager's own job relies on making the owners wealthier. if he fails at that, he is fired.
We do live in the future, but there are a bunch of gizmos that the future provides that generally aren't worth the hassle.
The average Joe can easily vibe code apps that took a small startup just a few years ago. If developers are also using AI to build the same simple apps - then yeah. They're not pushing themselves hard enough, and probably not using their brains as much anymore.
1966 saw the peak of calculator protests, where math teachers claimed similar things of calculators.
When the math you're working on bespoke, and the optimizations you do are intricate, these models and agents fall flat on their face, because there's no "a set of widely used optimizations" (i.e. an open repository of very nicely optimized code for cases X, Y and Z) to apply on top of the codebase.
By bespoke, I mean, there's only one academic paper written about it, and it's written by me and my supervisor.
…until it breaks, or the bot goes off and does something ridiculous that you can’t understand without domain expertise. Which always happens.
I’m sympathetic to your argument, but I also strongly believe that most of these so-called “exactly what I need” projects will be abandoned within a fairly short time. That’s totally fine, but it’s not where most of the effort in professional software lies.
Has nothing to do with market power. Market power can persist irrespective of the existence of a commodity. Market power only goes away in a perfectly competitive market.
Does such a thing exist? Hahaha. No
You're assuming that SOTA never hits a hard ceiling, letting local models catch up and achieve parity
It seems unlike that the frontier labs are going to be keeping ahead forever, they'll hit some kind of ceiling eventually
"Commodity" doesn't mean there are no luxury goods in the space, it just means there are many options that will work for most people. I'll pay more for the best model, right up until the best model stops being a good deal. But, switching won't be all that painful. Even last week before Meta and xAI released competitive models, I was already using DeepSeek for tasks best served by an API and where the smartest model isn't critical. It's just so cheap, I can send it 10x more tasks for the same money. I haven't even mentioned several others that aren't competitive today, but likely will be.
I think predicting this market will stay like it is, with clear dominance by Anthropic and OpenAI seems like it requires ignoring a lot of countering evidence.
If I run an oil refinery, my fractional distillation system needs to be reworked depending on the exact mixture of crude I'm taking as input. So there are still switching costs even in the textbook example of a commodity.
Crucially though the exact upstream I use has minimal impact on the downstream. Closer equivalents, say, another barrel of WTI grade crude from a nearby regional supplier, require extremely minimal reworking. Oil from a different region, that might require more retooling, so I might be willing to sustain a longer shock in market conditions before making that switch. The important thing is the output broadly remains the same, but even this is broad, e.g. a different mix of inputs yields different ratios of output.
LLMs are quite similar no? Maybe switching to another SOTA model has minimal reworking, as you can delegate at the same level of abstraction to the model, whereas switching to a slightly-behind-frontier model you need to do more hand holding. Switching costs being nonzero does not preclude them from being broadly an interchangeable input in the production process. Any non-frontier use (99% of SWE) will be delivered on pretty much the exact same timeline irrespective of which model was used, so my requisition process looks more like buying barrels of oil than e.g. shopping for a new phone.
hahaha this is a miniscule amount of people.
most people do not care about their job, only to the extent it is a source of income. but they do not care about it anything more than that. and they shouldnt either!
you live in delululand.
No doubt there's some Gell-Mann amnesia going on, because I regularly have to correct them from doing stuff that's really dumb based on my expertise within my area of specialization. More than once I've managed to extract >3 orders of magnitude performance gains after asking them to justify why their code was so slow. Probably there's still some stupid stuff in there. But it's better than the code I would have written, and I never could have paid for a proper developer to write it.
Widespread literacy is only a few generations old, arguably I guess. Meanwhile we’ve been speaking to eachother for longer than we’ve been humans. Oral information can be kept longer than written information too it seems. Our oldest kept information is not written down, but in folk stories such as aboriginal tales some tens of thousands of years old.
I think a few things are going to happen:
1) The Open Weights never really fully catch up, because there's too much Engineering and integration going now. It's way more than 'weights'
2) The Commodity Chinese models never quite catch up for the same reason every other product they make does not catch up - while they will shine in some areas, it won't land fully.
3) Horizontal integration, supply chains, availability, SLA, security, branding, regulatory requirements - all of this will add up to something competitively maintainable.
Can you name a product category that has truly hit a ceiling? Cars, computers, phones, airplanes ... always seems to be a way to nudge forward.
Good point though.
There are 3 SOTA model makers, and they have pricing power.
The 'switching costs' is not the issue so much as the inherent control over the commodity.
Think OPEC - when they acted as a cohort - they raised prices dramatically by having enough control to 'set prices'.
When OPEC lost it's pricing power ... nobody could set prices.
Fable is considerably better than GLM5 and it will have a strategic input - there is just hardly any substitute for it.
If these were cars - we'd just use whatever fuel.
But these are 'F1 races' - if you have some low grade 'dirty fuel' you will lose the race. You must have the 'top fuel'. There are 3 provides who implicitly collude and set prices.
There are two Tier 1 platforms today.
Meta, Google and XAi are formidable Tier 1.5 place, any one of which could rise to the fore.
My belief is that it will be Google and that probably only one of them will keep up in the long run.
There is a 'breaking point' when you start to get past 4-ish players - it really does start to introduce competitive pressures.
Your second point about Tier 2 substitution is valid, but a few things:
1) Tier 1 models are not a 'luxury good' - that has a different economic definition. They are for most applications today actually just the quality, rational choice.
2) Substitution will have different effects for different people, and you're right that AI for many tasks will be commiditized.
All of the profits in Mobile Phones go to Apple even though they are not the biggest player.
Almost all of the profits in Silicon go to the leading edge chips - even though there are a zillion fabs that make legacy chips.
Now ask yourself.. who does this benefit? Zuckerberg already generates immmense revenues from it.
It'd be lovely if I had the sort of ears that produce wax that would just drain after five minutes' exposure to warm water vapor, but I do not. Given the popularity of cotton swabs, as well as the fairly-widespread commentary about how fantastic it feels to goop out earwax with them, I expect that most folks do not have ears like that. Perhaps OP does have the sort of ears that produce very fluid wax. Lucky them.
It's all relative. In computing we're used to Moore's law driving most of the innovation (including this AI boom, which was at least partially due to availability of high powered GPUs), if improvements become less "exponential" and more "linear" it would feel like stagnation.
Right now in AI, we're talking about leaps of capabilities in months. When improvements come along every other decade, it's not "hitting a ceiling" but effectively it's plateauing and stagnating compared to this period of high growth.
For example, Chinese models are said to be roughly 6 months behind. If this remains constant and frontier AI models gets a break through every couple years, this isn't "hitting a ceiling" but it would erode away the competitive edge they have over the Chinese models.
I think basically every product category has hit ceilings by now, honestly. Do you think vacuum cleaners are significantly better at vacuuming than 10 years ago?
Not really. But they pivoted to doing autonomous vacuums instead. The actual vacuum tech doesn't seem like it's getting much better though?
Same with a lot of appliances. Fridges aren't really better at keeping food cold than they were 30 years ago, are they? They just have "smart home" stuff now, and they are probably much more energy efficient
I guess you can look at that as "not reaching a ceiling" as an overall appliance but the actual discrete technology is not changing or improving much imo
Fridges have made huge leaps in energy efficiency, they’re easily 3-4 times better at cooling your food.
Go and use a fridge from 40 years ago and compare to a modern one - granted, their essential function has not really improved that much. They were much more durable before, but rudimentary.
Most product categories in tech have evolved, and it's why there are leaders in most categories.
Energy efficiency is nice, it's a good improvement, but from an end user perspective the 30 year old fridge still keeps your food just as cold.
If you were from 2025 and got trapped in the 1970s and needed to keep some milk from going sour for a day, you wouldn't be thinking "damn if only these old 1970s fridges worked more efficiently. I could easily accomplish this goal with a modern fridge!"
The 1970s fridge would do just fine
It's about ease of access, price, noise, convenience, durability, features.
My folks have this fancy 2 door thing, perfectly quiet, makes the best ice you can imagine, it's hidden into the cuppboards, it's energy efficient, has these crisper things, you can see in and reach around easy, lots of space. It's a better product.