Thursday May 28 2026

Hacker Times

I think Anthropic and OpenAI have found product-market fit

Discussion

Just because API pricing would've been $2180.16 doesn't mean that's the value of those tokens. For starters, you personally probably wouldn't have paid that. But also, sales price isn't value. This is like saying, oh, I saw this bar of gold somewhere for $10000 but got it here for $1000! So I got $10000 worth of gold for $1000! - no, the value of that gold is determined by its weight, which wasn't even mentioned.

We have no market convergence on tokens yet (and it'll differ between LLMs), so it's impossible to say what value you got for your $200.

Loading replies...

recursive•about 20 hours ago

I'm willing to charge you $100k for those same tokens.

Does that mean you'll be saving $99k?

It sounds an awful lot like the mark-up to mark-down scheme where the price stays the same.

OrangeDelonge•about 24 hours ago

Large enterprises make deals and won’t be paying 2,180.16$ either. Just like with AWS

Loading replies...

altruios•about 24 hours ago

> If I had been paying API pricing it would have been $2,180.16

The point being made above is that API pricing is calculated... somehow... seemingly arbitrarily. Possibly untethered to the infrastructure costs entirely: which would be the basis of any 'value', however that holds the labor theory of value, which isn't accurate either. So how do you accurately price these tokens at all (other than through price-discovery: which is slow, messy and fuzzy)?

Loading replies...

prepend•about 6 hours ago

Lets say McDonalds charges $2000 for a BigMac. If they offer a deal and sell it to you for $200, did you save $1800?

Maybe if you spend $2000 on a BigMac. But it’s unlikely you would buy such a burger.

What is a hamburger worth? Don’t look to McDonalds to set the value.

pembrook•about 24 hours ago

API pricing drops DRAMATICALLY in enterprise agreements.

As with pretty much anything priced on volume/usage.

Enterprise deals are negotiated ad-hoc, the listed pricing is simply a jumping off point for the final negotiated discount.

If you’re going to give 20,000 employees Claude code you are not going to be spending $1B per year on Anthropic tokens as if you gave everyone an individual API key. Just as Anthropic isn’t paying AWS SES $10,000,000 to send 1 email update to their massive user base when the next Claude version drops.

Loading replies...

xnorswap•about 23 hours ago

Have you or I misunderstood the "teams" plan?

edit: I missed the "enterprise" feature matrix with the usual audit/compliance stuff to force the biggest enterprise customers onto enterprise plans. Otherwise the "teams" plan is much better value for any business.

orig-continued:

https://claude.com/pricing/team

Teams premium is "Everything in standard, plus more usage*"

And from my experience, it's a very generous usage, I've only hit the limits once or twice, and both times required multi-boxing agents.

I could single-window agentic development all day on opus-4.7 auto-mode without hitting limits.

If you're a business using claude, then that seems like the right plan, the enteprise/API plan seems more suited to where your product is built on top of the agent themselves, so seats/limits aren't really meaningful?

Loading replies...

regularfry•about 20 hours ago

> I suspect that AI will fail to pan out to the same extent for the same reason why outsourcing hasn't fully panned out

My mental model for that is that outsourcing fails where the work is being done organisationally far from the knowledge needed to do it. We know that's true of teams inside organisations, there's been a lot of research on how distance in the organisational tree negatively impacts productivity. Outsourcing is a pathological worst-case of that.

The promise (promise! We're not there yet!) of AI is that I can have a cross-functional team on my laptop. Organisational distance is zero. Where previously the outsourced team has to wait for the time zones to roll round so I can answer their blocking question when I get to my email STRICTLY AFTER I have had my coffee, now it's a prompt in a chat window with a button I can click to make a choice in 5 seconds. Delay is gone, cost of delay is gone.

> The problems that will come up will be and always have been ongoing maintenance. AI is great at writing new code without a brain behind it, but once you get to the point where you need to refactor code, you start really needing someone with coding experience to guide the AI or veto it's mistakes.

Oh, absolutely. That's a minefield. Today. It will be, right up until it isn't. There are ways to set up agents and projects right now that make a dramatic difference to how this part of the picture plays out, but those will sink into the harnesses as time goes on.

But also the big problem with maintenance and outsourced teams tends to be the commercial structure around the contract. You get a Build team, who Build the Thing and then: no more features for you, anything you want to add past the original spec costs extra. They hand over to the Run And Maintain team, who get to fix all the bugs that the Build team left but without the knowledge gained from building the thing, but are scaled and located to be absolutely as cheap as the supplier can get away with so probably don't have the skill, inclination, motivation, or permission to take on any restructuring to make the bug fixing easier and they're on the wrong end of the globe so there's a 24-hour latency on any queries. It's a terrible way to set teams up, but it looks good on paper.

Again, that's peculiar to outsourcing and completely goes away if I have the same team that built the thing own the thing long-term. That's true if it's humans or AI!

> I don't think that's really fixable even with a lot better AI. It's not something that ultimately comes out of the likes of github data.

No, it's a harness problem. You need to start from a maintainable point and keep standards in place. It'll take work to get the harnesses there and it's not ubiquitous. You might also need better models, but I've already personally seen big differences in outcomes between projects that took certain steps and others that didn't; it's nothing revolutionary, mostly stuff that works for humans also works for AIs but you need to know to ask for it.

> I'm not saying that AI isn't going to make things better, btw, I just don't think we'll see a 20x improvement. Probably more like 1.5 or 2x.

I think people radically underestimate the cost of delay. I don't know if 20x is realistic for the AI itself, but I think it's not impossible once the inefficiencies of having to go to other humans is factored in.

Loading replies...

thesz•about 21 hours ago

The difference between training and inference is 1) one have to keep intermediate results for backward pass in training and 2) computation for training double because of the backward pass.

Training is also done over batches, which increase memory requirements by several orders of magnitude. This is why training needs costly compute.

One of the ways out of this unfortunate situation is to use something like Stochastic Average Gradient Descent [1]. Examples there are mostly concerned with regularized logistic regression, which makes problem more or less convex. Neural networks are inherently non-convex. Still, maybe some ideas from there can be utilized in the context of neural networks, like use of estimated Lipshitz constant to derive curvature and appropriate learning step.

  [1] https://www.cs.ubc.ca/~schmidtm/Courses/540-W19/L12.pdf

regularfry•about 21 hours ago

What I think is happening is that the scale of thing you can hope to build at a below-corporate scale should radically grow. Corporate environments should suffer for this, being that inefficient.

> YCombinator was founded on that premise - small teams of founders and early employees could be orders of magnitudes more productive than the 1000s of corporate employees at their competitors.

I think this is still true, but the theory is:

1. You don't need YC-type funding to do YC-type business any more; 2. You don't need to scale the business past those small teams any more, you just buy more tokens.

For clarity YC still obviously has a place as an incubator, mentoring, and networking function. I just think that what was previously the inevitable conclusion that you have to hire all the people the second you hit PMF to keep up with scaling the business as you scale sales is no longer inevitable. If you didn't want to go that way before AI, you were a "lifestyle business" and not worth investing in. As more and more knowledge functions get capably implemented by AI, it's the preferred position: humans are vastly more expensive than tokens, so you want them doing the stuff the AI still can't do.

I don't think this necessarily translates to mass unemployment. I think it translates to masses of smaller businesses that are radically more efficient because the handoffs between business functions are tool calls, not emails to someone who doesn't want to help.

> The Internet was revolutionary because it let millions of people bring products to market without asking permission.

Think about it this way: if I am a small business owner but I think it makes sense to do something that previously only a team in a corporate environment could do but is now within the reach of AI, not only can I do it now, but I also don't have to ask anyone for permission! Who wins between the corporation and the small business in that scenario?

> AI doesn't really have that property - if anything, it makes things more centralized, with more gatekeepers, and so seems more likely to destroy economic value than add to it.

I think this will turn out to be backwards. I can see a version of this where the number of things you can do without needing to turn to a gatekeeper for help increases to the extent that the balance completely inverts.

The vast majority of businesses are small, and AI can give them tools which previously required corporate scale to make sense, without the inefficient hand-offs between busy, political humans. Which is also something that the internet did! Getting an advert in front of a national market pre-internet was Hard but sometimes you had to do it because your target market was "all Canadians who buy toothpaste" or whatever and that meant saturation-bombing the physical environment with physical billboard ads, posters, flyers, and so on. So you only did it if you were P&G-scale. Now you, personally, can do it, trivially, for better or worse.

Loading replies...

27th May 2026

Anthropic are strongly rumored to be about to have their first profitable quarter. Stories are circulating of companies surprised at how expensive their LLM bills are becoming from usage by their staff. I think this is because OpenAI and Anthropic have both found product-market fit.

Enterprise customers are now paying API prices

I currently subscribe to the $100/month Max plan from Anthropic and the $100/month Pro plan from OpenAI. If you are a heavy user of coding agents these plans are a fantastic deal. I just ran the ccusage tool on my laptop to get an estimate of how much I would have spent if I were to pay for API tokens in the past 30 days and got:

$1,199.79 for Anthropic Claude Code
$980.37 for OpenAI Codex

That’s $2,180.16 worth of tokens for $200—not bad at all! I’m a moderately heavy user of these tools, but I’m certainly not running agents every hour of the day and night.

I had assumed that companies making extensive use of agents were getting similar discounts. It turns out I could not have been more wrong about that.

I haven’t been able to track down the exact date, but at some point in the last six months Anthropic switched their Enterprise plan (originally “Claude seats include enough usage for a typical workday” back in August 2025) to $20/seat/month plus API pricing for usage. This story about the change from The Information is dated Apr 14, 2026, but cites an Anthropic spokesperson claiming that the pricing change occurred in November 2025. Existing customers are finding out about the change as they renew their contracts.

OpenAI made a similar pricing change in April. The Codex rate card (Internet Archive copy) currently says:

Note: On April 2, 2026, we updated Codex pricing to align with API token usage, instead of per-message pricing. This change was applicable to new and existing Plus, Pro, ChatGPT Business and new ChatGPT Enterprise plans.

On April 23, 2026, we made this update for all existing ChatGPT Enterprise plans as well, inclusive of Edu, Health, Gov, and ChatGPT for Teachers.

It’s a little harder to decode as they quote prices in “credits”, but as far as I can tell those credit costs are an exact match for the API token costs listed for those models.

All of which is to say that as of April 2026 the “Enterprise” cost for both OpenAI Codex and Anthropic Claude Code/Cowork is the same as the listed API price.

GPT-5.5 (released April 23rd) is 2x the API price of GPT-5.4. Opus 4.7 (April 16th) is around 1.4x the price of Opus 4.6 when you take their new tokenizer into account.

So April saw both leading model companies release new frontier models with a higher API price, and both companies now have measures to lock their enterprise customers (who tend to sign year-long deals) at those API prices, not the previous extreme discounts.

I think they’ve found product-market fit

Why these sudden aggressive moves on pricing? Both Anthropic and OpenAI are planning to IPO, but I suspect there’s a more important factor here: I think they’ve finally found product-market fit, with the coding/general-purpose agent products embodied by Claude Code/Cowork and Codex.

Tools like ChatGPT are wildly popular, but that wild popularity has been difficult to turn into revenue. In February OpenAI boasted more than 900 million weekly active users for ChatGPT, but only 50 million—5.6% of that—were paying consumer subscribers.

Charging $10-$20/month per user is an OK business, but you’d need 1-2 billion subscribers sticking around for four years to cover $1 trillion in infrastructure.

Companies spending $200+/month/user will get you there a whole lot faster—and as noted above, as a power-user I’m at ~$1,000/month in API costs per vendor already.

Coding agents really did change everything. These are tools which burn vastly more tokens, but are also quickly becoming daily drivers for the work carried out by extremely well-compensated professionals. Right now that’s still mostly software engineers, but a coding agent is a tool that can automate anything you can do by typing commands into a computer... so they are clearly applicable to a much wider set of skilled knowledge workers.

As I’ve discussed on this site at length, the models released in November 2025 elevated agents to being genuinely useful. We’ve had six months to get used to that idea now—it’s no wonder companies are beginning to spend real money on this technology.

You could argue that ChatGPT achieved product-market fit when it became the fastest-growing consumer app in history back in February 2023... but it certainly wasn’t making any actual money back then. Coding agents plus enterprise pricing marks the point when these companies start making very real revenue. Maybe even enough to start covering their costs!

And they’re ramping up

As further evidence that enterprise agents represent product-market fit for these companies, consider their open job listings.

OpenAI have 703 open jobs right now, of which I’d categorize 229 (32.6%) as relating to enterprise sales and support—account executives, “Go To Market”, “Forward Deployed Engineers” and the like.

Anthropic have 390 open jobs, 105 (26.9%) of which look enterprisey to me.

It’s pleasingly ironic that these AI labs have picked a business model with such a heavy demand on human labor—enterprise sales contracts don’t close themselves without a whole lot of humans in the mix!

(I ran this analysis by scraping their job sites with Claude Code, then having it use Datasette’s JSON API to pipe that data into Datasette Cloud where I used Datasette Agent for the analysis, exported here. Dogfood!)

The AI-failure stories around this are pretty thin

I started digging into this in response to a growing volume of stories claiming that large companies were sounding the alarm because their AI usage costs had grown so large.

The most widely cited of these stories appear quite overblown to me.

The most discussed has been Uber, based on this report where CTO Praveen Neppalli Naga indicated that Uber had “maxed out its full year AI budget just a few months into 2026”, mostly thanks to Claude Code.

Given that Claude Code only got really good in November it’s entirely unsurprising to me that a budget set in 2025 may have failed to predict demand for that tool in 2026!

That Uber story was further fueled by comments made by Uber’s COO, Andrew Macdonald, on the Rapid Response podcast. I tracked down the segment and there really isn’t much there. Here’s what Andrew said:

But then you sometimes go and talk to your senior engineering leaders and you’re saying, OK, how many projects that were on the cutting room floor got moved above the line because of the productivity gains because 25% of our code commits were via Claude Code last quarter?

That link is not there yet, right? I think maybe implicitly there’s more that is getting shipped. But it’s very hard to draw a line between one of those stats and, OK, now we’re actually producing like 25% more useful consumer features, right? And that line is hard to draw.

Somehow this fragment turned into headlines like Uber’s COO says it’s getting harder to justify the money spent on AI tokenmaxxing, because the market for stories about AI failures remains enormous.

The other popular story around this is Microsoft starts canceling Claude Code licenses, ostensibly to encourage their engineers to dogfood their own Copilot CLI agent instead—but The Verge reporter Tom Warren says “sources tell me the decision is also a financial one”, triggered by the June 30th end of Microsoft’s financial year.

I think both of these stories support my “product-market fit” hypothesis. The best advice I ever heard on pricing a product was that your customer should suck air through their teeth and then say yes. Uber’s budget overrun and Microsoft’s seat cancellations look like that effect playing out in practice.

We also know the labs are spending a lot

The big AI labs spend billions of dollars on both training and inference. Credible figures are hard to come by, but we did get one huge hint as to the figures involved from, oddly enough, the recent SpaceX S-1:

[...] in May 2026, we entered into Cloud Services Agreements with Anthropic PBC (“Anthropic”), an AI research and development public benefit corporation, with respect to access to compute capacity across COLOSSUS and COLOSSUS II. Pursuant to these agreements, the customer has agreed to pay us $1.25 billion per month through May 2029 [...]

The Anthropic announcement said that this deal meant they could “increase our usage limits for Claude Code and the Claude API”, heavily implying that Colossus is being used for inference, not model training.

Anthropic already have vast amounts of compute from other providers. The fact that they’re willing to spend $1.25 billion per month for extra capacity from just one of their vendors hints at how big these inference budgets have become.

API revenue is becoming less important

Over the past two years my impression has been that OpenAI made more of their income from subscription revenue while Anthropic made more from their API.

Anthropic’s API revenue was historically quite dependent on a small number of large API customers—this VentureBeat story from August 2025 quotes “sources familiar with the matter” suggesting that just Cursor and GitHub Copilot were responsible for $1.2 billion of the company’s then-$4 billion revenue.

Today Anthropic are rumored to hit $10.9 billion in the second quarter, potentially even operating at a profit for the first time.

This pivot-to-Enterprise suggests that the labs have realized that the real money lies in cutting out the middlemen. Anthropic’s Claude Code directly competes with Cursor and Copilot. No wonder Cursor are investing in their own models!

April is a new inflection point

I’ve called November 2025 the November inflection point because that was when GPT-5.1 and Opus 4.5, combined with their respective coding agent harnesses, got good—good enough that we’ve spent the last six months adapting to agent systems that can reliably get useful work done.

I think April 2026 is a new inflection point where the revenue implications of this have started to land, to the benefit of the frontier AI labs and with material impacts on the budgets of large companies.

We’ll know for sure how real this moment is when the S-1 documents for the upcoming Anthropic and OpenAI IPOs give us some real, audited numbers to get our teeth into.

aspenmartin•about 13 hours ago

Well I think there are several fairly stable trends that paint a pretty compelling picture:

- performance scales with compute very very reliably. We have “scaling laws” (and have for years) and they are almost miraculously stable and show no sign of being invalidated at all even at the very largest scales. There are some theoretical bases for this though I’m not as familiar with the details

- these scaling laws are on an unintuitive quantity (validation loss on pretraining datasets), so we can look at downstream performance. Benchmarks are a minefield of junk but there are many decent ones and enough variety of techniques and data sources and scoring methods etc that in aggregate they are useful. The single number that I think is the best summary statistic across the crazy (O(100k)) number of benchmarks is the “epoch capability index” (just some branding over a reasonably standard statistical model that was really well thought out and a great idea). The trends in this are extremely stable. Eyeballing the trend over time on their graph we’re getting basically a GPT-4 to GPT-5 level capability improvement every ~18 months

- coding agents are not limited by the quality of the human training data they’re trained on, this is such a massive misconception: human data is only a bootstrap to a reinforcement learning phase. This combined with the fact that we have verifiable rewards means it’s just a matter of when not if for any given level of reliability.

- the massive compute investment implies that the compute that we’re building over the next 2-3 years will 10x the effective compute for training models. That combined with various R&D contributions (historically which have been very significant and there is no shortage of wins here), better data curation and flywheels, richer data (wait until conversation capability gets good) means we have several orders of magnitude of runway that we know of, today.

In short I don’t see any compelling evidence to suggest all of the trends we observe in many different ways will end any time soon.

tomjakubowski•about 10 hours ago

Anthropic advertises volume discounts:

Volume discounts may be available for high-volume users. These are negotiated on a case-by-case basis.

* Standard tiers use the pricing shown in Model pricing

* Enterprise customers can contact sales for custom pricing

And there are discounts available through "Claude Platform on AWS":

Anthropic rates your token usage in USD at standard per-model, per-feature rates, applies any negotiated discount, converts the result to CCUs at $0.01 per CCU, and reports the CCU quantity to AWS Marketplace hourly. Your AWS bill shows a single CCU line item.

https://platform.claude.com/docs/en/about-claude/pricing

On the other hand, contrary to Anthropic's documentation, another source claims they've killed pre-existing API volume discounts for large enterprise customers as of April: https://itbrief.news/story/anthropic-shifts-enterprise-billi...

wongarsu•about 22 hours ago

But the feature list at https://claude.com/pricing#team-&-enterprise literally lists "tiered incentives on committed spend" and "non-standard terms" as perks of the sales-assisted Enterprise plan. Maybe "non-standard terms" could mean "we dance for you if you pay", but what would "tiered incentives on committed spend" mean besides "we can negotiate on price if you bring the volume"

asib•about 21 hours ago

> > The parent there may be making the same assumption I am, that large enterprise _never_ pays sticker price.

> I shared that assumption until yesterday, when I found out that it wasn't holding for LLM pricing from OpenAI and Anthropic.

This reads like GP saying "enterprise never pays sticker price" and you responding "I thought so too until I saw the sticker price".

Is there some info you have that you can't/didn't share? Your article doesn't offer anything beyond the above.

Loading replies...

mvanbaak•about 23 hours ago

large enterprises dont pay openai or anthropic, they get this thing called copilot and get a nice price there. At least on this side of the pond (eu)

mrandish•about 14 hours ago

> I found out that it wasn't holding for LLM pricing

You're correct. When a type of cloud service grows large enough and has a few competitive suppliers, enterprise pricing tends to coalesce with the large buyers paying around the same price for the same thing. While that might be lower than the publicly cited rate card, the private price similarly large customers pay ends up being similar.

One reason is that the very largest, long-term enterprise customers are so valuable, they can command MFN clauses ensuring no one else is paying substantially less for the same thing, then the rest of the rate card for smaller customers flows down from that. There's a strong disincentive for vendors to cheat or allow big disparities between similar classes of customers because the number of people involved on both sides of these deals is large enough that word will get around eventually.

Large scale enterprise sales and purchasing in a given sector tends to be rather circular. Account execs move to other vendors and call on the same customers, while purchasing agents can move to other customer firms. Personal relationships, reputation and credibility matter. Lying to screw a large customer over just to make one commission or quarterly quota can be a very bad long-term career move. Sometimes purchasing agents or executives quietly compare notes off-the-record with their peers from other similarly-sized firms. After dinner drinks at industry association meetings and trade shows can be quite productive in terms of verbally exchanging 'market insight' with peers.

When there are significant pricing differences, its usually due to different volume commitments, SLA/QoS guarantees, payment terms and other material factors which justify the difference. Source: been there, done that inside a top ten valley tech company. Once was in a meeting where a newly minted EVP tried to get a long-time senior account exec to pressure a huge customer by being semi-dishonest. The account exec schooled the EVP on the fact that the EVP could only make him unemployed for about 8 hours but that huge customer not wanting to work with him could make him unemployed forever. :-)

JohnBooty•about 14 hours ago

    would you be dissatisfied by Opus-4.6-level open-weight 
    models, just because Opus 4.8 will be out?

Well, I see what you mean, but two big concepts...

1A. Models get stale pretty quickly w.r.t. new developments that occur past their cutoff date. "But you can just keep them current by linking them to never documentation, etc!" Well, no, you sorta can't -- at least not in perpetuity. Those search results fill up your context window real quick. So that gets unsustainable real quick.

1B. Even when your context has plenty of free space, the results you get from "here's a link to the documentation for this new framework that released after your cutoff date" absolutely pales to the results you get from knowledge that is fully baked into the trained model as opposed to your context window. For one thing, that documentation link you pasted into your context might link to... a dozen code examples. Whereas if that was baked into the model itself, the model might have been trained on many thousands of examples in Github etc.

2. It's also a reality that most professional engineers have to keep up with their peers and competitors. We can maybe say it shouldn't be that way, but it is. So if $SOME_NEW_MODEL is significantly better than 4.6... and my peers and or competitors are using it, then yeah I might but really feeling the need to match them. And I'm not even necessarily talking about some kind of cutthroat dog-eat-dog stack-ranked workplace.

These limitations aren't relevant for all use cases or careers but they're hiiiiiiiighly relevant for professional software engineering.

scribble0242•about 15 hours ago

This worked for me with qwen3.6-36b-a3b even at a q4 quant. I ran pi in a docker container and it had to figure out how to install python as well. I used the same initial prompt you had without any additional. You talked about Qwen 3.6, but then said you tried Qwen 3.5 in lmstudio. Not sure if you meant Qwen 3.6. I ran with llama.cpp llama-server with the recommended settings from unsloth.

I'm not an expert in SQLLite so I can't say if this is 100% correct, but it seemed directionally similar to the conclusion from claude.

  ### TL;DR
  
  - Authorizer + EXPLAIN:  No — authorizer only sees SQLITE_INSERT, not VDBE opcodes
  - EXPLAIN opcode analysis alone:  Yes — Delete opcode at position 10 is the unique signature of INSERT OR REPLACE / REPLACE

I can't help but think the not-so-distant future will see language models expected on commodity personal computing devices.