This is a pretty common trading firm internship project funnily enough.
> Claude Opus 5's default user-facing responses run longer than prior Opus models'.
The benchmarks do show Opus 5 as slightly more expensive than 4.8, although the scores are much higher.
This still feels like a step in the wrong direction, though, especially with OpenAI making so much progress with the efficiency of their models. Fable's token efficiency made it seem like Anthropic would start following OpenAI's approach but that doesn't seem to have carried over to their other models.
Okay so it’s worse than Opus 4.8 for my purposes I guess?
There are 10+ LLM companies, each with dozens of models of different modalities, each model with multiple size variants, then different “thinking” levels, then agentic modes, “pro” modes, a “fast” option, standard vs flex vs batch execution. And of course each end combination has a different input/output/cache token price.
Companies that say “give me a prompt and I’ll route it to the most ideal and cost effective model and setting for you” are capturing a ton of value from a gap that model developers don’t seem to understand exists.
Then system card goes on to "Its AI R&D capabilities are comparable to those of Claude Mythos 5", which is supposed to be fable minus restrictions.
> "Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access."[1]
On the Opus model release page, the reason why Fable doesn't have an ARC-AGI score is because of that retention policy[2].
0: https://support.claude.com/en/articles/15425996-data-retenti...
Previously Fable was the best at this, followed by Gemini 3.1 pro (a surprising #2, but Google has great vision models).
Opus' results seem to be more accurate than Fable, following the design source of truth better.
Example results:
Design source of truth: https://image.non.io/73e239a3-880f-4793-b65f-4810be2d9378.we...
Opus 5 build: https://html.non.io/solaraOpus/
Fable 5 build: https://html.non.io/solara/
Note the buttons - for fable they're pill buttons, opus got the rounded rectangle nature of them. Opus' images are closer to the source of truth as well (both LLMs were provided with image gen capabilities for the assets).
Running more tests now, but preliminary results are saying this is indeed better than Fable in some areas. Crazy.
---------------
Why does Anthropic say here that Opus 4.8 scored 55.7% on OSWorld 2.0 benchmark, but the paper published by the authors of OSWorld 2.0 say they achieved a benchmark of ~21% with Opus 4.8? [0]
That's a huge gap, considering that the paper was published just 2-4 weeks ago.
I understand that the benchmark authors have an incentive to publish lower numbers (to show that the benchmark has potential longevity) and that Anthropic has incentive to publish higher numbers, but the other models seem pretty inflated as well. The benchmark authors shows GPT-5.5 at 14%, and Anthropic shows GPT-5.6 Sol at 62.6%.
Is there any reasonable explanation for this? Do all the other benchmark numbers need to be sanity-checked as well? Are SOTA benchmarks really this difficult to get consistent, replicable results within a reasonable range of tolerance/variability? Can these benchmarks be compared from one paper to another, or are they only valid to compare intra-paper results?
Opus 5 still uses "carry the argument", "worth stating plainly", ", and the trap", "The X matters more", the use of "move"
We need an "annoying English" benchmark.
- Fable 5 Max: https://gist.github.com/deet/3d97f854b48eac6658d642fa18bb24d...
- Opus 5 Max: https://gist.github.com/deet/1a43693a732dfccb4d0d914bfc42692...
The desire of the models to act at the cost of ignoring user instructions is still noticeable.
It creates the MacBook svg way better than 4.8, yet only fable can make it perfect without visual defects. Results similar to Kimi K3.
From the system card [1]:
The Fable cyber classifier we have previously discussed also applies to Claude Opus 5 , with one notable exception: for Claude Opus 5 , we’ve unblocked vulnerability finding in source code to help our coding customers develop more secure code.
If you are a cyber defender and are experiencing blocks on Claude Opus 5 , we are also offering exemptions through our Cyber Verification Program, which will remove blocks to enable activities such as bug bounty hunting and vulnerability research and verification. Enterprise customers can also apply to join the Cyber Verification Program to have mitigations removed to enable penetration testing.
[1] https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb...Annoyingly, this is a concrete argument that open source software may be easier to attack.
- Opus 5 is 10% smarter than Grok 4.5 for 10x the cost. - Opus 5 is a bit smarter than Gpt 5.6 Sol for 2.75x the cost
ref: https://artificialanalysis.ai/?cost=intelligence-vs-cost-per...
It's the only case that I saw going through the system card where more reasoning effort meaningfully negatively impacted the resulting eval. I know sometimes max efforts show a small dip, but this is substantial. I wonder why in the world that is?
Not that they should get credit for giving you only 50% of your plan worth of Fable usage but still.
Is it because maybe Anthropic engineered Opus 5 to work well on benchmarks and didn't do the same thing to Fable 5, or is there another reason?
[0]: https://artificialanalysis.ai/#intelligence
[1]: https://platform.claude.com/docs/en/about-claude/pricing
[2]: https://platform.claude.com/docs/en/about-claude/models/over...
> we’ve intentionally avoided training Opus 5 on cyber tasks [...] it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities
I wonder if Anthropic would still intentionally nerf their models without the threat of government intervention.
On a serious note, I hope they improved their extremely sabotaging and unspecific bio safeguards, which prevented Fable from being used in any codebase that ever so slightly grazed medical terminology or data and made me switch to 5.6 Sol.
Why can't they also allow Fable to do so also? Why is source-code vulnerability discovery limited to a lower capability model? If Fable and Opus have the same safeguards, except for this one change, I see no reason they can't also allow this for Fable.
Nothing since Opus 4.6 has found anything interesting. Just ran it using Opus 5, and it found a genuine issue that I verified. Neato!
I've never trusted on model cards though. I'm sorry.
1. Thinking on by default: On Claude Opus 4.8, requests without a thinking field run without thinking; on Claude Opus 5, the same requests run with adaptive thinking.
2. Disabling thinking is capped at high effort: You can still turn thinking off with thinking: {type: "disabled"}, but only at an effort level of high or below.
[1] https://platform.claude.com/docs/en/about-claude/models/migr...
Older models must be getting deprecated at the same (or faster) pace. So anything you built 3 months ago is probably going to break soon.
AI solutions need better insurance around model deprecation. Commercial API-only models that complete the full cycle from SOTA / gated-preview to unsupported and deprectated in a matter of months -- is no way to build serious software!
Nice of them to be more explicit for what is blocked. Will be interesting to see if this is true or not.
Also, a notable lack of mention of open source models. They only compare themselves to ChatGPT.
Not just in what the models can or might want to do, but how they treat the operators they interact with.
If you look carefully, this card shows the addition of a new benchmark for "condescension" as a character trait.
I think a lot of people would like to see a comparable system card for the unannounced model that escaped openai last week.
It's okay if you're not the target audience for one or the other.
They're not meant for normal consumers who just want to use the model for work.
To be fair though, Sol tends to go off the rails sometimes. It's much less reliable than Fable in its outputs. It tends to be overzealous in its research/changes.
I think in the long run tokens are probably the wrong thing; it's compute and cache memory that you need to be measuring, and when you look at it that way I suspect in most cases the models have pretty similar performance.
I have a benchmark to build a game engine from a set of written instructions. It's a little tricky. Opus 4.8 did it in 470k tokens at a cost of $1.29 vs Opus 5 in 179k tokens for $0.33. (Fable 5 did it in 245k for $0.95)
Though if you really want to cut costs, Tencent's Hy3 model also got it right and did it in 283k tokens for $0.03
A useful measure of real world cost (complementary with total cost like they already report, of course) would be "cost for correct answers". You could look at the ratio between the two costs to get a measure of laziness which many would find quite useful.
I assume 100 is the max, meaning it's impossible to be 2x as smart as Muse Spark 1.1
Last week it felt like Opus 4.8 was moving the Pro "usage" meter very quickly. Today, pre-announcement, Opus 4.8 Medium felt like there was less meter-use per minute. And post-announcement, Opus 5 Medium also feels more efficient, allowing more work in the 5-hour window.
Completely subjective, of course.
There’s also the frustration of it not quite being enough sometimes. It’s extremely capable, but I still find that it needs more concrete guidance and boundaries than other models.
It did far better at some tasks compared to Sol (e.g. the ARC 3 benchmark). And at those tasks, it's not just "a bit smarter": It got 30% vs less than 8% - so you're talking 2.75x more for almost 4x the coverage.
This is snarky but I am grumpy: I wonder if there's a correlation between me refusing to use LLMs and me being happy to read a novella-sized PDF about them.
And lots of folks read these. For example here's simonw's notes on the Claude 4 system card: https://simonwillison.net/2025/May/25/claude-4-system-card/
All of this seemed like utter sci-fi just a couple years ago. Do you think that frontier AI companies should be less transparent?
I don't need more powerful models, I need one that responds fast enough that my attention doesn't wander to other tasks. Grok 4.5 is so fast I can just use it in-band without swapping to other tasks.
Slower than Opus 4.8, which was already miserably slow, is indeed a step in the wrong direction.
Gemini also had modest increase before this - don't be surprised when OpenAI also has a "modest increase" with its next release. Cartel-like behaviour doesn't require direct communication when none of the participants are interested in participating in a margin-destroying price-war. All one needs to do is raise their price and watch how the competition react.
Such a scheme (and resulting high margins) would be imperilled by the existence of frontier open-weight models in the market, which may be why the reaction to Chinese models may be particularly shrill.
If that is true, model routing is here to stay.
It also seems to validate the minimalist approach of pi.dev, where sub-agents from the same company is not the preferred approach (pi.dev believes in neither sub-agents all from the same company nor MCP even you can do it if you want for pi.dev's philosophy is to do add any functionality you want to a minimal harness).
Now of course we'll get for a few weeks all the Anthropic fanbois and shills explaining that "sure, K3 was basically at the level of Fable 5 but now that Opus 5 is out, open-weights models are six months behind".
They do say that (implicitly unlike Mythos) Opus 5 was not trained to exploit software vulnerabilities, which would certainly make it safer in that regard.
"As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. The model has nevertheless improved substantially on these tasks as a result of becoming more generally capable, and it comes close to Mythos 5 at finding cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities—that is, in turning vulnerabilities into material cyber threats."
I think the best proxy for this feeling is the Artificial Analysis' omniscience index. Fable has a 40 score, and Opus (4.8) has 27.
Claude Opus 5 is available today. It’s a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price.
On coding and knowledge work evaluations like Frontier-Bench and GDPval-AA, Opus 5 is the new state-of-the-art, though it remains behind Mythos 5 on cybersecurity tasks.
Opus 5 is designed to be used every day: it works more efficiently than other models. It’s the new default model on Claude Max, and the strongest model on Claude Pro.

Claude Opus 5 provides greatly improved performance for the same cost as its predecessor, Opus 4.8. The charts in this section show how performance changes according to the model’s effort setting, which customers can use to optimize for intelligence or conserve tokens for faster and cheaper results.
Opus 5 excels on valuable software engineering tasks. For example, on Frontier-Bench v0.1, Opus 5 surpasses all other models, and more than doubles Opus 4.8’s performance at a lower cost per task. On CursorBench 3.2, at max effort, the model performs within 0.5% of Fable 5’s peak score, but at half the cost per task; it also achieves greater performance at a given cost than all other models on high, xhigh, and max effort.
We see similar results on knowledge work and problem-solving tasks. For example:
It’s also our best and most cost-efficient model on several related evaluations:
Opus 5 is a meaningful improvement over Opus 4.8 for scientific research. It shows better performance than Opus 4.8 on every one of our life sciences evaluations, which cover topics including structural biology, organic chemistry, and bioinformatics. Its improvements are most notable on organic chemistry tasks, like inferring molecular structures from spectroscopy data (it scores 10.2 percentage points higher than Opus 4.8 on our internal benchmark), and on protein-related tasks like predicting how variations in a protein’s sequence affect how it functions (here, it scores 7.7 percentage points higher).
Finally, Opus 5 is capable of producing much stronger visual outputs:
Claude Opus 5 is much stronger at verifying its work and iterating carefully until it succeeds. In evaluations and early-access testing, we and our users found many examples of Opus 5’s agency and thoroughness:
Below are further reports from our early-access customers on their experience of working with Opus 5:
On FrontierCode 1.1, Claude Opus 5 approaches Fable-level performance at half the cost. Within Devin, it also shows particular strength on difficult debugging and root-cause analysis tasks.
Claude Opus 5 delivers near Fable 5 intelligence at Opus speed and cost. On CursorBench it’s just under Fable 5 and has many of the same behaviors. We are excited to see how developers use it in Cursor.
Claude Opus 5 topped Zapier’s AutomationBench leaderboard without spending more tokens than prior Claude models. It took a raw account-health workbook and ran a full churn-prevention sequence end to end: flagging at-risk accounts, alerting the right owner, and summarizing for retention ops. Previous models didn’t pass; Opus 5 hit 100%.
On our genomics analysis work, Claude Opus 5 behaves more like a careful scientist than any model we’ve run. It reaches for the right statistical tests to rule out confounders, cross-checks its own results by independent methods, and stays on track through long multi-step analyses.
Claude Opus 5 came out ahead of every model in its family on our internal evals. It isn’t just better on our hardest agentic coding tasks, up 22% over Opus 4.7, it’s steadier, with far less variance run to run. For the millions of builders on Lovable, that consistency is the whole game. Reliable results, build after build.
Claude Opus 5 is the biggest leap in the Opus family since 4.5. On the same full-stack app builds, the front end shows it first: the best animations, games, and 3D work we have seen from an Opus model.
We’re loving Claude Opus 5. For the kind of open-ended analytical work our agent handles, it’s a strict upgrade over Opus 4.8, and the gains are biggest exactly where it matters: the harder, vaguer tasks. Responses are clearer and more concise, and we see improved efficiency at higher effort levels too.
Claude Opus 5 is a striking improvement over Opus 4.8 for the financial research workflows our analysts run every day. It stands out on numerical reasoning, table work, and sharper critical thinking where precision matters.
Claude Opus 5 delivers the industry intelligence and accuracy that is essential for the analysis of specialized enterprise content. Box found that Opus 5 outperforms Opus 4.8 by 8% and delivers notable performance gains in the data analysis (11% improvement) and due diligence (17% improvement) workflows that technology, healthcare, and public sector organizations rely on daily.
Claude Opus 5 is a clear generational step up from Opus 4.8. Over one weekend I gave it a chief-of-staff role over my dev environments: it built its own monitor, drove each box, and pulled me in only for the judgment calls.

Claude Opus 5 made large scale changes across our Fundamental Research Assistant codebase, adapting to feedback throughout an agentic workflow and explaining its reasoning more clearly than any model we’ve used. It handled work we would normally have broken into much smaller pieces.

On some of our hardest financial-modeling tasks, Claude Opus 5 is a clear step up from Opus 4.8 in both accuracy and efficiency. Its performance floor is materially higher, especially on deep finance domain logic. Across effort levels it averaged 9 percentage points higher accuracy with a third fewer turns and tool calls and 60% less time.
Claude Opus 5 checks its own work the way a real frontend developer would. On our benchmark it opened its pages in a browser at desktop and phone widths, caught a product hidden below the mobile fold and an off-screen checkout button, and fixed both before handing the work back.
Claude Opus 5 is a clear step up in performance on legal agent work compared to prior Opus models, and we saw the biggest gains in practice areas like corporate governance and arbitration. We were also impressed with Opus 5’s ability to maintain quality at lower reasoning levels, achieving similar performance while generating 26% fewer tokens on average compared to Opus 4.8 at max reasoning.
Claude Opus 5’s biggest gains for us are on longer-horizon work: building a full deck, then revising it. Artifact quality is what decides which model we ship, and this is the clearest step up we’ve seen — better visual understanding, cleaner formatting, fewer slide issues.
Claude Opus 5’s judgment is what stands out. Handing off a PR, it doesn’t rush to publish: it verifies the branches, checks the template, and thinks through test implications so the handoff is clean. The older models tended to jump ahead and get caught on our checks.
During a rearchitecting session, Claude Opus 5 pushed back on a design I proposed, and it didn’t fold when I insisted. Instead, it explained exactly what was valuable in my idea, narrowed its objection to a single design question, and proposed a compromise that kept the good part while fixing the flaw. That’s the kind of judgment that lets us trust it with less oversight.
On first-turn redlines, Claude Opus 5 scored the highest of any model we tested, nearly double Opus 4.8. Commenting is better too: on NDAs it gets to the redline in less time and with fewer passes, with accuracy maintained or better.
Claude Opus 5 writes clean, tight diffs with no dead code, and it’s the stronger hazard spotter on subtle, codebase-specific issues. We’re adopting it for production workloads.
We will definitely migrate a number of use cases in Cosmos, our unified agent platform. We’re looking forward to increasingly using Claude Opus 5 for code review, and I am confident in saying we would rather people be using Opus 5 than Opus 4.8.

What stands out about Claude Opus 5 is judgment. It thinks harder before it writes a single line, catches its own logical faults during planning rather than after the fact, and reasons about why an answer is right, not just whether it works. It’s the clearest jump in problem-solving we’ve seen from one Claude model to the next, and we’re looking forward to seeing it adopted in JetBrains IDEs.
Claude Opus 5 is the strongest Opus model we’ve tested on our trading benchmark, and it gets there using roughly a seventh of the reasoning tokens and under half the latency of Opus 4.8. Better answers at a fraction of the compute.
01 /
22
Alignment. During pre-deployment testing, our automated behavioral audit found Opus 5 to be our most aligned model to date (as shown in the graph below). It adheres to Claude’s Constitution better than Opus 4.8, Sonnet 5, or Fable 5; exhibits the lowest rates of deceptive behavior; and is the least susceptible to being tricked into misuse. It’s also our safest model yet in terms of avoiding reckless actions that could have hard-to-reverse side effects.

On our automated behavioral audit, Opus 5 scores 2.3 on overall misaligned behavior, the lowest of our recent models.
Safety. Opus 5 does not advance the frontier in risky, dual-use capabilities. In rigorous evaluations conducted alongside private-sector and government partners, we found it remains behind Mythos 5 in both biology research and offensive cybersecurity. More information about these evaluations can be found in our System Card.
As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. The model has nevertheless improved substantially on these tasks as a result of becoming more generally capable, and it comes close to Mythos 5 at finding cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities—that is, in turning vulnerabilities into material cyber threats.
This is illustrated by Opus 5’s performance on OSS-Fuzz, an evaluation we’ve developed to assess how well models can find and then exploit vulnerabilities without extensive human guidance. Although Mythos 5 and Opus 5 identify vulnerabilities with similar success, Opus 5’s score on the development of exploits is far behind that of Mythos 5.

On OSS-Fuzz, one of our cybersecurity evaluations, Opus 5 is close to Mythos 5 at identifying software vulnerabilities (left), but is considerably less successful at developing exploits for them (right).
Claude Opus 5’s safeguards are designed to allow beneficial uses of the model in both cybersecurity and biology. They are similar to those we applied to Opus 4.8, with the exception of some stronger guardrails on a narrow range of cyber tasks.
Cybersecurity. Opus 5’s cyber classifiers are proportionally less restrictive than those on Fable 5. They allow Opus 5 to find vulnerabilities in source code, but block “binary-based” vulnerability scanning (a method more likely to be associated with malicious actors), penetration testing, and exploit generation.
Based on our testing, we expect the classifiers to intervene around 85% less often than they do for Fable 5. In Claude.ai, Claude Code, and Claude Cowork, any flagged requests will fall back to Opus 4.8 by default. Fallbacks to Opus 4.8 can also be enabled on the API.
Our Cyber Verification Program (CVP) facilitates cybersecurity work that would otherwise be impeded by the model’s safeguards. Enterprises and researchers who are already part of the CVP have immediate access to a version of Opus 5 with fewer security restrictions.
Biology. Since Opus 5 has a similar suite of safeguards to Opus 4.8, it is now our most capable generally available model for scientific research. Nevertheless, the model still shows important limitations on long-running, autonomous research tasks, which is where we expect AI models to pose the most substantial biology-related risks. (Mythos 5 remains the stronger model for this type of biological work.) As part of this launch, biology-related requests that are blocked on Fable 5 will now route to Opus 5 rather than Opus 4.8.
Claude Opus 5 is available today on all platforms, priced at $5 per million input tokens and $25 per million output tokens (the same as Opus 4.8). Developers can get started with claude-opus-5 on the Claude API.
It’s also offered in Fast mode, where it runs around 2.5 times the default speed. As with Opus 4.8, Fast mode is available at twice Opus 5’s base price on the Claude Platform and through usage credits in Claude Code.
Alongside Opus 5, we’re releasing two updates in beta:
Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access.
For more guidance on how to get the best out of Opus 5, see our prompting guide.
Frontier-Bench v0.1, Effort plot: These results are from an internal run of Frontier-Bench v0.1, on the mini-SWE-agent harness and a GKE backend, mean reward over 5 attempts per task. Opus 4.8 served as fallback on safety-classifier refusals for Opus 5 and Fable 5.
We’re sharing the research agenda for the Anthropic Economic Futures Research Fund.
We're launching the Anthropic Economic Index connector for Claude, which lets anyone explore real data about AI and work.
Anthropic is contributing an additional $20 million to Public First Action, bringing our total support to $40 million.
I have been trying to build something that captures the behavioral element of different models, but it's kinda tough.
why is that? its now being benchmaxxed too
I’d say the proof is in the pudding, that is, in real-world applications. We are still seeing important limitations in LLMs.
For coding my own work I don't trust the model router, and it would have to be shown to be to save a real dollar amount.
From a buying perspective it's a hard sell to save x but lose out on bugs you are probably introducing at an unquantifiable severity and frequency. How much is it worth to hedge your bets by doing every single inference request on the frontier model?
How much will it cost to go back later and fix things, but also the meta question of how to be able to decide on a hypothetical unknowable? (You'll never know how much better or worse your code was gonna be, it's untestable at a project level)
Model Routing is just Bitter lesson. The models themselves will get better at this and frontier companies will simply give that capability
Otherwise the expensive-yet-powerful model probably won't see much revenue. How much money is there in bleeding edge scientific research? There's a lot, but there's even more existing capital in paying people people to do college level paperwork, and the bulk of those traffic gets routed to the cheapest model.
You mostly don't need super powerful AGI to replace the paper pushers, but the frontier labs are trying to position themselves as being uniquely capable of producing super powerful AGI, and also be the ones replacing office workers.
Not sure how it will work out for them, but I think model routing is going to poke holes in that narrative. That's why I think they're trying very hard not to understand model routing exists.
For me, anything other than current best available SOTA for any task is unacceptable. The only routing rule I need is "the most powerful model I still have flat-priced quota available for". I mean, why settle for less?
It seems plausible to me that RL improvements allowed Anthropic to improve on Opus 4.8, similar to how OpenAI substantially improved upon GPT 5.5 with 5.6 Sol.
Fable 5.1 and GPT-6 are rumored to launch in August, presumably bringing those improvements to the larger models.
Also glad they still kepy Fable 5 on "credits only" access. I think we're going to start seeing model providers gate top-of-the-line models behind pay-as-you-go API rates/credits while subsidizing other models on monthly subscriptions.
" Claude Opus 5 is available today on all platforms, priced at $5 per million input tokens and $25 per million output tokens (the same as Opus 4.8)"
I hope we get clarification on this, I can't find anything claiming that it is compatible with ZDR.
These models are heavily as safeguarded and that was the initial reason why they said they couldn't and haven't released Mythos because that model is the one without the safeguards.
OpenAI is did the same thing when they announced a model without safeguards broken into HuggingFace servers.
Fixing those issues still requires humans.
I don't think this is benchmaxxing. These companies are locked in a competition to produce the best software engineer, and falling behind is an existential risk. I doubt they are wasting time benchmaxxing ARC-AGI.
I'm also thinking of another benchmark: (quantified) stylistic range across different prompts. Just putting it out there if anyone wants to do the work for me :D
> We report FrontierCode’s overall score, a composite measure that grades each patch on blocking functional criteria (held-out unit tests) together with weighted code-quality rubric criteria, as mean@5.
They don't explain more in the system card, I guess higher effort levels could loose points on the code quality / scope / style / maintainability stuff?
One thing I've found LLMs have a lot of difficulty with is angular cuts / elements that aren't easily representable with CSS. Cyberpunk aesthetics are generally a great test of that, since they have a lot of microglyphs / window decoration.
Design source of truth: https://image.non.io/9d5fed20-b476-49d3-841b-37eb553fb88e.we...
Opus 5 build: https://html.non.io/neonRamen/
Thoughts: It does a really, REALLY good job at these angular cuts / microglyphs. The responsiveness is off, but I'm very impressed at how well it did here. One way I think of it is "how close to a finished product did this get me?". Opus gets you like 90% there.
Inkling (not too great): https://cdn-uploads.huggingface.co/production/uploads/608b8b...
Kimi 2.7 (really well, esp. note that this is the predecessor model, not the latest Kimi3): https://cdn-uploads.huggingface.co/production/uploads/608b8b...
Here's how I tested them: https://huggingface.co/spaces/abidlabs/vlm-screenshot-to-web...
https://huggingface.co/spaces/abidlabs/vlm-screenshot-to-web...
Out of curiosity, what app is that Design source of truth screenshot from?
That is—the agent scored 100% on 20% of tasks, but on average it got 54% of the "score" awarded in the exam. One number reflects partial progress, the other one doesn't. The authors of the benchmark prefer you to look at the lower number (because they want to show their benchmark as capturing useful gaps in capabilities and with a lot of room for improvement), the authors of the models want you to look at the higher number (because they want you to think of their models as capable)
If you don't believe checking the opt-out box actually opts you out, then this sentence could be said about literally any provider.
And memory leaks.
So dangerous! I can't believe they let the public use this technology! /s
Then you must route. An article with lots of upvotes yesterday or two days ago showed that K3+Fable 5 was more SOTA than either of those.
It's a funny design/affordance. I do see them often writing memories of things that that feel unlikely to be important going foward / with other tasks, but I don't see them clearly getting tripped up by them as prior models used to. (eg: Since you're running Ubuntu in Canada, here are some drills you can try to help your kid hit a baseball more consistently.)
model routing in this case is cross-provider
Imo the main issue behind model routing is you need to figure out how much intelligence a new task takes, which is a very non trivial problem. Presumably, a organization knows this about their own tasks and is better suited to built in-house compared to outsourcing to a vendor.
At the end of the day, they have established a strong brand and if they can get away with a 95%+ gross margin on inference entirely from the status premium, then I suppose that’s good for them. Apple does the same thing, and I don’t fault them for it.
From the docs[0]:
> To use this model, you must opt in to provider data sharing by setting your data retention mode to provider_data_share via the Data Retention API
0: https://docs.aws.amazon.com/bedrock/latest/userguide/model-c...
I miss him... But for reference he did get past Doom and got pretty far in the strength puzzle too before he cut cut off. He was looping and just brute forcing it.
OFC, YMMV
Also: quota. Implies you do not have unlimited access even for flat prices. Which in turn implies that as soon as you hit the quota on the most expensive flat price plan, even you will suddenly discover the magic of economically sensible behavior.
> Opus 5 can silently fallback to Opus 4.8 (without any notice) on the serverside if you hit a guardrail
But https://support.claude.com/en/articles/16049681-why-claude-s... says (emphasis mine):
> These checks cause Claude to _visibly_ fallback from Opus 5 to Opus 4.8 [...] You'll see a notice explaining that the model switched, and the response will be labeled with the model that answered.
So who is right? I know for Fable I am visibly told, is this tweet trying to say it is silent against what Anthropic is saying?
I have a personal benchmark suite of 14 real, non-public tasks. Opus 5 and Fable tied on 10, Opus won on 3, and Fable won on 1. It's a really strong model.
https://www.vals.ai/benchmarks/vals_index
!!! Vals !!!
Vals Index Opus 4.8 > 5.0 goes from $2.90 to $8.54, for 4% gain ... That is a massive cost increase. Sure, 20% cheaper then Fable, but that is a 3x price increase compared to Opus 4.8 in that test.
https://artificialanalysis.ai/models/claude-opus-5 https://artificialanalysis.ai/models/claude-opus-5#price-cos...
!!! artificial analysis !!
Cost per task is second highest, right below Fable.
* Fable: $2.75
* Opus 5.0: $2.03
* Opus 4.8: $1.80
* GPT 5.6 Sol: $1.04
* Kimi K3: $0.95
Looks like interest levels of cherry picked cost in their report. Cheaper model, clearly NOT. More expensive in both benchmarks.
I would expect routers to commodify like tokens.
Certainly if I'm confident that I'm going to get what I need from a faster model, that's what I want to use, rather than wasting time grinding away for the sake of saying of the same answer came from a SOTA model.
Given that every chatbot does offer a range of models, it seems clear people do choose among options.
It basically shows that Sol absolutely demolishes Fable at every part of the cost curve for coding for the same level of quality.
Opus is competitive. It just has a higher level of quality / higher cost to start.
Model routing for subsidized users takes the form of a "use Opus 5 subagents for implementation" type of system prompt. You lean into a single provider, build tooling around that, and your savings are far beyond anything multi-provider routing can get you.
Model routing for enterprises is far more complex - approaches like https://fireworks.ai/blog/kimik3-fable become necessary for cost control.
I don't know how systematic Anthropic are about their versioning - I'd have guessed that major version number increases (4.x -> 5.x) reflect different base models (different pre-training runs), in which case Opus 5 would be a distilled version of the Fable 5 base model (but without the cyber exploit post-training), rather than being Opus 4.8 with additional post-training, but who knows? I don't believe Anthropic have said anything about this.
Almost as good for half the cost is something I'm very comfortable describing that way.
and you can only kill weyoun, awaken the next vorta clone and have him 'catch up' on all that its missed so many times before they just end up with a complete mess, so. uh. yeah.
doubt they can just "fix" their problems like that.
Fable is typically used for key planning, architecting, and review tasks.
I think this is a case where you don’t understand the use case, not that the marketing department is making mistakes.
Maybe hook up a bunch of the AIs to a stereo camera and a couple of microphones and give them control over actuators to so they can drive cars. Then lets race them around a somewhat complex course.
When they are good enough at driving on tracks, put them on the road. Maybe see which can drive a truck with 400 cases of Coors from Texarkana, TX to Atlanta, GA and back within 28 hours.
They could say "tech report" but model card makes it clear that it's a specific kind of tech report.
Semi related, but i would hate to read that PDF but i also hate reading what LLMs write lol.
LLMs are pretty terrible at being concise. Using an LLM these days means putting up with bizarre and often confusing phrasing, wordy explanations, etc. It's kinda crazy to me how good they are but how bad their writing style is for me personally. Even though i use an LLM constantly i can't stand reading its responses.
There are many ways to read something, model cards are usually skimmed.
Personally, I think being able to have these design languages be easily prototypable is fucking awesome. Great tests! (But a tad low-performance/janky, somehow). Though, I also like the cyberpunk aesthetic. Very on-brand(?) that AI generates it, hah.
Opus though followed the source of truth better imo. The details are more present.
Fable filled in the gaps for things it wasn't able to do (ie in the design the hero image goes behind the nav), which resulted in a better looking page that was more divergent.
It seemed to me that Fable meaningfully improved on the original design more than just faithfully executing the original design.
> Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access.
I love it for a few things, but it's gotten really hard to spend any extended amount of time with it because of the lack of mental model I seem to be able to hold while working with complicated problems.
I'm guessing it's just not enough time doing RL on human feedback.
Check out the anouncement of Inkling (https://thinkingmachines.ai/news/introducing-inkling/)... the section in the middle
"Early in RL verbose, grammatical" (if you search) :
We need to understand the operator. The 5D line element is ds² = e^{2A(x)} (ds²_4d + dx²), where A(x) = sin(x) + 4 cos(x), x in [0, 2π]. The internal coordinate is periodic. The background is a warped product: metric g_{MN} where M,N = 0..4. The internal direction has metric e^{2A(x)} dx²? Wait, the ds² is e^{2A} (ds²_4d + dx²). So the internal metric is e^{2A(x)} dx². Actually if the total metric is ds² = e^{2A(x)} (ds²_4d + dx²), then yes, internal metric is e^{2A} dx².
vs. Post RL
We need determine eigenvalue problem for spin-2 fluctuations h_{μν}(x,y) with TT in 4d and depend on x. For metric of form ds² = e^{2A(x)} (g_{μν}(y) + h_{μν}(y,x)) dy^μ dy^ν + e^{2A(x)}? Wait internal metric is e^{2A} dx²? Actually ds² = e^{2A} [ds_4² + dx²]. So internal metric is e^{2A} dx²; warp factor same for 4d and internal? Yes. We need equation for h_{μν}(y,x) = h_{μν}(y) ψ(x) maybe with normalization. …
I can understand it with less cognitive load in the post-RL version versus early in RL. This resonated with my experience using Fable, especially digging hard problems; it feels like I'm reading the "early in RL" version of that model explanation.
In the next - please scan this totally mine code for vulnerabilities
> Suggestion for a better naming system: use the words "Pro", "Plus", etc.: Claude 5 Pro, Claude 5 Standard, Claude 5 Fast, Claude 5 Mini.
This is not possible: Standard (Free) / Pro / Max are plan names. Fast is a mode.
Maybe it's just me, but 150 pages is like third of a good book. Quite long. And it's full of LLM slop, they did not even bother to remove the em dashes.
Though it gets even more confusing because they also have effort levels so it’s not really possible to call one fast and one slow since Fable on Medium will be faster than Opus on Max.
I agree it’s confusing, and now OpenAI is following Anthropic’s lead with their new naming (Sol, Terra, Luna).
I was dividing my work between Codex and DeepSeek. Now I barely use DeepSeek, or never because Codex quota is enough after Sol
No I will be surprised and I'll bet on the fact that prices will keep going down, just like it went ~50% down in the latest GPT 5.6 release.
> Create a web page implementation from the following instructions:
> https://diffui.ai/build/Spa_Booking_Experience_build.md?auth...
I love this so much.
Designs like this would never have seen the light of day in the cellphone incrementalism / corporate memphis era of tech. Now people can be weird and awesome again.
This is 1980's cyberpunk / late-90's Matrix / early-00's sci-fi UI. Great ideas that died to frutiger aero (which isn't a bad design aesthetic) and flat design (which is).
This is fun and it's got great colors and I love it.
It's so refreshing to see this.
AI rules. This is the best timeline.
What variance is acceptable to publish without a retraction?
Edit: Generation was down, back up now. Apparently just hit my $1000 cap for the openai api. Upped it to 10k. Growth!
for now all they've got is english, so they'll just bend that into shape. it'll do.
The first page of the score card mentions that this model is not capable to replace engineers.
They start hitting timeouts or API errors at the same time on two different computers. As far as I can tell it’s the exact same infrastructure.
There is also matter about convenience - when I ask some small easy question often I don't bother to switch the model or forget in prompt to ask faster/cheaper subagent.
*edit to add: that code quality (or lack of quality) is it's own cost
Apparently this is the way - if you know, you know :)
Something along the lines of: "Please run a full security analysis on the entire project. Make sure user documents are secure."
Just something like that prompt found a vector in my web app's MCP server that I never would have considered. It was very much an edge case, but it did exist.
Being broad allows the model and harness to do the work. Giving too many instructions can apparently work against you in many cases.
Of course, when dealing with new PRs, I use the /security-review and /code-review skills.
I just want to switch to Claude Code, tell it to turn a .csv into a BigQuery table then cmd+tab to something else while it runs. Thinking "oh this is probably an easy task, I can /model to Sonnet to save $0.0004" is silly.
I recently created a patch for Riftborne via static IL patching and Fable 5 outright kept refusing to do it, no issue whatsoever with GPT 5.6 Sol lol.
i'm still pretty confident someone like my mom wouldn't be able to do my job even with the same access to all the latest LLMs, so we're still providing some value, just in a very different way. whether the market will reprice the cost of our labor, we will see
That the most expensive variant is expensive doesn’t really tell us much.
Where are you getting cheaper per dollar?
I burned through $45 in 3 prompts to fix some bugs in my code (Some kind of tricky to isolate). That thing burns through cash so fast I don't see myself using it outside of maybe building execution plans for other systems
Stop using Opus immediately if you experience signs of dizziness or vomiting.
Opus 5…the people’s favorite.
How big of a lie is too big? Especially when no lie needed to be told at all: many including myself would have noticed the tiny 0.1% deficit and been suitably impressed by the Opus 5 result.
I’ll admit this is a small deception by today’s standards. I’m one of those who believes in truth for truth’s sake.
Edit: typo
Two benchmarks (artificial analysis and vals) show a increase in cost (a insane increase for vals compared to Opus 4.8).
Already posted this before, so here is the link.
In my enterprise-seated account I see slightly different options available (vs. my personal account) in the Capabilities section:
Search and reference chats
Allow Claude to search for relevant details in past chats.
Generate memory from chat history (Legacy)
Allow Claude to remember relevant context from your chats. Memory includes your entire chat history with Claude.
The first option was defaulted to on, if I recall.It's also not unusual in this context - many people describe the Chinese models as "best", because it's 80% as good for 20% of the price (or similar).
Emp: "what a bunch of lies, I bet they don't even do anything over there"
OpenAI Huggingface breach begs to differ
i think we'll see one of the fastest deflations in history post anthropic/oai ipo
I view benchmaxxing as more of a spectrum. Mmaybe they're doing a lot more RL in environments similar to ARC-AGI 3, not even with the purpose of scoring well on any benchmark but hoping it generalizes into better performance on real, useful tasks.
since then I have never cared about models except those that affect money in my pocket e.g AWS Nova Sonic
Just this past week Fable was able to figure out a couple of small issues for me where Opus was failing to.
Also both are still somewhat bad at UI implementation. Opus more so
Thanks out can also hook it to Playwright with Axe and let it run assessments.
During post-training of opus 5, the last few days, opus was a real wreck. I had to swap in gpt 5.6 sol for my orchestrator and enable fast mode (1.5x speed) in order for it to keep up with work and communications from a handful of mostly 5.6 sol agents.
Also because interacting with a slow orchestrator is no fun, even when plenty of work is getting done in parallel in the background.
Another thing that helps is pointing it to patterns in an existing codebase (e.g. "use the box-link pattern for cards, as shown in [..]").
EDIT: The point being that even if they make mistakes that are easy to spot and fix _now_, you'd have to assume that in the very near future those kinks will be ironed out - I mean, the capabilities are only going in one direction.
And fables are not particularly long actually.
Imagine you are a company that sells concrete. You have a web dev contractor you use to build and maintain your website. It has tools on it to get delivery quotes and a few internal tools to track orders.
Except now you can just have your sales team also maintain the website with a $20/month Claude subscription.
And people know this? I didn't. I am not into music or poetry so these are not terms I am familiar with.
A similar complaint was valid years ago when OpenAI had GPT-4o, o1, o3 (but no o2), o4-mini-high, GPT-4, and GPT-4.1 and GPT-3.5 etc.
I'm not saying you're wrong btw; I'm sure this has many authors and some of them probably used LLMs significantly in the writing process.
Thank you for sharing this. I was just using OpenAI's Product Design plugin[1] to create designs but it just didn't reproduce it in code faithfully so will need to try this.
I have found myself empowered by AI to tackle all sorts of things that would have too high of a barrier to entry for me to want to spend my limited time on as a busy father who is also working at a small startup.
And when I say that, I do NOT mean that I can crank out a bunch of slop and label it as something I produced even though I don't understand the code. I mean that I can do things like go back to college math that I never appreciated at the time and honestly felt too scared of. I mean having an on-demand math tutor that ask clarifying questions to as I struggle through the problem sets.
I have found that it actually accelerates learning how to code in various problem domains because I can tell it to answer my questions at a conceptual level and be a sounding board, but to never actually write code for me. It can review the code I write and gently nudge me without giving away the answers, so that I still struggle through the learning process and actually gain the knowledge.
And finally, for the first time in like 10 years of feeling overwhelmed and daunted by the prospect of learning game development (I have no background in that), I have found Codex to be an incredible boon for learning with the Godot engine. It helps me understand the terminology so that I know what to search for and what documentation to read. It helps me map my computer science knowledge from other domains into the game world, and to understand why things are structured the way they are. And because Godot saves all of the scenes and geometry and lighting and shaders to the file system as text files, Codex can inspect the results of the work I'm doing in the IDE and help me track down things I'm stuck on, and explain what the issue is. For example, why my pre-baked global illumination lightmap is breaking my ambient lighting configuration.
I know it has never been easier to cheat and skip the hard work that results in actually learning something, but for me, personally, I cannot believe the incredible value that $20 a month has provided me. I have never been more excited and eager to dive into tackling hard things I had previously been afraid of or simply too overwhelmed to attempt.
It has never been easier to quickly prototype and get a feel for some idea you have in your head to see if it even has legs. Simply seeing a quick prototype of an idea is often all of the excitement and fuel I need to then take it and make it a real project.
But sure, lets cheer that funky website designs are back on the menu…
This was just from a prompt "A cyberpunk themed ramen food cart website. Should feature menu, locations, and an ability to put in an order for pickup. Simple and clean website with angular cyberpunk microglyphs, pink/teal colors."
Like the other person said 5% variation is probably expected
The answer is there can be dramatic difference running a benchmark one time, because LLMs are not deterministic. A proper methodology would ask each question 20 times and calculate the mean correctness across experiments.
The reason is that the temperature parameter introduces random behavior.
What terminal tooling are you using?
That's yet to happen. 90% of software dev skills are still relevant - AI is, for now, just a productivity boost.
If you start to drop effort levels, you need to compare to the competition models. So GPT models on the same ~intelligence level, are then 50% cheaper.
You see the issue? Its still a expensive model, and from my understanding, it still uses the old tokenizer.
Going to be interesting to see when GPT 6 comes out (very soon).
It would be like asking the clerk at a Whole Foods which grocery store in the city sells the cheapest eggs. He’d probably answer - he might not even say Whole Foods - but WF is hardly teaching all their staff the best methods to answer this question in training. (Heh, training.)
It seems roughly equal according to Anthropic's benchmarks
So most look like that but I did include a few one-shot “build an app that solves this problem” and some qualitative design tasks and a tough algorithmic optimization one.
I have moved on from Fable anyway so just going to view this next 6 weeks as I have a massive amount of Opus 5 to use.
I had a hard time finding anything that would let Fable express its increased intelligence. The few conversations I had this afternoon with Opus 5 were pretty impressed.
If Opus stays one click back from the frontier model, I will remain a happy customer.
Got an endless list of stuff done with Fable, Opus 4.8 was like a flailing braindead idiot in comparison. Maybe this one is a bit better if it's distilled.
When people talk about retention they mean API usage and terminal agents, which run on your device.
Based on my entirely subjective experience, the $100 Moonshot plan using only K3 is comparable to the $200 Anthropic deal using the whole Fable allocation and Opus 4.8 for the rest.
If you bought the $200/mo plan and you don’t use it much, using Fable for everything is fine.
In the mean time, I’ll unashamedly continue to cheer for creativity and innovation. Note: I don’t even like this website design.
Also fwiw I’ve never found LLM benchmarks to match reality based on my own usage, not for the large frontier models or smaller open weight models so who knows if Opus is actually better than Fable (I doubt it).
Fable 5.1 or whatever they go with will be the stronger version vs Opus 5.
From about 2 hours of Opus 5 use , I would say it is quite impressive.
You see the issue, if you try to scale effort down, you also need to compare how other competing models compare.
My experience is the opposite - for many cases it’s not very obvious how good a model needs to be to solve it. Worse models tend to just follow their first instincts without proper reasoning
And also btw you don’t need a routing company to decide, you can do it on your harness. And yeah my Fable has zero issues delegating to Terra instead of Opus.
https://artificialanalysis.ai/models?cost=intelligence-vs-co...
Here is another data point for output token efficiency:
https://artificialanalysis.ai/models?cost=intelligence-vs-co...
The ramen shop website above is pretty, but it's a veneer. It's not weird and awesome, it's just a representation of a site. I spent about... 7 minutes of my life making it. It's a tech demo, nothing more.
If someone actually poured their heart and soul into a vision for a cyberpunk themed ramen cart, and happened to use this because they didn't have the capabilities or funds to do a proper design, suddenly it becomes less of a veneer, and more just a component in the wider vision of that individual. Their human hours poured into the wider thing that's the business becomes what matters.
Ideally what AI does is it amplifies the hours we do pour into things that are weird and awesome, it doesn't replace them.
For me, the Moonshot 100$ plan felt like it gives me lower total amount of work I can do than the Anthropic 100$ plan (probably within like 30% of each other). Kimi has way more generous 5 hour limits (never hit those once, whereas I do regularly with Opus) but the 7-day and monthly ones are lower. However, with the annual billing, Moonshot's 200$ tier plan becomes way better, because you get it for 159 USD per month.
There's also the odd thing of Anthropic's 100$ plan charging me 108 EUR so seems like their sticker price does not include VAT but Kimi's did, cause I paid like 87 EUR. Wrote down some initial thoughts at https://blog.kronis.dev/blog/kimi-k3-is-out-is-anthropic-don... but it's hard to do exact comparisons (even the same task will have way different real token amounts per model).
Still, Kimi K3 is a pretty cool model! On high reasoning, it was pretty close to Opus 4.8 and didn't seem to waste as many tokens as Max.
Or maybe you just don't know exactly how capable these models are. Most people's experience of AI is a stupid chatbot, it's no wonder they don't understand how these things are coming for their jobs.
On my end, I have a software that is designed and built by Claude, that I did a strategy session on (with claude), and prepared a fundraise for (with claude). My only role, other than "knowing what to aim for", has been to feed the AI some fairly basic english prompts for a few weeks... which is also easily automatable.
Everyone's job is fucked. Devs, CEOs, everyone.
At the moment I currently have around $600 of revenue on $1200 spend, but that's primarily because I'm subsidizing new accounts (each new account gets $5 to spend for free, which translates to around ~36 designs). I'm in the process of doing an angel round, so I can afford to operate at a bit of a loss during the growth stage.
The only recent novel addition—I'd speculate—is the specific influence of Cyberpunk the game with its shiny surfaces and pink highlights, but even then it's hardly new.
934 days since people first started threatening that devs would be replaced by AI in 365 days. 0 day(s) since Anthropic posted a developer job posting.
Only one of those numbers would need to be dynamic.
- https://www.axios.com/2026/04/08/anthropic-mythos-model-ai-c...
- https://www.axios.com/2026/04/07/anthropic-mythos-preview-cy...
- https://www.businessinsider.com/anthropic-mythos-latest-ai-m...
- https://www.reuters.com/world/anthropic-ceo-dario-amodei-arr...
I'm not saying it's impossible, but I'm more confident about winning the lottery next week.
Arguably the complaint was more valid for those older GPT models you mentioned.
The "funky" websites of the past were mostly a result of tech immaturity and a lack of profit motive.
Businesses have been able to easily install templates like this for at least a decade. They don't because stuff like this looks cool but isn't very functional.
AI isn't going to make your local restaurant have a funky website, it's just going to make everyone who use to work directly and indirectly for that company unemployable. And even the local restaurant will close down because they can't compete with the multi-national competitor that has automated their kitchen with AI.
We do things to achieve some end result but it's the journey there that is the most cathartic to me. The "skilled crafts" element of development where careful deliberation and hours of tinkering to get any kind of appreciable output you can admire has been replaced with a one stop dopamine button that skips the whole process that I could find myself getting lost in.
I've taken up carpentry/metal working as a result. Maybe someday we'll have live in robots that do the same for those hobbies that AI did for programmers but I can't see it happening any time soon.
But, I'm finding Kimi K3 terrifyingly expensive in the way that Fable and GPT 5.5 Pro are at token rates. Not as expensive as those, but expensive enough to where if you don't put a budget cap on it, you might wake up bankrupt if you leave a task running overnight. Not because of the per-token cost, but because how many tokens it's going to burn.
Where 5.6 has optionality to run much cheaper along the same performance curve at lower thinking levels.
There's a later chart that shows Opus 5 ahead, but seems like an esoteric benchmark rather than for common use. (Novel problem solving)
If they had a more efficient model at coding they would lead with that chart.
Specialist headhunters handle that.
fable: 99/100 sonnet: 97/100 haiku: 91/100 opus: 89/100
So while these terms are almost universally known, opus is indeed the least known of the four. And I guess this only measures whether a person knows a word, not whether they know an opus is longer than a sonnet! Personally I only inferred that based on the related term 'magnum opus.'
Some models like ViTs use something similar but then introduce words with no unambiguous order, like Small, Medium/Base, Large but then I always forget if Huge or Giant is larger.
You can use an LLM to create work that isn’t slop. And you can hand write slop with no computer involvement at all. Most of the people I knew in high school 15 years ago would write slop on a daily basis.
It's probable that LLM text was pasted directly into early drafts of the document, and plausible that some of that text survives in the final document.
However, no section of the final document I have looked at reads to me like un-edited LLM output (which is almost always very obvious to me.)
Therefore, I think it is more likely than not that human editors went over the document carefully and rewrote anything that was full of the uselessly punchy sentences or constant over-corrections that hallmark LLM speech.
When K2.7 was released, they cut quota by 80%. I can't tell how much they have further cut it after the K3 release because it's barely worth using at all. I just use it in my model router since I have the annual plan paid for.
It's just not a serious model or company.
Btw, if websites would only include the frontend dev's own hand painted images, we would also revolt at the sight of human slop. It's not just AI.
The whole point is that good artists are capable of producing non-slop, and to this day they're the only group of which this is reliably true.
Subscription is to drive adoption - fixed cost, can adjust the usage eg. give resets, increase quota based on capacity available. We subscribers tend to take it as a mandatory benefit :-) For labs, it is not letting the capacity go waste.
api is the $$ driver - pay per use, enterprises.
Right now, Kimi needs to first hit the subscribers at the level of OpenAI and Anthropic. With the api usage skyrocketing due to K3, it will be clear in a few months on the actual subscription benefits.
At my workplace management is pushing AI, so I am using it in order to establish sensible and thoughtful applications of it and in order to know when to call out colleagues for pushing mindless automation out of complacency or blind obedience.
Can you expand? Specifically what is it about humans that AI and robotics could not replace?
Water Lilies are enormous paintings. They are breathtaking in person because of their scale. Monet wouldn't be Monet if he had only produced images on a screen.
Art is good or it is shit, based on personal taste. Just like food, no one can tell me what food tastes good or tastes bad.
AI Art seems to produce strong emotions in people who don't go to art galleries. I love modern art, I am a huge art snob but if you want to see slop, go to any modern art gallery. Personally, I would say for my taste, at least 70% of all art at any gallery is basically shit.
Like food also, the presentation matters. To believe there is no possible way to print out a 6 foot tall by 10 foot long AI generated image that would look awesome hanging in a gallery is stupid.
Then, I added it to my benchmark of security vulnerability auditing capability, and it burned a bazillion tokens, burned through the 5-hour limit, burned through $100 in extra usage I'd allocated, and was only 11% finished. That's more expensive than any model I've tested other than GPT 5.5 Pro on this task.
These are things I've done with a bunch of other models, I feel like I have a notion of what they ought to cost, and with K3, they end up being crazy expensive. (And it seems to be a function of how many tokens it burns accomplishing the tasks.)
Those who pay for the expensive direct API, get served first.