They can do the difficult small level optimization, the boring but tedious code but cannot be tasteful.
That means I'm more valuable and more productive. Good stuff
> Maybe add sunglasses? no.
> Maybe add water? no.
https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
This reminds me of one of the predictions from https://ai-2027.com/ . Only that there it's "OpenBrain" doing this, not the Chinese. And the authors of that paper were also slightly wrong about "Mid 2026: China Wakes Up": China woke up already a while ago. And:
> But China is falling behind on AI algorithms due to their weaker models. The Chinese intelligence agencies—among the best in the world—double down on their plans to steal OpenBrain’s weights.
No need to steal anything, they have already caught up.
And then there's this prediction for February 2027:
> Officials are most interested in its cyberwarfare capabilities: Agent-2 is “only” a little worse than the best human hackers
I think we're past that point now, too…
- if you're gonna order the rest of the bar chart by rank, order your model accordingly.
- if you're gonna highlight a winner in a table of benchmarks, don't highlight your entire model row in the table.
Etc etc
Or is it like bicycles? Unless your problem is named Tadej, you don't need a $13,000 bike.
They all suck.
They shoulda put their stick where they belong, not at far left.
It just makes comparison to Deepseek 90% of them time as Hy4 has nothing to show off.
link to source code?
There are no open source models, at least not useful ones (yet) [0]. Open weight is not the same as open source. The current "open weight" models are just opaque binary blobs you can run on your own computer instead of through a web API.
Imagine thinking that running a Photoshop binary on your own computer instead of through a SaaS web app means that it's "open source". Of course you think that's ridiculous.
Edit: someone else commented that as I was typing this, lol.
Whether the distillation has constituted "attacks" or has or will meet the bar of "stealing" IP is not super interesting to me, though.
The first column has both the Hy4 and Hy3 scores overlaid on one another (Hy4 is darker blue and the taller one), with both scores written below the top of the respective bar - maybe you're seeing that?
But, what bars are clearly off? I couldn't spot any.
Models are lossy compressed datasets you can pick up and amend (fine tune / continue training / alter) according to license they were released under.
Hy4 is released under OSI approved Apache License 2.0.
Super computers keep getting better but most people don't need them for most things.
The solution to that (to my mind) would be not a better model but a basic shift in architecture beyond the current paradigm and into a setup where agents have durable, plastic memories and undergo contextual individuation over time. But at that point agents start to become quasi-persons and not tools.
https://martinalderson.com/posts/watch-out-for-cache-read-co...
Btw I still haven't came across any decent model that is <$0.01/MTok cache costs apart from deepseek thru their official API (even with the price increases).
Seems like a bit of an opportunity for someone to take - drop cache read costs significantly.
wouldn't trust they dont do Capitalism like the rest of the AI field.
For large tasks like a web browser or a compiler, even expensive swarms of frontier LLMs have not been shown capable of producing codebases that actually work. (Anthropic built a C compiler with Opus 4.6 but it lacked optimizations and apparently hit a complexity wall.)
I also want to use LLMs for reverse engineering, but apparently it's pretty hit-or-miss, especially if you're forced to use open-source models to avoid restrictions.
Both animated and live action results would be acceptable.
Unfortunately most existing LLMs lack the capability to maintain context across tens of thousands of frames.
The other option is that you do understand those words the same way, and the people making these (now nonsensical) anti-AI claims simply aren’t talking about the same programs/models we are. Their idea of SOTA is when chatgpt.com launched.
If you took a point sample pre-Opus, and didn’t write a good prompt, of course you would think all AI programming was worthless slop.
If we create a stripped-down vocabulary with greater token density to use less resources and to resolve ambiguities earlier in the semantic process, aren't we creating NEWSPEAK and dragging along the worst aspects of it? The ambiguity and multi-valence of words is what creates more connections between words, increases the directionality of associations, and expands the potential subtlety and depth of meaning. By paring down (or requiring verifiability) we make it harder to say certain things, or at least make it harder to unintentionally say something that makes MORE or DEEPER sense than what we intended. If the token density becomes extreme, you're left with something like a calculator.
Maybe this is the ultimate path toward better coding? But the worse path toward better genuine thinking?
But the reality is, the weights are a useful artifact that you can use to create derivative works. So, dismissing it as a photoshop binary is as technically wrong as calling it open source.
There are two more points in favor of this kind of AI movie project: there's zero chance that anyone would greenlight a Hollywood budget for the Silmarillion, and it is beyond human capability to write that screenplay.
It's also interesting because, while coding agents are important and are a notable success, they are never going to be a multi trillion dollar business. And are there any other domains where LLMs have such a large impact?
edit: I do wish openrouter would let you sort providers by Cache Hit % and Cache cost. These are the only things that matter to me at this point when choosing a provider.
Like lobbying the US president to harm their competitors?
I think, also, like in the traditional film makers career, this process should be built iteratively, start with a fast food commercial, then do a music video, then you can probably do a short film. Continue to improve the process, and one day I’m sure the LLM film studio can make you any movie you want, provided you have enough tokens.
Real humans get non-primary information from word variation. It's reasonable to hypothesize that it has a role in thinking things, because it endures. Our languages need to breathe over time, and flourishing might be one of the aspects that allows that breathing space.
Maybe add a small cycling cap or helmet if it doesn’t obscure the head.
Be concise.
OR
Brief is best. OR
Eschew verbosity etc.Caching was always here, you don't need to do anything special to get it on a single user local backend running a base model or a chatbot in the first place. Among commercial providers, OpenAI adopted it in 4o first.
EDIT: Your username doesn't help, either.
-- William Strunk Jr. and E.B. White., The Elements of Style
Ranked among the top tier of open-source models, Hy4 preview is built for real-world productivity tasks, delivering outstanding performance across coding, office work, and scientific research

Tencent has released and open-sourced Tencent Hy4 preview, a next-generation large language model with 770B total parameters and 49B active parameters, and a context window exceeding 1M tokens. It demonstrates outstanding capabilities on real-world productivity tasks spanning coding, office work, and scientific research.
Hy4 preview is now available as an open-source model and can also be accessed globally through WorkBuddy and CodeBuddy, as well as Yuanbao, ima and other Tencent products. Users can try the model directly through these applications, or connect to it via API through Tencent Cloud TokenHub and OpenRouter.
Upon launch, Hy4 preview will be available for free on WorkBuddy and CodeBuddy for two weeks. Free access to Hy3 on both platforms has also been extended until September 30.
Hy4 preview was expanded significantly in model size, context length, and data volume, and the advances in both pre-training and post-training have led to a major leap in overall intelligence, placing the model among the top tier of open-source models.

Hunyuan continuously works in deep co-design with products such as CodeBuddy and WorkBuddy, optimizing the real-world user experience across productivity scenarios. In a blind evaluation conducted internally by Tencent involving 163 experts and 203 engineering tasks, Hy4 preview scored an average of 2.99 out of 4.00, slightly ahead of GLM-5.3 (2.92/4.00) and Kimi K3 (2.94/4.00).
Designed for productivity, Hy4 preview was developed using high-quality training data co-created with Tencent experts across software engineering, gaming, finance, security, and other domains, as well as through deep co-design with products such as WorkBuddy. This has helped drive significant improvements across a wide range of real-world productivity tasks.
In software engineering, Hy4 preview delivers stronger understanding, planning, debugging, and validation capabilities for long-context development tasks, while also enhancing the visual quality and interaction experience of front-end development.
In office productivity and analytical scenarios, the model demonstrates a significantly stronger understanding of complex working environments and enhanced financial analysis capabilities. It has also been optimized for data analysis and cross-document collaboration, supporting the full workflow from information processing through to the creation of documents, spreadsheets, and presentations.
In game development, Hy4 preview can generate a playable prototype from a single natural-language request, and work effectively with game engines. Developers can then continue refining complex game projects through multi-turn interactions.
In scientific research, Hy4 preview demonstrates stronger capabilities in understanding, reasoning through and solving complex research problems, with notable improvements across areas including AI research and development, molecular dynamics simulation, condensed-matter physics and fundamental mathematics.
Notably, Hy4 preview also contributed to its own development process, participating for the first time in the automated optimization of training methods, data strategies, evaluation frameworks, and low-level operators. The model proposed approaches, ran experiments, and iterated based on the results, with the resulting code, logs, and feedback feeding into subsequent rounds of exploration. This established an early-stage recursive self-improvement loop.
Hy4 preview has also autonomously analyzed bottlenecks in its inference system and carried out multiple rounds of optimization on areas such as operator fusion and communication optimization. These improvements increased end-to-end throughput by 31.8% compared with the baseline, with consistent gains across different context lengths and concurrency levels. This demonstrates the model’s ability to autonomously optimize its own inference infrastructure.
Hy4 preview continues to offer cost efficiency, helping make advanced AI more widely accessible. API pricing is set at USD 0.834 per million input tokens, USD 2.501 per million output tokens and USD 0.042 per million tokens for cache hits.
Through a preview-first approach, followed by official releases, Hunyuan continuously incorporates real-world feedback into its research and development process, enabling its models to improve by solving real-world problems. The next batch of models in the Hy4 series is expected to roll out soon.
I thought you had to actively manage caches, do you not?
This behavior makes it so you don't benefit much from the caching, unless you pin it to a single provider.
These cache Hit % are accurate, I've done a ton of testing of this myself. The cache hit % is one of the most important metrics as far as estimating cost. There are many providers with cheap cache reads, but have an effective cache hit % of 30%, making their cheaper cache pricing meaningless compared to another provider who charges more but has a 85% cache hit percentage.
[0]: https://openrouter.ai/deepseek/deepseek-v4-flash-0731?endpoi...
scroll down on the provider/model card and you'll see a field called cache hit %, its different for every provider/model.
I don't use routing on openrouter, I strictly use models with a single provider and no fallback, at least for use with harnesses its pretty dumb to route requests to multiple providers you are busting your cache every other request and increasing costs by 20-50%.
I'm not sure it's wholey accurate to say they "randomize" the provider, rather my assumption based on usage is that it's something like cheapest-ish/responded to the request within some reasonable-ish time/etc algorithm that chooses the provider on each request - which seems, remarkably questionable in terms of optimizing for user experience or hidden user costs.
> This behavior makes it so you don't benefit much from the caching, unless you pin it to a single provider.
I so very much recommend this approach. My avenues that automate llm calls to openrouter are setup to make api reqs to openrouter to determine best price/response/etc and then pin the request to that (and, preferably, a fallback if there's reasonable difference between #1 and #2) provider for that session. Otherwise you're going to have a bad time.
I'd imagine this could make things interesting in cases where one provider is offering different quants than the others and openrouter is just swapping you back and forth on a long agentic session.
I don't believe this is correct? AFAIK once it routes you to a provider for a given conversation that choice is sticky unless you hit technical difficulties. (It's more complicated than that, they recently added named routing strategies that you can append to the model name.)
IMO the relevant metric is cache TTL which isn't typically published AFAIK.
Fine-tuning is great for really small models on specific applications, but it's not something that can essentially improve a more generic model.
That said, there seems to be a fine line in quantization+finetuning that could recover performance. It's just hard to get a hold of it (I feel it in some models, but it's hard to say yet; lots of small labs working on this RN).