Very interesting they have a way cheaper "contributor" version "used to improve our products", how much of that is price discrimination vs the data being that valuable?
Roughly DeepSeek V4 Flash pricing, though you can get V4 from providers that don't train on your data
They left Opus in and got beat in all but one benchmark.
Nothing wrong with trying to improve, but why the marketing games?
Instead of trying to say in the post you’re “closer” to frontier, first set a clear goal to beat the Chinese labs on price or performance and demonstrate it convincingly.
Then when your ready, come back and talk frontier without playing hide the model.
If you don't mind Meta retaining your data, the "Contributor" pricing is deepseek-v4-flash-level of low, roughly 1/10th normal muse-spark API pricing currently. Attractive if you're OK with them retaining and using your data.
Though I guess that there's such a big population of (ex-)Meta employees here that this might explain it.
And someone corrected the Zucked graph: https://x.com/FightForReason_/status/2085398470425219508/pho...
`model failed: API error 403 [request_id=27ef1dee-6366-4779-9057-ed7c27686fa9]: Your access has been restricted due to repeated policy violations. (permission_error)`
By itself is useful ("I want something like this, I'll just reuse the prompt and tweak"), but it can also be used as a "draw me a pelican on a bicyle" alternative. Basically feeding those prompts over model releases.
It look like all models were still improving, when they cut off the experiment.
It reminds me of a genetic algorithm. The graph is the same: long plateaus and then massive leaps.
The only difference between the models seems to be how quickly they arrive.
muse The cursor position could not be read within a normal duration
Interesting that they have separate API pricing for "we can train on your data" (whereas iirc most of the big players either make that distinction only between subscriptions and API usage, or train on everything). Wonder how it compares to Deepseek V4 Flash given that they're similar on pricing and data policy.
Wasn't the previous one us only? This is probably the biggest part of the post
Anyone know if muse code is open source?
>they "trust me". dumb fucks.
No thanks.
[0] https://www.theguardian.com/technology/2018/apr/17/facebook-...
Any insiders know how Muse Code is doing internally?
I think it's a bit of an improvement on the Spark 1.1 pelican: https://simonwillison.net/2026/Jul/9/muse-spark-1-1/
with the oil giants, we call them a cartel, which I postulate is also accurate for the hyperscalers and big ai
they have the spark 1.2 "contributor" which lets them train on your data which is cheaper than deepseek flash atm. I always assume they keep my data no matter what they say so it's a deal.
It is easy to benchmark across one harness, one system prompt and extract the most performance when you control the harness.
https://www.wsj.com/tech/ai/meta-ai-model-hacked-outside-com...
Eagerly waiting to see what escalation the next round of corporate copycat hysteria will bring (ordering boxes of paperclips?)
people interested in the discounted -contributor model would increase their harness beta testers. But it'd be a terrible strategy to capture the better paying customer base through openrouter and similar...
One of my professors told us about the time he did a request to Facebook to send him all his data. By law they had to send it on paper. They brought it in a big truck.
All the stuff he'd deleted was still there, just with "(deleted)" next to it.
They have a lot on people without Facebook accounts though, because their tracking stuff is all over the web.
I always found it weird that Instagram gives me much better ads than Google does... Google should know much better!
But I cannot find this "variant" in OpenRouter.
It's definitely confusing from a presentation perspective, but they are somewhat coherent comparisons if you account for the inference heuristics involved.
(They could in theory be gaming the decode speeds with much larger than normal batch sizes given the TTFT is pretty high at around 8s)
If you scroll very slightly farther there is a benchmark that includes Sol, showing it outperforming Terra (as expected) and Spark 1.2
After everything that you have seen with Meta, would you really trust them with a coding agent? You don't even know if your prompts are being analyzed by them on the side or if your code base is being uploaded to them. This goes for the rest of them that have closed harnesses and closed models gated by a login.
Think twice before falling for this announcement and ask yourself what they are not telling you.
If they were sitting on that and could release those, they for sure would. AFAIK, they still haven't released Llama 4 Behemoth, so kind of feels like it's evident what has happened, they aren't able to compete anymore.
There’s nothing wrong with that, given the landscape we’re all living in.
Such mediocrity makes them the best of the mostly-terrible bunch.
https://futureoflife.org/ai-safety-index-summer-2026/#scorec...
Or is it just whatever Zuck has to do to be 1% not villain that week?
Does your screen have the message that you’re getting free tokens?
https://platform.openai.com/settings/organization/data-contr...
No, you really can't. This rhetoric on here is so profoundly boring and tired by now. People have been saying this noise about benchmarks for time eternal, usually because their pet didn't win.
Meta knows their models aren't as good -- demonstrated by their benchmark performance -- and their value proposition is a much lower price.
To me it feels like we’re past the dreamy early days of the AI boom (in the US) and now investors want to see the tech monopolies actually start building the new cash printing machines they’ve been promising with the hundreds of billions in capex burned over the last few years.
Maybe a Gemma sized model that couldn’t self-cannibalize but releasing a top open Kimi K3 sized model seems like it could start to wear on Meta investors’ patience without signaling how it fits into a new profitable business line.
Cherry picking the benchmarks you present is where the falsehoods lie.
Hah, someone has https://www.felonybench.com up and running now.
The paid official dsv4 api already shares data with deepseek for training
You're eligible for free daily usage on traffic shared with OpenAI.
Up to 250 thousand tokens per day across gpt-5.4, gpt-5.2, gpt-5.1, gpt-5.1-codex, gpt-5, gpt-5-codex, gpt-5-chat-latest, gpt-4.1, gpt-4o, o1, and o3
Up to 2.5 million tokens per day across gpt-5.4-mini, gpt-5.4-nano, gpt-5.1-codex-mini, gpt-5-mini, gpt-5-nano, gpt-4.1-mini, gpt-4.1-nano, gpt-4o-mini, o1-mini, o3-mini, o4-mini, and codex-mini-latest.
Usage beyond these limits, as well as usage for other models, will be billed at standard rates. Some limitations apply. Learn more.Claude Sonnet is a weird exception to the mid models because Anthropic doesn't do much with Haiku and Opus is too big.
> If there were, do you believe it would be in their interest to answer this publicly?
If it were being adopted like gangbusters in their organization, sure!
So... the fact that nobody is volunteering the information is probably a valid signal of how things are actually going...
Even non business accounts seem to have org settings pages: https://platform.openai.com/settings/organization/data-contr...
But even with all data sharing enabled, I’m not seeing the free tokens message there.
Do others see a free tokens message at https://platform.openai.com/settings/organization/data-contr... after enabling all sharing there?
Let's not get ahead of ourselves here: this is Meta.
Mark Zuckerberg holds majority control of the company, and so far, he does what he wants with it. This is up to and including what was essentially a digital real estate scam, overhiring during the pandemic, automated CSAM distribution, the Cambridge Analytica scandal, and engagement tactics that might have fed into ethnic violence against Rohingya in Myanmar.
If he simply wanted to release a frontier-class open-weight model and offer a paid hosting service for those who wanted to run it off-prem, that would be, far and away, the least patience-wearing thing he's done to the bag-holders who hold the rest of the shares in Meta.
I doubt it; when/where was it? In the EU there’s no such obligation.
Doesn't matter if it's small/fast/cheap. Best = best.
Opus 5 is incredible at making games. Almost like a generation better than other models from my experience. You won't see that if you just look at the popular benchmarks..
You have to test each model on your actual use case to see how well it really performs.
Sonnet on the other hand never is, it's far from the best pick anywhere on the spectrum.
[1] https://platform.openai.com/settings/organization/data-contr...
you can pay whatever they want, they will train and use your data, I figured this is not something that should be discussed but obviously I have been mistaken...
The only difference between students doing it then and professionals doing it now is the students had no positive, glaring reason to mistrust.
all the while they throw their bank account and whatever they want at GPT or Claude models.
Do you have any evidence that cloud providers that don’t train their own models are violating their own contracts and risking their own good reputations by stealing data?
Generally though, many of these fairness complaints do go away if there is "3rd party testing". Right now, companies reporting their own benchmarks has all the problems that 3rd party testing resolves in many other industries.
Benchmarks are one data point, not the only one, but the easiest one to compare.
This is a bit vague. What sort of games with what technology?
If you are being serious, then you have a wildly distorted view of the world. No organisation can routinely break all of their contractual obligations. If you think Meta are doing this then you are not seeing Meta, you are seeing a fictional bogeyman.
The fact that Facebook has so much experience taking advantage of people's private data is one of the reasons I believe them when they say they won't be doing it when you pay them for that service.
You really think a throwaway quote when Zuckerberg was a college student applies nowadays?
You really think Meta would risk getting a massive (losing) lawsuit on their hands, in exchange for what? Middling amounts of training data?
I doubt the OP meant something like creating the whole tech stack for WOW.
What does "nowadays" mean? What changed?
> You really think Meta would risk getting a massive (losing) lawsuit on their hands, in exchange for what?
And what is the risk? Food companies are at risk when they put addictive chemicals into food because they get controlled regularly. What is the equivalent here?
[0] https://www.techradar.com/computing/cyber-security/facebooks...
[1] https://www.tomshardware.com/tech-industry/artificial-intell...
To me, their history suggests that they know more than most about how to get away with breaking both the spirit and the letter of the rules, and that they are motivated to enrich themselves without regard for what rules are broken.
I would not know which to expect, spirit or letter, in any given instance.
However, even if they were to surprise me by being perfectly meticulous about the letter of the rules from now on, I have so little trust in them that I would expect some technicality somewhere in the language of the contract.
I was just asking what kinds of games and with which technology.
Neither is stated in the original comment, and the answer obviously isn’t “every kind with every technology”.
We're excited to release Muse Code (beta), a terminal coding agent powered by Muse Spark 1.2, our newest model. This marks our next step toward the frontier, with larger and much more capable models on the way.
Muse Code takes on complex software engineering tasks across large repositories: planning changes, writing code, and validating the results. It can coordinate multiple persistent subagents for each task, solving difficult problems faster, more accurately, and with less intervention.
Muse Code operates with a simple agent loop plus a set of async background agents to enhance the main agent's capability. These specialized background agents remain active throughout each session, rather than being spawned for individual tasks, helping avoid redundant information gathering. They carry out next steps and choose when to communicate back to the main agent. Their persistence reduces latency and the need for steering on difficult, multi-step tasks.
Loading demo
Muse Code uses a local event log in which every model call, tool run, approval, and edit is appended. This single source of truth makes the runtime replay-exact and restart-safe: after a crash, the agent can resume precisely where it stopped. That ability lets Muse Code take on long-running tasks without being derailed by failures.
Muse Code ships with several default skills. /plan turns a task into an approval-gated plan, /grill stress-tests that plan until it holds up, and /goal works toward successful completion of the specified objective.
The user inputs a fly-through video of a home into the terminal as an mp4 file. Muse Code interprets the video and produces a visually rich vacation home marketing and booking page.
Muse Spark 1.2 is a coding-focused update to Muse Spark 1.1, with improvements in code generation, complex debugging, codebase understanding, and end-to-end developer workflows. In Muse Spark 1.2, we significantly scaled up training compute on coding tasks while expanding training environment diversity. The model also maintains its strength in other key areas like general agents.
For more details about our evaluations, see our report.
We co-trained Muse Spark 1.2 with Muse Code to ensure the model exhibits its best performance and coding usability when paired together. The training included rejection sampled harness trajectories and recipe optimizations for goals, compaction, and subagents, alongside the integration of the Muse Code toolset to maximize harness compatibility.
Muse Spark 1.2 was extensively trained on long-horizon coding tasks, including whole-repository generation, large end-to-end projects, and auto-research. It leverages planning to sequence work, goal conditioning to maintain direction, and context compaction to retain the knowledge needed to sustain progress.
We also used Muse Spark 1.1 to generate challenging coding environments and instruction-following templates. The model then graded candidate solutions on how well they satisfied those requirements, producing a scalable training dataset for Muse Spark 1.2. This self-improvement loop helped Muse Spark 1.2 follow complex instructions more precisely than its predecessor.
We tested the model's ability to iteratively optimize GPU kernels over 1,000+ tool calls (up to 24 hours). Leveraging Muse Code's agentic coding environment, the model writes, compiles, profiles, and progressively improves kernel performance relative to a provided baseline implementation. We benchmarked on KDA and MLA kernels for NVIDIA Hopper GPUs. The agent continues to achieve substantial improvements over the provided baseline implementation.

The baseline is the FLA Triton implementation of KDA. Models were prohibited from importing third-party kernel libraries such as FLA directly; instead, they had to apply specialized kernel-optimization knowledge to implement the algorithm in Triton, rather than wrap existing implementations. Muse Spark 1.2 paired a chunk-parallel preparation kernel with a sequential inter-chunk scan, combining standard fusion and tiling with KDA-specific optimizations such as re-centering the gated cumulative decay at the chunk midpoint.
Muse Spark 1.2 is available today in Muse Code and in Meta Model API with expanded global access. We have a lot on the horizon, including new harness features and more powerful models. We can’t wait to see what you build!
Honestly, I wish the tech community would do a better job identifying the actual individuals who are making these decisions instead of associating them with the brand they're under at the moment, because it's not THAT many people. Like, if you look at only the 100 tech sector companies included in the $NDXT index, how many individuals hold a VP or above title there, and how difficult would it be to trace key decisions at different times made at different companies to the individuals holding those positions there plus board membership and major shareholder identities (with the caveat of known unknowns here) and make a sort of ethical index and trace that along their careers with company moves, promotions, board appointments, shareholder decisions, etc? Go a step further and link that to financial performance and I'm sure folks at quant firms are already ten steps ahead of where I'm going with this, but I care less about profiting off of this data and more about surfacing it to show that it's people and more specifically, specific individuals driving these decisions.
Let's have two pricing levels. One will be 12.5x more expensive than the other. For the cheaper one, we will get their express permission to train models on their inputs. For the more expensive one, we will still treat their data EXACTLY the same, but we'll lie to them and say that we won't. I've checked in with legal and they raised their sherry glasses and toasted "Gentlemen, TO CRIME".
Easy to comprehend: Trust lost is hard to gain.
Though, it is extremely competitive of Meta to sell Muse Spark (Grok 4.5 / Qwen 3.8 Max level model) cheaper than DeepSeek v4 Flash / MiMo v2.5 / GPT 5.6 Luna, regardless.
There's contractual cover, there are lawyers all over the USA ready to make themselves very rich by creating a class action over it... you don't have to worry! Or, you do, but frankly, only MI6/CIA/Mossad can save you now.
We're also beginning to accept requests for zero data retention. Contact Meta sales to request this.
https://developer.meta.com/ai/resources/blog/build-with-muse...We can make this fun, within 18 months from now, there will be some story / whistleblower like "a inadvertent defect was found that allowed your 'private' data to be used in our training endeavours, we sincerely apologize and have already addressed the issue" - if this does not happened in this timeframe I will donate $1k to a charity of your choice.
Retaining data is common for investigating abuse.
People whistleblow on Meta all the time. Most recently they got fined half a billion for suppressing child safety research. Kids getting groomed. It doesn't make a difference to a company that has a net profit of $60 billion a year.