It’s also so inefficient, when they release the full performance numbers it’s not going to be good.
One example, it takes about 3.6x more tokens to finish the same work as Gemini Flash 3.6.
Am I just using it on tasks that makes it go on forever vs these benchmarks that are short&sweet, or something like that? I've been throwing bunch of identical prompts at different models at the same time, and when comparing hy3 and K3 I've never once had K3 reason less than hy3, as just one anecdotal data point.
Similar price? Doesn't make sense. Maybe they meant power, capability or speed?
The specific agent is focused on getting precise and on point answers about a codebase.
The starting point was nowhere near. E.g. asked why was X implemented in a certain way it would give bogus answers when the real answer was that there was no reason at all.
The benchmark included more than 50 questions or different difficulty.
But when the agent was improved in its prompt and rooting it was impossible to have it perform worse than closed source sota.
Just to say that the quality of the harness is as important as agents intelligence.
The subsidized subscription model won't last, API pricing "feels" closer to a true sustainable business model.
https://artificialanalysis.ai/models/deepseek-v4-flash?intel...
The ban on these open models is coming within weeks, if not days. As usual, the excuse will be "national security".
Daily reminder that improving your samplers from the garbage default top_p/top_k to min_p or subsequent methods dramatically improves the performance of these models, and makes most quantities like measured "verbosity" and subsequent calculations of "intelligence per token" meaningless
Daily reminder that no one, including within academic AI research, AI engineers, normies, etc takes LLM sampling seriously enough.
Deepseek v4 Pro prices with Opus 5 perf would be freaking unbelievable!!
This is probably a dream.
I claim the CCP will wise up within 2 years, possibly much much sooner, and ban their own companies from open sourcing to prevent the Americans from acquiring the capabilities.
Despite all the nonsense claims of China distilling US models, the reality is that the Americans absolutely do distill these free Chinese models, and distillation when full logprobs are available (i.e. you have access to the weights of the model) is an order of magnitude better than when you don't.
Yes, Chinese open weight models in the short term harm US closed source model providers bottom line. In the slightly longer term, "showing your hand" and publishing both the architecture innovations and the models weights will be too dangerous for the CCP to allow. This is triply true if they can release a model that beats the Americans on most benchmarks.
I've already warned investors that this is probably the closest open weight models will ever get to closed access.
It’s so cost effective I can offer a generous free tier since my goal isn’t to make money with it.
First, your direct comparison, Deepseek V4 Flash 0731 (max effort) $0.03 (rounded up) per task @ index 50.
OpenAI Luna:
* high effort $0.03 (rounded down) @ index 46
* xhigh effort $0.04 @ index 49
* max effort $0.07 @ index 51
So I would say a fair statement would be "OpenAI Luna between 2x and 3x the price of Deepseek Flash, what you get is 2 to 5 times faster inference"
The cheapest OpenAI model that beats it is OpenAI Luna (max effort) $0.07 @ index 51 (if you take the rounding out it summarizes to triple the price for similar performance), but still close to 3x faster.
And can SOMEONE please tell artificialanalysis that using dark blue for both Deepseek AND OpenAI is an especially unfortunate choice of colors, especially today?
People on HN downvote objectively correct information because they don't like it 24/7. There's a reason the creator of Zig left and gave the computer version of a middle finger on the way out to HN!
For simple tasks, they're already saturated, and you'd prefer the faster model, so that you can have a realtime/interactive-ish experience.
Or to put it bluntly, it's cheaper if you don't value your time. That goes for smaller models in general -- need more handholding, more correcting -- but the Chinese ones are slower on top of that.
As for speed, Sol on Low is faster than Luna on most settings.
Or a benchmark to benchmark benchmarks?
For example: no government contract to any company who uses even one vendor in it's entire chain of dependencies, who uses such open models.
They can extend this further by laying more conditions, such as: any company dealing in this-this field can only use models "officially" approved as "safe". Rest you can guess how easy it would be to get that "safe" rating for such open models.
One feature of the app is that all scripture is verified and what’s show to the user doesn’t come from the LLM at all and instead a trusted source.
I think exploring scripture this way does not alleviate you from struggling to learn and apply it. It hasn’t for me.
What's difficult and doesn't have to be with philosophy/ spirituality is to find relevant bits off situation, theme etc.
This app does that very well, LLMs are good at entity recognition.
It use the stock model, no new models requires.
Worth spend a few hours to try.
The DGX Spark requires a small hack to ignore the difference between sm120 vs sm121, but it does run on sm121.
If the full non-flash model follows up with the expected improvements, and at the price point they've been keeping, it puts the frontier labs in a tough position and it feels to me like like OpenAI is reaching deep into their pockets to try to head that off.
TFA link is a 404 though. I'm reading through https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731 instead
create a plan with SOTA, execute with this.
I'm not sure the outcome would be beneficial for the US as a whole here. But perhaps that is not their priority.
[0] https://petergpt.github.io/bullshit-benchmark/viewer/index.v...
Plus a size you can genuinely run at home: Unsloth lossless Q8 at 162GB.
So, are they planning to announce an optimized coding agent harness as well ? DSv4 flash is a fantastic model, and my daily driver. With reasonix or pi, I can code all day long and pay a few pennies for it. No token anxiety. Whereas the same model with fireworks/openrouter, with zdr thrown in, token costs ratchet up with no explanation. Likely that the model is subsidized for gathering usage data. I am waiting for the day I can run this locally.
Does the file hosting actually cost peanuts when you do it yourself and the cloud has shattered my understanding of what it actually costs to deliver so much data?
https://files.parasmittal.com/openai_aa_luna_dsflash.svg
1: https://openai.com/index/advancing-the-price-performance-fro...
maybe it is Fable level
High hopes for V4 Pro
At the scale of Huggingface, that still amounts to a lot load of money. Significantly less than if you did the same in AWS, but still a lot
That said, they do have a deal with AWS to make the data available in AWS ip space. Maybe they got some cheap hosting out of that too
"For the Code Agent tasks among the public benchmarks above, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness (to be released) as the agent framework, using the max reasoning effort level with temperature = 1.0, top_p = 0.95."
CDN costs are pennies compared to inference and training though, HuggingFace will just get another 100 million and be set.
But it makes little to no sense as long as API prices are what they are. Except for maybe privacy reasons.
I am using openrouter with zdr guardrail which routes to any provider that supposedly doesnt train on user data. I also use fireworks (directly, not via openrouter) which is a provider promising zdr and has a bunch of open weights models. My issue is that these zdr providers dont transparently disclose caching/tokens etc and so they end up being far more expensive than directly using DS.
owning a few GPUs is a lot cheaper than supercars.
I don't use it for local inference so much. I use it to learn.
I also use it as my daily driving Aarch64 development system.
Aside it's also very cool what else can be done with unified GPU memory, once you realize you have it...
Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR
Reasoning models are indicated by a lightbulb icon
Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.
Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR
Reasoning models are indicated by a lightbulb icon
Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.
Intelligence evaluations measured independently by Artificial Analysis · Higher is better
Agentic business operations
Reasoning models are indicated by a lightbulb icon
While model intelligence generally translates across use cases, specific evaluations may be more relevant for certain use cases.
AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.
Reasoning models are indicated by a lightbulb icon
AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.
Openness Index assesses model openness on a 0 to 100 normalized scale (higher is more open)
Reasoning models are indicated by a lightbulb icon
Artificial Analysis Intelligence Index · Weighted average cost (USD) per Artificial Analysis Intelligence Index task
Most attractive quadrant
Pareto line
Reasoning models are indicated by a lightbulb icon
Weighted average cost per Intelligence Index task. Each evaluation’s cost is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight.
Weighted average number of output tokens used to run one task in the Artificial Analysis Intelligence Index
Reasoning models are indicated by a lightbulb icon
The number of tokens required per Intelligence Index task. This is calculated by multiplying the output tokens per eval by the relative weights of each benchmark in the Intelligence Index, then dividing by task count (excluding repeats).
Weighted average cost (USD) per Artificial Analysis Intelligence Index task, segmented by token type. Lower is better
Reasoning models are indicated by a lightbulb icon
Weighted average cost per Intelligence Index task. Each evaluation’s cost is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight.
Cost (USD) to run all evaluations in the Artificial Analysis Intelligence Index
Reasoning models are indicated by a lightbulb icon
The cost to run the evaluations in the Artificial Analysis Intelligence Index, calculated using the model's input, cache hit, cache write, reasoning, and answer token prices and the number of tokens used across evaluations (excluding repeats).
Price (USD per M Tokens)
Reasoning models are indicated by a lightbulb icon
Price per token for cached prompts (previously processed), typically offering a significant discount compared to regular input price, represented as USD per million tokens. The values shown here are the cache hit price; cache write and cache storage are billed separately and vary by provider — see "Cache pricing by provider" for detail.
Context window: tokens limit · Higher is better
Reasoning models are indicated by a lightbulb icon
Larger context windows are relevant to RAG (Retrieval Augmented Generation) LLM workflows which typically involve reasoning and information retrieval of large amounts of data.
Comparison between total model parameters and parameters active during inference
Reasoning models are indicated by a lightbulb icon
The total number of trainable weights and biases in the model, expressed in billions. These parameters are learned during training and determine the model's ability to process and generate responses.