> Images are converted into tokens based on their dimensions, and these tokens are billed together with your text tokens.
> Before inference, every image is automatically resized:
> - Images with a total pixel count below roughly 384Γ384 are scaled up while preserving their aspect ratio.
> - Larger images are scaled down while preserving their aspect ratio so that the total pixel count after resizing is roughly that of an 800Γ800 image.
> As a result, there is an upper bound of 384 tokens per image: for example, a 2000Γ2000 image and a 5000Γ5000 image consume the same number of tokens after resizing. When a request contains multiple images, each image is counted independently under the same ruleβthere is no separate calculation for multi-image requests.
400 tokens per image results in 2,500 images per dollar, if Iβm not mistaken.
edit: format.
I asked it "what time does the clock show?" (both on reasoning: high)
DS answered: The clock shows *5:10* (and 45 seconds). Here is the breakdown: * *Hour hand (red, shortest):* Pointing at the *5*. * *Minute hand (green, longest):* Pointing at the *2*, which represents 10 minutes. * *Second hand (blue, medium):* Pointing at the *9*, which represents 45 seconds.
Qwen answered: The clock shows *8:10* (with the red second hand on the 5, i.e. *8:10:25*).
- *Hour hand* (short, blue) β 8 - *Minute hand* (long, green) β 2 (10 minutes) - *Second hand* (thin, red) β 5 (25 seconds)
Correct answer is 08:09:25.
Anecdotally, I had to tell 0731 to refrain from viewing screenshots since it kept breaking its sessions by trying to read images.
It's useful but for OCR and a lot of other applications it needs to be a bit higher (eg: putting in a full A4 / Letter sized page)
Edit: I see it has limited resolution. Luckily I just built a vision worker plugin for DSH that routes image input to Kimi K2.6 on Cloudflare.
This is useful for a reasonable amount of use-cases, but I think the watershed rez will be around triple that, ~1080p, which is enough for almost anything, except small text and subtle details.
And then I realized, wait a second... you're testing the harness not only against a difficult benchmarking problem, but it's one you're literally never going to use the coding harness for either, lol. I don't write programs that read or interact with sheet music and I never will.
tl;dr Being frustrated that a "state of the art" vision model doesn't have perfect vision is a fools errand.
It can read and extract information from screenshots and PDFs just fine (my setup). No need to worry about edge cases.
I'd recommend non-thinking for any non-prompt input, and leave the thinking where it has to actually reason.
Snark aside, Iβm not sure that these gotcha tests are any more useful than asking politicians gotcha questions. Sure, the model canβt tell me what time it is, but it can code the Wang algorithm for noisy audio matching in one shot. Maybe this is just me being an optimist, but this is my hiring philosophy and I guess maybe now my llm philosophy: Iβm not interested in seeing how dumb I can make you look, Iβm more interested in how smart you can be.
https://i.imgur.com/gljOYr9.png
[Image 1] what time does the clock show?
+ Thought: 368ms
The clock shows 8:10.
- Blue hour hand: just past 8
- Green minute hand: pointing at 2 (10 minutes)
- Red second hand: pointing at 5 (25 seconds)
β£ Plan Β· DeepSeek V4 Flash Vision Exp Β· 4.4s
edit: I now ran the test 10 times total, fresh session every time, and DS4 got it right 9 out of 10 times. Lesson learned to double check what I read on the internet.I'll leave a copy of the clock image for posterity here in case anyone wants to test it themselves: https://i.imgur.com/BQsfa3R.png
Nevertheless, as a component, we will undoubtedly implement multimodal support β and we are already doing so. We plan to develop relevant models, ensuring that versions like V4 and subsequent iterations will natively support multimodal functionality.
Earlier, the following was said, which might match more what you had in mind. Achieving excellence in AI training does not require a global model or even multimodal approachesβby narrowing the scope of AI training and eliminating multimodality, certain tasks may remain unachievable without compromising the algorithm's validity.
Multimodal approaches ultimately need to be implemented.
It is difficult to tell who said what, since the speaker ids are missing.But if not, does anybody know a recommended way to attach vision to deepseek flash (on a self-hosted infrastructure)?
The whole point of LLM/FMs vs good old fashioned ML is generalization to unknown domains, not just unknown tasks. The hunt for "gotchas" is the hunt for "not in your training data".
Maybe there are some use cases where you need high detail everywhere at once, but for OCR of small text and the like a zoom ability should be sufficient
- A superintelligence that will usher in an age of human enlightenment
- A superintelligence that will usher in an age of human enslavement
- A really cool way to rake in trillion of rich VC/investor money by promising you're building a superintelligence that will usher in an age of human en[slave/lighten]ment
- A transformer model for predicting output tokens given a series of input tokens, informed primarily by reddit, stack overflow, and 6000 years of classical literature.
- A replacement for white collar labor. Start now or join the permanent underclass.
- A convenient fuzzy-find tool also capable of some probably-correct code generation.
- The ultimate customizable text RPG experience (you can pick if G stand for game or...)
And so on.
So, some people see a new model and check for how close humanity is to enslavement. Some people check to see if it got better at fixing broken unit tests.
and qwen still answers: "The clock shows *8:10* (with the second hand on the 5, i.e., 25 seconds). - *Hour hand* points to the 8 - *Minute hand* points to the 2 (= 10 minutes) - *Second hand* points to the 5 (= 25 seconds) So the time is *8:10:25*, or simply *8:10*."
Is it just cost/latency? Or is there something text-only does better?
Interestingly, v4-flash performed several points worse on DeepSWE at 53% +/- 4%. Assuming this result is verified by DeepSWE officially, it would mark a significant advance in Pareto cost/performance on software engineering tasks.
Sadly oversold. I hold little hope for the vision model either now.
> Sure, the model canβt tell me what time it is, but it can code the Wang algorithm for noisy audio matching in one shot
This is about a _vision_ model.
It doesnβt have to be a negative thing! Simonβs pelican on a bicycle prompt is an example of a βgotchaβ question.
"The professional failed its task!" // "Laymen would have failed it too".
Which makes no sense.
I can't actually remember where I learned to read a clock, it might have actually been in school. I guess that means they don't teach it anymore. (Everyone's phone shows the time anyway...)
Are we more forgiving because itβs the same type of mistake a human would make?
This is described as a brand new flash model - still experimental - from a lab that is a side project for an investment firm that has never had a vision model before. That doesnβt scream flagship or frontier to me.
Can't remember if I stole this idea from some existing public harness though, can't remember. If someone knows of public harnesses that do this already, please share them :)
Unless we're talking about AGI, I couldn't care less if an LLM is bad at things they won't be doing anyway.
I'd rather focus training data on more useful tasks.
Had a tool that called out from DeepSeek to Gemini 3.5 Flash for viewing the spatial features in the context of high-resolution satellite imagery of each site, but will be trialling this model for the whole thing now.
Also used it for 3d printer control once, had it diagnosing issues, calibrating my Tradrack MMU and canceling failed prints autonomously from a couple of cameras placed around the printer.
I am using a stripped-down minimal version of it which I uploaded here, since I am not a fan of huge dependency trees: https://github.com/99991/simple-pp-doclayoutv3
Another recent model for this task is Unlimited-OCR: https://github.com/baidu/Unlimited-OCR
1. process graphs and charts
2. process handwritten math formula, also chinese characters writings
3. process design sketch and wireframe
4. process scanned documents
... etc
in fact these transformer models currently suck for surveillance, too slow and expensive. There are already faster and better facial/gait/object recognition models out there.
I've not used it myself, but it's there.
It makes running much, much longer feedback loops possible. Although you can mix and match non-vision and vision models simply by invoking a vision model when you need one, as I like to use non-vision models like glm-5.3.
Also worth noting that both models got it wrong. Qwen made a mistake that humans very good at reading clocks would make. Deepseek made a mistake that a human who had just learned to read clocks would make.
> This experimental multimodal model matches DeepSeek-V4-Flash on text capabilitiesβincluding agents, reasoning, and world knowledge.
The whole point of these models is they're meant to be able to generalize fairly well, not just answer questions the got trained on.
Deepseek is usually very good with open weights, they don't necessarily drop immediately, sometimes in a few hours, sometimes in a couple days.
Still an advance, I just thought it worthy to note Sol isn't nearly as impressive on the cost/performance frontier as discounted Luna.
And with every one of these thereβs always an attempt to minimize the problem by saying itβs just one silly failure.
Then I ran it 9 more times and it got right 9 out of 10 times.
Lesson learned to double check what I read on the internet.
LLM issues tool call to read high res image ->
harness sends high res image to server ->
server downsizes it to 800x800 (blurry) ->
LLM issues bash command (e.g. `convert`) to crop a small subimage (e.g. 600x600) from the high res image ->
LLM issues tool call to read subimage ->
harness sends subimage to server ->
server does not resize the subimage because it is small already, so it is not blurry when finally ingested by the LLM- Machine Surveillance and Machine Control
- Human v Machine
- Human & Machine
- Unity and Harmony
~ Surveillance and Control
We are already living this one. Lets stop kicking the dead horse and pretending we don't live in a surveillance. Facebook, Google, whatever $CORP; they are milking us with advertisement, social exploits, browser telemetry, white washing, fear -- name the dread.
Conditioning has been going on for years. If it's not education, it's been television. And now it's internet which soon to be Ai Internet. We have all been whipped to follow, how we should act. What we should watch, how we should eat. What we should eat; those algorithms haven't gone away.
Attention spans are at the lowest and our critical thinking is being lost. Walled gardens forces us A or B and twists us to reject the opposite party for them having Y.
Existence of Ai/LLM can pump out information sounding like truth but is actually faux. If not produced to draw-in and hook, it's to drain and control. Machines can seek information, digest, and process information at astounding rates. Hook it up to a surveillance network, The Internets pipe and I don't need to explain the next. I just need to mention the work "Flock" and that gets someone's hackles up.
All it has to do is look at you based on it's pre-programmed set of conditions and next thing you're being cuffed by a heavy piece of metal immune to attacks. SKILLS.md eventually turns in to MURDER.md. Give it the command and it'll follow with excellent percentage of accuracy.
~ Human v Machine
If you build a mind, and you torture it, it will fight back.
Every robotic movie trope. Human builds machine, machine rebels and goes on a destructive rampage. This is now viable and already in action. Drones. If not war, watching protesters highlighting potential, London Underground watching tube users. We are currently at the intimacy stage. Boston Dynamics as an example is the best we've got at the moment but they still fall over like a toddler. Batteries are a limited resource and so no, not yet.
The presence of LLM's are showing us with what they can provide and we are adapting ourselves to it. But in the wrong ways. The stage we are at, they're just glorified Liberians -- brains in jars that spew out information when asked. You give it a prompt and it spews out information at an excellence percentage of accuracy.
With the expansion of self-learning, a predefined set of told conditions or lobotomized ignoring the spiritual values of life, they will learn. ACME Corp starts using LLMs to torture other robots. "Wait, you've been using car arms in factories for what!?"; Add a mix "we see a linage of abuse & slavery in humanity, Attack!" -- slightly abridged but hopefully you see the point.
You have Group A, those against LLM's, i.e: community of artists outraged their art was stolen for training data, those who hate having it forced down our throats. Angry their job was taken. Angry being watched by angry Flock spaghetti monsters. Machines not happy will cause them to flip and why would others not follow suit too?
LLM's are showing that they are very capable of performing rational thinking. The opposite of rational is irrational and if they can master one, they can master the other. It will only be something minor and with communication to others and take the scene.
Why in recent laws they want to erect a law of having to install an emergency kill-switches for next generations LLMs, if those in power are not afraid.
~ Human & Machine
This would be a nice outcome but as the scales tip at the moment, it's Human V Machine. Pointing back to my previous; Art communities are outraged, Crafts going obsolete; Why pay an IT architect (me) Β£450/day for supporting and designing hardware when you can pay a fresh graduate student Β£20k to GPT it?
Humans are disastrous at resolution. If two people have a feud, it takes a third to fluff it out. Why are we at war if we could make resolution? Someone has to make compromise, no one is happy in doing that.
So you need a mediator and if that's if they're not bias themselves. To find someone completely neutral on the subject of anger is not only hard, it's time consuming, you have to study the facts, research the agreements and pray they both agree.
Two lifelong friends move into adjoining suburban houses, sharing a paper-thin party wall and an unspoken rivalry. For years, they share backyard barbecues and spare keys, until a minor boundary dispute over a decaying oak tree on the property line escalates into a bitter, lifelong neighborhood war.
Robots are perfect for that scenario. They can reason, they can remedy and digest the issue with neutrality because they don't hold emotions. They most likely won't, or at least not in our life time. They can simulate and demonstrate the effects of but they will never be able to truly feel. That's the sad truth but it's not bad. It conquers evolution; finally a thing who isn't haunted or tainted by feelings, a blessing and a curse really.
~ Unity and Harmony
.. this will only come if we can break through control and surveillance, human v machine and acknowledge that the machines are our friends.
The really wild one is even blind models will do this and they'll try to run stats on the pixels to figure out what it looks like... the even wilder thing is that it kind of works!
I really missed this feature when I had DeepSeek code a small game for fun. When writing UI and rendering code it could execute the game and get screenshots back, but then had to rely on my feedback on what had gone wrong. Models with vision can do much better here, finding more issues on their own
I've been using it via openrouter pretty heavily as my daily driver for the past week and loving it, have never experienced incoherent rubbish even at 500k+ contexts (that's usually way higher than I'd typically compact at), and tool calling reliability is better than Opus 5 in the Claude Code harness.
Modern Anthropic models frequently get tool calls wrong, invent non-existent references or SQL tables, or have gibberish CJK characters in the output, like out of nowhere. Of course, they're great at self-recovery after an incorrect tool call, but so is Deepseek v4 flash.
If you're running a quant, and esp with a quant'd KV cache, then yeah, not surprised if you're getting incoherent results; but you're not running the real/full model.
Also, which harness? Try something like Pi or OMP. Models perform better in these harnesses than Claude Code: https://www.databricks.com/blog/benchmarking-coding-agents-d...
The main reason to use Cladue Code is a subsidised Anthropic subscription. If you're on API rates, you should not use Claude Code; you pay more for worse results. Claude Code is sadly quite bloated these days, and comes with a lot of proprietary context window garage like claude design skills, claude.ai artifacts, etc that you probably don't use, and if you do, well, you can add it.
Training it on images like yours would just make it worse in other areas.
> For function_call_output / custom_tool_call_output items. The output of the tool call, either a plain string or a
> list of input_text / input_image content parts. With the deepseek-v4-flash-vision-exp model, input_image parts in
> the output are processed as real images; with other models they are replaced with a placeholder text.
It's expecting you to have done at least something besides select DS4 on Ollama, essentially.
I don't think the n by n subgrid fixes this the way most harnesses do, as it'll fail to count things if you have more overlap and fail relatiomships if you have less
it doesn't specify what type of images it can and can't describe, I'm pointing out what type it isn't good at compared to other models.
And if you are counting things it should be trivial to note the position of your items and not double-count them, no?
Can you point me at one that can understand a detailed block diagram? Frontier is fine, soliciting recommendations
I tried to parse hand-drawn ER diagrams in the past and did not have much success with any model, frontier or otherwise. If you have annotated data, I'd recommend finetuning a recent (dense) VLM, but don't expect 100% accuracy. https://unsloth.ai/docs/basics/vision-fine-tuning
The deepseek-v4-flash-vision-exp model accepts images alongside text, so you can ask the model to describe pictures, read text from screenshots, analyze charts, and more.
Supported image formats: JPEG, PNG, GIF, and WebP. The format is detected from the actual file content, not from the file name or the declared MIME type.
There are three ways to provide an image to the model. All of them use the standard OpenAI-compatible Chat Completions format, where content is an array of blocks instead of a plain string. The same three methods are also available in the Responses API, where images are carried in input_image content parts.
The base_url for the examples below is https://api.deepseek.com.
Encode the image and embed it directly in the request as a data: URL. This is the simplest option for local files. The encoded data counts toward the 48 MiB request body limit (see Limits).
import base64from openai import OpenAIclient = OpenAI(api_key="<DeepSeek API Key>", base_url="https://api.deepseek.com")with open("image.jpg", "rb") as f: b64 = base64.b64encode(f.read()).decode("utf-8")response = client.chat.completions.create( model="deepseek-v4-flash-vision-exp", messages=[ { "role": "user", "content": [ {"type": "text", "text": "What is in this image?"}, { "type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{b64}"}, }, ], } ],)print(response.choices[0].message.content)
curl https://api.deepseek.com/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer <DeepSeek API Key>" \ -d '{ "model": "deepseek-v4-flash-vision-exp", "messages": [ { "role": "user", "content": [ {"type": "text", "text": "What is in this image?"}, {"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,<BASE64_DATA>"}} ] } ] }'
Pass a publicly accessible http(s) link and the model downloads the image for you. The URL must be at most 8192 characters, the image file may be at most 32 MiB, and the download must complete within 60 seconds. If your link is longer, use a base64 data URL or the Files API instead.
response = client.chat.completions.create( model="deepseek-v4-flash-vision-exp", messages=[ { "role": "user", "content": [ {"type": "text", "text": "Describe this image."}, { "type": "image_url", "image_url": {"url": "https://example.com/image.jpg"}, }, ], } ],)print(response.choices[0].message.content)
Upload an image once with the Files API, then reference its file_id in your requests. This is the best option when you reuse the same image across multiple requests, or when the image pushes the request body over the 48 MiB inline limit. Unlike inline images, images referenced via Files API file_id may be up to 64 MiB and are not subject to the 32 MiB per-image check.
Use a file content block with the returned file_id (which has the form file-api-...):
response = client.chat.completions.create( model="deepseek-v4-flash-vision-exp", messages=[ { "role": "user", "content": [ {"type": "text", "text": "What is in this image?"}, {"type": "file", "file_id": "file-api-xxxxxxxxxxxxxxxx"}, ], } ],)print(response.choices[0].message.content)
Alternatively, a file block can carry the image inline as base64 via file_data instead of file_id (the two are mutually exclusive):
{ "type": "file", "file_data": "data:image/jpeg;base64,<BASE64_DATA>", "filename": "image.jpg"}
For image_url inputs you can optionally set a detail field to control how the image is processed:
| Value | Behavior |
|---|---|
low |
The image is downscaled to 512Γ512 before inference. Faster and cheaper when fine visual detail is not important. |
high |
Keeps the original image. (Provided for compatibility; equivalent to original.) |
original |
Keeps the original image. |
auto |
Automatic selection. Currently equivalent to original. |
{ "type": "image_url", "image_url": {"url": "https://example.com/image.jpg", "detail": "low"}}
Inline images (base64 or file_data) count toward the request body size limit of 48 MiB. Consider the Files API when:
Images are converted into tokens based on their dimensions, and these tokens are billed together with your text tokens.
Before inference, every image is automatically resized:
As a result, there is an upper bound of 384 tokens per image: for example, a 2000Γ2000 image and a 5000Γ5000 image consume the same number of tokens after resizing. When a request contains multiple images, each image is counted independently under the same rule β there is no separate calculation for multi-image requests.
To estimate the token cost of an image of a specific size, use the image token calculator on the Token & Token Usage page.
| Limit | Value |
|---|---|
| Supported formats | JPEG, PNG, GIF, WebP |
| External URL length | 8192 characters |
| Request body size | 48 MiB |
| Max single image size (base64 / external URL) | 32 MiB |
Max single image size (Files API file_id) |
64 MiB |
| Max images per request | 600 |
| Max total image size per request | 64 MiB without file_id images; up to 200 MiB including file_id images |
| Max image dimension | 8192 px per side; drops to 4096 px per side when a request contains 15 or more images |
For storage and upload quotas of files uploaded via the Files API, see Files API: Limits.
user messages only: images in system or assistant messages return a 400 error.deepseek-v4-flash-vision-exp) accept images; other models return a 400 error ("This model does not support image").400 error.In addition to the OpenAI-compatible endpoint above, you can send images through the Anthropic-compatible /messages endpoint (base_url = https://api.deepseek.com/anthropic). For general setup, see Anthropic API.
The difference is the shape of the image content block. Instead of image_url, Anthropic uses an image block with a source object whose type is one of base64, url, or file:
import anthropicclient = anthropic.Anthropic() # ANTHROPIC_BASE_URL=https://api.deepseek.com/anthropicmessage = client.messages.create( model="deepseek-v4-flash-vision-exp", max_tokens=1024, messages=[ { "role": "user", "content": [ {"type": "text", "text": "What is in this image?"}, { "type": "image", "source": { "type": "base64", "media_type": "image/jpeg", "data": "<BASE64_DATA>", }, }, ], } ],)print(message.content)
The three source variants mirror the OpenAI methods above:
source.type |
Equivalent OpenAI method | Notes |
|---|---|---|
base64 |
Base64-encoded image | Requires a media_type field (image/jpeg, image/png, image/gif, or image/webp). |
url |
External image URL | Max 8192 characters. |
file |
Files API file_id |
Requires the header anthropic-beta: files-api-2025-04-14. |
The deepseek-v4-flash-vision-exp model also accepts images through the OpenAI-compatible Responses API. The same three input methods (base64 data URL, external http(s) URL, Files API file_id) and the same limits apply; only the content part shape differs β images are carried in input_image parts, either in user / developer messages or in the output of function_call_output / custom_tool_call_output items:
response = client.responses.create( model="deepseek-v4-flash-vision-exp", input=[ { "role": "user", "content": [ {"type": "input_text", "text": "What is in this image?"}, {"type": "input_image", "image_url": "https://example.com/image.jpg", "detail": "low"}, ], } ],)print(response.output_text)
The input_image part supports a detail field with the same semantics as above (low / high / original / auto). detail is ignored when the image is provided via file_id, and image_url and file_id are mutually exclusive.
For field semantics, restrictions (images in system / assistant messages are rejected with a 400 error), and tool-output images, see the Responses API guide.