We need more like this as well as llama.app, which also has a native mac app.
Look it up, it’s a bit difficult to explain concisely in words but it is intuitive visually.
If we are thinking of intelligence and price, a model will be in the Pareto Frontier if there’s no cheaper model of the same or higher intelligence. Or if there’s no more intelligent model for that price or lower.
EDIT: See this chart from Artificial Analysis: https://artificialanalysis.ai/#intelligence-comparison-tabs
So for example DeepSeek V4 Pro can be considered a frontier model because there's no cheaper model that is as intelligent.
For any solution in the Pareto Frontier, there no "no-brainer" alternative, in the sense that there's no other option that is better in some way without giving up something else. It's the best of its "weight-class".
But (off on a tangent) I do think that there are multiple frontiers generally — and I also think the open weights, small local model frontier is by far the most important and exciting one.
I keep mucking about with what Gemma 4 12B can do and every time I do I find myself thinking that all the energies in the AI world are going in entirely the wrong direction, because it is small, clever, efficient and remarkable.
If all of that research money were to be spent on improving AI models that fit inside a 16GB RAM machine with unified memory, I think really important progress could be made.
I enjoy using the Qwen 3.6 models (and BottleCap's new fine tune of the 27B) but the small Gemma 4 models are impressive in a way that I think is going quite unreported.
So while I don't think this website should use the word "frontier" here, even referring to Qwen 3.6 27B which is weirdly close, I think it could.
Also, the title is "Run AI models locally on your Mac," not "Run frontier open models locally on your Mac."
I think it's likely the 26B QAT model won't fit in your machine — you may be able to fit one of the UD_Q3 or UD_Q2 variants but whether you'll be able to run other things you want at the same time, I don't know.
(The QAT models are "quantization aware training" — AIUI the model weights have been assigned during training to survive four-bit quantization with less loss.)
So what I would recommend trying is this model:
https://huggingface.co/unsloth/gemma-4-12B-it-qat-GGUF
Try UD Q4_K_M maybe.
(I don't think the M1 really gets much benefit from MLX, in case you were wondering, though I could be wrong)
My interest in this model is largely to really get to grips with what small models can actually do, especially with tool calling, because I think it helps comprehend what the value proposition of the cloud models is.
I have been very surprised by the quality and clarity of its answers. It's also helped me understand that much of a typical harness system prompt is likely to be unnecessary now; Gemma 4 seems to be pretty sensible out of the box.
You're absolutely not going to be able to get it to go off and build whole apps from a long prompt; it is not that good, but it does tool calling and thinking, and you should be able to explore pointing a coding harness at it if you turn on LM Studio or Unsloth Studio's API server. You could also use the Llama system tray app (formerly LlamaBarn) or just use llama-server from the llama.cpp distribution.
Probably Pi is going to be a better harness because it can have a minimal system prompt, though I've not tested it with Pi myself.
It seems to know PHP and SQL to a fairly decent depth (and I suspect JS and Python). It also has a unified vision model (it doesn't need a separate mmproj sidecar thingy) that is fairly fast, and it is quite impressive at image analysis.
So you could probably use it to generate image descriptions and tags, summarise text, generate wordpress snippets, that sort of thing.
It can capably answer questions like "Can you characterise this image and suggest further similar images I might like?" — I am currently using this to provoke me to take photos again.
Have a play with the E4B edge model, too — again, much more interesting than I expected.
brew install llama.cpp
llama-server -hf unsloth/gemma-4-12b-it-GGUF:UD-Q8_K_XLDon’t expect much for coding. But it’s great for general knowledge, rubber ducking, image classification…
The maintainer of mlx-vlm being behind this as well is the main thing kicking me over into trying it, even if it is incredibly young. I'm confused and unenthused to see it chomping on a whole GB of disk, but the Swift makes it feel much more refined even if it's not yet as feature rich. It automatically picked up the existing models I was using with mlx_vm directly, which was nifty.
And I find it curiously interesting when talking about photography. I've been finding it intriguing to ask it about my own photos and make suggestions about other images to research. I just showed it three of my own photos, and asked it to analyse them and recommend photographers I should research. It recommended someone amazing I have never heard of before. But it also recommended a 19th century British photographer who happens to be my lifelong photographic hero — someone whose broad characteristics inform what I do without me slavishly copying them. Bit of a jaw-dropping moment for it to have picked up their influence in subject matter that they would never have approached.
I'm still suggesting it more to people for them to see what the small-model future might look like, because it's so much more capable than one might expect.
Does it? I read: "Run AI models locally on your Mac."
Latest release · 100% open source · live now
Nativ puts frontier intelligence on your desk. Download and run open models on Apple Silicon — no accounts, no subscriptions, no cloud.
Universal · Apple Silicon (M1+)
Nativ, running on macOSReal app. Real local models. No cloud.

01 / CHAT
Talk to open models with streaming responses and per-message performance metrics.
CAPTURED LIVE ON THIS MAC · JUL 2026
/ SIX REASONS TO GO LOCAL
01CURATED LIBRARY
Run standout open models from Google, Cohere, and Liquid AI. Nativ recommends the right partner model for your hardware.
02CHAT
A clean interface with streaming, markdown, code highlighting, and image input. Every response is generated locally.
03TELEMETRY
Live tokens/sec, memory pressure, thermal state, and time-to-first-token. The details developers want.
04APPLE SILICON
Built on MLX-VLM and tuned for M-series unified memory and Metal — no wrappers, no translation layers.
05EVERY MODALITY
Chat with an LLM, caption an image, summarize video, autocomplete code, or transcribe and generate speech.
06FREE FOREVER
Not a subscription.
No accounts to create, no credits to buy, no data to sell. You own it end-to-end.
MITLICENSED
/ INTEGRATIONS
Connect the coding agents you already use to models running locally on your Mac. Nativ handles the endpoint; your workflow stays familiar.
01Pi 02Codex 03Claude Code 04Hermes 05OpenCode

LOCAL ENDPOINT One model server. Every coding agent.
MANIFESTO · REV 1.0
philosophy.txtUTF-8
$ cat philosophy.txt
The other “local AI” apps you’ve heard of? They’re proprietary shells built on top of open-source engines they don’t own. They keep the UI closed, add a paywall, and hope you don’t look under the hood.
We built in the open. The desktop app is open too. Every line. Every model loader. Every telemetry chart. You can read it, fork it, or send a pull request tonight.
No VC roadmap. No enterprise tier. No dark pattern that turns your prompts into training data. Just software made by researchers and hackers, for researchers and hackers.
MODELCREATORCONTEXTSIZETYPE
Explore the full model library →
YOUR MAC IS MORE CAPABLE THAN YOU THINK
Run it locally, in the open.