TL;DR
On July 15, 2026, China's Cyberspace Administration approved the first-ever batch of on-device generative AI filings for smartphones. Buried in that approval: Apple's China build of Apple Intelligence runs on Alibaba's Qwen, not Apple's own Foundation Models (Bloomberg, CNBC). Honor shipped a Qwen-powered on-device agent on its Robot Phone the same month (Gate US). This page verifies which small AI models run on your specific iPhone, Pixel, Galaxy, or laptop in 2026, what license each one carries, and which apps (Ollama, LM Studio, PocketPal AI) let you run them without writing code. Six of nine flagship small-model families we checked ship under Apache 2.0, MIT, or comparably permissive licenses. Last verified: August 1, 2026.
China approved on-device AI for phones on July 15, and the approval notice revealed something we did not expect from Cupertino: Apple's own China build of Apple Intelligence runs on a Chinese open-weight model, not Apple's model.
That's not a rumor. The Cyberspace Administration of China approved seven on-device generative AI filings that day, covering Apple, Huawei, Xiaomi, OPPO, vivo, Samsung, and ZTE/Nubia (DIGITIMES). Six of those seven ship a domestic model as their own on-device backend. Apple's China build is the outlier: it runs Alibaba's Qwen, confirmed by Bloomberg, CNBC, and the filing notice itself (Bloomberg, CNBC, TechNode).
Three days later, Alibaba and Honor jointly launched a Qwen-powered on-device agent, live on Honor's "Robot Phone," announced at WAIC 2026 (Gate US). I read that headline three times before I believed it.
Qwen isn't a China-only story anymore, either. Apple's own WWDC26 Core AI session confirmed a curated catalog of third-party open-source models optimized for Apple Silicon, and Qwen sits on that list next to Mistral and SAM3 (Apple Developer). A Caltech spinout called PrismML also compressed all 27 billion parameters of a Qwen model from roughly 54GB down to under 4GB, small enough to run on an iPhone 15 (CNBC).
So the question we get asked most isn't "is local AI real yet." It's "what runs on the phone in my pocket, right now." This page answers that with device-specific tables, not generic RAM tiers.
What's Inside
Which iPhone, Pixel, or Galaxy model runs local AI, and what it can do offline. Whether "on-device" really means your data never leaves your phone, or just usually doesn't. Which small models (Qwen, Gemma, Phi-4, Mistral) are genuinely open source versus proprietary with a friendly name. What Ollama, LM Studio, and PocketPal AI need from your hardware, and whether any of them are truly no-code. Why China's July 2026 approval of on-device AI filings matters even if you've never opened a Chinese app.
Six Pages Rank for "Best Local AI Model," and Here's What They Skip
I read the top-ranking pages for "best local AI model 2026" before writing a word of this one. Six of them, cover to cover.
Every single one gives you a generic RAM tier. "8GB RAM," "16GB RAM," "iPhone 15 Pro or better." Useful in theory, useless when you're standing in a Verizon store holding your actual phone.
fast.io covers Qwen3, Gemma 4, DeepSeek R1, and Devstral with full VRAM tables, written for engineers (fast.io). PC Build Advisor and AI Intel Report cover similar model lists for developers and enterprise teams, with partial hardware detail (PC Build Advisor, AI Intel Report). PromptQuorum frames some picks for "most users" but still resolves to RAM tiers (PromptQuorum). Yaps.ai is the closest thing to a normal-person guide we found, covering a dozen offline apps (Yaps.ai). iOS App Lists does per-app RAM and storage specs for iPhone apps specifically (iOS App Lists).
None of the six cross-references a specific device model against a specific on-device AI capability in one table. Not one.
What this means in practice: if you've ever read one of these lists and still didn't know whether your exact phone qualifies, that's not a reading comprehension problem. The information you needed wasn't there. We built the table below to fix that.
What Your Exact Phone or Laptop Can Run
I pulled Apple's own compatible-device list, Samsung's Galaxy AI page, and Google's AICore documentation, then matched each against the model families in the matrix further down this page.
| Your Device | On-Device AI You Get | Runs Locally? |
|---|---|---|
| iPhone 15 Pro/Pro Max, iPhone 16 (all), iPhone 17 series, iPhone Air, 17e | Apple Intelligence core features, plus compressed Qwen (27B via PrismML) and Core AI catalog models | Yes, on-device-first (iClarified, CNBC) |
| iPhone 17 Pro/Pro Max/Air | Advanced Apple Intelligence tiers | Yes (iClarified) |
| Base iPhone 15/15 Plus (6GB RAM) | No Apple Intelligence, but small (<3B) models via PocketPal or Off Grid if storage allows | Partial (fone.tips) |
| Pixel 8/8a | Gemini Nano, manual enable via Developer Options | Yes, for supported ML Kit tasks (secondary sources: Android Police-type reporting) |
| Pixel 9 and newer | Gemini Nano, pre-enabled | Yes (secondary sources) |
| Samsung Galaxy S24 and newer | Call Screening, Now Nudge, Now Brief, Scam Detection on-device. Circle to Search and Creative Studio go to the cloud | Mixed (Samsung, PromptQuorum) |
| Any Android phone, 6GB+ RAM | PocketPal AI or Off Grid running Qwen3.5 (0.8B-2B), Phi, or Gemma 2 | Yes, fully local (PocketPal GitHub, Show HN) |
| Any Android phone, 4GB+ RAM, including $200-300 budget phones | Off Grid running Qwen3.5-0.8B, about 500MB | Yes, works in airplane mode (Show HN) |
| Mac (Apple Silicon), 16GB RAM | LM Studio or Ollama running 7B-8B models comfortably | Yes (LM Studio docs, Ollama docs) |
| Mac (Apple Silicon), 32GB RAM | 13B-14B models, plus Mistral Small (~24B) | Yes (Mistral) |
| Windows/Linux PC, 24GB+ VRAM | 27B-32B models (Gemma 3 27B, Qwen3-32B) | Yes (fast.io tier tables, Ollama docs) |
| Intel Mac | LM Studio doesn't support Intel Macs. Jan and Ollama still work | Depends on tool (LM Studio docs) |
| Honor "Robot Phone" | Qwen-powered on-device agent, factory-shipped | Yes (Gate US) |
What this means in practice: if you're carrying an iPhone 15 Pro or newer, a Pixel 9, or a Galaxy S24 or newer, some form of on-device AI is already running on your phone right now, whether you turned it on or not.
The Verified Small-Model Matrix
I cross-checked model cards, license pages, and vendor blogs for every small model shippable on consumer hardware in 2026. Skip the marketing pages. This is what's confirmed.
| Model | Params | Min RAM/VRAM | License | Runs On | What Stays Local |
|---|---|---|---|---|---|
| Qwen3.5-0.8B | 0.8B | <2GB VRAM | Apache 2.0 | Any modern smartphone, 4GB+ RAM | Fully local |
| Phi-4-mini-instruct | 3.8B | Modest consumer CPU/GPU | MIT | Laptops, higher-end phones | Fully local |
| Gemma 3n E2B | ~1.91B effective | Phones, tablets, laptops | Gemma license (permissive, custom) | Phones, tablets, laptops | Fully local |
| Qwen3.5-4B | 4B | Fits iPhone 15 Pro/16 via 4-bit GGUF | Apache 2.0 | iPhone, laptops | Fully local |
| Mistral Small (~24B) | ~24B | 32GB RAM Mac | Apache 2.0 | Mac, high-end PC | Fully local |
| Qwen3.6-27B | 27B | 18GB RAM | Apache 2.0 | Laptop/desktop | Fully local |
| Gemini Nano-1/Nano-2 | 1.8B / 3.25B | ~1GB storage, via AICore | Proprietary | Pixel 8+/9+, Galaxy S24+ | On-device for supported ML Kit tasks, broader Gemini features route to cloud |
| Apple Foundation Models | Undisclosed | 8GB+ RAM, Apple Intelligence-eligible device | Proprietary, OS-bundled | iPhone 15 Pro+/16/17, eligible iPad/Mac | On-device by default, overflow goes to Private Cloud Compute |
| Llama 3.2 1B/3B | 1.23B / 3.21B | Edge/ARM-class hardware | Llama 3.2 Community License (custom, not OSI-approved) | ARM/MediaTek/Qualcomm edge devices | Local for on-device deployment |
Sources: Alibaba Cloud blog, Hugging Face Phi-4-mini card, Gemma 3n docs, MindStudio, Mistral official, Wikipedia: Qwen, Apple Foundation Models docs, Apple PCC security blog.
I pulled the license field from every model card myself rather than trusting a summary blog. The smallest usable model in the set, Qwen3.5-0.8B, needs less than 2GB of VRAM and runs on basically any modern smartphone under Apache 2.0 (Alibaba Cloud). Phi-4-mini is Microsoft's smallest shipping model at 3.8B parameters, and it's MIT licensed, meaning you can use, modify, and redistribute it freely (Hugging Face). At the upper end of laptop-friendly, Mistral Small runs on a 32GB Mac at roughly 150 tokens per second, also under Apache 2.0 (Mistral).
What this means in practice: three model families you've probably heard of, Qwen, Phi, and Mistral, are all Apache 2.0 or MIT licensed. You can download them, run them, and never send a single prompt to a company server.
Who Owns Their Model? The License Count
Of the nine flagship small-model families verified above, six carry Apache 2.0, MIT, or comparably permissive licenses: the Qwen3.5 family, Phi-4-mini, Gemma 3 and 3n, Qwen3.6-27B, and Mistral Small. Three are proprietary or restricted: Apple's Foundation Models, Gemini Nano, and Llama 3.2, whose Community License isn't OSI-approved because of a clause requiring separate licensing above 700 million monthly active users.
That's two-thirds of the small on-device models we checked, shipping under genuinely open terms. Not a single one of the six competitor pages we audited computes that ratio.
What this means in practice: "open source" gets used loosely in AI marketing. Apache 2.0 and MIT mean you can inspect, modify, and redistribute the model. A Gemma-style permissive license is close but has extra terms. Proprietary means you're trusting the vendor's word, and nothing else.
Private Cloud Compute vs Gemini Nano vs Samsung's Split: What "On-Device" Really Means
"On-device" gets used as a single marketing word, and every major vendor means something slightly different by it.
Apple's version has a name and a public paper trail. Private Cloud Compute handles requests too big for the phone itself. Apple states the data is end-to-end encrypted, stays on PCC nodes only until the response returns, and is never available to Apple staff, even those with administrative access, and publishes production PCC software images for outside researchers to inspect within 90 days (Apple security blog).
Google's version is more fragmented. Gemini Nano runs fully on-device for supported ML Kit tasks through Android's AICore. But Google's broader Gemini assistant features, including Circle to Search, route to the cloud, and we couldn't find one official Google page that draws the same clear on-device/cloud line that Apple's PCC page draws for PCC.
Samsung splits it feature by feature. Call Screening, Now Nudge, Now Brief, and Scam Detection stay on-device. Creative Studio, Circle to Search, and Gemini agents go to the cloud. Photo Assist does both: local segmentation, then a cloud generation step (Samsung, PromptQuorum).
What this means in practice: if privacy is the whole reason you want local AI, "on-device" isn't a yes-or-no checkbox. Check the specific feature. Apple's Private Cloud Compute pairs a real technical commitment with public documentation. Samsung tells you which of four features stay local. Google hasn't published an equivalent breakdown for Gemini Nano yet, so treat any specific claim about what leaves the device as unverified beyond what ML Kit states per API.
Running Local AI Without Writing a Line of Code
You don't need a terminal to run a model on your own machine anymore. Four tools get you there, each with a real hardware floor.
Ollama grew from about 100,000 monthly downloads in Q1 2023 to 52 million in Q1 2026, a roughly 520x jump, while staying fully MIT licensed (dev.to, GitHub). It needs 8GB RAM for 7B models, 16GB for 13B, 32GB for 33B (Ollama docs). Here's the honest nuance most guides blur: Ollama is CLI-first, not no-code by default, though desktop and mobile clients now exist.
LM Studio is the one we'd point a genuinely non-technical friend toward. Full GUI, 16GB RAM recommended, runs on Apple Silicon, not Intel Macs (LM Studio).
Jan says it outright on its own homepage: "Do I need coding skills? No. You point and click, then start chatting" (Jan.ai). It's AGPL-3.0, needs 8GB RAM for 3B models, and runs entirely on CPU if you don't have a GPU.
PocketPal AI is the phone version. Open source, iOS and Android, needs 6GB RAM for smaller models and 8GB for what its own Play Store listing calls "the good stuff." It ships with Qwen, Phi, and Gemma 2 pre-configured (GitHub).
What this means in practice: if you own a Mac with 16GB of RAM or an Android phone with 6GB, you already have enough hardware to run a genuinely capable model this weekend, no coding required. And if running a model locally leaves you wanting more than a chat window, the next rung on the ladder is wiring one into a task that runs without you, which is exactly what my first-AI-agent walkthrough covers.
FAQ
What's the best local AI model to run in 2026?
There's no single best model. It depends on your device. Qwen3.5-0.8B runs on almost any phone under Apache 2.0. Mistral Small needs a 32GB Mac. Start with whatever your hardware supports in the matrix above rather than chasing a leaderboard number.
Can I run AI offline on my phone without an internet connection?
Yes, on several paths. Apps like Off Grid and PocketPal AI run fully local models like Qwen3.5-0.8B in airplane mode on Android phones with as little as 4GB of RAM (Show HN). On iPhone, Apple Intelligence itself is on-device-first, and only overflows to Apple's Private Cloud Compute when your phone can't handle the request.
Is local AI without coding possible, or is that marketing?
It's real, and Jan's own team says it plainly: point and click, then start chatting (Jan.ai). LM Studio and PocketPal AI work the same way. Ollama is the one exception on this list. It's genuinely powerful and genuinely CLI-first, so if a terminal makes you nervous, start with Jan or LM Studio instead.
Does Apple Intelligence in China really run on a Chinese AI model?
Yes, confirmed by Bloomberg, CNBC, and the Cyberspace Administration of China's own July 15, 2026 filing notice. Apple's China build runs on Alibaba's Qwen rather than Apple's own Foundation Models (Bloomberg, CNBC). Neither company has disclosed which specific Qwen checkpoint powers it.
Does "on-device AI" mean my data never leaves my phone?
Not automatically, and this is the question we'd want answered before trusting any vendor's marketing. Apple publishes a specific, auditable answer through Private Cloud Compute. Samsung breaks it down feature by feature on its own Galaxy AI page. Google's Gemini Nano runs on-device for supported ML Kit tasks, but we found no single official Google page mapping every feature's data boundary the way Apple and Samsung do. Check the specific feature, not the marketing headline.
We built this page because six other guides made us hunt for an answer that should've taken ten seconds: does my phone do this. If you've tested one of these apps on your own device and it did something the tables above didn't predict, tell us. That's exactly the mismatch this page exists to close.
And if your phone turns out to be running a Chinese open-weight model without you ever noticing, welcome to 2026. We'd love to hear about it 🤗