Flutter Gemma is now Flutter Edge AI. Same models, new name — 2.0 moves RAG into flutter_edge_ai_rag. Moving from flutter_gemma →

On-device LLMs in your Flutter app.

No servers. No cloud. Just Dart.

$ flutter pub add flutter_edge_ai flutter_edge_ai_litertlm

runs in your browser

6 platforms · multimodal · private · MIT
150/160 pub.dev points | DeepWiki — AI-indexed codebase

Everything you need for on-device AI

🧠

Multimodal

Vision + audio input with Gemma 4, Gemma3n, FastVLM

📞

Function Calling

Models call your Dart functions — structured tool use on-device

💭

Thinking Mode

See the reasoning chains of Gemma 4, DeepSeek R1 & Qwen3

🤖

Agent Skills

Give the model SKILL.md skills it runs itself — JS, native intents, MCP

🎤

Speech (STT + TTS)

Transcribe audio and synthesize speech on-device — selectable models, LiteRT C API, fully offline

🔍

On-device RAG

Independent RAG with pluggable qdrant-edge or sqlite-vec storage

⚡

Hardware Acceleration

CPU/GPU across native targets, plus supported Qualcomm Android and Intel Windows NPUs

📱

Built-in AI

Gemini Nano, Apple Foundation Models, Phi Silica & Chrome Prompt API — zero model download

🔌

Pluggable Runtimes

Choose inference, embedding and speech runtimes independently — register only what you need

📦

Opt-in Packages

Core + 11 opt-in packages — agent skills, speech, RAG & more; ship only what you use

🧑‍💻

Package Skills

Claude Code, Codex, Cursor & Copilot learn the flutter_edge_ai API from skills shipped in the package

See it in action

Real on-device features — recorded on a phone, not a server.

Thinking mode

Watch the model reason step by step before it answers — fully on-device.

▶ Try it live

Function calling

The model calls your Dart functions with structured arguments — no server in the loop.

▶ Try it live

Multimodal with Gemma 4

Send an image and chat about it — vision and audio input, running locally with Gemma 4.

Learn more →

On-device agent skills

Give the model a set of SKILL.md skills and watch it pick and run them through the tool-calling loop — text, JS, native intents, MCP.

Learn more →

Platform support matrix

Platform Vision Audio Embeddings NPU
Android ✅ ✅ ✅ ✅
iOS ✅ ✅ ✅ ❌
Web ✅ ❌ ✅ ❌
macOS ✅ ✅ ✅ ❌
Windows ✅ ✅ ✅ ✅
Linux ✅ ✅ ✅ ❌

NPU support covers supported Qualcomm Snapdragon devices on Android and Intel LunarLake/PantherLake on Windows. iOS GPU runs on Metal on device; the Simulator is CPU-only (256 MB Metal allocation cap).

5 minutes to on-device inference

Register your engines once, install a model, create a chat session — then generate. The same Dart API across all six platforms.

dart
await FlutterEdgeAi.initialize(
  inferenceEngines: [LiteRtLmEngine(), MediaPipeEngine()],
  embeddingBackends: [LiteRtEmbeddingBackend()],
  embeddingTokenizers: [GemmaEmbeddingTokenizers()],
);

await FlutterEdgeAi.installModel(
  modelType: ModelType.gemma4,
  fileType: ModelFileType.litertlm, // the declared type picks the engine
).fromNetwork('https://.../gemma-4-E2B-it.litertlm').install();

final model = await FlutterEdgeAi.getActiveModel(maxTokens: 2048);
final chat = await model.createChat();
await chat.addQueryChunk(Message.text(text: 'Hello!', isUser: true));
final response = await chat.generateChatResponse();

Need step-by-step setup? Read the full guide →

Supported models

All models run entirely on-device. Pick by capability, size, or platform support.

Gemma 4 E2B

Next-gen multimodal — text, image & audio

Gemma 4 E4B

Next-gen multimodal — higher capacity

Gemma3n E2B/E4B

Multimodal chat — image & audio; function calling on E4B .litertlm

FastVLM 0.5B

Fast vision-language on desktop

Phi-4 Mini

Reasoning & instruction following

DeepSeek R1

Reasoning & code generation

Qwen3 0.6B

Compact multilingual with thinking

Qwen 2.5

Multilingual chat

Gemma 3 1B

Balanced text — all platforms

Gemma 3 270M

LoRA fine-tuning base

FunctionGemma 270M

On-device function calling

SmolLM 135M

Ultra-compact for edge devices

Why on-device?

🔒

Privacy

Data never leaves the device

✈️

Offline

Works with no network connection

💸

Zero cost

No API bills, no rate limits

⚡

Low latency

No round-trip to a server

Ship AI that never leaves the device.

Open source, MIT licensed, maintained by the community. Star the repo and help spread on-device AI for Flutter.