Function calling lets a model request that your app run an external function — for example, changing the UI, querying a database, or calling another service — and then continue the conversation with the result.
Supported models#
Models with function calling support#
- Gemma 4 (E2B, E4B) — full support (native function-call tokens).
-
Gemma3n E4B — function calling on the downloadable E4B
.litertlmbuild; not on E2B or the MediaPipe.taskbuilds. - FunctionGemma 270M — Google's specialized function-calling model.
- DeepSeek R1 — function calling + thinking mode.
- Qwen models (0.5B, 0.6B, 1.5B) — full support.
- Phi-4 Mini — advanced reasoning with function calling.
Models without function calling support#
- Gemma 3 270M — text generation only.
- Gemma 3 1B — text generation only.
- SmolLM 135M — text generation only.
- LFM2.5 230M — text generation only.
- SmolLM3 3B — text generation with reasoning, no function calling.
- Phi-4 Mini Reasoning — reasoning model, no function calling.
- FastVLM 0.5B — vision model, no function calling.
- Qwen2-VL 2B — vision model, no function calling.
- SmolVLM2 500M — vision model, no function calling.
- LLaVA-OneVision 0.5B — vision model, no function calling.
If you pass tools with supportsFunctionCalls: false, the chat logs a warning
and does not inject them — the model still works normally for text generation.
Pass supportsFunctionCalls: true for models that support it.
Built-in AI (OS models)#
Built-in AI models — Gemini Nano (Android), Apple Foundation
Models (iOS/macOS), Phi Silica (Windows), and browser Prompt APIs (Web) — also
support function calling, but it is prompt-based: these OS models don't expose a usable
structured tool API, so core InferenceChat weaves the tool definitions into
the prompt and parses the calls back out of the model's text. You declare tools
the same way (see below). Gemini Nano handles single-turn tool calls; multi-turn
agent chaining is not supported on Web.
Declaring tools#
Describe each function as a Tool — a name, a description, and a
JSON-Schema parameters map — then pass the list to createChat
(or
openChat) together with supportsFunctionCalls: true. createSession
also
takes tools:, and the two passthrough families read them there, through the
runtime's own tool path: Gemma 4, and FunctionGemma on a .litertlm. For every
other model it is the chat that puts the tools into the prompt and parses the
calls back out, so use the chat API:
final tools = [
const Tool(
name: 'change_background_color',
description: 'Change the app background color.',
parameters: {
'type': 'object',
'properties': {
'color': {'type': 'string', 'description': 'A CSS color name.'},
},
'required': ['color'],
},
),
];
final chat = await model.createChat(
tools: tools,
supportsFunctionCalls: true,
toolChoice: ToolChoice.auto,
);
ToolChoice controls whether the model may call a tool:
ToolChoice.auto(default) — the model decides.ToolChoice.required— the model must respond with a function call.-
ToolChoice.none— the model must not call any tool. Where the SDK renders the declarations this takes them out of the prompt; on a.litertlmGemma 4 or FunctionGemma the runtime already holds them, so the model can still emit a call — and with parsing off it reaches the stream as raw text.
ToolChoice.required reaches the model only where the SDK writes the
declarations itself. Gemma 4, and FunctionGemma on a .litertlm, hand them to
the runtime instead, and the runtime's tool payload carries no tool_choice —
so for those two required behaves as auto, silently. A .task
FunctionGemma gets the old "not supported" warning, and the JSON-format
families (Qwen, DeepSeek, Phi-4 Mini) do get a "you must call a function"
instruction written into their prompt.
Who renders the declarations#
On a .litertlm, both Gemma 4 and FunctionGemma go through LiteRT-LM's own tool
path: the declarations travel to the runtime as structured data, the call comes
back parsed, and the result of a turn goes back as one role-tool message that
continues the same model turn. Since flutter_gemma 1.8.4 with
flutter_gemma_litertlm 1.7.1 that is true for FunctionGemma too — before them,
its tool results were sent as an ordinary user message, and the model answered
them by repeating the call it had just made. Both halves are needed: core decides
the wire format, the engine sends it.
Where a call comes back as text rather than structured tool_calls — the web
SDK, or a .litertlm exported without the FunctionGemma model type, whose
runtime opens no tool-call channel — flutter_edge_ai parses that text itself, so
your code still receives a FunctionCallResponse.
Two consequences worth knowing:
-
Nothing in your code changes.
createChat(tools: ..., supportsFunctionCalls: true)and the loop stay the same; the wire format is chosen frommodelTypeand the file type together. -
ToolChoice.nonecannot take the declarations back out, because the runtime holds them. It stops the SDK from suppressing tool-call text, which is why a call made undernonecan reach the bubble as raw markup.
On .task models through MediaPipe there is no native tool path, so
FunctionGemma keeps the text wire format the SDK renders itself.
FunctionGemma is also an action model, and that shows in the turn after the
result: google/mobile-actions, the corpus it is tuned on, does not contain a
single row where the assistant writes a sentence after a tool result, so it
often ends its turn at the call. Render the tool's own result in your UI rather
than waiting for prose, and reach for Gemma 4 when you want the model to talk
about what came back.
Handling function calls#
When the model wants to call a function, the response stream emits a
FunctionCallResponse with the function name and arguments. Execute it, then send
a Message.toolResponse(...) back to the model:
chat.generateChatResponseAsync().listen((response) {
if (response is TextResponse) {
// Regular text token
print('Text token: ${response.token}');
} else if (response is FunctionCallResponse) {
// Model wants to call a function
print('Function: ${response.name}');
print('Arguments: ${response.args}');
_handleFunctionCall(response);
}
});
A model can also request several calls at once — the stream then emits a
ParallelFunctionCallResponse carrying a calls list of FunctionCallResponses.
generateChatResponseWithTools (below) handles this internally; a manual
listener must handle it too:
} else if (response is ParallelFunctionCallResponse) {
for (final call in response.calls) {
_handleFunctionCall(call);
}
}
Send the function result back to the model so it can continue:
final toolMessage = Message.toolResponse(
toolName: 'change_background_color',
response: {'status': 'success', 'color': 'blue'},
);
await chat.addQueryChunk(toolMessage);
final followUp = await chat.generateChatResponse();
Or drive the whole loop in one call (recommended)#
InferenceChat.generateChatResponseWithTools runs that whole cycle for you — it
streams the reply, and whenever the model calls a tool it invokes your
onToolCall, feeds the result back as a Message.toolResponse, and continues
until the model produces a final call-free answer (bounded by maxToolTurns).
You only implement the tools; the parse → execute → feed-back loop is handled.
// Stage the user message first, as for generateChatResponseAsync.
await chat.addQueryChunk(
Message.text(text: 'Make the background blue', isUser: true),
);
final stream = chat.generateChatResponseWithTools(
onToolCall: (call) async {
// Run whatever tool the model asked for; return its result map.
return switch (call.name) {
'change_background_color' => {'status': 'success', 'color': call.args['color']},
_ => {'error': 'unknown tool ${call.name}'},
};
},
maxToolTurns: 8, // safety cap on tool round-trips
isCancelled: () => false, // optional: return true to stop (e.g. barge-in)
);
await for (final response in stream) {
if (response is TextResponse) print(response.token); // final-answer tokens
}
This is the same driver the voice loop uses: VoiceSession.fromChat(…, onToolCall:)
runs function calls inside a spoken turn through it (see
Speech → Tool calling in the voice loop).
Platform support#
Function calling is supported on Android, iOS, Web, and Desktop. For Gemma 4,
the native function-call tokens are routed through the LiteRT-LM SDK chat-template
path (use ModelType.gemma4), so Gemma 4 function calling needs a .litertlm
model: with MediaPipe .task (including the web -web.task
builds) and with ONNX
the tools do not reach a Gemma 4 model. See Capabilities.
With the Built-in AI engine (flutter_edge_ai_builtin_ai)
function calling is prompt-based rather than a native tool API — Gemini Nano
(Android), Apple Foundation Models (iOS/macOS), and Phi Silica (Windows) handle
single-turn tool calls; on Web multi-turn agent chaining is not supported. Tool
declarations are deliberately not handed to the OS runner as well, which would
run two competing tool loops for one turn. Apple's native tool calling stays
reachable through flutter_local_ai's own LocalAiSession API, outside
flutter_edge_ai's chat loop.
Function calling works on the web .litertlm path, with one upstream caveat:
constrained decoding is left off there (LiteRT-LM#2434
aborts the next turn after a tool call when it is on), so the call's JSON shape
is not grammar-enforced. See Troubleshooting.
See Models for the correct ModelType per
model family.
Writing this with a coding assistant? dart run skills@ get --all installs
flutter-edge-ai-function-calling, the skill that teaches it declaring tools, handling
FunctionCallResponse, and the built-in tool loop.
