Every Flutter team we work with in 2026 has the same two tools open: an editor with an AI coding agent in the side panel, and a simulator that they still alt-tab to a hundred times a day. Both of those habits are now worth re-engineering. The Dart and Flutter toolchain has grown first-party support for agent workflows — the Dart MCP server — and a first-party way to render a widget without booting a full app flow — widget previews.
Used carelessly, an agent on a Flutter codebase is a very fast way to generate plausible, broken Dart. Used with the right tooling wired up, it closes the loop: the agent can read your analyzer errors, run your tests, hot reload the running app, and see whether its change actually compiled. This tutorial is about building that loop, and about the guardrails that keep it from quietly degrading your codebase.
Why a generic AI assistant underperforms on Flutter
A chat model with no tools is working blind. It cannot see that flutter analyze is throwing 14 errors, it does not know which packages are actually in your pubspec.lock, and it has no way to tell that the widget it just wrote overflows on a 360 dp screen. So it guesses — and the guesses skew toward the Flutter of two or three years ago, because that is what the public corpus is thickest with.
The practical symptoms are familiar to anyone who has reviewed agent output on a Flutter repo:
- Deprecated APIs:
WillPopScopeinstead ofPopScope,textScaleFactorinstead ofTextScaler,MaterialState*instead ofWidgetState*,RaisedButton,accentColor,Color.value. - Old state-management idioms — legacy
ChangeNotifierProviderwiring in a Riverpod 3 codebase, or hand-rolledStatefulWidgetcontrollers next to your existing notifiers. - Imaginary package APIs, invented constructor parameters, and imports of packages that are not in your
pubspec.yaml. - Logic that compiles but violates architecture: network calls inside
build(), repositories bypassed,BuildContextused across anawait.
None of these are fixed by a better prompt alone. They are fixed by giving the model feedback from your actual toolchain and by constraining it with rules checked by a machine, not by a reviewer's patience.
Part 1 — Wire up the Dart MCP server
MCP (Model Context Protocol) is the now-standard way an AI client talks to external tools. The Dart team ships an MCP server with the SDK, and it exposes the Dart/Flutter toolchain as callable tools: analysis, package management, test running, hot reload, runtime inspection of the widget tree, and DevTools data.
Check that your SDK is recent enough, then start it:
dart --version # need a reasonably current Dart 3.9+ SDK
dart mcp-server --help
Most clients are configured with a small JSON block. The shape varies slightly by client, but the server invocation is the same:
{
"mcpServers": {
"dart": {
"command": "dart",
"args": ["mcp-server"]
}
}
}
Some editors (recent VS Code Dart/Flutter extensions, and several agent CLIs) can register this for you — look for a "Dart MCP server" setting before hand-editing config. If your agent runs in a container, make sure the Flutter SDK and your PATH are present inside the container, not just on the host.
Once connected, the tools your agent gains are the interesting part:
| Capability | What the agent can now do |
|---|---|
| Static analysis | Read real analyzer diagnostics with file/line, instead of guessing at correctness |
| Pub tooling | Add, remove, and resolve dependencies; see what versions actually resolved |
| Run tests | Execute flutter test and read failures |
| Hot reload | Push its edit into the already-running app |
| Runtime inspection | Query the widget tree and app state of the running app |
| Docs/search over pub.dev | Look up real package APIs instead of inventing them |
That last pair is what turns the agent from an autocomplete into something that can self-correct. A loop of edit → analyze → fix → hot reload → inspect removes most of the deprecated-API and hallucinated-parameter class of error before a human ever sees the diff.
A grounding prompt that actually constrains
Give the agent a project file it reads on every session (AGENTS.md, CLAUDE.md, .github/copilot-instructions.md — whichever your client supports). Keep it short and specific; long prose gets ignored.
# Flutter agent rules — AviaryApps house style
- Target: Flutter 3.4x stable, Dart 3.9+. Material 3 only.
- State: Riverpod 3 code-gen notifiers. Never add `provider`, `get_it`,
`bloc`, or a new state library without asking.
- Navigation: go_router. All routes declared in `lib/routing/routes.dart`.
- Networking lives in `lib/data/`. Widgets never call Dio/http directly.
- Use `PopScope`, `WidgetState*`, `TextScaler`, `Color.withValues(...)`.
Never `WillPopScope`, `MaterialState*`, `textScaleFactor`, `withOpacity`.
- No new dependency without checking pub.dev for maintenance + license.
- After every change: run the analyzer and `flutter test`. Do not report
done while either fails.
- Every new widget gets a `@Preview` and every bug fix gets a test.
Two notes from doing this on client codebases. First, rules the toolchain can enforce beat rules written in English: move as many as you can into analysis_options.yaml so the agent's own analyze step catches them.
include: package:flutter_lints/flutter.yaml
linter:
rules:
- deprecated_member_use_from_same_package
- use_build_context_synchronously
- avoid_print
- prefer_const_constructors
- require_trailing_commas
analyzer:
errors:
deprecated_member_use: error # hard-fail old APIs the agent likes
unused_import: error
Second, pin the context. If your agent can read a pubspec.lock and the analyzer output, it stops arguing with reality about which package version you are on.
Part 2 — Widget previews: a render loop measured in seconds
Flutter's widget preview feature lets you annotate a function returning a widget and render it in a preview surface without launching the app and navigating to that screen. For a design-system package or a deeply nested screen behind three logins, this is the difference between a two-second iteration and a two-minute one.
import 'package:flutter/material.dart';
import 'package:flutter/widget_previews.dart';
import 'package:myapp/ui/order_card.dart';
import 'package:myapp/models/order.dart';
@Preview(name: 'OrderCard — pending')
Widget orderCardPending() => _wrap(
OrderCard(order: Order.sample(status: OrderStatus.pending)),
);
@Preview(name: 'OrderCard — delivered, long title')
Widget orderCardLongTitle() => _wrap(
OrderCard(
order: Order.sample(
status: OrderStatus.delivered,
title: 'Refrigerated pallet delivery — bay 14, gate C, after hours',
),
),
);
@Preview(name: 'OrderCard — dark, large text', textScaleFactor: 1.8)
Widget orderCardAccessible() => _wrap(
OrderCard(order: Order.sample()),
brightness: Brightness.dark,
);
Widget _wrap(Widget child, {Brightness brightness = Brightness.light}) {
return MaterialApp(
debugShowCheckedModeBanner: false,
theme: ThemeData(brightness: brightness, useMaterial3: true),
home: Scaffold(body: Center(child: Padding(
padding: const EdgeInsets.all(16), child: child))),
);
}
Then run the previewer alongside your editor:
flutter widget-preview start
The API surface here is still settling between releases — check flutter widget-preview --help and the annotation's parameters on the version you are on (name, size, text scale, theme/brightness, wrapper and localization hooks are the ones worth knowing). If your SDK is older, the same discipline works with a small preview/main.dart harness that renders a gallery of states; you lose the hot-attach niceties but keep the fast loop.
Why previews and agents belong in the same article
Previews are how you make an agent's UI work verifiable. The rule "every new widget ships with at least three previews — empty, loaded, and error/overflow" gives both the agent and the reviewer a cheap, visual pass/fail. Pair it with golden tests and the preview states become the golden fixtures:
testWidgets('OrderCard pending matches golden', (tester) async {
await tester.pumpWidget(orderCardPending());
await expectLater(
find.byType(OrderCard),
matchesGoldenFile('goldens/order_card_pending.png'),
);
});
Now a UI regression introduced by an agent (or a human) fails in CI with a picture attached, instead of being discovered in a store review.
Part 3 — The loop, end to end
Here is the workflow we actually ask teams to adopt on an agent-assisted Flutter engagement.
- Start the app and attach the agent. Run the app on a device or simulator. With the Dart MCP server connected, the agent can hot reload into that running instance and inspect the widget tree.
- Write the task as a contract, not a wish. "Add a
DeliveryWindowPickerwidget inlib/ui/orders/, Riverpod notifier inlib/state/, three previews, one widget test. Do not add dependencies." Scope beats eloquence. - Let it iterate against the analyzer. The agent edits, analyzes, fixes its own diagnostics, runs tests. You do not review intermediate states.
- Look at the previews. Empty, loaded, error, 1.8× text scale, dark mode, 320 dp width. Most layout sins are visible here in ten seconds.
- Review the diff like a human wrote it — more carefully. See the checklist below.
- Commit with the prompt in the message body. When something is wrong three weeks later, knowing what was asked for is worth more than knowing who approved it.
CI is the backstop
Agent-assisted or not, none of this is safe without a gate. A minimal workflow:
name: flutter-ci
on: [pull_request]
jobs:
verify:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: subosito/flutter-action@v2
with: { channel: stable, cache: true }
- run: flutter pub get
- run: dart format --output=none --set-exit-if-changed .
- run: flutter analyze --fatal-infos
- run: flutter test --coverage
# fail the build if a new dependency sneaked in
- run: git diff --exit-code pubspec.yaml pubspec.lock
That last step is deliberately blunt. Dependency creep is the most expensive habit agents have: a one-line convenience package that drags in a native plugin, breaks your iOS build, and nobody can say why it is there.
The review checklist for agent-written Flutter
We hand clients this list. It is shorter than it could be because every item has actually bitten somebody.
- Dependencies: any new package? Is it maintained, licensed compatibly, and does it have the platforms you ship?
- Deprecated APIs: analyzer set to error on
deprecated_member_usecatches this; confirm it is on. - Disposal: every
AnimationController,TextEditingController,StreamSubscription,Timercreated is disposed. - Async + context: no
BuildContextused after anawaitwithout amountedcheck. - Rebuild scope: is the whole screen rebuilding where a
Consumer/selectwould do? Agents over-rebuild. - Error and empty states: did it only build the happy path? It usually did.
- Accessibility: semantic labels on icon-only buttons, 48 dp touch targets, does it survive 2.0× text scale?
- Secrets and config: no API keys inline, no hardcoded base URLs, no logging of tokens.
- Tests: does the new test actually fail if you revert the fix? Delete the implementation and check.
- Platform code: changes to
ios/,android/,Podfile, Gradle files deserve a human rewrite, not a skim.
What we do not let agents drive
Hard-won boundaries, from engagements where we learned them the hard way:
- Release and signing config. Keystores, provisioning, notarization, store metadata. One bad edit here costs a release cycle.
- Data migrations. Drift/sqflite schema migrations touch user data that has no undo.
- Security-critical paths. Auth flows, token storage, certificate pinning, payment entitlement checks. Review every line, and write the tests yourself.
- Architecture decisions. An agent will happily introduce a fourth way of managing state because the file it was looking at did it that way.
Does it actually pay off?
On our own work the honest accounting is: large wins on boilerplate (model classes, copyWith, serialization, repository scaffolding, test fixtures, ARB localisation plumbing, one-off migration scripts), real wins on exploratory refactors where the analyzer can verify the result, and roughly break-even on genuinely novel UI or concurrency work — where the review cost eats the typing savings.
The teams that get the large wins are the ones who invested the afternoon in the setup: MCP server connected, lint rules set to error, previews for the design system, CI that fails loudly. The teams that got burned treated the agent as a junior developer who never needs a code review.
Takeaways
- Connect the Dart MCP server so your agent can analyze, test, hot reload, and inspect — feedback beats prompting.
- Encode house rules in
analysis_options.yamland a short agent rules file; machine-checked beats English. - Make
deprecated_member_usean error. It eliminates the single largest category of stale AI output. - Give every widget previews covering empty / loaded / error / large-text / dark / narrow, and promote them to golden tests.
- Gate everything behind CI, including a check that
pubspec.lockdid not change unexpectedly. - Keep signing, migrations, auth, and architecture in human hands.
Need a Flutter team that already works this way? AviaryApps supplies senior Flutter consultants and developers — for full app builds, for modernisation work, or as a subcontracted delivery team alongside your own engineers. Get in touch and tell us what you are building.