+1 (415) 480-3939

On-Device ML in Flutter: Camera Streams, ML Kit, LiteRT, and Small Local Models

Almost every Flutter engagement that starts with "we want AI in the app" ends up in the same place: not everything should go to a server. Scanning a document, reading a barcode, blurring a face, classifying a photo, or summarising text a user typed — these are latency-sensitive, privacy-sensitive, and often need to work in a warehouse with no signal. That is on-device inference, and in 2026 the Flutter tooling for it is finally boring enough to ship.

Our other AI posts cover calling hosted models and streaming their output. This one is the opposite half of the problem: a camera feed on the device, a model running locally, and a frame budget you cannot blow. The stack is camera + Google ML Kit for the solved problems, LiteRT (the runtime formerly called TensorFlow Lite) for your own models, and the on-device small models exposed through Firebase AI Logic / Gemini Nano for text.

Pick the lowest tier that solves the problem

Teams lose weeks by starting at the bottom of this table instead of the top.

TierUse whenCost to ship
ML Kit / Vision APIsbarcodes, text (OCR), faces, pose, object detection, translationhours; no model to train or host
LiteRT with a stock modelclassification or detection over a known label setdays; model conversion + preprocessing
LiteRT with a custom modelyour domain, your labels, your dataweeks; a data pipeline is the real project
On-device LLM (Gemini Nano class)short text summarise/classify/rewrite, offlinedays; device-gated, needs a cloud fallback
Hosted modelanything large, or anything that changes weeklydays; network + cost + privacy review

The consulting answer we give most often is: use ML Kit until it demonstrably cannot do the job, because it removes the entire model lifecycle from the project.

Dependencies

dependencies:
  camera: ^0.11.0+2
  google_mlkit_barcode_scanning: ^0.13.0
  google_mlkit_text_recognition: ^0.14.0
  tflite_flutter: ^0.11.0        # LiteRT bindings
  image: ^4.2.0
  permission_handler: ^11.3.1

Two platform notes that cause the first build failure on every project: ML Kit's iOS pods push the minimum deployment target up (set platform :ios, '15.5' in the Podfile and match it in Xcode), and the unbundled Android models download on first use, so your first scan on a fresh device needs connectivity unless you bundle them in the manifest.

The camera stream is the hard part, not the model

startImageStream fires at the sensor frame rate — 30 or 60 times a second. Inference takes 20–200 ms. If you await inside the callback without a guard, you queue frames, memory climbs, and the app dies on a mid-range Android device in about forty seconds. This single-slot pattern is the fix:

class FrameThrottler {
  bool _busy = false;
  DateTime _last = DateTime.fromMillisecondsSinceEpoch(0);
  final Duration minGap;

  FrameThrottler({this.minGap = const Duration(milliseconds: 120)});

  /// Runs [task] only if the previous run finished and the cool-down elapsed.
  /// Frames arriving while busy are dropped, deliberately.
  Future<void> maybeRun(Future<void> Function() task) async {
    if (_busy) return;
    final now = DateTime.now();
    if (now.difference(_last) < minGap) return;
    _busy = true;
    try {
      await task();
    } finally {
      _last = DateTime.now();
      _busy = false;
    }
  }
}

Dropping frames is not a compromise. A barcode scanner running at 8 fps feels identical to one running at 30 fps, and uses a quarter of the battery.

Future<void> _start() async {
  _controller = CameraController(
    _backCamera,
    ResolutionPreset.medium,              // 720p is plenty for detection
    enableAudio: false,
    imageFormatGroup: Platform.isAndroid
        ? ImageFormatGroup.nv21           // what ML Kit wants on Android
        : ImageFormatGroup.bgra8888,      // and on iOS
  );
  await _controller.initialize();
  await _controller.startImageStream((CameraImage image) {
    _throttler.maybeRun(() => _process(image));
  });
}

Converting a CameraImage without copying it three times

The conversion from CameraImage to ML Kit's InputImage is where most sample code allocates a new buffer per frame. Keep it flat and reuse the metadata:

InputImage? _toInputImage(CameraImage image, CameraDescription camera, int deviceOrientation) {
  final rotation = _rotationOf(camera, deviceOrientation);
  if (rotation == null) return null;

  final format = InputImageFormatValue.fromRawValue(image.format.raw);
  if (format == null) return null;

  return InputImage.fromBytes(
    bytes: image.planes.first.bytes,       // nv21 / bgra8888 are single-plane here
    metadata: InputImageMetadata(
      size: Size(image.width.toDouble(), image.height.toDouble()),
      rotation: rotation,
      format: format,
      bytesPerRow: image.planes.first.bytesPerRow,
    ),
  );
}

Rotation is the bug that survives to production, because it only shows up on one axis of one platform. Android reports sensor orientation that must be combined with the current device orientation; iOS handles it for you unless the app locks orientation. Test landscape-left explicitly — it is the case everyone forgets.

Tier 1: ML Kit, end to end

final _scanner = BarcodeScanner(formats: [BarcodeFormat.code128, BarcodeFormat.qrCode]);

Future<void> _process(CameraImage image) async {
  final input = _toInputImage(image, _backCamera, _deviceOrientation);
  if (input == null) return;

  final barcodes = await _scanner.processImage(input);
  if (barcodes.isEmpty) return;

  final value = barcodes.first.rawValue;
  if (value == null || value == _lastAccepted) return;   // de-dupe consecutive reads

  _lastAccepted = value;
  HapticFeedback.mediumImpact();
  await _controller.stopImageStream();                   // stop the moment you have it
  if (mounted) Navigator.of(context).pop(value);
}

@override
void dispose() {
  _scanner.close();          // native resources are NOT garbage collected
  _controller.dispose();
  super.dispose();
}

Restricting formats is a real performance setting, not documentation: scanning for every symbology costs several milliseconds per frame. And close() matters — leaked detectors show up as a slow, confusing memory climb after the user opens the scanner a dozen times.

Tier 2: your own model with LiteRT

Custom models add three responsibilities: preprocessing that exactly matches training, an isolate so inference never touches the UI thread, and a versioning story.

class Classifier {
  Classifier._(this._interpreter, this._labels);
  final Interpreter _interpreter;
  final List<String> _labels;

  static Future<Classifier> load() async {
    final options = InterpreterOptions()..threads = 2;
    // Hardware delegates: big wins, but verify per device — some Android GPU
    // drivers are slower than CPU for small models, and a few simply crash.
    if (Platform.isAndroid) options.addDelegate(XNNPackDelegate());
    if (Platform.isIOS) options.addDelegate(GpuDelegate());

    final interpreter = await Interpreter.fromAsset('assets/models/defects_v3.tflite', options: options);
    final labels = (await rootBundle.loadString('assets/models/defects_v3.txt')).split('\n');
    return Classifier._(interpreter, labels);
  }

  ({String label, double score}) run(Float32List input) {
    final output = List.filled(_labels.length, 0.0).reshape([1, _labels.length]);
    _interpreter.run(input.reshape([1, 224, 224, 3]), output);
    final scores = (output[0] as List).cast<double>();
    var best = 0;
    for (var i = 1; i < scores.length; i++) if (scores[i] > scores[best]) best = i;
    return (label: _labels[best], score: scores[best]);
  }
}

Run it off the UI isolate. Isolate.run is fine for one-shot work; for a camera stream, keep a long-lived worker so the interpreter is loaded once:

// In the worker isolate.
final classifier = await Classifier.load();
await for (final job in receivePort.cast<_Job>()) {
  job.reply.send(classifier.run(job.pixels));
}

Two rules that save re-training cycles. Preprocess identically to training — same resize algorithm, same channel order, same normalisation (/255, or mean/std, or [-1,1]; guessing wrong yields a model that is confidently wrong, never obviously broken). And always ship a confidence floor with a human-readable fallback: below ~0.6, show "not sure — tap to enter manually" rather than a wrong answer stated firmly.

Tier 3: on-device text models

Small on-device LLMs now handle summarise / classify / rewrite over short text without a round trip. They are genuinely useful and genuinely limited: availability is gated by device, RAM, and sometimes a model download the OS controls. Treat availability as a runtime question:

Future<String> summarise(String notes) async {
  if (await OnDeviceModel.isAvailable()) {
    try {
      return await OnDeviceModel.summarise(notes, maxOutputTokens: 120);
    } on ModelUnavailableException {
      // fall through
    }
  }
  return _cloud.summarise(notes);   // same interface, different cost and privacy profile
}

Keep the two paths behind one repository interface, feature-flag the on-device path, and log which path served each request. When a stakeholder asks "is the AI working", the answer needs to be a number, not a feeling.

Budgets and how to defend them

Attach numbers to this work in the first sprint, or the feature ships and quietly drains batteries:

  • Latency: time from frame captured to result rendered. Aim under 150 ms for interactive scanning; measure with Timeline.timeSync and read it in DevTools.
  • Sustained frame rate: the preview must hold 60 fps while inference runs. If it does not, the work is on the wrong isolate.
  • Memory ceiling: watch RSS across a five-minute scanning session. Flat is pass; a slope is a leaked detector or a queued frame buffer.
  • Thermals and battery: a ten-minute run on a mid-range Android device, not a flagship. Throttled devices halve their frame budget, and that is where dropped-frame logic earns its place.

Put the first three in an integration test that fails CI on regression — the same discipline as any other performance budget.

Permissions, privacy, and the store review

On-device inference is a privacy selling point, so claim it accurately:

  • Request camera permission in context, with a one-line reason, and handle permanent denial by deep-linking into Settings.
  • If frames never leave the device, say so in the UI and in your store listing — and make sure that stays true when someone later adds a cloud fallback.
  • Declare data handling honestly in Google Play's Data Safety form and in the iOS privacy manifest. "Processed on device, not collected" is a valid answer only when it is the real one.
  • If a cloud fallback exists, disclose it, and let the user turn it off.

Face and biometric-adjacent features attract extra scrutiny in both stores and under regional biometric-privacy law. Bring legal in before the sprint, not before the release.

Where projects go wrong

The failures we get called in to fix are rarely model quality:

  1. Inference on the UI isolate. The demo looked fine on a flagship; the pilot ran on €200 Android handsets.
  2. No frame throttling, so memory grows until the OS kills the app.
  3. Model files shipped as assets with no version field, making a bad model unfixable without a store release. Give models a version and a remote update path — the same discipline as any other shipped artefact.
  4. No confidence floor, so users are told the wrong thing with total conviction and stop trusting the feature.
  5. Testing only in a bright office. Test in the warehouse, the loading bay, and at dusk.

Where to start

Build the throttled camera pipeline first with ML Kit barcode scanning behind it — it is a day's work and it proves the hard parts (rotation, formats, lifecycle, permissions, budgets). Swap in a LiteRT model only when the business problem genuinely needs one, and keep the interface between pipeline and model narrow enough that the swap is a one-file change.

If you would like a second pair of eyes on an on-device ML feature — scoping a pilot, reviewing a frame pipeline that jitters, or staffing senior Flutter developers alongside your team — get in touch and we will walk through it with you.