⚙️ MLOps & Engineering · Lecture 10 of 15

Edge AI and TinyML: Running Models on Phones and Microcontrollers

Running models on-device brings privacy, offline operation, low latency and low cost. We cover the edge hardware spectrum, the optimisation pipeline, TensorFlow Lite, ONNX Runtime and TinyML on microcontrollers, and field-deployment lessons.

Not every AI system can rely on a fast internet connection and a cloud server. A health worker in a remote area needs a diagnostic aid that works offline; a sensor in a field runs for months on a battery; a user's photos should not leave their phone. Edge AI runs inference directly on devices — phones, embedded computers, cameras, microcontrollers. TinyML pushes this to tiny microcontrollers with kilobytes of memory and milliwatts of power.

Why run on the edge?#

BenefitExplanation
PrivacyRaw data (images, audio, health signals) stays on the device
Offline operationWorks without connectivity — essential in many rural and crisis settings
LatencyNo network round-trip; real-time response
CostNo per-request server cost; scales with devices
BandwidthSend only results or alerts, not raw video
EnergyLocal inference can use less energy than transmitting data

Trade-offs: limited compute and memory, heterogeneous hardware, harder updates and monitoring, and model theft risk (the model file is on the device).

The hardware spectrum#

Device classMemoryTypical models
Smartphones (CPU/GPU/NPU)GBsMobileNet/EfficientNet, small LLMs (quantised), on-device speech
Single-board computers / edge accelerators (Raspberry Pi, Jetson, Coral)1–16 GBDetection, segmentation, small transformers
Microcontrollers (Arm Cortex-M, ESP32)32 KB – 1 MB RAMKeyword spotting, anomaly detection, tiny vision, gesture recognition

Many phones include neural processing units (NPUs) that accelerate quantised models dramatically.

The optimisation pipeline#

  1. Choose an efficient architecture (MobileNet, EfficientNet-Lite, small transformers, or tiny CNNs for microcontrollers).
  2. Train (often with transfer learning).
  3. Compress: quantisation (INT8 is standard; sometimes lower), pruning, knowledge distillation.
  4. Convert to an on-device format: TensorFlow Lite / LiteRT, ONNX, Core ML, ExecuTorch.
  5. Benchmark on the target device: latency, memory, energy, accuracy on realistic data.
  6. Deploy and update through app updates or over-the-air model delivery, with versioning.

TensorFlow Lite with full-integer quantisation#

python
import numpy as np
import tensorflow as tf

model = tf.keras.models.load_model("crop_disease.keras")      # a trained Keras model

def representative_data():
    for img in np.load("calibration_images.npy")[:200]:       # ~100–500 real samples
        yield [img[None].astype(np.float32)]

converter = tf.lite.TFLiteConverter.from_keras_model(model)
converter.optimizations = [tf.lite.Optimize.DEFAULT]
converter.representative_dataset = representative_data         # calibrate activation ranges
converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8]
converter.inference_input_type = tf.uint8
converter.inference_output_type = tf.uint8
open("crop_disease_int8.tflite", "wb").write(converter.convert())

interpreter = tf.lite.Interpreter(model_path="crop_disease_int8.tflite")
interpreter.allocate_tensors()
print(interpreter.get_input_details()[0]["shape"], interpreter.get_input_details()[0]["dtype"])

Always compare the quantised model's accuracy with the float model on held-out data — especially per class.

ONNX Runtime#

ONNX is an open format supported by many frameworks. Export from PyTorch and run with ONNX Runtime on CPUs, GPUs and mobile, with execution providers for NPUs:

python
import torch, onnxruntime as ort, numpy as np
from torchvision.models import mobilenet_v3_small

m = mobilenet_v3_small(weights="DEFAULT").eval()
torch.onnx.export(m, torch.randn(1, 3, 224, 224), "mnv3.onnx", input_names=["image"],
                  output_names=["logits"], dynamic_axes={"image": {0: "batch"}}, opset_version=17)
sess = ort.InferenceSession("mnv3.onnx", providers=["CPUExecutionProvider"])
print(sess.run(None, {"image": np.random.rand(1, 3, 224, 224).astype(np.float32)})[0].shape)

TinyML on microcontrollers#

TinyML models fit in tens to hundreds of kilobytes. Classic applications: keyword spotting ("Hey device"), vibration-based predictive maintenance, anomaly detection on sensor data, simple person detection, and wildlife or gunshot sound detection for conservation.

Workflow: train a tiny model (often a small CNN on spectrograms or sensor windows), quantise to INT8, convert with TensorFlow Lite for Microcontrollers (or tools like Edge Impulse), and compile into firmware as a C array. Constraints drive design: no operating system, no dynamic memory allocation, fixed "tensor arena" memory, and careful power management (often running inference only when a cheap trigger fires).

Field-deployment lessons#

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

⚙️ MLOps & Engineering

Feature Stores and Training–Serving Consistency

Feature stores manage features as shared, versioned assets available offline for training and online for serving. We explain training–serving skew, point-in-time correctness, online vs offline stores, and when a feature store is worth it.

Advanced⏱ 4 min#250
⚙️ MLOps & Engineering

GPUs and Hardware for Machine Learning

Understanding hardware helps you train faster and cheaper. We explain why GPUs suit deep learning, the roles of memory capacity and bandwidth, precision and tensor cores, estimating requirements, and choosing between local, cloud and free resources.

Intermediate⏱ 5 min#252
⚙️ MLOps & Engineering

CI/CD for Machine Learning: Testing and Automating ML Systems

Continuous integration and delivery bring software-engineering discipline to ML. We cover the testing pyramid for ML — code, data and model tests — quality gates, continuous training, and a practical GitHub Actions workflow.

Intermediate⏱ 5 min#249