Not every AI system can rely on a fast internet connection and a cloud server. A health worker in a remote area needs a diagnostic aid that works offline; a sensor in a field runs for months on a battery; a user's photos should not leave their phone. Edge AI runs inference directly on devices — phones, embedded computers, cameras, microcontrollers. TinyML pushes this to tiny microcontrollers with kilobytes of memory and milliwatts of power.
Why run on the edge?#
| Benefit | Explanation |
|---|---|
| Privacy | Raw data (images, audio, health signals) stays on the device |
| Offline operation | Works without connectivity — essential in many rural and crisis settings |
| Latency | No network round-trip; real-time response |
| Cost | No per-request server cost; scales with devices |
| Bandwidth | Send only results or alerts, not raw video |
| Energy | Local inference can use less energy than transmitting data |
Trade-offs: limited compute and memory, heterogeneous hardware, harder updates and monitoring, and model theft risk (the model file is on the device).
The hardware spectrum#
| Device class | Memory | Typical models |
|---|---|---|
| Smartphones (CPU/GPU/NPU) | GBs | MobileNet/EfficientNet, small LLMs (quantised), on-device speech |
| Single-board computers / edge accelerators (Raspberry Pi, Jetson, Coral) | 1–16 GB | Detection, segmentation, small transformers |
| Microcontrollers (Arm Cortex-M, ESP32) | 32 KB – 1 MB RAM | Keyword spotting, anomaly detection, tiny vision, gesture recognition |
Many phones include neural processing units (NPUs) that accelerate quantised models dramatically.
The optimisation pipeline#
- Choose an efficient architecture (MobileNet, EfficientNet-Lite, small transformers, or tiny CNNs for microcontrollers).
- Train (often with transfer learning).
- Compress: quantisation (INT8 is standard; sometimes lower), pruning, knowledge distillation.
- Convert to an on-device format: TensorFlow Lite / LiteRT, ONNX, Core ML, ExecuTorch.
- Benchmark on the target device: latency, memory, energy, accuracy on realistic data.
- Deploy and update through app updates or over-the-air model delivery, with versioning.
TensorFlow Lite with full-integer quantisation#
import numpy as np
import tensorflow as tf
model = tf.keras.models.load_model("crop_disease.keras") # a trained Keras model
def representative_data():
for img in np.load("calibration_images.npy")[:200]: # ~100–500 real samples
yield [img[None].astype(np.float32)]
converter = tf.lite.TFLiteConverter.from_keras_model(model)
converter.optimizations = [tf.lite.Optimize.DEFAULT]
converter.representative_dataset = representative_data # calibrate activation ranges
converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8]
converter.inference_input_type = tf.uint8
converter.inference_output_type = tf.uint8
open("crop_disease_int8.tflite", "wb").write(converter.convert())
interpreter = tf.lite.Interpreter(model_path="crop_disease_int8.tflite")
interpreter.allocate_tensors()
print(interpreter.get_input_details()[0]["shape"], interpreter.get_input_details()[0]["dtype"])Always compare the quantised model's accuracy with the float model on held-out data — especially per class.
ONNX Runtime#
ONNX is an open format supported by many frameworks. Export from PyTorch and run with ONNX Runtime on CPUs, GPUs and mobile, with execution providers for NPUs:
import torch, onnxruntime as ort, numpy as np
from torchvision.models import mobilenet_v3_small
m = mobilenet_v3_small(weights="DEFAULT").eval()
torch.onnx.export(m, torch.randn(1, 3, 224, 224), "mnv3.onnx", input_names=["image"],
output_names=["logits"], dynamic_axes={"image": {0: "batch"}}, opset_version=17)
sess = ort.InferenceSession("mnv3.onnx", providers=["CPUExecutionProvider"])
print(sess.run(None, {"image": np.random.rand(1, 3, 224, 224).astype(np.float32)})[0].shape)TinyML on microcontrollers#
TinyML models fit in tens to hundreds of kilobytes. Classic applications: keyword spotting ("Hey device"), vibration-based predictive maintenance, anomaly detection on sensor data, simple person detection, and wildlife or gunshot sound detection for conservation.
Workflow: train a tiny model (often a small CNN on spectrograms or sensor windows), quantise to INT8, convert with TensorFlow Lite for Microcontrollers (or tools like Edge Impulse), and compile into firmware as a C array. Constraints drive design: no operating system, no dynamic memory allocation, fixed "tensor arena" memory, and careful power management (often running inference only when a cheap trigger fires).