Learning Goals 5 min
"AI" usually means "data uploaded to a cloud server with a GPU". Edge AI moves that inference onto small devices — phones, doorbells, sensors. With TensorFlow Lite Micro a Nano 33 BLE Sense can run a real neural network and recognise gestures. By the end of this lesson you will:
- Compare cloud AI vs edge AI and explain why each fits different problems.
- Identify the resource constraints of edge devices (RAM, flash, no GPU, milliwatt power budget).
- Name three real edge AI uses: keyword spotting ("Hey Google"), gesture recognition, image classification.
Warm-Up 10 min
No new hardware today — conceptual lesson. Bring a Nano 33 BLE Sense if you have it; you'll meet its sensors.
Where AI runs today
| Location | Hardware | Model size | Latency |
|---|---|---|---|
| Cloud (OpenAI, Google) | A100 GPU clusters | ~1 TB (LLMs) | 200 ms – seconds + network |
| Phone (Siri / Google Lens) | Mobile NPU | ~100 MB | ~50 ms |
| Edge AI (Nano 33 BLE Sense) | nRF52840 ARM Cortex-M4 | ~50 KB | ~10 ms |
Edge AI runs the smallest models on the smallest chips. The trade-off: 50 KB of model vs 1 TB. You can't run ChatGPT on a Nano. But you CAN classify gestures, detect keywords ("hey"), recognise objects from a tiny camera. The use cases just have to be small enough.
New Concept · Edge vs cloud trade-offs 25 min
Why edge over cloud
- Privacy: data never leaves the device. Camera frames, microphone audio, biometric sensors — all stay local.
- Latency: no round trip to the cloud. 10 ms decisions instead of 200+ ms.
- Reliability: works offline. The doorbell that says "Amazon delivery" doesn't care if your WiFi is down.
- Cost: no per-API-call fee. One small board, free inference for life.
- Power: edge inference can be milliwatts. Cloud inference is kilowatts.
Why cloud over edge
- Model size: GPT-class models simply don't fit on edge hardware.
- Updates: cloud models update centrally; edge models need OTA flashes.
- Collective data: cloud sees all users; edge sees one.
Edge AI examples in the wild
| Product | Model | Chip class |
|---|---|---|
| Alexa "wake word" | ~50 KB keyword spotter | ARM Cortex-M4 |
| Apple Face ID | Custom NN on the Neural Engine | Apple A-series NPU |
| Ring doorbell person-detect | ~200 KB image classifier | NPU on the doorbell SoC |
| Smart-watch heart-rhythm detection | Tiny LSTM | ARM Cortex-M |
| Industrial "is this machine failing" vibration analyser | ~10 KB autoencoder | STM32 |
The hardware: Nano 33 BLE Sense
Arduino's edge-AI poster board. Includes:
- nRF52840 ARM Cortex-M4 + 1 MB flash + 256 KB RAM.
- IMU (LSM9DS1 — 9-axis: accel + gyro + magnetometer).
- Microphone (MP34DT05).
- Temperature + humidity sensor (HTS221).
- Pressure sensor (LPS22HB).
- Gesture / proximity / light (APDS-9960).
- BLE radio.
That's a wearable's worth of sensors on one small board. The right brain for gesture / sound / breath / etc. recognition.
The Edge AI workflow
- Collect training data on the device (acceleration during "wave" gesture, etc.).
- Train a small model on a laptop or in the cloud (Google Colab, Edge Impulse).
- Convert the trained model to TensorFlow Lite Micro (a tiny inference runtime).
- Deploy the model + runtime onto the device as a C array.
- Run the model on live sensor data in your Arduino sketch.
We'll work through this end-to-end over L04-32 to L04-36.
Worked Example · Run the bundled gesture-classifier demo 20 min
If you have a Nano 33 BLE Sense
- Install the "Arduino_TensorFlowLite" library (Tools → Manage Libraries).
- File → Examples → Arduino_TensorFlowLite → magic_wand.
- Upload to your Nano 33 BLE Sense.
- Open Serial Monitor at 9600 baud.
- Hold the board flat. Move it in a clear "W", "O", or "ring" gesture.
- The serial monitor prints which gesture it recognised (with confidence).
That's a real neural network classifying real-time IMU data on one small board with no internet. The TensorFlow Lite runtime occupies ~80 KB of flash; the model itself is ~20 KB.
If you don't have one
You can train + simulate models on Google Colab and watch the inference. The hardware is needed only for the actual on-device deployment. We'll provide simulation paths in L04-33.
Basic 5 min
Goal: Work out which models fit on which board. A model needs flash for the TFLM runtime (about 80 KB) plus the model itself. It also needs RAM to work in.
| Board | Flash | RAM |
|---|---|---|
| Arduino UNO | 32 KB | 2 KB |
| Nano 33 BLE Sense | 1024 KB | 256 KB |
| Model | Model size | RAM to run |
|---|---|---|
| A · gesture classifier | 20 KB | 8 KB |
| B · person detector | 200 KB | 100 KB |
| C · chatbot | 1 TB | 16 GB |
For each model and each board, write fits or does not fit. Where it fits, give the percentage of flash and RAM it uses.
It works if you have six answers, and every percentage is rounded to a whole number.
Challenge 1 5 min
Goal: Put numbers on two reasons for edge AI: bandwidth and latency. Show your working for each part.
- A smart speaker sends all its audio to the cloud. Audio is 16,000 samples per second, 2 bytes per sample. How many bytes does it send in one day? Give the answer in GB.
- A smarter speaker spots its wake word on the device. It only sends 4 seconds of audio per command, 20 commands a day. How many bytes per day now? How many times less is that?
- A car travels at 20 m/s. A cloud pedestrian detector answers in 250 ms. An edge detector answers in 10 ms. How far does the car move while waiting for each answer?
It works if your three answers have units, and part 2 is about a thousand times smaller than part 1.
Challenge 2 5 min
Goal: Time one tiny neural network layer on your own board. The function below is one layer: 16 inputs feeding 8 neurons. That is 128 multiply-adds.
const int INPUTS = 16;
const int OUTPUTS = 8;
const int RUNS = 100;
float inputs[INPUTS];
float weights[OUTPUTS][INPUTS];
float outputs[OUTPUTS];
void denseLayer() {
for (int o = 0; o < OUTPUTS; o++) {
float sum = 0.0;
for (int i = 0; i < INPUTS; i++) {
sum += inputs[i] * weights[o][i];
}
outputs[o] = sum > 0 ? sum : 0;
}
}
void setup() {
Serial.begin(9600);
for (int i = 0; i < INPUTS; i++) {
inputs[i] = random(100) / 100.0;
for (int o = 0; o < OUTPUTS; o++) {
weights[o][i] = random(100) / 100.0;
}
}
// your timing code here
}
void loop() {
}- Add code in
setup()that callsdenseLayer()RUNStimes and times it withmicros(). - Print the time for one layer, in microseconds.
- The first layer of the ARD-L04-34 model has 450 inputs and 32 neurons. Use your result to estimate how long that layer takes on your board.
It works if pressing reset three times gives three times within 5% of each other.
Challenge 3 · Decide on the edge 10 min
This sketch behaves like a cloud device. It streams every MPU6050 reading out of the board and lets something else decide. It also counts the bytes it sends. Serial.print() returns how many bytes it wrote.
#include <Adafruit_MPU6050.h>
#include <Adafruit_Sensor.h>
#include <Wire.h>
Adafruit_MPU6050 mpu;
const unsigned long SAMPLE_MS = 10;
const unsigned long REPORT_MS = 10000;
unsigned long lastSample = 0;
unsigned long lastReport = 0;
unsigned long bytesSent = 0;
void setup() {
Serial.begin(115200);
Wire.begin();
mpu.begin();
}
void loop() {
unsigned long now = millis();
if (now - lastSample < SAMPLE_MS) return;
lastSample = now;
sensors_event_t a, g, t;
mpu.getEvent(&a, &g, &t);
bytesSent += Serial.print(a.acceleration.x);
bytesSent += Serial.print(",");
bytesSent += Serial.print(a.acceleration.y);
bytesSent += Serial.print(",");
bytesSent += Serial.println(a.acceleration.z);
if (now - lastReport >= REPORT_MS) {
lastReport = now;
Serial.print("# bytes in last 10 s: ");
Serial.println(bytesSent);
bytesSent = 0;
}
}- Upload it. Write down the bytes sent in 10 seconds.
- Change it into an edge device. The board works out the size of the acceleration itself. It sends one line,
SHAKEand the time, only when the size passes 2 g (19.6 m/s²). - Send at most one
SHAKEline every 500 ms, so one shake gives one line. - Keep the byte counter. Compare the two numbers.
It works if the board sends almost no bytes while still, and one SHAKE line for each sharp shake.
Recap 5 min
Edge AI = small models on small chips. Wins on privacy, latency, offline reliability, cost. Loses on model size and central updates. Nano 33 BLE Sense is Arduino's edge AI flagship — Cortex-M4 + 9-axis IMU + mic + climate + light. Tomorrow we learn ML basics; the rest of Cluster F applies them.
- Edge AI
- Running inference on the data-generating device, not in the cloud.
- Inference
- Running a trained model to make a prediction. Distinct from training (which produces the model).
- TensorFlow Lite Micro
- Google's ML runtime for microcontrollers. ~16 KB code footprint. Runs on Cortex-M0+ and up.
- NPU (Neural Processing Unit)
- Specialised AI chip. Phones (Apple Neural Engine, Google Tensor) include them; Arduino-class don't.
- Model size
- How many bytes the trained weights occupy. Tiny models: < 100 KB. Big models: GB+.
- Wake word
- A small phrase ("Hey Google") detected on-device before activating cloud-based interpretation.
- Keyword spotting
- Detecting one of a small vocabulary of words. Doable on edge.
- Anomaly detection
- Spotting "this isn't the normal pattern". Small autoencoder models excel.
Extra Mission 5 min
Part 1 — Design an edge AI product
Pick one kind of edge AI: gesture recognition (accelerometer), keyword spotting (microphone) or vibration checking (a machine that might be failing). On paper, design one product that uses it.
Your design must include:
- A name, the one job it does, and who uses it.
- The sensor it reads and the board it runs on.
- A guess at the model size, and a check that it fits (use your Basic working).
- One reason it must run on the edge, not in the cloud.
- What it does when it is not sure.
Part 2 — Make it
You cannot train a model yet, so build a rule-based version of your product's detector. Use a sensor you already have, such as the MPU6050. The board must decide by itself and light an LED when the event happens. Like Challenge 3, it sends only events, never the raw stream.
Bring back next class: your uploaded sketch, a photo of the working circuit and a Serial Monitor screenshot showing a few events. Set up a free Google Colab account before ARD-L04-34.