Edge AI on drones
UAT 322 Artificial Intelligence and Autonomous Unmanned Aircraft Systems Integration Laboratory
Lesson
By the end of this module you will be able to
- Explain the steps of deploying an AI model to an edge device, from ONNX export to hardware optimisation
- Calculate INT8 quantization and its error bound
- Measure frame-to-result latency with p95 and compare it with the per-frame budget
- Write a model card stating test conditions, limitations and stop criteria
Why this matters
An object detector that is highly accurate on a desktop computer may be too slow once moved to an onboard board running on a few tens of watts. If a person or obstacle is detected half a second late, a drone flying at 8 m/s has already moved 4 m. Deploying AI on edge devices (edge AI) is therefore a trade-off between speed, accuracy, memory and power, which must be measured, not read from marketing.
From model to board
- Train and evaluate on a powerful machine (UAT 315)
- Export to an interchange format such as ONNX, and check the outputs match the original model
- Optimise for the hardware, for example ONNX Runtime on a CPU or TensorRT on an NVIDIA GPU, and reduce numeric precision with quantization
- Measure on the real board: speed, accuracy, memory, temperature and power
For popular boards such as the Jetson Orin Nano Super, NVIDIA states 67 TOPS, an INT8 sparse figure that says nothing about the speed of our model. The figures used for decisions must come from measuring the real model on the real board.
Quantization
Quantization stores a model’s weights and values as 8-bit integers instead of 32-bit floats, saving 4× memory and computing faster on supporting hardware. ONNX Runtime uses the linear mapping , where is the scale and the zero-point (the integer representing 0.0). Jacob et al. (2018) set out the integer-only arithmetic now widely used.
Example 1 Converting FP32 weights to INT8
import numpy as np
rng = np.random.default_rng(348)
w = rng.normal(0, 0.8, 1000).astype(np.float32) # 1000 hypothetical weights
lo, hi = float(w.min()), float(w.max())
scale = (hi - lo) / 255
zero_point = int(round(-lo / scale))
q = np.clip(np.round(w / scale) + zero_point, 0, 255).astype(np.uint8)
restored = scale * (q.astype(np.float32) - zero_point)
err = np.abs(restored - w)
print(f"range [{lo:.3f}, {hi:.3f}], scale {scale:.5f}, zero-point {zero_point}")
print(f"max error {err.max():.5f} (half a step = {scale / 2:.5f}); memory {w.nbytes} -> {q.nbytes} bytes")
range [-2.329, 2.248], scale 0.01795, zero-point 130
max error 0.00896 (half a step = 0.00897); memory 4000 -> 1000 bytes
The largest error is no more than half a scale step, and memory falls 4×. Yet thousands of small errors together can change detections, so re-evaluate accuracy after every quantization. ONNX Runtime offers two kinds: static, which uses sample data to find scales ahead of time (recommended for CNNs), and dynamic, which computes activation scales at run time.
Measuring latency properly
The time the model spends computing (inference) is not the time from receiving a frame to having a result (end-to-end), which also includes preprocessing and postprocessing such as filtering overlapping boxes. Measure after a warm-up and report p95, the value 95% of samples do not exceed, because the mean hides the unusually slow frames.
Example 2 FP32 versus INT8 at a 30 fps budget
Continuing from Example 1 (with the same random generator), simulate 400 frames of each stage, in milliseconds.
def p95(x):
s = np.sort(x)
return s[int(np.ceil(0.95 * len(s))) - 1] # nearest-rank
n = 400
pre = rng.normal(4.0, 0.5, n)
fp32 = rng.gamma(9, 4.6, n)
int8 = rng.gamma(9, 2.0, n)
post = rng.normal(3.0, 0.4, n)
budget = 1000 / 30
for name, inference in (("FP32", fp32), ("INT8", int8)):
e2e = pre + inference + post
print(f"{name}: inference median {np.median(inference):.1f} ms | end-to-end mean {e2e.mean():.1f}, "
f"p95 {p95(e2e):.1f} ms | over {budget:.1f} ms: {(e2e > budget).mean():.0%} | "
f"distance at 8 m/s during p95: {8 * p95(e2e) / 1000:.2f} m")
FP32: inference median 40.8 ms | end-to-end mean 49.8, p95 73.5 ms | over 33.3 ms: 90% | distance at 8 m/s during p95: 0.59 m
INT8: inference median 17.7 ms | end-to-end mean 25.3, p95 36.1 ms | over 33.3 ms: 10% | distance at 8 m/s during p95: 0.29 m
INT8 is much faster, but its p95 still slightly exceeds the 33.3 ms budget; about 10% of frames are late. Options are to reduce image size, process fewer frames, or accept it and add safety distance. Decisions must always consider accuracy after quantization too.
The model card
Mitchell et al. (2019) proposed the model card, a short document accompanying every model version. For drones it should include at least:
- Version and provenance: model version, training data set and software versions
- Test conditions: board, image size, runtime, numeric precision, warm-up runs and sample count
- Results: accuracy by condition (daylight, dusk, haze) and p95 latency
- Limitations and out-of-scope uses, such as “not tested at night”
- Stop and rollback criteria, such as reverting to the previous version if field accuracy falls below a threshold
Module lab
Lab: measuring a detector on the real board
- Export the detector from UAT 315 to ONNX and check its outputs match the original model on the same images.
- Apply static quantization with ONNX Runtime, using images from real work as calibration data, and evaluate accuracy before and after.
- On the companion board, measure model-only and end-to-end times after at least 10 warm-up runs, over 400 runs, saved as CSV.
- Compute p95 with the code in Example 2, compare with the per-frame budget, and record board temperature and power during the test.
- Write a one-page model card with stop criteria.
Common mistakes
Watch out
- Using the manufacturer’s TOPS instead of real measurements
- Measuring only inference time and calling it system latency
- Measuring without warm-up, or reporting only the mean
- Quantizing without re-evaluating accuracy
- Measuring on a PC and drawing conclusions for the drone’s board
Summary
- The steps are train, export to ONNX, optimise for the hardware, then measure on the real board
- INT8 quantization uses , cutting memory 4×, but accuracy must be re-evaluated
- Latency must be measured end to end after warm-up and reported as p95 against the per-frame budget
- A model card states test conditions, results, limitations and stop criteria
Check your understanding
- Weights range from −1.0 to 1.55. What is the INT8 scale?
- With scale 0.02 and zero-point 50, what real value does represent?
- What is the per-frame budget at 25 fps?
- Twenty times are 1 to 20 ms. What is the nearest-rank p95?
- Why must accuracy be re-evaluated after quantization?
Answers
- ms
- Rank , so 19 ms
- Many small rounding errors together can change the detections
Key formulas
| Linear quantization | |
| Nearest-rank p95 | |
| Per-frame budget |
Key references
- Microsoft. Quantize ONNX models. ONNX Runtime documentation. link
- ONNX Runtime. Documentation. link
- NVIDIA. NVIDIA TensorRT documentation. link
- Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., & Kalenichenko, D. (2018). Quantization and training of neural networks for efficient integer-arithmetic-only inference. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 2704–2713). link
- NVIDIA. (2024). NVIDIA Jetson Orin Nano developer kit gets a "super" boost. NVIDIA Technical Blog. link
- Warden, P., & Situnayake, D. (2020). TinyML: Machine learning with TensorFlow Lite on Arduino and ultra-low-power microcontrollers. O'Reilly Media.
- Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., & Gebru, T. (2019). Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency (pp. 220–229). ACM. link
Further reading
Study the assigned knowledge units in advance, review media and take the module quiz
Deploying AI models to edge devices
Edge benchmarking and deployment planning
Computer vision and the path to edge AI
In class / field
Intensive lab and field practice recorded in a lab notebook
Learning evidence: Lab notebook signed by the instructor