Module 4/5 · Weeks 10–12 · 27 h

Edge AI on drones

UAT 322 Artificial Intelligence and Autonomous Unmanned Aircraft Systems Integration Laboratory

About 90 minDraft, awaiting reviewLast updated 27 September 2026

Lesson

By the end of this module you will be able to

  1. Explain the steps of deploying an AI model to an edge device, from ONNX export to hardware optimisation
  2. Calculate INT8 quantization and its error bound
  3. Measure frame-to-result latency with p95 and compare it with the per-frame budget
  4. Write a model card stating test conditions, limitations and stop criteria

Prerequisites: UAT 315 (AI and computer vision) · UAT 322 module 1

Why this matters

An object detector that is highly accurate on a desktop computer may be too slow once moved to an onboard board running on a few tens of watts. If a person or obstacle is detected half a second late, a drone flying at 8 m/s has already moved 4 m. Deploying AI on edge devices (edge AI) is therefore a trade-off between speed, accuracy, memory and power, which must be measured, not read from marketing.

From model to board

  1. Train and evaluate on a powerful machine (UAT 315)
  2. Export to an interchange format such as ONNX, and check the outputs match the original model
  3. Optimise for the hardware, for example ONNX Runtime on a CPU or TensorRT on an NVIDIA GPU, and reduce numeric precision with quantization
  4. Measure on the real board: speed, accuracy, memory, temperature and power

For popular boards such as the Jetson Orin Nano Super, NVIDIA states 67 TOPS, an INT8 sparse figure that says nothing about the speed of our model. The figures used for decisions must come from measuring the real model on the real board.

Quantization

Quantization stores a model’s weights and values as 8-bit integers instead of 32-bit floats, saving 4× memory and computing faster on supporting hardware. ONNX Runtime uses the linear mapping , where is the scale and the zero-point (the integer representing 0.0). Jacob et al. (2018) set out the integer-only arithmetic now widely used.

Example 1 Converting FP32 weights to INT8

import numpy as np

rng = np.random.default_rng(348)
w = rng.normal(0, 0.8, 1000).astype(np.float32)        # 1000 hypothetical weights
lo, hi = float(w.min()), float(w.max())
scale = (hi - lo) / 255
zero_point = int(round(-lo / scale))
q = np.clip(np.round(w / scale) + zero_point, 0, 255).astype(np.uint8)
restored = scale * (q.astype(np.float32) - zero_point)
err = np.abs(restored - w)
print(f"range [{lo:.3f}, {hi:.3f}], scale {scale:.5f}, zero-point {zero_point}")
print(f"max error {err.max():.5f} (half a step = {scale / 2:.5f}); memory {w.nbytes} -> {q.nbytes} bytes")
range [-2.329, 2.248], scale 0.01795, zero-point 130
max error 0.00896 (half a step = 0.00897); memory 4000 -> 1000 bytes

The largest error is no more than half a scale step, and memory falls 4×. Yet thousands of small errors together can change detections, so re-evaluate accuracy after every quantization. ONNX Runtime offers two kinds: static, which uses sample data to find scales ahead of time (recommended for CNNs), and dynamic, which computes activation scales at run time.

The upper line is FP32 values from −2.329 to 2.248; the lower line INT8 values from 0 to 255. Arrows map the minimum to 0, 0.0 to 130 and the maximum to 255. Below: scale 0.01795 and zero-point 130
Figure 1 Mapping FP32 weights to INT8

Measuring latency properly

The time the model spends computing (inference) is not the time from receiving a frame to having a result (end-to-end), which also includes preprocessing and postprocessing such as filtering overlapping boxes. Measure after a warm-up and report p95, the value 95% of samples do not exceed, because the mean hides the unusually slow frames.

Example 2 FP32 versus INT8 at a 30 fps budget

Continuing from Example 1 (with the same random generator), simulate 400 frames of each stage, in milliseconds.

def p95(x):
    s = np.sort(x)
    return s[int(np.ceil(0.95 * len(s))) - 1]            # nearest-rank


n = 400
pre = rng.normal(4.0, 0.5, n)
fp32 = rng.gamma(9, 4.6, n)
int8 = rng.gamma(9, 2.0, n)
post = rng.normal(3.0, 0.4, n)
budget = 1000 / 30
for name, inference in (("FP32", fp32), ("INT8", int8)):
    e2e = pre + inference + post
    print(f"{name}: inference median {np.median(inference):.1f} ms | end-to-end mean {e2e.mean():.1f}, "
          f"p95 {p95(e2e):.1f} ms | over {budget:.1f} ms: {(e2e > budget).mean():.0%} | "
          f"distance at 8 m/s during p95: {8 * p95(e2e) / 1000:.2f} m")
FP32: inference median 40.8 ms | end-to-end mean 49.8, p95 73.5 ms | over 33.3 ms: 90% | distance at 8 m/s during p95: 0.59 m
INT8: inference median 17.7 ms | end-to-end mean 25.3, p95 36.1 ms | over 33.3 ms: 10% | distance at 8 m/s during p95: 0.29 m

INT8 is much faster, but its p95 still slightly exceeds the 33.3 ms budget; about 10% of frames are late. Options are to reduce image size, process fewer frames, or accept it and add safety distance. Decisions must always consider accuracy after quantization too.

Two stacked bars, FP32 and INT8, each split into preprocessing, inference and postprocessing. FP32 totals 49.8 ms on average and INT8 25.3 ms. A pink dashed line at 33.3 ms is the 30 fps budget
Figure 2 Mean time per stage against the per-frame budget (ms)

The model card

Mitchell et al. (2019) proposed the model card, a short document accompanying every model version. For drones it should include at least:

  • Version and provenance: model version, training data set and software versions
  • Test conditions: board, image size, runtime, numeric precision, warm-up runs and sample count
  • Results: accuracy by condition (daylight, dusk, haze) and p95 latency
  • Limitations and out-of-scope uses, such as “not tested at night”
  • Stop and rollback criteria, such as reverting to the previous version if field accuracy falls below a threshold

Module lab

Lab: measuring a detector on the real board

  1. Export the detector from UAT 315 to ONNX and check its outputs match the original model on the same images.
  2. Apply static quantization with ONNX Runtime, using images from real work as calibration data, and evaluate accuracy before and after.
  3. On the companion board, measure model-only and end-to-end times after at least 10 warm-up runs, over 400 runs, saved as CSV.
  4. Compute p95 with the code in Example 2, compare with the per-frame budget, and record board temperature and power during the test.
  5. Write a one-page model card with stop criteria.

Common mistakes

Watch out

  • Using the manufacturer’s TOPS instead of real measurements
  • Measuring only inference time and calling it system latency
  • Measuring without warm-up, or reporting only the mean
  • Quantizing without re-evaluating accuracy
  • Measuring on a PC and drawing conclusions for the drone’s board

Summary

  • The steps are train, export to ONNX, optimise for the hardware, then measure on the real board
  • INT8 quantization uses , cutting memory 4×, but accuracy must be re-evaluated
  • Latency must be measured end to end after warm-up and reported as p95 against the per-frame budget
  • A model card states test conditions, results, limitations and stop criteria

Check your understanding

  1. Weights range from −1.0 to 1.55. What is the INT8 scale?
  2. With scale 0.02 and zero-point 50, what real value does represent?
  3. What is the per-frame budget at 25 fps?
  4. Twenty times are 1 to 20 ms. What is the nearest-rank p95?
  5. Why must accuracy be re-evaluated after quantization?
Answers
  1. ms
  2. Rank , so 19 ms
  3. Many small rounding errors together can change the detections

Key formulas

Linear quantization
Nearest-rank p95
Per-frame budget

Key references

  1. Microsoft. Quantize ONNX models. ONNX Runtime documentation. link
  2. ONNX Runtime. Documentation. link
  3. NVIDIA. NVIDIA TensorRT documentation. link
  4. Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., & Kalenichenko, D. (2018). Quantization and training of neural networks for efficient integer-arithmetic-only inference. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 2704–2713). link
  5. NVIDIA. (2024). NVIDIA Jetson Orin Nano developer kit gets a "super" boost. NVIDIA Technical Blog. link
  6. Warden, P., & Situnayake, D. (2020). TinyML: Machine learning with TensorFlow Lite on Arduino and ultra-low-power microcontrollers. O'Reilly Media.
  7. Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., & Gebru, T. (2019). Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency (pp. 220–229). ACM. link

Further reading

Study the assigned knowledge units in advance, review media and take the module quiz

In class / field

Intensive lab and field practice recorded in a lab notebook

Learning evidence: Lab notebook signed by the instructor

Module quiz

This is a formative self-check, not a graded exam

Knowledge domain: Artificial intelligence and computer vision · Sensors and embedded systems