Module 5/5 · Weeks 13–15 · 27 h

Model evaluation and edge

UAT 306 Artificial Intelligence for UAS

About 85 minDraft, awaiting reviewLast updated 28 September 2026

Lesson

By the end of this module you will be able to

  1. Measure latency correctly, excluding warm-up and reporting percentiles
  2. Separate model time from end-to-end time from image capture to result
  3. Explain INT8 quantization and its effect on size and error
  4. Plan deployment on a companion computer with acceptance criteria

Prerequisites: UAT 306 Modules 1–4 · UAT 315 Module 5 (AI in practice)

Why this matters

A model that is accurate on a PC may be too slow on a companion computer, or too large for its memory. The drone knowledge base’s unit on benchmarking and edge deployment stresses that model inference time and time from image capture to result are different measures; you must state what is timed, the warm-up, the number of runs, the machine, the runtime and the image size; and p95 means 95% of samples finish within that time. UAT 315 computed how far the drone moves during processing; this module focuses on how to measure and how to make models smaller.

Measuring time correctly

The first few runs of a model are often much slower than normal, because weights are loaded, memory is allocated and code is compiled. So discard a warm-up period before measuring, and report percentiles, not just the mean, because the drone’s system must cope with slow cases. The MLPerf Inference rules report latency percentiles for some scenarios, for example the 90th percentile for single stream and the 99th percentile for multistream.

Example 1 The effect of including warm-up

205 latency measurements on a companion computer; the first five are warm-up (simulated data in milliseconds).

import numpy as np

rng = np.random.default_rng(9)
lat = np.r_[rng.normal(180, 20, 5), rng.gamma(20, 1.6, 200) + 10]    # ms
WARMUP = 5
print("first runs:", np.round(lat[:WARMUP]).astype(int))
for name, x in (("including warm-up", lat), ("after warm-up", lat[WARMUP:])):
    p50, p95 = np.percentile(x, [50, 95])
    print(f"{name:<18} n={len(x)}: mean {x.mean():5.1f}, p50 {p50:5.1f}, p95 {p95:5.1f}, max {x.max():6.1f} ms")
p95 = np.percentile(lat[WARMUP:], 95)
print(f"frame rate the model can sustain for 95% of frames: {1000 / p95:.1f} fps")
first runs: [164 185 147 193 203]
including warm-up  n=205: mean  46.1, p50  43.0, p95  59.3, max  202.9 ms
after warm-up      n=200: mean  42.8, p50  42.8, p95  56.4, max   63.4 ms
frame rate the model can sustain for 95% of frames: 17.7 fps

The median barely changes and the mean and p95 rise slightly, but the maximum rises more than threefold because of warm-up. Reporting everything together makes the model look slower than it is, while reporting only the mean after warm-up can hide the slow cases that make the system fall behind. This is still only model time; image reading, resizing and sending results must also be measured to get the end-to-end time.

Latency for each of 205 runs: the first five pink points are high at about 150 to 200 milliseconds and the remaining blue points sit at about 30 to 60 milliseconds, with a gold dashed line at the p95 after warm-up
Figure 1 Per-run latency and warm-up

Shrinking the model with INT8

Quantization replaces 32-bit floating-point weights with 8-bit integers. Jacob et al. (2018) use the relation between the real value , the integer , the scale and the zero point ; symmetric quantization uses . The ONNX Runtime documentation describes symmetric and asymmetric, per-tensor and per-channel quantization. Size falls by about four times, and many kinds of hardware compute INT8 faster, but a large outlier makes the scale coarse, so most values lose more precision.

Example 2 One outlier and INT8 error

10,000 weights from a normal distribution plus one outlier of 0.6. Compare using the true maximum with clipping at ±0.2 before computing the scale.

import numpy as np

rng = np.random.default_rng(1)
w = rng.normal(0, 0.05, 10000)
w[0] = 0.6                                      # outlier
for name, src in (("scale from max |w|", w), ("clip at +/-0.2 first", np.clip(w, -0.2, 0.2))):
    s = np.abs(src).max() / 127
    q = np.clip(np.round(src / s), -127, 127)
    err = q * s - w
    print(f"{name:<22} scale {s:.5f}: mean |error| {np.abs(err[1:]).mean() * 1e3:.3f}e-3, "
          f"outlier error {abs(err[0]):.3f}")
print(f"size: float32 {w.size * 4 / 1024:.1f} KiB -> int8 {w.size / 1024:.1f} KiB")
scale from max |w|     scale 0.00472: mean |error| 1.186e-3, outlier error 0.000
clip at +/-0.2 first   scale 0.00157: mean |error| 0.392e-3, outlier error 0.400
size: float32 39.1 KiB -> int8 9.8 KiB

Using the true maximum makes the error of most weights about three times larger than clipping first, but clipping badly distorts the outlier. There is no single answer: always evaluate model accuracy on the validation set after quantization, and set an acceptance criterion such as AP falling by no more than an agreed amount.

Two groups of bars: the mean error of most weights is higher when the scale comes from the true maximum than when clipping first; the outlier error is much higher when clipping first
Figure 2 INT8 error for two scale choices

Deployment plan for the drone

A good plan states the machine and runtime, the input image size, the maximum end-to-end p95 latency, the minimum accuracy accepted after shrinking the model, power use and heat, and what the system does when processing falls behind, such as skipping frames or storing images for processing after the flight.

Module lab

Lab: benchmarking and shrinking the model

  1. Measure the latency of the model from Module 4 on a PC and on the companion computer, at least 200 runs after warm-up
  2. Report p50, p95 and maximum with Example 1, stating the machine, runtime and image size
  3. Measure end-to-end time from camera capture to result and compare with model time
  4. Export the model to ONNX and apply INT8 quantization, comparing size, time and AP
  5. Write a deployment plan for the drone with acceptance criteria

Common mistakes

Watch out

  • Including warm-up in the measurement
  • Reporting only the mean without percentiles
  • Measuring only model time and calling it end-to-end time
  • Quantizing without re-evaluating accuracy
  • Not stating the machine, runtime and image size, so results cannot be compared

Summary

  • Discard warm-up and report p50, p95 and maximum with the measurement conditions
  • Model time and end-to-end time are different measures
  • INT8 cuts size by about four times, but outliers make the scale coarse; always re-evaluate accuracy
  • A deployment plan needs criteria for time, accuracy and power, and a response when processing falls behind

Check your understanding

  1. With p95 of 50 ms, about what frame rate can be sustained for 95% of frames?
  2. How much memory do one million float32 weights use, and how much in INT8?
  3. With a maximum absolute value of 1.27, what is the symmetric scale?
  4. Why should warm-up not be included?
  5. Why measure end-to-end time as well as model time?
Answers
  1. fps
  2. About 4 MB and 1 MB
  3. The first runs are abnormally slow because of loading and allocation, which distorts the mean and maximum
  4. Reading images, resizing and sending results also take time, possibly more than the model itself

Key formulas

Symmetric quantization
Highest sustainable frame rate

Key references

  1. Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., & Kalenichenko, D. (2018). Quantization and training of neural networks for efficient integer-arithmetic-only inference. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 2704–2713). link
  2. Microsoft. Quantize ONNX models. ONNX Runtime documentation. link
  3. MLCommons. MLPerf inference rules. link
  4. Prince, S. J. D. (2023). Understanding deep learning. MIT Press. link

Further reading

Study the assigned knowledge units in advance, review media and take the module quiz

In class / field

Lab or field practice from worksheets with a safety checklist

Learning evidence: Checked worksheets and quiz results

Module quiz

This is a formative self-check, not a graded exam

Knowledge domain: Artificial intelligence and computer vision