Model evaluation and edge
UAT 306 Artificial Intelligence for UAS
Lesson
By the end of this module you will be able to
- Measure latency correctly, excluding warm-up and reporting percentiles
- Separate model time from end-to-end time from image capture to result
- Explain INT8 quantization and its effect on size and error
- Plan deployment on a companion computer with acceptance criteria
Why this matters
A model that is accurate on a PC may be too slow on a companion computer, or too large for its memory. The drone knowledge base’s unit on benchmarking and edge deployment stresses that model inference time and time from image capture to result are different measures; you must state what is timed, the warm-up, the number of runs, the machine, the runtime and the image size; and p95 means 95% of samples finish within that time. UAT 315 computed how far the drone moves during processing; this module focuses on how to measure and how to make models smaller.
Measuring time correctly
The first few runs of a model are often much slower than normal, because weights are loaded, memory is allocated and code is compiled. So discard a warm-up period before measuring, and report percentiles, not just the mean, because the drone’s system must cope with slow cases. The MLPerf Inference rules report latency percentiles for some scenarios, for example the 90th percentile for single stream and the 99th percentile for multistream.
Example 1 The effect of including warm-up
205 latency measurements on a companion computer; the first five are warm-up (simulated data in milliseconds).
import numpy as np
rng = np.random.default_rng(9)
lat = np.r_[rng.normal(180, 20, 5), rng.gamma(20, 1.6, 200) + 10] # ms
WARMUP = 5
print("first runs:", np.round(lat[:WARMUP]).astype(int))
for name, x in (("including warm-up", lat), ("after warm-up", lat[WARMUP:])):
p50, p95 = np.percentile(x, [50, 95])
print(f"{name:<18} n={len(x)}: mean {x.mean():5.1f}, p50 {p50:5.1f}, p95 {p95:5.1f}, max {x.max():6.1f} ms")
p95 = np.percentile(lat[WARMUP:], 95)
print(f"frame rate the model can sustain for 95% of frames: {1000 / p95:.1f} fps")
first runs: [164 185 147 193 203]
including warm-up n=205: mean 46.1, p50 43.0, p95 59.3, max 202.9 ms
after warm-up n=200: mean 42.8, p50 42.8, p95 56.4, max 63.4 ms
frame rate the model can sustain for 95% of frames: 17.7 fps
The median barely changes and the mean and p95 rise slightly, but the maximum rises more than threefold because of warm-up. Reporting everything together makes the model look slower than it is, while reporting only the mean after warm-up can hide the slow cases that make the system fall behind. This is still only model time; image reading, resizing and sending results must also be measured to get the end-to-end time.
Shrinking the model with INT8
Quantization replaces 32-bit floating-point weights with 8-bit integers. Jacob et al. (2018) use the relation between the real value , the integer , the scale and the zero point ; symmetric quantization uses . The ONNX Runtime documentation describes symmetric and asymmetric, per-tensor and per-channel quantization. Size falls by about four times, and many kinds of hardware compute INT8 faster, but a large outlier makes the scale coarse, so most values lose more precision.
Example 2 One outlier and INT8 error
10,000 weights from a normal distribution plus one outlier of 0.6. Compare using the true maximum with clipping at ±0.2 before computing the scale.
import numpy as np
rng = np.random.default_rng(1)
w = rng.normal(0, 0.05, 10000)
w[0] = 0.6 # outlier
for name, src in (("scale from max |w|", w), ("clip at +/-0.2 first", np.clip(w, -0.2, 0.2))):
s = np.abs(src).max() / 127
q = np.clip(np.round(src / s), -127, 127)
err = q * s - w
print(f"{name:<22} scale {s:.5f}: mean |error| {np.abs(err[1:]).mean() * 1e3:.3f}e-3, "
f"outlier error {abs(err[0]):.3f}")
print(f"size: float32 {w.size * 4 / 1024:.1f} KiB -> int8 {w.size / 1024:.1f} KiB")
scale from max |w| scale 0.00472: mean |error| 1.186e-3, outlier error 0.000
clip at +/-0.2 first scale 0.00157: mean |error| 0.392e-3, outlier error 0.400
size: float32 39.1 KiB -> int8 9.8 KiB
Using the true maximum makes the error of most weights about three times larger than clipping first, but clipping badly distorts the outlier. There is no single answer: always evaluate model accuracy on the validation set after quantization, and set an acceptance criterion such as AP falling by no more than an agreed amount.
Deployment plan for the drone
A good plan states the machine and runtime, the input image size, the maximum end-to-end p95 latency, the minimum accuracy accepted after shrinking the model, power use and heat, and what the system does when processing falls behind, such as skipping frames or storing images for processing after the flight.
Module lab
Lab: benchmarking and shrinking the model
- Measure the latency of the model from Module 4 on a PC and on the companion computer, at least 200 runs after warm-up
- Report p50, p95 and maximum with Example 1, stating the machine, runtime and image size
- Measure end-to-end time from camera capture to result and compare with model time
- Export the model to ONNX and apply INT8 quantization, comparing size, time and AP
- Write a deployment plan for the drone with acceptance criteria
Common mistakes
Watch out
- Including warm-up in the measurement
- Reporting only the mean without percentiles
- Measuring only model time and calling it end-to-end time
- Quantizing without re-evaluating accuracy
- Not stating the machine, runtime and image size, so results cannot be compared
Summary
- Discard warm-up and report p50, p95 and maximum with the measurement conditions
- Model time and end-to-end time are different measures
- INT8 cuts size by about four times, but outliers make the scale coarse; always re-evaluate accuracy
- A deployment plan needs criteria for time, accuracy and power, and a response when processing falls behind
Check your understanding
- With p95 of 50 ms, about what frame rate can be sustained for 95% of frames?
- How much memory do one million float32 weights use, and how much in INT8?
- With a maximum absolute value of 1.27, what is the symmetric scale?
- Why should warm-up not be included?
- Why measure end-to-end time as well as model time?
Answers
- fps
- About 4 MB and 1 MB
- The first runs are abnormally slow because of loading and allocation, which distorts the mean and maximum
- Reading images, resizing and sending results also take time, possibly more than the model itself
Key formulas
| Symmetric quantization | |
| Highest sustainable frame rate |
Key references
- Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., & Kalenichenko, D. (2018). Quantization and training of neural networks for efficient integer-arithmetic-only inference. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 2704–2713). link
- Microsoft. Quantize ONNX models. ONNX Runtime documentation. link
- MLCommons. MLPerf inference rules. link
- Prince, S. J. D. (2023). Understanding deep learning. MIT Press. link
Further reading
Study the assigned knowledge units in advance, review media and take the module quiz
Evaluating an object-detection model
Edge benchmarking and deployment planning
In class / field
Lab or field practice from worksheets with a safety checklist
Learning evidence: Checked worksheets and quiz results