Module 5/5 · Weeks 13–15 · 27 h

Experiments and interpreting test results

UAT 106 Statistics and Data Analytics for Technology

About 90 minDraft, awaiting reviewLast updated 26 September 2026

Lesson

By the end of this module you will be able to

  1. Separate requirement, acceptance criterion, procedure and evidence, and order UAV test stages
  2. Design experiments using randomisation, replication and blocking
  3. Analyse a 2² factorial experiment for main effects and interaction
  4. Evaluate Type A and Type B measurement uncertainty, and report pass rates with confidence intervals

Prerequisites: UAT 106 modules 3–4

Why this matters

Good data comes not from clever analysis but from good experimental design before anyone flies. Fly all of brand A in the cool morning and all of brand B in the hot afternoon, and the measured difference may be the temperature, not the brand, and no analysis can separate them. This module brings the whole course together into a testing process that others can check.

From requirement to evidence

The drone knowledge hub’s unit on UAV testing and evidence analysis separates four things.

What to writeExample
RequirementEndurance no less than 20 minutes with a 1.0 kg payload at 25–35 °C
Acceptance criterionThe one-sided 95% lower bound of the mean exceeds 20 minutes, from at least 8 flights
ProcedureStart timing when the skids leave the ground, stop when the battery reaches 20%, hover at 10 m
EvidenceEvery flight log, the results table, the analysis code and the weather

Write the acceptance criterion before seeing results. Testing moves from low risk to high risk, with each stage answering a different question; simulation results are not evidence about the real aircraft.

Four boxes in a row: simulation SITL, bench test, integration test and controlled field test. A band below reads every stage: requirement, acceptance criterion, procedure, evidence
Figure 1 Stages of UAV testing

Principles of experimental design

Box, Hunter and Hunter’s Statistics for Experimenters stresses three principles:

  • Randomisation: randomise the run order so factors that change over time, such as temperature or pilot skill, do not mix with the factor being studied
  • Replication: repeat each condition several times with independent units to estimate variability
  • Blocking: if testing spans several days, run every condition each day and account for the day in the analysis
import numpy as np

plan = [(payload, prop) for payload in (0.5, 1.0) for prop in ("standard", "efficient") for _ in range(3)]
rng = np.random.default_rng(2026)
order = rng.permutation(len(plan))
for run, idx in enumerate(order[:5], start=1):
    print(run, plan[idx])
print("...", len(plan), "runs in total")
1 (1.0, 'standard')
2 (0.5, 'standard')
3 (1.0, 'efficient')
4 (1.0, 'standard')
5 (1.0, 'efficient')
... 12 runs in total

The 2² factorial experiment

Changing one factor at a time misses something important: factors can act together. A factorial experiment tests every combination of factor levels at once. Here there are two factors with two levels each, payload 0.5 or 1.0 kg and a standard or efficient propeller, with three repeats per corner, 12 flights in all. The data is in factorial.csv.

A square showing a factorial experiment. The horizontal axis is payload, 0.5 and 1.0 kilograms. The top half is the efficient propeller and the bottom half the standard propeller. The corner means are 24.82 at top left, 20.39 at top right, 23.16 at bottom left and 20.05 at bottom right
Figure 2 A 2² factorial experiment with the mean at each corner (min)
import pandas as pd

fx = pd.read_csv("factorial.csv")
cell = fx.groupby(["payload_kg", "propeller"])["endurance_min"].mean()
print(cell.round(3).to_string())

light_std, light_eff = cell[(0.5, "standard")], cell[(0.5, "efficient")]
heavy_std, heavy_eff = cell[(1.0, "standard")], cell[(1.0, "efficient")]
payload_effect = (heavy_std + heavy_eff) / 2 - (light_std + light_eff) / 2
prop_effect = (light_eff + heavy_eff) / 2 - (light_std + heavy_std) / 2
interaction = ((heavy_eff - heavy_std) - (light_eff - light_std)) / 2
print(f"payload effect {payload_effect:+.2f} min   propeller effect {prop_effect:+.2f} min   interaction {interaction:+.2f} min")
payload_kg  propeller
0.5         efficient    24.817
            standard     23.160
1.0         efficient    20.393
            standard     20.053
payload effect -3.76 min   propeller effect +1.00 min   interaction -0.66 min

The main effect of payload is a drop of about 3.8 minutes going from 0.5 to 1.0 kg, and the efficient propeller adds about 1 minute on average. But the negative interaction says the efficient propeller helps a lot with a light payload (1.66 minutes) and little with a heavy one (0.34 minutes). Testing the propeller at 0.5 kg alone would overstate its benefit for heavy delivery missions. These differences still need comparing with the spread between repeats before calling them significant.

Measurement uncertainty

Every measurement is uncertain. The GUM (JCGM 100:2008) describes two ways to evaluate uncertainty:

  • Type A: statistical analysis of repeated observations, such as the standard error of a mean
  • Type B: any other means, such as instrument resolution or a calibration certificate. If an instrument reads to 0.1 s, the true value is equally likely anywhere within ±0.05 s of the reading, so a rectangular distribution gives

Example 1 Uncertainty of a hover time

Five testers time the same hover with stopwatches reading to 0.1 s.

import math
import statistics

readings = [1212.4, 1212.9, 1212.1, 1212.6, 1212.5]
mean = statistics.mean(readings)
u_a = statistics.stdev(readings) / math.sqrt(len(readings))
u_b = 0.05 / math.sqrt(3)
u_c = math.sqrt(u_a ** 2 + u_b ** 2)
print(f"mean {mean:.2f} s   u_A {u_a:.3f} s   u_B {u_b:.3f} s   u_c {u_c:.3f} s")
print(f"result {mean:.2f} s ± {2 * u_c:.2f} s (k = 2)")
mean 1212.50 s   u_A 0.130 s   u_B 0.029 s   u_c 0.134 s
result 1212.50 s ± 0.27 s (k = 2)

The uncertainty from human timing (Type A) is much larger than the stopwatch resolution (Type B), so a finer stopwatch would barely help; the method must change, for example by taking times from the flight controller log. When reporting expanded uncertainty under the GUM, always state the coverage factor ; this example uses .

Reporting pass rates within their limits

The drone knowledge hub’s unit on exercises and evaluating simulated missions warns that pass criteria must be defined before testing, and runs that did not finish must be reported separately, never deleted. A pass rate from a limited number of trials is uncertain too.

from scipy import stats

passed, total = 80, 100
ci = stats.binomtest(passed, total).proportion_ci(confidence_level=0.95)
print(f"pass rate {passed / total:.0%}   95% CI (Clopper-Pearson) [{ci.low:.3f}, {ci.high:.3f}]")
ci_small = stats.binomtest(8, 10).proportion_ci(confidence_level=0.95)
print(f"8 of 10 runs: 95% CI [{ci_small.low:.3f}, {ci_small.high:.3f}]")
pass rate 80%   95% CI (Clopper-Pearson) [0.708, 0.873]
8 of 10 runs: 95% CI [0.444, 0.975]

An 80% pass rate from 100 runs has a confidence interval of about 71% to 87%, but 8 out of 10 runs gives about 44% to 97%. The same 80% means very different things. And the pass rate applies only to the model and conditions tested, not to the chance of success in the field.

What a good test report contains

Test ID and requirement, the acceptance criterion written before testing, environment, settings and firmware version, raw data with units, analysis method and assumptions, sample size, confidence intervals and uncertainty, excluded data with reasons, and conclusions that go no further than the data.

Module lab

Lab: designing and analysing a propeller test

  1. Write a test plan using the table in this lesson (requirement, acceptance criterion, procedure, evidence) for comparing two propellers, and have a classmate review it before opening the data.
  2. Analyse factorial.csv for main effects, interaction and the standard deviation between repeats at each corner, and judge which effects clearly exceed the repeat-to-repeat variation.
  3. Draw an interaction plot (payload on the horizontal axis, one line per propeller) and explain what non-parallel lines mean.
  4. Evaluate both Type A and Type B uncertainty for endurance timing in your own test, and write a complete one-page test report.

Common mistakes

Watch out

  • Setting the acceptance criterion after seeing results
  • Not randomising the run order, so time-varying factors mix with the factor studied
  • Changing one factor at a time, which hides interactions
  • Reporting a measurement without uncertainty, or without stating
  • Silently removing failed runs, or reporting a pass rate without the number of runs

Summary

  • Separate requirement, acceptance criterion, procedure and evidence; write criteria before seeing results; test from low to high risk
  • Randomisation, replication and blocking make experimental results trustworthy
  • Factorial experiments reveal both main effects and interactions, which one-factor-at-a-time testing cannot
  • Report both Type A and Type B uncertainty, and give pass rates with the number of runs and a confidence interval

Check your understanding

  1. Why randomise the flight order when testing two battery brands?
  2. The four corner means are A−B− = 20, A+B− = 24, A−B+ = 22, A+B+ = 30. What is the main effect of A?
  3. From question 2, what is the AB interaction?
  4. A scale reads to 0.01 kg. What is the Type B uncertainty from its resolution?
  5. A system passes 20 of 20 runs. Can we conclude that “it always passes, 100%”?
Answers
  1. So that time-varying factors, such as temperature or pilot skill, do not mix with the brand effect
  2. kg, so kg
  3. No: the 95% confidence interval of the pass rate still has a lower bound of about 83%, and the result holds only for the conditions tested

Key formulas

Main effect of factor A in a 2² factorial
AB interaction
Type B uncertainty from a rectangular distribution
Combined and expanded uncertainty

Key references

  1. Box, G. E. P., Hunter, J. S., & Hunter, W. G. (2005). Statistics for experimenters: Design, innovation, and discovery (2nd ed.). Wiley. link
  2. Montgomery, D. C., & Runger, G. C. (2018). Applied statistics and probability for engineers (7th ed.). Wiley. link
  3. JCGM. (2008). Evaluation of measurement data — Guide to the expression of uncertainty in measurement (JCGM 100:2008). BIPM. link
  4. JCGM. (2012). International vocabulary of metrology — Basic and general concepts and associated terms (VIM, JCGM 200:2012). BIPM. link
  5. NIST/SEMATECH. e-Handbook of statistical methods. National Institute of Standards and Technology. link
  6. The SciPy community. Statistical functions (scipy.stats), SciPy 1.18 documentation. link
  7. Wasserstein, R. L., & Lazar, N. A. (2016). The ASA statement on p-values: Context, process, and purpose. The American Statistician, 70(2), 129–133. link

Further reading

Study the assigned knowledge units in advance, review media and take the module quiz

In class / field

Lab or field practice from worksheets with a safety checklist

Learning evidence: Checked worksheets and quiz results

Module quiz

This is a formative self-check, not a graded exam

Knowledge domain: Mathematics, physics and statistics · Installation, maintenance and testing · Automation, robotics and swarms · Aircraft, structures and design