Module 3/5 · Weeks 7–9 · 27 h

Confidence intervals and hypothesis tests

UAT 106 Statistics and Data Analytics for Technology

About 90 minDraft, awaiting reviewLast updated 26 September 2026

Lesson

By the end of this module you will be able to

  1. Calculate the standard error and a confidence interval for the mean using the t distribution
  2. Set hypotheses and the test direction from a requirement before seeing data
  3. Test one mean and compare two groups with Welch's t-test, and interpret p-values correctly
  4. Explain type I and type II errors, and separate the confidence interval of a mean from the spread of individual flights

Prerequisites: UAT 106 modules 1–2

Why this matters

A maker states that a drone flies “no less than 20 minutes”. A test unit flies four times and gets a mean of 22.6 minutes. Is that enough to confirm the requirement? Four more flights would give a different mean. Inferential statistics says how uncertain a sample figure is and how far the data supports a conclusion. The drone knowledge hub’s unit on UAV testing and evidence analysis, and its lab guide L14, use these ideas directly.

Standard error and confidence intervals

A sample mean is not the population mean. The standard error (SE) measures how far the sample mean can stray. The confidence interval (CI) for the mean, as given in the NIST statistics handbook, is . It uses the t distribution instead of the normal because is estimated from the sample; with small , is much larger than 1.96.

Example 1 Three flights of 20, 22 and 24 minutes

This example comes from the UAV testing knowledge unit of the drone knowledge hub.

import math
from scipy import stats

x = [20, 22, 24]
n = len(x)
mean = sum(x) / n
s = math.sqrt(sum((v - mean) ** 2 for v in x) / (n - 1))
se = s / math.sqrt(n)
t_crit = stats.t.ppf(0.975, df=n - 1)
print(f"mean {mean}  s {s}  SE {se:.6f}  t {t_crit:.8f}")
print(f"95% CI [{mean - t_crit * se:.2f}, {mean + t_crit * se:.2f}] min")
ci = stats.ttest_1samp(x, popmean=0).confidence_interval(confidence_level=0.95)
print(f"scipy check [{ci.low:.2f}, {ci.high:.2f}]")
mean 22.0  s 2.0  SE 1.154701  t 4.30265273
95% CI [17.03, 26.97] min
scipy check [17.03, 26.97]

The interval runs from 17.03 to 26.97 minutes because there are only three flights and . Using 1.96 instead would shrink it to less than half its proper width ().

“95%” means that if the same experiment were repeated many times, intervals built this way would contain the true mean 95% of the time. It does not mean a 95% chance that the next flight falls inside the interval.

Testing a hypothesis from a requirement

A hypothesis test starts from a null hypothesis , which we look for evidence against, and an alternative hypothesis . Before seeing any data, choose the direction to match the requirement and the significance level , such as 0.05.

Five boxes in a row: write the requirement; set H0, H1 and alpha before data; collect independent flights; compute t and p; report with CI and limits. Below, a pink box reads never change the test direction after seeing results
Figure 1 Steps of a hypothesis test

The requirement “mean endurance no less than 20 minutes with a 1.0 kg payload” needs evidence that the mean is greater than 20, so we set and , a one-sided test.

import pandas as pd

flights = pd.read_csv("endurance.csv")
a_1kg = flights.query("battery == 'A' and payload_kg == 1.0")["endurance_min"]
print("flights:", a_1kg.tolist())
res = stats.ttest_1samp(a_1kg, popmean=20, alternative="greater")
print(f"mean {a_1kg.mean():.3f}  t {res.statistic:.3f}  p {res.pvalue:.4f}  df {res.df}")
se = a_1kg.std() / math.sqrt(len(a_1kg))
print(f"one-sided 95% lower bound {a_1kg.mean() - stats.t.ppf(0.95, len(a_1kg) - 1) * se:.2f} min")
ci = stats.ttest_1samp(a_1kg, popmean=20).confidence_interval(0.95)
print(f"two-sided 95% CI [{ci.low:.2f}, {ci.high:.2f}] min")
flights: [23.92, 22.04, 22.47, 21.83]
mean 22.565  t 5.447  p 0.0061  df 3
one-sided 95% lower bound 21.46 min
two-sided 95% CI [21.07, 24.06] min

The p-value is the probability of a result at least this extreme if were true. Here p = 0.006, below 0.05, so we reject : the data supports a mean above 20 minutes, and the one-sided 95% lower bound is about 21.46 minutes.

An endurance axis from 18 to 25 minutes. A pink dashed line at 20 minutes is the requirement. The top row shows four individual flights between about 21.8 and 23.9 minutes. The bottom row shows the 95% confidence interval of the mean from 21.07 to 24.06 minutes with the mean marked in the middle
Figure 2 The CI of the mean versus individual flights

A passing mean does not mean every flight passes

A confidence interval describes uncertainty in the mean. The question “will 99% of flights exceed 20 minutes?” needs a tolerance interval, a different method in the NIST handbook that needs much more data. Conclusions from four flights also hold only for the battery, payload and weather tested.

Comparing two groups with Welch’s t-test

Next question: “does brand B fly longer than brand A with a 0.5 kg payload?” Use Welch’s t-test, which does not assume the two groups have equal spread. In SciPy you must pass equal_var=False, because the default assumes equal variances.

half = flights.query("payload_kg == 0.5")
a = half.query("battery == 'A'")["endurance_min"]
b = half.query("battery == 'B'")["endurance_min"]
res = stats.ttest_ind(b, a, equal_var=False)
ci = res.confidence_interval(0.95)
print(f"mean B - A = {b.mean() - a.mean():.2f} min   t {res.statistic:.3f}   p {res.pvalue:.3f}   df {res.df:.2f}")
print(f"95% CI of the difference [{ci.low:.2f}, {ci.high:.2f}] min")
mean B - A = 0.70 min   t 1.015   p 0.353   df 5.36
95% CI of the difference [-1.04, 2.44] min

With p = 0.35, above 0.05, we do not reject , but we must not conclude that “the brands are equal”. The confidence interval for the difference, −1.04 to +2.44 minutes, shows that four flights per group are consistent both with B being about a minute worse and with it being almost two and a half minutes better. The honest answer is “the data cannot tell yet”.

Two kinds of error and how to use p-values

true false
Reject Type I error (chance )Correct (power)
Do not reject CorrectType II error
import numpy as np

rng = np.random.default_rng(116)
null_true = rng.normal(20, 0.9, size=(10_000, 4))
p = stats.ttest_1samp(null_true, 20, axis=1, alternative="greater").pvalue
print(f"H0 true: rejected in {np.mean(p < 0.05):.1%} of 10 000 experiments")
for n in (4, 8, 16):
    real = rng.normal(21, 0.9, size=(10_000, n))
    p = stats.ttest_1samp(real, 20, axis=1, alternative="greater").pvalue
    print(f"true mean 21 min, n = {n:>2}: power {np.mean(p < 0.05):.1%}")
H0 true: rejected in 5.0% of 10 000 experiments
true mean 21 min, n =  4: power 52.1%
true mean 21 min, n =  8: power 87.7%
true mean 21 min, n = 16: power 99.5%

When is true, a test at wrongly rejects about 5% of the time, as designed. When the true mean is 21 minutes, four flights detect it in only about half the experiments (52%). More flights raise the power, so plan the number of flights before testing.

The American Statistical Association (ASA) issued a statement on p-values in 2016. Its key principles include: a p-value does not measure the probability that a hypothesis is true; conclusions should not rest only on whether a p-value passes 0.05; reporting must be full and transparent; and a p-value does not measure the size of an effect. Always report the confidence interval of the difference alongside it.

Module lab

Lab: interpreting endurance tests

  1. Following lab guide L14 of the drone knowledge hub, write the requirement, hypotheses, direction and on the worksheet before opening endurance.csv.
  2. Test the “no less than 20 minutes” requirement for every brand and payload pair. Report the mean, SD, n, p-value and one-sided lower bound.
  3. Compare brands A and B at all three payloads with Welch’s t-test, and write conclusions that do not go beyond the data.
  4. Use simulation to find the smallest number of flights that gives 80% power when the true mean is 21 minutes and SD 0.9 minutes.

Common mistakes

Watch out

  • Choosing one-sided or two-sided after seeing results so the p-value passes
  • Using 1.96 with small samples: use the t value at degrees of freedom
  • Concluding “equal” because was not rejected: finding no difference is not evidence of no difference
  • Thinking the p-value is the chance that is true: it is computed assuming is already true
  • Using the CI of the mean to certify every flight: that needs a tolerance interval

Summary

  • SE , and the CI of the mean uses the value at degrees of freedom; small samples give wide intervals
  • Set , , the direction and from the requirement before seeing data
  • Compare two groups with Welch’s t-test, and report the CI of the difference alongside the p-value
  • is the chance of a type I error, power grows with more data, and a p-value says nothing about the size or importance of an effect

Check your understanding

  1. A sample of 9 flights has minutes. What is the SE?
  2. For the requirement “noise no more than 70 dBA”, how should be set?
  3. What do you conclude from p = 0.03 at ?
  4. A test gives p = 0.40 and the author writes “proven to be no different”. What is wrong?
  5. Why must SciPy’s ttest_ind be given equal_var=False for a Welch’s t-test?
Answers
  1. minutes
  2. dBA, because we need evidence that noise is below the limit ()
  3. Reject ; the data supports at the 0.05 level, and the confidence interval should be reported too
  4. Not rejecting is not evidence of equality; the data may simply be too small. Report the confidence interval of the difference
  5. Because the default of ttest_ind is equal_var=True, which assumes the two groups have equal variance

Key formulas

Standard error of the mean
Two-sided confidence interval for the mean
One-sample t statistic
Welch t statistic

Key references

  1. Montgomery, D. C., & Runger, G. C. (2018). Applied statistics and probability for engineers (7th ed.). Wiley. link
  2. NIST/SEMATECH. Confidence limits for the mean (section 1.3.5.2). e-Handbook of statistical methods. link
  3. NIST/SEMATECH. Two-sample t-test for equal means (section 1.3.5.3). e-Handbook of statistical methods. link
  4. NIST/SEMATECH. Tolerance intervals for a normal distribution (section 7.2.6.3). e-Handbook of statistical methods. link
  5. The SciPy community. Statistical functions (scipy.stats), SciPy 1.18 documentation. link
  6. Wasserstein, R. L., & Lazar, N. A. (2016). The ASA statement on p-values: Context, process, and purpose. The American Statistician, 70(2), 129–133. link
  7. Diez, D. M., Çetinkaya-Rundel, M., & Barr, C. D. (2019). OpenIntro statistics (4th ed.). OpenIntro. link

Further reading

Study the assigned knowledge units in advance, review media and take the module quiz

In class / field

Lab or field practice from worksheets with a safety checklist

Learning evidence: Checked worksheets and quiz results

Module quiz

This is a formative self-check, not a graded exam

Knowledge domain: Mathematics, physics and statistics