Module 4/5 · Weeks 10–12 · 27 h

Safety and effectiveness evaluation

UAT 494 Unmanned Aircraft Systems and Automation Technology Capstone Project II

About 90 minDraft, awaiting reviewLast updated 28 September 2026

Lesson

By the end of this module you will be able to

  1. Estimate reliability from test flight hours with MTBF and the exponential model
  2. Calculate how many failure-free tests are needed to demonstrate reliability
  3. Measure usability with the SUS questionnaire
  4. Evaluate effectiveness against the existing method and review risks after testing

Prerequisites: UAT 494 Module 3 · UAT 493 Module 5 (FMEA)

Why this matters

A system that met its requirements in a few tests may still not be ready for others to use. The committee and the users will go on to ask how often it fails, whether it is easy to use, and whether it is really better than the current method. This module gives numerical tools for those questions.

Reliability

When the failure rate is constant, the NIST handbook gives and estimates MTBF as total operating time divided by the number of failures. To demonstrate reliability with tests that have no failures at all, use (Luko, 1997). If tests all succeed, the 95% upper bound on the failure rate is about , the “rule of three” (Hanley and Lippman-Hand, 1983).

Example 1 From test flight hours to reliability

The team flew 42 test hours with 2 failures that aborted the mission. One canal inspection round takes about 20 minutes (hypothetical data).

import math

T_HOURS, FAILURES, MISSION_MIN = 42.0, 2, 20
mtbf = T_HOURS / FAILURES
r_mission = math.exp(-(MISSION_MIN / 60) / mtbf)
print(f"MTBF {mtbf:.1f} h, reliability of one {MISSION_MIN}-min mission {r_mission:.3f}")
print(f"expected aborts in 100 missions: {100 * (1 - r_mission):.1f}")

for r_target, conf in ((0.95, 0.90), (0.95, 0.95), (0.99, 0.95)):
    n = math.ceil(math.log(1 - conf) / math.log(r_target))
    print(f"to show R >= {r_target} with {conf:.0%} confidence: {n} missions with zero failures")
print(f"30 missions with no failure -> failure rate below about {3 / 30:.0%} (rule of three)")
MTBF 21.0 h, reliability of one 20-min mission 0.984
expected aborts in 100 missions: 1.6
to show R >= 0.95 with 90% confidence: 45 missions with zero failures
to show R >= 0.95 with 95% confidence: 59 missions with zero failures
to show R >= 0.99 with 95% confidence: 299 missions with zero failures
30 missions with no failure -> failure rate below about 10% (rule of three)

From the data so far, a 20-minute mission succeeds about 98% of the time. But demonstrating 95% with 90% confidence needs 45 missions without a single failure, which is beyond the project’s time, so the report must state the confidence level the data actually supports.

A reliability curve R of t against continuous flight time from 0 to 10 hours: a blue line falls from 1.0 to about 0.62 at 10 hours, and a pink point near the vertical axis marks the 20-minute flight at about 0.98
Figure 1 Exponential reliability with an MTBF of 21 hours

Is it easy to use?

The System Usability Scale (SUS) by Brooke (1996) is a 10-item questionnaire on a 1–5 scale. Odd items score (response − 1), even items score (5 − response), and the sum is multiplied by 2.5 to give 0–100. Sauro (2011) reports an average of about 68 across 500 studies, and Bangor and colleagues (2008) give guidance for interpreting scores.

Example 2 SUS from five irrigation staff, and effectiveness

import statistics as st

responses = [                      # 10 items, scale 1–5
    [4, 2, 4, 2, 4, 1, 4, 2, 3, 2],
    [5, 1, 4, 2, 4, 2, 5, 1, 4, 3],
    [3, 2, 3, 3, 4, 2, 4, 2, 3, 2],
    [4, 2, 4, 1, 4, 2, 4, 2, 4, 2],
    [4, 3, 3, 2, 3, 2, 4, 3, 3, 2],
]

def sus(r):
    return 2.5 * sum((v - 1) if i % 2 == 0 else (5 - v) for i, v in enumerate(r))

scores = [sus(r) for r in responses]
print("SUS per user:", scores, f"mean {st.mean(scores):.1f}, lowest {min(scores)}")

manual = {"minutes per round": 180, "people": 2}
drone = {"minutes per round": 45, "people": 2}
saved = manual["minutes per round"] * manual["people"] - drone["minutes per round"] * drone["people"]
print(f"staff time saved per 2 km round: {saved} person-minutes ({saved / (manual['minutes per round'] * manual['people']):.0%})")
SUS per user: [75.0, 82.5, 65.0, 77.5, 62.5] mean 72.5, lowest 62.5
staff time saved per 2 km round: 270 person-minutes (75%)

The mean is slightly above the reference average, but two users scored below 68, so interview them to find where they struggled. Five users is still a small sample, and the time saving must state that it excludes preparing flight permission documents and reviewing reports.

A SUS score bar from 0 to 100 with a gold dashed line at 68 labelled Sauro average 68 and a blue point at 72.5 labelled project 72.5
Figure 2 The project's SUS score against the average

Reviewing risk after testing

Test results are new evidence that must feed back into the FMEA from UAT 493. For example, if GNSS dropouts near tall trees happen more often than expected, raise the occurrence score and add controls. Before operating the system for real outside the test area, assess it with an operational risk framework such as JARUS SORA and under the CAAT rules.

Module lab

Lab: post-test evaluation

  1. Total all test flight hours and failures and compute MTBF with Example 1
  2. Compute the number of rounds needed to demonstrate reliability, and write down the level the data actually supports
  3. Have at least five real users try the system, complete the SUS questionnaire and a short interview
  4. Measure time against the existing method systematically, including hidden time such as preparation and report review
  5. Review the FMEA with the test data and record what changed

Common mistakes

Watch out

  • Claiming high reliability from a few tests
  • Scoring SUS incorrectly by not reversing the even items
  • Letting the development team answer SUS themselves
  • Comparing effectiveness while leaving out time that does not favour the new system
  • Not updating the FMEA with the test results

Summary

  • MTBF = total time / number of failures, and R(t) = e^(−t/MTBF) when the failure rate is constant
  • Demonstrating reliability needs many tests,
  • SUS gives a 0–100 score compared with a reference average of about 68
  • Evaluate effectiveness completely and feed test results back into the risk review

Check your understanding

  1. 30 flight hours with 3 failures. What is the MTBF?
  2. MTBF 20 hours and a 1-hour mission. What is the reliability?
  3. How many failure-free tests demonstrate R ≥ 0.90 at 90% confidence?
  4. What SUS score results if every item is answered 3?
  5. 60 tests with no failures. What is the approximate 95% upper bound on the failure rate?
Answers
  1. hours
  2. , so 22 tests

Key formulas

MTBF and reliability
Number of failure-free tests
SUS score

Key references

  1. NIST/SEMATECH. Exponential distribution (section 8.1.6.1). e-Handbook of statistical methods. link
  2. NIST/SEMATECH. Estimating the exponential parameter (section 8.4.5.1). e-Handbook of statistical methods. link
  3. Luko, S. N. (1997). Attribute reliability and the success run: A review (SAE Technical Paper 972753). SAE International. link
  4. Hanley, J. A., & Lippman-Hand, A. (1983). If nothing goes wrong, is everything all right? Interpreting zero numerators. JAMA, 249(13), 1743–1745. link
  5. Brooke, J. (1996). SUS: A 'quick and dirty' usability scale. In P. W. Jordan, B. Thomas, B. A. Weerdmeester, & I. L. McClelland (Eds.), Usability evaluation in industry (pp. 189–194). Taylor & Francis. link
  6. Brooke, J. (2013). SUS: A retrospective. Journal of Usability Studies, 8(2), 29–40. link
  7. Sauro, J. (2011, February 3). Measuring usability with the System Usability Scale (SUS). MeasuringU. link
  8. Bangor, A., Kortum, P. T., & Miller, J. T. (2008). An empirical evaluation of the System Usability Scale. International Journal of Human–Computer Interaction, 24(6), 574–594. link
  9. Joint Authorities for Rulemaking on Unmanned Systems. (2024). JARUS guidelines on Specific Operations Risk Assessment (SORA), main body, edition 2.5 (JAR-DEL-SRM-SORA-MB-2.5). link

Further reading

Study the assigned knowledge units in advance, review media and take the module quiz

In class / field

Team project work, advisor meetings and progress presentations

Learning evidence: Project milestone deliverables

Module quiz

This is a formative self-check, not a graded exam

Knowledge domain: Law, safety and risk