Module 1/5 · Weeks 1–3 · 27 h

AI in drone work

UAT 306 Artificial Intelligence for UAS

About 85 minDraft, awaiting reviewLast updated 28 September 2026

Lesson

By the end of this module you will be able to

  1. Frame a drone AI problem so that it is measurable, and state the cost of each type of error
  2. Compute positive predictive value (precision) when the target is rare
  3. Choose a decision threshold from the costs of misses and false alarms
  4. Explain where AI can support decisions and where people must check

Prerequisites: UAT 315 (AI, data analysis and computer vision) · UAT 106 (statistics)

Why this matters

UAT 315 laid the foundations of AI and computer vision. This course goes deeper into building AI systems that work in practice with drone data. The first step, often skipped, is framing the problem: what to detect, what decision the result supports, and which kind of error costs more. The drone knowledge base’s unit on AI and drones walks through one object-detection project from data rights, dataset splitting and evaluation to PC benchmarking. The whole course uses the case of detecting cracks on the concrete surface of a bridge from drone images. The numbers are hypothetical.

Framing a measurable problem

The same problem can be framed several ways: classifying whole images as cracked or not, detecting crack locations as boxes, or segmenting crack pixels. Each needs different labels and different measures. Before choosing, ask what the user will do with the result. If the result filters images for an engineer to review, missing a crack (a false negative) costs far more than a false alarm (a false positive), because a missed crack may never be seen again, while a false alarm costs only a few seconds of the engineer’s time.

When the target is rare

A detector’s sensitivity and specificity do not say “when it reports a crack, how often is there really one?” That value is the positive predictive value (PPV, or precision), and it depends on the prevalence of the target, as explained by Altman and Bland (1994) for medical tests, which follow the same principle.

Example 1 A good detector with low precision

A detector with sensitivity 0.95 and specificity 0.98 is used on 1,000 images.

SE, SP, N = 0.95, 0.98, 1000
for prevalence in (0.20, 0.05, 0.01):
    tp = SE * prevalence * N
    fp = (1 - SP) * (1 - prevalence) * N
    ppv = tp / (tp + fp)
    print(f"crack in {prevalence:4.0%} of images: {tp:5.1f} true alarms, {fp:4.1f} false alarms, precision {ppv:.2f}")
crack in  20% of images: 190.0 true alarms, 16.0 false alarms, precision 0.92
crack in   5% of images:  47.5 true alarms, 19.0 false alarms, precision 0.71
crack in   1% of images:   9.5 true alarms, 19.8 false alarms, precision 0.32

The same detector has high precision when cracks are common, but when cracks appear in only 1% of images, most alerts are false. Accuracy figures advertised by vendors must be questioned for the prevalence of the dataset they were measured on, and evaluated at the prevalence of the real job.

Precision against crack prevalence from 0 to 20 percent: a blue curve rising from low values, with three pink points at 1 percent precision 0.32, 5 percent 0.71 and 20 percent 0.92
Figure 1 Precision of the same detector by prevalence

Choosing a threshold from costs

Most detectors output a confidence score, and we choose the threshold. A low threshold catches every crack but raises many false alarms; a high one raises few false alarms but misses more. The best threshold depends on the cost of each error, not the program’s default of 0.5.

Example 2 The threshold with the lowest total cost

A validation set of 1,000 images contains 60 with cracks. The team judges that missing one crack costs as much as reviewing 50 false-alarm images (simulated scores).

import numpy as np

rng = np.random.default_rng(11)
pos = rng.normal(0.72, 0.12, 60).clip(0, 1)     # scores of images with cracks
neg = rng.normal(0.35, 0.15, 940).clip(0, 1)    # scores of normal images
C_FN, C_FP = 50, 1
for t in (0.3, 0.4, 0.5, 0.6, 0.7):
    fn, fp = int((pos < t).sum()), int((neg >= t).sum())
    tp = len(pos) - fn
    cost = C_FN * fn + C_FP * fp
    print(f"threshold {t}: missed {fn:>2}, false alarms {fp:>3}, recall {tp / len(pos):.2f}, "
          f"precision {tp / max(tp + fp, 1):.2f}, cost {cost}")
threshold 0.3: missed  0, false alarms 607, recall 1.00, precision 0.09, cost 607
threshold 0.4: missed  0, false alarms 363, recall 1.00, precision 0.14, cost 363
threshold 0.5: missed  2, false alarms 155, recall 0.97, precision 0.27, cost 255
threshold 0.6: missed  9, false alarms  45, recall 0.85, precision 0.53, cost 495
threshold 0.7: missed 26, false alarms  11, recall 0.57, precision 0.76, cost 1311

A threshold of 0.5 gives the lowest total cost of the values tried, even though precision is low, because missing a crack is so much more expensive. If the cost ratio changes, so does the best threshold. The team must agree the costs with the user first and then choose the threshold on the validation set, not the test set.

Bars of total cost at thresholds 0.3 to 0.7: 607, 363, 255, 495 and 1311, with the 0.5 bar in green as the lowest
Figure 2 Total cost by threshold

AI supports decisions; it does not make them

In structural inspection, AI fits best as a screening tool that reduces the number of images people must review. Deciding whether a structure is safe remains the job of a qualified engineer. A good system shows its confidence, marks locations in the image for checking, and records the model version used, so that questions can be traced back.

Module lab

Lab: framing an AI problem for a real job

  1. Choose one inspection job and write what the user will do with the AI’s results
  2. Choose a problem type (classification, detection or segmentation) with reasons
  3. Estimate the prevalence of the target in the real job and compute the expected precision with Example 1
  4. Agree the cost ratio of misses to false alarms with the user and choose a threshold with Example 2
  5. Write the scope of what the AI may decide and what people must check

Common mistakes

Watch out

  • Collecting data before deciding what the results are for
  • Using accuracy on imbalanced data, which looks better than it is
  • Trusting precision from a dataset whose prevalence differs from the real job
  • Using the default threshold of 0.5 without considering costs
  • Letting AI decide safety in place of experts

Summary

  • Frame the problem from what the user will do with the result, and state the cost of each type of error
  • Precision depends on prevalence; a good detector can raise mostly false alarms when the target is rare
  • Choose the threshold that minimises total cost, on the validation set
  • AI is a screening tool and assistant; the final decision belongs to experts

Check your understanding

  1. With sensitivity 0.9, specificity 0.95 and prevalence 10%, what is the precision?
  2. Why can accuracy mislead on imbalanced data?
  3. As the threshold is lowered, how do recall and precision usually change?
  4. If misses are far more expensive than false alarms, which way should the threshold move?
  5. Why choose the threshold on the validation set rather than the test set?
Answers
  1. If cracks appear in 1% of images, a model that always answers “no crack” scores 99% accuracy
  2. Recall rises and precision usually falls
  3. Lower, to catch more
  4. The test set must be kept for the final evaluation; using it to choose values makes the result look better than it is

Key formulas

Positive predictive value (precision) from prevalence
Total cost at threshold t

Key references

  1. Géron, A. (2022). Hands-on machine learning with Scikit-Learn, Keras, and TensorFlow (3rd ed.). O'Reilly.
  2. Prince, S. J. D. (2023). Understanding deep learning. MIT Press. link
  3. Altman, D. G., & Bland, J. M. (1994). Diagnostic tests 2: Predictive values. BMJ, 309(6947), 102. link

Further reading

Study the assigned knowledge units in advance, review media and take the module quiz

In class / field

Lab or field practice from worksheets with a safety checklist

Learning evidence: Checked worksheets and quiz results

Module quiz

This is a formative self-check, not a graded exam

Knowledge domain: Artificial intelligence and computer vision · Programming and digital technology · Surveying, mapping and geoinformatics