AI in drone work
UAT 306 Artificial Intelligence for UAS
Lesson
By the end of this module you will be able to
- Frame a drone AI problem so that it is measurable, and state the cost of each type of error
- Compute positive predictive value (precision) when the target is rare
- Choose a decision threshold from the costs of misses and false alarms
- Explain where AI can support decisions and where people must check
Why this matters
UAT 315 laid the foundations of AI and computer vision. This course goes deeper into building AI systems that work in practice with drone data. The first step, often skipped, is framing the problem: what to detect, what decision the result supports, and which kind of error costs more. The drone knowledge base’s unit on AI and drones walks through one object-detection project from data rights, dataset splitting and evaluation to PC benchmarking. The whole course uses the case of detecting cracks on the concrete surface of a bridge from drone images. The numbers are hypothetical.
Framing a measurable problem
The same problem can be framed several ways: classifying whole images as cracked or not, detecting crack locations as boxes, or segmenting crack pixels. Each needs different labels and different measures. Before choosing, ask what the user will do with the result. If the result filters images for an engineer to review, missing a crack (a false negative) costs far more than a false alarm (a false positive), because a missed crack may never be seen again, while a false alarm costs only a few seconds of the engineer’s time.
When the target is rare
A detector’s sensitivity and specificity do not say “when it reports a crack, how often is there really one?” That value is the positive predictive value (PPV, or precision), and it depends on the prevalence of the target, as explained by Altman and Bland (1994) for medical tests, which follow the same principle.
Example 1 A good detector with low precision
A detector with sensitivity 0.95 and specificity 0.98 is used on 1,000 images.
SE, SP, N = 0.95, 0.98, 1000
for prevalence in (0.20, 0.05, 0.01):
tp = SE * prevalence * N
fp = (1 - SP) * (1 - prevalence) * N
ppv = tp / (tp + fp)
print(f"crack in {prevalence:4.0%} of images: {tp:5.1f} true alarms, {fp:4.1f} false alarms, precision {ppv:.2f}")
crack in 20% of images: 190.0 true alarms, 16.0 false alarms, precision 0.92
crack in 5% of images: 47.5 true alarms, 19.0 false alarms, precision 0.71
crack in 1% of images: 9.5 true alarms, 19.8 false alarms, precision 0.32
The same detector has high precision when cracks are common, but when cracks appear in only 1% of images, most alerts are false. Accuracy figures advertised by vendors must be questioned for the prevalence of the dataset they were measured on, and evaluated at the prevalence of the real job.
Choosing a threshold from costs
Most detectors output a confidence score, and we choose the threshold. A low threshold catches every crack but raises many false alarms; a high one raises few false alarms but misses more. The best threshold depends on the cost of each error, not the program’s default of 0.5.
Example 2 The threshold with the lowest total cost
A validation set of 1,000 images contains 60 with cracks. The team judges that missing one crack costs as much as reviewing 50 false-alarm images (simulated scores).
import numpy as np
rng = np.random.default_rng(11)
pos = rng.normal(0.72, 0.12, 60).clip(0, 1) # scores of images with cracks
neg = rng.normal(0.35, 0.15, 940).clip(0, 1) # scores of normal images
C_FN, C_FP = 50, 1
for t in (0.3, 0.4, 0.5, 0.6, 0.7):
fn, fp = int((pos < t).sum()), int((neg >= t).sum())
tp = len(pos) - fn
cost = C_FN * fn + C_FP * fp
print(f"threshold {t}: missed {fn:>2}, false alarms {fp:>3}, recall {tp / len(pos):.2f}, "
f"precision {tp / max(tp + fp, 1):.2f}, cost {cost}")
threshold 0.3: missed 0, false alarms 607, recall 1.00, precision 0.09, cost 607
threshold 0.4: missed 0, false alarms 363, recall 1.00, precision 0.14, cost 363
threshold 0.5: missed 2, false alarms 155, recall 0.97, precision 0.27, cost 255
threshold 0.6: missed 9, false alarms 45, recall 0.85, precision 0.53, cost 495
threshold 0.7: missed 26, false alarms 11, recall 0.57, precision 0.76, cost 1311
A threshold of 0.5 gives the lowest total cost of the values tried, even though precision is low, because missing a crack is so much more expensive. If the cost ratio changes, so does the best threshold. The team must agree the costs with the user first and then choose the threshold on the validation set, not the test set.
AI supports decisions; it does not make them
In structural inspection, AI fits best as a screening tool that reduces the number of images people must review. Deciding whether a structure is safe remains the job of a qualified engineer. A good system shows its confidence, marks locations in the image for checking, and records the model version used, so that questions can be traced back.
Module lab
Lab: framing an AI problem for a real job
- Choose one inspection job and write what the user will do with the AI’s results
- Choose a problem type (classification, detection or segmentation) with reasons
- Estimate the prevalence of the target in the real job and compute the expected precision with Example 1
- Agree the cost ratio of misses to false alarms with the user and choose a threshold with Example 2
- Write the scope of what the AI may decide and what people must check
Common mistakes
Watch out
- Collecting data before deciding what the results are for
- Using accuracy on imbalanced data, which looks better than it is
- Trusting precision from a dataset whose prevalence differs from the real job
- Using the default threshold of 0.5 without considering costs
- Letting AI decide safety in place of experts
Summary
- Frame the problem from what the user will do with the result, and state the cost of each type of error
- Precision depends on prevalence; a good detector can raise mostly false alarms when the target is rare
- Choose the threshold that minimises total cost, on the validation set
- AI is a screening tool and assistant; the final decision belongs to experts
Check your understanding
- With sensitivity 0.9, specificity 0.95 and prevalence 10%, what is the precision?
- Why can accuracy mislead on imbalanced data?
- As the threshold is lowered, how do recall and precision usually change?
- If misses are far more expensive than false alarms, which way should the threshold move?
- Why choose the threshold on the validation set rather than the test set?
Answers
- If cracks appear in 1% of images, a model that always answers “no crack” scores 99% accuracy
- Recall rises and precision usually falls
- Lower, to catch more
- The test set must be kept for the final evaluation; using it to choose values makes the result look better than it is
Key formulas
| Positive predictive value (precision) from prevalence | |
| Total cost at threshold t |
Key references
Further reading
Study the assigned knowledge units in advance, review media and take the module quiz
In class / field
Lab or field practice from worksheets with a safety checklist
Learning evidence: Checked worksheets and quiz results