Module 4/5 · Weeks 10–12 · 27 h

Datasets and evaluation

UAT 315 Artificial Intelligence, Data Analytics and Computer Vision for Unmanned Aircraft Systems

About 90 minDraft, awaiting reviewLast updated 27 September 2026

Lesson

By the end of this module you will be able to

  1. Explain data leakage when frames from one flight appear in both training and test sets, and split by group with GroupShuffleSplit
  2. Calculate the confusion matrix, precision, recall, F1 and IoU, and read them in the context of the task
  3. Explain the difference between COCO mAP50 and mAP50–95
  4. Choose a confidence threshold by the harm of each error, and write a data card

Prerequisites: UAT 315 modules 2–3

Why this matters

A team reports that its victim-detection model is 99% accurate, yet flying over an unfamiliar forest, the system misses almost everyone. The cause is not the model but how the data was split and how results were measured. A flawed evaluation misleads the whole team and can lead commanders to deploy a system that is not ready. The drone knowledge hub’s units on checking a dataset before training AI and on image quality and AI metrics underpin this module.

Leakage from frames of the same flight

Frames from the same flight are very alike. Split randomly by frame, and near-identical frames end up in both training and test sets; the model scores well by recognising the scene, not by learning what we want. This is data leakage. The fix is to split by group, keeping each flight in one set. If the model must work in new areas, split by area as well.

Two columns. On the left, a random split by frame: dots of the same colour, from the same flight, are scattered across train, validation and test. On the right, a split by flight: each colour sits in one set only. Below: same colour means frames from the same flight
Figure 1 Splitting by flight to prevent leakage

Example 1 Inflated scores from a random split

Synthetic data from 40 flights of 30 frames each. Each flight has its own scene “signature”, and the whole flight shares one label, with only a weak real signal related to the label. The model is a 1-nearest-neighbour (1-NN) classifier, which predicts from the most similar training example.

import numpy as np
from sklearn.model_selection import GroupShuffleSplit, train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score

rng = np.random.default_rng(345)
flights, frames = 40, 30
signature = rng.normal(0, 1, (flights, 5))
flight_label = rng.integers(0, 2, flights)
X = np.repeat(signature, frames, axis=0) + rng.normal(0, 0.1, (flights * frames, 5))
X[:, 0] += 0.3 * np.repeat(flight_label, frames)
y = np.repeat(flight_label, frames)
groups = np.repeat(np.arange(flights), frames)

knn = KNeighborsClassifier(n_neighbors=1)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.25, random_state=0)
print(f"random split by frame:  accuracy {accuracy_score(y_te, knn.fit(X_tr, y_tr).predict(X_te)):.3f}")
train_idx, test_idx = next(GroupShuffleSplit(test_size=0.25, random_state=0).split(X, y, groups))
print(f"split by flight:        accuracy {accuracy_score(y[test_idx], knn.fit(X[train_idx], y[train_idx]).predict(X[test_idx])):.3f}")
print("flights shared between train and test:", len(set(groups[train_idx]) & set(groups[test_idx])))
random split by frame:  accuracy 1.000
split by flight:        accuracy 0.300
flights shared between train and test: 0

The random split scores 100% because every test frame has a twin from the same flight in the training set. Split by flight, the score falls to just 30%, no better than guessing in a two-class problem (with only 10 test flights, this figure swings a lot). The second figure is the real ability on new flights, which here is almost none.

The confusion matrix and metrics

A confusion matrix counts four outcomes: TP (present, and the system says present), FP (absent but flagged, a false alarm), FN (present but missed) and TN (absent, and the system says absent). From these we compute:

  • Precision: of what the system flagged, how much was right
  • Recall: of what was really there, how much the system found
  • F1: the harmonic mean of the two

This example uses the counts from the drone knowledge hub’s unit on computer vision and edge AI: TP = 8, FP = 2, FN = 4.

from sklearn.metrics import confusion_matrix, precision_recall_fscore_support

y_true = np.array([1] * 12 + [0] * 8)
y_pred = np.array([1] * 8 + [0] * 4 + [1] * 2 + [0] * 6)
tn, fp, fn, tp = confusion_matrix(y_true, y_pred).ravel()
p, r, f1, _ = precision_recall_fscore_support(y_true, y_pred, average="binary")
print(f"TP {tp}  FP {fp}  FN {fn}  TN {tn}")
print(f"precision {p:.4f}  recall {r:.4f}  F1 {f1:.4f}")
TP 8  FP 2  FN 4  TN 6
precision 0.8000  recall 0.6667  F1 0.7273

In a search for victims, a recall of 0.67 means missing one person in three, which is far worse than a false alarm that sends someone to check. A model’s confidence score is not accuracy either: a score of 0.9 on one image does not mean the model is right 90% of the time on all images.

IoU and mAP

In object detection, predicted boxes must first be matched one-to-one with true boxes, and a match only counts as TP when its IoU reaches the threshold. Two boxes on the same object count as only one TP.

A blue box, the truth A, and a dashed pink box, the prediction B, partly overlap. The overlap is shaded gold and labelled 60. On the right, three boxes: intersection equals 60 px squared; union equals 100 plus 100 minus 60 equals 140; IoU equals 60 over 140, about 0.43, less than 0.5
Figure 2 Intersection over union of two boxes

AP (average precision) summarises how precision and recall trade off as the confidence threshold changes, and mAP averages AP over classes. COCO evaluation uses AP averaged over IoU thresholds from 0.50 to 0.95 in steps of 0.05 as its primary metric, while AP50 uses a single IoU of 0.50, which is much more lenient. The two numbers cannot be compared directly.

iou_thresholds = np.round(np.arange(0.50, 0.96, 0.05), 2)
print("COCO IoU thresholds:", iou_thresholds.tolist(), "count", len(iou_thresholds))
ap_per_class = {"person": 0.6, "vehicle": 0.8}
print("mAP =", round(sum(ap_per_class.values()) / len(ap_per_class), 2))
COCO IoU thresholds: [0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, 0.95] count 10
mAP = 0.7

The per-class AP values are hypothetical figures from the knowledge hub. AP cannot be computed from counts at a single threshold; it needs an evaluator that ranks every confidence score. When reporting, name the metric precisely, such as “mAP50–95 on a test set split by area”.

Choosing a threshold by harm

scores = np.array([0.95, 0.91, 0.88, 0.84, 0.80, 0.74, 0.69, 0.62, 0.55, 0.48,
                   0.44, 0.39, 0.33, 0.28, 0.22, 0.18, 0.12, 0.09, 0.05, 0.03])
labels = np.array([1, 1, 1, 0, 1, 1, 0, 1, 1, 0, 1, 0, 1, 0, 0, 1, 0, 0, 0, 0])
for threshold in (0.7, 0.5, 0.3):
    pred = (scores >= threshold).astype(int)
    p, r, f1, _ = precision_recall_fscore_support(labels, pred, average="binary")
    print(f"threshold {threshold}: alerts {pred.sum():>2}  precision {p:.2f}  recall {r:.2f}  F1 {f1:.2f}")
threshold 0.7: alerts  6  precision 0.83  recall 0.50  F1 0.62
threshold 0.5: alerts  9  precision 0.78  recall 0.70  F1 0.74
threshold 0.3: alerts 13  precision 0.69  recall 0.90  F1 0.78

A lower threshold finds more real objects but raises more false alarms. No threshold is best for every task: choose it from the harm of each kind of error and how many alerts people can check, and choose it on the validation set, not the test set.

Data card

A data card documents a dataset: the source and usage rights of the images, how they were collected and labelled, the definition of each class, how the data was split and by what unit, the number of examples per class, the conditions covered and not covered (for example, no night images), and known limits. Anyone using the model should be able to see which situations its results apply to.

Class activity

Activity: auditing a model report

  1. The instructor hands out a mock report claiming “98% accuracy”. Each group lists at least five questions to ask before believing it, such as how the data was split, which metric was used and where the test set came from.
  2. Do the drone knowledge hub’s exercise on checking a dataset before training AI: find the five kinds of error in the manifest before running the checking program.
  3. For a victim search and for tree counting, choose a threshold from the table in this lesson, with reasons.
  4. Write a one-page data card for an image set your group might collect in future.

Common mistakes

Watch out

  • Splitting randomly by frame when frames come from the same flight or area
  • Reporting accuracy alone in detection tasks where real objects are rare
  • Comparing mAP50 with mAP50–95 as if they were the same number
  • Choosing the confidence threshold on the test set
  • Having no data card, so users never learn the model has not seen night images or other areas

Summary

  • Frames from one flight must stay in one set, or scores will be inflated
  • Precision, recall and F1 answer different questions; choose metrics by the harm of each error
  • Detection matches boxes with IoU, and mAP50 and mAP50–95 are different metrics
  • Choose thresholds on the validation set, and document data with a data card

Check your understanding

  1. With TP = 30, FP = 10 and FN = 20, what are precision and recall?
  2. If FP drops to 0 while TP = 8 and FN = 4, what is F1?
  3. Why does a random split by frame give inflated scores?
  4. How many IoU levels does COCO mAP50–95 average over?
  5. Which task should favour recall over precision: “search for victims” or “choose good-looking images for publicity”?
Answers
  1. Precision and recall
  2. Near-identical frames from the same flight land in both training and test sets, so the model recognises the scene
  3. Ten levels: 0.50, 0.55, …, 0.95
  4. Searching for victims, because missing a person who needs help is far worse than a false alarm

Key formulas

Precision
Recall
F1
mAP

Key references

  1. scikit-learn developers. Cross-validation: Evaluating estimator performance (scikit-learn 1.9). link
  2. scikit-learn developers. Metrics and scoring: Quantifying the quality of predictions (scikit-learn 1.9). link
  3. COCO Consortium. Detection evaluation. Common Objects in Context. link
  4. Ultralytics. Models supported by Ultralytics (YOLO11, YOLO26). Ultralytics documentation. link
  5. Géron, A. (2025). Hands-on machine learning with Scikit-Learn and PyTorch. O'Reilly. link
  6. National Institute of Standards and Technology. (2023). Artificial intelligence risk management framework (AI RMF 1.0) (NIST AI 100-1). link

Further reading

Study the assigned knowledge units in advance, review media and take the module quiz

AvailableD11

Checking datasets before AI training

ฝึกตรวจข้อมูลรั่วจากเที่ยวบินเดียวกัน ภาพซ้ำ และข้อมูลไม่ครบด้วย CSV และ Python ก่อนเริ่มเทรนโมเดลตรวจจับวัตถุ
AvailableD18D11

Checking image quality and reading AI metrics

แยกภาพที่ยังตัดสินไม่ได้ ฝึก precision และ recall จากข้อมูลสมมติ และอธิบายขอบเขตของผลประเมิน
AvailableD11

Preparing and splitting datasets

โมเดลเรียนได้เท่าที่ข้อมูลและ labels สนับสนุน ต้องกำหนดวัตถุที่จะตรวจ ขอบเขตภาพ และสิทธิ์ก่อนรวบรวม แยกข้อมูลตามกลุ่มที่อาจคล้ายกัน เช่น เที่ยวบินหรือสถานที่ ก่อนใช้ validation เลือกค่าและพัก test ไว้
AvailableD11

Training and evaluating object detection

Object detection ต้องทำนายทั้งชนิดและตำแหน่งวัตถุ IoU คือพื้นที่กรอบที่ทับกันหารพื้นที่รวม ใช้ช่วยจับคู่คำทำนายกับคำตอบ ส่วน mAP ของระบบประเมินยังรวมการจัดอันดับและหลายเงื่อนไข ไม่เท่ากับ accuracy ระดับภาพ

In class / field

Lecture, case discussion and in-class problem solving

Learning evidence: Quiz results and submitted exercises

Module quiz

This is a formative self-check, not a graded exam

Knowledge domain: Artificial intelligence and computer vision · Management, innovation and professional practice · Law, safety and risk