Module 3/5 · Weeks 7–9 · 27 h

Datasets and labelling

UAT 306 Artificial Intelligence for UAS

About 85 minDraft, awaiting reviewLast updated 28 September 2026

Lesson

By the end of this module you will be able to

  1. Write labelling guidelines that make different annotators agree
  2. Measure annotator agreement with Cohen's kappa and box IoU
  3. Split a dataset by flight while keeping class proportions similar
  4. Check the dataset before training for leakage and wrong labels

Prerequisites: UAT 306 Modules 1–2 · UAT 315 Module 4 (datasets and evaluation)

Why this matters

The drone knowledge base’s unit on preparing and splitting datasets stresses that a model learns only what the data and labels support: define the objects, the image scope and the rights before collecting, and separate data by groups that may be similar, such as flights or sites. UAT 315 showed that random splits give over-optimistic scores. This module goes deeper into two questions: how trustworthy are the labels, and how to split by flight while keeping class proportions similar.

Labelling guidelines

If two annotators label the same image differently, the model learns that confusion too. Good guidelines define each class with images of what does and does not belong, the smallest object to label, how to draw boxes (tight to the object, with or without shadow), and what to do when unsure. The drone knowledge base’s CVAT unit shows how to label, review and export in a training format.

Example 1 How well do two annotators agree?

Two annotators classify 100 image patches as crack, none or stain, and draw boxes for the same three cracks (simulated data). Cohen’s kappa (Cohen, 1960) removes the agreement expected by chance.

a = ["crack"] * 40 + ["none"] * 50 + ["stain"] * 10
b = ["crack"] * 34 + ["none"] * 6 + ["none"] * 44 + ["crack"] * 4 + ["stain"] * 2 + ["stain"] * 8 + ["none"] * 2
n = len(a)
labels = sorted(set(a) | set(b))
po = sum(x == y for x, y in zip(a, b)) / n
pe = sum((a.count(c) / n) * (b.count(c) / n) for c in labels)
print(f"observed agreement {po:.2f}, chance agreement {pe:.2f}, kappa {(po - pe) / (1 - pe):.2f}")

def iou(p, q):                                  # box (x1, y1, x2, y2) in pixels
    ix = max(0, min(p[2], q[2]) - max(p[0], q[0]))
    iy = max(0, min(p[3], q[3]) - max(p[1], q[1]))
    inter = ix * iy
    union = (p[2] - p[0]) * (p[3] - p[1]) + (q[2] - q[0]) * (q[3] - q[1]) - inter
    return inter / union

pairs = [((100, 100, 200, 160), (104, 98, 206, 158)), ((50, 50, 90, 80), (60, 55, 110, 85)),
         ((300, 200, 340, 260), (330, 240, 380, 300))]
for k, (p, q) in enumerate(pairs, 1):
    print(f"box {k}: IoU between annotators {iou(p, q):.2f}")
observed agreement 0.86, chance agreement 0.42, kappa 0.76
box 1: IoU between annotators 0.85
box 2: IoU between annotators 0.38
box 3: IoU between annotators 0.04

Agreement of 86% looks high, but after removing chance agreement, kappa is about 0.76 (kappa of 1 means agreement on every image, and 0 means agreement no better than chance). Boxes 2 and 3 have low IoU, meaning the annotators understood the extent of the crack differently. Revise the guidelines on drawing boxes before labelling the rest.

Three pairs of boxes: the first blue and orange pair almost exactly overlap, IoU 0.85; the second overlaps partly, IoU 0.38; the third overlaps only at a small corner, IoU 0.04
Figure 1 IoU of boxes from two annotators

Splitting by flight while keeping class proportions

Images from the same flight are very similar and must stay in the same split as a whole (a group split). But splitting by flight alone can give some splits unusually few or many cracks. scikit-learn provides StratifiedGroupKFold, which keeps both groups and class proportions; this example writes a simple version of the idea to show the mechanism.

Example 2 Splitting 12 flights into train, validation and test

Each flight has a different number of images and proportion of images with cracks (simulated data). Sort the flights by crack rate, then deal them into splits with the repeating pattern train, train, validation, train, train, test.

flights = {"F01": (118, 0.138), "F02": (135, 0.245), "F03": (137, 0.160), "F04": (66, 0.107),
           "F05": (110, 0.060), "F06": (124, 0.220), "F07": (77, 0.228), "F08": (103, 0.130),
           "F09": (131, 0.201), "F10": (94, 0.246), "F11": (134, 0.243), "F12": (89, 0.234)}   # images, crack rate
pattern = ["train", "train", "val", "train", "train", "test"]
split = {"train": [], "val": [], "test": []}
for i, f in enumerate(sorted(flights, key=lambda k: flights[k][1])):
    split[pattern[i % len(pattern)]].append(f)

total = sum(n for n, _ in flights.values())
for name, fs in split.items():
    imgs = sum(flights[f][0] for f in fs)
    cracks = sum(flights[f][0] * flights[f][1] for f in fs)
    print(f"{name:<5} flights {', '.join(fs):<32} images {imgs:>4} ({imgs / total:.0%}), crack rate {cracks / imgs:.3f}")
assert not (set(split["train"]) & set(split["test"]))
train flights F05, F04, F01, F03, F06, F07, F11, F02 images  901 (68%), crack rate 0.180
val   flights F08, F12                         images  192 (15%), crack rate 0.178
test  flights F09, F10                         images  225 (17%), crack rate 0.220

Train and validation have very similar crack rates, test is slightly higher, and no flight is in both train and test. With few flights, the split is coarse, so neither image shares nor class rates hit their targets exactly. In practice, collect data from many flights and sites, and record in the dataset documentation how the split was made.

Three horizontal bars for train, validation and test, each divided into cells by flight with widths proportional to image count; the end of each bar shows the crack rates 0.180, 0.178 and 0.220
Figure 2 Flight-grouped split preserving class ratio

Checking the dataset before training

The drone knowledge base’s unit on checking datasets before AI training practises finding leakage from the same flight, duplicate images and incomplete data with CSV and Python. The minimum checklist is: duplicate or near-duplicate images across splits, boxes outside the image or with zero size, images without labels that should have them, class proportions in each split, and image usage rights.

Module lab

Lab: a trustworthy dataset

  1. Write one page of labelling guidelines for cracks, with images of what does and does not count
  2. Have two classmates label the same 50 images in CVAT, and compute kappa and IoU with Example 1
  3. Revise the guidelines where they disagree, relabel and measure again
  4. Split the dataset by flight with Example 2 or StratifiedGroupKFold and report class proportions
  5. Check the dataset against the minimum checklist and write the dataset documentation

Common mistakes

Watch out

  • Labelling without guidelines and hoping the model will sort it out
  • Reporting raw agreement without removing chance
  • Splitting randomly by image when images from the same flight are similar
  • Splitting by flight without checking class proportions
  • Not recording image usage rights

Summary

  • Clear labelling guidelines make labels consistent, and label consistency caps model accuracy
  • Cohen’s kappa removes chance agreement; IoU measures box agreement
  • Split by flight or site and keep class proportions similar in every split
  • Check the dataset before every training run and record it in the dataset documentation

Check your understanding

  1. With = 0.9 and = 0.5, what is kappa?
  2. What is the IoU of boxes (0,0,10,10) and (5,0,15,10)?
  3. Why might 90% agreement still not be good enough?
  4. Why split by flight?
  5. What does a kappa of 0 mean?
Answers
  1. An overlap of 50 over a union of 150, so
  2. If one class is very common, high agreement can arise by chance; look at kappa
  3. Images from the same flight are very similar; if they end up in both train and test, scores look better than they are
  4. The annotators agree no better than chance

Key formulas

Cohen's kappa
IoU of two boxes

Key references

  1. Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46. link
  2. scikit-learn developers. Cross-validation: Evaluating estimator performance (scikit-learn 1.9). link
  3. Géron, A. (2022). Hands-on machine learning with Scikit-Learn, Keras, and TensorFlow (3rd ed.). O'Reilly.

Further reading

Study the assigned knowledge units in advance, review media and take the module quiz

In class / field

Lab or field practice from worksheets with a safety checklist

Learning evidence: Checked worksheets and quiz results

Module quiz

This is a formative self-check, not a graded exam

Knowledge domain: Artificial intelligence and computer vision