Datasets and labelling
UAT 306 Artificial Intelligence for UAS
Lesson
By the end of this module you will be able to
- Write labelling guidelines that make different annotators agree
- Measure annotator agreement with Cohen's kappa and box IoU
- Split a dataset by flight while keeping class proportions similar
- Check the dataset before training for leakage and wrong labels
Why this matters
The drone knowledge base’s unit on preparing and splitting datasets stresses that a model learns only what the data and labels support: define the objects, the image scope and the rights before collecting, and separate data by groups that may be similar, such as flights or sites. UAT 315 showed that random splits give over-optimistic scores. This module goes deeper into two questions: how trustworthy are the labels, and how to split by flight while keeping class proportions similar.
Labelling guidelines
If two annotators label the same image differently, the model learns that confusion too. Good guidelines define each class with images of what does and does not belong, the smallest object to label, how to draw boxes (tight to the object, with or without shadow), and what to do when unsure. The drone knowledge base’s CVAT unit shows how to label, review and export in a training format.
Example 1 How well do two annotators agree?
Two annotators classify 100 image patches as crack, none or stain, and draw boxes for the same three cracks (simulated data). Cohen’s kappa (Cohen, 1960) removes the agreement expected by chance.
a = ["crack"] * 40 + ["none"] * 50 + ["stain"] * 10
b = ["crack"] * 34 + ["none"] * 6 + ["none"] * 44 + ["crack"] * 4 + ["stain"] * 2 + ["stain"] * 8 + ["none"] * 2
n = len(a)
labels = sorted(set(a) | set(b))
po = sum(x == y for x, y in zip(a, b)) / n
pe = sum((a.count(c) / n) * (b.count(c) / n) for c in labels)
print(f"observed agreement {po:.2f}, chance agreement {pe:.2f}, kappa {(po - pe) / (1 - pe):.2f}")
def iou(p, q): # box (x1, y1, x2, y2) in pixels
ix = max(0, min(p[2], q[2]) - max(p[0], q[0]))
iy = max(0, min(p[3], q[3]) - max(p[1], q[1]))
inter = ix * iy
union = (p[2] - p[0]) * (p[3] - p[1]) + (q[2] - q[0]) * (q[3] - q[1]) - inter
return inter / union
pairs = [((100, 100, 200, 160), (104, 98, 206, 158)), ((50, 50, 90, 80), (60, 55, 110, 85)),
((300, 200, 340, 260), (330, 240, 380, 300))]
for k, (p, q) in enumerate(pairs, 1):
print(f"box {k}: IoU between annotators {iou(p, q):.2f}")
observed agreement 0.86, chance agreement 0.42, kappa 0.76
box 1: IoU between annotators 0.85
box 2: IoU between annotators 0.38
box 3: IoU between annotators 0.04
Agreement of 86% looks high, but after removing chance agreement, kappa is about 0.76 (kappa of 1 means agreement on every image, and 0 means agreement no better than chance). Boxes 2 and 3 have low IoU, meaning the annotators understood the extent of the crack differently. Revise the guidelines on drawing boxes before labelling the rest.
Splitting by flight while keeping class proportions
Images from the same flight are very similar and must stay in the same split as a whole (a group split). But splitting by flight alone can give some splits unusually few or many cracks. scikit-learn provides StratifiedGroupKFold, which keeps both groups and class proportions; this example writes a simple version of the idea to show the mechanism.
Example 2 Splitting 12 flights into train, validation and test
Each flight has a different number of images and proportion of images with cracks (simulated data). Sort the flights by crack rate, then deal them into splits with the repeating pattern train, train, validation, train, train, test.
flights = {"F01": (118, 0.138), "F02": (135, 0.245), "F03": (137, 0.160), "F04": (66, 0.107),
"F05": (110, 0.060), "F06": (124, 0.220), "F07": (77, 0.228), "F08": (103, 0.130),
"F09": (131, 0.201), "F10": (94, 0.246), "F11": (134, 0.243), "F12": (89, 0.234)} # images, crack rate
pattern = ["train", "train", "val", "train", "train", "test"]
split = {"train": [], "val": [], "test": []}
for i, f in enumerate(sorted(flights, key=lambda k: flights[k][1])):
split[pattern[i % len(pattern)]].append(f)
total = sum(n for n, _ in flights.values())
for name, fs in split.items():
imgs = sum(flights[f][0] for f in fs)
cracks = sum(flights[f][0] * flights[f][1] for f in fs)
print(f"{name:<5} flights {', '.join(fs):<32} images {imgs:>4} ({imgs / total:.0%}), crack rate {cracks / imgs:.3f}")
assert not (set(split["train"]) & set(split["test"]))
train flights F05, F04, F01, F03, F06, F07, F11, F02 images 901 (68%), crack rate 0.180
val flights F08, F12 images 192 (15%), crack rate 0.178
test flights F09, F10 images 225 (17%), crack rate 0.220
Train and validation have very similar crack rates, test is slightly higher, and no flight is in both train and test. With few flights, the split is coarse, so neither image shares nor class rates hit their targets exactly. In practice, collect data from many flights and sites, and record in the dataset documentation how the split was made.
Checking the dataset before training
The drone knowledge base’s unit on checking datasets before AI training practises finding leakage from the same flight, duplicate images and incomplete data with CSV and Python. The minimum checklist is: duplicate or near-duplicate images across splits, boxes outside the image or with zero size, images without labels that should have them, class proportions in each split, and image usage rights.
Module lab
Lab: a trustworthy dataset
- Write one page of labelling guidelines for cracks, with images of what does and does not count
- Have two classmates label the same 50 images in CVAT, and compute kappa and IoU with Example 1
- Revise the guidelines where they disagree, relabel and measure again
- Split the dataset by flight with Example 2 or
StratifiedGroupKFoldand report class proportions - Check the dataset against the minimum checklist and write the dataset documentation
Common mistakes
Watch out
- Labelling without guidelines and hoping the model will sort it out
- Reporting raw agreement without removing chance
- Splitting randomly by image when images from the same flight are similar
- Splitting by flight without checking class proportions
- Not recording image usage rights
Summary
- Clear labelling guidelines make labels consistent, and label consistency caps model accuracy
- Cohen’s kappa removes chance agreement; IoU measures box agreement
- Split by flight or site and keep class proportions similar in every split
- Check the dataset before every training run and record it in the dataset documentation
Check your understanding
- With = 0.9 and = 0.5, what is kappa?
- What is the IoU of boxes (0,0,10,10) and (5,0,15,10)?
- Why might 90% agreement still not be good enough?
- Why split by flight?
- What does a kappa of 0 mean?
Answers
- An overlap of 50 over a union of 150, so
- If one class is very common, high agreement can arise by chance; look at kappa
- Images from the same flight are very similar; if they end up in both train and test, scores look better than they are
- The annotators agree no better than chance
Key formulas
| Cohen's kappa | |
| IoU of two boxes |
Key references
- Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46. link
- scikit-learn developers. Cross-validation: Evaluating estimator performance (scikit-learn 1.9). link
- Géron, A. (2022). Hands-on machine learning with Scikit-Learn, Keras, and TensorFlow (3rd ed.). O'Reilly.
Further reading
Study the assigned knowledge units in advance, review media and take the module quiz
Preparing and splitting datasets
Checking datasets before AI training
CVAT Community
In class / field
Lab or field practice from worksheets with a safety checklist
Learning evidence: Checked worksheets and quiz results