Module 2/5 · Weeks 4–6 · 27 h

ML fundamentals

UAT 315 Artificial Intelligence, Data Analytics and Computer Vision for Unmanned Aircraft Systems

About 90 minDraft, awaiting reviewLast updated 27 September 2026

Lesson

By the end of this module you will be able to

  1. Distinguish supervised from unsupervised learning, and classification from regression
  2. Split data into training, validation and test sets and explain the role of each
  3. Train and compare logistic regression and a decision tree with scikit-learn
  4. Explain underfitting and overfitting, and why accuracy alone can mislead

Prerequisites: UAT 315 module 1 · UAT 106 module 3

Why this matters

Machine learning (ML) lets a computer find rules from examples, instead of people writing every rule. Many drone tasks suit ML, such as predicting a failing motor from vibration logs or classifying flooded areas in images. But a model that scores well on data it has seen may not work at all on new data. This module lays the principles that make ML results trustworthy.

Kinds of learning

  • Supervised learning has the correct answer (the label) for each training example. It divides into classification, when the answer is a category such as normal or faulty motor, and regression, when the answer is a number such as remaining flight time.
  • Unsupervised learning has no labels; the model finds structure itself, such as grouping similar flights or finding flights that differ from the rest.
Features x, vibration and current, go into the model f of x and theta to give a prediction y hat, normal or faulty. It is compared with the true label y to give an error, the loss, which is used to update theta in the model
Figure 1 Supervised learning

Training adjusts the model’s parameters so predictions come as close as possible to the true labels, measuring the gap with a loss function.

Three data sets

This example uses synthetic data from 400 flights. The features are mean vibration (g) and hover current (A); the label is whether the motor is faulty. The data splits into three sets with different jobs:

  • Training set: used to fit parameters
  • Validation set: used to choose the model and settings, such as tree depth
  • Test set: kept aside to evaluate once, after every choice is made
import numpy as np
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(345)
n = 400
faulty = rng.random(n) < 0.2
vibration = np.where(faulty, rng.normal(0.17, 0.05, n), rng.normal(0.12, 0.04, n))
current = np.where(faulty, rng.normal(20, 2.5, n), rng.normal(18, 2.5, n))
X = np.column_stack([vibration, current])
y = faulty.astype(int)

X_rest, X_test, y_rest, y_test = train_test_split(X, y, test_size=0.2, stratify=y, random_state=345)
X_train, X_val, y_train, y_val = train_test_split(X_rest, y_rest, test_size=0.25, stratify=y_rest, random_state=345)
print(f"train {len(y_train)}  validation {len(y_val)}  test {len(y_test)}   faulty share {y.mean():.1%}")
train 240  validation 80  test 80   faulty share 21.8%

stratify=y keeps the share of faulty motors similar in every set, which matters when one class is rare.

A first model and the accuracy trap

Logistic regression is a basic classification model: it combines features linearly and turns the result into a probability with a sigmoid function. Before looking at results, always have a baseline.

from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, recall_score
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

baseline = np.zeros_like(y_val)
print(f"always 'healthy': accuracy {accuracy_score(y_val, baseline):.3f}  recall {recall_score(y_val, baseline):.3f}")

model = make_pipeline(StandardScaler(), LogisticRegression()).fit(X_train, y_train)
pred = model.predict(X_val)
print(f"logistic regression: accuracy {accuracy_score(y_val, pred):.3f}  recall {recall_score(y_val, pred):.3f}")
pred_03 = (model.predict_proba(X_val)[:, 1] >= 0.3).astype(int)
print(f"threshold 0.3:       accuracy {accuracy_score(y_val, pred_03):.3f}  recall {recall_score(y_val, pred_03):.3f}")
always 'healthy': accuracy 0.787  recall 0.000
logistic regression: accuracy 0.950  recall 0.765
threshold 0.3:       accuracy 0.900  recall 0.824

A model that always answers “healthy” is almost 79% accurate while finding no faulty motor at all (recall = 0). That is the accuracy trap when one class is rare. Logistic regression is much better, and lowering the probability threshold from 0.5 to 0.3 finds more faulty motors at the cost of more false alarms. In maintenance, where missing a failing motor may bring a drone down, higher recall is usually worth it. make_pipeline makes StandardScaler learn its statistics from the training set only, as in module 1.

Underfitting and overfitting

Three panels of the same data points. In the first a nearly flat straight line misses the data: underfitting. In the second a smooth curve follows the trend: a good fit. In the third a jagged line passes through every point: overfitting
Figure 2 Underfitting, a good fit and overfitting
  • Underfitting: the model is too simple and does poorly on both training and validation data
  • Overfitting: the model memorises details and noise in the training set, doing very well there but worse on new data

A decision tree splits data with a sequence of yes/no questions; its depth controls complexity.

from sklearn.tree import DecisionTreeClassifier

for depth in (1, 2, 3, 4, 6, 10, None):
    tree = DecisionTreeClassifier(max_depth=depth, random_state=0).fit(X_train, y_train)
    print(f"max_depth {str(depth):>4}: train {accuracy_score(y_train, tree.predict(X_train)):.3f}"
          f"   validation {accuracy_score(y_val, tree.predict(X_val)):.3f}")
max_depth    1: train 0.867   validation 0.887
max_depth    2: train 0.879   validation 0.863
max_depth    3: train 0.904   validation 0.925
max_depth    4: train 0.929   validation 0.925
max_depth    6: train 0.954   validation 0.912
max_depth   10: train 0.992   validation 0.887
max_depth None: train 1.000   validation 0.887

As the tree deepens, training accuracy climbs to 1.000, but validation accuracy peaks at depth 3 to 4 and then falls. Unlimited depth is clear overfitting, so we choose depth from the validation set, not the training set.

Example 1 The final evaluation on the test set

After choosing depth 3 from the validation set, retrain on training plus validation data and evaluate on the test set once.

final = DecisionTreeClassifier(max_depth=3, random_state=0).fit(X_rest, y_rest)
test_pred = final.predict(X_test)
print(f"test accuracy {accuracy_score(y_test, test_pred):.3f}   test recall {recall_score(y_test, test_pred):.3f}")
test accuracy 0.800   test recall 0.412

The test result is much worse than validation: accuracy falls from 0.925 to 0.800, and only 41% of faulty motors are found. This happens often in real work. One reason is that the validation set holds only 17 faulty motors, so figures from such a small set swing widely, and picking the best setting on the validation set also flatters that set a little. The test figure is the one to report honestly. Going back to adjust the depth after seeing it would make the test set no longer independent. The right response is to collect more data, use cross-validation, and prepare a new test set for the next round.

Class activity

Activity: designing an ML task for maintenance

  1. Define a problem, such as predicting whether a battery should be retired. State whether it is classification or regression, supervised or not, and where the labels come from.
  2. Discuss how the harm of missing a real case (FN) differs from a false alarm (FP) in that task, and choose the main metric.
  3. Change the threshold in the example code to 0.2 and 0.5, record accuracy and recall, and choose a threshold with reasons.
  4. Explain to a classmate why the test set must never be used to choose tree depth.

Common mistakes

Watch out

  • Reporting accuracy alone on data where one class is rare
  • Having no baseline, so you cannot tell whether the model beats a simple guess
  • Choosing the model on the training set, which rewards overfitting
  • Using the test set many times until it becomes another validation set
  • Always using a 0.5 threshold without weighing the harm of each kind of error

Summary

  • Supervised ML learns from labelled examples, as classification or regression
  • The training set fits parameters, the validation set chooses the model, and the test set evaluates once
  • Always compare with a baseline; accuracy misleads when classes are unbalanced
  • Too much complexity overfits; choose complexity from validation results

Check your understanding

  1. Predicting remaining flight time in minutes is which kind of task?
  2. Of 1,000 flights, 50 have faulty motors. How accurate is a model that always answers “healthy”?
  3. A model scores 1.00 accuracy on training data but 0.80 on validation data. What does that tell you?
  4. There are 20 truly faulty motors and the model finds 15. What is the recall?
  5. Why use stratify when splitting data where one class is rare?
Answers
  1. Supervised regression
  2. , while finding no faulty motor at all
  3. The model overfits: it memorises the training set but works poorly on new data
  4. To keep the share of the rare class similar in every set; otherwise some sets may have almost no examples of it

Key formulas

Sigmoid of logistic regression
Accuracy
Recall

Key references

  1. Géron, A. (2025). Hands-on machine learning with Scikit-Learn and PyTorch. O'Reilly. link
  2. Géron, A. (2022). Hands-on machine learning with Scikit-Learn, Keras, and TensorFlow (3rd ed.). O'Reilly.
  3. Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. MIT Press. link
  4. scikit-learn developers. Metrics and scoring: Quantifying the quality of predictions (scikit-learn 1.9). link
  5. scikit-learn developers. Cross-validation: Evaluating estimator performance (scikit-learn 1.9). link
  6. Montgomery, D. C., & Runger, G. C. (2018). Applied statistics and probability for engineers (7th ed.). Wiley. link

Further reading

Study the assigned knowledge units in advance, review media and take the module quiz

In class / field

Lecture, case discussion and in-class problem solving

Learning evidence: Quiz results and submitted exercises

Module quiz

This is a formative self-check, not a graded exam

Knowledge domain: Artificial intelligence and computer vision