Module 2/5 · Weeks 4–6 · 27 h

Machine learning fundamentals

UAT 306 Artificial Intelligence for UAS

About 90 minDraft, awaiting reviewLast updated 28 September 2026

Lesson

By the end of this module you will be able to

  1. Write logistic regression and train it with gradient descent from scratch
  2. Explain the cross-entropy loss and the role of the learning rate
  3. Use k-fold cross-validation to choose model complexity
  4. Distinguish underfitting from overfitting using training and validation error

Prerequisites: UAT 306 Module 1 · UAT 104 (programming)

Why this matters

Libraries such as scikit-learn and PyTorch let you train a model in a few lines, but without understanding what happens inside, you cannot fix a model that will not learn, or one that learns too much to work on new data. The drone knowledge base’s unit on machine learning fundamentals covers supervised and unsupervised learning, data splitting, overfitting and model evaluation, and the textbooks by Géron and Prince explain them in detail. This module writes everything in NumPy to see the real mechanism.

Logistic regression

Logistic regression predicts a class probability with the sigmoid of a weighted sum of features. Training finds the weights that minimise cross-entropy, which heavily penalises confident wrong predictions. Gradient descent moves the weights against the gradient one step at a time; the step size is the learning rate . Too small and learning is slow; too large and it jumps over the minimum.

Example 1 Separating cracks from stains with two features

Two hypothetical features from image patches: the length-to-width ratio of the mark and its edge strength (200 simulated examples).

import numpy as np

rng = np.random.default_rng(5)
n = 200
elong = np.r_[rng.normal(2, 0.8, n // 2), rng.normal(4, 0.8, n // 2)]   # length-to-width ratio
edge = np.r_[rng.normal(1, 0.6, n // 2), rng.normal(2.5, 0.6, n // 2)]  # edge strength
y = np.r_[np.zeros(n // 2), np.ones(n // 2)]                            # 1 = crack
X = np.c_[np.ones(n), (elong - elong.mean()) / elong.std(), (edge - edge.mean()) / edge.std()]

w, lr = np.zeros(3), 0.5
for it in range(1, 501):
    p = 1 / (1 + np.exp(-X @ w))
    w -= lr * X.T @ (p - y) / n
    if it in (1, 10, 100, 500):
        loss = -np.mean(y * np.log(p + 1e-12) + (1 - y) * np.log(1 - p + 1e-12))
        print(f"iteration {it:>3}: loss {loss:.3f}, accuracy {((p > 0.5) == y).mean():.3f}")
print("weights (bias, elongation, edge):", np.round(w, 2))
iteration   1: loss 0.693, accuracy 0.500
iteration  10: loss 0.241, accuracy 0.980
iteration 100: loss 0.102, accuracy 0.980
iteration 500: loss 0.083, accuracy 0.980
weights (bias, elongation, edge): [0.09 4.25 3.11]

The loss falls quickly at first and then slows, and accuracy settles after a few dozen iterations. Both feature weights are positive, meaning long, narrow marks with sharp edges tend to be cracks. Standardising the features (subtracting the mean and dividing by the standard deviation) before training makes gradient descent converge faster.

Loss against training iteration from 1 to 500 on a log axis: a blue line falling from 0.69 to about 0.08
Figure 1 Loss during training

Overfitting and cross-validation

A model that is too complex memorises noise in the training set, so its training error is low but its error on new data is high: overfitting. A model that is too simple is wrong on both: underfitting. k-fold cross-validation splits the data into parts, trains times leaving one part out for validation each time, and averages the results. It chooses complexity without touching the test set (scikit-learn cross-validation documentation).

Example 2 Choosing a polynomial degree with 5-fold

A hypothetical relationship between crack depth and distance from the beam edge, 30 points (simulated data).

import numpy as np

rng = np.random.default_rng(2)
x = np.linspace(0, 1, 30)
y = 5 * x + 2 * np.sin(5 * x) + rng.normal(0, 0.6, 30)
folds = np.array_split(rng.permutation(30), 5)
for degree in (1, 3, 9):
    tr_err, va_err = [], []
    for f in folds:
        tr = np.setdiff1d(np.arange(30), f)
        c = np.polyfit(x[tr], y[tr], degree)
        tr_err.append(np.sqrt(np.mean((np.polyval(c, x[tr]) - y[tr]) ** 2)))
        va_err.append(np.sqrt(np.mean((np.polyval(c, x[f]) - y[f]) ** 2)))
    print(f"degree {degree}: train RMSE {np.mean(tr_err):.2f}, validation RMSE {np.mean(va_err):.2f}")
degree 1: train RMSE 1.09, validation RMSE 1.22
degree 3: train RMSE 0.55, validation RMSE 0.61
degree 9: train RMSE 0.46, validation RMSE 0.74

Degree 1 is wrong on both training and validation data: underfitting. Degree 9 has the lowest training error but a higher validation error: overfitting. A middle degree gives the lowest validation error. The same principle applies to neural networks: watch the validation loss, not the training loss.

Paired bars for degrees 1, 3 and 9: grey training RMSE falling with degree, and blue validation RMSE lowest at degree 3
Figure 2 Training and validation error by complexity

Module lab

Lab: training by hand and comparing with a library

  1. Run Example 1 with learning rates of 0.01, 0.5 and 5 and observe the loss
  2. Train the same data with scikit-learn’s LogisticRegression and compare weights and accuracy
  3. Run Example 2 for degrees 1 to 12 and plot both errors
  4. Reduce the data to 15 points and see whether the best degree changes
  5. Write a summary of how to tell whether your own model underfits or overfits

Common mistakes

Watch out

  • Not standardising features, so training converges slowly or not at all
  • Setting the learning rate too high, so the loss oscillates or explodes
  • Looking only at training error
  • Using the test set to choose complexity
  • Increasing complexity when data is scarce

Summary

  • Logistic regression predicts probabilities with a sigmoid and is trained by minimising cross-entropy with gradient descent
  • The learning rate sets the step size; standardising features speeds convergence
  • Underfitting is wrong on both sets; overfitting is right on the training set but wrong on validation
  • k-fold cross-validation chooses complexity without touching the test set

Check your understanding

  1. If , what probability is predicted?
  2. If the model predicts but the answer is 0, what is that example’s loss?
  3. Very low training error and high validation error indicate what problem?
  4. How many times does 5-fold cross-validation train a model?
  5. Why should the test set not be used to choose the degree?
Answers
  1. 0.5
  2. Overfitting
  3. 5 times
  4. The test result would look better than it is, because that data has already been used to decide

Key formulas

Logistic regression
Cross-entropy and its gradient
Weight update

Key references

  1. Géron, A. (2022). Hands-on machine learning with Scikit-Learn, Keras, and TensorFlow (3rd ed.). O'Reilly.
  2. Prince, S. J. D. (2023). Understanding deep learning. MIT Press. link
  3. scikit-learn developers. Cross-validation: Evaluating estimator performance (scikit-learn 1.9). link

Further reading

Study the assigned knowledge units in advance, review media and take the module quiz

In class / field

Lab or field practice from worksheets with a safety checklist

Learning evidence: Checked worksheets and quiz results

Module quiz

This is a formative self-check, not a graded exam

Knowledge domain: Artificial intelligence and computer vision