Machine learning fundamentals
UAT 306 Artificial Intelligence for UAS
Lesson
By the end of this module you will be able to
- Write logistic regression and train it with gradient descent from scratch
- Explain the cross-entropy loss and the role of the learning rate
- Use k-fold cross-validation to choose model complexity
- Distinguish underfitting from overfitting using training and validation error
Why this matters
Libraries such as scikit-learn and PyTorch let you train a model in a few lines, but without understanding what happens inside, you cannot fix a model that will not learn, or one that learns too much to work on new data. The drone knowledge base’s unit on machine learning fundamentals covers supervised and unsupervised learning, data splitting, overfitting and model evaluation, and the textbooks by Géron and Prince explain them in detail. This module writes everything in NumPy to see the real mechanism.
Logistic regression
Logistic regression predicts a class probability with the sigmoid of a weighted sum of features. Training finds the weights that minimise cross-entropy, which heavily penalises confident wrong predictions. Gradient descent moves the weights against the gradient one step at a time; the step size is the learning rate . Too small and learning is slow; too large and it jumps over the minimum.
Example 1 Separating cracks from stains with two features
Two hypothetical features from image patches: the length-to-width ratio of the mark and its edge strength (200 simulated examples).
import numpy as np
rng = np.random.default_rng(5)
n = 200
elong = np.r_[rng.normal(2, 0.8, n // 2), rng.normal(4, 0.8, n // 2)] # length-to-width ratio
edge = np.r_[rng.normal(1, 0.6, n // 2), rng.normal(2.5, 0.6, n // 2)] # edge strength
y = np.r_[np.zeros(n // 2), np.ones(n // 2)] # 1 = crack
X = np.c_[np.ones(n), (elong - elong.mean()) / elong.std(), (edge - edge.mean()) / edge.std()]
w, lr = np.zeros(3), 0.5
for it in range(1, 501):
p = 1 / (1 + np.exp(-X @ w))
w -= lr * X.T @ (p - y) / n
if it in (1, 10, 100, 500):
loss = -np.mean(y * np.log(p + 1e-12) + (1 - y) * np.log(1 - p + 1e-12))
print(f"iteration {it:>3}: loss {loss:.3f}, accuracy {((p > 0.5) == y).mean():.3f}")
print("weights (bias, elongation, edge):", np.round(w, 2))
iteration 1: loss 0.693, accuracy 0.500
iteration 10: loss 0.241, accuracy 0.980
iteration 100: loss 0.102, accuracy 0.980
iteration 500: loss 0.083, accuracy 0.980
weights (bias, elongation, edge): [0.09 4.25 3.11]
The loss falls quickly at first and then slows, and accuracy settles after a few dozen iterations. Both feature weights are positive, meaning long, narrow marks with sharp edges tend to be cracks. Standardising the features (subtracting the mean and dividing by the standard deviation) before training makes gradient descent converge faster.
Overfitting and cross-validation
A model that is too complex memorises noise in the training set, so its training error is low but its error on new data is high: overfitting. A model that is too simple is wrong on both: underfitting. k-fold cross-validation splits the data into parts, trains times leaving one part out for validation each time, and averages the results. It chooses complexity without touching the test set (scikit-learn cross-validation documentation).
Example 2 Choosing a polynomial degree with 5-fold
A hypothetical relationship between crack depth and distance from the beam edge, 30 points (simulated data).
import numpy as np
rng = np.random.default_rng(2)
x = np.linspace(0, 1, 30)
y = 5 * x + 2 * np.sin(5 * x) + rng.normal(0, 0.6, 30)
folds = np.array_split(rng.permutation(30), 5)
for degree in (1, 3, 9):
tr_err, va_err = [], []
for f in folds:
tr = np.setdiff1d(np.arange(30), f)
c = np.polyfit(x[tr], y[tr], degree)
tr_err.append(np.sqrt(np.mean((np.polyval(c, x[tr]) - y[tr]) ** 2)))
va_err.append(np.sqrt(np.mean((np.polyval(c, x[f]) - y[f]) ** 2)))
print(f"degree {degree}: train RMSE {np.mean(tr_err):.2f}, validation RMSE {np.mean(va_err):.2f}")
degree 1: train RMSE 1.09, validation RMSE 1.22
degree 3: train RMSE 0.55, validation RMSE 0.61
degree 9: train RMSE 0.46, validation RMSE 0.74
Degree 1 is wrong on both training and validation data: underfitting. Degree 9 has the lowest training error but a higher validation error: overfitting. A middle degree gives the lowest validation error. The same principle applies to neural networks: watch the validation loss, not the training loss.
Module lab
Lab: training by hand and comparing with a library
- Run Example 1 with learning rates of 0.01, 0.5 and 5 and observe the loss
- Train the same data with scikit-learn’s
LogisticRegressionand compare weights and accuracy - Run Example 2 for degrees 1 to 12 and plot both errors
- Reduce the data to 15 points and see whether the best degree changes
- Write a summary of how to tell whether your own model underfits or overfits
Common mistakes
Watch out
- Not standardising features, so training converges slowly or not at all
- Setting the learning rate too high, so the loss oscillates or explodes
- Looking only at training error
- Using the test set to choose complexity
- Increasing complexity when data is scarce
Summary
- Logistic regression predicts probabilities with a sigmoid and is trained by minimising cross-entropy with gradient descent
- The learning rate sets the step size; standardising features speeds convergence
- Underfitting is wrong on both sets; overfitting is right on the training set but wrong on validation
- k-fold cross-validation chooses complexity without touching the test set
Check your understanding
- If , what probability is predicted?
- If the model predicts but the answer is 0, what is that example’s loss?
- Very low training error and high validation error indicate what problem?
- How many times does 5-fold cross-validation train a model?
- Why should the test set not be used to choose the degree?
Answers
- 0.5
- Overfitting
- 5 times
- The test result would look better than it is, because that data has already been used to decide
Key formulas
| Logistic regression | |
| Cross-entropy and its gradient | |
| Weight update |
Key references
Further reading
Study the assigned knowledge units in advance, review media and take the module quiz
In class / field
Lab or field practice from worksheets with a safety checklist
Learning evidence: Checked worksheets and quiz results