Object detection and tracking
UAT 307 Computer Vision and Perception Technology
Lesson
By the end of this module you will be able to
- Explain tracking-by-detection
- Write a constant-velocity Kalman filter to track object position
- Associate boxes between frames by IoU and count identity switches
- Evaluate tracking with MOTA
Why this matters
A detector says where objects are in this frame, but many drone jobs need to know where the same object went, such as counting cars through a junction, following people in a search, or monitoring animals in a reserve. The drone knowledge base’s unit on intelligent video analytics covers detection, tracking, counting and real-time alerts. A widely used approach is tracking-by-detection, as in SORT (Bewley et al., 2016), which predicts positions with a constant-velocity Kalman filter and associates them with detections by IoU using the Hungarian method. This module’s case is a drone watching a car park, with simulated data.
Predicting position with a Kalman filter
The Kalman filter from UAT 305 can track objects in images. The state is position and velocity in the image; the predict step moves the position by the velocity, and the update step uses the detector’s box. When the detector misses some frames, for example a car hidden by a tree, the filter keeps predicting, but its uncertainty grows with every frame without data.
Example 1 Tracking a car hidden for three frames
The car moves 3 pixels per frame horizontally and 1.5 vertically; the detector has about 2 pixels of error and misses frames 5–7.
import math
import numpy as np
dt = 0.1
F = np.array([[1, 0, dt, 0], [0, 1, 0, dt], [0, 0, 1, 0], [0, 0, 0, 1]]) # [x, y, vx, vy]
Hm = np.eye(2, 4)
Q, R = np.eye(4) * 0.5, np.eye(2) * 4.0
truth = [(10 + 3 * k, 50 + 1.5 * k) for k in range(12)]
rng = np.random.default_rng(6)
meas = [(x + rng.normal(0, 2), y + rng.normal(0, 2)) for x, y in truth]
missing = {5, 6, 7} # occluded frames
x = np.array([meas[0][0], meas[0][1], 0.0, 0.0])
P = np.diag([4.0, 4.0, 100.0, 100.0])
for k in range(1, 12):
x, P = F @ x, F @ P @ F.T + Q
if k not in missing:
S = Hm @ P @ Hm.T + R
K = P @ Hm.T @ np.linalg.inv(S)
x = x + K @ (np.array(meas[k]) - Hm @ x)
P = (np.eye(4) - K @ Hm) @ P
err = math.dist(x[:2], truth[k])
print(f"frame {k:>2} {'predict only' if k in missing else 'updated '} error {err:4.2f} px, position sigma {math.sqrt(P[0, 0]):.2f} px")
frame 1 updated error 3.41 px, position sigma 1.52 px
frame 2 updated error 2.30 px, position sigma 1.46 px
frame 3 updated error 1.95 px, position sigma 1.46 px
frame 4 updated error 1.31 px, position sigma 1.44 px
frame 5 predict only error 1.75 px, position sigma 2.00 px
frame 6 predict only error 2.57 px, position sigma 2.58 px
frame 7 predict only error 3.51 px, position sigma 3.17 px
frame 8 updated error 2.69 px, position sigma 1.77 px
frame 9 updated error 2.63 px, position sigma 1.47 px
frame 10 updated error 0.93 px, position sigma 1.35 px
frame 11 updated error 0.66 px, position sigma 1.29 px
While hidden, the uncertainty and error grow each frame. When detections return, the filter pulls the position back within a few frames. The growing uncertainty is used to widen the search region for re-associating the same object when it reappears. If it stays hidden too long, the system should close the track rather than keep predicting.
Association and MOTA
With several objects, boxes in a new frame must be associated with existing tracks. A wrong association swaps the identities of two objects (an ID switch), which breaks counting. The MOTA metric from CLEAR MOT (Bernardin and Stiefelhagen, 2008) combines three kinds of error, misses (FN), false detections (FP) and identity switches (IDSW), divided by the total number of real objects.
Example 2 Two cars passing each other
Cars A and B approach each other over five frames. Compare the tracker’s output with ground truth, treating IoU of 0.5 or more as a match (simulated data).
def iou(p, q):
ix = max(0, min(p[2], q[2]) - max(p[0], q[0]))
iy = max(0, min(p[3], q[3]) - max(p[1], q[1]))
inter = ix * iy
return inter / ((p[2] - p[0]) * (p[3] - p[1]) + (q[2] - q[0]) * (q[3] - q[1]) - inter)
gt = {1: [("A", (10, 10, 30, 40)), ("B", (60, 10, 80, 40))], 2: [("A", (14, 10, 34, 40)), ("B", (56, 10, 76, 40))],
3: [("A", (18, 10, 38, 40)), ("B", (52, 10, 72, 40))], 4: [("A", (22, 10, 42, 40)), ("B", (48, 10, 68, 40))],
5: [("A", (26, 10, 46, 40)), ("B", (44, 10, 64, 40))]}
tracks = {1: [(1, (11, 11, 31, 41)), (2, (59, 10, 79, 40))], 2: [(1, (15, 10, 35, 41)), (2, (55, 11, 75, 40))],
3: [(1, (19, 10, 39, 40))], 4: [(2, (22, 11, 42, 40)), (1, (47, 10, 67, 41))],
5: [(2, (27, 10, 47, 40)), (1, (43, 10, 63, 40)), (3, (90, 50, 110, 80))]}
last, fn, fp, idsw, total = {}, 0, 0, 0, 0
for f in sorted(gt):
total += len(gt[f])
used = set()
for gid, gbox in gt[f]:
best = max(((iou(gbox, tb), tid) for tid, tb in tracks[f] if tid not in used), default=(0, None))
if best[0] >= 0.5:
used.add(best[1])
if gid in last and last[gid] != best[1]:
idsw += 1
print(f"frame {f}: {gid} switched from track {last[gid]} to track {best[1]}")
last[gid] = best[1]
else:
fn += 1
fp += sum(1 for tid, _ in tracks[f] if tid not in used)
print(f"FN {fn}, FP {fp}, IDSW {idsw}, objects {total} -> MOTA {1 - (fn + fp + idsw) / total:.2f}")
frame 4: A switched from track 1 to track 2
frame 4: B switched from track 2 to track 1
FN 1, FP 1, IDSW 2, objects 10 -> MOTA 0.60
The tracker misses car B in one frame, adds one extra box, and swaps the identities of both cars as they pass close together. MOTA is therefore only 0.60 even though almost every car is detected in every frame. Counting with this output would send each car the wrong way. Using the Kalman filter’s velocity for association, and appearance features of the objects, reduces identity switches.
Module lab
Lab: tracking cars in drone video
- Run the detector from UAT 306 on every frame of a car-park video and save the boxes
- Write a tracker from Example 1 with IoU association
- Label true tracks for 100 frames and compute MOTA with Example 2
- Try different IoU thresholds for association and different numbers of missing frames before closing a track
- Count cars crossing a set line and compare with counting by eye
Common mistakes
Watch out
- Evaluating only per-frame detection without looking at identity switches
- Predicting forever when an object has been gone a long time
- Setting the IoU threshold too high for fast-moving objects
- Not compensating for the drone’s own motion, so every object seems to move
- Counting objects by track ID without checking for switches
Summary
- Tracking-by-detection predicts positions with a Kalman filter and associates them with detections
- Uncertainty grows during occlusion and is used to widen the search region and decide when to close tracks
- Identity switches break counts and path analysis
- MOTA combines FN, FP and IDSW, divided by the number of real objects
Check your understanding
- An object is at x = 100 px moving at 20 px/s. Where is it predicted after 0.5 s?
- FN 5, FP 3, IDSW 2, with 50 real objects in total. What is MOTA?
- Why does uncertainty grow when there are no detections?
- What is an ID switch?
- Why compensate for the drone’s motion when tracking ground objects?
Answers
- px
- The predict step adds every frame, but no update step reduces the uncertainty
- The same real object is associated with a different track ID from the previous frame
- The whole image shifts as the drone moves, so stationary objects appear to move
Key formulas
| Constant-velocity model | |
| MOTA |
Key references
- Bewley, A., Ge, Z., Ott, L., Ramos, F., & Upcroft, B. (2016). Simple online and realtime tracking. In 2016 IEEE International Conference on Image Processing (pp. 3464–3468). link
- Bernardin, K., & Stiefelhagen, R. (2008). Evaluating multiple object tracking performance: The CLEAR MOT metrics. EURASIP Journal on Image and Video Processing, 2008, 246309. link
- Szeliski, R. (2022). Computer vision: Algorithms and applications (2nd ed.). Springer. link
- Corke, P. (2023). Robotics, vision and control: Fundamental algorithms in Python. Springer. link
Further reading
Study the assigned knowledge units in advance, review media and take the module quiz
In class / field
Lab or field practice from worksheets with a safety checklist
Learning evidence: Checked worksheets and quiz results