Module 4/5 · Weeks 10–12 · 27 h

VIO and SLAM

UAT 307 Computer Vision and Perception Technology

About 90 minDraft, awaiting reviewLast updated 28 September 2026

Lesson

By the end of this module you will be able to

  1. Distinguish VIO, VIN and SLAM by the roles of camera, IMU and map
  2. Compute depth from a stereo camera and its depth error
  3. Explain the scale problem of a single camera and estimate scale from an IMU
  4. Choose perception sensors to suit range and environment

Prerequisites: UAT 307 Modules 1–3 · UAT 366 Module 2 (SLAM and sensor fusion)

Why this matters

The drone knowledge base’s unit on VIN, VIO and SLAM explains the roles of the camera, IMU and data fusion before choosing navigation tools. VIO (visual-inertial odometry) estimates continuous motion from a camera and IMU. SLAM builds a map and localises within it at the same time, and can close loops to reduce drift. UAT 366 used SLAM ideas indoors; this module focuses on depth and scale, the two things that most often make visual navigation go wrong.

Stereo depth

Two cameras separated by a baseline see the same point at different image positions. The difference is the disparity , and depth is , where is the focal length in pixels. Because is inversely proportional to , a small disparity error causes large depth errors at long range, growing with the square of range, following the relation that Keselman et al. (2017) use to describe Intel RealSense stereo depth cameras.

Example 1 A small stereo camera on a drone

Focal length 640 pixels, baseline 95 mm and disparity error 0.25 pixel (hypothetical values).

F_PX, BASE, DD = 640, 0.095, 0.25               # pixels, m, pixels
for z in (2, 5, 10, 20):
    d = F_PX * BASE / z
    dz = z * z / (F_PX * BASE) * DD
    print(f"range {z:>2} m: disparity {d:5.2f} px, depth error about {dz:.3f} m ({dz / z:.1%})")
range  2 m: disparity 30.40 px, depth error about 0.016 m (0.8%)
range  5 m: disparity 12.16 px, depth error about 0.103 m (2.1%)
range 10 m: disparity  6.08 px, depth error about 0.411 m (4.1%)
range 20 m: disparity  3.04 px, depth error about 1.645 m (8.2%)

At 2 m the depth is off by only a couple of centimetres, but at 20 m it is off by more than a metre. Small stereo cameras therefore suit close-range obstacle avoidance; for depth at long range, increase the baseline or resolution, or use LiDAR.

Depth error against range from 0 to 20 metres: a blue curve rising quadratically, with pink points at 2, 5, 10 and 20 metres of about 0.02, 0.10, 0.41 and 1.64 metres
Figure 1 Stereo depth error versus range

Single-camera scale

A single camera knows the direction of motion but not the scale: a real house far away and a model house nearby can look the same. VIO solves this with an IMU, which measures acceleration in real units. Comparing visual displacement (unitless) with IMU-integrated displacement over short periods gives a scale factor. There must be enough acceleration for the IMU to measure clearly, so long periods at constant speed make scale hard to estimate.

Example 2 Estimating scale with least squares

Displacements over six short periods from the IMU (metres) and from images (unitless), with a true scale of 3.2 (simulated data).

import numpy as np

imu = np.array([0.52, 1.05, 1.48, 2.10, 2.47, 3.03])                 # m
vis = imu / 3.2 + np.random.default_rng(8).normal(0, 0.01, 6)        # unitless
s = (vis @ imu) / (vis @ vis)
resid = imu - s * vis
print("visual displacements:", np.round(vis, 3))
print(f"estimated scale {s:.3f} (true 3.2), residual RMS {np.sqrt(np.mean(resid ** 2)):.3f} m")
print(f"a 10-unit visual path would be {10 * s:.1f} m")
visual displacements: [0.145 0.315 0.449 0.653 0.749 0.945]
estimated scale 3.250 (true 3.2), residual RMS 0.034 m
a 10-unit visual path would be 32.5 m

The estimated scale is close to the true value, and the residual indicates estimate quality. A scale error of just 5% makes a 100 m path wrong by 5 m, so VIO systems estimate scale continuously throughout the flight, and some use a rangefinder or barometer to confirm it.

Three boxes: the camera gives motion direction without scale; the IMU gives metric acceleration but drifts; arrows from both lead to a VIO box giving pose with scale, and an arrow continues to a SLAM box that adds a map and loop closure
Figure 2 Roles of camera, IMU, VIO and SLAM

Choosing perception sensors

  • Stereo gives dense depth at short range and works when there is enough texture and light
  • LiDAR gives accurate range farther out, independent of light, but is heavier and more expensive
  • A single camera with an IMU is light and cheap but needs enough motion to find scale
  • All of them struggle with smooth textureless surfaces, reflective surfaces, water and fog

Module lab

Lab: depth and scale

  1. Measure stereo depth at 1–10 m, compare with true range and with Example 1
  2. Change the background to a smooth, textureless wall and observe where depth disappears
  3. Use recorded camera and IMU data to estimate scale with Example 2
  4. Compare scale during acceleration with scale during constant-speed flight
  5. Write a proposal for choosing perception sensors for a drone inspecting the underside of a bridge

Common mistakes

Watch out

  • Trusting stereo depth at long range
  • Using a single camera with no source of scale
  • Flying at constant speed throughout, so VIO cannot find scale
  • Relying on visual perception over water or smooth ground without a backup
  • Confusing VIO with SLAM and assuming it closes loops

Summary

  • VIO estimates motion from camera and IMU; SLAM adds a map and loop closure
  • Stereo depth , with error growing as the square of range
  • A single camera does not know scale; an IMU can estimate it when there is enough acceleration
  • Choose sensors for the job’s range, lighting and surfaces

Check your understanding

  1. With = 700 px, = 0.1 m and = 14 px, what is the depth?
  2. If range doubles, by what factor does depth error grow?
  3. Why does a single camera not know scale?
  4. A 3% scale error on a 200 m path gives what error?
  5. How does SLAM differ from VIO?
Answers
  1. m
  2. 4 times
  3. A large distant object and a small near object produce the same image
  4. 6 m
  5. SLAM builds a map and closes loops when returning to a place, so it reduces accumulated drift

Key formulas

Stereo depth
Depth error
Least-squares scale

Key references

  1. Szeliski, R. (2022). Computer vision: Algorithms and applications (2nd ed.). Springer. link
  2. Corke, P. (2023). Robotics, vision and control: Fundamental algorithms in Python. Springer. link
  3. Hartley, R., & Zisserman, A. (2004). Multiple view geometry in computer vision (2nd ed.). Cambridge University Press. link
  4. Keselman, L., Woodfill, J. I., Grunnet-Jepsen, A., & Bhowmik, A. (2017). Intel RealSense stereoscopic depth cameras (arXiv:1705.05548). link

Further reading

Study the assigned knowledge units in advance, review media and take the module quiz

In class / field

Lab or field practice from worksheets with a safety checklist

Learning evidence: Checked worksheets and quiz results

Module quiz

This is a formative self-check, not a graded exam

Knowledge domain: Control, autopilot and navigation · Sensors and embedded systems