COSC428 (2026)

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/116

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 5:00 AM on 6/5/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

117 Terms

1
New cards

What are the strengths of the CIE colour space?

Colours are perceptually uniform; conceptually easier to mix colours in this space.

2
New cards

What are the weaknesses of the CIE colour space?

Challenging to use for computer vision because CIE is based on human perception and was originally intended for humans to subjectively compare colours rather than image processing. Some coordinates don't represent real colours.

3
New cards

What are the applications of the CIE colour space?

Colour temperature of lighting for photographers; subjectively comparing food colours (e.g. chips).

4
New cards

What are the strengths of the RGB colour space?

Used to represent colour by media such as displays and cameras — immediately available to computer vision algorithms. Unit cube so all possible RGB values are realisable, which simplifies range checking of red, green and blue values.

5
New cards

What are the weaknesses of the RGB colour space?

Not all colours are perceptually uniform — thus it doesn't make sense to calculate colour differences in RGB. Different RGB values are needed to produce the same colour on different displays (device specific).

6
New cards

What are the applications of the RGB colour space?

Computer graphics.

7
New cards

What are the strengths of the HSV colour space?

More useful than RGB for analysing colour, e.g. simplifying colour range checking. Simple transformation of RGB.

8
New cards

What are the weaknesses of the HSV colour space?

Not perceptually uniform; device specific.

9
New cards

What are the applications of the HSV colour space?

Used by artists (e.g. Photoshop) as it is more natural to think about colour in terms of hue and saturation. Computer vision based colour analysis.

10
New cards

How does a camera's colour space differ from the human retina?

Camera uses RGB colour space with evenly distributed CCD elements (25% red, 50% green, 25% blue). The human retina resembles the CIE colour space with non-uniform distribution of red, green and blue cones.

11
New cards

How does camera sensitivity differ from the human eye?

Camera has much lower dynamic range but wider spectral resolution (sensitive to infra-red and ultra-violet) and a much higher frame rate. The eye has a 10^8:1 dynamic range but brain processing limits it to ~100 colours and ~16-32 shades of B&W.

12
New cards

How does camera resolution differ from the human eye?

Camera can have higher spatial resolution potential. The eye is equivalent to a foveal 6.5 Mpixel/3-colour narrow-angle camera combined with a peripheral 100 Mpixel monochrome wide-angle camera, but limited to 1-3 cm resolution at 20 m.

13
New cards

How does a Laplacian of Gaussian filter sharpen an image?

The filter subtracts the low frequencies (blurred image) from the original image, leaving the high frequencies (edges) remaining as a sharpened image with higher contrast.

14
New cards

Why does a sharpened image appear to have more content than the original?

Although there is less content (no low frequencies) in the sharpened image, the accentuated high-frequency edges give the illusion of more content because there appear to be more edges and human perception is sensitive to edges.

15
New cards

What is erosion in morphological image processing?

Removes outside pixels of a region/blob (and internal holes/regions), usually using a convolution kernel in an AND operation. Removes small details such as thin lines and noise points, widens gaps, and shrinks a region.

16
New cards

What is dilation in morphological image processing?

Adds pixels to the outside of a region/blob using a convolution kernel in an OR operation. Enlarges a region/blob, thickens lines, fills small holes.

17
New cards

What is morphological opening?

Erode then dilate. Removes small details such as thin lines, spurs and noise. Smooths jagged edges without changing the size of the original object.

18
New cards

What is morphological closing?

Dilate then erode. Closes/fills in small gaps/holes and preserves thin lines without changing the size of the original object.

19
New cards

Describe the prediction step of the Kalman filter.

Predict the state in the current frame based on state in previous frames. The new state is predicted by multiplying the old state by a known constant and adding zero-mean noise.

20
New cards

Describe the data association step of the Kalman filter.

Calculate the state from the current frame considering kinematic models and error minimisation.

21
New cards

Describe the correction step of the Kalman filter.

If measurement error (Gaussian noise) is low, use the measured state from the current frame; otherwise use a higher weighting on the predicted state.

22
New cards

How does a Kalman filter provide a smoothed estimate?

Run two Kalman filters — one moving forward and one backward in time. Combine state estimates by viewing the backward filter's prediction as yet another measurement for the forward filter.

23
New cards

What are two advantages of a Particle Filter over a Kalman Filter?

1) Can predict multiple positions simultaneously. 2) Supports multi-modal and non-Gaussian distributions.

24
New cards

What is supervised learning? Give a computer vision example.

Learns a function from inputs to outputs based on labelled data pairs. Example: road sign classification from images using a labelled dataset.

25
New cards

What is weakly supervised learning? Give a computer vision example.

Utilises label-data examples where the labels are weaker (noisy, limited or imprecise). Example: an object detector trained on images with class labels but no bounding box annotations.

26
New cards

What is semi-supervised learning? Give a computer vision example.

Supervised learning with only a small amount of labelled data but a larger set of unlabelled data. Example: train on supervised data, use model to assign labels to unsupervised data, retrain on all data and iterate.

27
New cards

What is self-supervised / unsupervised learning? Give a computer vision example.

Uses known properties of the data to provide a supervision signal. Example: use an auxiliary task like image completion to learn a mapping from an image to a feature vector for similarity metrics.

28
New cards

What is the difference between a loss function and an evaluation metric?

A loss function is used to train a model — it must be differentiable but need not be human-interpretable. An evaluation metric measures model performance — it need not be differentiable but must be understandable and comparable across models/tasks.

29
New cards

Why is percentage accuracy a poor metric for object detection?

It is ill-defined without specifying precision, recall, confidence threshold, and spatial matching threshold. A model may be accurate in classification but poor in localisation. A better metric is mean Average Precision (mAP) which captures both precision and recall.

30
New cards

Give two reasons you might NOT use self/unsupervised learning.

1) Small dataset or small-scale problem where supervised approaches are sufficient. 2) Very large computational requirements needed to train self-supervised models.

31
New cards

What three properties make image correspondence a good candidate for self-supervised learning?

1) Correspondence is equivariant to image transformations. 2) Correct correspondences can be verified (e.g. geometric verification). 3) Hand-labelling pixel correspondences is extremely laborious or impossible.

32
New cards

Why is the unscented Kalman filter better than the standard Kalman filter and particle filter?

It improves approximations for non-linear systems while still assuming Gaussian distributions. It provides a balance between the low computational cost of the Kalman filter and the high performance of the particle filter, and is easier to initialise.

33
New cards

What are the four probability distributions used by a particle filter?

1) Prior density — state values in previous frame. 2) Process density — kinematic models. 3) Observation density — image in previous frame. 4) State density — state values in next frame.

34
New cards

In camera projection, what do points, lines, planes, polygons and polyhedra map to?

Points → points. Lines → lines. Planes → whole image or half-planes. Polygons → polygons. Polyhedra → polygons.

35
New cards

In degenerate projection through the focal point, what do lines, planes, polygons and polyhedra map to?

Lines → points. Planes → lines. Polygons → lines. Polyhedra → planes.

36
New cards

In a spectral power distribution graph, which light source peaks at 400-500 nm and which rises toward red?

A clear sky peaks at 400-500 nm (blue dominant). A tungsten light bulb rises from low blue across the spectrum to dominant red/yellow-white.

37
New cards

Which filter — averaging, Gaussian or median — removes noise without blurring edges, and why?

The median filter. It replaces each pixel with the median of its neighbourhood, so it eliminates noise and small objects while preserving sharp edges.

38
New cards

Match low-pass, band-pass and high-pass filters to their effect on a striped image.

Low-pass → blurs sharp edges (removes high frequencies). High-pass → sharpens image by removing gradual low-frequency changes. Band-pass → removes both sharp edges and gradual changes.

39
New cards

What is the aperture problem in computer vision?

If only part of a moving object is visible through the camera aperture, the components of its motion vector cannot be unambiguously determined — only the component perpendicular to an edge can be measured.

40
New cards

What are four sources of gross outliers in feature matching for 3D reconstruction?

1) Self-similarities (repeated textures). 2) Occlusion boundaries. 3) Shadows. 4) Feature motion / motion blur / measurement noise.

41
New cards

Describe the Gaussian image pyramid and its use.

Progressively blurred and subsampled versions of the image. Adds scale invariance to fixed-size algorithms.

42
New cards

Describe the Laplacian image pyramid and its use.

Shows the information added in the Gaussian pyramid at each spatial scale. Useful for noise reduction and image coding.

43
New cards

Describe the Wavelet/QMF image pyramid and its use.

A bandpassed representation that is complete but has aliasing and some non-oriented subbands.

44
New cards

Describe the Steerable pyramid and its use.

Shows image components at each scale and orientation separately with non-aliased subbands. Good for texture and feature analysis.

45
New cards

What are the advantages and disadvantages of fiducial marker tracking?

Pros: Cheap, easy way to roughly estimate 6DOF pose. Cons: Environment must be instrumented; sensitive to image noise, blur, poor lighting, vignetting; jittering of overlaid graphics; fails with partial occlusion.

46
New cards

What are the advantages and disadvantages of markerless tracking?

Pros: Natural targets catch less attention, are potentially everywhere, and work if partially in view. Cons: More difficult and computationally intensive than marker tracking; requires storing a keypoint database.

47
New cards

What is the difference between invariance and equivariance in a machine learning model?

Invariant: outputs do not change when inputs are transformed. Equivariant: outputs change proportionately to the transformation applied to the inputs.

48
New cards

Give four invariances relevant to image classification.

Brightness, contrast, translation, scale, viewpoint, rotation, focus (any four).

49
New cards

Why is it useful for an object detection model to use invariant or equivariant operations?

If the operations match the invariances of the target problem, the dimensionality of the learning problem is reduced. E.g. CNN convolutions are equivariant to translation, so object translation is already accounted for.

50
New cards

What invariance does a typical CNN detector lack, and what is the consequence?

A CNN is not invariant to changes in object scale. Therefore more training data showing objects at different scales is required.

51
New cards

What is the problem with using skin colour as a manifold in RGB for face localisation?

While skin colours occupy a small 3D manifold in RGB, many non-skin colours also fall in that region (e.g. wood, sand). Lighting changes shift the manifold, causing false positives and missed detections.

52
New cards

Why is it problematic to only partially annotate objects in training images?

Most object detection frameworks treat unannotated objects as implicit negative examples. Leaving sheep unannotated while annotating only a few contradicts the positive examples, making training fail or perform poorly.

53
New cards

If a model's loss function oscillates and does not decrease after several epochs, what should you do?

Stop training — no improvement will come with more time. Investigate causes: a bug in training code, bad hyperparameters (learning rate too high/low), or corrupt data/preprocessing.

54
New cards

What is the Abbe number and what does it indicate?

The Abbe number (V) quantifies how dispersive a glass material is: V = (n_d - 1) / (n_F - n_C). A high Abbe number means low dispersion and less chromatic aberration. Crown glass (~60) disperses less than flint glass (~35).

55
New cards

What is longitudinal chromatic aberration and where does it appear in an image?

Different wavelengths (colours) focus at different distances along the optical axis. Blue focuses closer to the lens, red further away. Affects the entire image including the centre, appearing as colour halos around objects.

56
New cards

What is lateral (transverse) chromatic aberration and where does it appear?

Different colours are magnified by different amounts, causing them to land at different heights on the focal plane. Only affects the edges/corners of the image as colour fringing along high-contrast edges. Easier to correct in post-processing than longitudinal CA.

57
New cards

What is an achromatic doublet and how does it reduce chromatic aberration?

A converging crown glass element cemented to a diverging flint glass element. They have different dispersion curves so their chromatic errors partially cancel, while overall focusing power is maintained. Brings two wavelengths to the same focal point.

58
New cards

What is an apochromat lens?

A more advanced lens design that brings three wavelengths to the same focal point. Used in high-end microscope and camera lenses. More expensive than achromatic doublets.

59
New cards

What are the four types of vignetting?

1) Optical vignetting — lens barrel blocks oblique rays at wide apertures. 2) Mechanical vignetting — physical obstructions like lens hoods or stacked filters. 3) Natural (cos⁴) vignetting — fundamental physics; illuminance falls as cos⁴(θ). 4) Pixel vignetting — oblique light partially misses photosite wells on digital sensors.

60
New cards

State the cos⁴ law of natural vignetting and explain each contributing factor.

E(θ) = E₀·cos⁴(θ). Four cos(θ) terms: (1) reduced projected aperture area, (2) longer path length to sensor, (3) reduced solid angle of lens seen from image point, (4) angle of incidence on sensor surface.

61
New cards

How does stopping down (increasing f-number) affect optical vignetting?

Stopping down greatly reduces optical vignetting because the cone of light from off-axis points narrows, making it less likely to be clipped by the lens barrel. Natural (cos⁴) vignetting is unaffected by aperture.

62
New cards

What is the thick lens lensmaker's equation and what does the extra term represent?

1/f = (n-1)[1/R₁ - 1/R₂ + (n-1)d / (n·R₁·R₂)]. The extra term (n-1)d/(n·R₁·R₂) is the thickness correction — it increases 1/f, meaning a thicker lens has a shorter focal length (focal point moves closer).

63
New cards

What are the two principal planes in a thick lens and why do they matter?

H (front principal plane) and H' (rear principal plane). Focal length is measured from these planes, not the physical surfaces. They allow a thick lens to be treated like a thin lens for ray tracing, but the image distance and object distance must be measured from H and H' respectively.

64
New cards

What is the difference between a low-pass and a high-pass spatial filter?

A low-pass filter retains low spatial frequencies (smooth regions, gradual changes) and removes high frequencies (edges, fine detail) — blurring the image. A high-pass filter does the opposite, retaining edges and removing smooth regions — sharpening the image.

65
New cards

What is a Gaussian filter and what is its main advantage over a simple averaging filter?

A Gaussian filter weights pixels by a Gaussian function of their distance from the centre. Its main advantage is that it does not introduce ringing artefacts (it has no side lobes in the frequency domain) and produces a smooth, natural blur compared to the box-like averaging filter.

66
New cards

What is the median filter and when is it preferred over Gaussian smoothing?

The median filter replaces each pixel with the median value of its neighbourhood. It is preferred when the noise is impulsive (salt-and-pepper) or when edge preservation is important, since it removes noise without blurring edges.

67
New cards

What is the Nyquist theorem and why does it matter for image sampling?

A signal must be sampled at least twice the highest frequency it contains to be reconstructed without aliasing. For images, if the sampling rate (pixel density) is too low relative to the spatial frequency of features, aliasing artefacts appear (e.g. moiré patterns).

68
New cards

What is aliasing in images and how can it be prevented?

Aliasing occurs when high spatial frequencies are undersampled and appear as lower-frequency artefacts. Prevention: apply a low-pass (anti-aliasing) filter before downsampling to remove frequencies above the Nyquist limit.

69
New cards

What are the three main steps of the Canny edge detector?

1) Smooth image with a Gaussian filter to reduce noise. 2) Compute gradient magnitude and direction. 3) Apply non-maximum suppression and hysteresis thresholding to thin edges and remove weak/false edges.

70
New cards

What is non-maximum suppression in the context of edge detection?

After computing the gradient magnitude, pixels that are not local maxima along the gradient direction are suppressed (set to zero). This thins the edges to single-pixel width.

71
New cards

What is hysteresis thresholding in the Canny detector?

Two thresholds are used: high and low. Pixels above the high threshold are definite edges. Pixels between the thresholds are only kept if they are connected to a definite edge. Pixels below the low threshold are discarded. This reduces noise while preserving continuous edges.

72
New cards

What is the Sobel operator and what does it compute?

The Sobel operator uses two 3×3 kernels to compute the image gradient in the x and y directions. The gradient magnitude (sqrt(Gx²+Gy²)) estimates edge strength and the angle (atan2(Gy,Gx)) gives edge orientation.

73
New cards

What is the Laplacian operator and how is it used for edge detection?

The Laplacian computes the second derivative of image intensity. Edges correspond to zero-crossings of the Laplacian. It is sensitive to noise so is typically combined with Gaussian smoothing (Laplacian of Gaussian, LoG).

74
New cards

What properties should a good image feature (keypoint) have?

Repeatability (detected under different conditions), distinctiveness (unique descriptor), locality (small region, robust to occlusion), quantity (enough features for matching), accuracy (precise location), efficiency (fast to compute).

75
New cards

What is the SIFT descriptor and why is it scale and rotation invariant?

Scale-Invariant Feature Transform. Detects keypoints at multiple scales using a Difference of Gaussians. Assigns a dominant orientation based on local gradients. The descriptor is a 128-dimensional histogram of gradients computed relative to the keypoint's scale and orientation, making it invariant to both.

76
New cards

What is the Harris corner detector?

Detects corners by analysing the second-moment matrix (structure tensor) of local image gradients. A corner is a point where intensity changes significantly in multiple directions. The Harris response R = det(M) - k·trace²(M) is large and positive at corners.

77
New cards

What is RANSAC and what is it used for in computer vision?

Random Sample Consensus. An iterative algorithm for robustly fitting a model to data containing outliers. Randomly samples a minimal set of points to fit a model, counts inliers (points within a threshold), and repeats. Used in homography estimation, fundamental matrix computation, and pose estimation.

78
New cards

What is the epipolar constraint and why is it useful?

Given a point in one image, its correspondence in the other image must lie on a specific line (the epipolar line) in the second image. This reduces the 2D correspondence search to a 1D search, greatly reducing computation and false matches.

79
New cards

What is the fundamental matrix F?

A 3×3 rank-2 matrix encoding the epipolar geometry between two uncalibrated cameras. For corresponding points x and x', x'ᵀFx = 0. It can be estimated from point correspondences alone without knowing camera intrinsics.

80
New cards

What is the essential matrix E and how does it differ from the fundamental matrix?

Like the fundamental matrix but for calibrated cameras (intrinsics known). E = Kᵀ F K' where K and K' are the intrinsic matrices. It encodes the relative rotation and translation between cameras and has only 5 degrees of freedom.

81
New cards

What is disparity in stereo vision and how does it relate to depth?

Disparity is the horizontal pixel difference between corresponding points in left and right images. Depth Z = f·B/d, where f is focal length, B is baseline (distance between cameras) and d is disparity. Larger disparity → closer object.

82
New cards

What is Structure from Motion (SfM)?

A technique that recovers 3D structure and camera poses simultaneously from a sequence of 2D images taken from different viewpoints. Requires feature matching across views, estimation of relative camera poses, and triangulation of 3D points, followed by bundle adjustment to minimise reprojection error.

83
New cards

What is optical flow?

The apparent motion of brightness patterns in an image sequence due to relative motion between the camera and the scene. It represents the velocity field of pixels between consecutive frames.

84
New cards

State the optical flow constraint equation and explain its terms.

Iₓu + Iᵧv + Iₜ = 0, where Iₓ, Iᵧ are spatial image gradients, Iₜ is the temporal gradient, and (u, v) is the optical flow vector. This one equation has two unknowns — the aperture problem.

85
New cards

What assumptions does the Lucas-Kanade optical flow method make?

1) Brightness constancy: pixel intensity does not change between frames. 2) Small motion: displacement is small enough for linearisation. 3) Spatial coherence: neighbouring pixels have the same flow. It solves the overconstrained system using a least-squares fit over a local window.

86
New cards

What is the Horn-Schunck optical flow method and how does it differ from Lucas-Kanade?

A global method that adds a smoothness regularisation term over the entire image, penalising large flow gradients. Unlike Lucas-Kanade (local window), it produces a dense, globally smooth flow field but is more sensitive to discontinuities at motion boundaries.

87
New cards

What is a convolutional neural network (CNN) and why is it suited to image tasks?

A CNN uses layers of learnable convolutional filters applied across the spatial dimensions of an image. Weight sharing across positions makes it equivariant to translation, greatly reducing parameters. Pooling layers add limited scale and spatial invariance. Hierarchical layers learn increasingly abstract features.

88
New cards

What is the difference between object detection and image classification?

Image classification assigns a single label to the entire image. Object detection both classifies and localises multiple objects in the image, outputting class labels and bounding boxes for each detected instance.

89
New cards

What is mean Average Precision (mAP) and why is it used for object detection?

mAP is the mean of Average Precision (area under precision-recall curve) across all object classes. It simultaneously captures precision and recall at multiple detection thresholds, providing a single comprehensive metric that accounts for both false positives and false detections.

90
New cards

What is non-maximum suppression (NMS) in object detection?

After a detector produces many overlapping bounding boxes, NMS removes redundant boxes. It iteratively selects the highest-scoring box, then removes all other boxes with IoU above a threshold. This ensures each object is detected only once.

91
New cards

What is Intersection over Union (IoU) and how is it used in object detection evaluation?

IoU = Area of Overlap / Area of Union between a predicted and ground-truth bounding box. A detection is counted as a true positive if IoU ≥ a threshold (commonly 0.5). Used in NMS and in computing Average Precision.

92
New cards

What is transfer learning and why is it commonly used in computer vision?

Using a model pre-trained on a large dataset (e.g. ImageNet) as a starting point for a new task. The pre-trained weights encode general visual features (edges, textures, shapes) that are useful for many tasks, allowing good performance with far less labelled data and training time.

93
New cards

What is data augmentation and why is it used in training CNNs?

Artificially expanding the training set by applying random transformations to existing images (flipping, rotation, cropping, colour jitter, etc.). It improves generalisation by exposing the model to more variation and reducing overfitting.

94
New cards

What is overfitting and how can it be detected and prevented?

Overfitting occurs when a model learns the training data too well but fails to generalise to unseen data. Detected by a large gap between training and validation loss/accuracy. Prevention: more data, data augmentation, dropout, regularisation (L1/L2), early stopping, or a simpler model.

95
New cards

What is batch normalisation and what problem does it solve?

Normalises activations within each mini-batch to have zero mean and unit variance, then applies learnable scale and shift parameters. Solves internal covariate shift (changing input distributions during training), allowing higher learning rates and more stable, faster training.

96
New cards

What is the Hough transform and what is it used for?

A feature extraction technique that detects parameterised shapes (lines, circles, ellipses) in images. Each edge pixel votes for all parameter combinations consistent with it. The accumulator cell with the most votes corresponds to the most supported shape.

97
New cards

How does the Hough transform detect lines?

Each edge pixel (x, y) votes in the parameter space (ρ, θ) for all lines passing through it: ρ = x·cos(θ) + y·sin(θ). Peaks in the accumulator correspond to lines in the image.

98
New cards

How does the Hough transform detect circles?

Each edge pixel votes for all circles of a given radius that pass through it. The centre (a, b) of a circle with radius r satisfies (x-a)² + (y-b)² = r². Peaks in the 3D accumulator (a, b, r) correspond to circles.

99
New cards

What is background subtraction and what is it used for?

A technique that models the background of a scene and subtracts it from each frame to isolate moving foreground objects. Used in surveillance, traffic monitoring, and activity detection.

100
New cards

What is the Gaussian Mixture Model (GMM) approach to background subtraction (MOG2)?

Each pixel's intensity over time is modelled as a mixture of Gaussians. Pixels well-explained by background Gaussians are classified as background; others as foreground. The model adapts over time to handle gradual lighting changes.