1/77
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
RGB Colour Space
Three axis 'cube' each axis goes from 0 to 255
HSV Colour Space
Hue, saturation, value
CIE Colour Space
Based on human colour perception
Camera colour space
Evenly distributed colour space
Camera Sensitivity
CCD elements (25% red 50% green 25% blue), to approximate equal sensitivity to red, green, and blue
Lower dynamic range
Wider spectral resolution (CCD elements being sensitive to IR and UV)
Camera resolution
Potential for much higher frame rate, potentially higher resolution.
Human photoreceptors colour space
CIE colour space
Human photoreceptors sensitivity
Eyes have red, green, blue cones, not very sensitive to low light levels. Blind spot at back of eye.
Human photoreceptors resolution
Equivalent to 6.5Mpixel 3 colour camera with a narrow angle lense combined with a peripheral sensitive 100Mpixel monochrome camera with a wide angle lens.
Impact of varying the Kernel size in Canny edge
Small detects fine features, large detects large scale edges
Impact of varying threshold in Canny edge
Faint or strong edges
Canny edge detection
Method to find edges in an image derivative from Gaussian
Non-maximum suppression
Check if pixel is local maximum along gradient direction
Hough transform
Method to detect lines in an image, uses Hough space, every edge point 'votes' for all lines that could pass through it
Classification
Is there a __ present yes or no
Object detection
x y coordinates, bounding boxes, how many
Dense segmentation
What class are all the pixels
Instance segmentation
What class and instance of class are all the pixels
Dense
Most or all of the values are non-zero
Sparse
Most values are zero
Erosion
Getting rid of pixels around edge
Dilation
Adding pixels around edge
Binary open
Erosion then dilation, removes small details such as thin lines, spurs and noise. Smooths jagged edges
Binary close
Dilation then erosion, closes/fills in small gaps/holes and preserves thin lines
Lucas Kanade
Integrate gradients over a patch, looks at pixels around to see if it's a key point
A good local image feature to track should: ...
- Satisfy brightness constancy
- Have sufficient texture variation
- Correspond to a 'real' surface patch
- Not deform too much over time
Harris detector
Captures the structure of the local neighbourhood using an auto-correlation matrix, where 2 strong eigenvalues of this matrix indicate a good local feature
Harris gives a measure of the quality of a feature because the best feature points can be thresholded on the eigenvalues
SIFT
Threshold image gradients are sampled over 16x16 array of locations in a scale space (at 8 different scales/gaussians)
An array of orientation histograms is created at each location (e.g. 8 orientations x 4x4 histogram array = 128 dimensions)
Because SIFT is based on a vector of angles, it is computationally efficient
Harris vs SIFT
Both are illumination and rotation invariant, not deformation invariant
SIFT is scale invariant, sampled at different scales
Harris is not
Translation invariant for x & y motion but not for z
Kalman filter
Estimates the state of a dynamic system over time from noisy measurements
For linear transitions and gaussian distributions, exact solutions obtained
Particle filter
Probabilistic algorithm used for state estimation in nonlinear and non-Gaussian systems
Each particle represents a possible state, and the algorithm updates their weights based on how well each particle matches the actual measurement.
Unscented Kalman Filter
Improves particle and Kalman approximations of non-linear systems, still assumes most common gaussian distributions, balance between low computational effort of Kalman and high performance of particle, easier to initialise
RANSAC
Random sample consensus, estimates parameters of a model from data that contains significant outliers
Randomly sample min number of data points to estimate the model, fit a model to this sample, evaluate how many points are inliers and outliers, repeat, select model with highest number of inliers
Homography
Relates relative pose of 2 cameras (2 frames of a moving camera) viewing a planar scene, estimate from feature correspondences using RANSAC
Essential matrix
Relates relative pose of 2 cameras viewing a 3D scene, estimate from feature correspondences using RANSAC
Bundle adjustment
Initialise RANSAC (for E), estimate a set of 3D points and camera posed which minimises reprojection error
Tracking Prediction
Predict the state in the current frame based on the state of the previous frames. Here the new state is predicted by multiplying the old state by a known constant, and then adding zero-mean noise. Therefore, predicted mean for the new state is constant times mean for old state (also the old variance is a normal random variable where the variance is multiplied by the square of a constant and variance of noise is added)
Tracking Data Association
Calculate the state from the current frame considering kinematic models and error minimisation
Tracking Correction
If the measurement error (Gaussian noise) is low use on the measured state from the current frame, otherwise, use a higher weighting on the predicted state
Two advantages of a particle filter over a Kalman filter
Predict multiple positions
Multi-modal and non-Gaussian
Loss function
A loss function is used to train a model. It needs to be differentiable, but it is not crucial that it is understandable or comparable between tasks/models
Evaluation metric
Used to measure the performance of a model and is often nit differentiable, but it needs to be understandable and comparable
Is the percentage accuracy a good evaluation metric for object detection
No. Even though it is easy to understand it does not capture the range of possibilities. It is ill-defined and ambiguous as to its meaning
A model may be accurate in classification but not accurate in terms of localization. A model may also have low precision (accuracy) but high recall and vice versa.
When would you not use self-supervised or unsupervised learning
Little amount of data, small scale problem
Large amount of computation often required
Why is image correspondence good for self/unsupervised learning?
Correspondence is equivalent to image transformations, images can be transformed without changing the correspondence
Correct correspondences can be verified - for example matching image features can be filtered by geometric verification
More generally, images contain redundant information - for example an auxiliary task such as image completion can be used to learn representations for matching
Image correspondence can be generated, for example with a known or estimated depth or model, synthetic data can be synthesized.
Background subtraction
usually refers to the first frame, or some derivative of it, being the reference frame
Difference Algorithm
usually refers to the difference between two adjacent frames where in this case, the previous frame is the reference frame
Ghosting
the second image of the moving object appearing as an artifact of a difference algorithm
Foreground aperture
refers to a hole appearing in the moving object as an artifact of a difference algorithm
Double difference algorithm
take difference between frame before and current frame, and frame after and current frame, then take commonality between two to get just feature in current frame
Structured light camera disadvantages
Can't work in direct sunlight
Can't work closer than 0.5m
Can't work further than 3.5m
Motion Blur
Low accuracy with greater distance
Structured light camera advantages
Cheap
Can work in the dark
Time of flight camera disadvantages
Can't work in direct sunlight
Limited range due to low intensity infra-red light
Time of flight camera advantages
Accuracy independent of distance
Can work in the dark
Stereo camera disadvantages
Noisy in low light
Accuracy decreases with distance
Need extensive calibration
Gaps in image regions without features
Stereo camera advantages
Potential for highest resolution
Works well in sunlight
Works for motion
LIDAR disadvantages
Low resolution
Low frame rate
Has moving parts
Expensive
LIDAR advantages
Good range
Accuracy independent of distance
Fiducial marker advantages
Less computationally intensive
High accuracy and robustness
Easy to detect
Natural feature advantages
No need for marker
Natural features catch less attention
Natural features work even if partially showing
Can use existing data
Fiducial marker disadvantages
Requires placing physical markers in the scene.
Limited to where markers are visible.
Can fail if markers are occluded or damaged.
Natural feature disadvantages
Sensitive to lighting, blur, motion, or lack of texture.
Requires robust feature detectors (e.g., SIFT, ORB).
Can struggle in feature-poor environments (e.g., white walls).
How is Canny edge detector good at responding to edge not noise
The Canny edge detector used a filter based on the first derivative of a Gaussian, because Canny is susceptible to noise present on raw unprocessed image data, so to begin with, the raw image is convolved with a Gaussian filter. The result is a slightly blurred version of the original which is not affected by a single noisy pixel to any significant degree.
How is Canny edge detector good at detecting edges near the true edge
Uses hysteresis to improve localisation (checks that maximum value of gradient value is sufficiently large). Uses a high threshold to start edge and a low threshold to continue them.
How is Canny edge detector good at providing one response per edge
Uses non-maximum suppression for thinning (checks if pixel is local maximum along gradient direction).
How do two correctly matched features enable finding depth values in stereo camera
The 'x' distance between a matching pair of points is called the disparity. The larger the disparity, the closer that point is to the camera based on triangulation (but this is not linear).
How do two correctly matched features enable finding optical flow points in Lucas Kanade
The Lucas-Kanade method integrates gradients over a patch to find features good enough to track using the Harris detector.
A constant velocity is assumed for all pixels within an image patch. Optical flow is the measure of the movement that feature points undergo in successive frames.
How can depth be calculated from optical flow
Relative depth can be calculated from the velocity of optical flow points - which is larger when the distance of the camera to the tracked points is less.
So absolute depth could be determined if the camera velocity is known (or distance between camera locations for two successive frames).
When is a loss function and an evaluation metric the same
Stereo matching or optical flow: mean squared error
Segmentation: Dice or F1 score
When is a loss function and an evaluation metric different
Classification: cross entropy/percent accuracy
Object detection: box regression + cross entropy / mean average precision
Six image transformations and deformations
Translation
Rotation
Illumination
Blur
Noise
Partial Occlusion
Non-rigid deformation (warping, barrel, pin cushion)
Human spectral resolution
400 (violet) - 700nm (red)
Human dynamic range
10^8:1 (ratio between lowest perceptible light intensity to highest)
Human spatial resolution at 20m
1-3cm
Human radiometric resolution
100 colours, 16-32 shades black and white
Accuracy
how often does the model correctly predict the outcome / how often does the model perform across all object classes
Precision
When the model predicted the positive class, what percentage of predictions were correct
Recall
When the ground truth was positive, what percentage of predictions were correctly identified as positive