1/174
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
What is a sensory gap?
Between the object in the world and the information in the computational description derived from a recording of that scene
What is a semantic gap?
Lack of coincidence between the information that one can extract from the visual data and the interpretation that the same data can have for a user in the given situation
How do you recover 3D information?
Motion, stereo, texture, shading, contour, time of flight
What are the 4 stages of CV processing?
Simulate human image processing, pre processing, early processing, late processing.
What are 4 examples of low level image processing
Image compression, Noise reduction, edge detection
segmentation
What is spectral resolution?
the ability of a sensor to define fine wavelength intervals
Human eye - spectral resolution
400 (violet) - 700mm (red)
What is dynamic range (eye)
Difference between the lowest perceptible light intensity and the highest intensity we can tolerate without glare.
Human eye - dynamic range
10^8 : 1
What is spatial resolution?
The ability to distinguish two separate objects close together
Human eye - spatial resolution
1-3cm @ 20m
What is radiometric resolution?
The ability to describe very small energy changes
Human eye - radiometric resolution
16-32 shades b&w 100 colours
What does the eye work?
Light enters eye, focused by cornea and lens and strikes the retina at the back of the eye. 2 cells that are stimulated: rods and cones
What are cones?
Provide vision in bright light, color vision, sharp images. Closely packed. Shorter and thicker
What are rods?
retinal receptors that detect black, white, and gray; necessary for peripheral and twilight vision, when cones don't respond. 100 megapixel camera. Scotopic - most active in the short wavelength
How many cones are in the human eye?
6.5 million
How many rods are in the human eye?
100 million
What is the distribution of rods and cones in the retina?
They are not evenly distributed.
Where is the density of cones greatest in the eye?
At the fovea
What is the fovea known for?
It is the region of sharpest vision.
Why is there a blindspot in the eye?
No receptors because this is where the ganglion cells leave the eye to form the optic nerve.
Why specify colour numerically?
Representation is commercially valuable
CIE colour space
A standardised colour space where any colour can be represented as (xy) coordinates.
What's wrong with CIE?
Doesn't describe the way light interacts in reality. Instead it is simple in the way that it models how the human eye perceives.
What is HSV?
hue, saturation, value.
Why use lenses?
To collect more light
Pinhole camera
Perfect in focus image with lenses.
0.35MM is the sweet spot
Vignetting
Light at different wave lengths. Follows different paths, some wavelengths are defocused.
Pros of multi-plane calibration (Checkerboard)
Only requires a plane. Dont have to know positions/orientations
What does adding a lens do?
Focus light onto the film
Depth of field
Changing aperture size affects it.
Focal length
Adjusts the zoom
Exposure time
How long an image is exposed
ISO
sensitive of the "film". Proportional to noise.
How is distortion of an image caused?
Imperfect lenses
What is image filtering?
Modify the pixels in an image based on some function of the local neighbourhood of the pixels.
Linear functions (filters)
Replace each pixel by a linear combinations of it's neighbours.
What is a convolution kernal?
The prescription for the linear comb. (linear function filter)
What is a box filter?
Averaging filter, blurring filter
What does a box filter do?
Replaces each pixel with an average of its neighborhood
What does a box filter achieve?
Smoothing effect (remove sharp features)
What happens when the pixel offsets is positive?
Shifts image to the left
What happens when the middle left value is 1 in a box filter?
Shifts image left by 1 pixel
What happens when pixel offsets are evenly distributed?
Blur
Implication of smoothing and noise
Implies that smoothing suppresses noise, for appropriate noise models.
What would 2 offset - 0.33 (3 total bars) achieve?
Sharping. We took away information but it made it better for humans.
What are edge points?
Points of sharp change in an image
Edge detection
Convert a 2D image into a set of curves
What causes edges?
Surface normal discontinuity, Depth surface colour, illumination
How can you tell that a pixel is on an edge?
Change to black and white
What is an edge?
Where change occurs
Sobel filter
Used in image processing/computer vision.
Creates and image that emphasises edges.
Image gradiant
Points in the direction of rapid change in intensity
Can be used for edge strength
How can we differentiate a digital image
Reconstruct a continuous image, then take grad
Take discrete derivative (finite difference)
Optimal edge detection - Canny
Assume: Linear filtering
Additive Gaussian nose
What should a good edge detector have
Good detection: filter for edge not noise
Good localisation: detected edge near true edge
Single response: one per edge
Detection/localisation trade off
More smoothing improve detections
And hurts localization
Non-maxiumum suppression
Check if pixel is local maximum along grad direction
Predicating the next edge point
Marked point is edge point. Create tangent to the edge curve (Norm) use this to predict next points.
Hysteresis
Check that maximum value of grad value is large.
Choice of sigma in guassian kernal size
Large to detect large scale edges
Small to detect fine features
Finding lines in images
Search for the line at every possible position/orientation
Use voting scheme: Hough transform
How to do Hough transform
Connection between image (x, y) and Hough (m, b) spaces
Line in image to point in Hough space
Finding Corners Intuition
Right at corner, grad is ill defined
Near corner, grad has two different values
How to detect corners
Filter image, compute magnitude, construct C in a window, find eigen values, if both are large means corner
Whats a good feature?
Satisfies brightness constancy, has sufficient texture variation, does not have too much texture variation, corresponds to a "real" surface patch, does not deform too much over time.
Feature distortion
Feature may change shape over time, need a distortion model to really make this work???
Harris detector
A corner detection algorithm. It IDs corners by doing local neighbourhood of each pixel. Used for feature detection.
What is Harris detector invariant to?
Rotation and illumination changes
Advantages of invariant local features
Locality
Distinctiveness
Quantity
Efficiency
Extensibility
SIFT vector formation
Thresholded images are sampled over 16x16 in scale space
Create array of orientation of histograms (For image gradients)
8 orientations x 4x4 histogram array = 128 diamensions
Erosion
Shrinks an object
Dilation
Expands an object
Open
Erosion then dilation
Close
Dilation then erosion
What does opening do?
Smoothes regions
What doe closing do?
Fills gaps
Greyscale erode?
Output at a point is minimum of image pixel and structuring element pixel
Greyscale Dilate
Output is maximum of image and structuring element
Skeleton
Reduces regions of binary image to lines one pixel thick
What does skeleton preserve
Shape, continuity
Thinning algorithm
Repeatedly thin image
Retain end points and connections
Distance Transform algorithm
Skeleton lies along discontinuities
Sort of local maxima or ridges
Application of Thinning and distance transform
Shape representation, maintaining topology
Character recognition
Tracking applications
Motion capture, recognition from motion, surveillance, targeting
Tracking and recursive estimation
Need a real-time efficient algorithm. The task is at each time point, recompute estimate of position
recursive estimation
Estimate position of a tracked object at each time from all previous data.
What is the first main issue in tracking?
Prediction: determining what set of measurements predict for the ith frame and finding p(x|y).
What is the second main issue in tracking?
Data association: using p(x|y) to identify measurements that tell us about the object's state.
What is the third main issue in tracking?
Correction: computing p(x|y) after obtaining yi.
How do we simplify tracking assumptions
Only the immediate past matter, measurements depend only on the current state
Kalman filter
the best estimate of the current position can be obtained by predicting the position using the initial position and the time that has passed and combining this estimate with the noisy measurements of the sensors
Prediction for 1D Kalman filter
Old state * constant + noise
Kalman filter smoothing
We don't have the best estimate of state - what abou the future?
Run two filters, one forward and back in time. Combine these estimates
n-D Kalman filter
Multiple estimate at prior time with forward model and propagate covariance through model and add new noise
Particle filter
Predict multiple positions, multi-model and non-gaussian
Particle filter - specification
Prior density, p(xt-1), Process density p(xt | xt-1), observation density p(zt-1 | xt-1)
Particle filter - Prior density
Joint angles in prev frame
Particle filter - Process density
Kinematic and body models