1/57
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Image sub-sampling definition
throw away every other row and column to create a ½ size image
Image sub-sampling effects (2)
aliasing/artifacts appear because damaging certain frequencies.
sampling no longer double highest frequency
How to limit artifacts on Image sub-sampling
gaussian (low-pass) prefiltering
Gaussian (lowpass) Pre-filtering
filter the image, then subsample
for every image reduction of s=0.5, smooth by sigma = 1
sigma inversely proportional to image reduction
sigma = 1 / 2s
Image Pyramid Definition
a collection of representations of an image, each layer of the pyramid is half width and half height of the previous layer
Gaussian Pyramid Definition
each pyramid layer is smoothed by a Gaussian filter and resampled to get next layer
What happens to details in Gaussian Pyramid?
they get smoothed out as we move higher levels because only low frequency info remains
What is preserved at the higher levels in Gaussian Pyramid?
mostly large uniform regions in original image
How would you reconstruct the original image from the image at the upper level in Gaussian Pyramid?
not possible because each layer removes high frequency information
What is an Image/Gaussian Pyramid good for? (3)
coarse-to-fine search
search over translations (efficient localization based on searching coarse scales first
search over scale (template matching, find face at different scales)
pre-computation
need to access image at different blur levels
useful for mip-mapping (texture mapping at different resolutions)
deep learning
cnn already produces a pyramid - each pooling stage halves the resolution
diffusion models and super-resolution work coarse-to-fine by construction
For generic images, what would be good features to detect for?
Edges
Origin of Edges (4)
surface normal discontinuity
depth discontinuity
surface color discontinuity
illumination discontinuity
Information theory view
edges encode change and change is what is hard to predict - so edges encode an image efficiently
Edge Detection Definition
convert a 2d image into a set of curves
extracts major features of the scene
more compact than pixels
What to edges look like in images as functions?
steep cliffs
Edge Detection Basic Idea
look for a neighborhood with strong signs of change
Problems with Edge Detection (2)
neighborhood size
how to detect change
Edges in terms of image intensity function
place of rapid change
Differential Operators Definition
some operation that when applied to the image returns some derivatives
How to model differential operators?
as masks/kernels which when applied to the image yields the image gradient function.
then threshold this gradient function to select the edge pixels
Image Gradient Definition (2)
points in the direction of the most rapid increase in intensity.
edge strength is given by gradient magnitude
Image Gradient Equation/Points (3)
∇f = [ (∂f / ∂x), (∂f / ∂y) ]
horizontal: ∇f = [ (∂f / ∂x), 0 ]
vertical: ∇f = [ 0, (∂f / ∂y) ]
Gradient Direction Equation
θ = tan-1 ( (∂f / ∂y) / (∂f / ∂x) )
Edge Strength equation
||∇f|| = sqrt{ (∂f / ∂x)2 + (∂f / ∂y)2 }
For a 2D function, f(x, y), the partial derivative is:
(∂f(x, y) / ∂x) = lim{ε→0} ( f(x+ε, y) - f(x, y) ) / ε
For a 2D function, f(x, y), partial derivative approximation:
using finite differences (smallest step of 1)
(∂f(x, y) / ∂x) ~= f(x+1, y) - f(x, y)
the discrete gradient is
average of “left” and “right” derivative
Sobel Operator Matrixes
Sx = 1/8 [ [ -1 0 1 ], [ -2 0 2 ], [ -1 0 -1]]
Sy = 1/8 [ [ 1 2 1 ], [ 0 0 0 ], [ -1 -2 -1]]
Sobel Operator Gradient Funtions
gx = response to mask Sx
gy = response to mask Sy
Gradient: ∇I = [gx, gy]T
Gradient Magnitude: g = (gx2 + gy2)1/2
Gradient Direction = θ = atan2(gy, gx)
Prewitt Mask Matrix
Sx = [ -1 0 1 ], [ -1 0 1 ], [ -1 0 -1]]
Sy = [ [ 1 1 1 ], [ 0 0 0 ], [ -1 -1 -1]]
Roberts Mask Matrix
Better for Diagonal
Sx = [ [ 0 1 ], [ -1 0 ]]
Sy = [ [ 1 0 ], [ 0 -1 ]]
Why would it be difficult to find the edge of after a derivative operator?
a noisy image’s high frequencies would be emphasized
How to find edge of noisy image?
smooth first, then take derivative.
Measures the rate of change of pixel intensity.
Locates an edge where the first derivative has a peak (local maximum or minimum)
f = signal
h = kernel
h * f = convolution
(∂ / ∂x) (h * f) = differentiation
Derivative Theorem of convolution
saves 1 operation
f = signal
(∂ / ∂x) h = derivative of Gaussian kernel
((∂ / ∂x) h) * f = convolution
2nd derivative of Gaussian
Measures the change in the rate of change of pixel intensity (acceleration of intensity).
Locates an edge at the exact point where the signal crosses zero.
f = signal
(∂2 / ∂x2) h = 2nd derivative of Gaussian kernel
((∂2 / ∂x2) h) * f = convolution
Why to use 2D Derivative of Gaussian filter?
preferable because smoothing and calculating derivative at same time
Effects of sigma on derivatives
apparent structures differ depending on scale parameter
lager values: larger scale edges detected
smaller values: finer features detected
Criteria for optimal edge detector: (3)
good detection: the optimal detector should minimize the probability of false positives (detecting spurious edges caused by noise), as well as that of false negatives (missing real edges)
Good localization: the edges detected should be as close as possible to the true edges
Single response: the detector should return one point only for each true edge point; that is, minimize the number of local maxima around the true edge
Primary edge detection steps
smoothing - suppress noise
edge enhancement - filter for contrast
edge localization - determine which local maxima from filter output are actually edges vs. noise
Canny edge detector Steps
filter image with derivative of gaussian
find magnitude and orientation of gradient
non-maximum suppression
hysteresis
Hysteresis Definition
linking and thresholding
define two thresholds (low and high)
use the high threshold to start edge curves and low threshold to continue them
Non-maximum Suppression Definition
check if pixel is local maximum along gradient
can require checked interpolated pixels p and r
thin multi pixel wide ridges down to single pixel width
about Canny edge detector (3)
(all edge detectors) cannot a shadow from an object
some edge detectors can classify cause of edge
canny used as structural condition for image generation
Single 2D Edge Detection laplacian filter equation
hsigma = gaussain equation
∇2 f = Laplacian operator = (∂² f / ∂x²) + (∂² f / ∂y²)
∇2 hsigma(u, v) = Laplacian of Gaussian
Difficulty of line fitting (3) and how to improve
extra edge points
some parts of lines missing
noise in measured edge points
Voting approaches, such as the Hough transform, make it possible to find likely model parameters without searching all combinations of feature
Hough Space Points
a line in image (set of points (x,y)) corresponds to point in Hough Space (m, b)
image to hough
(x, y) to (m, b) such than y = mx + b
Hough Spaces Lines
Line in hough space, point in image space
point intersection of hough lines = line that passes though both points in image spce
set of points (m, b) to point (x, y) such than b = -xm + y
Hough Algorithm (basic)
let each edge point in image space vote for a set of possible parameters in Hough space
accumulate votes in discrete set of bins
parameters with most votes indicate line in image space
time complexity is linear
Issues with (m, b) parameter space: (2)
(m, b) can take on infinite values and undefined for vertical lines
Polar Representation for Lines
Point in image space = sinusoid segment in Hough space
d = perpendicular distance from line to origin
θ = angle the perpendicular makes with the x axis
x cos θ - y sin θ = d
Impacts of Noise on Hough Space (2)
false positives/negatives
small sparse low count segments
Extensions for Noise on Hough Transform (3)
use image gradient or theta (reduces degrees of freedom
give more votes for stronger edges
change sampling of (d, θ) to give more/less resolution
Hough Transform circles
every point on circle in image to circle in hough space
intersection of circles = radius in image space
(a, b) = center
r = radius
(x - a)2 + (y - b)2 = r2
Hough Transform unknown radius circles
map becomes 3d space with (a, b, r)
for unknown radius and gradient direction, cones in hough space
for unknow radius and known gradient direction, line
Practical tips Hough Transform Voting (4)
Minimize irrelevant tokens first (take edge points with significant gradient magnitude)
Choose a good grid / discretization
Too coarse: large votes obtained when too many different lines correspond to a single bucket
Too fine: miss lines because some points that are not exactly collinear cast votes for different buckets
Vote for neighbors, also (smoothing in accumulator array)
Utilize direction of edge to reduce free parameters by 1
Hough Transform Pros (3)
All points are processed independently, so can cope with occlusion
Some robustness to noise: noise points unlikely to contribute consistently to any single bin
Can detect multiple instances of a model in a single pass
Hough Transform Cons (3)
Complexity of search time increases exponentially with the number of model parameters
Non-target shapes can produce spurious peaks in parameter space
Quantization: hard to pick a good grid size