1/57
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
What are the three main components of a local feature pipeline?
Detection (identify the interest points)
description (extract a feature vector around each interest point)
matching (determine correspondence between descriptors in two views).
What is the purpose of the detection step?
To identify interest points that can be found again in other images, running the detection procedure independently on each image.
What is the purpose of the description step?
To extract a feature vector around each interest point that is distinctive and has some invariance to geometric and photometric differences between views.
What is the purpose of the matching step?
To determine the correspondence between descriptors in two views.
What is correspondence?
Matching points, patches, edges, or regions across images.
What are the four characteristics of good local features?
Repeatability,
saliency (each feature is distinctive),
compactness and efficiency (many fewer features than image pixels)
locality (a small area of the image, so robust to clutter and occlusion).
What does repeatability mean?
The same feature can be found in several images despite geometric and photometric transformations.
What is the goal of an interest operator (detector) versus a descriptor?
Detector: find at least some of the same points in both images, even though it runs independently on each image.
Descriptor: reliably determine which point goes with which, with invariance to geometric and photometric differences.
Why are corners good interest points?
You can easily recognize them by looking through a small window: shifting the window in any direction gives a large change in intensity.
How does intensity change when you shift a window over a flat region, an edge, and a corner?
Flat: no change in all directions.
Edge: no change along the edge direction.
Corner: significant change in all directions.
What is the formula for the window shift error E(u,v)?
E(u,v) = Σ over (x,y) of w(x,y) [ I(x+u, y+v) − I(x,y) ]², the change in appearance of window w(x,y) for the shift [u,v].
What are E(0,0) and E for a flat patch?
E(0,0) = 0 (no shift means no change),
flat patch gives E = 0 for every shift.
What two forms can the window function w(x,y) take?
1 inside the window and 0 outside, or a Gaussian.
Why can't we just compute E(u,v) directly for every shift?
It is too slow: about O(window_width² × shift_range² × image_width²), which is roughly 5.2 billion operations in the slide example.
What approximation is made to simplify E(u,v)?
For small shifts, intensity changes almost linearly: I(x+u, y+v) ≈ I(x,y) + Ix·u + Iy·v. Substituting and expanding the square gives a quadratic form.
What is the second moment matrix M?
M = Σ w(x,y) [ Ix², IxIy ; IxIy, Iy² ], a 2×2 matrix of image derivatives averaged over the neighborhood of a point.
What are the steps to compute M for a window?
For each pixel compute Ix, Iy, then Ix², Iy², IxIy. Sum each (weighted by w) over the window. Put Σ w Ix² and Σ w Iy² on the diagonal and Σ w IxIy in both off-diagonal entries.
What do the eigenvalues of M say when both are small?
λ1 and λ2 both small means E is almost constant in all directions: a flat region.
What do the eigenvalues of M say when one is much larger than the other?
λ1 >> λ2 (or λ2 >> λ1) means an edge: E increases in only one direction.
What do the eigenvalues of M say when both are large?
λ1 and λ2 both large with λ1 ~ λ2 means a corner: E increases in all directions.
How are det(M) and trace(M) related to the eigenvalues?
det(M) = λ1·λ2 and trace(M) = λ1 + λ2.
What is the Harris corner response function?
R = det(M) − α·trace(M)² = λ1λ2 − α(λ1 + λ2)²
What is the typical range for α in the Harris response?
0.04 to 0.06.
What is the sign of R for a corner, an edge, and a flat region?
Corner: R > 0 (and large).
Edge: R < 0.
Flat: |R| is small (R near 0).
Why does the Harris response use det and trace instead of eigenvalues?
The determinant and trace of a 2×2 matrix take only a few multiplications and an addition, and together they encode both eigenvalues, so no eigenvalue is ever computed.
What are the three high-level steps of the Harris corner detector?
1) Compute M for each image window to get a cornerness score.
2) Find points with large response (R > threshold).
3) Take the local maxima of R (non-maximum suppression).
What are the detailed steps of the Harris detector implementation?
1) Compute image derivatives Ix, Iy (optionally blur first).
2) Compute Ix², Iy², IxIy.
3) Gaussian-filter each: g(Ix²), g(Iy²), g(IxIy).
4) Compute cornerness R = g(Ix²)g(Iy²) − [g(IxIy)]² − α[g(Ix²) + g(Iy²)]².
5) Non-maximum suppression.
What does the ellipse [u v] M [u v]ᵀ = const tell you about M?
M = R⁻¹ diag(λ1, λ2) R
The half-axis lengths are λ^(−1/2).
The larger eigenvalue gives the shorter axis, which is the direction of fastest change.
The smaller eigenvalue gives the longer axis, which is the direction of slowest change.
Is Harris invariant to image rotation? scale changes? Why?
Yes. Under rotation the ellipse rotates but its shape (the eigenvalues) stays the same.
No, so we need a way to detect interest points at the right scale.
How can scale-invariant detection choose corresponding region sizes if each image is processed independently?
Use a rule both images can follow on their own: find the scale that gives a local maximum of some function f in both position and scale (the characteristic scale).
What is the characteristic scale?
The scale that produces the peak filter response; we search over scales to find it.
How is LoG used for edge detection versus blob detection?
Edges: look for zero-crossings of the LoG response.
Blobs: look for extrema (maxima or minima) of the response.
What is the Difference of Gaussians (DoG)?
DoG = G(x,y,kσ) − G(x,y,σ), used as a kernel in scale-invariant detection.
What do the LoG and DoG kernels have in common?
Both kernels are invariant to scale and rotation.
How many neighbors does a DoG extremum compare against, and what must be true?
26: 8 in the current image plus 9 in the scale above and 9 in the scale below. The point is selected only if it is larger than all 26 neighbors.
What implementation recipe is given for scale-invariant detection?
For each level of the Gaussian pyramid compute the feature response (for example Harris or Laplacian). Then for each level, if the point is a local maximum and also a maximum across scale, save the scale and location (x, y, s).
Which keypoints are removed after finding DoG extrema?
Those with low contrast or poorly localized along an edge
How does SIFT decide whether a keypoint is poorly localized along an edge?
It computes the ratio of the eigenvalues of C (as with Harris corners, principal curvatures) and checks whether it is greater than a threshold.
Why was SIFT developed?
The Harris operator is not invariant to scale and correlation is not invariant to rotation. Lowe wanted a detector invariant to scale and rotation, and a descriptor robust to typical viewing variations.
What are the four steps of SIFT feature generation?
1) Scale-space extrema detection (DoG pyramid)
2) Keypoint localization.
3) Orientation assignment.
4) Keypoint descriptor.
What parameters does the SIFT pyramid use?
Difference of Gaussian pyramid with 3 scales per octave, downsampling by a factor of 2 for each octave.
What is a problem with using raw pixel intensities as the descriptor?
It is sensitive to photometric transformations, such as changes in absolute intensity values.
What does using image gradients (pixel differences) fix, and what problem remains?
It makes the feature invariant to absolute intensity values, but it is still sensitive to geometric transformations such as scale, translation, rotation, and deformation.
Why does SIFT assign an orientation to each keypoint?
To find a local orientation (the dominant gradient direction) and compute derivatives relative to it, so the descriptor is invariant to rotation.
How is the SIFT canonical orientation assigned?
Create a histogram of local gradient directions at the selected scale and assign it at the peak of the smoothed histogram.
What are the details of the SIFT orientation histogram?
36 bins (10 degree increments), with Gaussian-weighted voting using a window based on a Gaussian of 1.5 times the scale.
What happens when other orientation histogram peaks are near the maximum?
The highest peak and any peaks above 80% of the highest are also considered for calculating dominant orientations (additional orientations for that keypoint).
What four values define a SIFT keypoint after orientation assignment?
Position (x, y), scale, and orientation, which gives stable 2D coordinates invariant to those factors.
What is the difference between keypoint detection and keypoint description?
Detection assigns location, scale, and orientation.
Description computes a vector for each keypoint.
How is the SIFT descriptor computed on the window around a keypoint?
First rotate the window to line up with the keypoint's orientation and resize it to its scale, so the patch looks the same regardless of the image's rotation or scale.
Then take gradients inside it, weighting each by a Gaussian centered on the keypoint (variance set to half the window size), so gradients near the center count more and the influence fades smoothly toward the edge.
How is the SIFT descriptor structured, and why this size?
A 4×4 array of gradient orientation histograms weighted by magnitude, with 8 orientations each: 8 × (4×4) = 128 dimensions. This gives some sensitivity to spatial layout, but not too much.
How does SIFT keep the descriptor smooth across bin boundaries?
A Gaussian weight plus interpolation: a given gradient contributes to 8 bins (4 in space times 2 in orientation).
How does SIFT reduce the effect of illumination in the descriptor?
Normalize the 128-dimensional vector to unit length, clamp values greater than 0.2 to avoid excessive influence of high gradients, then renormalize.
How is scale space generated for feature detection?
A Gaussian pyramid processes one octave at a time, resampling and blurring to produce a set of scales, then subtracting to get DoG images
What is the general idea of key point localization in scale space?
Find a robust extremum (maximum or minimum) both in space and in scale, using a pyramid of resample, blur, and subtract.
How are scale-invariant interest points chosen from the filter responses?
As local maxima in both position and scale of the squared filter response maps, giving a list of (x, y, σ).
What is the key difference in scale handling between Harris and SIFT detection?
Harris works at a single fixed window scale and is not scale invariant.
SIFT uses a DoG pyramid to find extrema across both position and scale , so it returns (x, y, scale) for each keypoint.
Why do we consider regions of different sizes around a point for scale invariance?
Regions of corresponding sizes will look the same in both images, but each image must choose its sizes independently, so we pick the scale that gives a local maximum of a function f.