Object Perception, Gestalt Principles, and Bayesian Inference Notes
Core Challenges in Object Perception
The Inverse Projection Problem: This challenge arises because the same pattern of light falling on the retina can be created by multiple different combinations of objects in the physical world. The retinal image is essentially ambiguous, requiring the brain to work backward to determine the most likely 3D source.
Occlusion and Blurriness: Objects in the world are frequently hidden (occluded) or appear blurry.
Occlusion Example: In a cluttered environment, you might only see the tip of a pencil or a small part of a pair of glasses sitting behind a computer. Despite seeing only the "least glasses-like" part of the object, humans immediately recognize the object as a whole piece of eyewear rather than just a fragment.
Blurriness Example: A blurry image of Barack Obama remains recognizable to those with prior knowledge of his face. However, for an individual lacking that prior knowledge, identifying the subject would be significantly more difficult.
Viewpoint Invariance: The same physical object can produce vastly different patterns of neuronal firing on the retina depending on the angle or distance from which it is viewed.
Example: As you walk around a chair or a desk, the signal pattern sent to the brain changes constantly. Humans do not perceive these as new or different objects; we Maintain object constancy despite the shifting perspective.
The Role of Prior Knowledge and AI Development
Importance of Prior Knowledge: Humans navigate perceptual ambiguity by using a lifetime of accumulated information. We make inferences so quickly that we often don't realize how difficult the task is until we encounter something with no context.
Cultural Specifics: Stimuli common in one culture (e.g., specific objects from Eastern or Western cultures) may be impossible for someone from another culture to identify. Without prior knowledge, one cannot even guess the size or intended use of an object.
The Evolution of Machine Perception:
Past Limitations: Only 8 years ago, computers were significantly worse at classifying objects. For example, a mobile app used for a class at Union College frequently misidentified objects: it thought a ukulele was a microwave or a wooden spoon, and identified a football as a bottle cap.
The CAPTCHA System: Originally developed in the 1990s by Luis von Ahn, CAPTCHAs (Completely Automated Public Turing test to tell Computers and Humans Apart) were designed to ensure users were human. The creator eventually used the human effort to digitize books by having people transcribe blurry words. Later, Google used CAPTCHAs (clicking squares with street lamps, fire hydrants, or bicycles) to train AI for Google Maps.
Modern AI (GenAI): Contemporary Large Language Models (LLMs) and tools like Google Lens have overcome these hurdles by "stealing" and processing massive amounts of data from the internet. They can now classify objects and provide location data as effectively as humans.
Historical Examples of AI Misclassification (Student Submissions):
Two different plants identified as grasshoppers or wine bottles.
Non-bobsled objects identified as bobsleds.
A plastic bag or rhinoceros beetle identified as a tarantula.
A prison identified as ice cream, a stingray, or a guillotine.
Roommates identified as punching bags.
Stuffed animals identified as Arabian camels, fur coats, or Dalmatians.
A Tide Pod identified as a nipple.
Philosophical and Psychological Approaches to Perception
Grouping and Segregation: Perception involves knowing what an object is and what it isn't. This requires grouping related visual signals together and separating them from the background (e.g., seeing a table as distinct from the wall behind it).
Structuralism (Wilhelm Wundt):
Core Principle: Perception arises from the sum of its individual sensations.
Definition: "The whole is equal to the sum of its parts."
Analogy: Pointillism in art, where individual dots of color are added together to create a recognizable image.
Gestalt Psychology:
Etymology: "Gestalt" is a German word meaning configuration.
Key Researchers: Max Wertheimer (primary), along with other German researchers.
Core Principle: The brain perceives things that are more than or different from the sum of sensory parts.
Challenges to Structuralism:
Apparent Movement: The perception of motion when no physical motion occurs. Max Wertheimer used a stroboscope to flash lights back and forth quickly, creating the illusion of a single light moving. Flipbooks are another example of static images creating fluid movement.
Illusory Contours: Perceiving edges or shapes that do not exist in the stimulus.
Example: The Kanizsa Square (Pac-Man shapes arranged to make the brain "see" a square in the center).
Gestalt Principles of Perceptual Organization
Good Continuation: We see lines as following the smoothest, easiest path even if they go out of view.
Example: Coiled rope is seen as one continuous strand rather than many tiny pieces.
Closure: The tendency to see whole objects even when they are incomplete.
Example: The World Wildlife Fund (WWF) panda logo, which lacks a top closing line yet is seen as a complete animal.
Simplicity (Prägnanz/Good Figure): We perceive the most likely or simple explanation.
Example: The Olympic Rings are seen as five overlapping circles rather than nine separate, complex shapes.
Similarity: Grouping items that look alike in color, shape, or orientation.
Example: Dots in a grid may look like columns if they share colors vertically.
Proximity: Things that are close together are grouped.
Example: Spots on a whiteboard or sets of candles are grouped based on physical closeness.
Common Fate: Elements that move together are grouped together as a single object.
Example: A flock of birds or a school of fish. In the lab, a hidden image of a dog or a bird made of random lines is only visible when the lines move together.
Common Region: Grouping items that occupy the same enclosed area, which can often override proximity.
Uniform Connectedness: The tendency to group items that are physically connected by a shared property (like a line connecting two dots).
The Visual Pathway Overview
The Optical Process:
1. Light enters through the Cornea (first focusing element).
2. Light passes through the Pupil.
3. The Lens further focuses light onto the Retina (back of the eye).
Neural Transmission:
4. Photoreceptors in the retina perform transduction.
5. Signals exit via the Optic Nerve.
6. Signals reach the Optic Chiasm (where information from visual fields crosses).
7. Signals travel to the LGN (Lateral Geniculate Nucleus) in the Thalamus and the Superior Colliculus in the midbrain.
Cortical Processing:
8. Information reaches the V1 (Primary Visual Cortex/V1 visual cortex).
9. Dorsal Pathway: The "Where/How" pathway leading to the Parietal Lobe.
Ventral Pathway: The "What" pathway leading to the Temporal Lobe (mnemonic: "If someone has a temper, you need a vent").
Inference and Bayesian Perception
Likelihood Principle (Hermann von Helmholtz): We perceive the object that is most likely to have caused the pattern of stimuli we received.
Example: Seeing two overlapping rectangles instead of one rectangle and a weird "L" shape.
Bayesian Inference (Thomas Bayes): A statistical approach to perception that compares hypotheses against each other.
Formula Components:
Prior Probability: How likely a hypothesis is in general, regardless of the current observation.
Consistency (Likelihood): How well the hypothesis explains the specific observation currently being made.
Posterior Distribution: The final "answer" or perception determined by multiplying the Prior and the Consistency.
The Relative Relationship:
Medical Diagnosis Example (The Cough):
If you hear a cough, your brain weighs:
Cold: High prior (common), high consistency (coughs are symptoms).
Lung Disease: Low prior (rare), high consistency.
Heartburn: Moderate prior (common), low consistency (doesn't usually cause coughs).
The brain selects the "Cold" because it has the highest combined value.
Visual Example (Monkey and Hippo):
Observation: A monkey and a hippo appearing the same size on the retina.
Hypothesis A (Same distance/size): High consistency, but low prior (monkeys aren't hippo-sized).
Hypothesis B (Small cat closer): Low consistency (it looks like a monkey, not a cat).
Hypothesis C (Small monkey is closer): High prior (monkeys are small) and High consistency (close items look larger). This becomes the perceived reality.