Chapter 4 Perception Notes

Major Functions of the Visual Perceptual System

Perception is the process through which we interpret the world around us using our senses. Far from being a passive reception of sensory data, perception actively organizes and translates sensory information into meaningful experiences. The visual perceptual system, in particular, plays a crucial role in how we navigate and understand our environment. Its primary functions include localization, recognition, and maintaining constancies.

  1. Localization
    Localization refers to the ability to determine the position of objects in space. This function is essential for interacting with our surroundings, as illustrated when you reach for a cup of coffee. Your perceptual system precisely locates the cup in relation to your hand. This process involves several sub-steps:

  • Separating Objects from the Background: The initial step is to distinguish objects from their surroundings. For instance, identifying the cup as a distinct entity separate from the table.

  • Organizing Objects into Meaningful Groups: This involves grouping related items together to form a coherent scene. For example, perceiving the cup as part of a set of items arranged on the table, rather than as an isolated object.

  1. Recognition
    Recognition is the ability to identify what an object is, allowing us to categorize and understand its purpose. This process relies on recognizing specific features and matching them to stored representations in our memory. For instance, you recognize a chair by its characteristic features, such as four legs, a flat seat, and a back. Recognition allows us to understand that the object is a chair and to infer its purpose, such as a place to sit.

  2. Constancies
    Perceptual constancies are mechanisms that enable us to perceive objects as stable and consistent, despite changes in viewing conditions. These constancies help us maintain a stable and reliable view of the world.

  • Size Constancy: Size constancy refers to the perception of an object as having a constant size, regardless of its distance from the observer. For example, a car appears to be the same size whether it is far away or up close. This is because our perceptual system takes distance into account when estimating size. Without this our perception of size would be constantly in flux.

  • Lightness Constancy: Lightness constancy is the ability to perceive an object as having a constant level of lightness, regardless of the amount of illumination. For instance, a white shirt looks white whether you are indoors or outside, even though the amount of light reflecting off the shirt differs in each environment. Because our perceptual system adapts to changes in illumination an object's perceived lightness remains stable.

  • Color Constancy: Color constancy is the perception of an object as having a constant color, despite changes in the spectrum of light illuminating it. A red apple still looks red whether it is in sunlight or under a lamp, even though the wavelengths of light reaching our eyes differ in each situation. The function ensures that we perceive objects as having consistent colors.

  • Shape Constancy: Shape constancy refers to the perception of an object as having a constant shape, regardless of the viewing angle. A door is still perceived as rectangular, even when it is opened and the retinal image changes. This allows us to recognize objects from different perspectives.

  • Location Constancy: Location constancy is the understanding that fixed objects maintain constant locations, even as the retinal image changes with movement. For example, knowing your house hasn't moved even if you drive to another city allows us to navigate effectively and maintain a stable sense of place.

Historically, perception was thought to rely on the five classical senses: sight, hearing, touch, taste, and smell. Modern theory broadens this perspective by emphasizing five perceptual systems: visual, auditory, haptic (touch), savor (taste and smell), and gravity/orientation. The gravity/orientation system, which involves the vestibular system and kinesthetic cues, helps us maintain balance and orientation in space. The vestibular system is located in the inner ear and coordinates with muscles and eyes to discern body position relative to gravity (Bartley, 1980). Kinesthetic cues involve sensory information from muscles, tendons, and joints that provide awareness of body position and movement.

Gestalt Principles and Figure–Ground Organization

The brain naturally organizes visual information into meaningful patterns or "gestalts." These principles, developed by Gestalt psychologists, explain how we perceive and organize visual elements into unified wholes.

  • Law of Pragnanz: Also known as the law of good Gestalt, it states that among various possible organizations, we perceive the simplest, most stable form. For instance, seeing a square instead of four lines, even if there is a space between each line. Max Wertheimer was one of the main proponents of this law.

Figure–ground organization is the process of distinguishing between the figure (the object of focus) and the ground (the background). This fundamental aspect of perception allows us to focus on relevant information while relegating the rest to the background.

  • Rubin’s Criteria: Edgar Rubin described several criteria that characterize the separation between figure and ground:

    • The figure has a definite shape, while the ground is formless.

    • The ground appears to continue behind the figure, creating a sense of depth.

    • The figure seems closer and has a specific location in space, while the ground recedes.

    • The figure is more dominant and memorable than the ground.

Ambiguous figure-ground relationships can reverse, as seen in the famous Rubin vase, where you can alternately see two faces or a vase. This illustrates the dynamic nature of perceptual organization.

Gestalt laws of grouping describe how elements are perceptually bound together, influencing how we perceive relationships between objects:

  • Proximity: Elements that are near each other are seen as a unit. For example, seeing three groups of two dots instead of six individual dots (xx xx xx).

  • Similarity: Similar elements are grouped together. For example, seeing a group of circles among a group of squares (OOOOO [] [] [] [] []).

  • Good Continuation: Elements arranged on a line or curve are seen as a unit. For example, tracing a winding road on a map where the road is perceived as a single continuous entity.

  • Closure: Incomplete figures are perceived as complete. For example, recognizing a circle even with a small gap, where our minds fill in the missing information to perceive a complete circular shape.

  • Common Fate: Elements moving in the same direction are seen as a unit. An example of this is a flock of birds flying together where all the birds are perceived as a single unit.

Depth Perception: Monocular and Binocular Cues

Depth perception allows us to perceive the three-dimensional world and judge distances. This ability relies on two types of cues: monocular and binocular.

Monocular Cues (Require One Eye)

Monocular cues, also known as pictorial cues, can be perceived with only one eye. Artists commonly use these cues to create a sense of depth in two-dimensional images.

  • Relative Size: Smaller instances of familiar objects are seen as farther away. For example, seeing a small car and assuming it's farther away.

  • Superimposition: When one object covers another, the occluding object appears nearer. For example, seeing a stack of books and knowing the books at the bottom are in the back.

  • Relative Height: Objects higher in the visual field are seen as farther away. For example, seeing mountains high in the sky and knowing they're farther away.

  • Linear Perspective: Parallel lines converge in the distance. For example, seeing train tracks meet far away, creating a sense of depth.

  • Motion Parallax: During movement, closer objects move faster across the retina than farther objects. For example, when driving, nearby trees appear to move by faster than distant mountains.

  • Texture Gradient (Gibson): As a textured surface recedes, texture elements become smaller and more dense, providing direct depth information. Gibson’s direct perception theory posits that depth can be perceived directly from the ground via texture gradients without needing elaborate inference. This is associated with James J. Gibson.

Binocular Cues (Require Both Eyes)

Binocular cues rely on the slightly different views from each eye to provide depth information.

  • Parallax: Relative motion of objects as the observer moves. Closer objects shift more than distant ones. For example, holding up your finger and looking at it with one eye at a time, noticing how much it shifts compared to background objects.

  • Disparity: The retinal images differ between the two eyes, and the brain uses this difference to compute depth. This is the principle behind 3D movies, where each eye sees a slightly different image.

  • Convergence: The inward turning of the eyes when focusing on near objects. The degree of inward turn signals distance. You can slightly feel this when very focused on something close to your face.

Object Recognition: Theories and Mechanisms

Object recognition is the process of assigning a percept to a category, allowing us to identify and understand objects in our environment. Several theories explain how we achieve this.

  • Template Theory: We recognize objects by comparing them to stored templates of previously learned patterns. This theory struggles with the variability of object appearances (size, orientation) and unseen objects. Because real-world objects vary greatly template theory has limitations.

  • Prototype Theory: Instead of exact templates, we store abstract representations or prototypes of a category. New stimuli are matched to these prototypes. While more flexible than template theory, it lacks details on the matching process and context effects.

  • Feature Theory: Objects are broken down into basic features, and recognition occurs by detecting combinations of these features. This approach accounts for variability but underemphasizes context. Hubel and Wiesel’s work on feature detectors in the visual cortex supports this theory.

  • Marr’s Computational Theory:

Marr’s computational theory of vision (Marr, 1982) proposes a sequence of representations:

  • Primal Sketch: Initial description of the retinal image with basic features like lines and edges.

  • 2.5D Sketch: Description of surfaces and depth from the observer’s viewpoint.

  • Object Descriptions: Higher-level representations of object identities.

In addition, recognition involves differing processing methods:

  • Bottom-Up Processing: Driven by sensory input and feature extraction (e.g., edges, lines). This processing starts with the raw sensory data and builds up to a complete perception.

  • Top-Down Processing: Guided by stored knowledge, expectations, and context, helping to resolve ambiguities in the visual input. For example, you can quickly recognize a friend's face in a crowd because you already know what they look like. This process relies on higher-level cognitive processes to interpret sensory information.

Perceptual Constancies

Perceptual constancies help stabilize our perception despite varying sensory inputs. These mechanisms allow us to perceive objects as having stable properties, even when the sensory information changes.

  • Size Constancy: The perceived size remains constant despite changes in retinal size due to distance. For example, a building still appears to be the same size whether observing close to it or far away.

  • Lightness Constancy: An object appears equally light despite varying illumination. For instance, a white piece of paper looks white whether in a brightly lit room or a dimly lit room.

  • Color Constancy: Object color remains roughly the same despite changes in illumination. A red car still looks red in sunlight or shade.

  • Shape Constancy: Object shape is perceived as constant even as the retinal image changes. A door opening is still perceived as a rectangle even though the image on our retina is a trapezoid.

  • Location Constancy: Fixed objects maintain constant locations even as the retinal image changes with movement. For example, knowing your house hasn't moved even if you drive to another city.