Geographic Information Systems: Raster and Vector Data Models
Fundamentals of Spatial Representation and GIS Modeling
Purpose of GIS Models:
The primary purpose of Geographic Information Systems (GIS) is to accurately represent the physical real world within a digital computer environment.
Computers possess no inherent knowledge of real-world environments. Every feature, boundary, and spatial relationship must be explicitly defined and provided by the user.
Incomplete data models lead to system misrepresentations. For example, failing to explicitly code the existence of a building results in a model that assumes empty space.
Essential Components of Spatial Models:
Geographic Entities: Representation of real-world objects using basic vector geometry, including points, lines/polylines, and polygons (e.g., bridle roads, main roads, administrative buildings such as A2 polygons).
Nodes: Specific coordinate junction points that establish connectivity and topology between lines and polygons.
Spatial Reference & Coordinates: Spatial data requires explicit position parameters defined along an -axis and -axis to establish location within geographic space.
Scale: The proportional relationship between distance on the model/map and the corresponding distance on the ground.
Attribute Tables: tabular databases linked directly to spatial geometries, providing essential non-spatial details, characteristics, and classifications for every feature.
Manual Vector Model Construction:
Building a manual vector model of a real-world setting (e.g., taxisanian campus) requires extensive labor and hours of manual input.
Every vertex and point must be created manually and connected strategically so the database can distinguish that Point A is structurally distinct from Point B.
A GIS model does not inherently understand context; it only processes instructions regarding points, lines, polylines, polygons, and directional vectors.
Vector vs. Raster Data Structures
Vector Data Structure:
Geometry: Formed using discrete points, lines/polylines, and polygons defined by explicit coordinate pairs (, ).
Structure: Points connect to form polylines, and closed sequences of polylines create polygons.
Complexity: High structural complexity. Vector data maintains explicit topological relationships, making digitizing and editing labor-intensive. A single misplaced coordinate or node can disrupt the spatial model.
Attribute Organization: Complex and non-uniform attribute tables designed to store detailed, feature-specific properties.
Raster Data Structure:
Geometry: Composed of a continuous grid array or matrix of rows and columns.
Grid Elements: Individual grid units are referred to interchangeably as pixels, cells, bits, boxes, or grids.
Spatial Coverage: Every individual cell represents a uniform, discrete unit of geographic space on Earth.
Dimensions: Standard raster models consist of a two-dimensional () array of rectangular cells, but can extend into three-dimensional () structures via multi-layered grid stacks.
Attribute Representation: Each pixel contains a single, uniform numerical value that represents the dominant attribute or condition of the space it covers.
Coordinate Uniformity: Grid origins and cell dimensions maintain uniform distance steps along the and axes.
Spatial Resolution, Cell Scaling, and Grid Geometry
Spatial Resolution Definitions:
Fine (High) Spatial Resolution: Characterized by smaller individual pixel dimensions. Each cell covers a small physical footprint on the ground, enabling high detail, distinct feature differentiation, and sharp visual clarity.
Coarse (Low) Spatial Resolution: Characterized by larger individual pixel dimensions. Each cell covers a vast physical footprint on the ground, forcing feature generalization, loss of detail, and visual pixelation.
Impact of Cell Size Dimensions:
Resolution: Captures high ground detail, allowing individual micro-features (such as lithotripsy units or narrow structures) to be distinctly resolved.
Resolution: Each cell represents an area measuring exactly by on the physical ground surface.
Resolution: Each cell represents an area measuring by on the physical ground surface.
Resolution: Each cell represents an area measuring by on the physical ground surface.
Resolution: Features inside the cell footprint are generalized and blended into a single pixel average value.
Mathematical Scaling of Raster Cell Counts:
Decreasing resolution does not scale dataset size linearly; changes in spatial resolution affect cell count exponentially by a squared spatial factor.
When doubling cell side length (e.g., transitioning from a cell size to a cell size), four separate fine-resolution cells ( grid) are merged into a single coarse cell.
The resulting data volume decreases by a factor of .
Grid Aggregation Comparison: An area mapped by individual fine cells requires only cells ( layout) when aggregated to coarse resolution.
Computational and Cost Trade-offs:
High Spatial Resolution Advantages: Provides fine spatial detail and clear feature boundaries.
High Spatial Resolution Disadvantages: Requires massive computational storage, high acquisition cost, and intense processing power that can overwhelm computing systems.
Low Spatial Resolution Advantages: Dramatically reduces file sizes and computational overhead, lowering processing costs.
Low Spatial Resolution Disadvantages: Causes significant spatial generalization and loss of critical spatial boundary data.
Optimal Resolution Selection: Analysts must strike a balance based on application requirements. Low spatial resolution is suitable for macro-scale analyses (e.g., tracking global phenomena, atmospheric movement, or large-scale temporal shifts), while fine resolution is required for local engineering or urban planning.
Raster Attribute Tables and Data Measurement Scales
Attribute Table Structure in Raster Models:
Raster attribute tables are structured simply and uniformly compared to vector tables. Every cell in a given raster layer follows the exact same schema:
Column 1 (ID/Location): Sequential cell identification numbers (e.g., Cell 1, Cell 2, Cell 3, Cell 4, Cell 5, Cell 6…).
Column 2 (Value): Numeric pixel code or value assigned to the cell (e.g., ).
Column 3 (Area/Resolution): Fixed spatial area corresponding to the cell size (e.g., ).
Strict Uniformity Rule: The spatial resolution of a raster dataset is strictly uniform across the entire layer. A single raster layer cannot contain varying cell sizes (e.g., a cell cannot exist within a standard layer).
Computer Processing of Value Mapping:
Computers process raw numeric matrix values ( or ).
GIS analysts define value lookup tables to translate numeric codes into thematic information (e.g., assigning Value to represent a specific subtype of tree, such as conifers).
The computer executes quantitative spatial queries by counting matching pixel values across the array. For example, querying a raster for pixel value scans the matrix and returns the exact count (e.g., total pixels found).
Data Measurement Scales in Raster Datasets:
Nominal Data: Qualitative categories with no numerical ranking or order. Each cell contains an arbitrary code representing discrete classes (e.g., Value = Water, Value = Agriculture, Value = Residential, Value = Industrial; or unordered forest species types).
Ordinal Data: Ranked or ordered categories expressing relative qualitative levels without fixed numerical intervals (e.g., quality of life classified per cell as "Low Quality", "Moderate Quality", or "High Quality").
Interval Data: Quantitative numeric scales with arbitrary zero points, where differences between values are equal and meaningful (e.g., temperature recorded in Celsius or Fahrenheit, or voltage potential differences).
Ratio Data: Quantitative numeric scales featuring a true, non-arbitrary absolute zero point, allowing ratio comparisons and mathematical aggregation (e.g., measured rainfall amounts in millimeters or inches per pixel).
Visualizing Continuous Data and Graduated Colors:
Continuous surface phenomena (e.g., elevation, slope, pollution concentrations, population density) rely on numeric values mapped to color ramps.
Elevation Mapping Example: Red values assigned to low elevations; dark green values assigned to high elevations.
Graduated Color Schemes: Color hue intensity varies dynamically according to attribute magnitude. For instance, in a pollution dataset, cells with pollution indices between and display as light gray, whereas extreme pollution zones ranging from to render in dark gray or intense hues.
Spatial Generalization and Elevation Surface Modeling
Feature Overlap and Generalization:
When multiple distinct physical features (such as multiple light poles) fall within the geographic extent of a single coarse raster cell, the cell generalizes the data into a single averaged value.
To resolve individual features, spatial resolution must be increased (reducing individual cell size).
System-Wide Consequence: Decreasing cell size to resolve features in one location forces a grid-wide cell reduction, exponentially multiplying the total cell count across the entire raster model.
Center-Point Value Storage Rule:
Raster attribute values are mathematically anchored and referenced at the exact geometric center point of each grid cell.
Topographic Impact on Elevation Models:
In low-resolution elevation grids, localized topographic variations (e.g., sharp peak spikes or dramatic ravine dips located at cell corners) are lost, as the entire cell records only the single generalized value calculated at its central point (e.g., values , , or ).
Flat Terrain Performance: In flat, uniform landscapes (e.g., North Texas plains), coarse spatial resolution causes minimal error because micro-topographic variations between cell centers are negligible (e.g., minor variations like are trivial).
Variable/Mountainous Terrain Performance: In highly variable, mountainous terrain, coarse spatial resolution results in severe modeling errors, missing intermediate peaks, valleys, and critical slope transitions due to center-point spatial averaging.
The Mixed Pixel Problem and Decision Rules
Definition of the Mixed Pixel Problem:
Occurs when a single raster cell spans two or more distinct land cover types or spatial features in the real world (e.g., a boundary pixel along a coastline containing both grass and water).
Predominantly impacts coarse spatial resolution datasets where pixel sizes are larger than the geographic features being mapped.
Forces a classification ambiguity where the system requires specific rules to resolve which single value to assign to the pixel.
Mixed Pixel Classification Rules:
Winner Takes All (Majority Rule):
The feature covering the largest proportional area within the boundaries of the cell determines the cell's final classification.
Example: In a mixed cell containing grass cover and open water, the cell is classified entirely as Grass.
Dominance Rule (Feature Dominance):
Assigns the cell to a specific class if a designated priority feature is present anywhere within the cell boundary, regardless of whether it covers a majority of the area.
Water Dominant Example: If open water is present even along the outer lip of a cell containing mostly grass, the cell is classified entirely as Water.
Oil Spill Mapping Application: Essential for environmental hazard tracking. Using an "Oil Dominates" rule ensures that any pixel containing even trace amounts of oil is classified as Oil, guaranteeing full spatial boundary detection for containment efforts.
Edges / Mixed Class Method (The Lazy Method):
Avoids forcing ambiguous boundary pixels into primary feature categories by assigning all multi-feature boundary cells to a distinct third classification named "Edges" or "Mixed".
Example: Pure land cells = Class 1 (Grass); pure water cells = Class 2 (Water); boundary pixels containing both = Class 3 (Edges).
Remote Sensing and Structural Comparisons
Remote Sensing Origins of Raster Data:
The majority of raster data models originate from Remote Sensing—the science of acquiring geospatial information and images about Earth features from a distance without making direct physical contact.
Continuous orbital Earth-observation satellites capture vast datasets of raster imagery constantly.
Structural Trade-Off Summary: Vector vs. Raster:
Raster Model Characteristics:
Simple grid structure based on row/column numeric matrices.
Simple, standardized, uniform attribute tables across all cells.
High storage footprint as resolution increases.
Structurally resilient: a single corrupted pixel or incorrect numeric value affects only its local grid position without damaging the overall geometric integrity of the dataset.
Vector Model Characteristics:
Highly complex geometric structure relying on explicit mathematical coordinates, polylines, polygons, and topological connections.
Highly detailed, variable attribute tables.
Structurally sensitive: a single missing coordinate, misaligned node, or topological error can invalidate geometry across the model.