Spatial Data Analysis

Geographic latitude: angel between equatorial plane and the normal to the ellipsoid at the point of interest and indication of north-south coordinate

  • Data Cleaning

    • Structure

    • Outliers (functional: every time you know your data is incorrect or statistical)

    • Labels and units

    • Very rarely using basic GIS program to do this usually need an extra software add-on

  • Spatial rectification

    • Co-locations

      • Data record where two points have the same coordinates

      • Delete co-located data records

      • Average co-located data records

    • Spatial shifts

      • Cant put antenna right above the sensor

  • Data standardization

    • Detrending

      • The values you are trying to measure are changing while you are measuring

      • Removes trends

    • Removing bias

      • Removes bias

  • Data Rasterization

    • Interpolation

    • With ultra high density data you do more averaging than interpolation

    • Defining the raster is a very large challenge, what are the potential uses, and what is the data usage points

  • Temporal dynamics (seeing how natures resistance changes over time)

    • Raster algebra

    • Temporal statistics (used in global warming models, looking at temp of ocean over 50 years time)

  • Classification

    • Contouring

    • Classification

    • Clustering

    • Artificial intelligence

  • Yield data

    • Lots of effort in data filtering

    • Remove about 25% from every file

  • Data filtering

    • What to remove (yield data example)

      • Header up points as as well as start and end pass delays

      • Points with yield values or individual sensor measurements exceeding the possible range

      • Outliers based on descriptive statistics

      • Points with detectible misplacement (e.g., co-aligned points)

      • Points that do not agree with a predefined statistical estimate based on local neighborhood statistics

      • Only leave what is seen as appropriate

  • Yield Map Analysis

    • What do you do with multiple years of data (different plants, different management)?

      • Standardize, bring data to special characteristics

        • If your yield is above or below average then try identify property of the side

      • Combine data temporally

      • Can do comparison group into management zones

  • Statistical parameters

    • Mean (avg)

    • Variance

    • Standard deviation

    • Coefficient of variation

  • Data (yield) Normalization

    • You get relative values

  • Temporal statistics

  • Simple yield classification

  • Finding the source of yield variability

    • Combine maps that show one feature with those with other features

  • Data Fusion

    • Soil Eca

    • Soil pH

    • Soil Reflectance

    • Topography

    • Lab data

  • Data Clustering

    • Supervised and unsupervised classification

    • Sampling an/or management zone delineation

      • To identify homogenous areas

      • Some times sampling and management zones are different

        • Management zones are for doing specific things with the land

    • Hierarchical clustering

      • One way to establish multiple clusters around a specific area

    • Overall goal

      • To develop a robust algorithm that could handle multilayer data and produce an unspecified number of spatially contiguous field partitions while relying solely on the information embedded within the specified dataset

      • To locate one representative location for each partition for point-based analysis

    • Can do it with interpolated data and another type of data

    • Many methods used to identify this