Spatial Data Analysis
Geographic latitude: angel between equatorial plane and the normal to the ellipsoid at the point of interest and indication of north-south coordinate
Data Cleaning
Structure
Outliers (functional: every time you know your data is incorrect or statistical)
Labels and units
Very rarely using basic GIS program to do this usually need an extra software add-on
Spatial rectification
Co-locations
Data record where two points have the same coordinates
Delete co-located data records
Average co-located data records
Spatial shifts
Cant put antenna right above the sensor
Data standardization
Detrending
The values you are trying to measure are changing while you are measuring
Removes trends
Removing bias
Removes bias
Data Rasterization
Interpolation
With ultra high density data you do more averaging than interpolation
Defining the raster is a very large challenge, what are the potential uses, and what is the data usage points
Temporal dynamics (seeing how natures resistance changes over time)
Raster algebra
Temporal statistics (used in global warming models, looking at temp of ocean over 50 years time)
Classification
Contouring
Classification
Clustering
Artificial intelligence
Yield data
Lots of effort in data filtering
Remove about 25% from every file
Data filtering
What to remove (yield data example)
Header up points as as well as start and end pass delays
Points with yield values or individual sensor measurements exceeding the possible range
Outliers based on descriptive statistics
Points with detectible misplacement (e.g., co-aligned points)
Points that do not agree with a predefined statistical estimate based on local neighborhood statistics
Only leave what is seen as appropriate
Yield Map Analysis
What do you do with multiple years of data (different plants, different management)?
Standardize, bring data to special characteristics
If your yield is above or below average then try identify property of the side
Combine data temporally
Can do comparison group into management zones
Statistical parameters
Mean (avg)
Variance
Standard deviation
Coefficient of variation
Data (yield) Normalization
You get relative values
Temporal statistics
Simple yield classification
Finding the source of yield variability
Combine maps that show one feature with those with other features
Data Fusion
Soil Eca
Soil pH
Soil Reflectance
Topography
Lab data
Data Clustering
Supervised and unsupervised classification
Sampling an/or management zone delineation
To identify homogenous areas
Some times sampling and management zones are different
Management zones are for doing specific things with the land
Hierarchical clustering
One way to establish multiple clusters around a specific area
Overall goal
To develop a robust algorithm that could handle multilayer data and produce an unspecified number of spatially contiguous field partitions while relying solely on the information embedded within the specified dataset
To locate one representative location for each partition for point-based analysis
Can do it with interpolated data and another type of data
Many methods used to identify this
