Spatial Description, The Exhaustive Data Set

Contour Maps

  • Figure 4.2 displays a computer-generated contour map of 100 V data points.
  • Contour lines are spaced at 10 ppm intervals, ranging from 0 to 140 ppm.
  • The program considers aesthetic aspects, such as using downhill tick marks to indicate depressions.
  • Contouring clarifies features initially observed in the data posting, like the north-south trough and local maximums.
  • Highlights features not obvious from data posting alone.
  • Closeness of contour lines in the southeastern corner indicates a steep gradient.
  • Draws attention to the proximity of the highest data value (145 ppm) to the lowest (0 ppm).
  • Automatic contouring typically requires interpolating data to a regular grid which results in:
    • Interpolated values are generally less variable than original data, smoothing the contoured surface.
    • A smoother surface is aesthetically pleasing but may understate variability, potentially misleading from a quantitative standpoint.
  • Contour maps, like conditional expectation curves, are treated as helpful qualitative displays with questionable quantitative significance.

Symbol Maps

  • For large, regularly gridded datasets, posting all data values may not be feasible, and contour maps might obscure local details.
  • Symbol maps are an alternative: each location is represented by a symbol indicating the data value's class.
  • Symbols are chosen to visually convey the relative ordering of classes through their visual density.
  • Convenient when a line printer is available but not a plotting device.
  • Scale on symbol maps is often distorted due to differences in horizontal and vertical character spacing on line printers.
  • Symbol map example: Digits 0-9 denote ten classes of V values on a 10 x 10 grid.
  • An alternative to symbol maps is grayscale maps, where symbols are replaced with shades of gray, offering an excellent visual summary of the data.

Indicator Maps

  • Indicator maps are symbol maps using only two symbols: a black box and a white box.
  • Data points are assigned to one of two classes based on whether values are above or below a certain threshold.
  • A series of indicator maps can be very informative.
  • Indicator maps show more detail than contour maps and avoid symbol discrimination issues of conventional symbol maps.
  • Figures 4.5a-i: A series of nine indicator maps corresponding to class boundaries from the symbol map in Figure 4.3.
  • Each map shows data locations in white (V value less than the threshold) and black (V value greater than or equal to the threshold).
  • The series of indicator maps reveals the transition from low values aligned north-south to high values grouped in the southeast corner.
  • Different indicator maps highlight different spatial features.
    • Figure 4.5f: Best image of the north-south trough.
    • Figure 4.5g: Best image of the local maximums.
  • Usefulness becomes apparent when exploring large datasets, e.g., in Chapter 5.

Moving Window Statistics

  • In earth science data analysis, focus is often on anomalies (high-grade veins) or impermeable layers.
  • Contour maps help locate areas with anomalous average values, but anomalies in variability are also significant.
  • Heteroscedasticity: Anomalies in the variability, which may have practical implications.
    • In mining, erratic ore grades can cause problems at the mill as metallurgical processes benefit from low variability.
    • In petroleum reservoirs, large fluctuations in permeability can hamper secondary recovery processes.
  • Calculating summary statistics within moving windows investigates anomalies in both average value and variability.
  • The area is divided into local neighborhoods (windows) of equal size, and summary statistics are calculated for each window.
  • Rectangular windows are commonly used for computational efficiency.
  • Window size depends on the average spacing between data locations and the overall dimensions of the area.
  • Balance between large windows for reliable statistics and small windows for local detail.
  • Overlapping windows: A compromise where adjacent neighborhoods share some data.
  • Figure 4.6: Example of an overlapping moving window calculation using a 4 x 4 m2m^2 window, resulting in 16 data points per neighborhood.
  • The window is moved by 2 m each time, overlapping half of the previous window, creating 16 windows in a 10 x 10 m2m^2 area.
  • Overlapping windows may not be necessary for large, regularly gridded datasets, but useful for smaller or irregularly spaced datasets.
  • Irregularly spaced data: Determine the number of data points required within each window for reliable calculation of summary statistics.
  • If there are too few data within a particular window, it is better to ignore that window in subsequent analysis than to incorporate an unreliable statistic.
  • With enough data: commonly used summary statistics such as mean and standard deviation can be calculated, with the former measuring the average value and the latter variability.
  • The median and interquartile range may be used if the local means are heavily influenced by a few erratic high values.
  • Figure 4.7: Shows the means and standard deviations within 4 x 4 m2m^2 windows.
  • Windows overlap by 2 m, giving 16 local neighborhood means and standard deviations.
  • Values are posted with the mean above the "+" sign (window center) and the standard deviation below.
  • For larger areas, contour maps showing means and standard deviations would be more informative.
  • The posting of moving window means and standard deviations reveals local changes in both average value and variability across the area.
  • The windows with a high average value correspond to the highs we can see on the contour map (Figure 4.2).
  • The local changes in variability, however, have not been captured by any of our previous tools.
  • Southeastern corner: Highest local standard deviations due to very low values in the trough being adjacent to some of the highest values in the entire area.
  • Western edge: Very low standard deviation reflecting very uniform V values in that region.
  • The standard deviations vary much more across the area than the means.
  • Uniformity in local means does not necessarily indicate generally well-behaved data values.
  • Even though the mean values are quite similar, ranging from 83.9 to 106.7 ppm, the standard deviations can be quite different, ranging from 9.1 to 41.3 ppm.

Proportional Effect

  • Anomalies in local variability impact the accuracy of estimates.
  • Areas with uniform values provide better prospects for accurate estimates.
  • High variability leads to poorer chances for accurate local estimates.
  • The estimates from any reasonable method will benefit from low variability and suffer from high variability.
  • Four relationships between the local average and local variability (Figures 4.8a-d):
    • Figure 4.8a: Constant average and variability.
    • Figure 4.8b: Trend in average, constant variability.
    • Figure 4.8c: Constant average, trend in variability.
    • Figure 4.8d: Change in both average and variability.
  • Cases with constant variability (Figures 4.8a-b) are most favorable for estimation.
  • It is preferable to be in a situation where the local variability is related to the local average and is, therefore, somewhat predictable (Figure 4Ad).
  • A scatterplot of local means and the local standard deviations is a good way to check for a relationship between the two.
  • Such a relationship is generally referred to as a proportional effect.
  • Normally distributed values: There is usually no proportional effect; in fact, the local standard deviations are roughly constant [3].
  • Lognormally distributed values: A scatterplot of local means versus local standard deviations will show a linear relationship between the two.
  • Figure 4.9: Shows a scatterplot of the local means versus the local standard deviations from the 16 local neighborhoods shown in Figure 4.7.
  • The correlation, coefficient is only 0.27, which is quite low.
  • There is no apparent relationship between the mean and the standard deviation for our 100 selected values.

Spatial Continuity

  • Spatial continuity exists in most earth science data sets.
  • Data points close to each other are more likely to have similar values than data points far apart.
  • Values don't appear randomly located; low values tend to be near low values, and high values tend to be near high values.
  • Zones of anomalously high values are of interest where the tendency of high values to be near other high values is very obvious.
  • A single very low value surrounded by high ones usually raises our suspicion.
  • Tools used to describe the relationship between two variables can also describe the relationship between the value of one variable and the value of the same variable at nearby locations.
  • The scatterplot can display spatial continuity, and the same statistics can summarize spatial continuity.

H-Scatterplots

  • An h-scatterplot displays all possible pairs of data values separated by a certain distance in a particular direction.
  • Vector notation is used to describe such pairs.
  • The letters 2 and y are used to refer to coordinates on a graphical display.
  • The letters x and y will be used to refer to coordinates that have spatial significance.
  • The location of any point can be described by a vector, as can the separation between any two points. Vector notation is convenient when describing pairs of values separated by a certain distance in a particular direction.
  • The location of the point at (x<em>i,y</em>i)(x<em>i,y</em>i) can be written as t<em>it<em>i, with the bold font of the t reminding us that it is a vector. Similarly, the location of the point at (x</em>j,y<em>j)(x</em>j,y<em>j) can be written as t</em>jt</em>j. The separation between point i and point j is t<em>jt</em>it<em>j - t</em>i, which can also be expressed as the coordinate pair (x<em>jx</em>i,y<em>jy</em>i)(x<em>j - x</em>i,y<em>j - y</em>i).
  • The symbol h<em>ijh<em>{ij} refers to the vector going from point i to point j and h</em>jih</em>{ji} for the vector from point j to i.
  • On h-scatterplots, the x-axis is labeled V(t) and the y-axis is labeled V(t + h).
  • The x-coordinate of a point corresponds to the V value at a particular location, and the y-coordinate corresponds to the V value a distance and direction h away.
  • If h=(0,1), each data location is paired with the data location whose easting is the same and whose northing is 1 m larger (i.e., the data location 1 m to the north).
  • Figure 4.11a shows the 90 pairs of data locations separated by exactly 1 m in a northerly direction for our 10 x 10 m2m^2 area.
  • If h=(l,l), each data location has been paired with the data location whose easting is 1 m larger and whose northing is also 1 m larger (i.e., the data location at a distance of 2\sqrt{2} m in a northeasterly direction).
  • Figure 4.11b shows the 81 pairs of data locations separated by 2\sqrt{2} m in a northeasterly direction for our 10 x 10 m2m^2 area.
  • The shape of the cloud of points on an h-scatterplot indicates how continuous the data values are over a certain distance in a particular direction.
  • If data values at locations separated by h are very similar, the pairs will plot close to the line x=yx = y, a 45-degree line passing through the origin.
  • As data values become less similar, the cloud of points on the h-scatterplot becomes fatter and more diffuse.
  • For the 100 selected V values, the change in continuity in a northerly direction is captured in Figures 4.12a-d, which show the four h-scatterplots for data values located 1 m apart to 4 m apart in a northerly direction.
  • Looking at these four h-scatterplots we see that from h=(0,1) to h=(0,4) the cloud gets progressively fatter as the points spread out, away from the 45-degree line.
  • Although the 45-degree line passes through the cloud of points on each h-scatterplot, the cloud is not symmetric about this line.
  • Direction is important in constructing an h-scatterplot; a pair of data values, u<em>iu<em>i and u</em>ju</em>j, appears on the h-scatterplot only once as the point (u<em>i,u</em>j)(u<em>i,u</em>j) and not again as (v<em>j,v</em>i)(v<em>j,v</em>i).
  • On four h-scatterplots in Figure 4.12 there are some points that do not plot close to the rest of the cloud.
  • Whenever we summarize h-scatterplots, we should be aware that our summary statistics may be influenced considerably by a few aberrant data values.
  • It is often worth checking how the summary statistics change if certain data are removed.

Correlation Functions, Covariance Functions, and Variograms

  • Quantitative summary of the information contained on an h-scatterplot.
  • Essential feature of an h-scatterplot is the fatness of the cloud of points.
  • As the cloud of points gets fatter, we expect the correlation coefficient to decrease.
  • Table 4.1: Correlation coefficients for each of the four scatterplots in Figures 4.12a-d.
  • As expected, the correlation coefficient steadily decreases and is therefore a useful index for our earlier impression that the cloud of points was getting fatter.
  • The relationship between the correlation coefficient of an h-scatterplot and h is called the correlation function or correlogram.
  • The correlation coefficient depends on h which, being a vector, has both a magnitude and a direction.
  • To graphically display the correlation function, one could use a contour map showing the correlation coefficient of the h-scatterplot as a function of both magnitude and direction.
  • In Figure 4.13a we use the data from Table 4.1 to show how the correlation coefficient decreases with increasing distance in a northerly direction.
  • Similar plots for other directions would give us a good impression of how the correlation coefficient varies as a function of both separation distance and direction.
  • An alternative index for spatial continuity is the covariance.
  • Table 4.1: Covariance for each of our h-scatterplots.
  • These also steadily decrease in a manner very similar to the correlation coefficient [5].
  • The relationship between the covariance of an h-scatterplot and h is called the covariance function [6].
  • Figure 4.13b shows the covariance function in a northerly direction.
  • Another plausible index for the fatness of the cloud is the moment of inertia about the line z = y, which can be calculated from the following:
    • moment of intertia=<em>i=12n(x</em>iyi)2nmoment \ of \ intertia = {\sum<em>{i=1}^{2n} (x</em>i - y_i)^2 \over n} (4.4)
      • It is half of the average squared difference between the z and y coordinates of each pair of points on the h-scatterplot, the factor f being a consequence of the fact that we are interested in the perpendicular distance of the points from the 45-degree line.
  • All of the points on the h-scatterplot for h = (0,O) will fall exactly on the line x = y since each value will be paired with itself.
  • As |h| increases, the points will drift away from this line and the moment of inertia about the 45-degree line is therefore a natural measure of the fatness of the cloud.
  • Unlike the other two indices of spatial continuity, the moment of inertia increases as the cloud gets fatter.
  • Table 4.1: Moment of inertia increases as the correlation coefficient and covariance decrease.
  • The relationship between the moment of inertia of an h-scatterplot and h is traditionally called the semivariogram [7].
  • The semivariogram, or simply the variogram, in a northerly direction is shown in Figure 4.13~.
  • The three statistics proposed for summarizing the fatness of the cloud of points on an h-scatterplot are all sensitive to aberrant points.
  • It is important to evaluate the effect of such points on our summary statistics.
  • Table 4.2: Compare the correlation coefficients from Table 4.1 to the correlation coefficients calculated with the 19 ppm value removed.
  • A single erratic value can have a significant impact on the correlation coefficient of an h-scatterplot.
  • The covariance and the moment of inertia are similarly affected.
  • The sensitivity of our summary statistics to aberrant points requires that we pay careful attention to the effect of any erratic values.
  • In practice, the correlation function, the covariance function, and the variogram often do not clearly describe the spatial continuity because of a few unusual values.
  • If the shape of any of these functions is not well defined it is worth examining the appropriate h-scatterplots to determine if a few points are having an undue effect.
  • H-scatterplots contain much more information than any of the three summary statistics.
  • Bypassing the h-scatterplots and going directly to either p(h), C( h) or Y( h) to describe spatial continuity is quite common.
  • Covariance function C(h)C(h), can be calculated from the following:
  • C(h)=1N(h)<em>i=1N(h)[v(t</em>i)m<em>h][v(t</em>i+h)m+h]C(h) = {1 \over N(h)} \sum<em>{i=1}^{N(h)} [v(t</em>i) - m<em>{-h}] [v(t</em>i + h) - m_{+h}] (4.2)
    • The data values are v1,. . . , vn;
    • The summation is over only the N(h) pairs of data whose locations are separated by h.
    • mhm_{-h} is the mean of all the data values whose locations are -h away from some other data location:
  • m<em>h=1N(h)</em>i=1N(h)v(ti)m<em>{-h} = {1 \over N(h)} \sum</em>{i=1}^{N(h)} v(t_i)
    • m+hm_{+h} is the mean of all the data values whose locations are +h away from some other data location:
  • m<em>+h=1N(h)</em>i=1N(h)v(ti+h)m<em>{+h} = {1 \over N(h)} \sum</em>{i=1}^{N(h)} v(t_i + h)
    • The values of m<em>hm<em>{-h} and m</em>+hm</em>{+h} are generally not equal in practice.
  • Correlation function, p( h), is the covariance function standardized by the appropriate standard deviations:
  • rho(h)=C(h)sigma<em>hsigma</em>+hrho(h) = {C(h) \over sigma<em>{-h} sigma</em>{+h}}
    • σh\sigma_{-h} is the standard deviation of all the data values whose locations are -h away from some other data location:
    • σ<em>h=</em>i=1N(h)[v(t<em>i)m</em>h]2N(h)\sigma<em>{-h} = {\sqrt{{\sum</em>{i=1}^{N(h)} [v(t<em>i) - m</em>{-h}]^2} \over N(h)} }
    • σ+h\sigma_{+h} is the standard deviation of all the data values whose locations are +h away from some other data location:
    • σ<em>+h=</em>i=1N(h)[v(t<em>i+h)m</em>+h]2N(h)\sigma<em>{+h} = {\sqrt{{\sum</em>{i=1}^{N(h)} [v(t<em>i + h) - m</em>{+h}]^2} \over N(h)} }
    • Like the means, the standard deviations, σ<em>h\sigma<em>{-h} and σ</em>+h\sigma</em>{+h}, are usually not equal in practice.
  • The variogram, Y( h), is half the average squared difference between the paired data values:
    • γ(h)=12N(h)<em>i=1N(h)[v(t</em>i)v(ti+h)]2\gamma(h) = {1 \over 2N(h)} \sum<em>{i=1}^{N(h)} [v(t</em>i) - v(t_i +h) ]^2 (4.8)
  • The values of p(h), C(h) and Y(h) are unaffected if we switch all of the i and j subscripts in the preceding equations.
  • γ(h)=γ(h)\gamma(h) = \gamma(-h) (4.11) - This result entails that the variogram calculated for any particular direction will be identical to the variogram calculated in the opposite direction. The correlation function and the covariance function share this property. For this reason, opposite directions are commonly combined when describing spatial continuity.

Cross h-Scatterplots

  • Extension of the idea of an h-scatterplot to that of a cross h-scatterplot.
  • Instead of pairing the value of one variable with the value of the same variable at another location, pairing values of different variables at different locations.
  • The crosses h-scatterplot between V and U for various values of h is shown in Figure 4.14.
  • The %-coordinate of each point is the V value at a particular data location, and the y-coordinate is the U value at a separation distance |h| to the north.
  • The scatterplot of V values versus the U values in Figure 3.4a can be thought of as a cross h-scatterplot for h = (0,O).
  • The 2-coordinate of each point corresponded to the V value at a particular location and the y-coordinate to the U value at the same location.
  • A comparison between the four scatterplots in Figure 4.14 shows that the relationship between the two variables becomes progressively weaker as | hl increases.
  • Describing the spatial continuity between variables is possible by using the correlation coefficient and the covariance.
  • Table 4.3: Correlation coefficient, the covariance, and variogram values of the four scatterplots shown in Figure 4.14.
  • Description of the relationship between statistics of a cross h-scatterplot and h using the cross-covariance and cross-correlation function.
  • The cross-covariance function between two variables can be calculated from the following equation:
    • C<em>uv(h)=1N(h)</em>i=1N(h)[u(t<em>i)m</em>uh][v(t<em>i+h)m</em>v+h]C<em>{uv}(h) = {1 \over N(h)} \sum</em>{i=1}^{N(h)} [u(t<em>i) - m</em>{u-h}] [v(t<em>i + h) - m</em>{v+h}] (4.12)
    • The data values of the first variable are u1,. . . , un and the data values of the second variable are v1,. . . , vn.
    • As in Equation 4.2, the summation is over only those pairs of data whose locations are separated by h.
    • muhm_{u-h} is the mean value of the first variable over those data locations which are -h away from some other v - type data location:
      • m<em>uh=1N(h)</em>i=1N(h)u(ti)m<em>{u-h} = {1 \over N(h)} \sum</em>{i=1}^{N(h)} u(t_i)
      • mv+hm_{v+h} is the mean value of the second variable over those locations that are +h away from some other u - type data location:
      • m<em>v+h=1N(h)</em>i=1N(h)v(ti+h)m<em>{v+h} = {1 \over N(h)} \sum</em>{i=1}^{N(h)} v(t_i + h)
  • The cross-correlation function is given by the equation
    • ρ<em>uv(h)=C</em>uv(h)σ<em>uhσ</em>v+h\rho<em>{uv}(h) = {C</em>{uv}(h) \over \sigma<em>{u-h} \sigma</em>{v+h}}
  • σuh\sigma_{u-h} is the standard deviation of the first variable at locations that are -h away from some other data location:
    • σ<em>uh=</em>i=1N(h)[u(t<em>i)m</em>uh]2N(h)\sigma<em>{u-h} = {\sqrt{{\sum</em>{i=1}^{N(h)} [u(t<em>i) - m</em>{u-h}]^2} \over N(h)} } (4.16)
  • σv+h\sigma_{v+h} is the standard deviation of the second variable at locations that are +h away from some other data location:
    • σ<em>v+h=</em>i=1N(h)[v(t<em>i+h)m</em>v+h]2N(h)\sigma<em>{v+h} = {\sqrt{{\sum</em>{i=1}^{N(h)} [v(t<em>i + h) - m</em>{v+h}]^2} \over N(h)} } (4.17)
  • Variogram extended to a cross-variogram equation:
    • γ<em>uv(h)=12N(h)</em>i=1N(h)[u(t<em>i)v(t</em>i+h)]2\gamma<em>{uv}(h) = {1 \over 2N(h)} \sum</em>{i=1}^{N(h)} [u(t<em>i) - v(t</em>i + h)]^2 (4.18)
  • In Figure 4.15a-c, cross correlation function, the cross covariance function, and the cross variogram between our 100 selected V and U values are shown.
  • If calculating cross correlation and covariance function in the reverse direction will result in different values.
  • Reversing the z and y values on a cross h-scatterplot entails switching not only the direction of h but also the order of the variables.
  • The cross variogram, however, is the same if we reverse the direction; that is γ<em>uv(h)=γ</em>vu(h)\,\gamma<em>{uv}(h) = \,\gamma</em>{vu}(-h).

The Exhaustive Data Set

  • The chapter deals with the description of the 78,000 data in the exhaustive data set, and the next two chapters will deal with the 470 data in the sample data set.
  • Encountering difficulties peculiar to the description of large, very dense data sets.
  • Overcoming the practical difficulties that such data sets pose.

Distribution of V

  • Table 5.1 Summarizes the distribution of the 78,000 V values.
  • The data values span several orders of magnitude, from 0 to 1,G31 ppm.
  • They are also strongly positively skewed; this makes it difficult to construct a single informative frequency table and its corresponding histogram.
  • The frequency table for the 50 ppm class interval given in Table 5.1 manages to cover most of the distribution. The corresponding