Practical Statistics for Data Scientists - Notes

Practical Statistics for Data Scientists - Notes

Practical Statistics for Data Scientists

  • Copyright © 2017 Peter Bruce and Andrew Bruce

  • First Edition: May 2017

Dedication

  • Dedicated to Victor G. Bruce, Nancy C. Bruce, John W. Tukey, Julian Simon, and Geoff Watson, inspiring figures in math and statistics.

Preface

  • Aimed at data scientists familiar with R and some statistics.

  • Highlights statistics' contributions to data science.

  • Acknowledges limitations of traditional statistics instruction.

  • Goals:

    • Present key statistical concepts relevant to data science in an accessible format.

    • Explain the importance and utility of these concepts from a data science perspective.

What to Expect

  • Data science is a fusion of statistics, computer science, IT, and domain-specific fields.

Conventions Used in This Book

  • Italic: Indicates new terms, URLs, email addresses, filenames, and file extensions.

  • Constant width: Used for program listings and references to program elements.

  • Constant width bold: Shows commands to be typed literally by the user.

  • Constant width italic: Shows text to be replaced with user-supplied values.

Using Code Examples

  • Supplemental material available for download.

  • Example code can be used in programs and documentation without permission, except for significant reproduction.

  • Attribution is appreciated but not required.

Safari® Books Online

  • Safari Books Online is an on-demand digital library for technology and business content.

  • Offers a range of plans for various sectors, providing access to numerous books, training videos, and prepublication manuscripts.

  • Available from multiple publishers.

How to Contact Us

  • O’Reilly Media, Inc. address provided.

  • Contact details for comments, technical questions, and general inquiries included.

  • Book web page: http://bit.ly/practicalStatsforDataScientists

Acknowledgments

  • Acknowledges contributions from various individuals and organizations, including Gerhard Pilcher, Anya McGuirk, Wei Xiao, Jay Hilfiger, Shannon Cutt, Kristen Brown, Rachel Monaghan, Eliahu Sussman, Ellen Troutman-Zaig, Marie Beaugureau, Ben Bengfort, Galit Shmueli, Elizabeth Bruce, and Deborah Donnell.

Exploratory Data Analysis

  • Statistics developed mostly in the past century.

  • Probability theory was developed in the 17th to 19th centuries.

  • Statistics is an applied science focused on data analysis and modeling.

  • Modern statistics traces its roots to Francis Galton and Karl Pearson.

  • R. A. Fisher pioneered experimental design and maximum likelihood estimation.

  • Exploratory Data Analysis (EDA) is a new area of ​​statistics.

  • Classical statistics focused on inference.

  • John W. Tukey called for a reformation of statistics, proposing data analysis as a new discipline.

  • Exploratory data analysis was established with Tukey’s 1977 book.

/

/

Elements of Structured Data
  • Data comes from various sources.

  • Challenge: harnessing raw data into actionable information.

  • Raw data must be processed into a structured form for statistical analysis.

  • Key terms for data types:

    • Continuous: Data that can take on any value in an interval (interval, float, numeric).

    • Discrete: Data that can take on only integer values ​​(integer, count).

    • Categorical: Data that can take on only categories (enums, enumerated, factors, nominal, polychotomous).

    • Binary: Special case of categorical data with two categories (dichotomous, logical, indicator, boolean).

    • Ordinal: Categorical data with explicit ordering (ordered factor).

  • Two basic types: numeric and categorical.

  • Data type is important for visual display, data analysis, and statistical models.

  • Data science software uses data types to improve computational performance.

  • Explicit identification of data as categorical offers advantages.

  • Explicit identification of data as categorical acts as signal of how statistical procedures should behave.

  • Explicit identification of data as categorical allows storage and indexing optimizations.

  • Explicit identification of data as categorical allows possible values enforcement in the software.

  • Default behavior of R is to automatically convert text columns into factors.

Rectangular Data
  • Typical frame of reference is a rectangular data object.

  • Key terms:

    • Data frame: Rectangular data structure.

    • Feature: A column in the table (attribute, input, predictor, variable).

    • Outcome: The variable to be predicted (dependent variable, response, target, output).

    • Records: A row in the table (case, example, instance, observation, pattern, sample).

  • Rectangular data: two-dimensional matrix with rows indicating records and columns indicating features.

  • There is a mix of measured/counted data and categorical data.

  • Traditional database tables have one or more columns designated as an index.

  • In Python (pandas), rectangular data structure is a DataFrame object.

  • In R, the basic rectangular data structure is a data.frame object.

  • Terminology for rectangular data can be confusing.

  • Statisticians and data scientists use different terms for the same thing.

Nonrectangular Data Structures
  • Time series data: successive measurements of the same variable.

  • Spatial data structures: used in mapping and location analytics.

  • Graph (or network) data structures: used to represent physical, social, and abstract relationships.

  • Focus of the book is on rectangular data.

Estimates of Location

  • Getting a typical value for each feature.

  • Estimating where most of the data is located.

  • Metrics often used include the mean, weighted mean, median, weighted median, and trimmed mean.

  • Key terms:

    • Mean: The sum of all values divided by the number of values (average).

    • Weighted mean: The sum of all values times a weight divided by the sum of the weights (weighted average).

    • Median: The value such that one-half of the data lies above and below (50th percentile).

    • Weighted median: The point at which 50% of weights lie above and below the sorted data.

    • Trimmed mean: The average of all values after dropping a fixed number of extreme values (truncated mean).

    • Robust: Not sensitive to extreme values (resistant).

    • Outlier: A data value that is very different from most of the data (extreme value).

  • Mean, easy to compute, may not always be the best measure.

  • Two goals underlie the use of weighted mean:

    • Some values are intrinsically more variable, and observations are given a lower weight.

    • The data does not equally represent different groups.

  • The median is less sensitive to outliers; it is more robust.

  • The trimmed mean is a compromise between the mean and the median.

  • <br>u<br>u N (or n) refers to the total number of records or observations. In statistics it is capitalized if it is referring to a population, and lowercase if it refers to a sample from a population.

  • trimmedmean=∑<em>i=p+1n−px</em>(i)n−2ptrimmed_mean = \frac{\sum<em>{i=p+1}^{n-p} x</em>{(i)}}{n-2p}

  • weightedmean=∑<em>i=1nw</em>ix<em>i∑</em>i=1nwiweighted_mean = \frac{\sum<em>{i=1}^{n} w</em>i x<em>i}{\sum</em>{i=1}^{n} w_i}

Median and Robust Estimates
  • The median is the middle number in the sorted data.

  • Using the median is better when looking at typical household incomes in areas like Medina neighborhood compared to Bill Gates house in Lake Washington area.

  • The median is a robust estimate of location, since it is not influenced by outliers.

Example: Location Estimates of Population and Murder Rates
  • Table 1-2 containing population and murder rates (per 100,000) for each state.

  • Calculations:

    • Mean population: 6,162,876

    • Trimmed mean (trim=0.1): 4,783,697

    • Median population: 4,436,370

    • Weighted mean murder rate: 4.445834

    • Weighted median murder rate: 4.4

Key Points
  • The sample mean is a good measure of central tendency, but it is not robust.

  • Other metrics, such as the median and trimmed mean are more robust.

Estimates of Variability

  • Variability (dispersion) measures whether data values are tightly clustered or spread out.

  • Variability lies at the heart of statistics.

  • Key terms for variability metrics:

    • Deviations: The difference between the observed values and the estimate of location (errors, residuals).

    • Variance: The sum of squared deviations from the mean divided by n – 1 (mean-squared-error).

    • Standard deviation: The square root of the variance (l2-norm, Euclidean norm).

    • Mean absolute deviation: The mean of the absolute value of the deviations from the mean (l1-norm, Manhattan norm).

    • Median absolute deviation from the median: The median of the absolute value of the deviations from the median.

    • Range: The difference between the largest and the smallest value in a data set.

    • Order statistics: Metrics based on the data values sorted from smallest to biggest (ranks).

    • Percentile: The value such that P percent of the values take on this value or less and (100–P) percent take on this value or more (quantile).

    • Interquartile range: The difference between the 75th percentile and the 25th percentile (IQR).

  • Different metrics for different ways to measure variability.

  • The variance is an average of the squared deviations.

  • The standard deviation is the square root of the variance.

  • The median absolute deviation from the median (MAD) is robust to outliers.

  • Statistics based on sorted (ranked) data are referred to as order statistics.

  • A common measurement of variability is the difference between the 25th percentile and the 75th percentile, called the interquartile range (or IQR).

Standard Deviation and Related Estimates
  • The best-known estimates for variability are the variance and the standard deviation, which are based on squared deviations.

  • The standard deviation is much easier to interpret than the variance since it is on the same scale as the original data.

  • The variance and standard deviation are especially sensitive to outliers.

  • A robust estimate of variability is the median absolute deviation from the median or MAD.

  • MAD=median(∣xi−m∣)MAD = median(\mid x_i - m \mid)

Estimates Based on Percentiles
  • A different approach to estimating dispersion is based on looking at the spread of the sorted data.

  • Statistics based on sorted (ranked) data are referred to as order statistics.

  • A common measurement of variability is the difference between the 25th percentile and the 75th percentile, called the interquartile range (or IQR).

Example: Variability Estimates of State Population
  • Table 1-3: population and murder rates for each state.

  • Calculations:

    • Standard deviation: 6,848,235

    • IQR: 4,847,308

    • MAD: 3,849,870

Exploring the Data Distribution

  • Useful to explore how data is distributed overall.

  • Key terms:

    • Boxplot: A plot to visualize the distribution of data (box and whiskers plot).

    • Frequency table: A tally of the count of numeric data values ​​that fall into a set of intervals (bins).

    • Histogram: A plot of the frequency table (bins on x-axis, count/proportion on y-axis).

    • Density plot: Smoothed version of the histogram, often based on a kernel density estimate.

Percentiles and Boxplots
  • Percentiles are valuable to summarize the tails of the area.

  • Table 1-4 displays percentiles of the murder rate by state.

  • Boxplots are based on percentiles and give a way to visualize the distribution of data.

Frequency Table and Histograms
  • A frequency table of a variable divides the variable range into equally spaced segments.

  • A histogram is a way to visualize a frequency table.

  • This gives us range of 37,253,956 – 563,626 = 36,690,330, which we must divide up into equal size bins — let’s say 10 bins.

Both frequency tables and percentiles summarize the data by creating bins.
Equal-count bins and equal-size bins provide different perspectives.
It is important to include the empty bins; the fact that there are no values in those bins is useful information

Density Estimates
  • A density plot shows the distribution of data values as a continuous line.

  • A density plot can be thought of as a smoothed histogram.

  • A key distinction from the histogram is scale of the y-axis.

  • The scale of the y-axis is a proportion rather than counts.
    Data for the histogram: hist(state[["Murder.Rate"]], freq=FALSE)
    Density estimate code: lines(density(state[["Murder.Rate"]]), lwd=3, col="blue")

Exploring Binary and Categorical Data

  • For categorical data, simple proportions or percentages describe

  • Key terms:

    • Mode: The most commonly occurring category or value in a data set.

    • Expected value: Average value based on a category’s occurrence probability when categories can be associated with a numeric value.

    • Bar charts: Frequency or proportion per category plotted as bars.

    • Pie charts: Frequency or proportion per category plotted as wedges in a pie.

  • Bar charts are a common visual tool for displaying a single categorical variable, often seen in the popular press.

  • Categories are listed on the x-axis, and frequencies or proportions on the y-axis.

  • Pie charts are an alternative to bar charts.

  • As an alternative, converting numeric data to categorical data is an important and widely used step in data analysis since it reduces the complexity (and size) of the data

The mode is the value — or values in case of a tie — that appears most often in the data.
The expected value is really a form of weighted mean

Expected Value Example
  • A marketer for a new cloud technology, for example, offers two levels of service, one priced at $300/month and another at $50/month. This data can be summed up in a single expected value.

  • Multiply outcome by probability of occurring

  • Sum these values

  • 0.05∗300+0.15∗50+0.8∗0=22.50.05 * 300 + 0.15 * 50 + 0.8 * 0 = 22.5

What Data Scientists Need to Know About The Mode

Categorical data is typically summed up in proportions, and can be visualized in a bar chart. Categories might represent distinct things, levels of a factor variable, or numeric data that has been binned. Expected value is the sum of values times their probability of occurrence, often used to sum up factor variable levels.

Correlation

  • Examining correlation among predictors, and between predictors and a target variable.

  • Variables X and Y are said to be positively correlated if high values ​​of X go with high values ​​of Y, and low values x of X go with low values ​​of Y.

  • If high values ​​of x go with low values ​​of Y, and vice versa, the variables are negatively correlated.

  • Key Terms for Correlation

    • Correlation coefficient: A metric that measures the extent to which numeric variables are associated with one another (ranges from –1 to +1).

    • Correlation matrix: A table where the variables are shown on both rows and columns, and the cell values ​​are the correlations between the variables.

    • Scatterplot: A plot in which the x-axis is the value of one variable, and the y-axis is the value of another.

  • More useful is a standardized variant: the correlation coefficient, which gives an estimate of the correlation between two variables that always lies on the same scale

  • Variables can have an association that is not linear, in which case the correlation coefficient may not be a useful metric.

  • A table of correlations is commonly plotted to visually display the relationship between multiple variables.

  • Using the package corrplot R, we can create correlation matrix

  • ETFs for the S&P 500 and the Dow Jones Index have a high correlation. Similary, the QQQ and the XLK, composed mostly of technology companies, are postively correlated. Defensive ETFs, such as those tracking gold prices, oil prices, or market volatility tend to be negatively correlated with the other ETFs.

  • ρ=∑<em>i=1n(x</em>i−xˉ)(y<em>i−yˉ)(n−1)s</em>xsy\rho = \frac{\sum<em>{i=1}^{n}(x</em>i - \bar x)(y<em>i - \bar y)}{(n-1)s</em>xs_y}

Scatterplots

  • The standard way to visualize the relationship between two measured data variables is with a scatterplot.

  • The x-axis represents one variable, the y-axis another, and each point on the graph is a record.

Basic Ideas for Correlation

The correlation coefficient measures the extent to which two variables are associated with one another. When high values of v1 go with high values ​​of v2, v1 and v2 are positively associated. When high values of v1 are associated with low values ​​of v2, v1 and v2 are negatively associated. The correlation coefficient is a standardized metric so that it always ranges from –1 to +1. A correlation coefficient of 0 indicates no correlation, but be aware that random arrangements of data will produce both positive and negative values ​​for the correlation coefficient just by chance.

Exploring Two or More Variables

  • Familiar estimators look at variables one at a time (univariate analysis).

  • Correlation analysis is an important method that compares two variables (bivariate analysis).

  • Key terms

    • Contingency tables: tally of counts between two or more categorical variables

    • Hexagonal binning: a plot of two numeric variables with the records binned into hexagons

    • Contour plots: a plot showing the density of two numeric variables like a topographical map

    • Violin plots: similar to boxplot but showing the density estimate

Lik univariate analysis, bivariate analysis involves both computing summary statistics and producing visual displays.
The appropriate type of bivariate or multivariate analysis depends on the nature of the data: numeric versus categorical

Visual Inspection
Hexagonal Binning and Contours (Plotting Numeric versus Numeric Data)
  • Scatterplots are fine when there is a relatively small number of data values.

  • For data sets with hundreds of thousands or millions of records, a scatterplot will be too dense, so we need a different way to visualize the relationship.

Two Categorical Variables
  • A useful way to summarize two categorical variables is a contingency table.

  • Contingency tables can look at just counts, or also include column and total percentages.

  • Table 1-8 table of loan grade and status show the contingency table between the grade of a personal loan and the outcome of that loan.

Categorical and Numeric Data
  • Boxplots are a simple way to visually compare the distributions of a numeric variable grouped according to a categorical variable.

Visualizing Multiple Variables
  • The types of charts used to compare two variables are readily extended to more variables through the notion of conditioning.

  • A cluster of homes that have higher tax-assessed value per square foot.

Graphs in Statistics

Hexagonal binning and contour plots are useful tools that permit graphical examination of two numeric variables at a time, without being overwhelmed by huge amounts of data. Contingency tables are the standard tool for looking at the counts of two categorical variables. Boxplots and violin plots allow you to plot a numeric variable against a categorical variable.

Summary

With the development of exploratory data analysis, statistics set a foundation that was a precursor to the field of data science. The key idea of ​​EDA is that the first and most important step in any project based on data is to look at the data.