Data Analytics Notes - Complete Business Data Analytics (McMaster University)

Data Analytics Notes - Complete Business Data Analytics (McMaster University)

Assignments and Exam Structure

  • Assignments through A2L constitute 28% of course grade.

  • Midterm examination makes up 32% of course grade.

  • Final exam accounts for 40% of overall grading.

  • Importance of tutorials emphasized.

  • Frequent use of Excel in course content.

Data Variables

  • All variables classified as either qualitative or quantitative.

    • Time series: Data collected at regular intervals over time (e.g., daily temperature readings).

    • Cross-sectional data: Data collected at a specific point in time.

    • Understanding the difference between population (entire group) and sample (subset of the population).

Chapter 5: Descriptive Statistics

5.1 Histograms
  • Definition: Similar to a bar chart where bins (groups on horizontal axis) represent the count as bar heights.

  • Characteristic: No gaps between bars unless there's a gap in data.

  • Bins for categorical variables: One bar per category.

  • Bins for quantitative variables: Equal width chosen for data slicing.

  • Calculation for number of bins:

    • Rule: Number of bins = extlog2(n)ext{log}_2(n) or 2extnumberofbins=n2^{ ext{number of bins}} = n.

  • Endpoint determination: Choose round and interpretable numbers for bin limits.

  • Types of histograms:

    • Relative frequency histogram: Uses percentages rather than raw counts to represent data in bins.

  • Stem-and-leaf displays: Combines characteristics of histograms and individual value displays.

    • Example: Data of 2.1 represented as 2 | 1; with values 2.06, 2.22, 2.44, 3.28, 3.34 displayed as 2 | 124 and 3 | 33.

5.2 Distribution Shape
  • Characteristics of distribution to describe:

    • Shape: Focus on modes, symmetry, and outliers.

    • Modes:

    • Unimodal: One mode.

    • Bimodal: Two modes.

    • Multimodal: Three or more modes.

    • Uniformly distributed: No mode.

    • A distribution is symmetric if both sides can be folded along the middle and match.

    • Skewness: If one tail is longer than the other, it indicates skewed distribution.

    • Outliers: Data points that stand apart, significantly affecting methods reliant on mean and standard deviation.

5.3 Central Tendency Measures
  • Mean: The most common measure.

    • Formula for mean:

    • yˉ=racextTotaln=racextstyleextΣyb\bar{y} = rac{ ext{Total}}{n} = rac{ extstyle ext{Σ}y}{b} (where bb represents the count).

    • Sensitive to outliers, especially in skewed distributions.

  • Median:

    • Not affected by the shape; ideal measure for skewed data.

    • Steps to find median:

    1. n/2n/2: Middle index.

    2. Average the encompassing numbers if n/2n/2 is an integer, or round up if not.

  • Geometric Mean:

    • Useful for rates (e.g., interest rates): G=(a<em>1imesa</em>2imes…imesan)1/nG = (a<em>1 imes a</em>2 imes … imes a_n)^{1/n}.

5.4 Measuring Spread
  • Range: Difference between maximum and minimum values.

    • Formula: Range = max - min.

    • Influenced by outliers, so consider quartiles for robust analysis.

  • Interquartile Range (IQR):

    • IQR = Q<em>3−Q</em>1Q<em>3 - Q</em>1, focusing on middle fifty percent of the data.

  • Variance (s2s^2): Average of squared deviations from the mean.

    • Sample variance: s2=racextstyleextΣ(y−yˉ)2n−1s^2 = rac{ extstyle ext{Σ}(y - \bar{y})^2}{n - 1} .

    • Population variance: σ2=racextstyleextΣ(y−μ)2nσ^2 = rac{ extstyle ext{Σ}(y - μ)^2}{n}.

  • Standard Deviation (SD):

    • Derived from variance; measures spread around the mean:

    • Sample SD: s=ext√racextstyleextΣ(y−yˉ)2n−1s = ext{√} rac{ extstyle ext{Σ}(y - \bar{y})^2}{n - 1}.

    • Population SD: σ=ext√racextstyleextΣ(y−μ)2nσ = ext{√} rac{ extstyle ext{Σ}(y - μ)^2}{n}.

  • Coefficient of Variation (CV):

    • Measures relative variability: CV=racSDMeanCV = rac{SD}{Mean} .

5.5 Reporting Measures
  • Necessary to report:

    • Shape: Clearly indicate if the distribution is skewed; report both median and IQR.

    • Use mean and SD for symmetric distributions.

    • Report outliers separately and check effects thereof.

    • Pair median with IQR, and mean with SD for meaningful interpretation.

5.6 Adding Measures of Centre and Spread
  • Adding Means: Possible, but medians are not.

  • Calculating Variance: Only allow additive variances under uncorrelation condition; final SD derived afterward by square root.

5.7 Grouped Data
  • Midpoints will be used to estimate measures for data ranges.

    • Mean: Multiply midpoint by percent of data in range.

    • Variance: Formula adapted accordingly and standard deviation derived from variance.

5.8 Five-Number Summary and Boxplots
  • Five-Number Summary: Provides a quick summary (minimum, Q1, median, Q3, maximum).

  • Creating Boxplots:

    1. Extend vertical axis to encompass total data.

    2. Indicate lower and upper quartiles and median with horizontal lines, forming a box.

    3. Establish fences at 1.5 IQR above and below quartiles.

    4. Draw whiskers to extreme data points within fences.

    5. Display outliers with distinct symbols outside fences.

5.9 Percentiles
  • Computing Percentiles:

    1. Order data.

    2. Multiply desired percentile by total observations; round appropriately to find relevant position.

5.10 Comparing Data Groups
  • Histograms suitable for one/two distributions; boxplots optimal for more distributions.

  • Side-by-side comparison facilitates assessment of medians, spreads and overall ranges

5.11 Outliers Management
  • Assess outliers contextually; identify potential data entry errors or genuine deviations.

5.12 Standardizing Data
  • Standardization allows easier comparison of different variables via z-scores.

    • Calculation: z=y−yˉsz = \frac{y - \bar{y}}{s}.

5.13 Time Series Plots
  • Time series plots display data values over time; trends should be interpreted cautiously.

  • Stationary time series: No change over time.

5.14 Transforming Skewed Data
  • Apply transformations (e.g., logarithmic) to better summarize skewed distributions.

Chapter 6: Bivariate Data Analysis

6.1 Scatterplots
  • Scatterplots visualize relationships between two quantitative variables.

  • Important characteristics include:

    • Direction: Positive (upper left to lower right) or negative (upper right to lower left).

    • Form: Linear or nonlinear.

    • Strength: How tightly clustered data points are; outliers should be identified and assessed.

6.2 Role of Variables in Scatterplots
  • Assign independent variable to x-axis and dependent to y-axis for bivariate analysis.

6.3 Correlation
  • Correlation calculation evaluates strength and direction.

    • Formula: r=Σ(z<em>xz</em>y)n−1r = \frac{\Sigma(z<em>x z</em>y)}{n-1}.

  • Correlation won’t indicate causation; it's simply a measure of association.

6.4 Straightening Scatterplots
  • Logarithmic or square root transformations can help clarify non-linear relationships.

6.5 Lurking Variables
  • Lurking variables can affect conclusions drawn from correlation.

6.6 Rival Hypotheses
  • Attempt to identify and discuss rival hypotheses before concluding from data analyses.

6.7 Regression Analysis
  • Test slope using t-distributions under assumptions mentioned.

6.8 Moderation and Mediation
  • Evaluate moderation effects and mediation relationships in two-variable interactions.

Chapter 7: Hypothesis Testing

7.1 Set Hypotheses
  • Start with null hypothesis (H0) and alternative hypothesis (HA).

7.2 Determine Significance Levels
  • Select α (alpha level) to define significance threshold.

7.3 Conduct Statistical Test
  • Use p-values or critical values to determine if to reject null hypothesis.

7.4 Report Findings
  • Communicate results within the business context.

7.5 Power of a Test
  • Power: Probability of correctly rejecting a false null hypothesis.

  • Formula for power calculation is given based on sample design.

Chapter 8: Comparing Two Populations

8.1 Set Hypotheses for Differences
  • Null (H0): No difference between two means. Alternative (HA): Some difference.

8.2 Conduct t-Tests
  • Implement t-tests to evaluate significance of mean differences.

8.3 Conclusion Drawing
  • Discuss implications of findings based on hypothesis tests.

Other Content

  1. Random Variables and Probability Distributions.

  2. Testing Hypotheses About Proportions.

  3. Inference for Population Means.

  4. Linear Regression Models.

  5. Calculating Confidence Intervals.

  6. Analysis of Variance (ANOVA).

Conclusion

  • Thorough understanding of descriptive statistics, data relationships, hypothesis testing, and regression analysis is necessary for effective data analytics.