Comprehensive Notes on Inferential Statistics

Fundamental Definitions and Branches of Statistics

  • Definition of Statistics: Statistics is a branch of mathematics concerned with four primary activities:

    • Collection of data.

    • Tabulation or presentation of data.

    • Analysis of data.

    • Interpretation of data.

  • Branches of Statistics:

    • Descriptive Statistics: Methods concerned with collecting and describing a data set to yield meaningful information.

      • Examples: Mean (describing the center) and Standard Deviation (describing the spread around the mean).

    • Inferential Statistics: Methods concerned with analyzing a subset of data (samples) to lead to predictions or inferences about the entire set of data (population).

      • Examples: ztestz-test, ttestt-test, chisquaredtestchi-squared\,test, and Analysis of Variance (ANOVAANOVA).

Data Classification and Variables

  • Data Definition: Raw materials of statistical investigation that arise whenever measurements are made or observations are recorded.

    • Examples: Eye color, grades, scores, number of students.

  • Types of Data According to the Variable:

    • Quantitative Data (Numerical Data): Collected on variables that can be measured numerically.

      • Examples: Number of children, age, time, salary, exam scores.

    • Qualitative Data (Categorical Data): Collected on variables that cannot assume a numerical value but are divided into two or more nonnumeric categories.

      • Examples: Gender, eye color, religious status, academic rank, occupation, religion, learning style, and days of the week.

Characteristics and Scales of Measurement

  • Characteristics of Measurement:

    • Classification: Numbers used to group or sort responses.

    • Order: Numbers are ordered; one can be greater than, less than, or equal to another.

    • Distance: Differences between numbers are ordered and meaningful.

    • Origin: The number series has a unique origin indicated by the number zero (00).

  • Four Scales of Measurement:

    • Nominal: Classification only. No order, no distance, no unique origin.

      • Examples: Gender, religion, learning style.

    • Ordinal: Classification and order. No distance or origin.

      • Examples: Academic rank, educational attainment, interest level, income class (Lower, Middle, Upper).

    • Interval: Classification, order, and distance. No unique origin (zero is arbitrary).

      • Examples: Temperature (C^{\circ}C or F^{\circ}F), IQ scores, calendar years (e.g., 20092009, 20102010), clock time, grades.

    • Ratio: Classification, order, distance, and unique origin (absolute zero exists).

      • Examples: Monthly salary, weight, score in a test.

Data Collection and Gathering Methods

  • Sources of Data:

    • Primary Data: Gathered directly from primary sources (e.g., data from an interview).

    • Secondary Data: Gathered from existing secondary sources (e.g., government publications, mass media).

  • Methods of Gathering Data:

    • Observation: Purposeful, systematic watching or listening to a phenomenon.

      • Participant Observation: The observer participates in group activities.

      • Nonparticipant Observation: The observer does not get involved.

    • Interview: Person-to-person interaction with a specific purpose.

      • Unstructured: Follows a general framework or interview guide.

      • Structured: Follows a predetermined set of questions (interview schedule).

    • Questionnaire: A written list of questions recorded by respondents.

      • Open-ended: Possible responses are not provided.

      • Closed-ended: Respondents select from provided categories.

Introduction to Sampling and Population

  • Key Terms:

    • Population: The totality of observations statisticians are concerned with. The size is denoted by NN. A numerical characteristic of a population is a parameter (e.g., Population Mean: μ\mu).

    • Sample: A subset of the population. The size is denoted by nn. A numerical characteristic of a sample is a statistic (e.g., Sample Mean: Xˉ\bar{X}).

  • Sampling Categories:

    • Probability (Random) Sampling: Every member has an equal and independent chance of being included. Generally used for quantitative research.

    • Non-probability (Non-random) Sampling: Not all members have equal chances. Generally used for qualitative research.

  • Specific Sampling Strategies:

    • Probability Types: Simple Random, Systematic, Stratified Random (Proportionate/Disproportionate), Cluster (Single-stage/Multi-stage).

    • Non-probability Types: Convenience, Purposive (Extreme Case, Heterogeneous, Homogeneous, Critical Case, Typical Case), Quota, Snowball, and Self-Selection.

Presentation and Periodic Classification of Data

  • Three Ways of Presenting Data:

    1. Textual: In paragraph form.

    2. Tabular: Organized in rows and columns.

    3. Graphical: Visual representation.

  • Classification by Time:

    • Cross-Section Data: Collected on different elements at the same point in time.

    • Time-Series Data: Collected on the same element at different points in time (e.g., enrollment data over five years).

Research Question Types and Corresponding Statistical Tools

  1. Descriptive Questions: Identify responses to a single variable.

    • Tools: Frequency/Percentage, Mean/Median/Mode, Range/Variance/Standard Deviation, Skewness/Kurtosis.

  2. Relationship Questions: Determine the degree and magnitude of relationship between two or more variables.

    • Tools: Pearson's r, Spearman rho, Chi-squared test of independence, Regression Analysis, Discriminant Analysis.

  3. Comparison Questions: Find how two or more groups differ on an independent variable regarding outcome variables.

    • Tools: ttestt-test (for two groups), ANOVAANOVA (for three or more groups).

Graphical Presentation Tools

  • Bar Graph: Bars representing frequencies or percentages of categories.

  • Pie Graph: A circle divided into portions representing percentages. Best for five or fewer categories.

  • Line Graph: Shows relationships between sets of quantities over time (trends).

  • Pictogram/Pictograph: Uses symbols to represent values.

Measures of Central Tendency

  • Mean (xˉ\bar{x}): The average or sum of all values divided by the number of values.

    • Formula: xˉ=xn\bar{x} = \frac{\sum x}{n}

  • Median (MdMd): The middle term after arranging data in order.

    • If NN is odd: Median is the (N+12)th(\frac{N+1}{2})^{th} term.

    • If NN is even: Median is the mean of the two middlemost observations.

  • Mode (MoMo): The value with the highest frequency.

  • Advanced Mean Concepts:

    • Weighted Mean: Used when values have different levels of importance.

      • Formula: xˉw=(w×x)w\bar{x}_w = \frac{\sum (w \times x)}{\sum w}

    • Grand Mean: The mean of multiple group means.

      • Formula: Xˉgrand=nixˉini\bar{X}_{grand} = \frac{\sum n_i \bar{x}_i}{\sum n_i}

  • Comparison of Measures:

    • Stability: Mean is the most stable; Mode is the least stable.

    • Sensitivity: Mean is very sensitive to outliers; Median and Mode are not influenced by extreme values.

    • Mathematical Manipulation: Mean can be subjected to numerous computations; Mode is a terminal statistic.

Measures of Position (Quantiles)

  • Quartiles (QQ): Divide distribution into four parts (25%each25\%\,each).

    • Q1Q_1: Separates lower 25%25\% from upper 75%75\%.

  • Deciles (DD): Divide distribution into ten parts (10%each10\%\,each).

    • D3D_3: Separates lower 30%30\% from upper 70%70\%.

  • Percentiles (PP): Divide distribution into 100100 parts (1%each1\%\,each).

    • P73P_{73}: Separates lower 73%73\% from upper 27%27\%.

Measures of Variability (Spread)

  • Range: Difference between High and Low values (HLH - L).

  • Variance (s2s^2): Mean of squared deviations from the mean.

    • Sample Variance Formula: s2=(xxˉ)2n1s^2 = \frac{\sum (x - \bar{x})^2}{n - 1}

  • Standard Deviation (ss): Positive square root of variance.

    • Sample SD Formula: s=(xxˉ)2n1s = \sqrt{\frac{\sum (x - \bar{x})^2}{n - 1}}

  • Coefficient of Variation (CVCV): Ratio of Standard Deviation to Mean, expressed as a percentage.

    • Formula: CV=(sxˉ)×100%CV = (\frac{s}{\bar{x}}) \times 100\%

    • Higher CVCV indicates higher variability.

Distribution Shape: Skewness and Kurtosis

  • Skewness: Measures asymmetry.

    • Formula: Skewness=3(MeanMedian)s\text{Skewness} = \frac{3(\text{Mean} - \text{Median})}{s}

    • Symmetric (Normal): Skewness=0\text{Skewness} = 0 (Mean = Median = Mode).

    • Positively Skewed (Right): Mean > Median > Mode (tail points right).

    • Negatively Skewed (Left): Mean < Median < Mode (tail points left).

  • Kurtosis: Measures steepness/peakedness.

    • Mesokurtic: Normal (Kurtosis=3\text{Kurtosis} = 3).

    • Platykurtic: Flatter than normal (\text{Kurtosis} < 3).

    • Leptokurtic: More peaked than normal (\text{Kurtosis} > 3).

Hypothesis Testing and Errors

  • Steps in Objective Procedure:

    1. State H0H_0 (Null) and H1H_1 (Alternative).

    2. Select statistical test.

    3. Specify significance level (α\alpha).

    4. Specify sampling distribution.

    5. Define region of rejection.

    6. Compute test value.

    7. Make decision (reject or accept H0H_0).

    8. Make conclusion.

  • Hypothesis Types:

    • Null Hypothesis (H0H_0): Expresses non-significance of difference/relationship.

    • Alternative Hypothesis (H1H_1): Affirmative statement predicting outcome (can be directional or non-directional).

  • Errors in Decision Making:

    • Type I Error (α\alpha): Rejecting H0H_0 when it is actually true (False Positive).

    • Type II Error (β\beta): Accepting H0H_0 when it is actually false (False Negative).

Correlation and Association Studies

  • Pearson's Product Moment Correlation (rr): Measures strength of linear relationship between quantitative data. Range: 1r1-1 \leq r \leq 1.

  • Chi-Squared Test (χ2\chi^2): Determines association between nominal variables (non-parametric).

    • Formula: χ2=(OE)2E\chi^2 = \sum \frac{(O - E)^2}{E}, where OO is observed and EE is expected frequency.

Comparing Means (t-Tests and ANOVA)

  • Two-Sample t-Test: Used to compare means of two independent groups.

    • Excel Procedure: Data > Data Analysis > t-Test: Two-Sample Assuming Equal Variances.

  • Analysis of Variance (ANOVA): Used for comparing means of three or more groups.

    • Excel Procedure: Data > Data Analysis > Anova: Single Factor.

  • Interpreting P-values for All Tests:

    • p0.05p \geq 0.05: Non-significant.

    • 0.01 \leq p < 0.05: Significant.

    • p < 0.01: Highly Significant.

Computing Statistics Using MS Excel

  • Activating Analysis ToolPak:

    1. Go to Excel Options.

    2. Select Add-Ins.

    3. Click Go at the bottom.

    4. Check Analysis ToolPak and click OK.

Questions & Discussion

  • Q: Which variable is not categorical? (Age, Gender, Test choice, Marital status)

    • A: Age (it is quantitative).

  • Q: What is the best measure of location for skewed data?

    • A: Median.

  • Q: Can variance or SD be negative?

    • A: No. It can be zero (if all values are identical) or positive.

  • Q: If a student score is at the 73rd73^{rd} percentile out of 500500 students, what does it mean?

    • A: 73%73\% (365365 people) scored below the mark; 27%27\% (135135 people) scored above the mark.

  • Q: How is the Grand Mean of a class calculated if 1515 boys score 7575 and 2525 girls score 8585?

    • A: Xˉ=15(75)+25(85)40=81.25\bar{X} = \frac{15(75) + 25(85)}{40} = 81.25.