Statistics and Field Experimentation I Study Notes
Meaning and Philosophical Origins of Statistics
Conceptual Definitions: * The term "statistics" carries varied meanings depending on the context of the user: * Environmental Protection Agency (EPA) Administrator: Statistics refer to information regarding the quantity and type of pollutants released into the environment. * Farm Manager: Statistics consist of daily reports on production, inventory levels, and absenteeism. * University Student: Statistics are the grades achieved in all courses taken during a specific examination period. * Broadly, statistics involve the collection of numerical information called Data, the analysis of that data, and the making of meaningful decisions based on that analysis.
Historical Etymology: * The word "STATISTIK" originates from the Italian word "STATISTA," which translates to "STATESMAN." * It was first utilized by Gittifield Achenwall (1719–1772), a Professor at Marlborough and Gottingen. * It was introduced to England by Dr EAW Ziminarman. * The term was popularized by Sir John Sinclair in his seminal work titled STATISTICS ACCOUNT OF SCOTLAND (1791–1799).
Ancient Statistical Practices: * Quantitative information was in use long before the term was coined. Official government statistics are as old as recorded history. * The Old Testament contains several accounts of census-taking. * Ancient Babylonians, Egyptians, Romans, and Israelites gathered detailed information regarding population and resources.
Classifications of Statistical Methods
Descriptive (Deductive) Statistics: * This involves the collection of data and the tabulation of results using tables, charts, or graphs to make data manageable and meaningful. * Examples: * A Professor computing an average grade for a course taught last semester to describe the class performance. * Recording daily meteorological data for a weather station (e.g., from Jan–Dec 1999) is an example of Secondary Data (data already collected by someone else). * Primary Data is data collected directly by a researcher for their specific purpose, such as a researcher going to a farm to collect info on production and practices.
Inferential (Inductive) Statistics: * This is a more complex approach used to make predictions about a large group (the Population or Universe) based on a smaller subset (the Sample). * Collecting data from an entire population (e.g., the average height of all students at the University of Ibadan) is often impossible due to constraints in time, money, and labor. Therefore, a representative sample is selected. * Generalizations or inferences about the population are made from this sample, provided the sample is truly representative.
Data Types and Variables
Quantitative Variables: * Discrete: Variables that take on specific, distinct values (e.g., number of children, number of cars). * Continuous: Variables that can take any value within a range (e.g., height, weight, temperature).
Qualitative Variables: * Nominal: Categories with no inherent order (e.g., gender, eye color). * Ordinal: Categories with a logical rank or order (e.g., socio-economic status: low, medium, high).
Construction of Frequency Distributions
The Frequency Distribution is the first step in processing a raw data set. It organizes data into classes and similar categories so that the number of observations in each category can be seen at a glance.
Case Study: Age at First Marriage of 50 Women in Ibadan: * Data: 20, 23, 21, 27, 24, 34, 19, 30, 25, 28, 17, 28, 18, 40, 19, 41, 23, 31, 23, 29, 24, 38, 21, 27, 22, 35, 23, 39, 17, 34, 18, 23, 23, 26, 19, 20, 25, 44, 24, 31, 16, 20, 22, 27, 24, 29, 23, 28, 18, 25. * Formulas for Grouping: * Range = * Range = * Width of Class Interval () = * With a maximum value of 44 and minimum of 16, . If using 7 intervals, width is approximately 4.
Frequency Table Components: * Relative Frequency (): Shows the proportion of the distribution in a class interval. * * Cumulative Frequency (): The running total of frequencies through the classes. * Mid-class (Class Mark): The average of the upper and lower limits of the class interval.
Graphical Representation of Data
Histogram: A graph of frequency distribution where each class interval is represented by a block/rectangle with no space between bars. Frequency is on the vertical axis, and class intervals are on the horizontal axis.
Frequency Polygon: A line graph created by connecting the mid-points of the histogram bars. It starts and ends at the horizontal axis (requiring extra classes for extrapolation).
Cumulative Frequency Curve (Ogive): A graph showing the cumulative totals of frequencies.
Relative Frequency Curve: Similar in shape to the Ogive but plots relative frequency proportions on the vertical axis.
Skewness and Kurtosis
Skewness: Measures the asymmetry of the probability distribution. * Positively Skewed: The tail of the curve tilts toward the right. * Negatively Skewed: The tail of the curve tilts toward the left.
Kurtosis: Measures the "peakedness" of the curve. Different curves can have the same central location and dispersion but different degrees of kurtosis.
Measures of Central Tendency
Arithmetic Mean (): * Ungrouped Data: * * Grouped Data: * * Coding/Assumed Mean Method: * * Where is the class mark assigned the code zero, is the width of the interval, and is the assigned code.
Weighted Mean (): * Used to account for the relative importance of each value to the overall total. * * Example Application: Calculating average labor cost per hour for different products using skilled and unskilled labor.
Geometric Mean (): * Used for determining the average growth rate or rate of change over time. * * Compound Interest/Growth Formula: * * Example: Rate of production increases by in Year 1 and in Year 2. The average growth rate is calculated by finding the of the growth factors (1.25 and 1.40).
Median: * The middle-most item in an ordered array. * Grouped Data Formula: * * Where is the lower class limit of the median class, is the cumulative frequency before the median class, and is the frequency of the median class.
Mode: * The value that occurs most frequently. * Grouped Data Formula: * * Where is the difference between the modal class frequency and the frequency of the preceding class, and is the difference between the modal class frequency and the frequency of the succeeding class.
Measures of Dispersion
Distance Measures: * Range: . It is limited as it only considers extremes. * Interfractile Range: Measures the difference between two fractiles (e.g., between the 1st and 2nd thirds of income). Fractiles describe proportions lying above or below a location: deciles (10 parts), quintiles (5 parts), quartiles (4 parts).
Average Deviation Measures: * Mean Deviation (): * * * Variance (): * * * Standard Deviation (): The square root of the variance.
Coefficient of Variation (): * A relative measure of dispersion, used to compare the variability of datasets with different units or scales. * * Higher indicates lower efficiency or higher variability.
Chebyshev's Theorem: * Devised by P.L. Chebyshev (1821–1894). * States that regardless of distribution shape, at least of values fall within from the mean, and at least fall within . * For bell-shaped (Normal) curves: within , within , and within .
Probability Theory and Set Operations
Foundations: Facilitated by set theory, developed by GEORGI CANTOUK (1845).
Set Definitions: * Subset: Set A is a subset of B () if all elements in A are in B. * Proper Subset: B contains at least one element not in A. * Universal Set (): The set containing all objects of interest. * Empty/Null Set (): A set with no elements.
Set Operations: * Union (): Contains all elements in A, B, or both. * Intersection (): Elements common to both A and B. * Disjoint Sets: Sets with no common elements (). * Complement (): Elements in the Universal set that are not in A.
Laws of Set Operations: * Idempotent (), Commutative (), Associative, Distributive, De Morgan’s Law ().
Concepts of Probability
Definitions: * Experiment: A process like tossing a coin or rolling a die. * Sample Space (): Set of all possible outcomes (e.g., for two coins, ). * Event: A subset of the sample space.
Types of Probability: * Subjective: Based on personal belief, evidence, or educated guesses. * Objective: Divided into Classical () and Relative Frequency. * Classical: . Outcomes are equally likely. * Relative Frequency: Probability determined by the proportion of times an event occurs in the long run.
Probability Rules: * * - If , the event is impossible. If , it is certain. * - .
Event Relationships: * Mutually Exclusive: Events that cannot occur simultaneously (). * Independent: The occurrence of one does not affect the other (). * Conditional Probability: The probability of B occurring given that A has already occurred. *
Bayes' Theorem: * Deals with revising prior probabilities after additional information (likelihoods) becomes available, resulting in posterior probabilities. *
Combinatorial Analysis and Distributions
Fundamental Principle: If one event can occur in ways and another in ways, both can occur in sequence in ways.
Permutations (): Arranging objects out of where order matters. * * With Repetition: (Example: arrangements of the word "STATISTICS").
Combinations (): Selection where order does not matter. *
Binomial (Bernoulli) Distribution: * Discovered by James Bernoulli. Used for independent trials with two outcomes (success/failure). * * Mean = , Variance = .
Poisson Distribution: * Discovered by S.D. Poisson. Used for "rare events" where is large and is small ( stays stable). * * Where . Mean and Variance both equal .
Normal Distribution: * A continuous distribution represented by the bell-shaped curve. * * Standardized variable: .
Sampling Theory and Techniques
Population vs Sample: * Population: All possible observations of a type. Measures are called Parameters (e.g., ). * Sample: A subset of the population. Measures are called Statistics (e.g., ).
Sampling Methods: 1. Convenience Sampling: Using readily available observations. 2. Judgment Sampling: Sampler deliberately selects items based on knowledge. 3. Random Sampling: Every element has an equal chance of selection. * Simple Random: Ballot method or Table of random numbers. * Systematic: Selection at regular intervals. Raising factor . * Stratified: Divide population into strata (low internal difference, high external difference) then sample randomly from each. * Cluster: Divide population into heterogeneous blocks (clusters) then select clusters and sample within them (two-stage).
Basic Estimators in Sampling: * Sample Mean: . * Standard Error of the Mean (): * , where is the sampling fraction. * The Finite Population Correction (FPC) is ignored if . * Confidence Interval for Mean: . * Proportions: . Variance of proportion .
Principles of Hypothesis Testing
Definitions: * Statistical Hypothesis: A numerical statement about a population parameter that needs verification. * Null Hypothesis (): Stated in the negative or "no difference" form (e.g., "There is no difference in achievement between boys and girls"). * Research/Alternative Hypothesis (): Predicts an outcome in a positive form or specifies a difference.
Errors in Testing: * Type I Error: Rejecting when it should have been accepted (). * Type II Error: Accepting when it should have been rejected.
Testing Parameters: * Level of Significance (): Usually () or (). * One-tailed Test: directional (e.g., "New drug is better"). * Two-tailed Test: non-directional (e.g., "There is a difference").
Z-Tests (Large Samples, ): * Single Mean: . * Difference of Means: .
T-Tests (Small Samples, n < 30): * Requires Degrees of Freedom ( or ). * Formula for difference of means uses pooled variance ().
Chi-Square () Analysis
Measures the discrepancy between Observed () and Expected () frequencies. * * (for one-way) or (for contingency tables).
Yate's Correction: Applied when . *
Correlation and Regression
Correlation (): Investigates the degree to which variables move together. * Positive: Both move in the same direction. * Negative: Move in opposite directions. * Zero: No linear relationship. * Spearman Rank Correlation (): .
Regression: Measures the amount of change in a dependent variable () caused by a unit change in an independent variable (). * Model: * . * .
Analysis of Variance (ANOVA)
Derived by Sir Ronald Fisher (1923), the F-ratio () is used to compare three or more means simultaneously. *
ANOVA Designs: 1. One-way (Completely Randomized Design - CRD): Tests effects of one factor. 2. Two-way (Randomized Complete Block Design - RCBD): Adds a second factor (blocking/education) to reduce error. * With Replications: Allows testing of the interaction term between factors. 3. Latin Square Design: Efficient when there are two sources of external variation (e.g., fertility gradients in two directions). Each treatment appears once in each row and column.
The ANOVA Table: Summarizes Source, Sum of Squares (), Degrees of Freedom (), Mean Squares (), and the Calculated F-ratio () against Critical F values ().