First Year Higher Secondary Statistics Study Guide
Definitions, Scope, and Limitations of Statistics
Introduction to Statistics:
In the modern era of computers and information technology, the importance of statistics is recognized across all disciplines.
Originally evolved as a science of statehood, but now finds applications in Agriculture, Economics, Commerce, Biology, Medicine, Industry, Education, and Planning.
Origin and Growth:
The word ‘Statistics’ is derived from the Latin word Status, which means a political state.
It is a comparatively recent branch of the scientific method, with continuous research into its mathematical theory globally.
Verbatim Definition - Croxton and Cowden:
"Statistics may be defined as the science of collection, presentation, analysis and interpretation of numerical data from the logical analysis."
Verbatim Definition - Horace Secrist:
"Statistics may be defined as the aggregate of facts affected to a marked extent by multiplicity of causes, numerically expressed, enumerated or estimated according to a reasonable standard of accuracy, collected in a systematic manner, for a predetermined purpose and placed in relation to each other."
Bowley’s Definitions:
"Statistics are numerical statement of facts in any department of enquiry placed in relation to each other."
"Statistics may be called the science of counting."
"Statistics may be rightly called the scheme of averages."
Note: These are considered incomplete as they ignore aspects like interpretation and analysis.
Functions of Statistics:
Condensation: Reducing huge masses of data into manageable observations (e.g., using averages and ranges).
Comparison: Using classification and tabulation to compare data across regions or sources.
Forecasting: Predicting future trends (e.g., rainfall, business profit) using Time Series and Regression analysis.
Estimation: Drawing inferences about a population from a sample (Estimation Theory).
Tests of Hypothesis: Formulating and testing statements about population distributions (e.g., drug efficacy).
Scope of Statistics:
Industry: Uses Control Charts and inspection plans to maintain quality levels.
Commerce: Market surveys and demand forecasting are essential for managing stock and competition.
Agriculture: Analysis of Variance (ANOVA), developed by Professor R.A. Fisher, tests the significance of differences in crop yields under various fertilizers.
Economics: Alfred Marshall stated, "Statistics are the straw only which I like every other economist have to make the bricks."
Planning: Indispensable for government policy formulation regarding production, consumption, and income-expenditure.
Medicine: Uses the to compare the efficiency of different drugs.
Limitations:
Qualitative Data: It cannot directly study honesty, beauty, or poverty unless reduced to numerical terms.
Individuals: It deals only with aggregates, not individual items.
Approximations: Statistical laws are not as exact as physical sciences; they are true only on average.
Misuse: Prof. King observes, "Statistics are like clay of which one can make a God or Devil as one pleases."
Introduction to Sampling Methods
Population and Census:
Population (Universe): The complete set of observations investigated.
Finite Population: Consists of a reachable number of units (e.g., workers in a factory).
Infinite Population: Uncountable units (e.g., stars in the sky).
Census Method: Every element of the population is surveyed. It is accurate but costly and time-consuming.
Sampling Terminology:
Sample: A finite subset of individuals chosen from the population.
Sample Size (): The number of units in the sample.
Sampling Unit: The individuals to be sampled (e.g., a family head for an income survey).
Sampling Frame: A list or map identifying each sampling unit (e.g., voters list).
Parameters vs. Statistics:
Parameters: Characteristics of a population (Mean , Standard Deviation , size ).
Statistics: Characteristics of a sample (Mean , s.d. , size ).
Principles of Sampling:
Statistical Regularity: A large number of units chosen at random will likely possess the characteristic of the whole group.
Inertia of Large Numbers: Accuracy increases with sample size.
Optimisation: Achieving maximum information with minimum cost and time.
Types of Sampling:
Probability (Random): Selection based on known probabilities (e.g., Simple Random Sampling).
Non-Probability (Non-Random): Based on personal judgment or quotas.
Simple Random Sampling (SRS): Every item has an equal probability. Methods include the Lottery Method and the Table of Random Numbers (Tippett’s, Fisher and Yates’, or Kendall and Smith’s tables).
Stratified Random Sampling: Dividing heterogeneous populations into homogeneous strata. Proportional allocation uses .
Systematic Sampling: Selecting every element, where .
Collection of Data, Classification, and Tabulation
Nature of Data:
Time Series: Collected over time (e.g., household annual expenditure).
Spatial Data: Relates to geographical locations (e.g., district-wise rainfall).
Spacio-Temporal Data: Relates to both time and space (e.g., population of states across census years).
Primary Data Collection Methods:
Direct Personal Interviews: High response rate and accurate, but costly.
Indirect Oral Interviews: Interviewing third parties (e.g., for theft or murder cases).
Questionnaires: A series of questions mailed to respondents. A good questionnaire should have short, simple, logical, and non-sensitive questions.
Schedules: Similar to questionnaires but filled by trained enumerators during face-to-face contact.
Secondary Data:
Data collected by others (published reports, journals, government records).
Sources: IMF, UN, Central/State governments, trade bodies, business journals.
Classification and Tabulation:
Classification Types: Chronological (Time), Geographical (Region), Qualitative (Attributes like sex/literacy), and Quantitative (Measurable variables like height/weight).
Table Structure: Table Number, Title, Captions (vertical headings), Stubs (horizontal headings), Body, Footnotes, and Source.
Frequency Distribution
Sturges' Rule:
To determine the number of class intervals (): .
The width of class interval (): .
Discrete vs. Continuous:
Discrete: Variables with definite differences (e.g., number of children).
Continuous: Variables take any fractional value (e.g., weights in kgs).
Classification Methods:
Exclusive Method: Upper limit of one class is the lower limit of the next (e.g., 0-10, 10-20).
Inclusive Method: Both limits are included (e.g., 10-19, 20-29).
Open-end Classes: Limits missing at the start or end (e.g., "Below 2000" or "Above 8000").
Cumulative Frequencies:
Less than: Running total from top down.
More than: Running total from bottom up.
Diagrammatic and Graphical Representation
One-Dimensional Diagrams:
Line, Simple Bar, Multiple Bar (comparing subsets), Sub-divided Bar (components), and Percentage Bar.
Area/Two-Dimensional Diagrams:
Rectangles, Squares, and Pie Diagrams (Sector calculation: ).
Graphs:
Histogram: Rectangles representing frequencies; width is class interval.
Frequency Polygon: Midpoints of histogram rectangles joined by straight lines.
Frequency Curve: Smooth freehand curve through the polygon points.
Ogive: Cumulative frequency curve (used to find Median).
Lorenz Curve: Measures socio-economic inequality (e.g., wealth distribution).
Measures of Central Tendency
Arithmetic Mean ():
Ungrouped: .
Grouped (Assumed Mean Method): , where .
Weighted Mean: .
Geometric Mean (G.M.):
Defined as the root of the product of observations.
Formula: .
Harmonic Mean (H.M.):
Reciprocal of the arithmetic average of the reciprocals of observations.
Formula: .
Median ():
The middle value dividing the distribution into two equal parts.
Continuous distribution formula: .
Mode ():
The most frequent value.
Formula: .
Empirical Relationship: .
Measures of Dispersion, Skewness, and Kurtosis
Absolute Measures:
Range: .
Quartile Deviation (Q.D.): .
Mean Deviation (M.D.): .
Standard Deviation (S.D. or ): .
Relative Measures:
Coefficient of Variation (C.V.): . Lower C.V. means higher consistency.
Moments:
Arithmetic mean of various powers of deviations from the actual mean ().
, .
Skewness:
Lack of symmetry.
Symmetrical: .
Positive: .
Negative: .
Kurtosis:
Measures peakedness.
Mesokurtic: .
Leptokurtic: .
Platykurtic: .
Correlation and Regression
Correlation Coefficient ():
Varies between and .
Karl Pearson’s Formula: .
Spearman’s Rank Correlation: .
Regression Analysis:
Predicting a dependent variable () based on an independent variable ().
Regression Line of on : , where .
Regression Line of on : , where .
Geometric Property: .
Index Numbers
Classification:
Price Index: Measures change in price level.
Quantity Index: Measures volume of production/consumption.
Weighted Aggregate Indices:
Laspeyre’s: Uses base year weights (). .
Paasche’s: Uses current year weights (). .
Fisher’s Ideal Index: Geometric mean of Laspeyre and Paasche. .
Tests of Consistency:
Time Reversal Test: .
Factor Reversal Test: .
Note: Fisher’s Ideal Index satisfies both tests.
Consumer Price Index (Cost of Living):
Aggregate Expenditure Method: Identical to Laspeyre’s.
Family Budget Method: , where .