First Year Higher Secondary Statistics Study Guide

Definitions, Scope, and Limitations of Statistics

  • Introduction to Statistics:

    • In the modern era of computers and information technology, the importance of statistics is recognized across all disciplines.

    • Originally evolved as a science of statehood, but now finds applications in Agriculture, Economics, Commerce, Biology, Medicine, Industry, Education, and Planning.

  • Origin and Growth:

    • The word ‘Statistics’ is derived from the Latin word Status, which means a political state.

    • It is a comparatively recent branch of the scientific method, with continuous research into its mathematical theory globally.

  • Verbatim Definition - Croxton and Cowden:

    • "Statistics may be defined as the science of collection, presentation, analysis and interpretation of numerical data from the logical analysis."

  • Verbatim Definition - Horace Secrist:

    • "Statistics may be defined as the aggregate of facts affected to a marked extent by multiplicity of causes, numerically expressed, enumerated or estimated according to a reasonable standard of accuracy, collected in a systematic manner, for a predetermined purpose and placed in relation to each other."

  • Bowley’s Definitions:

    • "Statistics are numerical statement of facts in any department of enquiry placed in relation to each other."

    • "Statistics may be called the science of counting."

    • "Statistics may be rightly called the scheme of averages."

    • Note: These are considered incomplete as they ignore aspects like interpretation and analysis.

  • Functions of Statistics:

    • Condensation: Reducing huge masses of data into manageable observations (e.g., using averages and ranges).

    • Comparison: Using classification and tabulation to compare data across regions or sources.

    • Forecasting: Predicting future trends (e.g., rainfall, business profit) using Time Series and Regression analysis.

    • Estimation: Drawing inferences about a population from a sample (Estimation Theory).

    • Tests of Hypothesis: Formulating and testing statements about population distributions (e.g., drug efficacy).

  • Scope of Statistics:

    • Industry: Uses Control Charts and inspection plans to maintain quality levels.

    • Commerce: Market surveys and demand forecasting are essential for managing stock and competition.

    • Agriculture: Analysis of Variance (ANOVA), developed by Professor R.A. Fisher, tests the significance of differences in crop yields under various fertilizers.

    • Economics: Alfred Marshall stated, "Statistics are the straw only which I like every other economist have to make the bricks."

    • Planning: Indispensable for government policy formulation regarding production, consumption, and income-expenditure.

    • Medicine: Uses the ttestt-\text{test} to compare the efficiency of different drugs.

  • Limitations:

    • Qualitative Data: It cannot directly study honesty, beauty, or poverty unless reduced to numerical terms.

    • Individuals: It deals only with aggregates, not individual items.

    • Approximations: Statistical laws are not as exact as physical sciences; they are true only on average.

    • Misuse: Prof. King observes, "Statistics are like clay of which one can make a God or Devil as one pleases."

Introduction to Sampling Methods

  • Population and Census:

    • Population (Universe): The complete set of observations investigated.

    • Finite Population: Consists of a reachable number of units (e.g., workers in a factory).

    • Infinite Population: Uncountable units (e.g., stars in the sky).

    • Census Method: Every element of the population is surveyed. It is accurate but costly and time-consuming.

  • Sampling Terminology:

    • Sample: A finite subset of individuals chosen from the population.

    • Sample Size (nn): The number of units in the sample.

    • Sampling Unit: The individuals to be sampled (e.g., a family head for an income survey).

    • Sampling Frame: A list or map identifying each sampling unit (e.g., voters list).

  • Parameters vs. Statistics:

    • Parameters: Characteristics of a population (Mean μ\mu, Standard Deviation σ\sigma, size NN).

    • Statistics: Characteristics of a sample (Mean xˉ\bar{x}, s.d. ss, size nn).

  • Principles of Sampling:

    • Statistical Regularity: A large number of units chosen at random will likely possess the characteristic of the whole group.

    • Inertia of Large Numbers: Accuracy increases with sample size.

    • Optimisation: Achieving maximum information with minimum cost and time.

  • Types of Sampling:

    • Probability (Random): Selection based on known probabilities (e.g., Simple Random Sampling).

    • Non-Probability (Non-Random): Based on personal judgment or quotas.

    • Simple Random Sampling (SRS): Every item has an equal probability. Methods include the Lottery Method and the Table of Random Numbers (Tippett’s, Fisher and Yates’, or Kendall and Smith’s tables).

    • Stratified Random Sampling: Dividing heterogeneous populations into homogeneous strata. Proportional allocation uses nN=c\frac{n}{N} = c.

    • Systematic Sampling: Selecting every KthK^{th} element, where K=NnK = \frac{N}{n}.

Collection of Data, Classification, and Tabulation

  • Nature of Data:

    • Time Series: Collected over time (e.g., household annual expenditure).

    • Spatial Data: Relates to geographical locations (e.g., district-wise rainfall).

    • Spacio-Temporal Data: Relates to both time and space (e.g., population of states across census years).

  • Primary Data Collection Methods:

    • Direct Personal Interviews: High response rate and accurate, but costly.

    • Indirect Oral Interviews: Interviewing third parties (e.g., for theft or murder cases).

    • Questionnaires: A series of questions mailed to respondents. A good questionnaire should have short, simple, logical, and non-sensitive questions.

    • Schedules: Similar to questionnaires but filled by trained enumerators during face-to-face contact.

  • Secondary Data:

    • Data collected by others (published reports, journals, government records).

    • Sources: IMF, UN, Central/State governments, trade bodies, business journals.

  • Classification and Tabulation:

    • Classification Types: Chronological (Time), Geographical (Region), Qualitative (Attributes like sex/literacy), and Quantitative (Measurable variables like height/weight).

    • Table Structure: Table Number, Title, Captions (vertical headings), Stubs (horizontal headings), Body, Footnotes, and Source.

Frequency Distribution

  • Sturges' Rule:

    • To determine the number of class intervals (KK): K=1+3.322log10(N)K = 1 + 3.322\log_{10}(N).

    • The width of class interval (CC): C=Range1+3.322log10(N)C = \frac{\text{Range}}{1 + 3.322\log_{10}(N)}.

  • Discrete vs. Continuous:

    • Discrete: Variables with definite differences (e.g., number of children).

    • Continuous: Variables take any fractional value (e.g., weights in kgs).

  • Classification Methods:

    • Exclusive Method: Upper limit of one class is the lower limit of the next (e.g., 0-10, 10-20).

    • Inclusive Method: Both limits are included (e.g., 10-19, 20-29).

    • Open-end Classes: Limits missing at the start or end (e.g., "Below 2000" or "Above 8000").

  • Cumulative Frequencies:

    • Less than: Running total from top down.

    • More than: Running total from bottom up.

Diagrammatic and Graphical Representation

  • One-Dimensional Diagrams:

    • Line, Simple Bar, Multiple Bar (comparing subsets), Sub-divided Bar (components), and Percentage Bar.

  • Area/Two-Dimensional Diagrams:

    • Rectangles, Squares, and Pie Diagrams (Sector calculation: Component ValueTotal×360\frac{\text{Component Value}}{\text{Total}} \times 360^{\circ}).

  • Graphs:

    • Histogram: Rectangles representing frequencies; width is class interval.

    • Frequency Polygon: Midpoints of histogram rectangles joined by straight lines.

    • Frequency Curve: Smooth freehand curve through the polygon points.

    • Ogive: Cumulative frequency curve (used to find Median).

    • Lorenz Curve: Measures socio-economic inequality (e.g., wealth distribution).

Measures of Central Tendency

  • Arithmetic Mean (xˉ\bar{x}):

    • Ungrouped: xˉ=xn\bar{x} = \frac{\sum x}{n}.

    • Grouped (Assumed Mean Method): xˉ=A+fdN×c\bar{x} = A + \frac{\sum fd}{N} \times c, where d=xAcd = \frac{x-A}{c}.

    • Weighted Mean: xˉw=wxw\bar{x}_{w} = \frac{\sum wx}{\sum w}.

  • Geometric Mean (G.M.):

    • Defined as the nthn^{th} root of the product of nn observations.

    • Formula: G.M.=Antilog(logxn)\text{G.M.} = \text{Antilog}\left( \frac{\sum \log x}{n} \right).

  • Harmonic Mean (H.M.):

    • Reciprocal of the arithmetic average of the reciprocals of observations.

    • Formula: H.M.=n1x\text{H.M.} = \frac{n}{\sum \frac{1}{x}}.

  • Median (MdM_d):

    • The middle value dividing the distribution into two equal parts.

    • Continuous distribution formula: Md=L+N2mf×cM_d = L + \frac{\frac{N}{2} - m}{f} \times c.

  • Mode (M0M_0):

    • The most frequent value.

    • Formula: M0=L+(f1f02f1f0f2)×cM_0 = L + \left( \frac{f_1 - f_0}{2f_1 - f_0 - f_2} \right) \times c.

    • Empirical Relationship: Mode=3Median2Mean\text{Mode} = 3\text{Median} - 2\text{Mean}.

Measures of Dispersion, Skewness, and Kurtosis

  • Absolute Measures:

    • Range: LSL - S.

    • Quartile Deviation (Q.D.): Q3Q12\frac{Q_3 - Q_1}{2}.

    • Mean Deviation (M.D.): fDN\frac{\sum f|D|}{N}.

    • Standard Deviation (S.D. or σ\sigma): σ=x2n(xn)2\sigma = \sqrt{\frac{\sum x^2}{n} - \left(\frac{\sum x}{n}\right)^2}.

  • Relative Measures:

    • Coefficient of Variation (C.V.): σxˉ×100\frac{\sigma}{\bar{x}} \times 100. Lower C.V. means higher consistency.

  • Moments:

    • Arithmetic mean of various powers of deviations from the actual mean (μr\mu_r).

    • μ1=0\mu_1 = 0, μ2=Variance\mu_2 = \text{Variance}.

  • Skewness:

    • Lack of symmetry.

    • Symmetrical: Mean=Median=Mode\text{Mean} = \text{Median} = \text{Mode}.

    • Positive: Mean>Median>Mode\text{Mean} > \text{Median} > \text{Mode}.

    • Negative: Mode>Median>Mean\text{Mode} > \text{Median} > \text{Mean}.

  • Kurtosis:

    • Measures peakedness.

    • Mesokurtic: β2=3\beta_2 = 3.

    • Leptokurtic: β2>3\beta_2 > 3.

    • Platykurtic: β2<3\beta_2 < 3.

Correlation and Regression

  • Correlation Coefficient (rr):

    • Varies between 1-1 and +1+1.

    • Karl Pearson’s Formula: r=Cov(x,y)σxσyr = \frac{\text{Cov}(x,y)}{\sigma_x \sigma_y}.

    • Spearman’s Rank Correlation: rs=16D2n(n21)r_s = 1 - \frac{6\sum D^2}{n(n^2-1)}.

  • Regression Analysis:

    • Predicting a dependent variable (YY) based on an independent variable (XX).

    • Regression Line of YY on XX: YYˉ=byx(XXˉ)Y - \bar{Y} = b_{yx}(X - \bar{X}), where byx=rσyσxb_{yx} = r \frac{\sigma_y}{\sigma_x}.

    • Regression Line of XX on YY: XXˉ=bxy(YYˉ)X - \bar{X} = b_{xy}(Y - \bar{Y}), where bxy=rσxσyb_{xy} = r \frac{\sigma_x}{\sigma_y}.

    • Geometric Property: r=byx×bxyr = \sqrt{b_{yx} \times b_{xy}}.

Index Numbers

  • Classification:

    • Price Index: Measures change in price level.

    • Quantity Index: Measures volume of production/consumption.

  • Weighted Aggregate Indices:

    • Laspeyre’s: Uses base year weights (p0q0p_0q_0). P01L=p1q0p0q0×100P_{01}^L = \frac{\sum p_1q_0}{\sum p_0q_0} \times 100.

    • Paasche’s: Uses current year weights (q1q_1). P01P=p1q1p0q1×100P_{01}^P = \frac{\sum p_1q_1}{\sum p_0q_1} \times 100.

    • Fisher’s Ideal Index: Geometric mean of Laspeyre and Paasche. P01F=L×P×100P_{01}^F = \sqrt{L \times P} \times 100.

  • Tests of Consistency:

    • Time Reversal Test: P01×P10=1P_{01} \times P_{10} = 1.

    • Factor Reversal Test: P01×Q01=Value RatioP_{01} \times Q_{01} = \text{Value Ratio}.

    • Note: Fisher’s Ideal Index satisfies both tests.

  • Consumer Price Index (Cost of Living):

    • Aggregate Expenditure Method: Identical to Laspeyre’s.

    • Family Budget Method: PWW\frac{\sum PW}{\sum W}, where P=Price RelativeP = \text{Price Relative}.