Unit I Descriptive Statistics (Univariate): Fundamentals and Data Methodology and Measurement Scales and Errors

Etymology and Historical Background of Statistics

  • The word Statistics is derived from the Latin term "Status" or the Italian term "Statista."

  • Both of these root words translate to "Political State" or "a Government."

  • Historically, the application of statistics was restricted to the needs of statecraft. Rulers and kings required systematic data regarding land, agriculture, commerce, and the population of their territories.

  • These data points were used to assess military potential, calculate state wealth, determine taxation levels, and manage other administrative functions of government.

Statistics in the Indian Context

  • Dr. Prasanta Chandra Mahalanobis is celebrated as the Father of Statistics in India.

  • National Statistics Day is observed on June 29th29^{th} in honor of his birthday.

  • Other distinguished Indian statisticians credited with significant contributions include:

    • Dr. C.R. Rao

    • Dr. P.C. Desai

    • Dr. P.V. Sukhatame

Definition and Scope of Statistics

  • Statistics is defined as a branch of science that encompasses the following processes regarding data:

    • Collection

    • Classification

    • Tabulation

    • Representation

    • Analysis

    • Interpretation

Main Divisions of Statistics

  • Mathematical, Theoretical, or Pure Statistics: This division includes the Theory of Probability, Statistical Inference (also known as Inductive Statistics), and Descriptive Statistics.

  • Applied Statistics: This division includes practical applications such as Statistical Quality Control, Vital Statistics, Design of Experiment, Index Numbers, Time Series, and Operations Research.

Descriptive Statistics versus Statistical Inference

  • Descriptive Statistics:

    • This branch focuses on methods that are limited to describing the specific sample being studied.

    • It does not involve making generalizations or conclusions about the larger population from which the sample was drawn.

  • Statistical Inference or Inductive Statistics:

    • This branch involves methods capable of making generalizations about a population based on observations made from a sample.

    • A sample is defined as a part of a population selected to obtain information about that population.

Concepts of Statistical Population

  • A Statistical Population (or simply Population) refers to the entire collection of individuals or objects that are the subject of a specific study.

  • A population can consist of living things, non-living things, or even a set of results generated from an experiment.

  • Examples of Populations:

    • A study on bulb quality: The entire production of bulbs from a factory during a specific period constitutes the population.

    • A study on primary education levels: All children enrolled in primary schools form the population.

Classifications of Population

  • By Count:

    • Finite Population: A population where the individuals can be counted (also known as a countable population).

    • Infinite Population: A population where it is impossible to count every unit contained within it (uncountable).

  • By Nature:

    • Real Population: A population consisting of concrete, tangible objects.

    • Hypothetical Population: A population consisting of imaginary or theoretical objects, such as the infinite series of outcomes from repeatedly throwing a dice or a coin.

  • By Study Selection:

    • Target Population: The entire group about which researchers want to gather information (e.g., all families living in rented houses in a large city).

    • Sampled Population (Studied Population): The specific population from which the sample is actually drawn. This may differ from the target population if parts of the group are ignored to reduce costs or if certain individuals refuse to participate.

Types of Data: Attributes and Variables

  • Attribute (Qualitative Characteristic):

    • A characteristic that belongs to an individual which cannot be measured numerically but can only be identified by its presence or absence.

    • Examples include blindness, honesty, and beauty.

    • Populations are generally divided into two classes based on attributes: those possessing the attribute and those not possessing it.

  • Variable (Quantitative Characteristic):

    • A characteristic that can be measured and assigned different numerical values.

    • Examples include height, weight, and the number of rooms in a building.

    • Discrete Variable: A variable that can only take finite or denumerable values, such as the number of students in a class or the number of rooms in a house.

    • Continuous Variable: A variable that can assume any value within a specific interval or set of intervals, such as height and weight.

Measurement in Science

  • Measurement is the process of assigning numbers to various attributes of people, objects, or concepts.

  • Examples of measurement include:

    • Assigning a number of inches to represent a person's height.

    • Assigning a number equal to the number of residents to represent the size of a city.

The Four Scales of Measurement

  • Nominal Measurement Scale:

    • This is a crude form of measurement where numbers serve only as labels to categorize observations into mutually exclusive and exhaustive groups.

    • There is no relative ordering; numbers are arbitrary (e.g., Male = 00, Female = 11).

    • Examples: Blood Type (A, B, AB, O), Eye Color (Blue, Brown, Green), Marital Status, and Football Jersey Numbers (Player #1010 is not "twice" Player #55).

    • Statistical uses: Classification, calculating frequencies, percentages, mode, and contingency tables (e.g., Chi-Square tests).

  • Ordinal Measurement Scale:

    • This scale indicates the rank-ordering of participants but provides only relative positions.

    • Differences between ranks are not consistent or measurable (e.g., the time gap between 1st1^{st} and 2nd2^{nd} place might be 0.010.01 seconds, while the gap between 2nd2^{nd} and 3nd3^{nd} is 55 seconds).

    • Examples: Likert Scales (Strongly Disagree to Strongly Agree), Education Level (High School to Ph.D.), Socioeconomic Status, and Race Finish Positions.

    • Statistical uses: Non-parametric methods, Spearman rank correlation, median, percentile ranks, and ordinal regression.

  • Interval Measurement Scale:

    • This scale provides quantitative information with equal distances between units across all levels of the scale.

    • It lacks an absolute (true) zero point; zero is arbitrary and does not represent the total absence of the property.

    • Addition and subtraction are valid, but ratios are not (e.g., 60F60^{\circ}F is not twice as hot as 30F30^{\circ}F).

    • Examples: Temperature in Fahrenheit and Celsius, and IQ scores.

    • Statistical uses: Mean, Standard Deviation, Pearson Correlation, t-tests, ANOVA, and linear regression.

  • Ratio Measurement Scale:

    • This scale possesses equal intervals and a fixed, meaningful zero point representing the total absence of the property.

    • All arithmetic operations, including division and ratios, are valid (e.g., "twice as much" is mathematically accurate).

    • Examples: Time (ten hours is twice as long as five hours), Height, Weight (0kg0\,kg means no weight), Income (Rs.100,000Rs.\,100,000 is twice Rs.50,000Rs.\,50,000), and the Kelvin Temperature Scale (0K0\,K is absolute zero).

    • Statistical uses: All parametric methods, geometric mean, coefficient of variation, and econometric modeling.

Comparison Summary of Measurement Scales

  • Nominal: No order, no equal intervals, no true zero point. Operations permitted: == and \neq. Central tendency: Mode. Examples: Hair color, Zip code.

  • Ordinal: Order matters, no equal intervals, no true zero point. Operations permitted: =,,<,>=, \neq, <, >. Central tendency: Median and Mode. Example: Customer satisfaction rating.

  • Interval: Order matters, equal intervals exist, no true zero point. Operations permitted: +,+, -. Central tendency: Mean, Median, Mode. Example: Temperature in Celsius (C^{\circ}C).

  • Ratio: Order matters, equal intervals exist, true zero point exists. Operations permitted: +,,×,÷+, -, \times, \div. Central tendency: Mean, Median, Mode. Examples: Monthly salary, Weight.

Sources of Data: Primary and Secondary

  • Primary Data:

    • Original data collected specifically for the current purpose by the investigator for the first time.

    • Example: Population census reports are primary data for the census organization that collects, compiles, and publishes them.

    • Methods of collection: Personal Investigation, through investigators, via telephone, or through emails and Google Forms.

  • Secondary Data:

    • Second-hand information already collected by another organization for a different purpose and made available for the current study.

    • These data are not "pure" as they have undergone treatment or processing at least once.

    • Example: The Economics Survey of India is secondary data because it is compiled from multiple organizations like the Bureau of Statistics, Board of Revenue, and banks.

    • Methods of collection: Published or unpublished theses, Government reports, journals, magazines, weeklies, and dailies.

Instruments for Data Collection: Questionnaire and Schedule

  • Questionnaire:

    • A written or printed tool consisting of a list of questions with provided choices of answers.

    • Usually sent to respondents via post or mail for them to complete independently.

  • Schedule:

    • A formal list of questions filled in by officially appointed investigators.

    • Investigators personally visit respondents, ask the questions, and record the answers in the provided space.

  • Comparison:

    • Economy: Questionnaires are economical; schedules are expensive.

    • Response Rate: Non-response is higher in questionnaires; schedules ensure higher response rates.

    • Time: Questionnaires are a prolonged process; schedules provide information in a timely manner.

    • Literacy: Questionnaires require literate respondents; schedules can be used for both literate and illiterate populations.

    • Accuracy: Questionnaires risk getting incomplete or wrong information; schedules minimize this risk through investigator oversight.

The Process of Editing Data

  • Editing involves examining collected data to discover errors or mistakes before presentation.

  • Editing Primary Data: Focused on four pillars:

    1. Consistency:Comparing questions designed to be mutually confirmatory.

    2. Uniformity: Ensuring all units of information are the same (e.g., all currency in Rs.Rs., all heights in cmcm).

    3. Completeness: Handling incomplete questionnaires by either completing them or excluding them.

    4. Accuracy: A difficult task requiring experience to verify the correctness of responses.

  • Editing Secondary Data: Requires checking for:

    • Suitability for the current study.

    • Adequacy to fulfill research requirements.

    • Reliability of the source.

    • Accuracy of the data, including the time and conditions under which it was originally collected.

    • Comparison with other sources for validation.

Quantitative Errors in Measurement

  • Absolute Error: The magnitude of the difference between the true (real) value and the individual measured (estimated) value.

    • Absolute Error=True ValueEstimated Value\text{Absolute Error} = | \text{True Value} - \text{Estimated Value} |

  • Relative Error: The ratio of the absolute error to the true value. It is a unit-free measure used for comparisons.

    • Relative Error=Absolute ErrorTrue Value\text{Relative Error} = \frac{\text{Absolute Error}}{\text{True Value}}

  • Percentage Error: The relative error expressed as a percentage.

    • Percentage Error=Relative Error×100%\text{Percentage Error} = \text{Relative Error} \times 100\,\%

Practical Example of Measurement Errors

  • Given a measurement of 25.54mm25.54\,mm and a true value of 26.00mm26.00\,mm:

    • Absolute Error: 26.0025.54=0.46mm| 26.00 - 25.54 | = 0.46\,mm

    • Relative Error: 0.4626.00=0.017\frac{0.46}{26.00} = 0.017

    • Percentage Error: 0.017×100%=1.77%0.017 \times 100\,\% = 1.77\,\%