Fundamentals of Biostatistics - Unit 1 Notes

Origin and Basic Meanings of Statistics

  • The word statistics is derived from various European languages, each sharing a core meaning related to a "political state":

    • Latin word: STATUS.

    • Italian word: STATISTA.

    • German word: STATISTIK.

  • In common parlance, the term "Statistics" is often used synonymously with "data."

  • The fundamental purpose of data in this context extends beyond simple collection; it encompasses a comprehensive methodology for handling, analyzing, and drawing valid inferences (interpretation) from numerical information.

Dual Senses of Statistics: Plural and Singular

  • The term "statistics" is generally employed in two distinct senses: the plural sense and the singular sense.

  • Plural Sense: This refers to statistics as numerical data or statistical data itself. It constitutes the raw facts and figures collected for study.

  • Singular Sense: This refers to statistics as a science or a branch of knowledge. It involves the techniques or methods used for collecting, classifying, presenting, analyzing, and interpreting data.

Definitions and Features in the Plural Sense

  • Definition by Bowley: "Statistics are numerical statements of facts in any department of enquiry placed in relation to each other."

  • Definition by Yule and Kendall: "By Statistics we mean quantitative data affected to a market extent by multiplicity of causes."

  • Key Features of Statistics in the Plural Sense:

    • Aggregate of Facts: Statistics do not refer to a single isolated figure; they represent a collection of many facts.

    • Numerically Expressed: Qualitative observations must be converted into numerical form to be considered statistics.

    • Affected by Multiplicity of Causes: Statistical data are influenced by numerous factors rather than a single cause.

    • Standard of Accuracy: Data must be enumerated or estimated according to a reasonable and predetermined standard of accuracy.

    • Systematic Collection: Data should be gathered in a planned and orderly manner.

    • Predetermined Purpose: Statistics are collected with a specific goal or objective in mind.

    • Relational Placement: Data points are placed in relation to one another to facilitate comparison.

Definitions and Features in the Singular Sense (Statistics as a Science)

  • Definition by Croxton and Cowdon: "Statistics may be defined as a collection, presentation, analysis and interpretation of numerical data."

  • Definition by Lovitt: "Statistics is the science which deals with the collection, classification and tabulation of numerical facts as a basis for the explanation, description and comparison of phenomena."

  • Key Features of Statistics as a Science:

    • Collection of data: The initial gathering of facts.

    • Organization of data: Arranging data into logical patterns or classifications.

    • Presentation of data: Using tables, charts, or graphs for clarity.

    • Analysis of data: Applying statistical tools and tests to examine the data.

    • Interpretation of data: Drawing conclusions and explaining findings.

Applications and Importance of Biostatistics

  • Broad Applications:

    • Clinical medicine and Medical science.

    • Community medicine and Public health.

    • Pharmacy and Nursing.

    • Genetical statistics.

    • Agriculture.

    • Demography.

    • Psychology.

  • Major Functions and Importance:

    • Factual Presentation: Biostatistics presents complex facts through figures. For example, comparing the population of Nepal in 2017 to 2007 is more logical when stated as a growth from 11 billion to 1.21.2 billion rather than just stating it "increased."

    • Simplification of Complexity: It uses tools like averages, diagrams, and graphs to make complex, non-comprehensive data readily understandable.

    • Facilitation of Comparison: Tools such as ratios, percentages, and coefficients allow researchers to draw valid conclusions. For instance, stating a drug's recovery rate is 33 days is only useful when compared against a standard drug's rate.

    • Forecasting Future Trends: Time series analysis allows hospital administrators and health officers to predict variables like the number of outpatients expected next year.

    • Policy Formulation: Vital statistics like birth rates, maternal mortality rates, and child death rates are used to create national health policies and family planning programs.

    • Understanding Relationships: Statistical tools like correlation, regression, chi-square tests, relative risk, and odds ratios determine links between different characteristics.

    • Hypothesis Testing: Inferential statistics (e.g., ZZ test, tt test, FF test, Chi-square test) are essential for testing scientific hypotheses.

    • Quality Control: Use of graphical control methods such as PP charts, CC charts, SS charts, and UU charts ensures consistency and lack of error in laboratory and production activities.

    • Handling Uncertainty: Probability tools help determine the likelihood of events, such as the chance of a wrong diagnosis or the risk of heart disease in smokers.

    • Managing Biological Variation: Statistics provide tools like Range and Standard Deviation (SDSD) to study variations in human characteristics like height, weight, blood pressure, and urea levels.

Data Collection and Classification in Pharmacy

  • Data Collection: The process of gathering info for research, which in pharmacy includes drug efficacy, patient responses, or epidemiological studies.

  • Best Practices:

    • Units of measurement must be clearly defined.

    • Records should be correct, complete, clear, concise, and easy to comprehend.

  • Types of Data:

    • Primary Data: Collected directly from experiments or surveys. Example: Measuring the blood pressure of patients on a new antihypertensive drug.

    • Secondary Data: Collected from existing sources like journals, databases, or hospital records. Example: Meta-analysis of previous clinical trials.

  • Classification of Data: The process of grouping data based on shared characteristics to make it easier to interpret.

    • Qualitative (Categorical): Descriptive and non-numerical (e.g., Gender, Drug Type, Side Effect Type).

    • Quantitative (Numerical): Measurable in numbers (e.g., Blood pressure, drug concentration).

  • Methods of Classification:

    • Discrete Classification: Data in distinct categories. Example: Grouping drugs as antibiotics, analgesics, or antihypertensives.

    • Continuous Classification: Data grouped into class intervals or ranges. Example: Blood glucose levels categorized as 7010070\text{--}100, 101120101\text{--}120, and 121140mg/dL121\text{--}140\,mg/dL.

Presentation of Data: Tabulation

  • Principles of Presentation:

    • Should arouse reader interest and be concise without losing detail.

    • Must be in a simple form for quick impressions and suggest solutions to problems.

  • Methods of Presentation: Tabulation and Graphical Presentation.

  • Tabulation of Data: The arrangement of data in rows and columns.

    • Simple/One-way Table: Contains data for a single characteristic (e.g., a Frequency distribution).

    • Two-way Table: Contains data based on two characteristics.

    • Manifold Table: Contains data for three or more characteristics.

  • Standard Principles for Tables:

    • Must be numbered and contain a brief, top-aligned title.

    • Rows and columns must be clearly defined with specified units of measurement.

    • The number of class intervals should be balanced (not too many or too few) and must not overlap.

    • Standard codes and symbols should be explained in footnotes.

    • Sources for secondary data must be cited at the bottom.

Constructing a Frequency Distribution

  • Step-by-Step Procedure:

    1. Sort Data: Arrange in ascending order.

    2. Calculate Range: Identify the minimum and maximum values.

    3. Decide Intervals: Choose a number of intervals that summarizes without being too broad or too detailed.

    4. Determine Intervals: Create non-overlapping, all-encompassing intervals.

    5. Tally: Count observations falling into each interval.

  • Example Construction:

    • Given 2020 observations ranging from 1.21.2 to 9.89.8.

    • Intervals chosen: 55 intervals with a width of 22 (020\text{--}2, 242\text{--}4, 464\text{--}6, 686\text{--}8, 8108\text{--}10).

Graphical Presentation of Data

  • Introduction: Graphics complement tables to summarize data attractively. They should be self-explanatory with clear captions and indices.

  • Bar Diagram: Mutually exclusive discrete data. The length of the bar indicates frequency. Categories on one axis, frequency on the other.

    • Simple Bar Diagram: Single variable for different categories.

    • Multiple/Compound Bar Chart: Each observation has multiple values (e.g., Percentage of males vs. females across different countries).

    • Component/Proportional Bar Chart: Subdivides a single bar into sections representing the relative proportion of parts within a total (e.g., energy intake divided into Protein, Fat, and Carbohydrate).

  • Histogram: Similar to bar charts but bars are adherent (no gaps), used for continuous data and class frequency tables.

  • Frequency Polygon: Created by connecting the midpoints of the tops of histogram rectangles. A smoothed frequency polygon is a Frequency Curve (Normal Curve).

  • Pie Diagram: A circle representing 100%100\% of frequency, divided into segments representing proportional compositions.

  • Stem-and-Leaf Plot: Data visualization based on place value.

    • Stem: Represents leading digits.

    • Leaf: Represents the last digit.

    • Key: Essential for reading (e.g., 25=252 \mid 5 = 25).

  • Whiskers Box-Plot: Displays a five-number summary: Minimum score, First Quartile (Q1Q_1), Median (Q2Q_2), Third Quartile (Q3Q_3), and Maximum score.

  • General Graph Principles:

    • Variables usually on the X-axis; frequencies on the Y-axis.

    • Lines should not be extrapolated beyond the range of actual values.

    • Graphs are aids for visualization, not final statistical analysis.

Statistical Data Series and Class Intervals

  • Individual Series: Data points given independently for each individual. No frequency. Example: Weight of 55 children: 3,3.5,2.5,4,23, 3.5, 2.5, 4, 2.

  • Discrete Series (Ungrouped): Tabular frequency distribution for discrete variables (XX and ff).

  • Continuous Series (Grouped): Numerical variables taking any value within a range.

    • Exclusive (Overlapping): Upper limit of one class is the lower limit of the next (e.g., 102010\text{--}20, 203020\text{--}30).

    • Inclusive (Non-overlapping): Both limits are included in the interval; overlapping is avoided (e.g., 101910\text{--}19, 202920\text{--}29).

    • Open-end Classes: A limit is missing at the first or last interval.

Measures of Central Tendency

  • Overview: These provide a "typical" value representing the center of the distribution where about 50%50\% of observations lie above and 50%50\% below.

  • Arithmetic Mean (xˉ\bar{x}): The mathematical average.

    • Individual Series: xˉ=Xn\bar{x} = \frac{\sum X}{n}.

    • Discrete Series: xˉ=fXN\bar{x} = \frac{\sum fX}{N}.

    • Continuous Series: xˉ=fxN\bar{x} = \frac{\sum fx}{N} (where xx is the midpoint).

    • Assumed Mean Method: xˉ=A+fdN\bar{x} = A + \frac{\sum fd}{N} (where AA is assumed mean and d=xAd = x - A).

    • Advantages: Rigidly defined, based on all observations, suitable for further math analysis.

    • Disadvantages: Highly affected by extreme values, cannot be computed for open-ended classes or qualitative data.

  • Median (MdMd): The middle value in an ordered set; a positional average.

    • Individual Series: If nn is odd, n+12\frac{n+1}{2}th item; if even, average of the two middle terms.

    • Discrete Series: Value corresponding to cumulative frequency (c.f.c.f.) equal to or just greater than N+12\frac{N+1}{2}.

    • Continuous Series: Md=L+N2c.f.f×iMd = L + \frac{\frac{N}{2} - c.f.}{f} \times i.

    • Advantages: Robust (not affected by extreme values), appropriate for qualitative data (intelligence, honesty) and open-ended classes.

  • Mode (MoMo): The value occuring most frequently.

    • Analysis: Determined by inspection. If multiple, it is "bimodal" or "multimodal."

    • Continuous Formula: Mo=L+f1f02f1f0f2×iMo = L + \frac{f_1 - f_0}{2f_1 - f_0 - f_2} \times i.

    • Relationship: For moderately skewed distributions: Mode=3Median2MeanMode = 3\,Median - 2\,Mean.

Partition Values: Quartiles, Deciles, and Percentiles

  • Quartiles (Qi,i=1,2,3Q_i, i = 1, 2, 3): Divide data into 4 equal parts.

    • Q1Q_1 (Lower): 25%25\% below it.

    • Q2Q_2: Median (50%50\% below it).

    • Q3Q_3 (Upper): 75%75\% below it.

  • Deciles (Dj,j=19D_j, j = 1 \dots 9): Divide data into 10 equal parts.

  • Percentiles (Pk,k=199P_k, k = 1 \dots 99): Divide data into 100 equal parts.

  • Advantages: Useful for dispersion and skewness; Easy to determine for individual/discrete series.

  • Disadvantages: Not based on all observations; Requires data rearrangement.

Measures of Dispersion

  • Definition: The state of data being spread or the extent to which numerical data vary around an average. It indicates homogeneity vs. heterogeneity.

  • Absolute Measures (Same units as original data):

    • Range: Difference between Largest (LL) and Smallest (SS) value (Range=LSRange = L - S).

    • Mean Deviation (MDMD): Arithmetic mean of absolute deviations from a central tendency (MD=XxˉNMD = \frac{\sum |X - \bar{x}|}{N}).

    • Variance (σ2\sigma^2): The average of squared differences from the mean.

    • Standard Deviation (σ\sigma): The positive square root of variance (σ=σ2\sigma = \sqrt{\sigma^2}).

    • Quartile Deviation (QDQD): Half the distance between the third and first quartile (QD=Q3Q12QD = \frac{Q_3 - Q_1}{2}).

  • Relative Measures (Unitless ratios, often percentages):

    • Coefficient of Range: LSL+S\frac{L - S}{L + S}.

    • Coefficient of Variation (CVCV): σxˉ×100\frac{\sigma}{\bar{x}} \times 100.

    • Coefficient of SD: σxˉ\frac{\sigma}{\bar{x}}.

    • Coefficient of QD: Q3Q1Q3+Q1\frac{Q_3 - Q_1}{Q_3 + Q_1}.

Standard Deviation and Variance Details

  • Standard Deviation (SDSD): Often called the "root mean square deviation." It is the most robust measure for inferential statistics.

    • Population SD: σ=(Xμ)2N\sigma = \sqrt{\frac{\sum(X - \mu)^2}{N}}.

    • Sample SD: S=(Xxˉ)2n1S = \sqrt{\frac{\sum(X - \bar{x})^2}{n - 1}}.

  • Merits of SD: Summarizes deviation in one figure; differentiates between real differences vs. chance; helps calculate standard error.

  • Demerits of SD: Lengthy calculations; gives extra weight to extreme values.

Standard Error (SE)

  • Overview: Measures the dispersion of sample means around the true population mean. It indicates the precision of a sample estimate.

  • Standard Error of the Mean: SExˉ=snSE_{\bar{x}} = \frac{s}{\sqrt{n}}.

  • Standard Error of a Proportion: SEp=pqnSE_p = \sqrt{\frac{pq}{n}}.

  • Comparison: Standard deviation describes variability within one sample, whereas standard error describes variability of means across multiple samples.

Scales of Measurement

  • Nominal Scale: Mutual exclusive categories without magnitude. Arithmetics restricted to counting (freq, percentage, chi-square). Example: Gender, Blood group, Urban/Rural.

  • Ordinal Scale: Categories with a meaningful natural order/rank but unequal intervals. Appropriate central tendency: Median. Example: Pain (Mild, Moderate, Severe), Cancer stages (I, II, III, IV).

  • Interval Scale: Continuous scale with defined distances between measurements. No absolute zero (zero is arbitrary). Addition/subtraction allowed. Example: Temperature (C^{\circ}C or F^{\circ}F).

  • Ratio Scale: Has all properties of interval scale plus a true absolute zero. All mathematical and statistical operations are usable. Example: Blood pressure, serum cholesterol, height, weight, pulse rate.