Fundamentals of Biostatistics - Unit 1 Notes
Origin and Basic Meanings of Statistics
The word statistics is derived from various European languages, each sharing a core meaning related to a "political state":
Latin word: STATUS.
Italian word: STATISTA.
German word: STATISTIK.
In common parlance, the term "Statistics" is often used synonymously with "data."
The fundamental purpose of data in this context extends beyond simple collection; it encompasses a comprehensive methodology for handling, analyzing, and drawing valid inferences (interpretation) from numerical information.
Dual Senses of Statistics: Plural and Singular
The term "statistics" is generally employed in two distinct senses: the plural sense and the singular sense.
Plural Sense: This refers to statistics as numerical data or statistical data itself. It constitutes the raw facts and figures collected for study.
Singular Sense: This refers to statistics as a science or a branch of knowledge. It involves the techniques or methods used for collecting, classifying, presenting, analyzing, and interpreting data.
Definitions and Features in the Plural Sense
Definition by Bowley: "Statistics are numerical statements of facts in any department of enquiry placed in relation to each other."
Definition by Yule and Kendall: "By Statistics we mean quantitative data affected to a market extent by multiplicity of causes."
Key Features of Statistics in the Plural Sense:
Aggregate of Facts: Statistics do not refer to a single isolated figure; they represent a collection of many facts.
Numerically Expressed: Qualitative observations must be converted into numerical form to be considered statistics.
Affected by Multiplicity of Causes: Statistical data are influenced by numerous factors rather than a single cause.
Standard of Accuracy: Data must be enumerated or estimated according to a reasonable and predetermined standard of accuracy.
Systematic Collection: Data should be gathered in a planned and orderly manner.
Predetermined Purpose: Statistics are collected with a specific goal or objective in mind.
Relational Placement: Data points are placed in relation to one another to facilitate comparison.
Definitions and Features in the Singular Sense (Statistics as a Science)
Definition by Croxton and Cowdon: "Statistics may be defined as a collection, presentation, analysis and interpretation of numerical data."
Definition by Lovitt: "Statistics is the science which deals with the collection, classification and tabulation of numerical facts as a basis for the explanation, description and comparison of phenomena."
Key Features of Statistics as a Science:
Collection of data: The initial gathering of facts.
Organization of data: Arranging data into logical patterns or classifications.
Presentation of data: Using tables, charts, or graphs for clarity.
Analysis of data: Applying statistical tools and tests to examine the data.
Interpretation of data: Drawing conclusions and explaining findings.
Applications and Importance of Biostatistics
Broad Applications:
Clinical medicine and Medical science.
Community medicine and Public health.
Pharmacy and Nursing.
Genetical statistics.
Agriculture.
Demography.
Psychology.
Major Functions and Importance:
Factual Presentation: Biostatistics presents complex facts through figures. For example, comparing the population of Nepal in 2017 to 2007 is more logical when stated as a growth from billion to billion rather than just stating it "increased."
Simplification of Complexity: It uses tools like averages, diagrams, and graphs to make complex, non-comprehensive data readily understandable.
Facilitation of Comparison: Tools such as ratios, percentages, and coefficients allow researchers to draw valid conclusions. For instance, stating a drug's recovery rate is days is only useful when compared against a standard drug's rate.
Forecasting Future Trends: Time series analysis allows hospital administrators and health officers to predict variables like the number of outpatients expected next year.
Policy Formulation: Vital statistics like birth rates, maternal mortality rates, and child death rates are used to create national health policies and family planning programs.
Understanding Relationships: Statistical tools like correlation, regression, chi-square tests, relative risk, and odds ratios determine links between different characteristics.
Hypothesis Testing: Inferential statistics (e.g., test, test, test, Chi-square test) are essential for testing scientific hypotheses.
Quality Control: Use of graphical control methods such as charts, charts, charts, and charts ensures consistency and lack of error in laboratory and production activities.
Handling Uncertainty: Probability tools help determine the likelihood of events, such as the chance of a wrong diagnosis or the risk of heart disease in smokers.
Managing Biological Variation: Statistics provide tools like Range and Standard Deviation () to study variations in human characteristics like height, weight, blood pressure, and urea levels.
Data Collection and Classification in Pharmacy
Data Collection: The process of gathering info for research, which in pharmacy includes drug efficacy, patient responses, or epidemiological studies.
Best Practices:
Units of measurement must be clearly defined.
Records should be correct, complete, clear, concise, and easy to comprehend.
Types of Data:
Primary Data: Collected directly from experiments or surveys. Example: Measuring the blood pressure of patients on a new antihypertensive drug.
Secondary Data: Collected from existing sources like journals, databases, or hospital records. Example: Meta-analysis of previous clinical trials.
Classification of Data: The process of grouping data based on shared characteristics to make it easier to interpret.
Qualitative (Categorical): Descriptive and non-numerical (e.g., Gender, Drug Type, Side Effect Type).
Quantitative (Numerical): Measurable in numbers (e.g., Blood pressure, drug concentration).
Methods of Classification:
Discrete Classification: Data in distinct categories. Example: Grouping drugs as antibiotics, analgesics, or antihypertensives.
Continuous Classification: Data grouped into class intervals or ranges. Example: Blood glucose levels categorized as , , and .
Presentation of Data: Tabulation
Principles of Presentation:
Should arouse reader interest and be concise without losing detail.
Must be in a simple form for quick impressions and suggest solutions to problems.
Methods of Presentation: Tabulation and Graphical Presentation.
Tabulation of Data: The arrangement of data in rows and columns.
Simple/One-way Table: Contains data for a single characteristic (e.g., a Frequency distribution).
Two-way Table: Contains data based on two characteristics.
Manifold Table: Contains data for three or more characteristics.
Standard Principles for Tables:
Must be numbered and contain a brief, top-aligned title.
Rows and columns must be clearly defined with specified units of measurement.
The number of class intervals should be balanced (not too many or too few) and must not overlap.
Standard codes and symbols should be explained in footnotes.
Sources for secondary data must be cited at the bottom.
Constructing a Frequency Distribution
Step-by-Step Procedure:
Sort Data: Arrange in ascending order.
Calculate Range: Identify the minimum and maximum values.
Decide Intervals: Choose a number of intervals that summarizes without being too broad or too detailed.
Determine Intervals: Create non-overlapping, all-encompassing intervals.
Tally: Count observations falling into each interval.
Example Construction:
Given observations ranging from to .
Intervals chosen: intervals with a width of (, , , , ).
Graphical Presentation of Data
Introduction: Graphics complement tables to summarize data attractively. They should be self-explanatory with clear captions and indices.
Bar Diagram: Mutually exclusive discrete data. The length of the bar indicates frequency. Categories on one axis, frequency on the other.
Simple Bar Diagram: Single variable for different categories.
Multiple/Compound Bar Chart: Each observation has multiple values (e.g., Percentage of males vs. females across different countries).
Component/Proportional Bar Chart: Subdivides a single bar into sections representing the relative proportion of parts within a total (e.g., energy intake divided into Protein, Fat, and Carbohydrate).
Histogram: Similar to bar charts but bars are adherent (no gaps), used for continuous data and class frequency tables.
Frequency Polygon: Created by connecting the midpoints of the tops of histogram rectangles. A smoothed frequency polygon is a Frequency Curve (Normal Curve).
Pie Diagram: A circle representing of frequency, divided into segments representing proportional compositions.
Stem-and-Leaf Plot: Data visualization based on place value.
Stem: Represents leading digits.
Leaf: Represents the last digit.
Key: Essential for reading (e.g., ).
Whiskers Box-Plot: Displays a five-number summary: Minimum score, First Quartile (), Median (), Third Quartile (), and Maximum score.
General Graph Principles:
Variables usually on the X-axis; frequencies on the Y-axis.
Lines should not be extrapolated beyond the range of actual values.
Graphs are aids for visualization, not final statistical analysis.
Statistical Data Series and Class Intervals
Individual Series: Data points given independently for each individual. No frequency. Example: Weight of children: .
Discrete Series (Ungrouped): Tabular frequency distribution for discrete variables ( and ).
Continuous Series (Grouped): Numerical variables taking any value within a range.
Exclusive (Overlapping): Upper limit of one class is the lower limit of the next (e.g., , ).
Inclusive (Non-overlapping): Both limits are included in the interval; overlapping is avoided (e.g., , ).
Open-end Classes: A limit is missing at the first or last interval.
Measures of Central Tendency
Overview: These provide a "typical" value representing the center of the distribution where about of observations lie above and below.
Arithmetic Mean (): The mathematical average.
Individual Series: .
Discrete Series: .
Continuous Series: (where is the midpoint).
Assumed Mean Method: (where is assumed mean and ).
Advantages: Rigidly defined, based on all observations, suitable for further math analysis.
Disadvantages: Highly affected by extreme values, cannot be computed for open-ended classes or qualitative data.
Median (): The middle value in an ordered set; a positional average.
Individual Series: If is odd, th item; if even, average of the two middle terms.
Discrete Series: Value corresponding to cumulative frequency () equal to or just greater than .
Continuous Series: .
Advantages: Robust (not affected by extreme values), appropriate for qualitative data (intelligence, honesty) and open-ended classes.
Mode (): The value occuring most frequently.
Analysis: Determined by inspection. If multiple, it is "bimodal" or "multimodal."
Continuous Formula: .
Relationship: For moderately skewed distributions: .
Partition Values: Quartiles, Deciles, and Percentiles
Quartiles (): Divide data into 4 equal parts.
(Lower): below it.
: Median ( below it).
(Upper): below it.
Deciles (): Divide data into 10 equal parts.
Percentiles (): Divide data into 100 equal parts.
Advantages: Useful for dispersion and skewness; Easy to determine for individual/discrete series.
Disadvantages: Not based on all observations; Requires data rearrangement.
Measures of Dispersion
Definition: The state of data being spread or the extent to which numerical data vary around an average. It indicates homogeneity vs. heterogeneity.
Absolute Measures (Same units as original data):
Range: Difference between Largest () and Smallest () value ().
Mean Deviation (): Arithmetic mean of absolute deviations from a central tendency ().
Variance (): The average of squared differences from the mean.
Standard Deviation (): The positive square root of variance ().
Quartile Deviation (): Half the distance between the third and first quartile ().
Relative Measures (Unitless ratios, often percentages):
Coefficient of Range: .
Coefficient of Variation (): .
Coefficient of SD: .
Coefficient of QD: .
Standard Deviation and Variance Details
Standard Deviation (): Often called the "root mean square deviation." It is the most robust measure for inferential statistics.
Population SD: .
Sample SD: .
Merits of SD: Summarizes deviation in one figure; differentiates between real differences vs. chance; helps calculate standard error.
Demerits of SD: Lengthy calculations; gives extra weight to extreme values.
Standard Error (SE)
Overview: Measures the dispersion of sample means around the true population mean. It indicates the precision of a sample estimate.
Standard Error of the Mean: .
Standard Error of a Proportion: .
Comparison: Standard deviation describes variability within one sample, whereas standard error describes variability of means across multiple samples.
Scales of Measurement
Nominal Scale: Mutual exclusive categories without magnitude. Arithmetics restricted to counting (freq, percentage, chi-square). Example: Gender, Blood group, Urban/Rural.
Ordinal Scale: Categories with a meaningful natural order/rank but unequal intervals. Appropriate central tendency: Median. Example: Pain (Mild, Moderate, Severe), Cancer stages (I, II, III, IV).
Interval Scale: Continuous scale with defined distances between measurements. No absolute zero (zero is arbitrary). Addition/subtraction allowed. Example: Temperature ( or ).
Ratio Scale: Has all properties of interval scale plus a true absolute zero. All mathematical and statistical operations are usable. Example: Blood pressure, serum cholesterol, height, weight, pulse rate.