Biostatistics and Data Analysis in Dental Hygiene
Introduction to Biostatistics and Data Analysis
- Definition of Biostatistics: The application of data analysis and interpretation specifically within the context of healthcare research.
- Role of Technology: Most contemporary computations are performed using dedicated computer programs. Researchers must possess the knowledge to input data correctly and accurately interpret the generated results.
- Scope of Data Analysis:
- Involves the application of various statistical tests to organize, describe, summarize, and analyze gathered data.
- Aimed at answering researchers' specific questions or testing formulated hypotheses.
- Requires critical thinking to explain the meaning and practical application of findings.
- Involves identifying factors that could have influenced results and drawing inferences from sample data to the broader population.
- Applications in Dental Hygiene:
- Demonstrating patient response to dental hygiene therapies.
- Testing clinical products and specific treatment regimens.
- Determining the specific oral health needs of target populations.
- Evaluating oral health treatments, prevention strategies, and educational programs.
- Importance of Research Knowledge for the Dental Hygienist:
- Understanding the epidemiology of diseases.
- Developing and practicing evidence-based therapies.
- Implementing effective public health programs.
- Practicing evidence-based dentistry to ensure clinical decisions are supported by data.
Potential Causes of Invalid Research Results
- Low Sample Size: Utilizing an insufficient number of subjects can undermine the statistical power of the study.
- Duration Issues: Studies that are too short in duration may not capture the true effects of an intervention.
- Tool Errors: Using incorrect or uncalibrated measurement instruments.
- Procedural Errors: Utilizing incorrect procedures during the data collection or intervention phase.
- Statistical Misapplication: Applying incorrect statistical tests to analyze the gathered data.
Categorization and Types of Data
- General Definition: Data refers to the specific information collected by a researcher.
- Qualitative Data:
- Definition: Data that reflects the quality or nature of variables rather than numerical counts.
- Ordering: Can often be rank-ordered (e.g., a list of patient responses regarding what they liked most vs. least).
- Categorical Variables: Variables with no numerical representation (e.g., hair color, ratings of "good" or "bad").
- Dichotomous Variables: Variables that place subjects into exactly two distinct groups (e.g., female/male, yes/no, pass/fail).
- Quantitative Data:
- Definition: Data represented by numbers and expressed as counts, percentages, and means.
- Examples: Pocket depths, number of teeth with sealants.
- Continuous Variables: Variables derived from a large or infinite number of measures. These can be expressed in fractions (e.g., height, weight).
- Discrete Variables: Measurements consisting of distinct and separate units expressed in whole numbers only (e.g., the number of children, DMFT score).
Scales of Measurement
- Nominal Scale:
- Organizes data into mutually exclusive categories.
- Categories have no inherent rank or numerical value.
- Dental Hygiene Example: Patient classification (Pediatric, Adult, Geriatric) or hair color (brown, black, blonde, gray).
- Ordinal Scale:
- Organizes data into mutually exclusive categories that possess a specific rank order.
- However, the difference between the ranks is not necessarily equal in value and has no true numerical meaning.
- Dental Hygiene Example: Shades of tooth whiteness (A1, A2, B1, B2) or the difficulty level of dental hygiene patients categorized as 0, 1, 2, 3, and 4.
- Interval Scale:
- Possesses all characteristics of the ordinal scale, but includes equal distances between units of measurement.
- Critically, it has no absolute zero point, meaning values can be negative.
- Example: Temperature measured on the Fahrenheit scale.
- Dental Hygiene Example: Periodontal probe measurements (though often debated, primarily categorized here for distance equality).
- Ratio Scale:
- Possesses all characteristics of the interval scale plus an absolute zero point.
- The presence of a 0 score indicates the total absence of the variable.
- All arithmetic operations can be applied to ratio data.
- Examples: Money, height, weight, number of teeth, or number of participants enrolled in a community program.
Descriptive vs. Inferential Statistics
- Descriptive Statistics:
- Procedures used to summarize, organize, and describe quantitative data.
- Utilizes tables and graphs for visualization.
- Includes measures of central tendency, measures of dispersion, and frequency distribution tables.
- Inferential Statistics:
- Procedures used to make inferences or generalizations about a larger population based on data collected from a sample.
- This process is also known as making statistical decisions.
Measures of Central Tendency and Practice Examples
- Mean: The arithmetic average of a data set. Symbolized as xˉ or M.
- Median: The exact midpoint of the data set when arranged in order.
- Mode: The value that occurs with the greatest frequency.
- Bimodal: The data set contains two modes.
- Multimodal: The data set contains more than two modes.
- Practice Array 1: Data set (2,2,3,4,5)
- Median: 3
- Mode: 2
- Mean: 3.2
- Practice Array 2: Data set (2,2,3,3,4,5,6)
- Median: 3
- Mode: 2 and 3
- Array Type: Bimodal
- Mean: 3.6
Measures of Dispersion (Variability)
- Definition: Measurements that describe how much variation is present within a group of data.
- Range: Calculated by subtracting the lowest score from the highest score. It is the simplest but least helpful measure because it is heavily affected by extreme outliers.
- Variance: Represents the average squared distance of each individual score from the calculated mean.
- Standard Deviation (SD): The square root of the variance. It is considered the most common and useful measurement of dispersion.
Normal Distribution and the Empirical Rule
- Gaussian Distribution: Also known as the normal distribution, it forms the theoretical foundation for statistical comparisons.
- Characteristics: A symmetrical, unimodal, bell-shaped curve where the Mean, Median, and Mode are exactly equal.
- The Empirical Rule (Standard Normal Distribution):
- Approximately 68% of data points fall within one standard deviation (1SD) of the mean.
- Approximately 95% of data points fall within two standard deviations (2SD) of the mean.
- Approximately 99.7% of data points fall within three standard deviations (3SD) of the mean.
- Curve Width: The larger the numerical value of the standard deviation, the wider and flatter the distribution curve appears.
Skewed Distributions
- Definition: An asymmetrical distribution of scores where the curve is distorted or pushed to one side, often caused by small sample sizes, non-random sampling, or extreme outlier scores.
- Positive Skew: The curve is shifted toward the left, meaning there are more scores in the lower range. The "tail" of the curve points toward the higher values (right).
- Negative Skew: The curve is shifted toward the right, meaning there are more scores in the higher range. The "tail" of the curve points toward the lower values (left).
- Identification: Skewness is identified by comparing the mean and the median.
Frequency Distributions and Graphing Techniques
- Frequency Distribution Tables: Show the number of times each score or item occurs. Data should be understandable without additional explanation.
- Ungrouped: Data presented in ascending or descending order with the frequency for each individual score.
- Grouped: Data presented in ranges (intervals) with the frequency of scores falling into those ranges.
- Cumulative: The frequency of occurrence of scores up to and including a specific value.
- Bar Graph: Used for categorical data. The length of the bar corresponds to frequency. A cluster bar graph can compare multiple groups.
- Histogram: Similar to a bar graph but with bars touching (side by side). Used for interval, ratio, or ordinal data treated as continuous.
- Frequency Polygon: A line graph representing continuous frequency data, created by connecting the midpoints of histogram bars.
- Scattergram: A plot showing the relationship between two variables, illustrating how one variable's level changes as the other changes.
- Pie Chart: Represents parts of a whole (represented as percentages). Generally more acceptable for lay audiences than for technical scientific publications.
Correlation and the Correlation Coefficient (r)
- Definition: Correlation measures the strength and direction of a relationship or association between two variables.
- The r-value: Signifies the correlation coefficient, ranging from −1 to +1.
- Strong Relationship: Values closer to −1 or +1. (r=−0.90 is as strong as r=+0.90).
- Weak Relationship: Values closer to 0.
- Perfect Relationship: Exact values of +1 or −1.
- No Relationship: A value of 0.
- Positive Correlation: A direct relationship. As variable x increases, variable y also increases; as x decreases, y decreases.
- Negative Correlation: An inverse relationship. As variable x increases, variable y decreases; as x decreases, y increases.
Inferential Statistical Tests and Decision Making
- Null Hypothesis: Statistical decisions are made to either reject or accept the null hypothesis based on probability.
- Probability Value (p-value):
- Evaluates the possibility that the results occurred by chance without intervention.
- Threshold: The commonly accepted p-value in oral health research is equal to or smaller than 0.05.
- Significance: If p≤0.05, results are statistically significant. If p>0.05, results are not statistically significant.
- T-test: A statistical test used to determine the hypothetical difference between two mean scores.
- ANOVA (Analysis of Variance): A test used to compare the statistical differences between three or more mean scores.
Research Quality: Validity, Reliability, and Calibration
- Validity: The degree to which an instrument measures what it is intended to measure.
- Internal Validity: Refers to the accuracy of the study results; affected by study length, sample size, and statistical accuracy.
- External Validity: Refers to the accuracy of generalizing the study results to the total population; affected by how well the sample represents that population.
- Reliability: The extent to which a measurement method performs consistently, yielding the same results in the same population upon repetition.
- Calibration: The process of establishing a relationship between a measuring device and the units of measurement to ensure consistency among examiners.
- Inter-examiner Reliability: Consistency between different examiners. Calibration increases this.
- Intra-examiner Reliability: Consistency within the same examiner over time.
Topics Excluded from Assessment
- For examination purposes, the following specific concepts are not tested from this chapter:
- Degrees of Freedom
- Type I or II Error
- Spearman Rank-Order Correlation
- Pearson Product-Moment Correlation Coefficient
- Regression Analysis
- Confidence Intervals
- Parametric and Nonparametric Statistics