Biostatistics and Data Analysis in Dental Hygiene

Introduction to Biostatistics and Data Analysis

  • Definition of Biostatistics: The application of data analysis and interpretation specifically within the context of healthcare research.
  • Role of Technology: Most contemporary computations are performed using dedicated computer programs. Researchers must possess the knowledge to input data correctly and accurately interpret the generated results.
  • Scope of Data Analysis:
    • Involves the application of various statistical tests to organize, describe, summarize, and analyze gathered data.
    • Aimed at answering researchers' specific questions or testing formulated hypotheses.
    • Requires critical thinking to explain the meaning and practical application of findings.
    • Involves identifying factors that could have influenced results and drawing inferences from sample data to the broader population.
  • Applications in Dental Hygiene:
    • Demonstrating patient response to dental hygiene therapies.
    • Testing clinical products and specific treatment regimens.
    • Determining the specific oral health needs of target populations.
    • Evaluating oral health treatments, prevention strategies, and educational programs.
  • Importance of Research Knowledge for the Dental Hygienist:
    • Understanding the epidemiology of diseases.
    • Developing and practicing evidence-based therapies.
    • Implementing effective public health programs.
    • Practicing evidence-based dentistry to ensure clinical decisions are supported by data.

Potential Causes of Invalid Research Results

  • Low Sample Size: Utilizing an insufficient number of subjects can undermine the statistical power of the study.
  • Duration Issues: Studies that are too short in duration may not capture the true effects of an intervention.
  • Tool Errors: Using incorrect or uncalibrated measurement instruments.
  • Procedural Errors: Utilizing incorrect procedures during the data collection or intervention phase.
  • Statistical Misapplication: Applying incorrect statistical tests to analyze the gathered data.

Categorization and Types of Data

  • General Definition: Data refers to the specific information collected by a researcher.
  • Qualitative Data:
    • Definition: Data that reflects the quality or nature of variables rather than numerical counts.
    • Ordering: Can often be rank-ordered (e.g., a list of patient responses regarding what they liked most vs. least).
    • Categorical Variables: Variables with no numerical representation (e.g., hair color, ratings of "good" or "bad").
    • Dichotomous Variables: Variables that place subjects into exactly two distinct groups (e.g., female/male, yes/no, pass/fail).
  • Quantitative Data:
    • Definition: Data represented by numbers and expressed as counts, percentages, and means.
    • Examples: Pocket depths, number of teeth with sealants.
    • Continuous Variables: Variables derived from a large or infinite number of measures. These can be expressed in fractions (e.g., height, weight).
    • Discrete Variables: Measurements consisting of distinct and separate units expressed in whole numbers only (e.g., the number of children, DMFT score).

Scales of Measurement

  • Nominal Scale:
    • Organizes data into mutually exclusive categories.
    • Categories have no inherent rank or numerical value.
    • Dental Hygiene Example: Patient classification (Pediatric, Adult, Geriatric) or hair color (brown, black, blonde, gray).
  • Ordinal Scale:
    • Organizes data into mutually exclusive categories that possess a specific rank order.
    • However, the difference between the ranks is not necessarily equal in value and has no true numerical meaning.
    • Dental Hygiene Example: Shades of tooth whiteness (A1A1, A2A2, B1B1, B2B2) or the difficulty level of dental hygiene patients categorized as 00, 11, 22, 33, and 44.
  • Interval Scale:
    • Possesses all characteristics of the ordinal scale, but includes equal distances between units of measurement.
    • Critically, it has no absolute zero point, meaning values can be negative.
    • Example: Temperature measured on the Fahrenheit scale.
    • Dental Hygiene Example: Periodontal probe measurements (though often debated, primarily categorized here for distance equality).
  • Ratio Scale:
    • Possesses all characteristics of the interval scale plus an absolute zero point.
    • The presence of a 00 score indicates the total absence of the variable.
    • All arithmetic operations can be applied to ratio data.
    • Examples: Money, height, weight, number of teeth, or number of participants enrolled in a community program.

Descriptive vs. Inferential Statistics

  • Descriptive Statistics:
    • Procedures used to summarize, organize, and describe quantitative data.
    • Utilizes tables and graphs for visualization.
    • Includes measures of central tendency, measures of dispersion, and frequency distribution tables.
  • Inferential Statistics:
    • Procedures used to make inferences or generalizations about a larger population based on data collected from a sample.
    • This process is also known as making statistical decisions.

Measures of Central Tendency and Practice Examples

  • Mean: The arithmetic average of a data set. Symbolized as xˉ\bar{x} or MM.
  • Median: The exact midpoint of the data set when arranged in order.
  • Mode: The value that occurs with the greatest frequency.
    • Bimodal: The data set contains two modes.
    • Multimodal: The data set contains more than two modes.
  • Practice Array 1: Data set (2,2,3,4,52, 2, 3, 4, 5)
    • Median: 33
    • Mode: 22
    • Mean: 3.23.2
  • Practice Array 2: Data set (2,2,3,3,4,5,62, 2, 3, 3, 4, 5, 6)
    • Median: 33
    • Mode: 22 and 33
    • Array Type: Bimodal
    • Mean: 3.63.6

Measures of Dispersion (Variability)

  • Definition: Measurements that describe how much variation is present within a group of data.
  • Range: Calculated by subtracting the lowest score from the highest score. It is the simplest but least helpful measure because it is heavily affected by extreme outliers.
  • Variance: Represents the average squared distance of each individual score from the calculated mean.
  • Standard Deviation (SD): The square root of the variance. It is considered the most common and useful measurement of dispersion.

Normal Distribution and the Empirical Rule

  • Gaussian Distribution: Also known as the normal distribution, it forms the theoretical foundation for statistical comparisons.
  • Characteristics: A symmetrical, unimodal, bell-shaped curve where the Mean, Median, and Mode are exactly equal.
  • The Empirical Rule (Standard Normal Distribution):
    • Approximately 68%68\% of data points fall within one standard deviation (1SD1\,SD) of the mean.
    • Approximately 95%95\% of data points fall within two standard deviations (2SD2\,SD) of the mean.
    • Approximately 99.7%99.7\% of data points fall within three standard deviations (3SD3\,SD) of the mean.
  • Curve Width: The larger the numerical value of the standard deviation, the wider and flatter the distribution curve appears.

Skewed Distributions

  • Definition: An asymmetrical distribution of scores where the curve is distorted or pushed to one side, often caused by small sample sizes, non-random sampling, or extreme outlier scores.
  • Positive Skew: The curve is shifted toward the left, meaning there are more scores in the lower range. The "tail" of the curve points toward the higher values (right).
  • Negative Skew: The curve is shifted toward the right, meaning there are more scores in the higher range. The "tail" of the curve points toward the lower values (left).
  • Identification: Skewness is identified by comparing the mean and the median.

Frequency Distributions and Graphing Techniques

  • Frequency Distribution Tables: Show the number of times each score or item occurs. Data should be understandable without additional explanation.
    • Ungrouped: Data presented in ascending or descending order with the frequency for each individual score.
    • Grouped: Data presented in ranges (intervals) with the frequency of scores falling into those ranges.
    • Cumulative: The frequency of occurrence of scores up to and including a specific value.
  • Bar Graph: Used for categorical data. The length of the bar corresponds to frequency. A cluster bar graph can compare multiple groups.
  • Histogram: Similar to a bar graph but with bars touching (side by side). Used for interval, ratio, or ordinal data treated as continuous.
  • Frequency Polygon: A line graph representing continuous frequency data, created by connecting the midpoints of histogram bars.
  • Scattergram: A plot showing the relationship between two variables, illustrating how one variable's level changes as the other changes.
  • Pie Chart: Represents parts of a whole (represented as percentages). Generally more acceptable for lay audiences than for technical scientific publications.

Correlation and the Correlation Coefficient (rr)

  • Definition: Correlation measures the strength and direction of a relationship or association between two variables.
  • The rr-value: Signifies the correlation coefficient, ranging from 1-1 to +1+1.
    • Strong Relationship: Values closer to 1-1 or +1+1. (r=0.90r = -0.90 is as strong as r=+0.90r = +0.90).
    • Weak Relationship: Values closer to 00.
    • Perfect Relationship: Exact values of +1+1 or 1-1.
    • No Relationship: A value of 00.
  • Positive Correlation: A direct relationship. As variable xx increases, variable yy also increases; as xx decreases, yy decreases.
  • Negative Correlation: An inverse relationship. As variable xx increases, variable yy decreases; as xx decreases, yy increases.

Inferential Statistical Tests and Decision Making

  • Null Hypothesis: Statistical decisions are made to either reject or accept the null hypothesis based on probability.
  • Probability Value (pp-value):
    • Evaluates the possibility that the results occurred by chance without intervention.
    • Threshold: The commonly accepted pp-value in oral health research is equal to or smaller than 0.050.05.
    • Significance: If p0.05p \le 0.05, results are statistically significant. If p>0.05p > 0.05, results are not statistically significant.
  • T-test: A statistical test used to determine the hypothetical difference between two mean scores.
  • ANOVA (Analysis of Variance): A test used to compare the statistical differences between three or more mean scores.

Research Quality: Validity, Reliability, and Calibration

  • Validity: The degree to which an instrument measures what it is intended to measure.
    • Internal Validity: Refers to the accuracy of the study results; affected by study length, sample size, and statistical accuracy.
    • External Validity: Refers to the accuracy of generalizing the study results to the total population; affected by how well the sample represents that population.
  • Reliability: The extent to which a measurement method performs consistently, yielding the same results in the same population upon repetition.
  • Calibration: The process of establishing a relationship between a measuring device and the units of measurement to ensure consistency among examiners.
    • Inter-examiner Reliability: Consistency between different examiners. Calibration increases this.
    • Intra-examiner Reliability: Consistency within the same examiner over time.

Topics Excluded from Assessment

  • For examination purposes, the following specific concepts are not tested from this chapter:
    • Degrees of Freedom
    • Type I or II Error
    • Spearman Rank-Order Correlation
    • Pearson Product-Moment Correlation Coefficient
    • Regression Analysis
    • Confidence Intervals
    • Parametric and Nonparametric Statistics