Introduction to Biostatistics Flashcards
Course and Academic Context
College: College of Medical Laboratory Science
Course Code & Title: BIOE211 – Introduction to Biostatistics
Overview of Biostatistics
Etymology & Fundamental Terms:
BIO: Refers to life.
STATISTICS: Refers to the science dealing with the collection, organization, analysis, and interpretation of numerical data.
Definition of Biostatistics:
The application of statistical methods to the life sciences, including biology, medicine, and public health.
A subfield of statistics that focuses specifically on the analysis, interpretation, and application of quantitative data in biological, medical, and public health research.
Integrates mathematical and statistical concepts with domain knowledge from biological sciences to enhance the understanding of human health and directly inform decision-making in healthcare settings.
Main Areas of Statistics
Mathematical Statistics:
Concerns the development of new methods of statistical inference.
Requires detailed knowledge of abstract mathematics for its theoretical formulation and implementation.
Applied Statistics:
Involves applying the methods established in mathematical statistics to specific, concrete subject areas.
Biostatistics as Applied Statistics: Biostatistics is a specialized branch of applied statistics that applies these statistical methods to biological and medical problems.
Philippine Statistics Authority (PSA)
Role: The central statistical authority of the Philippine government.
Functions:
Collects, compiles, analyzes, and publishes statistical information regarding economic, social, demographic, political, and general affairs of the Philippine population.
Enforces civil registration functions across the country.
Sub-Areas of Statistics
Descriptive Statistics:
Used to summarize and describe the main features of a dataset, providing an overview of data without drawing conclusions or making generalizations beyond the immediate set.
Includes methods of collecting, classifying, graphing, and averaging data strictly to describe its properties or characteristics.
Inferential Statistics:
Also referred to as Statistical Inference or Inductive Statistics.
Demands a higher degree of critical judgment and relies on advanced mathematical models to test the significance of observed results.
Concerned with drawing conclusions, generalizations, or inferences about a larger population based on organized sample data.
Population vs. Sample
Population (Universe):
Consists of all members of the specified group about which a researcher or analyst wants to draw conclusions.
Sample:
A representative portion or subset of individuals or items selected from a larger population.
Chosen for the purpose of conducting observations, performing experiments, or gathering data to make valid inferences or generalizations back to the full population.
Fundamentals of Data
Definition: A collection of facts, figures, or observations used for analysis, interpretation, or decision-making.
Function in Research: Data serves as the raw material that allows scientists and analysts to derive insights, test hypotheses, and validate theories.
Primary Categories:
Qualitative data
Quantitative data
Classifications of Data:
Based on Origin
Based on Structure
Based on Measurement Scale
Categories of Data
Qualitative Data:
Non-numerical information that describes qualities, attributes, or characteristics of a subject, event, or phenomenon.
Analysis involves identifying patterns, themes, or relationships and can be more subjective than quantitative analysis.
Sub-categories:
Nominal Data: Categories or labels that possess no intrinsic order, rank, or preference.
Qualitative Variable Examples & Categories:
Gender: Male, Female
Automobile Ownership: Yes, No
Type of Life Insurance Owned: Term, Endowment, Straight-life, Others, None
Ordinal Data: Categories that maintain a clear order or ranking, though the distances/intervals between the ranks are not equal or measurable.
Qualitative Variable Examples & Categories:
Student Class Designation: Freshman, Sophomore, Junior, Senior
Product Satisfaction: Unsatisfied, Neutral, Satisfied, Very Satisfied
Movie Classification: G, PG, PG-13, R-18, X
Faculty Rank: Professor, Associate Prof., Assistant Prof., Instructor
Student Grades: 1.00, 1.25, 1.50, 1.75, 2.00, …
Quantitative Data:
Numerical information that can be measured or counted.
Analysis utilizes statistical methods and techniques to describe data and draw inferences; typically more objective and precise than qualitative analysis.
Sub-categories:
Discrete Data: Consists of countable, whole numbers.
Continuous Data: Consists of measurements that can assume any value within a given continuous range, including fractions and decimals.
Classifications of Data
Based on Origin:
Primary Data: Gathered directly from the original source or subjects of the study using methods such as interviews, surveys, direct experiments, or field observations.
Secondary Data: Sourced from data that was previously collected by other entities and made available for reuse (e.g., government statistics, published research findings, company reports).
Based on Structure:
Structured Data: Pre-organized in a specific format (e.g., tables, spreadsheets, databases) allowing direct processing and analysis by computers.
Unstructured Data: Lacks a predefined structure or format (e.g., plain text, images, audio files, video recordings); requires advanced processing techniques for extraction, analysis, and interpretation.
Measurement Scales and Levels of Measurement
Measurement Scale Definitions:
Nominal Data: Used to differentiate classes or categories purely for identification or classification purposes.
Ordinal Data: Used in ranking items, though without measurable or standardized distances between individual ranks.
Interval Data: Numerical data possessing a consistent scale and equal distances/intervals between consecutive values, but lacking a true or absolute zero point.
Ratio Data: Numerical data with a consistent scale, equal intervals between values, and a true/absolute zero point (e.g., height, weight, age).
Characteristics and Properties of Levels of Measurement:
Nominal Level:
Indicates a distinction.
Ordinal Level:
Indicates a distinction.
Indicates the direction of the distinction (less than or more than).
Interval Level:
Indicates a distinction.
Indicates the direction of the distinction.
Indicates the amount of distinction (in equal intervals).
Ratio Level:
Indicates a distinction.
Indicates the direction of the distinction.
Indicates the amount of distinction.
Indicates an absolute zero.
Classification of Variables
Variable Definition: In research, a variable is any characteristic or attribute that can take on different values or categories.
Types of Variables:
Independent Variable (Explanatory Variable):
Controlled or manipulated by the researcher to determine its effect on the dependent variable.
Represents the presumed cause of change in an experimental setup.
Dependent Variable (Outcome Variable):
Expected to change as a direct result of manipulating the independent variable.
Represents the outcome or response that is measured and observed.
Control Variable:
Held constant by the researcher to minimize its potential impact on the dependent variable.
Serves to reduce the influence of confounding variables and increase the internal validity of a study.
Confounding Variable:
An unmanaged variable that may influence the relationship between the independent and dependent variables.
Obscures the true effect of the independent variable on the dependent variable, potentially leading to spurious correlations or incorrect conclusions.
Additional structural classifications include Categorical, Continuous, and Discrete variables.
Data Collection Definition and Purpose
Definition: The systematic process of gathering and measuring information on variables of interest in an organized and consistent manner to answer specific research questions, test hypotheses, or evaluate outcomes.
Main Purpose: To obtain accurate, reliable, and relevant information that can be analyzed and interpreted to generate actionable insights, support evidence-based decision-making, or inform policy development.
Methods of Data Collection
Surveys and Questionnaires:
Structured or semi-structured instruments designed to collect data from a sample of individuals or organizations by administering questions or recording agreement levels with various statements.
Interviews:
One-on-one or group conversations conducted between a researcher and participants to elicit detailed, in-depth information concerning their experiences, opinions, feelings, or attitudes toward the research topic.
Observations:
Systematic watching, recording, and analyzing of behaviors, events, or interactions as they occur in their natural settings, without environmental manipulation or interference by the researcher.
Experiments:
Designs where researchers actively manipulate one or more independent variables under strictly controlled conditions to observe and measure the specific outcome on a dependent variable.
Secondary Data Analysis:
The collection and secondary evaluation of pre-existing data originally gathered by other parties, such as non-governmental organizations, government agencies, research institutes, or private corporations.
Case Studies:
Intensive, in-depth examinations of a single case or a small set of cases, incorporating multiple data sources such as artifacts, documents, interviews, direct observations, or audiovisual materials.
Content Analysis:
Systematic examination and interpretation of the content of text, images, or audiovisual materials to extract patterns, themes, or meanings relevant to the research context.