1/69
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Cross-Sectional Study
Data are observed, measured, and collected at one point in time, not over a period of time.
Retrospective (case-control) study
Data are collected from a past time period by going back in time (through examination of records, interviews, and so on)
Prospective (Longitudinal/Cohort) study
data are collected in the future from groups that share common factors (such groups are called cohorts)
Bad Experimental Design
Treat all women subjects and give the men a placebo. (Problem: We don’t know if effects are due to sex or to treatment.)
Completely Randomized Experimental Design
Assign subjects to different treatment groups through a process of random selection. Use randomness to determine who gets the treatment and who gets the placebo
Randomized Block Design
A block is a group of subjects that are similar, but blocks differ in ways that might affect the outcome of the experiment. Use the following procedure:
1. Form blocks (or groups) of subjects with similar characteristics.
2. Randomly assign treatments to the subjects within each block.
Randomized Block Design
A design uses the same basic idea as stratified sampling, but this design are used when designing experiments, whereas stratified sampling is used for surveys
Matched Pairs Design
Compare two treatment groups (such as treatment and placebo) by using subjects matched in pairs that are somehow related or have similar characteristics, as in the following cases.
Get measurements from the same subjects before and after some treatment
Rigorously Controlled Design
Carefully assign subjects to different treatment groups, so that those given each treatment are similar in the ways that are important to the experiment.
This can be extremely difficult to implement, and often we can never be sure that we have accounted for all of the relevant factors
Random Sampling Error
occurs when the sample has been selected with a random method, but there is a discrepancy between a sample result and the true population result; such an error results from chance sample fluctuations
Non-sampling Error
is the result of human error, including such factors as wrong data entries, computing errors, questions with biased wording, false data provided by respondents, forming biased conclusions, or applying statistical methods that are not appropriate for the circumstances
Non-Random Sampling Error
is the result of using a sampling method that is not random, such as using a convenience sample or a voluntary response sample
Biostatistics
branch of statistics responsible for interpreting the scientific data that is generated in the health sciences, including the public health sphere.
Biostatisticians
Responsible to consider the variables in subjects (in public health, subjects are usually patients, communities, or populations), to understand them, and to make sense of different sources of variation
Objective of Biostatistics
to disentangle the data received and make valid inferences that can be used to solve problems in public health.
Application of Biostatistics
chronic diseases managing & thriving
Human growth development
Genetics and environment
Clinical medicine
Health economics
Role of Biostatisticians
take complex, mathematical findings of clinical trials and research-related data and translate them into valuable information that is used to make public health decisions.
these professionals use mathematics to enhance science and bridge the gap between theory and practice
required to develop statistical methods for clinical trials, observational studies, longitudinal studies, and genomics
Informatics
known as bioinformatics, a science that relies on the basic disciplines of science, mathematics, probability and statistics, and computer science to build a solid statistical foundation for making advances, improvements, and even breakthroughs in public health and medicine.
Health Informatics
deals with the resources, devices, and methods required for the effective storage, use, and retrieval of information,
Public Health Informatics
public health informatics includes the application of informatics in public health areas, such as surveillance, prevention, preparedness, and health promotion.
System Analyst
write and troubleshoot the software used by biostatisticians and researchers. Their work may also include conducting their own research, designing databases, and developing algorithms for processing and analyzing information.
Responsibilities of System Analyst
Incorporating bioinformatics/biostatistics into efficient and automated data analysis tools
Developing and tracking quality workflow metrics for detecting variants and sequences
Working with scientists and researchers to develop project plan
For Medical Professionals
crucial for designing, analyzing, and interpreting medical research studies.
Understanding biostatistics is essential for comprehending and applying research findings
Statistical literacy helps identify errors in statistical methods and reporting in medical literature
Statistical Competence in Research
Developing, teaching, and assessing statistical competencies is vital for medical researchers.
Statistical competencies define the statistical knowledge and skills required for medical research.
Research focuses on identifying and developing statistical competencies for medical research learners
A Cornerstone of Medical Research
Biostatistics is the application of statistics to human health and disease.
It covers a wide range of fields within medicine and health sciences.
Biostatisticians collaborate with researchers in study design, data analysis, Interpretation, and reporting.
Biostatistics plays a critical role in identifying causes of diseases and evaluating treatment effectivene
Cochran’s Formula
Developed by statistician William G. Cochran in 1963, this formula is designed to calculate a representative sample size for a large or unknown population based on a desired level of precision, confidence level, and estimated proportion of an attribute in the population
Confidence Level (Z)
The critical value corresponding to your desired Z = 1.96 confidence level (e.g., for a 95% confidence level)
Estimated Proportion
The expected proportion of the population that p has a particular attribute. If unknown, (50%) is used because it p = 0.5 provides the maximum possible sample size (most conservative estimate).
Complementary Proportion
q = 1− p
Margin of Error
The acceptable level of precision or error bound (e.g., e = 0.05 ±5 %for )
Uses of Cochran’s Formula
conducting rigorous academic research, thesis work, or publication- grade studies.
You know (or can estimate) the variance/proportion ( ) in the population. p
You need to specify a precise confidence level (e.g., 99% or 90% rather than just a fixed default).
The total population size is unknown, infinitely large, or very large.
Uses of Slovin’ Formula
You are performing preliminary surveys, exploratory research, or rapid field studies.
You have zero information about the population's variability or variance ( ). p N
The population ( ) is finite and explicitly known.
Strict mathematical precision is less critical than speed and ease of computation
Frequency Distribution
shows how data are partitioned among several categories (or classes) by listing the categories along with the number (frequency) of data values in each of them.
Lower Class Limits
are the smallest numbers that can belong to each of the different classes.
Upper Class Limits
are the largest numbers that can belong to each of the different classes
Class Boundaries
are the numbers used to separate the classes, but without the gaps created by class limits
Class Midpoints
are the values in the middle of the classes.
Class Width
is the difference between two consecutive lower class limits (or two consecutive lower class boundaries) in a frequency distribution.
Histograms
a graph consisting of bars of equal width drawn adjacent to each other (unless there are gaps in the data). The horizontal scale represents classes of quantitative data values, and the vertical scale represents frequencies. The heights of the bars correspond to frequency values.
Relative Frequency Histogram
A relative frequency histogram has the same shape and horizontal scale as a histogram, but the vertical scale uses relative frequencies (as percentages or proportions) instead of actual frequencies
Histogram Uses
Visually displays the shape of the distribution of the data
Shows the location of the center of the data
Shows the spread of the data
Identifies outliers
Normal Distribution
Bell-shaped
Many statistical methods require that sample data come from a population having a distribution approximately normal
Aside from histograms, they are very helpful to assess normalit
Uniform Distribution
The different possible values occur with approximately the same frequency, so the heights of the bars in the histogram are approximately uniform.
The graph on the right depicts outcomes of last digits of weights from a large sample of randomly selected subjects, and such a graph is helpful in determining whether the subjects were actually weighed or whether they reported their weights
Normal Quantile Plot
Normal Distribution
Not a Normal Distribution
Normal Distribution
The population distribution is normal if the pattern of the points in the normal quantile plot is reasonably close to a straight line, and the points do not show some systematic pattern that is not a straight-line pattern.
Not a Normal Distribution
The population distribution is not normal if the normal quantile plot has either or both of these two conditions:
• The points do not lie reasonably close to a straight-line pattern.
• The points show some systematic pattern that is not a straight-line pattern
Dot Plot
consists of a graph of quantitative data in which each data value is plotted as a above a horizontal scale of values. Dots representing equal values are stacked.
Displays the shape of the distribution of data.
It is usually possible to recreate the original list of data values
Dot Plot
consists of a graph of quantitative data in which each data value is plotted as a above a horizontal scale of values. Dots representing equal values are stacked.
Displays the shape of the distribution of data.
It is usually possible to recreate the original list of data values
Time-Series Graph
which are quantitative data that have been collected at different points in time, such as monthly or yearly. An advantage of a time-series graph is that it reveals information about trends over time.
Reveals information about trends over tiMe
Bar Graph
uses bars of equal width to show frequencies of categories of categorical (or qualitative) data.
The bars may or may not be separated by small gaps.
Features:
■ Shows the relative distribution of categorical data so that it is easier to compare the different categories
Pareto Charts
is a bar graph for categorical data, with the added stipulation that the bars are arranged in descending order according to frequencies, so the bars decrease in height from left to right.
Features:
■ Shows the relative distribution of categorical data so that it is easier to compare the different categories
■ Draws attention to the more important categories
Pie Charts
a very common graph that depicts categorical data as slices of a circle, in which the size of each slice is proportional to the frequency count for the category. Although pie charts are very common, they are not as effective as Pareto charts.
Features:
■ Shows the distribution of categorical data in a commonly used forma
Frequency Polygon
uses line segments connected to points located directly above class midpoint values.
Relative Frequency Polygon
uses relative frequencies (proportions or percentages) for the vertical scale.
An advantage of relative frequency polygons is that two or more of them can be combined on a single graph for easy comparison
Boxplots
uses the relationships among the median, upper quartile, and lower quartile to describe the skewness of a distribution.
Graphs that Deceive
commonly used to mislead people, and we really don’t want statistics students to be among those susceptible to such deceptions. Graphs should be constructed in a way that is fair and objective. The readers should be allowed to make their own judgments, instead of being manipulated by misleading graphs
Nonzero Vertical Axis
involves using a vertical scale that starts at some value greater than zero to exaggerate differences between groups.
TIP: Always examine a graph carefully to see whether a vertical axis begins at some point other than zero so that differences are exaggerated
Pictographs
They are often misleading. Data that are one-dimensional in nature (such as budget amounts) are often depicted with two-dimensional objects (such as dollar bills) or three-dimensional objects (such as stacks of coins, homes, or barrels). Artists can create false impressions that grossly distort differences
Correlation
exists between two variables when the values of one variable are somehow associated with the values of the other variable.
Linear Correlation
exists between two variables when there is a correlation and the plotted points of paired data result in a pattern that can be approximated by a straight line
Scatterplot
a plot of paired (x, y) quantitative data with a horizontal x-axis and a vertical y-axis. The horizontal axis is used for the first variable (x), and the vertical axis is used for the second variable (y)
Linear Correlation Coefficient
denoted by r, and it measures the strength of the linear association between two variables
Regression Line
Given a collection of paired sample data, the regression line (or line of best fit or least-squares line) is the straight line that “best” fits the scatterplot of the data.