Biostatistics and Data Analysis Review

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/22

flashcard set

Earn XP

Description and Tags

Comprehensive practice flashcards covering data types, descriptive statistics, probability theory, diagnostics, normal distributions, correlation, and parametric vs. non-parametric hypothesis testing.

Last updated 9:52 PM on 10/7/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

23 Terms

1
New cards

What distinguishes continuous quantitative data from discrete quantitative data?

Continuous data is measured and can take any value within a range including decimals (e.g., height or weight), whereas discrete data is counted and typically takes whole-number values (e.g., number of siblings).

2
New cards

How do nominal and ordinal categorical data differ?

Nominal data consists of categories with no natural order (such as eye color or blood type), while ordinal data consists of categories that follow a meaningful order (such as pain level or mild, moderate, severe).

3
New cards

What is the difference between a parameter and a statistic?

A parameter is a numerical value describing a population (represented by size NN), whereas a statistic is a numerical value describing a sample (represented by size nn).

4
New cards

Why does the formula for sample variance (s2s^2) use n−1n - 1 in the denominator instead of NN?

Sample variance uses n−1n - 1 in the denominator to reduce bias when estimating the population parameter, whereas population variance (σ2\sigma^2) divides by the entire population size NN.

5
New cards

What mathematical condition must be met for two events AA and BB to be statistically independent?

Two events are independent if and only if P(A∩B)=P(A)P(B)P(A \cap B) = P(A)P(B).

6
New cards

In diagnostic testing, how are Sensitivity and Specificity defined?

Sensitivity is the probability of testing positive given that the person has the disease (P(T∣D)P(T \mid D)), while Specificity is the probability of testing negative given that the person does not have the disease (P(Tc∣Dc)P(T^c \mid D^c)).

7
New cards

How are Positive Predictive Value (PPV) and Negative Predictive Value (NPV) defined?

PPV is the probability that a person has the disease given a positive test (P(D∣T)P(D \mid T)), whereas NPV is the probability that a person does not have the disease given a negative test (P(Dc∣Tc)P(D^c \mid T^c)).

8
New cards

What are the four risks of categorical thinking identified in research?

The four risks are Discrimination (favoring or disadvantaging based on category), Compression (treating individuals within a category as more similar than they are), Amplification (exaggerating differences between categories), and Fossilization (treating categories as fixed and permanent).

9
New cards

What empirical rule describes the distribution of data within standard deviations of a normal distribution?

The 68–95–99.768\text{--}95\text{--}99.7 rule states that approximately 68%68\% of observations fall within 11 standard deviation of the mean, 95%95\% fall within 22 standard deviations, and 99.7%99.7\% fall within 33 standard deviations.

10
New cards

How is the pp-value interpreted when using the Shapiro-Wilk test for normality?

If p<0.05p < 0.05, the data differs significantly from a normal distribution and normality cannot be assumed. If p≥0.05p \ge 0.05, there is no significant evidence that the data differs from normal.

11
New cards

What condition must be satisfied according to the Central Limit Theorem (CLT) for sample means to be normally distributed?

Regardless of the population's underlying distribution, randomly selected sample means approximate a normal distribution when the sample size reaches n≥30n \ge 30.

12
New cards

How do precision and accuracy differ in statistical estimation?

Precision is how close a sample estimate is to the true population value (measured by variability and standard error), while accuracy asks whether the calculated confidence interval actually contains the true population parameter.

13
New cards

What is the formula for Standard Error (SESE), and what happens as sample size (nn) increases?

The formula is SE=snSE = \frac{s}{\sqrt{n}}. As sample size increases, variability decreases, SESE decreases, and precision increases.

14
New cards

Why are Median and IQR preferred over Mean and SD when summarizing non-normal data?

Median and IQR are preferred because the mean and standard deviation are heavily influenced by extreme outliers, whereas the median is resistant to them.

15
New cards

What are the primary differences in requirements between Pearson's and Spearman's correlation?

Pearson's correlation is a parametric test requiring continuous variables, normal distribution, and a linear relationship. Spearman's correlation is non-parametric, requires continuous or ordinal variables, does not require normality, and only requires a monotonic relationship.

16
New cards

What do the null hypothesis (H0H_0) and alternative hypothesis (HaH_a) represent?

H0H_0 represents the current belief or baseline assumption of no difference or no effect, while HaH_a represents the challenge indicating a significant difference or directional effect.

17
New cards

When is a Z-test used instead of a T-test?

A Z-test is used when the population standard deviation (σ\sigma) is known. A T-test is used when the population standard deviation is unknown and the sample standard deviation (ss) must be used.

18
New cards

What is the decision rule in hypothesis testing when comparing the pp-value to the significance level (α\alpha)?

If p<αp < \alpha, reject H0H_0. If p≥αp \ge \alpha, fail to reject H0H_0.

19
New cards

How do you choose between a one-sample t-test, independent t-test, and paired t-test?

Use a one-sample t-test to compare a sample mean to a population mean, an independent t-test to compare two separate independent groups, and a paired t-test when the same participants are measured twice (e.g., before and after).

20
New cards

What are the non-parametric alternatives to the independent samples t-test and paired t-test?

The Mann-Whitney-Wilcoxon test (Wilcoxon rank-sum) is the non-parametric alternative for independent groups, and the Wilcoxon signed-rank test is the alternative for paired groups.

21
New cards

When should Fisher's exact test be selected over a Chi-squared test?

Fisher's exact test is selected when working with categorical data with small sample sizes where the Chi-squared assumption of having expected cell counts of at least 55 is violated.

22
New cards

What is the difference between a Type I error and a Type II error?

A Type I error is rejecting a true null hypothesis (H0H_0) (false positive, probability α\alpha), while a Type II error is failing to reject a false null hypothesis (H0H_0) (false negative, probability β\beta).

23
New cards

What is statistical power and how is it calculated?

Statistical power is the probability of correctly rejecting a false null hypothesis, calculated as Power=1−β\text{Power} = 1 - \beta.