1/22
Comprehensive practice flashcards covering data types, descriptive statistics, probability theory, diagnostics, normal distributions, correlation, and parametric vs. non-parametric hypothesis testing.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
What distinguishes continuous quantitative data from discrete quantitative data?
Continuous data is measured and can take any value within a range including decimals (e.g., height or weight), whereas discrete data is counted and typically takes whole-number values (e.g., number of siblings).
How do nominal and ordinal categorical data differ?
Nominal data consists of categories with no natural order (such as eye color or blood type), while ordinal data consists of categories that follow a meaningful order (such as pain level or mild, moderate, severe).
What is the difference between a parameter and a statistic?
A parameter is a numerical value describing a population (represented by size N), whereas a statistic is a numerical value describing a sample (represented by size n).
Why does the formula for sample variance (s2) use n−1 in the denominator instead of N?
Sample variance uses n−1 in the denominator to reduce bias when estimating the population parameter, whereas population variance (σ2) divides by the entire population size N.
What mathematical condition must be met for two events A and B to be statistically independent?
Two events are independent if and only if P(A∩B)=P(A)P(B).
In diagnostic testing, how are Sensitivity and Specificity defined?
Sensitivity is the probability of testing positive given that the person has the disease (P(T∣D)), while Specificity is the probability of testing negative given that the person does not have the disease (P(Tc∣Dc)).
How are Positive Predictive Value (PPV) and Negative Predictive Value (NPV) defined?
PPV is the probability that a person has the disease given a positive test (P(D∣T)), whereas NPV is the probability that a person does not have the disease given a negative test (P(Dc∣Tc)).
What are the four risks of categorical thinking identified in research?
The four risks are Discrimination (favoring or disadvantaging based on category), Compression (treating individuals within a category as more similar than they are), Amplification (exaggerating differences between categories), and Fossilization (treating categories as fixed and permanent).
What empirical rule describes the distribution of data within standard deviations of a normal distribution?
The 68–95–99.7 rule states that approximately 68% of observations fall within 1 standard deviation of the mean, 95% fall within 2 standard deviations, and 99.7% fall within 3 standard deviations.
How is the p-value interpreted when using the Shapiro-Wilk test for normality?
If p<0.05, the data differs significantly from a normal distribution and normality cannot be assumed. If p≥0.05, there is no significant evidence that the data differs from normal.
What condition must be satisfied according to the Central Limit Theorem (CLT) for sample means to be normally distributed?
Regardless of the population's underlying distribution, randomly selected sample means approximate a normal distribution when the sample size reaches n≥30.
How do precision and accuracy differ in statistical estimation?
Precision is how close a sample estimate is to the true population value (measured by variability and standard error), while accuracy asks whether the calculated confidence interval actually contains the true population parameter.
What is the formula for Standard Error (SE), and what happens as sample size (n) increases?
The formula is SE=ns. As sample size increases, variability decreases, SE decreases, and precision increases.
Why are Median and IQR preferred over Mean and SD when summarizing non-normal data?
Median and IQR are preferred because the mean and standard deviation are heavily influenced by extreme outliers, whereas the median is resistant to them.
What are the primary differences in requirements between Pearson's and Spearman's correlation?
Pearson's correlation is a parametric test requiring continuous variables, normal distribution, and a linear relationship. Spearman's correlation is non-parametric, requires continuous or ordinal variables, does not require normality, and only requires a monotonic relationship.
What do the null hypothesis (H0) and alternative hypothesis (Ha) represent?
H0 represents the current belief or baseline assumption of no difference or no effect, while Ha represents the challenge indicating a significant difference or directional effect.
When is a Z-test used instead of a T-test?
A Z-test is used when the population standard deviation (σ) is known. A T-test is used when the population standard deviation is unknown and the sample standard deviation (s) must be used.
What is the decision rule in hypothesis testing when comparing the p-value to the significance level (α)?
If p<α, reject H0. If p≥α, fail to reject H0.
How do you choose between a one-sample t-test, independent t-test, and paired t-test?
Use a one-sample t-test to compare a sample mean to a population mean, an independent t-test to compare two separate independent groups, and a paired t-test when the same participants are measured twice (e.g., before and after).
What are the non-parametric alternatives to the independent samples t-test and paired t-test?
The Mann-Whitney-Wilcoxon test (Wilcoxon rank-sum) is the non-parametric alternative for independent groups, and the Wilcoxon signed-rank test is the alternative for paired groups.
When should Fisher's exact test be selected over a Chi-squared test?
Fisher's exact test is selected when working with categorical data with small sample sizes where the Chi-squared assumption of having expected cell counts of at least 5 is violated.
What is the difference between a Type I error and a Type II error?
A Type I error is rejecting a true null hypothesis (H0) (false positive, probability α), while a Type II error is failing to reject a false null hypothesis (H0) (false negative, probability β).
What is statistical power and how is it calculated?
Statistical power is the probability of correctly rejecting a false null hypothesis, calculated as Power=1−β.