Introductory Applied Statistics for the Life Sciences Flashcards

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/44

flashcard set

Earn XP

Description and Tags

Flashcards covering Lectures 02, 04, 05, 06, and 07 on descriptive statistics, probability foundations, discrete distributions, continuous distributions, and scatterplots/regression.

Last updated 5:55 PM on 10/5/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

45 Terms

1
New cards

What is the primary difference between descriptive statistics and model probability?

Descriptive statistics describe what happened in the observed sample data (such as the observed proportion p^\hat{p}), whereas model probability describes what results would be plausible under a specified chance model (such as p=0.20p = 0.20 for uniform guessing).

2
New cards

What is a random process in probability theory?

A random process is a process whose outcome is uncertain before it occurs, even when the underlying conditions are specified.

3
New cards

How are a sample space and an event defined?

A sample space (SS) is the set of all possible outcomes of a random process. An event is a collection of one or more outcomes from the sample space.

4
New cards

What are the fundamental probability rules for any event AA and sample space SS in a finite setting?

For any event AA, 0≤P(A)≤10 \le P(A) \le 1, and for the sample space SS, P(S)=1P(S) = 1. Under finite sample spaces, P(A)=0P(A) = 0 means the event is impossible, and P(A)=1P(A) = 1 means it is certain.

5
New cards

How is a probability interpreted under the frequentist perspective?

A probability can be interpreted as a long-term relative frequency of an outcome over many comparable repetitions of a random process.

6
New cards

How is the probability of an event AA calculated when a finite sample space SS has equally likely outcomes?

P(A)=number of outcomes in Anumber of outcomes in SP(A) = \frac{\text{number of outcomes in } A}{\text{number of outcomes in } S}. Equal likelihood must be justified by the model and does not apply automatically.

7
New cards

What is the Complement Rule, and when is it particularly useful?

The Complement Rule states that P(Ac)=1−P(A)P(A^c) = 1 - P(A), where AcA^c is the complement of AA (the event that AA does not occur). It is particularly useful when calculating "at least one" probabilities by using "none."

8
New cards

What is the General Addition Rule, and why is the intersection subtracted?

The General Addition Rule states P(A or B)=P(A)+P(B)−P(A and B)P(A \text{ or } B) = P(A) + P(B) - P(A \text{ and } B). The intersection P(A and B)P(A \text{ and } B) is subtracted because outcomes contained in both events were counted twice when adding P(A)P(A) and P(B)P(B).

9
New cards

What are disjoint (mutually exclusive) events, and what is their addition rule?

Disjoint events cannot occur together, meaning P(A and B)=0P(A \text{ and } B) = 0. For disjoint events, the addition rule simplifies to P(A or B)=P(A)+P(B)P(A \text{ or } B) = P(A) + P(B).

10
New cards

How is conditional probability P(A∣B)P(A|B) defined, and what determines its denominator?

Conditional probability is defined as P(A∣B)=P(A and B)P(B)P(A|B) = \frac{P(A \text{ and } B)}{P(B)} for P(B)>0P(B) > 0. The event following the word "given" (event BB) determines the restricted sample space and denominator.

11
New cards

Why does the direction of conditioning matter when comparing P(A∣B)P(A|B) and P(B∣A)P(B|A)?

P(A∣B)P(A|B) and P(B∣A)P(B|A) generally differ because they restrict attention to different groups, resulting in different denominators even though their numerators P(A and B)P(A \text{ and } B) are identical.

12
New cards

What is the General Multiplication Rule for any two sequential events AA and BB?

P(A and B)=P(A)⋅P(B∣A)P(A \text{ and } B) = P(A) \cdot P(B|A), which multiplies the probability of the first event by the probability of the second event given that the first has occurred.

13
New cards

What does it mean for two events AA and BB to be independent?

Events AA and BB are independent if knowing that one occurred does not change the probability of the other occurring, satisfying P(A∣B)=P(A)P(A|B) = P(A), P(B∣A)=P(B)P(B|A) = P(B), and P(A and B)=P(A)P(B)P(A \text{ and } B) = P(A)P(B).

14
New cards

Why can two disjoint events with positive probabilities never be independent?

If event AA occurs, a disjoint event BB becomes impossible (P(B∣A)=0P(B|A) = 0), which changes the probability of BB from its original positive value P(B)P(B). Therefore, disjoint events with positive probabilities are always dependent.

15
New cards

What is a random variable, and how do uppercase and lowercase letters represent it?

A random variable is a rule that assigns one numerical value to each possible outcome of a random process. Uppercase XX denotes the random variable (uncertain before observation), while lowercase xx denotes one realized or observed numerical value.

16
New cards

How do discrete and continuous random variables differ?

Discrete random variables have possible values that can be listed or counted (such as counts or indicators), whereas continuous random variables can take any value within a numerical interval (such as measurements).

17
New cards

What is a Probability Mass Function (PMF), and what conditions must it satisfy?

A PMF pX(x)=P(X=x)p_X(x) = P(X = x) gives the probability of each exact value for a discrete random variable XX. It must satisfy pX(x)≥0p_X(x) \ge 0 for all xx, ∑pX(x)=1\sum p_X(x) = 1, and pX(x)=0p_X(x) = 0 for impossible values.

18
New cards

What defines a Bernoulli random variable?

A Bernoulli random variable X∼Bernoulli(p)X \sim \text{Bernoulli}(p) records a single binary outcome where X=1X = 1 if the event of interest ("success") occurs with probability pp, and X=0X = 0 with probability 1−p1 - p.

19
New cards

What are the four BINS conditions required for a Binomial distribution?

BINS stands for: Binary outcomes (success or failure on each trial), Independent trials, Number of trials fixed in advance (nn), and Same success probability (pp) on every trial.

20
New cards

What is the formula for exact probabilities of a Binomial random variable X∼Binomial(n,p)X \sim \text{Binomial}(n, p)?

P(X=k)=(nk)pk(1−p)n−kP(X = k) = \binom{n}{k} p^k (1 - p)^{n-k}, where (nk)=n!k!(n−k)!\binom{n}{k} = \frac{n!}{k!(n - k)!} is the binomial coefficient counting combinations of kk successes in nn trials.

21
New cards

In R, what are the distinct functions dbinom(), pbinom(), and qbinom() used for?

dbinom(x,size,prob)\text{dbinom}(x, \text{size}, \text{prob}) calculates exact point probability P(X=x)P(X = x); pbinom(q,size,prob)\text{pbinom}(q, \text{size}, \text{prob}) calculates cumulative left-tail probability P(X≤q)P(X \le q); and qbinom(p,size,prob)\text{qbinom}(p, \text{size}, \text{prob}) calculates the smallest count corresponding to a cumulative probability pp.

22
New cards

How do you express P(X≥16)P(X \ge 16) and P(14≤X≤17)P(14 \le X \le 17) using R commands for X∼Binomial(20,0.75)X \sim \text{Binomial}(20, 0.75)?

P(X≥16)=1−P(X≤15)P(X \ge 16) = 1 - P(X \le 15), written as 1−pbinom(15,20,0.75)1 - \text{pbinom}(15, 20, 0.75). P(14≤X≤17)=P(X≤17)−P(X≤13)P(14 \le X \le 17) = P(X \le 17) - P(X \le 13), written as pbinom(17,20,0.75)−pbinom(13,20,0.75)\text{pbinom}(17, 20, 0.75) - \text{pbinom}(13, 20, 0.75).

23
New cards

What are the mean, variance, and standard deviation formulas for Bernoulli and Binomial distributions?

For Bernoulli(p)\text{Bernoulli}(p): μ=p\mu = p, σ2=p(1−p)\sigma^2 = p(1 - p), σ=p(1−p)\sigma = \sqrt{p(1 - p)}. For Binomial(n,p)\text{Binomial}(n, p): μ=np\mu = np, σ2=np(1−p)\sigma^2 = np(1 - p), σ=np(1−p)\sigma = \sqrt{np(1 - p)}.

24
New cards

How is the expected value E(X)E(X) of a discrete random variable defined and interpreted?

E(X)=μX=∑x⋅P(X=x)E(X) = \mu_X = \sum x \cdot P(X = x). It represents the long-run average center of the probability distribution over many repetitions and does not need to be a value XX can actually take.

25
New cards

Why is the probability of a continuous random variable taking any exact single point equal to zero (P(X=x)=0P(X = x) = 0)?

Continuous variables model exact measurements where an exact point has zero area under the probability density curve. Recorded measurements like 120 mmHg120\,\text{mmHg} are rounded numbers representing a small interval (119.5≤X<120.5119.5 \le X < 120.5).

26
New cards

Why does including or excluding interval endpoints not change probabilities for continuous random variables?

Because P(X=a)=0P(X = a) = 0 and P(X=b)=0P(X = b) = 0 for a continuous random variable XX, P(a<X<b)=P(a≤X≤b)=P(a<X≤b)=P(a≤X<b)P(a < X < b) = P(a \le X \le b) = P(a < X \le b) = P(a \le X < b).

27
New cards

What is a Probability Density Function (PDF), and how is probability calculated from it?

A PDF f(x)f(x) is a continuous curve satisfying f(x)≥0f(x) \ge 0 and total area under the curve equal to 1. Probability over an interval P(a≤X≤b)P(a \le X \le b) equals the area under f(x)f(x) from aa to bb. Curve height f(x)f(x) represents density, not point probability.

28
New cards

What parameters specify a Normal distribution X∼N(μ,σ2)X \sim N(\mu, \sigma^2), and how do they affect the curve?

A Normal distribution is specified by mean μ\mu (center) and variance σ2\sigma^2 (spread). Changing μ\mu shifts the curve horizontally without altering spread. Changing standard deviation σ\sigma alters spread; larger σ\sigma makes the curve wider and flatter.

29
New cards

What is a z-score, and how is it calculated for a population model X∼N(μ,σ2)X \sim N(\mu, \sigma^2)?

A z-score z=x−μσz = \frac{x - \mu}{\sigma} measures how many population standard deviations an observation xx lies above or below the population mean μ\mu. Standardization changes the scale to Z∼N(0,1)Z \sim N(0, 1) while preserving probability areas.

30
New cards

What is the 68–95–99.7 Rule for Normal distributions?

For a Normal distribution N(μ,σ2)N(\mu, \sigma^2), approximately 68%68\% of observations fall within μ±1σ\mu \pm 1\sigma, approximately 95%95\% fall within μ±2σ\mu \pm 2\sigma, and approximately 99.7%99.7\% fall within μ±3σ\mu \pm 3\sigma.

31
New cards

What do dnorm(), pnorm(), and qnorm() return in R for a Normal distribution?

dnorm(x,mean,sd)\text{dnorm}(x, \text{mean}, \text{sd}) returns the density curve height at xx (not probability); pnorm(q,mean,sd)\text{pnorm}(q, \text{mean}, \text{sd}) returns the cumulative left-tail probability P(X≤q)P(X \le q); and qnorm(p,mean,sd)\text{qnorm}(p, \text{mean}, \text{sd}) returns the value cutoff with cumulative left-tail probability pp.

32
New cards

What criteria make categories for a categorical variable well-defined?

Categories must be mutually exclusive (each case belongs to only one category) and exhaustive (every case can be placed into some category).

33
New cards

What is the difference between frequency, relative frequency, and sample proportion p^\hat{p}?

Frequency is the raw count xx of cases in a category. Relative frequency is that count expressed as a proportion of total cases xn\frac{x}{n}. The sample proportion p^=xn\hat{p} = \frac{x}{n} satisfies 0≤p^≤10 \le \hat{p} \le 1.

34
New cards

In a two-way contingency table, what are joint, marginal, and conditional proportions?

Joint proportions describe combinations of categories using the overall sample total nn as the denominator. Marginal proportions describe one variable's row or column total using nn as the denominator. Conditional proportions restrict attention to a specific row or column group, using that group's total as the denominator.

35
New cards

How should a difference between two sample proportions p^1−p^2=0.40\hat{p}_1 - \hat{p}_2 = 0.40 be reported and interpreted?

It should be reported as a difference of 0.400.40 or 40 percentage points40\text{ percentage points} (not a 40%40\% increase). It indicates the observed proportion in group 1 was 40 percentage points higher than in group 2.

36
New cards

What roles do explanatory and response variables play on a scatterplot?

The explanatory variable (XX) is placed on the horizontal axis and is used to predict or explain changes in the response variable (YY), which is placed on the vertical axis.

37
New cards

What four characteristics should be described when examining a scatterplot?

  1. Direction (positive or negative); 2. Form (linear or nonlinear); 3. Strength (how closely points follow the overall pattern); and 4. Unusual features (outliers, clusters, gaps, changing variability, or ranges with little data).
38
New cards

What are the key properties of the sample correlation coefficient rr?

rr measures direction and strength of linear association (−1≤r≤1-1 \le r \le 1). It is unitless, invariant to linear conversions of scale, symmetric (rX,Y=rY,Xr_{X,Y} = r_{Y,X}), and NOT resistant to outliers.

39
New cards

When is the sample correlation coefficient rr undefined?

Correlation is undefined if either variable has zero variability (no variation), such as when all points form a perfectly horizontal or vertical line.

40
New cards

In the simple linear regression model y^=b0+b1x\hat{y} = b_0 + b_1 x, how are slope b1b_1 and intercept b0b_0 interpreted?

Slope b1b_1 represents the predicted change in the response variable y^\hat{y} per 1-unit increase in the explanatory variable xx, on average. Intercept b0b_0 is the predicted response when x=0x = 0 (scientifically meaningful only if x=0x = 0 is within the observed data range).

41
New cards

How is a residual ee defined, and what does the least-squares regression line minimize?

A residual is e=y−y^e = y - \hat{y} (observed response minus predicted response). The least-squares regression line minimizes the sum of squared residuals (∑e2\sum e^2).

42
New cards

What is the difference between interpolation and extrapolation in regression prediction?

Interpolation is predicting a response for an xx-value within the range of observed explanatory data (well-supported). Extrapolation is predicting for an xx-value outside the observed xx range (unreliable, as the linear pattern may not continue).

43
New cards

What is the Coefficient of Determination (R2R^2), and how is it interpreted?

R2=r2R^2 = r^2 is the proportion of observed response variability that is accounted for by the fitted linear relationship with the explanatory variable xx.

44
New cards

What ideal pattern should be seen in a residual plot (y^\hat{y} vs ee) to confirm a linear model is plausible?

Residuals should show no clear pattern: points randomly scattered around the zero horizontal reference line with roughly equal vertical spread throughout, with no curves, fan shapes, or distinct clusters.

45
New cards

Why does a strong correlation or high R2R^2 between two variables not prove that XX causes YY?

Correlation and regression describe mathematical association in observed data. Establishing causation depends on study design (such as a well-designed randomized experiment with random assignment) rather than the magnitude of rr or R2R^2.