1/44
Flashcards covering Lectures 02, 04, 05, 06, and 07 on descriptive statistics, probability foundations, discrete distributions, continuous distributions, and scatterplots/regression.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
What is the primary difference between descriptive statistics and model probability?
Descriptive statistics describe what happened in the observed sample data (such as the observed proportion p^), whereas model probability describes what results would be plausible under a specified chance model (such as p=0.20 for uniform guessing).
What is a random process in probability theory?
A random process is a process whose outcome is uncertain before it occurs, even when the underlying conditions are specified.
How are a sample space and an event defined?
A sample space (S) is the set of all possible outcomes of a random process. An event is a collection of one or more outcomes from the sample space.
What are the fundamental probability rules for any event A and sample space S in a finite setting?
For any event A, 0≤P(A)≤1, and for the sample space S, P(S)=1. Under finite sample spaces, P(A)=0 means the event is impossible, and P(A)=1 means it is certain.
How is a probability interpreted under the frequentist perspective?
A probability can be interpreted as a long-term relative frequency of an outcome over many comparable repetitions of a random process.
How is the probability of an event A calculated when a finite sample space S has equally likely outcomes?
P(A)=number of outcomes in Snumber of outcomes in A. Equal likelihood must be justified by the model and does not apply automatically.
What is the Complement Rule, and when is it particularly useful?
The Complement Rule states that P(Ac)=1−P(A), where Ac is the complement of A (the event that A does not occur). It is particularly useful when calculating "at least one" probabilities by using "none."
What is the General Addition Rule, and why is the intersection subtracted?
The General Addition Rule states P(A or B)=P(A)+P(B)−P(A and B). The intersection P(A and B) is subtracted because outcomes contained in both events were counted twice when adding P(A) and P(B).
What are disjoint (mutually exclusive) events, and what is their addition rule?
Disjoint events cannot occur together, meaning P(A and B)=0. For disjoint events, the addition rule simplifies to P(A or B)=P(A)+P(B).
How is conditional probability P(A∣B) defined, and what determines its denominator?
Conditional probability is defined as P(A∣B)=P(B)P(A and B) for P(B)>0. The event following the word "given" (event B) determines the restricted sample space and denominator.
Why does the direction of conditioning matter when comparing P(A∣B) and P(B∣A)?
P(A∣B) and P(B∣A) generally differ because they restrict attention to different groups, resulting in different denominators even though their numerators P(A and B) are identical.
What is the General Multiplication Rule for any two sequential events A and B?
P(A and B)=P(A)⋅P(B∣A), which multiplies the probability of the first event by the probability of the second event given that the first has occurred.
What does it mean for two events A and B to be independent?
Events A and B are independent if knowing that one occurred does not change the probability of the other occurring, satisfying P(A∣B)=P(A), P(B∣A)=P(B), and P(A and B)=P(A)P(B).
Why can two disjoint events with positive probabilities never be independent?
If event A occurs, a disjoint event B becomes impossible (P(B∣A)=0), which changes the probability of B from its original positive value P(B). Therefore, disjoint events with positive probabilities are always dependent.
What is a random variable, and how do uppercase and lowercase letters represent it?
A random variable is a rule that assigns one numerical value to each possible outcome of a random process. Uppercase X denotes the random variable (uncertain before observation), while lowercase x denotes one realized or observed numerical value.
How do discrete and continuous random variables differ?
Discrete random variables have possible values that can be listed or counted (such as counts or indicators), whereas continuous random variables can take any value within a numerical interval (such as measurements).
What is a Probability Mass Function (PMF), and what conditions must it satisfy?
A PMF pX(x)=P(X=x) gives the probability of each exact value for a discrete random variable X. It must satisfy pX(x)≥0 for all x, ∑pX(x)=1, and pX(x)=0 for impossible values.
What defines a Bernoulli random variable?
A Bernoulli random variable X∼Bernoulli(p) records a single binary outcome where X=1 if the event of interest ("success") occurs with probability p, and X=0 with probability 1−p.
What are the four BINS conditions required for a Binomial distribution?
BINS stands for: Binary outcomes (success or failure on each trial), Independent trials, Number of trials fixed in advance (n), and Same success probability (p) on every trial.
What is the formula for exact probabilities of a Binomial random variable X∼Binomial(n,p)?
P(X=k)=(kn)pk(1−p)n−k, where (kn)=k!(n−k)!n! is the binomial coefficient counting combinations of k successes in n trials.
In R, what are the distinct functions dbinom(), pbinom(), and qbinom() used for?
dbinom(x,size,prob) calculates exact point probability P(X=x); pbinom(q,size,prob) calculates cumulative left-tail probability P(X≤q); and qbinom(p,size,prob) calculates the smallest count corresponding to a cumulative probability p.
How do you express P(X≥16) and P(14≤X≤17) using R commands for X∼Binomial(20,0.75)?
P(X≥16)=1−P(X≤15), written as 1−pbinom(15,20,0.75). P(14≤X≤17)=P(X≤17)−P(X≤13), written as pbinom(17,20,0.75)−pbinom(13,20,0.75).
What are the mean, variance, and standard deviation formulas for Bernoulli and Binomial distributions?
For Bernoulli(p): μ=p, σ2=p(1−p), σ=p(1−p). For Binomial(n,p): μ=np, σ2=np(1−p), σ=np(1−p).
How is the expected value E(X) of a discrete random variable defined and interpreted?
E(X)=μX=∑x⋅P(X=x). It represents the long-run average center of the probability distribution over many repetitions and does not need to be a value X can actually take.
Why is the probability of a continuous random variable taking any exact single point equal to zero (P(X=x)=0)?
Continuous variables model exact measurements where an exact point has zero area under the probability density curve. Recorded measurements like 120mmHg are rounded numbers representing a small interval (119.5≤X<120.5).
Why does including or excluding interval endpoints not change probabilities for continuous random variables?
Because P(X=a)=0 and P(X=b)=0 for a continuous random variable X, P(a<X<b)=P(a≤X≤b)=P(a<X≤b)=P(a≤X<b).
What is a Probability Density Function (PDF), and how is probability calculated from it?
A PDF f(x) is a continuous curve satisfying f(x)≥0 and total area under the curve equal to 1. Probability over an interval P(a≤X≤b) equals the area under f(x) from a to b. Curve height f(x) represents density, not point probability.
What parameters specify a Normal distribution X∼N(μ,σ2), and how do they affect the curve?
A Normal distribution is specified by mean μ (center) and variance σ2 (spread). Changing μ shifts the curve horizontally without altering spread. Changing standard deviation σ alters spread; larger σ makes the curve wider and flatter.
What is a z-score, and how is it calculated for a population model X∼N(μ,σ2)?
A z-score z=σx−μ measures how many population standard deviations an observation x lies above or below the population mean μ. Standardization changes the scale to Z∼N(0,1) while preserving probability areas.
What is the 68–95–99.7 Rule for Normal distributions?
For a Normal distribution N(μ,σ2), approximately 68% of observations fall within μ±1σ, approximately 95% fall within μ±2σ, and approximately 99.7% fall within μ±3σ.
What do dnorm(), pnorm(), and qnorm() return in R for a Normal distribution?
dnorm(x,mean,sd) returns the density curve height at x (not probability); pnorm(q,mean,sd) returns the cumulative left-tail probability P(X≤q); and qnorm(p,mean,sd) returns the value cutoff with cumulative left-tail probability p.
What criteria make categories for a categorical variable well-defined?
Categories must be mutually exclusive (each case belongs to only one category) and exhaustive (every case can be placed into some category).
What is the difference between frequency, relative frequency, and sample proportion p^?
Frequency is the raw count x of cases in a category. Relative frequency is that count expressed as a proportion of total cases nx. The sample proportion p^=nx satisfies 0≤p^≤1.
In a two-way contingency table, what are joint, marginal, and conditional proportions?
Joint proportions describe combinations of categories using the overall sample total n as the denominator. Marginal proportions describe one variable's row or column total using n as the denominator. Conditional proportions restrict attention to a specific row or column group, using that group's total as the denominator.
How should a difference between two sample proportions p^1−p^2=0.40 be reported and interpreted?
It should be reported as a difference of 0.40 or 40 percentage points (not a 40% increase). It indicates the observed proportion in group 1 was 40 percentage points higher than in group 2.
What roles do explanatory and response variables play on a scatterplot?
The explanatory variable (X) is placed on the horizontal axis and is used to predict or explain changes in the response variable (Y), which is placed on the vertical axis.
What four characteristics should be described when examining a scatterplot?
What are the key properties of the sample correlation coefficient r?
r measures direction and strength of linear association (−1≤r≤1). It is unitless, invariant to linear conversions of scale, symmetric (rX,Y=rY,X), and NOT resistant to outliers.
When is the sample correlation coefficient r undefined?
Correlation is undefined if either variable has zero variability (no variation), such as when all points form a perfectly horizontal or vertical line.
In the simple linear regression model y^=b0+b1x, how are slope b1 and intercept b0 interpreted?
Slope b1 represents the predicted change in the response variable y^ per 1-unit increase in the explanatory variable x, on average. Intercept b0 is the predicted response when x=0 (scientifically meaningful only if x=0 is within the observed data range).
How is a residual e defined, and what does the least-squares regression line minimize?
A residual is e=y−y^ (observed response minus predicted response). The least-squares regression line minimizes the sum of squared residuals (∑e2).
What is the difference between interpolation and extrapolation in regression prediction?
Interpolation is predicting a response for an x-value within the range of observed explanatory data (well-supported). Extrapolation is predicting for an x-value outside the observed x range (unreliable, as the linear pattern may not continue).
What is the Coefficient of Determination (R2), and how is it interpreted?
R2=r2 is the proportion of observed response variability that is accounted for by the fitted linear relationship with the explanatory variable x.
What ideal pattern should be seen in a residual plot (y^ vs e) to confirm a linear model is plausible?
Residuals should show no clear pattern: points randomly scattered around the zero horizontal reference line with roughly equal vertical spread throughout, with no curves, fan shapes, or distinct clusters.
Why does a strong correlation or high R2 between two variables not prove that X causes Y?
Correlation and regression describe mathematical association in observed data. Establishing causation depends on study design (such as a well-designed randomized experiment with random assignment) rather than the magnitude of r or R2.