1/107
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Individual (in statistics)
An object or person described by a data set.
Variable
A characteristic measured or recorded for individuals.
Categorical variable
A variable that places individuals into groups or categories.
Quantitative variable
A numerical variable for which arithmetic operations make sense.
Discrete quantitative variable
A numerical variable with countable possible values.
Continuous quantitative variable
A numerical variable that can take any value within a given interval.
Distribution
The values a variable takes and how often they occur.
Frequency
The number of times a specific value or category occurs in a data set.
Relative frequency
Frequency divided by the total number of observations.
Cumulative relative frequency
The proportion of observations at or below a given value.
Type of data displayed in a bar graph
Categorical data.
Type of data displayed in a pie chart
Categorical data shown as parts of a whole.
Total sum of all slices in a pie chart
100%
Type of data displayed in a histogram
Quantitative data grouped into continuous intervals.
Do bars touch in a histogram?
Yes, because the intervals represent a continuous numerical scale.
Do bars touch in a bar graph?
No, because the categories are separate and distinct.
Key difference between a histogram and a bar graph
Histograms display quantitative data across intervals, while bar graphs display categorical data.
Bin (in a histogram)
An interval used to group quantitative observations.
Effect of changing histogram bin width
It can change the apparent shape and display of the distribution.
Dotplot
A display showing each individual observation as a dot above its corresponding value on a number line.
Main advantage of a dotplot
Individual observations and unusual values remain clearly visible.
Primary feature preserved by a stem-and-leaf plot
The original quantitative data values.
In the stemplot key 4∣7=47, what is the stem?
4
In the stemplot key 4∣7=47, what is the leaf?
7
Five-number summary
A summary consisting of Minimum, Q1, Median, Q3, and Maximum.
Can the exact mean be determined from a boxplot?
Usually no, because individual data values are not preserved.
CUSS acronym (describing quantitative distributions)
Center, Unusual features, Shape, Spread.
When to use CUSS
When asked to describe the distribution of a quantitative variable.
Center (in CUSS)
A typical or middle value, such as the mean or median.
Unusual features (in CUSS)
Outliers, gaps, clusters, or other notable patterns.
Shape (in CUSS)
The overall form of a distribution (e.g., symmetric, skewed, unimodal, bimodal, uniform).
Spread (in CUSS)
The variability of the data, measured by range, IQR, or standard deviation.
Symmetric distribution
A distribution whose two sides are approximately mirror images of each other.
Relationship between mean and median in a symmetric distribution
They are approximately equal (mean≈median).
Right-skewed distribution
A distribution with a long tail extending to the right.
Relationship between mean and median in a right-skewed distribution
The mean is usually greater than the median (mean>median).
Left-skewed distribution
A distribution with a long tail extending to the left.
Relationship between mean and median in a left-skewed distribution
The mean is usually less than the median (mean<median).
Direction the mean is pulled in skewed distributions
Toward the long tail.
Unimodal distribution
A distribution with one main peak.
Bimodal distribution
A distribution with two main peaks.
Multimodal distribution
A distribution with multiple main peaks.
Uniform distribution
A distribution where values occur with approximately equal frequency across their range.
Cluster
A group of observations concentrated close together.
Gap
An interval containing few or no observations between data points.
Outlier
An observation located unusually far from the rest of the data.
Mean
The sum of all observations divided by the total number of observations.
Median
The middle observation when data values are arranged in order.
Finding the median with an even number of observations
Calculate the average of the two middle observations.
Mode
The value that occurs most frequently in a data set.
Can a data set have more than one mode?
Yes, a data set can be bimodal or multimodal.
Range formula
Range=Maximum−Minimum
First quartile (Q1)
The value with about 25% of observations at or below it.
Second quartile (Q2)
The median; the value with about 50% of observations at or below it.
Third quartile (Q3)
The value with about 75% of observations at or below it.
Interquartile Range formula (IQR)
IQR=Q3−Q1
Portion of data described by IQR
The spread of the middle 50% of the observations.
Standard deviation
A measure of the typical distance of observations from the mean.
Indication of a larger standard deviation
Greater variability or spread around the mean.
Can standard deviation be negative?
No, it is always greater than or equal to zero.
Condition when standard deviation equals zero
When every observation in the data set has the exact same value.
Variance
The square of the standard deviation (s2 or σ2).
Resistant statistic
A summary statistic that is not heavily influenced by extreme values or outliers.
Resistant measures of center and spread
Median and IQR.
Non-resistant measures of center and spread
Mean and standard deviation.
Best measures of center and spread for symmetric data without outliers
Mean and standard deviation.
Best measures of center and spread for skewed data or data with outliers
Median and IQR.
Lower outlier fence formula
Q1−1.5×IQR
Upper outlier fence formula
Q3+1.5×IQR
Condition for a low potential outlier
A value that falls below Q1−1.5×IQR.
Condition for a high potential outlier
A value that exceeds Q3+1.5×IQR.
Are outlier fences required to be actual data values?
No, they are calculated cutoff threshold values.
If Q1=10 and Q3=30, what is IQR?
20 (since 30−10=20)
If IQR=20, what is 1.5×IQR?
30
If Q1=10 and IQR=20, what is the lower outlier fence?
−20 (calculated as 10−30=−20)
If Q3=30 and IQR=20, what is the upper outlier fence?
60 (calculated as 30+30=60)
z-score
A measure of how many standard deviations an observation lies from the mean.
Meaning of a positive z-score
The observation is above the mean.
Meaning of a negative z-score
The observation is below the mean.
Meaning of z=0
The observation is equal to the mean.
Meaning of z=2
The observation is 2 standard deviations above the mean.
Meaning of z=−1.5
The observation is 1.5 standard deviations below the mean.
Usefulness of z-scores in data comparison
They allow comparison of relative positions across different distributions or scales.
If mean=50, SD=10, and x=70, what is z?
2 (calculated as 1070−50=2)
If mean=100, SD=20, and x=80, what is z?
−1 (calculated as 2080−100=−1)
80th percentile
The value such that about 80% of observations fall at or below it.
Does scoring at the 80th percentile mean getting 80% correct?
No, it describes relative position within a group, not raw percent correct.
Density curve
A mathematical model describing the overall pattern of a continuous distribution.
Total area under a density curve
1 (or 100%)
Effect of the median on the area under a density curve
It divides the total area under the curve into two equal halves of 0.5 each.
Normal distribution
A symmetric, unimodal, bell-shaped continuous probability distribution.
Parameter determining the center of a Normal distribution
The mean (μ).
Parameter determining the spread of a Normal distribution
The standard deviation (σ).
68-95-99.7 Empirical Rule
In a Normal distribution, approximately 68%, 95%, and 99.7% of observations fall within 1, 2, and 3 standard deviations of the mean.
Percent of data within 1 standard deviation of the mean in a Normal distribution
Approximately 68%
Percent of data within 2 standard deviations of the mean in a Normal distribution
Approximately 95%
Percent of data within 3 standard deviations of the mean in a Normal distribution
Approximately 99.7%
Percent of data lying between the mean and 1σ above the mean in a Normal distribution
Approximately 34%
Percent of data lying between 1σ and 2σ above the mean in a Normal distribution
Approximately 13.5%
Effect on spread measures when adding a constant to every value
Spread measures (range, IQR, standard deviation) remain unchanged.