AP Stats Final

Here are the questions and answers from the Unit 1 tests, separated by document:

Unit 1 - Practice Test A-1 (5).pdf

Question 1: School administrators collect data on students attending the school. Which of the following variables is quantitative?

(A) class (freshman, sophomore, junior, senior)

(B) grade point average

(C) whether the student is in AP classes

(D) whether the student has taken the SAT

(E) none of these

Answer: (B) grade point average

Explanation: A quantitative variable is a variable that can be measured numerically. Grade point average is a numerical measurement, while the other options represent categories or groups.

Question 2: Which of the following variables would most likely follow a Normal model?

(A) weights of adult male elephants

(B) heights of singers in a co-ed choir

(C) family income

(D) scores on an easy test

(E) all of these

Answer: (A) weights of adult male elephants

Explanation: A Normal model is a bell-shaped and symmetrical distribution. Weights of adult male elephants are likely to follow a Normal distribution as they tend to cluster around a central value with some variation. The heights of singers in a co-ed choir are likely to have two peaks due to the differences in average heights between men and women. Family income is often skewed due to a small number of very high earners. Scores on an easy test may be skewed towards higher scores.

Question 3: A professor has kept records on grades that students have earned in his class. If he wants to examine the percentage of students earning the grades A, B, C, D, and F during the most recent term, which type of plot would be best to use?

(A) boxplot

(B) timeplot

(C) dotplot

(D) pie chart

(E) histogram

Answer: (D) pie chart

Explanation: A pie chart is the best choice for representing the percentage distribution of a categorical variable. It allows for a clear visual comparison of the proportions of each grade category.

Question 4: Which is true of the data shown in the histogram?

(A) I only

(B) II only

(C) I and II

(D) I and III

(E) II and III

Answer: (C) I and II

Explanation: The histogram shows a distribution that is roughly symmetrical, indicating that the mean and median are likely to be approximately equal.

Question 5: Two sections of a class took the same quiz. Section A had 15 students who had a mean score of 80, and Section B had 20 students who had a mean score of 90. Overall, what was the approximate mean score for all of the students on the quiz?

(A) 84

(B) 85

(C) 86

(D) the mean cannot be determined.

(E) none of these

Answer: (A) 84

Explanation: To find the approximate mean score for all students, you need to consider the weighted average of the two sections.

[(15 students 80 points) + (20 students 90 points)] / (15 students + 20 students) = approximately 84

Question 6: Your Stats teacher tells you your test score was the 3rd quartile for the class. Which is true?

(A) none of these

(B) I only

(C) II only

(D) III only

(E) II and III

Answer: (B) I only

Explanation: The 3rd quartile (Q3) represents the 75th percentile of a distribution. This means that you scored higher than 75% of the students in the class. You don't know what your exact score was, what the mean score was, or if the distribution is nearly Normal.

Question 7: Suppose that a Normal model described student scores in a history class. Parker has a standardized score (z-score) of +2.5. This means that Parker…

(A) is 2.5 standard deviations above average for the class.

(B) is 2.5 standard deviations above the mean for the class.

(C) has a score that is 2.5 times the average for the class.

(D) has a standard deviation of 2.5

(E) None of the above.

Answer: (A) is 2.5 standard deviations above average for the class.

Explanation: A z-score of +2.5 means that Parker's score is 2.5 standard deviations above the mean, which represents the average for the class.

Question 8: The advantage of making a stem-and-leaf display instead of a dotplot is that a stem-and-leaf display…

(A) presents the shape of the distribution better than a dotplot.

(B) shows the area principle.

(C) preserves the individual data values.

(D) A stem-and-leaf display is for quantitative data, while a dotplot shows categorical data

(E) None of these

Answer: (C) preserves the individual data values.

Explanation: A stem-and-leaf display retains the original data values, allowing you to see each individual data point. A dotplot can lose some detail, especially with large datasets, where multiple data points might be represented by a single dot.

Question 9: The five-number summary of credit hours for 24 students in a statistic class is:

MinQ1MedianQ3Max

13.0

15.0

16.5

18.0

22.0

Which statement is true?

(A) There are no outliers in the data.

(B) There is at least one high outlier in the data.

(C) There is at least one low outlier in the data.

(D) There are both low and high outliers in the data.

(E) None of the above.

Answer: (A) There are no outliers in the data.

Explanation: To determine if there are outliers, you need to calculate the interquartile range (IQR) and use the 1.5*IQR rule.

  • IQR = Q3 - Q1 = 18.0 - 15.0 = 3.0

  • 1.5 IQR = 1.5 3.0 = 4.5

  • Lower fence = Q1 - 1.5*IQR = 15.0 - 4.5 = 10.5

  • Upper fence = Q3 + 1.5*IQR = 18.0 + 4.5 = 22.5

Since all data points fall within the lower and upper fences (10.5 - 22.5), there are no outliers.

Question 10: Which of the following summaries are changed by adding a constant to each data value?

(A) I only

(B) III only

(C) I and II

(D) II and III

(E) I, II, and III

Answer: (C) I and II

Explanation:

  • The mean and median will both change if a constant is added to each data value. The mean will increase by the constant, and the median will also increase by the constant.

  • The standard deviation, however, is a measure of spread, and it remains unchanged when a constant is added to each data point.

Unit 1 - Practice Test B-1 (5).pdf

Question 1: The SPCA collects the following data about the dogs they house. Which is categorical?

(A) breed

(B) age

(C) weight

(D) number of days housed

(E) veterinary costs

Answer: (A) breed

Explanation: A categorical variable places individuals into categories or groups. Breed is a category, while age, weight, number of days housed, and veterinary costs are all numerical variables.

Question 2: Which of those variables about German Shepherds is most likely to be described by a Normal model?

(A) number of days housed

(B) age

(C) weight

(D) pie chart

(E) they want to show the trend

Answer: (C) weight

Explanation: A Normal model describes a symmetrical, bell-shaped distribution. The weight of German Shepherds is likely to follow a Normal distribution as it would cluster around a central value with some variation.

Question 3: The SPCA has kept these data records for the past 20 years. If they want to make the trend in the number of dogs they have housed, what kind of plot should they make?

(A) boxplot

(B) timeplot

(C) bar graph, D) pie chart

(E) histogram

Answer: (B) timeplot

Explanation: A timeplot displays data over time, showing any trends or patterns. This would be the appropriate plot to visualize how the number of dogs housed has changed over the past 20 years.

Question 4: The veterinary bills for the dogs are summarized in the ogive shown. Estimate the IQR of these expenses.

(A) $50

(B) $75

(C) $100

(D) $150

(E) $200

Answer: (C) $100

Explanation: The IQR is the difference between the third quartile (Q3) and the first quartile (Q1). An ogive represents a cumulative frequency graph.

  • Q1 is the value corresponding to 25% on the cumulative frequency axis, which is about $75 on the ogive.

  • Q3 is the value corresponding to 75% on the cumulative frequency axis, which is about $175 on the ogive.

Therefore, the IQR is approximately $175 - $75 = $100.

Question 5: Last weekend police ticketed 18 men whose mean speed was 72 miles per hour, and 30 women whose mean speed was 64 mph. Overall, what was the mean speed of all the people ticketed?

(A) none of these

(B) It cannot be determined.

(C) 69 mph

(D) 68 mph

(E) 67 mph

Answer: (D) 68 mph

Explanation: To calculate the overall mean speed, we need a weighted average based on the number of men and women ticketed.

[(18 men 72 mph) + (30 women 64 mph)] / (18 men + 30 women) = 68 mph

Question 6: Which is true of the data shown in the histogram?

(A) I only

(B) II only

(C) III only

(D) II and III only

(E) I, II, and III

Answer: (B) II only

Explanation: The histogram shows a distribution that is skewed to the right, which means it has a longer tail on the right side. In right-skewed distributions:

  • The mean is usually greater than the median.

  • The median is a better measure of center for skewed data because it is less influenced by extreme values than the mean.

Question 7: The best estimate of the standard deviation of the mens' weights displayed in this dotplot is

(A) 10

(B) 15

(C) 25

(D) 35

(E) 40

Answer: (B) 15

Explanation: Standard deviation is a measure of the spread of data. Observing the dot plot, most data points are within 15 units of the center, making 15 a reasonable estimate for the standard deviation.

Question 8: If we want to discuss any gaps and clusters in a data set, which of the following should not be chosen to display the data set?

(A) histogram

(B) stem-and-leaf

(C) dotplot

(D) boxplot

(E) any of these would work

Answer: (D) boxplot

Explanation: A boxplot summarizes data using the five-number summary and doesn't show individual data points, so gaps and clusters might not be evident. The other options (histogram, stem-and-leaf, and dotplot) all display the distribution of individual data points, making it easier to identify gaps and clusters.

Question 9: Suppose a Normal model describes the number of pages printer ink cartridges last. If we keep track of page counts for the ink cartridges at a company's office, which must be true?

(A) none

(B) I only

(C) II only

(D) II and III

(E) I, II, and III

Answer: (B) I only

Explanation: A Normal model describes a symmetrical, bell-shaped distribution. In a Normal distribution:

  • Approximately 68% of the data falls within one standard deviation of the mean, which is the basis for statement I.

  • Statement II refers to the Empirical Rule (68-95-99.7 rule), which applies to Normal distributions. However, this rule states that approximately 95% of data falls within two standard deviations of the mean, not three.

  • Statement III is incorrect. While a Normal distribution is symmetrical, there can still be some variation in the data, so not all page counts will be within 2 standard deviations of the mean.

Question 10: Suppose that a Normal model describes fuel economy (miles per gallon) for automobiles and that a Saturn has a standardized score (z-score) of +2.2. This means that Saturns …

(A) get 2.2 miles per gallon.

(B) get 2.2 times the gas mileage of the average car.

(C) get 2.2 mpg more than the average car.

(D) have a standard deviation of 2.2 mpg.

(E) achieve fuel economy that is 2.2 standard deviations better than the average car.

Answer: (E) achieve fuel economy that is 2.2 standard deviations better than the average car.

Explanation: A z-score measures how many standard deviations a data point is away from the mean. A positive z-score indicates a value above the mean. Therefore, a z-score of +2.2 for a Saturn means its fuel economy is 2.2 standard deviations better than the average fuel economy for all automobiles.

Unit 1 - Practice Test C-1 (2).pdf

Question 1: We collect these data from 50 male students. Which variable is categorical?

(A) eye color

(B) head circumference

(C) number of homework sets last week

(D) number of cigarettes smoked daily

(E) number of TV sets at home

Answer: (A) eye color

Explanation: A categorical variable places individuals into categories or groups. Eye color is a category, while the other options are all numerical variables.

Question 2: Which of those variables is most likely to be bimodal?

(A) eye color

(B) number of hours worked last week

(C) number of hours of homework last week

(D) number of cigarettes smoked daily

(E) number of TV sets at home

Answer: (B) number of hours worked last week

Explanation: A bimodal distribution has two peaks. The number of hours worked last week is most likely to be bimodal because there might be two groups of students: those who work part-time and those who work full-time, creating two peaks in the distribution.

Question 3: Which of those variables is most likely to follow a Normal model?

(A) eye color

(B) number of hours worked last week

(C) number of hours of homework last week

(D) number of cigarettes smoked daily

(E) number of TV sets at home

Answer: (B) number of hours worked last week

Explanation: A Normal model describes a symmetrical, bell-shaped distribution. The number of hours worked is likely to follow a Normal distribution, with most students working a moderate number of hours and fewer students working very few or very many hours.

Question 4: The mean number of hours worked for the 30 males was 6, and for the 20 females was 9. The overall mean number of hours worked is…

(A) 6.5

(B) 7.2

(C) 7.5

(D) none of these.

(E) cannot be determined.

Answer: (B) 7.2

Explanation: To find the overall mean, we need a weighted average based on the number of males and females:

[(30 males 6 hours) + (20 females 9 hours)] / (30 males + 20 females) = 7.2 hours

Question 5: We might choose to display data with a stemplot rather than a boxplot because a stemplot…

(A) I only

(B) II only

(C) III only

(D) I and III

(E) II and III

Answer: (E) II and III

Explanation:

  • A stemplot displays the shape of the data distribution, showing potential skewness, symmetry, and clusters.

  • It also reveals the actual data sets, preserving the individual data points, unlike a boxplot, which only shows summary statistics.

A stemplot may not be better at revealing large data sets, especially if the data range is wide, as it might become too cumbersome to display.

Question 6: Which is true of the data whose distribution is shown?

(A) I only

(B) II only

(C) III only

(D) I and II

(E) I, II, and III

Answer: (C) III only

Explanation:

  • The data is clearly skewed to the right, as the tail extends longer on the right side. In a right-skewed distribution:

    • The mean is usually greater than the median.

    • The median is a better measure of the center as it is less affected by extreme values.

Question 7: The standard deviation of the data displayed in standard deviation…

(A) I only

(B) II only

(C) I and II

(D) I and III

(E) I, II, and III

Answer: (C) I and II

Explanation:

  • The standard deviation measures the spread or variability of data. It is sensitive to outliers, meaning outliers can inflate the standard deviation.

  • A larger standard deviation indicates greater spread or variability in the data.

The standard deviation does not directly tell us about the normality of a data set.

Question 8: Suppose that a Normal model describes the acidity (pH) of rainwater, and that water tested after last week's storm had a z-score of 1.8. This means that the acidity of the rainwater…

(A) had a pH of 1.8.

(B) varied with a standard deviation of 1.8

(C) had a pH 1.8 higher than average

(D) had a pH that was 1.8 standard deviations higher than average rainfall

(E) had a pH 1.8 times that of average rainwater

Answer: (D) had a pH that was 1.8 standard deviations higher than average rainfall

Explanation: A z-score of 1.8 indicates that the rainwater's acidity (pH) was 1.8 standard deviations above the average pH of rainwater.

Question 9: The ages of people attending the opening show of a new movie are summarized in the ogive shown. Estimate the IQR of the ages.

(A) 13

(B) 21

(C) 30

(D) 37

(E) 49

Answer: (B) 21

Explanation: The IQR is the difference between the third quartile (Q3) and the first quartile (Q1).

  • Q1 (25th percentile) is about 15 on the ogive.

  • Q3 (75th percentile) is about 36 on the ogive.

Therefore, the IQR is approximately 36 - 15 = 21.

Question 10: Environmental researchers have attempted to reduce industrial air pollution. They want to see if there is any evidence that relates rain acidity (low pH for several ages. They produced a trend toward less acidic rainfall. They should display their data in an(n)…

(A) contingency table

(B) bar graph

(C) boxplot

(D) histogram

(E) timeplot

Answer: (E) timeplot

Explanation: A timeplot is the most appropriate display for visualizing changes in rain acidity over time. It will effectively show whether there is a trend toward less acidic rainfall.

Please let me know if you have any other questions!