Rucete ✏ AP Statistics In a Nutshell
2. Exploring Two-Variable Data — Practice Questions
This chapter introduces how to analyze relationships between two variables using two-way tables, scatterplots, correlation, regression, and residuals.
(Multiple Choice — Click to Reveal Answer)
1. A two-way table is primarily used to display
(A) one quantitative variable
(B) two categorical variables
(C) one categorical and one quantitative variable only
(D) two quantitative variables only
(E) time series data only
Answer
(B) — A two-way table displays the relationship between two categorical variables.
2. In a two-way table, the totals at the ends of rows and columns are called
(A) conditional frequencies
(B) residuals
(C) marginal frequencies
(D) leverage values
(E) regression coefficients
Answer
(C) — Row totals and column totals are marginal frequencies.
3. A conditional relative frequency is found by dividing
(A) a row total by the grand total
(B) a cell count by the grand total only
(C) a cell count by its row total or column total, depending on context
(D) a column total by the number of rows
(E) the grand total by a cell count
Answer
(C) — Conditional relative frequencies are based on a specific row or column category.
4. Segmented bar charts are most similar to
(A) histograms
(B) stem-and-leaf plots
(C) boxplots
(D) pie charts
(E) scatterplots
Answer
(D) — Segmented bar charts are like pie charts, but they use a rectangular bar instead of a circle.
5. If the conditional distributions differ across categories in a two-way table, the variables are said to be
(A) independent
(B) associated
(C) symmetric
(D) uniform
(E) linear
Answer
(B) — Different conditional distributions indicate an association between the variables.
6. A mosaic plot represents cell information mainly through
(A) point size only
(B) line slope only
(C) rectangle area
(D) circle radius only
(E) bar height only
Answer
(C) — In a mosaic plot, rectangle areas represent the frequencies or relative frequencies.
7. Simpson’s paradox occurs when
(A) a correlation is exactly zero
(B) a regression line has positive slope
(C) a pattern reverses when groups are combined
(D) two variables are perfectly associated
(E) an outlier is removed
Answer
(C) — Simpson’s paradox is when a conclusion reverses after combining several groups.
8. A scatterplot is used primarily to study the relationship between
(A) two categorical variables
(B) one categorical variable only
(C) two quantitative variables
(D) one quantitative variable only
(E) one row total and one column total
Answer
(C) — Scatterplots are used for two quantitative variables.
9. When larger values of one variable tend to be associated with larger values of another variable, the association is
(A) negative
(B) conditional
(C) segmented
(D) positive
(E) paradoxical
Answer
(D) — This is a positive association.
10. When larger values of one variable tend to be associated with smaller values of another variable, the association is
(A) positive
(B) negative
(C) marginal
(D) perfect
(E) categorical
Answer
(B) — This is a negative association.
11. When describing a scatterplot, which set of features should always be considered?
(A) Mean, median, range, mode
(B) Direction, unusual features, form, strength, and context
(C) Quartiles, IQR, variance, slope
(D) Row totals, column totals, percentages, and counts
(E) Shape, center, spread only
Answer
(B) — DUFS + context is the chapter’s guideline for scatterplots.
12. Correlation measures the strength and direction of a
(A) categorical relationship
(B) nonlinear relationship only
(C) linear relationship
(D) causal relationship
(E) segmented relationship
Answer
(C) — Correlation measures the strength and direction of a linear relationship.
13. Which statement is always true?
(A) Correlation implies causation
(B) A high correlation proves a cause-and-effect relationship
(C) Correlation does not imply causation
(D) Zero correlation means no relationship of any kind
(E) A negative correlation means no association
Answer
(C) — Correlation alone does not prove causation.
14. The value of a correlation coefficient r must lie between
(A) 0 and 1
(B) -1 and 1
(C) -100 and 100
(D) 0 and 100
(E) -10 and 10
Answer
(B) — Correlation always falls between -1 and 1.
15. If r is close to 1, the scatterplot shows
(A) strong negative linear association
(B) weak negative linear association
(C) strong positive linear association
(D) no association
(E) perfect nonlinear association
Answer
(C) — r close to 1 indicates strong positive linear association.
16. The coefficient of determination is
(A) r
(B) 1 − r
(C) r2
(D) √r
(E) 2r
Answer
(C) — The coefficient of determination is r2.
17. The coefficient of determination represents the percentage of variation in
(A) the x-variable explained by y only
(B) the y-variable explained by the linear model
(C) both variables equally, always
(D) the residuals explained by the line
(E) the slope explained by correlation
Answer
(B) — r2 is the percentage of variation in y explained by the linear model relating y and x.
18. In the least squares regression equation ŷ = a + bx, the slope b tells
(A) the predicted y when x = 1 only
(B) the amount predicted y changes for each 1-unit increase in x
(C) the value of r2
(D) the residual for the mean point
(E) the number of outliers
Answer
(B) — The slope is the change in predicted y for each one-unit increase in x.
19. In the least squares regression equation ŷ = a + bx, the y-intercept a is
(A) always meaningful
(B) the correlation coefficient
(C) the predicted y-value when x = 0
(D) the residual when x = 0
(E) the mean of x
Answer
(C) — The y-intercept is the predicted y-value when x = 0.
20. A residual is calculated as
(A) predicted minus observed
(B) observed minus predicted
(C) observed plus predicted
(D) y-intercept minus slope
(E) x minus y
Answer
(B) — Residual = observed − predicted.
21. If a point lies above the regression line, its residual is
(A) negative
(B) zero
(C) positive
(D) undefined
(E) equal to the slope
Answer
(C) — Observed is greater than predicted, so the residual is positive.
22. If a residual plot shows a clear curved pattern, this suggests that
(A) the linear model is definitely perfect
(B) a nonlinear model may fit better
(C) correlation must be zero
(D) the slope must be negative
(E) there are no outliers
Answer
(B) — A pattern in the residual plot suggests that a nonlinear model may be more appropriate.
23. A point with an x-value far from the mean of the x-values has
(A) no leverage
(B) high leverage
(C) zero residual automatically
(D) no chance of being influential
(E) no effect on correlation
Answer
(B) — Points far from the mean x-value have high leverage.
24. An influential point is one whose removal would substantially change the
(A) color of the scatterplot
(B) number of rows in a table
(C) regression results, such as slope, intercept, or correlation
(D) sample size only
(E) units of measurement
Answer
(C) — Influential points meaningfully change statistics such as slope, intercept, or correlation.
25. Which statement is correct about outliers and influential points?
(A) Every outlier is influential
(B) Every influential point is an outlier
(C) Outliers and influential points are always the same
(D) Outliers and high leverage points are often influential, but not always
(E) Neither can affect correlation
Answer
(D) — They are often influential, but not always.
26. In Example 2.1, 250 volunteers were studied. How many viewed baby animals?
(A) 50
(B) 65
(C) 85
(D) 95
(E) 100
Answer
(B) — The baby animals row total is 65.
27. In Example 2.1, what percentage of all participants had a high level of focus?
(A) 20%
(B) 24%
(C) 26%
(D) 34%
(E) 38%
Answer
(C) — 65 out of 250 had high focus, and 65/250 = 0.26 = 26%.
28. In Example 2.2, among those who viewed baby animals, what percentage had high focus?
(A) 7.7%
(B) 30.8%
(C) 38.5%
(D) 61.5%
(E) 76.9%
Answer
(D) — 40/65 = 0.615 = 61.5%.
29. In Example 2.2, among those who viewed adult animals, what percentage had medium focus?
(A) 17.6%
(B) 35.3%
(C) 40.0%
(D) 47.1%
(E) 55.0%
Answer
(D) — 40/85 = 0.471 = 47.1%.
30. In Example 2.2, among those who viewed tasty foods, what percentage had low focus?
(A) 10%
(B) 20%
(C) 35%
(D) 45%
(E) 55%
Answer
(E) — 55/100 = 55%.
31. In Example 2.5, Dr. Fixit’s overall survival rate was
(A) 68%
(B) 71.4%
(C) 76%
(D) 80%
(E) 88.2%
Answer
(C) — 190/250 = 0.76 = 76%.
32. In Example 2.5, Dr. Patch’s overall survival rate was
(A) 70.8%
(B) 76%
(C) 80%
(D) 87.6%
(E) 88.2%
Answer
(C) — 200/250 = 0.80 = 80%.
33. In Example 2.5, among patients in good condition, which surgeon had the higher survival rate?
(A) Dr. Patch
(B) Dr. Fixit
(C) They were equal
(D) Cannot be determined
(E) Neither surgeon had any survivors
Answer
(B) — Dr. Fixit had 60/68 = 88.2%, which was slightly higher than Dr. Patch’s 120/137 = 87.6%.
34. In Example 2.5, among patients in poor condition, which surgeon had the higher survival rate?
(A) Dr. Patch
(B) Dr. Fixit
(C) They were equal
(D) Cannot be determined
(E) Neither surgeon had any survivors
Answer
(B) — Dr. Fixit had 130/182 = 71.4%, slightly higher than Dr. Patch’s 80/113 = 70.8%.
35. Example 2.5 illustrates
(A) regression to the mean
(B) residual analysis
(C) Simpson’s paradox
(D) perfect correlation
(E) high leverage only
Answer
(C) — The overall comparison reverses the within-group comparisons.
36. In the high school survey table on page 10, what percentage of those surveyed were students?
Answer
50% — There were 250 students out of 500 total, so 250/500 = 50%.
37. In the same table, what percentage of all those surveyed were teachers who picked challenging?
Answer
25% — 125 out of 500 were teachers who picked challenging, so 125/500 = 25%.
38. In the same table, what percentage of administrators picked strict as most important?
Answer
50% — 25 out of 50 administrators picked strict, so 25/50 = 50%.
39. In the same table, what percentage of those picking enthusiastic were students?
Answer
71.4% — 150 of the 210 people who picked enthusiastic were students, so 150/210 = 0.714 ≈ 71.4%.
40. In the same table, which group was most likely to pick strict?
Answer
Administrators — Strict rates were 50/250 = 20% for students, 25/200 = 12.5% for teachers, and 25/50 = 50% for administrators.
41. In Example 2.9, if r = 0.84, find r2.
Answer
0.7056 — (0.84)2 = 0.7056.
42. In Example 2.9, interpret r2 as a percentage.
Answer
70.56% — 70.56% of the variation in Total Points Scored is explained by the linear relationship with Total Yards Gained.
43. In Example 2.11, the regression line was ŷ = -1.73 + 0.5492x. What is the predicted number of Facebook checks for x = 24 close friends?
Answer
11.45 — ŷ = -1.73 + 0.5492(24) = 11.45.
44. In Example 2.11, what does the slope 0.5492 mean in context?
Answer
For each additional close friend, the model predicts about 0.5492 more evening Facebook checks on average. — That is the contextual interpretation of the slope.
45. In Example 2.11, if r = 0.8836, what is r2 rounded to two decimals?
Answer
0.78 — (0.8836)2 ≈ 0.7807, which rounds to 0.78.
46. In Example 2.12, the regression line for calories burned on time was ŷ = 23.8 + 6.55x. Predict calories burned for x = 44 minutes.
Answer
312 — 23.8 + 6.55(44) = 312.0.
47. In Example 2.13, the regression line was ŷ = -0.03656 + 0.001582x. Predict the probability of dying for age 70.
Answer
0.074 — -0.03656 + 0.001582(70) = 0.07418, which rounds to 0.074.
48. In Example 2.15, the regression line was ŷ = 389.19 - 5.9776x. What is the predicted skin cancer mortality rate at latitude 39?
Answer
156.06 — 389.19 - 5.9776(39) = 156.0636, which rounds to 156.06.
49. In Example 2.15, the actual skin cancer mortality rate at latitude 39 was 162. Find the residual.
Answer
+5.94 — Residual = observed − predicted = 162 − 156.06 = +5.94.
50. In Example 2.15, the computer output gave R-Sq = 68.0%. What does this mean?
Answer
68.0% of the variation in skin cancer mortality rate is explained by the linear model relating skin cancer mortality rate and latitude. — This is the contextual meaning of R-Sq.
