Rucete ✏ AP Statistics In a Nutshell
2. Exploring Two-Variable Data — Practice Questions 3
This chapter introduces new practice questions on two-way tables, association, scatterplots, correlation, regression, residuals, leverage, influence, and transformations. :contentReference[oaicite:0]{index=0}
(Multiple Choice — Click to Reveal Answer)
1. Which display is most appropriate for two categorical variables?
(A) Histogram
(B) Boxplot
(C) Two-way table
(D) Stem-and-leaf plot
(E) Dotplot
Answer
(C) — A two-way table is used to organize two categorical variables. :contentReference[oaicite:1]{index=1}
2. The totals at the bottom of columns and ends of rows in a two-way table are called
(A) conditional frequencies
(B) marginal frequencies
(C) residuals
(D) leverage values
(E) z-scores
Answer
(B) — These totals are marginal frequencies. :contentReference[oaicite:2]{index=2}
3. A conditional relative frequency for a row is found by dividing
(A) each cell in the row by the grand total
(B) each cell in the row by the row total
(C) each row total by the grand total
(D) each cell in the row by the column total
(E) the row total by the number of columns
Answer
(B) — Row conditional relative frequencies divide each cell by its row total. :contentReference[oaicite:3]{index=3}
4. If the conditional distributions are the same across groups, the variables are most likely
(A) associated
(B) influential
(C) independent
(D) nonlinear
(E) residual
Answer
(C) — Similar conditional distributions suggest little or no association. :contentReference[oaicite:4]{index=4}
5. A mosaic plot differs from a segmented bar chart because a mosaic plot also shows differences in
(A) color only
(B) rectangle width
(C) line slope
(D) residual size
(E) correlation
Answer
(B) — In a mosaic plot, widths represent group sizes. :contentReference[oaicite:5]{index=5}
6. Simpson’s paradox refers to a situation in which
(A) a regression slope is zero
(B) the same graph can be drawn two ways
(C) a trend in separate groups reverses when groups are combined
(D) residuals add up to 100%
(E) correlation equals causation
Answer
(C) — That is the defining feature of Simpson’s paradox. :contentReference[oaicite:6]{index=6}
7. A scatterplot is used to study the relationship between
(A) two quantitative variables
(B) two categorical variables only
(C) one variable and one frequency table only
(D) one row total and one column total
(E) one qualitative variable only
Answer
(A) — Scatterplots display two quantitative variables. :contentReference[oaicite:7]{index=7}
8. When describing a scatterplot, DUFS stands for
(A) direction, unusual features, form, strength
(B) distribution, units, frequency, slope
(C) direction, uniformity, frequency, spread
(D) deviation, units, fit, scale
(E) data, units, form, symmetry
Answer
(A) — DUFS is the recommended checklist for scatterplots. :contentReference[oaicite:8]{index=8}
9. A scatterplot with points close to an upward-sloping line shows
(A) weak negative association
(B) strong positive association
(C) no association
(D) perfect nonlinear association
(E) two-way categorical association
Answer
(B) — Upward slope with tight clustering indicates strong positive association. :contentReference[oaicite:9]{index=9}
10. Correlation measures the strength of a
(A) nonlinear relationship only
(B) categorical relationship only
(C) linear relationship
(D) cause-and-effect relationship
(E) segmented bar chart relationship
Answer
(C) — Correlation is for linear relationships. :contentReference[oaicite:10]{index=10}
11. Which statement is true?
(A) Correlation proves causation
(B) Correlation does not imply causation
(C) A large r means x causes y
(D) A negative r means no association
(E) r can be greater than 1
Answer
(B) — Correlation alone never proves causation. :contentReference[oaicite:11]{index=11}
12. The possible values of the correlation coefficient r are
(A) from 0 to 1
(B) from -1 to 1
(C) from -100% to 100%
(D) any real number
(E) from 1 to 10
Answer
(B) — r always lies between -1 and 1. :contentReference[oaicite:12]{index=12}
13. If the explanatory and response variables are swapped, the value of r
(A) changes sign
(B) becomes 0
(C) stays the same
(D) becomes r²
(E) doubles
Answer
(C) — Correlation is unchanged when x and y are interchanged. :contentReference[oaicite:13]{index=13}
14. Changing the units of measurement for x or y will
(A) always change r
(B) never change r
(C) always make r positive
(D) always reduce r
(E) make r equal 1
Answer
(B) — Correlation is unit-free. :contentReference[oaicite:14]{index=14}
15. The coefficient of determination is
(A) r
(B) 2r
(C) r²
(D) √r
(E) 1 − r
Answer
(C) — The coefficient of determination is r². :contentReference[oaicite:15]{index=15}
16. If r = -0.60, then r² =
(A) -0.36
(B) 0.36
(C) 0.60
(D) 1.20
(E) -0.60
Answer
(B) — Squaring gives 0.36. :contentReference[oaicite:16]{index=16}
17. In context, r² is interpreted as the percentage of variation in
(A) x explained by x
(B) y explained by the linear model using x
(C) residuals explained by y
(D) x explained by residuals
(E) slope explained by intercept
Answer
(B) — r² describes the explained variation in the response variable. :contentReference[oaicite:17]{index=17}
18. In the least squares regression equation ŷ = a + bx, b is the
(A) correlation
(B) residual
(C) slope
(D) standard deviation of residuals
(E) coefficient of determination
Answer
(C) — b is the slope of the regression line. :contentReference[oaicite:18]{index=18}
19. In the least squares regression equation ŷ = a + bx, a is the
(A) slope
(B) y-intercept
(C) residual
(D) correlation
(E) x-mean
Answer
(B) — a is the y-intercept. :contentReference[oaicite:19]{index=19}
20. A least squares regression line always passes through
(A) (0, 0)
(B) (1, 1)
(C) (x̄, ȳ)
(D) (median x, median y)
(E) the point with the highest leverage
Answer
(C) — The regression line passes through the point of means. :contentReference[oaicite:20]{index=20}
21. A residual is calculated as
(A) predicted − observed
(B) observed − predicted
(C) x − y
(D) y − x̄
(E) slope − intercept
Answer
(B) — Residual equals observed minus predicted. :contentReference[oaicite:21]{index=21}
22. If a point lies above the regression line, its residual is
(A) negative
(B) positive
(C) zero
(D) undefined
(E) always an outlier
Answer
(B) — Above the line means observed is greater than predicted. :contentReference[oaicite:22]{index=22}
23. If a residual is negative, the model has
(A) underestimated the response
(B) overestimated the response
(C) no error
(D) a positive slope
(E) a perfect fit
Answer
(B) — Negative residual means observed is less than predicted, so the model overestimated. :contentReference[oaicite:23]{index=23}
24. A residual plot with a clear pattern suggests that
(A) a linear model is definitely best
(B) a nonlinear model may fit better
(C) correlation must be 1
(D) the response variable is categorical
(E) there are no unusual points
Answer
(B) — A visible pattern in the residual plot suggests nonlinearity. :contentReference[oaicite:24]{index=24}
25. A point has high leverage when its x-value is
(A) close to the mean of x-values
(B) far from the mean of x-values
(C) equal to the median of y-values
(D) above the regression line
(E) below the regression line
Answer
(B) — High leverage comes from an extreme x-value. :contentReference[oaicite:25]{index=25}
26. In the Cuteness Factor table, what fraction of all participants viewed adult animals?
(A) 65/250
(B) 85/250
(C) 90/250
(D) 95/250
(E) 100/250
Answer
(B) — The adult animals row total is 85 out of 250. :contentReference[oaicite:26]{index=26}
27. In the Cuteness Factor table, what percent of all participants had medium focus?
(A) 26%
(B) 34%
(C) 36%
(D) 38%
(E) 40%
Answer
(D) — 95 out of 250 had medium focus, so 95/250 = 38%. :contentReference[oaicite:27]{index=27}
28. Among those who viewed baby animals, what percent had low focus?
(A) 7.7%
(B) 17.6%
(C) 30.8%
(D) 35.3%
(E) 61.5%
Answer
(A) — 5/65 = 0.077 = 7.7%. :contentReference[oaicite:28]{index=28}
29. Among those who viewed adult animals, what percent had low focus?
(A) 17.6%
(B) 35.3%
(C) 41.7%
(D) 47.1%
(E) 55%
Answer
(B) — 30/85 = 0.353 ≈ 35.3%. :contentReference[oaicite:29]{index=29}
30. Among those who viewed tasty foods, what percent had high focus?
(A) 7.7%
(B) 10%
(C) 17.6%
(D) 30.8%
(E) 35%
Answer
(B) — 10/100 = 10%. :contentReference[oaicite:30]{index=30}
31. In the surgeon example, Dr. Fixit’s overall survival rate was
(A) 70.8%
(B) 71.4%
(C) 76%
(D) 80%
(E) 88.2%
Answer
(C) — 190 out of 250 survived, so 76%. :contentReference[oaicite:31]{index=31}
32. In the same example, Dr. Patch’s survival rate among patients in good condition was approximately
(A) 70.8%
(B) 71.4%
(C) 76.0%
(D) 87.6%
(E) 88.2%
Answer
(D) — 120/137 ≈ 87.6%. :contentReference[oaicite:32]{index=32}
33. In the same example, Dr. Fixit’s survival rate among patients in poor condition was approximately
(A) 70.8%
(B) 71.4%
(C) 76.0%
(D) 87.6%
(E) 88.2%
Answer
(B) — 130/182 ≈ 71.4%. :contentReference[oaicite:33]{index=33}
34. The surgeon example is important because it shows that
(A) overall rates always tell the full story
(B) grouped data can hide the effect of another variable
(C) regression is better than tables
(D) residuals cause paradoxes
(E) conditional percentages are never useful
Answer
(B) — Combining groups can hide an important third-variable effect. :contentReference[oaicite:34]{index=34}
35. In the teacher-characteristics table, what percent of the 500 people surveyed were teachers who chose enthusiastic?
(A) 5%
(B) 10%
(C) 15%
(D) 20%
(E) 25%
Answer
(B) — 50/500 = 10%. :contentReference[oaicite:35]{index=35}
36. In the same table, what percent of teachers chose challenging?
(A) 25%
(B) 37.5%
(C) 50%
(D) 62.5%
(E) 71.4%
Answer
(D) — 125/200 = 62.5%. :contentReference[oaicite:36]{index=36}
37. In the same table, what percent of students chose enthusiastic?
(A) 20%
(B) 30%
(C) 40%
(D) 50%
(E) 60%
Answer
(E) — 150/250 = 60%. :contentReference[oaicite:37]{index=37}
38. In Example 2.9, if r = 0.84, what percent of the variation in y is explained by the linear model?
(A) 29.44%
(B) 70.56%
(C) 84.00%
(D) 91.00%
(E) 94.30%
Answer
(B) — r² = 0.7056, or 70.56%. :contentReference[oaicite:38]{index=38}
39. In Example 2.11, the regression line is ŷ = -1.73 + 0.5492x. Predict the number of Facebook checks for x = 30 close friends.
Answer
14.75 — ŷ = -1.73 + 0.5492(30) = -1.73 + 16.476 = 14.746 ≈ 14.75. :contentReference[oaicite:39]{index=39}
40. In Example 2.11, if a student with 30 close friends actually checks Facebook 15 times, what is the residual?
Answer
+0.25 — Residual = observed − predicted = 15 - 14.75 ≈ +0.25. :contentReference[oaicite:40]{index=40}
41. In Example 2.11, if a student with 20 close friends actually checks Facebook 8 times, use the regression line to find the residual.
Answer
-1.25 — Predicted = -1.73 + 0.5492(20) = 9.254 ≈ 9.25, so residual = 8 - 9.25 = -1.25. :contentReference[oaicite:41]{index=41}
42. In Example 2.12, using ŷ = 23.8 + 6.55x, predict calories burned for 60 minutes.
Answer
416.8 — 23.8 + 6.55(60) = 23.8 + 393 = 416.8. :contentReference[oaicite:42]{index=42}
43. In Example 2.12, using ŷ = -0.829 + 0.1465x for minutes from calories, predict the number of minutes needed to burn 500 calories.
Answer
72.42 minutes — -0.829 + 0.1465(500) = -0.829 + 73.25 = 72.421. :contentReference[oaicite:43]{index=43}
44. In Example 2.13, using ŷ = -0.03656 + 0.001582x, predict the probability of dying for age 80.
Answer
0.0900 — -0.03656 + 0.001582(80) = -0.03656 + 0.12656 = 0.09000. :contentReference[oaicite:44]{index=44}
45. In Example 2.15, using ŷ = 389.19 - 5.9776x, predict the skin cancer mortality rate at latitude 40.
Answer
150.09 — 389.19 - 5.9776(40) = 389.19 - 239.104 = 150.086 ≈ 150.09. :contentReference[oaicite:45]{index=45}
46. In Example 2.15, if the observed rate at latitude 40 were 145, what would the residual be?
Answer
-5.09 — Residual = 145 - 150.09 = -5.09. :contentReference[oaicite:46]{index=46}
47. In Example 2.15, if R-Sq = 68.0% and the slope is negative, what is r to three decimal places?
Answer
-0.825 — r = -√0.68 ≈ -0.8246, so about -0.825. :contentReference[oaicite:47]{index=47}
48. In Example 2.16, which point is not a regression outlier even though its x-value and y-value may each look unusual by themselves?
Answer
(30, 0.5) — It is not a regression outlier because it still follows the overall linear pattern. :contentReference[oaicite:48]{index=48}
49. In Example 2.17, which point is influential?
(A) Point B only
(B) Point A only
(C) Both A and B equally
(D) Neither A nor B
(E) The graph does not say
Answer
(B) — Removing A changes the regression line a lot, so A is influential. :contentReference[oaicite:49]{index=49}
50. According to Example 2.21, what two signs suggest that a transformed model may be better than the original linear model?
Answer
A more random residual plot and/or an increase in r². — Either of these supports the transformed linear model as more appropriate. :contentReference[oaicite:50]{index=50}
