DS 2003 · Communicating with Data · Class Activity
Simpson's Paradox: Does School Spending Hurt Test Scores?
In 1994, columnist George Will looked at exactly this data — state-level spending per student vs. average SAT score — and concluded that pouring more money into schools doesn't help, and might even hurt. He wasn't lying about the trend line. He was missing a variable.
Data: per-pupil expenditure, teacher salary, and SAT scores for all 50 U.S. states,
1994–95 school year (National Center for Education Statistics / College Board, via
Deborah Guber, "Getting What You Pay For: The Debate Over Equity in Public School
Expenditures," Journal of Statistics Education 7(2), 1999). Real, unmodified state data.
1. The Headline Version
Each dot is one state: spending per student vs. average SAT score.
Slope: about −21 SAT points per additional $1,000 spent per student.
Read literally, that says higher-spending states do worse — exactly the
conclusion Will drew in his column. Before you believe it: what isn't shown here that
might differ from state to state?
2. The Same 50 States, One More Variable
Colored by the share of each state's students who even took the SAT.
Low participation (4–11%)
Mid participation (12–55%)
High participation (57–81%)
Within every participation-rate group, the slope is flat or slightly positive
— not negative. Spending isn't hurting scores. States differ enormously in
who takes the SAT: in low-participation states, only the strongest,
most college-bound students bother — inflating the average. In
high-participation states (often ones that push the SAT statewide, or where the ACT
isn't the regional norm), nearly everyone takes it, pulling the average down. That
single confound — participation rate — explains the "headline" trend.
Why this matters
This isn't a hypothetical statistics trick — it's a real case where a public, influential misreading of aggregate data fed directly into a real policy argument about school funding. Guber's 1999 paper was written specifically to correct it.
Discuss:
- Before you saw Panel 2, would you have thought to ask "who's in the denominator?" What made the pooled trend line in Panel 1 look so convincing?
- Is there a "real" effect of spending here at all, once you control for participation rate? What would you need to know to answer that more confidently?
- Where else might you encounter this exact shape of mistake — a pooled average hiding a confound that flips the story once you split the groups?