Skip to content

Measuring Success in Education: The Role of Effort on the Test Itself

icpsr_231439

Uri Gneezy, John A. List, Jeffrey A. Livingston, Xiangdong Qin, Sally Sadoff, Yang Xu (2019). AER: Insights

DOI: 10.1257/aeri.20180633

field value
venue AER: Insights
discipline education
paper 10.1257/aeri.20180633
replication data openICPSR 231439
data license CC BY 4.0 (LICENSE.txt (shipped in the openICPSR 231439 deposit; project page lists 'other') — maintainer-read 2026-07-20) — “Modified BSD License (https://opensource.org/licenses/BSD-3-Clause) - applies to all code, scripts, programs, and SOFTWARE. [...] Creative Commons Attribution 4.0 International Public License (https://creativecommons.org/licenses/by/4.0/) - applies to databases, images, tables, text, and any other objects. COPYRIGHT 2019 American Economic Association”
regression datasets (in this corpus) 3
graduated / eligible / exported 3 / 18 / 29 (policy claims_manual_regression_cap)
routed / gap / residue (at source) 18 / 4 / 0

Claims & tasks

Every claim the paper makes in its main body about a parameter, joined to the benchmark task (dataset · coefficient) that captures it. The effect, s.e. and p are the captured regression's own reported values (the paper's Stata figures, carrying the paper's SE method); the row shows the paper's verbatim quote (page · exhibit), so you can confirm the regression matches what the paper says. The prominence column is the within-study inclusion signal (⭐ headline → primary → secondary). The sensitivity column is the privacy level of the most-sensitive variable in the regression (🔴 high / 🟠 medium / ⚪ low).

# prominence sensitivity claim effect s.e. p task instance (dataset · coef)
C1 ⭐ primary 🟠 medium “US students improve performance substantially in response to incentives, while Shanghai students—who are top performers on assessments—do not.” — p. 1
“In contrast, the estimated effects of incentives in Shanghai are small in magnitude (−0.26 to −0.28 questions, or −0.09 standard deviations) and not statistically significant.” — p. 10
“In response to incentives, the performance of the Chinese students does not change while the scores of US students increase substantially.” — p. 3
Table 2, col 4 (Shanghai (4))
-0.275 0.264 0.298 Linear Regression reg_10 · t
n=656 · d=7
C2 ⭐ primary 🟠 medium “US students improve performance substantially in response to incentives, while Shanghai students—who are top performers on assessments—do not.” — p. 1
“The estimated treatment effect in the United States is an increase of 1.34 to 1.59 questions ( p < 0.01), which is equivalent to an effect size of approximately 0.24 to 0.28 standard deviations (we calculate standard deviations using the full sample).” — p. 10
“In response to incentives, the performance of the Chinese students does not change while the scores of US students increase substantially.” — p. 3
Table 2, col 1 (United States (1))
+1.588 0.401 1.2e-04 Linear Regression reg_7 · t
n=447 · d=5
C3 secondary 🔴 high “As shown in column 1 of panel A, incentives increase the overall probability that a US student answers a question by about 4 percentage points.” — p. 11
“Under incentives, US students attempt more questions (particularly toward the end of the test) and are more likely to answer those questions correctly.” — p. 3
Table 3, col 1 (United States All questions (1))
+0.037 0.017 0.031 Linear Regression reg_12 · t
n=11175 · d=14

Regression datasets

The 3 regression dataset(s) graduated into the benchmark corpus from this study's replication package — each backs one or more claims above. Expand each for its features, response, (sound, data-independent) public bounds, and reproduction grade.

icpsr_231439_reg_10 — score ~ t + sh_school2 + sh_school3 + sh_school4 + t2018 + age + female

Sensitivity proposal: 🟠 medium maximum across columns (advisory; reviewed manually). Reproduction grade:match (benchmark-eligible).

Features

name description type sensitivity bound lo bound hi estimate s.e.
_cons intercept continuous 1 1 +19.15 5.1
t treatment/incentive assignment indicator continuous ⚪ low 0.0 1.0 -0.2753 0.264
sh_school2 Shanghai school 2 fixed-effect indicator continuous ⚪ low 0.0 1.0 +2.572 0.409
sh_school3 Shanghai school 3 fixed-effect indicator continuous ⚪ low 0.0 1.0 +5.207 0.327
sh_school4 Shanghai school 4 fixed-effect indicator continuous ⚪ low 0.0 1.0 +4.59 0.451
t2018 survey wave/round indicator for 2018 continuous ⚪ low 0.0 1.0 -1.624 0.344
age student age in years continuous 🟠 medium 0.0 120.0 -0.07463 0.307
female indicator student is female continuous 🟠 medium 0.0 1.0 -0.5264 0.23

Response

name description type sensitivity bound lo bound hi estimate s.e.
score count of questions answered correctly on test continuous 🟠 medium

Public bounds (data-independent) sourced from: age in years, plausible human age range; binary indicator; sex/gender is medium-sensitivity; binary school fixed-effect dummy; binary wave/round fixed-effect indicator; randomized treatment arm indicator.

n = 656 samples.

icpsr_231439_reg_12 — qa ~ t + us_school1_reg + us_school1_hon + us_school2_reg + us_school2_hon + question + a…

Sensitivity proposal: 🔴 high maximum across columns (advisory; reviewed manually). Reproduction grade:match (benchmark-eligible).

Features

name description type sensitivity bound lo bound hi estimate s.e.
_cons intercept continuous 1 1 +1.14 0.236
t treatment/incentive assignment indicator continuous ⚪ low 0.0 1.0 +0.03717 0.017
us_school1_reg US school 1 regular track indicator continuous ⚪ low 0.0 1.0 +0.05893 0.0294
us_school1_hon US school 1 honors track indicator continuous ⚪ low 0.0 1.0 +0.03919 0.0324
us_school2_reg US school 2 regular track indicator continuous ⚪ low 0.0 1.0 +0.07609 0.034
us_school2_hon US school 2 honors track indicator continuous ⚪ low 0.0 1.0 +0.1879 0.0301
question question order/identifier on test categorical ⚪ low
age student age in years continuous 🟠 medium 0.0 120.0 -0.01157 0.014
agemissing indicator that age value is missing continuous ⚪ low 0.0 1.0 +0.09668 0.024
female indicator student is female continuous 🟠 medium 0.0 1.0 -0.01955 0.0143
black indicator student is Black race/ethnicity continuous 🔴 high 0.0 1.0 -0.05364 0.0247
asian indicator student is Asian race/ethnicity continuous 🔴 high 0.0 1.0 -0.05987 0.0334
hisp_white indicator Hispanic, white race continuous 🔴 high 0.0 1.0 -0.06238 0.0249
hisp_nw indicator Hispanic, non-white race continuous 🔴 high 0.0 1.0 -0.09044 0.0612
other indicator other/unclassified race category continuous 🔴 high 0.0 1.0 -0.01157 0.0404

Response

name description type sensitivity bound lo bound hi estimate s.e.
qa indicator question was attempted/answered continuous 🟠 medium 0.0 1.0

Public bounds (data-independent) sourced from: age in years, plausible human age range; binary indicator; race/ethnicity is a high-sensitivity attribute; binary indicator; sex/gender is medium-sensitivity; binary indicator; test-response outcome is medium-sensitivity; binary missingness flag, not a demographic value itself; binary school/track fixed-effect dummy; randomized treatment arm indicator.

n = 11,175 samples.

icpsr_231439_reg_7 — score ~ t + us_school1_reg + us_school1_hon + us_school2_reg + us_school2_hon

Sensitivity proposal: 🟠 medium maximum across columns (advisory; reviewed manually). Reproduction grade:match (benchmark-eligible).

Features

name description type sensitivity bound lo bound hi estimate s.e.
_cons intercept continuous 1 1 +5.607 0.826
t treatment/incentive assignment indicator continuous ⚪ low 0.0 1.0 +1.588 0.401
us_school1_reg US school 1 regular track indicator continuous ⚪ low 0.0 1.0 +2.32 0.844
us_school1_hon US school 1 honors track indicator continuous ⚪ low 0.0 1.0 +6.426 0.924
us_school2_reg US school 2 regular track indicator continuous ⚪ low 0.0 1.0 +8.795 1.02
us_school2_hon US school 2 honors track indicator continuous ⚪ low 0.0 1.0 +13.92 0.885

Response

name description type sensitivity bound lo bound hi estimate s.e.
score count of questions answered correctly on test continuous 🟠 medium

Public bounds (data-independent) sourced from: binary school/track fixed-effect dummy; randomized treatment arm indicator.

n = 447 samples.