Tutoring Efficacy, Household Substitution, and Student Achievement: Experimental Evidence from an After-School Tutoring Program in Rural China
icpsr_193091
Jere R. Behrman, C. Simon Fan, Naijia Guo, Xiangdong Wei, Hongliang Zhang, Junsen Zhang (2024). International Economic Review
DOI: 10.1111/iere.12668
| field | value |
|---|---|
| venue | International Economic Review |
| paper | 10.1111/iere.12668 |
| replication data | openICPSR 193091 |
| data license | CC BY 4.0 (openICPSR project page 193091 (maintainer-read 2026-07-20)) — “This work is licensed under a Creative Commons Attribution 4.0 International (CC BY 4.0) License.” |
| regression datasets (in this corpus) | 4 |
| graduated / eligible / exported | 4 / 24 / 32 (policy claims_manual_regression_cap) |
| routed / gap / residue (at source) | 38 / 31 / 0 |
Claims & tasks
Every claim the paper makes in its main body about a parameter, joined to the benchmark task (dataset · coefficient) that captures it. The effect, s.e. and p are the captured regression's own reported values (the paper's Stata figures, carrying the paper's SE method); the row shows the paper's verbatim quote (page · exhibit), so you can confirm the regression matches what the paper says. The prominence column is the within-study inclusion signal (⭐ headline → primary → secondary). The sensitivity column is the privacy level of the most-sensitive variable in the regression (🔴 high / 🟠 medium / ⚪ low).
| # | prominence | sensitivity | claim | effect | s.e. | p | task | instance (dataset · coef) |
|---|---|---|---|---|---|---|---|---|
| C1 | primary | 🟠 medium | “Taking column 2 with the full set of control variables as an example, the estimated coefficient λ̂ indicates that the tutees have an average gain of 0.136σ in math scores (significant at 1% level) compared to the control students from the same experimental classes.” — p. 10 “We find that the tutoring program significantly improved tutees’ endline math scores and the score gains were significantly larger for LBC.” — p. 2 “We also find that the tutoring program significantly improved tutees’ endline math scores, with the score gains being significantly larger for children without parents at home.” — p. 37 Table 2, col 2 (Math Scores (2)) |
+0.136 | 0.040 | 9.8e-04 | Linear Regression | reg_10 · assignedtuteen=1815 · d=5 |
| C2 | primary | 🟠 medium | Regardless of the empirical specifications used, the estimates of λ for reading scores are always insignificant and small in magnitude. p. 10 · Table 2, col 3 (Reading Scores (3)) |
+0.007 | 0.040 | 0.868 | Linear Regression | reg_11 · assignedtuteen=1815 · d=2 |
| C3 | primary | 🟠 medium | Regardless of the empirical specifications used, the estimates of λ for reading scores are always insignificant and small in magnitude. p. 10 · Table 2, col 4 (Reading Scores (4)) |
+0.017 | 0.040 | 0.673 | Linear Regression | reg_12 · assignedtuteen=1815 · d=5 |
| C4 | primary | 🟠 medium | “The point estimates show that the average treatment effect is 0.091σ for tutees living with both parents, 0.076σ for tutees living with a single parent, and 0.203σ for tutees living without both parents. Although the first two estimates are not significant at conventional levels, the last coefficient is significant at the 1% level.” — p. 11 “Nonetheless, for the dummy indicator for living with no parent, the facts that the T3−T1 difference is largest both quantitatively and statistically among all baseline characteristics and that the coefficient estimate on its interaction term with the treatment dummy is also significant (Table 3) seem to suggest that it is indeed the most relevant characteristic to define subgroups with heterogeneous treatment effects.” — p. 17 Table 3, col 1 ((1)) |
+0.203 | 0.046 | 3.4e-05 | Linear Regression | reg_13 · assignedtutee_bothabsent_rd1n=1815 · d=4 |
Regression datasets
The 4 regression dataset(s) graduated into the benchmark corpus from this study's replication package — each backs one or more claims above. Expand each for its features, response, (sound, data-independent) public bounds, and reproduction grade.
icpsr_193091_reg_10 — z_math_rd2 ~ assignedtutee + z_math_rd1 + female + oneabsent_rd1 + bothabsent_rd1
Sensitivity proposal: 🟠 medium maximum across columns (advisory; reviewed manually).
Reproduction grade: ✅ match (benchmark-eligible).
Features
| name | description | type | sensitivity | bound lo | bound hi | estimate | s.e. |
|---|---|---|---|---|---|---|---|
assignedtutee |
treatment assignment: student assigned as tutee | continuous | ⚪ low | -1.0 | 1.0 | +0.1356 | 0.0395 |
z_math_rd1 |
standardized baseline math test score | continuous | 🟠 medium | -12.0 | 12.0 | +0.5433 | 0.044 |
female |
student sex/gender indicator | continuous | 🟠 medium | -1.0 | 1.0 | +0.005172 | 0.0392 |
oneabsent_rd1 |
indicator one parent absent (migrant) at baseline | continuous | 🟠 medium | -1.0 | 1.0 | -0.001256 | 0.0468 |
bothabsent_rd1 |
indicator both parents absent (migrant) at baseline | continuous | 🟠 medium | -1.0 | 1.0 | +0.06111 | 0.0452 |
Response
| name | description | type | sensitivity | bound lo | bound hi | estimate | s.e. |
|---|---|---|---|---|---|---|---|
z_math_rd2 |
standardized endline math test score | continuous | 🟠 medium | -12.0 | 12.0 | — | — |
Public bounds (data-independent) sourced from: binary indicator; binary treatment indicator; standardized z-score, conventional clip range.
n = 1,815 samples.
icpsr_193091_reg_11 — z_chinese_rd2 ~ assignedtutee + z_chinese_rd1
Sensitivity proposal: 🟠 medium maximum across columns (advisory; reviewed manually).
Reproduction grade: ✅ match (benchmark-eligible).
Features
| name | description | type | sensitivity | bound lo | bound hi | estimate | s.e. |
|---|---|---|---|---|---|---|---|
assignedtutee |
treatment assignment: student assigned as tutee | continuous | ⚪ low | -1.0 | 1.0 | +0.006657 | 0.0399 |
z_chinese_rd1 |
standardized baseline Chinese test score | continuous | 🟠 medium | -12.0 | 12.0 | +0.6846 | 0.0433 |
Response
| name | description | type | sensitivity | bound lo | bound hi | estimate | s.e. |
|---|---|---|---|---|---|---|---|
z_chinese_rd2 |
standardized endline Chinese test score | continuous | 🟠 medium | -12.0 | 12.0 | — | — |
Public bounds (data-independent) sourced from: binary treatment indicator; standardized z-score, conventional clip range.
n = 1,815 samples.
icpsr_193091_reg_12 — z_chinese_rd2 ~ assignedtutee + z_chinese_rd1 + female + oneabsent_rd1 + bothabsent_rd1
Sensitivity proposal: 🟠 medium maximum across columns (advisory; reviewed manually).
Reproduction grade: ✅ match (benchmark-eligible).
Features
| name | description | type | sensitivity | bound lo | bound hi | estimate | s.e. |
|---|---|---|---|---|---|---|---|
assignedtutee |
treatment assignment: student assigned as tutee | continuous | ⚪ low | -1.0 | 1.0 | +0.01688 | 0.0398 |
z_chinese_rd1 |
standardized baseline Chinese test score | continuous | 🟠 medium | -12.0 | 12.0 | +0.6567 | 0.0448 |
female |
student sex/gender indicator | continuous | 🟠 medium | -1.0 | 1.0 | +0.2689 | 0.0448 |
oneabsent_rd1 |
indicator one parent absent (migrant) at baseline | continuous | 🟠 medium | -1.0 | 1.0 | -0.007597 | 0.0472 |
bothabsent_rd1 |
indicator both parents absent (migrant) at baseline | continuous | 🟠 medium | -1.0 | 1.0 | +0.02659 | 0.0453 |
Response
| name | description | type | sensitivity | bound lo | bound hi | estimate | s.e. |
|---|---|---|---|---|---|---|---|
z_chinese_rd2 |
standardized endline Chinese test score | continuous | 🟠 medium | -12.0 | 12.0 | — | — |
Public bounds (data-independent) sourced from: binary indicator; binary treatment indicator; standardized z-score, conventional clip range.
n = 1,815 samples.
icpsr_193091_reg_13 — z_math_rd2 ~ z_math_rd1 + assignedtutee_noabsent_rd1 + assignedtutee_oneabsent_rd1 + assi…
Sensitivity proposal: 🟠 medium maximum across columns (advisory; reviewed manually).
Reproduction grade: ✅ match (benchmark-eligible).
Features
| name | description | type | sensitivity | bound lo | bound hi | estimate | s.e. |
|---|---|---|---|---|---|---|---|
z_math_rd1 |
standardized baseline math test score | continuous | 🟠 medium | -12.0 | 12.0 | +0.5424 | 0.0434 |
assignedtutee_noabsent_rd1 |
interaction: treatment x no-parent-absent | continuous | ⚪ low | -1.0 | 1.0 | +0.09084 | 0.0682 |
assignedtutee_oneabsent_rd1 |
interaction: treatment x one-parent-absent | continuous | ⚪ low | -1.0 | 1.0 | +0.07622 | 0.0517 |
assignedtutee_bothabsent_rd1 |
interaction: treatment x both-parents-absent | continuous | ⚪ low | -1.0 | 1.0 | +0.2031 | 0.046 |
Response
| name | description | type | sensitivity | bound lo | bound hi | estimate | s.e. |
|---|---|---|---|---|---|---|---|
z_math_rd2 |
standardized endline math test score | continuous | 🟠 medium | -12.0 | 12.0 | — | — |
Public bounds (data-independent) sourced from: interaction of two binary indicators; standardized z-score, conventional clip range.
n = 1,815 samples.