2024-10-13-世界银行-后双重选择拉索在田间试验中的应用(英)_44页_1mb
报告摘要
Summary of PDS Lasso Performance in Field Experiments
Key Findings
-
Limited Impact of PDS Lasso
- PDS Lasso rarely selects control variables (median of 2 variables), and standard errors are marginally smaller than those from ANCOVA, but improvements are minimal on average.
- In over 75% of cases, treatment estimates and standard errors are similar to those obtained using ANCOVA (difference < 0.04 standard deviations).
-
Attrition Effects
- PDS Lasso is more likely to select controls in cases of high attrition. However, even here, the impact on treatment estimates remains small.
-
Failure to Select Key Predictors
- In some cases, PDS Lasso fails to select the lagged dependent variable—a common source of omitted variable bias, leading to larger standard errors.
-
Cross-Validation Risks
- Cross-validation tends to select more variables, increasing the risk of overfitting and inflating standard errors, especially in small samples.
- Recommended to use the default plug-in penalty for consistency.
Recommendations for Researchers
-
Amelioration Set
- Include lagged variables and randomization strata in the amelioration set to prevent their exclusion.
-
Control Variables Selection
- Avoid a “kitchen-sink” approach; limit inputted controls to reduce spurious associations and improve precision.
-
Handling Missing Data
- Address missing values (8 of 18 papers had sample size reduction due to this). Dummy out missingness and include an indicator, or correct the amelioration set.
-
Interactions with Treatment
- Include interacting variables in the amelioration set to avoid unintended selection as controls.
-
Implementation Checklist
- Pre-specify controls and amelioration sets.
- Avoid excessive controls and questionable missing-value handling.
- Do not use cross-validation unless well-justified.
Conclusion
PDS Lasso offers a structured approach to covariate selection but fails to significantly reduce standard errors on average. Researchers should implement it judiciously and prioritize robust model specification over optimism from machine learning.
展开完整摘要
试读结束,高清完整版pdf/doc/ppt,请点下载