世界银行-机器学习时代的贫困地图(英)-2023.5-37页_1mb
报告摘要
Poverty Mapping with Machine Learning: Methodology and Validation
The paper analyzes machine learning (ML) approaches for poverty mapping, which rely heavily on remotely-sensed data and are compared to traditional small area estimation methods. Key findings from simulation experiments using Mexican survey data include:
-
Validation Issues: The common R² validation metric is downward-biased when using direct estimates and fails to account for systematic biases, leading to incorrect model selection in many scenarios.
-
Performance Comparison: ML methods, particularly gradient boosting, perform as well as the gold standard (CensusEB unit-level traditional method) when using census-based covariates. However, they underperform significantly when relying on geo-referenced data alone, with substantial variability in bias and poor out-of-sample prediction for hard-to-reach populations.
-
Implications for Targeting: Even suboptimal methods improve poverty reduction through spatial targeting, but ML approaches with insufficient data do not outperform traditional methods in sensitivity, bias, or overall accuracy.
Recommendations: Use traditional small area estimation with updated census data for reliable assessments. Geo-referenced data should be supplementary, as it does not consistently enhance performance in data-scarce contexts.
Summary of Findings
This study demonstrates that while machine learning can rival traditional poverty mapping methods with appropriate data (e.g., census aggregates), its reliance on geo-referenced covariates often results in poor predictive accuracy and biased estimates. The paper underscores the need for rigorous validation using true poverty measures and highlights the importance of methodological choices in improving targeting efficiency.
试读结束,高清完整版pdf/doc/ppt,请点下载