EBA欧洲银行-Session-2-Slides-Predicting-Bank-Insolvencies-using-ML_19页_619kb
报告摘要
Summary of "Predicting bank insolvencies using Machine Learning techniques"
Core Content
This document presents a study on predicting bank insolvencies using various machine learning (ML) techniques, focusing on the Random Forests model. The authors, from the Bank of Greece, aim to develop a novel rating system for financial institutions to assess their risk of insolvency and to provide a robust framework for supervisory authorities to monitor the banking system effectively.
The study is based on US bank data from the FDIC (2008-2014) and European bank data from SNL (year-end 2015), and explores multiple statistical and ML models such as Logistic Regression, Linear Discriminant Analysis (LDA), Random Forests (RF), Support Vector Machines (SVM), Neural Networks (NN), and Conditional Inference Trees (CRF). The goal is to evaluate their predictive performance and discriminatory power in identifying banks at risk of insolvency.
Main Drivers of Bank Insolvency
The study identifies the following as the primary risk indicators:
- Profitability
- Capital Adequacy
- Asset Quality
- Liquidity
- Earnings
These drivers are evaluated using the CAMELS framework, which includes:
- Capital (Equity, Tier 1, Total risk-based capital)
- Asset Quality (Net charge-offs, loan loss allowances, noncurrent assets, etc.)
- Management (Noninterest income, efficiency ratios, etc.)
- Earnings (Return on Assets, Return on Equity, etc.)
- Liquidity (Net loans and deposits ratios, volatile liabilities)
- Sensitivity to Market Risk (Asset fair value)
Implementation and Data
Dataset Overview
- FDIC data: 175,649 records from 2008 to 2014, including 173,594 "Good" banks and 2,055 "Bad" banks.
- SNL data: 173 European banks, based on year-end 2015 accounting and regulatory data.
Variable Reduction Process
- The initial dataset includes 660+ covariates.
- After correlation and LASSO screening, the number of variables is reduced to 59.
- Further reduction using Random Forest Importance leads to 23 final predictors.
Performance Results
The models are evaluated using multiple performance metrics on different samples:
- In-sample
- Out-of-sample
- Out-of-time
Key Performance Metrics
| Metric | Logit | LDA | RF | SVM | NN | CRF |
|---|---|---|---|---|---|---|
| AUROC | 0.980 | 0.973 | 0.989 | 0.981 | 0.984 | 0.991 |
| G-mean | 0.898 | 0.884 | 0.921 | 0.898 | 0.923 | 0.914 |
| LR- | 0.183 | 0.209 | 0.139 | 0.184 | 0.137 | 0.156 |
| DP | 3,116 | 2,971 | 3,255 | 3,181 | 3,356 | 3,312 |
| BA | 0.902 | 0.889 | 0.923 | 0.902 | 0.925 | 0.916 |
| Youden | 0.804 | 0.778 | 0.846 | 0.804 | 0.851 | 0.833 |
| WBA1 | 0.943 | 0.936 | 0.953 | 0.944 | 0.955 | 0.951 |
| WBA2 | 0.861 | 0.842 | 0.893 | 0.860 | 0.895 | 0.881 |
Findings
- Random Forests outperform all other models in terms of discriminatory power and stability across different test samples.
- Neural Networks also perform well, especially in in-sample and out-of-time scenarios.
- The Capital Adequacy Ratio (CAR) and Leverage Ratio (LEV) show model-dependent importance in predicting bank failures.
- Metrics related to capital and earnings have the highest marginal contribution to the AUROC metric, indicating their significance in the prediction process.
Case Study: European Banks
The study applies the Random Forest model to 173 European banks to create an Early Warning System. A bank is classified as "High Risk" if its Probability of Default (PD) exceeds 25%.
Country-Level Results
| Country | High Risk Banks | Total Banks in Sample |
|---|---|---|
| AT | 0 | 5 |
| BA | 0 | 2 |
| BE | 1 | 2 |
| BG | 1 | 2 |
| CH | 0 | 17 |
| CY | 0 | 2 |
| CZ | 0 | 1 |
| DE | 2 | 8 |
| DK | 0 | 15 |
| ES | 1 | 8 |
| FI | 0 | 2 |
| FR | 3 | 11 |
| GB | 0 | 10 |
| GE | 0 | 1 |
| GR | 3 | 5 |
| HR | 2 | 3 |
| HU | 2 | 2 |
| IE | 0 | 3 |
| IT | 7 | 15 |
| LI | 0 | 1 |
| MD | 1 | 1 |
| MK | 0 | 2 |
| MT | 0 | 3 |
| NO | 0 | 12 |
| PL | 0 | 11 |
| PT | 1 | 2 |
| RO | 1 | 2 |
| RS | 0 | 1 |
| RU | 5 | 5 |
| SE | 1 | 3 |
| SK | 0 | 4 |
| TR | 7 | 9 |
| UA | 3 | 3 |
| Grand Total | 41 | 173 |
- Eurozone countries with prolonged macroeconomic deterioration show the highest number of high-risk banks.
- Stronger economies are associated with resilient banking systems.
- Countries regaining competitiveness have lower levels of risky banks.
- Non-Eurozone countries with strong economies show minimal or no high-risk banks.
- Cautious interpretation is advised for Eastern European and small countries due to limited sample size.
Our Contribution
- Extensive exploration of various statistical and ML techniques.
- First empirical application of Random Forests in assessing bank failures.
- Robust validation of models using multiple performance measures.
- Extended set of potential drivers (over 660 variables) and variable reduction process.
- Benchmarking with Moody’s rating methodology, showing a high positive concordance (Spearman’s Rho: 67%, Fisher correlation: 59%, Kendal’s Tau: 47%).
Conclusion
The study highlights the potential of machine learning techniques in predicting bank insolvencies, with Random Forests emerging as the most effective and stable model. It also underscores the importance of capital and earnings metrics in the prediction process and the value of integrating multiple risk drivers into a single measure. Supervisory authorities are encouraged to adopt these techniques to enhance their monitoring capabilities and improve risk assessment in the banking sector.
试读结束,高清完整版pdf/doc/ppt,请点下载