Loan Default Risk Classification Using Random Forest Based on Monthly Loan Performance Data

Credit risk, default, machine learning, random forest, monthly performance data

Authors

January 26, 2026
January 27, 2026

Downloads

The risk of loan default is a major challenge in the financial industry because it has a direct impact on the stability of lending institutions. As loan performance data becomes increasingly available, machine learning-based approaches are becoming a potential alternative to improve the accuracy of credit risk assessments. This study aims to analyze the performance of the Random Forest model in classifying default risk based on debtors' monthly performance data. The dataset used consists of more than one million observational records with variables representing credit characteristics and payment behavior. To avoid data leakage, only variables that were available before the default occurred are used. The model evaluation was carried out using Area Under the Curve (AUC) and confusion matrix metrics, accompanied by threshold adjustments to increase the sensitivity of the detection of risky debtors. The results showed that the Random Forest model achieved an AUC value of 0.737, which indicates a good discriminating ability. Threshold adjustments increase the detection of risky debtors with the consequence of increasing false positives, so this approach has the potential to support decision-making in credit risk management.