Research Article

Stacking Ensemble Approach for Predicting Loan Approval Using Machine Learning Techniques

DOI:

10.3791/68832

September 23rd, 2025

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study develops a stacking ensemble model integrating XGBoost, CatBoost (Gradient Boosting Model), LightGBM (Efficient Gradient Boosting Model), AdaBoost, and Extra Trees to predict loan approvals using Kaggle data. Achieving 98% accuracy, it identifies key predictors like income and credit score, promoting fair, efficient decisions on loan approval and/or rejection.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Digital lending and fintech innovations have upended established banking systems, changing financial inclusion and credit availability in nations around the world. This study examines how peer-to-peer (P2P) and digital lending platforms are changing, emphasizing how technologies like artificial intelligence and machine learning are changing the way loans are approved. A thorough study of the literature highlights the opportunities and problems in the digital lending ecosystem, such as algorithmic risk assessment, customer trust, financial exclusion, and regulatory loopholes. This paper suggests a strong machine learning approach that uses a stacking ensemble model to accurately forecast loan approvals in order to address these issues. The data was pre-processed using train-test partitioning, exploratory analysis, and label encoding using a publicly accessible Kaggle dataset that included applicant demographics, financial characteristics, and credit histories. With XGBoost serving as the meta-learner, the ensemble incorporates the Gradient Boosting Model, Efficient Gradient Boosting, AdaBoost, and Extra Trees classifiers as base learners. With an accuracy of 98%, the model was assessed using measures including accuracy, precision, recall, F1-score, and error metrics (MAE- Mean Absolute Error, MSE- Mean Squared Error, and RMSE- Root Mean Square Error). According to correlation studies, factors including assets, income, and CIBIL scores have a significant impact on loan approvals. Outperforming conventional methods, the model showed balance and generalization across both classes. The usefulness of these models for automated, data-driven credit determinations is emphasized in the paper's conclusion.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

In the latest phase of the banking industry's technology transformation, disruptive new financial service providers from outside the established banking system have entered the market1. BigTech (large tech companies that primarily focus on lending directly or with financial institutions) and FinTech (financial technology, including models like P2P lending and online credit alternatives to traditional banks) companies are making substantial inroads into the finance sector, posing a challenge to traditional banking despite banks' efforts to adapt to the digital landscape2. This rapid evolution signals a shift in the financial ecosystem, where non-traditional players are increasingly reshaping the way financial services are accessed and delivered3. The emergence of digital lending has a negative correlation with banking credit, suggesting that as new lenders enter the market, traditional banking may give way to alternative digital credit4. This transition was catalyzed further by the 2008 global financial crisis (GFC), which drastically lowered client trust in financial services and helped drive the expansion of financial technology or Fintech ventures5. Fintech is the term for the combination of technology and finance, which refers to the application of technology to provide financial solutions6. As Fintech matured, one of its most transformative applications was seen in the rise of P2P lending, also referred to as online lending services7. P2P lending's primary innovation is the direct matching of lenders and borrowers. Borrowers submit applications for small, unsecured loans, and lending platforms are used by several investors to evaluate and finance loan requests8. P2P lending functions similarly to a bank, but it uses the internet and cutting-edge technology to enable online lending and debt arrangements 9. The success and scalability of this model became evident with the launch of ZOPA.com, the first P2P platform in history, which debuted in the UK in 2005. Since then, online lending has grown significantly, reaching over $100 billion by 2015, and is expected to reach over $1 trillion in 202510. Digital lending, particularly in emerging economies, has further evolved with Fintech integration11. Fintech integration in digital lending enhances financial inclusion, particularly in emerging markets. Mobile payments and blockchain solutions enable P2P transactions and micro-lending, reducing barriers to financial services12. This paradigm shift is driven by the incorporation of technologies such as blockchain, artificial intelligence (AI), machine learning, and digital payment systems to create a more inclusive, efficient, and customer-centric financial environment13. Digital lending platforms use technology to expedite applications, save expenses, and enhance credit risk evaluation, allowing small firms and individuals to receive financing more quickly14. They use big data, blockchain, AI, and machine learning to improve borrower evaluation, lower costs, and promote financial inclusion15. Machine learning, in particular, has revolutionized risk management by leveraging alternative data sources16. It surpasses traditional credit assessment approaches by leveraging non-traditional data, enhancing borrower ratings, and forecasting economic developments17. This method reduces the risk of default by increasing the accuracy of borrower assessments and assisting in the prediction of shifts in the economy18. One of the most important effects of digital lending is its ability to address the difficulties of financial inclusion, particularly in emerging economies and marginalized areas19.

In order to forecast loan acceptance with high accuracy using a structured Kaggle dataset, this paper proposes a novel stacking ensemble model that combines the Gradient Boosting Model, the Efficient Gradient Boosting Model, the AdaBoost, the Extra Trees, and the XGBoost. To improve predictive adaptability and generalization, this method combines several advanced learners with XGBoost as a meta-classifier, in contrast to earlier research that frequently uses single models or conventional classifiers. The model performed well in both accepted and rejected loan classes, with an impressive accuracy rate of 98%. This methodological development provides a practicable and expandable way to automate loan approval decisions in digital lending settings, particularly in developing financial ecosystems.

The aim of this research is to create a strong stacking ensemble model for digital lending that accurately predicts loan acceptance by combining Gradient Boosting Model, Efficient Gradient Boosting Model, AdaBoost, Extra Trees, and XGBoost. Additionally, it seeks to examine how important demographic and financial variables (income, asset value, and CIBIL-Credit Information Bureau (India) Limited Score) affect loan choices, evaluate how well the ensemble model performs in comparison to more conventional models using classification and error metrics, and emphasize how ensemble approaches can increase efficiency, generalization, and fairness. The primary goal is to statistically analyze how applicant features influence loan approval and to evaluate the performance of ensemble learning algorithms.

P2P and digital lending continue to transform the financial landscape globally, presenting both opportunities and challenges.

Digital lending is rapidly transforming the global financial landscape, offering an alternative to traditional banking20. This global outlook underscores how regional contexts uniquely shape digital lending maturity. Digital lending is expanding but remains technologically immature, while automation and predictive scoring bring efficiency, and platforms still depend heavily on third-party systems for background checks, which limits robustness21. Despite its rapid expansion, financial exclusion remains a major issue globally, with an estimated 44% of adults in developing countries lacking access to formal financial services, necessitating urgent reform, better infrastructure, and digital literacy initiatives. Such limitations also appear in other converging sector highlights, ongoing challenges in data handling, and system integration22. As digital integration deepens, Security vulnerabilities across the Fintech space are escalating. To address these, a security-focused framework has been proposed to safeguard digital transactions23. Similar developments are observed in other emerging markets. In Kenya, while mobile money and digital loan apps have improved financial access, data privacy remains a persistent concern, and recent regulations have limited impact, suggesting that stronger enforcement mechanisms, formal audits, and clear development guidelines are needed24. This reflects a broader trend where regulatory frameworks often lag behind fintech innovation. The regulatory landscape of fintech is different from that of traditional banking. For example, unless loans are high-risk, law enforcement has less of an effect on interest rates in fintech25. Especially, there is a strong need for improved supervision, use of data analytics, and regulatory updates to curb illegal fintech growth and privacy breaches26. Beyond regulation, the success of digital lending also hinges on trust, so trust plays a critical role in lending decisions. Trust in barrowers is more influential than in intermediaries27.

A parallel evolution is visible in India's digital lending ecosystem28. The digital lending business is expanding rapidly, due to advances in fintech, helpful regulatory measures implemented by the Reserve Bank of India (RBI), and an increase in consumer trust following the COVID-19 outbreak29. However, with innovation comes risk. While unlicensed digital lending applications or platforms improve access, they pose severe consumer risks, such as harassment, high interest rates, and data misuse due to weak regulations. Strengthening consumer protection and accountability is therefore critical to promoting responsible financial inclusion30. The dangers of borrower defaults and fraudulent applications are substantial for digital lending; good consumer protection measures not only protect consumers but also positively influence financial performance as data security and transparency improve profitability indicators like Return on assets (ROA) and Return on equity (ROE)31. Globally, there is a considerable emphasis on operational improvements, with increased emphasis on enhancing loan origination systems, encouraging the use of mobile technology, and developing clear strategies to fulfil regulatory standards and consumer expectations32. To address these risks, advanced analytical and AI are increasingly being employed to predict high-risk lenders, outlier detection using indicators like failed loans, repayment duration, and credit scoring has proven effective33. Using the socio-technical model as a guide, we discovered that risks come from both stakeholders and the lack of interdependencies between platform design and organizational components34. Adoption of dynamic models like UTAUT2 dominates in explaining user adoption, with trust emerging as a key predictor of borrowing intent35. Machine learning-based fraud detection algorithms, such as Random Forest and SVM models, are also used36. According to the study's findings, machine learning models can adequately evaluate personal credit information and determine the likelihood of loan default; the deep neural network performed best (accuracy: 0.94)37. The study, which used Naïve Bayes with 94% accuracy, discovered that characteristics such as interest rate, repayment time, description, credit grade, loan history, gender, and credit score have a substantial impact on loan success38. Meanwhile, the probabilities of both prepayment and default risks exist, important occurrences that result in loan termination and loss of profit for creditors were predicted using multivariate logistic regression, and the model's overall accuracy was 76.63%39. According to the study, lending clubs' revenue can be increased with high accuracy of 68 % by utilizing an Efficient Gradient Boosting Model to forecast default risk on Digital lending platforms40. Simultaneously, more sophisticated AI models are evolving, such as deep multiview learning, which combine various variables (such as app usage and behavioural patterns) and perform better than conventional techniques, particularly in situations where historical data is limited41. Studies from China confirm that improving default predictions and financial inclusion, with models like Gradient Boosting Model and LGBM outperforming traditional credit-based evaluations42, system dynamic modelling also helps simulate interest rate fluctuations on P2P platforms, offering insight into borrower investor behavior under various conditions43. Efficient Gradient Boosting Model has been shown to improve default prediction and platform profitability40, while deep neural networks also outperform traditional models when proper training37, and stabilize digital markets through better risk management44 to ensure sustainability, regulatory technology is gaining traction, such as Robotic Process Automation assists financial institutions in aligning regulatory requirements with business plans, enhancing compliance and operational efficiency45. Table 1 summarizes key studies that explore the application of machine learning in digital lending and loan approval processes.

Access restricted. Please log in or start a trial to view this content.

Protocol

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Data collection

This study utilized the Loan Approval Prediction Dataset available on Kaggle. The dataset was extracted in February 2025 and consists of 4269 records aimed at evaluating loan data and forecasting loan approval outcomes. It includes 12 columns comprising detailed information on applicants' demographic profiles, such as employment status, dependents, self-employed, loan amount, loan term, CIBIL scores, financial background, and loan-specific attributes. The dataset was imported using the Pandas library and visually inspected using df.head () to understand its structure and quality.

Data pre-processing

During the data pre-processing phase, the first step involved removing the identifier column (loan_id) due to its lack of predictive value and potential to introduce noise into the model. The second step involved label encoding, where categorical variables such as education, self-employed, and loan_status were converted into numerical representations. This transformation was conducted using Label Encoder from the sklearn.preprocessing module. Specifically, education was encoded as 0 for Graduate and 1 for Not Graduate; self_employed as 0 for No and 1 for Yes, and loan_status, the target variable, as 0 for Not Approved and 1 for Approved. These conversions were necessary to ensure compatibility with machine learning models, which require numeric inputs, particularly for digital lending applications. The features were separated from the target variable using X=df.drop (["loan_status"], axis=1) and y=df ["loan_status]. This setup provided a comprehensive basis for examining the factors influencing loan approval decisions using historical loan records to train multiple ensemble machine learning models. These models were intended to improve overall accuracy and robustness by combining the predictive strengths of multiple classifiers.

The processed dataset was then split into training and testing subsets using the train_test_split function from sklearn.model_selection, with 80% of the data used for training and 20% reserved for testing. This ensured that the model was trained on a sufficiently large portion of the data while retaining a representative sample for performance evaluation. With the dataset cleansed, structured, and statistically explored, the foundation was laid for the implementation of a robust machine learning framework aimed at enhancing predictive accuracy in loan approval classification. Model development was conducted using four ensemble-based machine learning algorithms: Gradient Boosting Model, AdaBoost, Efficient Gradient Boosting Model, and Extra Trees Classifier. These were selected for their proven performance in classification tasks involving structured, tabular data. The Gradient Boosting Model Classifier, implemented from the Gradient Boosting Model library, was instantiated with default settings (iterations=1000, learning rate=0.1, depth=6, verbose=False). It was trained using. fit (x_train, y_train) and evaluated with .predict (X_test). Although the Gradient Boosting Model automatically handles categorical data encoding, this feature was not utilized as the data had already been label-encoded. The AdaBoost Classifier (Adaptive Boosting, which improves weak learners) was implemented using sklearn-ensemble. AdaBoost Classifier was configured with n_estimators=100 and learning_rate=1.0, using decision stumps as the default base estimator. It was trained and evaluated in a similar fashion, contributing robustness through iterative weighting of misclassified instances. The Efficient Gradient Boosting, implemented through the Efficient Gradient Boosting Model library (LGBMClassifier), was configured with n_estimators=100, learning_rate=0.1, and max_depth=-1 (unrestricted tree depth). This model, known for its speed and efficiency, particularly excels on large datasets with high-dimensional features using optimized gradient boosting decision trees.

Finally, the ExtraTrees Classifier from sklearn.ensemble was used with n_estimators=100 and criterion="gini" as the splitting strategy. Unlike Random Forest, Extra Trees introduces further randomness by selecting cut points at random, which helps to reduce model variance and improve generalization. The ensemble was conducted using scikit-learn's Stacking Classifier, which enhances generalization by aggregating predictions from the base learners. Each model was evaluated using standard classification metrics, including accuracy, precision, F1-score, error analysis, and the confusion matrix. These metrics were computed using functions from the sklearn.metrics module to ensure standardized performance comparison across all models.

The best-performing model (based on accuracy and F1-score) was saved for deployment using the Python library. dump(model, "best_model.pkl"), ensuring that the trained model could be reused without the need for retraining. To simulate a real-world application, a sample input array containing 11 features was created using NumPy and passed to the model .predict () function. For instance, the input vector [[0, 1, 1,4100000, 12200000, 8, 417, 2700000, 2200000, 8800000, 3300000]] returned a prediction of 1, indicating loan approval. All experimentation was conducted in a Python 3.10 environment using Google Notebook on Kaggle. Model development and evaluation were carried out using the scikit-learn (v1.3), Gradient Boosting Model, and Efficient Gradient Boosting Model libraries. All hyperparameters were documented explicitly, and defaults were clearly stated where applicable. The encoding procedures followed the approach described by Pedregosa and were implemented in scikit-learn46. This comprehensive and transparent methodology ensures that the experimental protocol is fully reproducible and adheres to rigorous academic standards in machine learning research.

The structure of the suggested methodology, encompassing the phase of data preparation feature section, model training and evaluation, are shown in Figure 1.

This research introduces a stacking ensemble learning framework that brings together the capabilities of four powerful classifiers: Gradient Boosting Model, AdaBoost, Efficient Gradient Boosting Model, and Extra Trees to predict loan approval decisions based on historical financial records. By combining both boosting and bagging strategies within stacked model architecture46. The approach effectively overcomes the individual shortcomings of these models, such as bias and variance, lending to enhanced prediction accuracy and model generalization. Each base learner contributes unique strengths Gradient Boosting Model is efficient with categorical variable, it is designed for handling high cordiality categorical features and internally performs target encoding using ordered boosting47. This avoids over fitting by ensuring that only past data is used in computing statistics. In the formula

Gradient boosting formula, ΣTt=1 ntht(x), model equation, machine learning concept.   ,

each ht (x) represents a decision tree trained on residuals of the previous model, and nt denotes the step-specific learning contribution. AdaBoost or Adaptive Boosting, adjusts the weight of each instance during training and focuses on previously misclassified data points48. In the formula
Ensemble learning equation \(H(x)=sign(\sum_{t=1}^{T}\alpha_{t}h_{t}(x))\), mathematical formula.

αt reflects the performance of the t-th weak learner ht(x), placing more emphasis on previously misclassified samples. Efficient Gradient Boosting Model Incorporates gradient-based one-side sampling (GOSS) and exclusive feature bundling for faster performance. Efficient Gradient Boosting offers high speed and performance on large-scale data49.

Gradient boosting formula ΣF(t)=Σ[l(yi, Ft-1(xi)+ft(xi))+Ω(ft)]; equation, data modeling.

ft(xi) represents the new decision tree added to minimize the loss l(•) while Ω(ft) is a regularization term . In contrast boosting algorithms, Extra Trees reduces variance by adding randomness in decision tree splits50. It relies on bagging principles but injects extra randomness during node splitting into its prediction rule

Ensemble averaging equation Σhm(x), ΣM=1, formula; statistical analysis method, research study.

Averages the outputs of M independently trained randomized trees. For each split, Extra trees selects random thresholds for features and chooses the best among them, thereby reducing variance and offering high diversity across trees, which improves generalization. These models are collectively integrated through a stacking classifier, which learns to optimally combine their outputs to decide whether a loan should be approved. The framework was evaluated with common classification metrics and tested with live input samples, demonstrating its practical relevance in digital lending environments51. These models are combined collectively using a stacking classifier, which learns to blend their outputs ideally to determine loan acceptance results. The model's performance was assessed using important classification measures such as accuracy, precision, recall, F1-score, and AUC-ROC, as well as a confusion matrix, to determine its capacity to reduce both Type I and Type II mistakes. To maintain class balance, a stratified 80:20 train-test split was utilized, with 5-fold cross-validation ensuring robustness and reducing sample variability. Furthermore, the model was evaluated on realistic loan applicant profiles that included information such as credit history, income, employment status, and loan amount, yielding binary judgments and probability ratings. This two-phase test demonstrates the model's efficacy, fairness, and practicality in real-time digital lending contexts. The novelty of this work lies in the hybrid ensemble design tailored for credit scoring, making it a robust, interpretable, and reproducible model for modern financial platforms52 .

Access restricted. Please log in or start a trial to view this content.

Results

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Feature correlation analysis

The feature correlation heatmap (Figure 2) gave useful information about the interrelationships between various attributes. Strong positive correlations were found between income, annual loan amount, and asset-related variables such as luxury assets value and bank asset value, demonstrating that an applicant's financial profile is import...

Access restricted. Please log in or start a trial to view this content.

Discussion

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The stacking ensemble model for loan approval prediction performs exceptionally well across various evaluation metrics, demonstrating great accuracy and reliability. The correlations heatmap revealed that financial indicators such as annual income, loan amount, and asset values are strongly interrelated, emphasizing their importance in loan evolution, whereas the CIBIL scores have a strong negative correlation with loan status, strengthening their role in creditworthiness assessment. The model's confusion matrix had a lo...

Access restricted. Please log in or start a trial to view this content.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The author declares no conflict of interest related to this research.

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This research was supported by VIT-AP University, Amaravati, India.

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
Kagglehttps://www.kaggle.com/
Pandashttps://pandas.pydata.org/
Model libraryIBMhttps://www.ibm.com

References

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,
  1. European Systemic Risk Board. Reports of the Advisory Scientific Committee. , Elsevier. (2012).
  2. Vives, X. The impact of FinTech on banking. Eur Econ. 2, 97-105 (2017).
  3. Jacobides, M. G., Drexler, M., Rico, J. Rethinking the future of financial services: A structural and evolutionary perspective on regulation. J Financ. , https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3078138 (2014).
  4. Cuadros-Solas, P. J., Cubillas, E., Salvador, C. Does alternative digital lending affect bank performance? Cross-country and bank-level evidence. Int Rev Financ Anal. 90, 102873(2023).
  5. Murinde, V., Rizopoulos, E., Zachariadis, M. The impact of the FinTech revolution on the future of banking: Opportunities and risks. Int Rev Financ Anal. 81, 102103(2022).
  6. Hurani, J., Abdel-haq, M. K., Camdzic, E. FinTech Implementation Challenges in the Palestinian Banking Sector. Int J Financial Stud. 12 (4), 122(2024).
  7. Sumit, A., Jian, Z. FinTech Lending and Payment Innovation: A Review. Asia-Pacific J Financ Stud. 1 (1), 11-15 (2020).
  8. Balyuk, T. FinTech Lending and Bank Credit Access for Consumers. Manage Sci. 69 (1), 555-575 (2023).
  9. Novaliando, M. A., Purwokerto, U. M. Legal Protection of Consumer Personal Data in the Case of Fintech Peer to Peer Lending. Proc Series Soc Sci Humanities. 14, 118-124 (2023).
  10. Huang, R. H. Online P2P Lending and Regulatory Responses in China Opportunities and Challenges. Eur Bus Organ Law Rev. 19 (1), 63-92 (2018).
  11. Puschmann, T. Fintech. Bus Inf Syst Eng. 59 (1), 69-76 (2017).
  12. Ebirim, G. U., Odonkor, B. Enhancing Global Economic Inclusion With Fintech Innovations and Accessibility. Financ Account Res J. 6 (4), 648-673 (2024).
  13. Sanyaolu, T. O., Adeleke, A. G., Azubuko, A. F., Osundare, O. F. Exploring fintech innovations and their potential to transform the future of financial services and banking. Int J Sch Res Sci Technol. 5 (1), 054-072 (2024).
  14. Omowole, B. M., Urefe, O., Mokogwu, C., Ewim, S. E. Integrating fintech and innovation in microfinance Transforming credit accessibility for small businesses Integrating fintech and innovation in microfinance Transforming credit accessibility for small businesses. Eur J Innov Manag. 27 (9), 562-581 (2024).
  15. Umavezi, J. U. Innovations in Lending-Focused FinTech Leveraging AI to Transform Credit Accessibility and Risk Assessment. IJCATR. 14 (1), 46-61 (2025).
  16. Leo, M., Sharma, S., Maddulety, K. Machine learning in banking risk management: A literature review. Risks. 7 (1), 2319-8656 (2019).
  17. Bazarbash, M. FinTech in Financial Inclusion: Machine Learning Applications in Assessing Credit Risk. IMF Work Pap. 2019 (109), 1(2019).
  18. Berg, T., Burg, V., Gombovi, A., Puri, M. On the Rise of the Fintech. Rev Financ Studies. 33 (7), 2845-2897 (2020).
  19. Gomber, P., Kauffman, R. J., Parker, C., Weber, B. W. On the Fintech Revolution Interpreting the Forces of Innovation , Disruption and Transformation in Financial Services. J Manag Informat Syst. 35 (1), 220-265 (2018).
  20. Anakpo, G., Xhate, Z., Mishi, S. The Policies, Practices, and Challenges of Digital Financial Inclusion for Sustainable Development The Case of the Developing Economy. FinTech. 2 (2), 327-343 (2023).
  21. Digital Lending High Level System Architecture in Indonesia. Sarungu, C. M. 2020 1st Int Conf Informat Technol Adv Mech Elect Eng, , 159-164 (2020).
  22. Allioui, H., Mourdi, Y. Exploring the Full Potentials of IoT for Better Financial Growth and Stability: A Comprehensive Survey. Sensors. 23 (19), 8015(2023).
  23. Fintech future business & Cyber vulnerabilities and challenges. Venkata, T., Rao, V. 2023 IEEE 8th Int. Conf. Softw. Eng. Comput. Syst, , 1-4 (2023).
  24. Flores, A. M., He, M., Wu, W., Munyaka, I. N. S. A License to Prey Investigating the Impact of Digital Loan App Regulations on Permission Requests and Privacy Policies in the Kenyan Market. 2024 IEEE Int Symp Technol Soc. , 1-5 (2024).
  25. Peng, H., Ji, J., Sun, H., Xu, H. Legal enforcement and fintech credit: International evidence. J Empir Financ. 72, 214-231 (2023).
  26. Suryono, R. R., Budi, I., Purwandari, B. Detection of fintech P2P lending issues in Indonesia. Heliyon. 7 (4), e06782(2021).
  27. Chen, D., Lai, F., Lin, Z. A trust model for online peer-to-peer lending a lender ' s perspective. Info Techno Manag. 15 (4), 239-254 (2014).
  28. Migozzi, J., Urban, M., Wójcik, D. You should do what India does': FinTech ecosystems in India reshaping the geography of finance. Geoforum. 151, 2023(2024).
  29. Asamani, A., Majumdar, J. An Empirical Study of Digital Lending in India and the Variables Associated with its Adoption. Administration Review. 21 (3), 1-13 (2024).
  30. Mark, T. Digital Lending in Emerging Economies: the Nexus Between Financial Innovation and Consumer Protection. Am. J. Financ. Account. 7 (1), 145-168 (2023).
  31. Imanuddin, I., Dewi Anggraeni, R. R., Fridayani, S. Construction of Consumer Protection Against Illegal Online Loan Transactions As a Means of IUS Constituendum in Indonesia. J IUS Kaji Huk dan Keadilan. 11 (3), 539-556 (2023).
  32. Katsamakas, E., Sanchez-Cartas, J. M. A computational model of the effects of borrower default on the stability of P2P lending platforms. Eurasian Econ Rev. 14 (3), 597-618 (2024).
  33. Li, H., Zhang, Y., Zhang, N., Jia, H. Detecting the Abnormal Lenders from P2P Lending Data. Procedia Comput Sci. 91, 357-361 (2016).
  34. Bao, T., Ding, Y., Gopal, R., Möhlmann, M. Throwing Good Money After Bad: Risk Mitigation Strategies in the P2P Lending Platforms. Inf Syst Front. 26 (4), 1453-1473 (2024).
  35. Mudjahidin, A. A., Hidayat,, Aristio, A. P. Conceptual model of use behavior for peer-to-peer lending in Indonesia. Procedia Comput Sci. 197 (2021), 215-222 (2021).
  36. Identifying Features for Detecting Fraudulent Loan Requests on P2P Platforms. Xu, J., Chen, D., Chau, M. 2016 IEEE Conf Intell Secur Informatics, , 79-84 (2016).
  37. Research on Personal Loan Default Assessment Based on Machine Learning. Liu, G. ITM Web of Conferences, 01012, 1-14 (2025).
  38. An application of Naive Bayes classification for credit scoring in e-lending platform. Vedala, R., Kumar, B. R. Proc 2012 Int Conf Data Sci Eng. ICDSE 2012, , 81-84 (2012).
  39. Li, Z., Li, K., Yao, X., Wen, Q. Predicting Prepayment and Default Risks of Unsecured Consumer Loans in Online Lending. Emerg Mark Financ Trade. 55 (1), 118-132 (2019).
  40. Ko, P. C., Lin, H., Do, T., Huang, Y. F. P2P Lending Default Prediction Based on AI and Statistical Models. Entropy. 24 (6), 1-23 (2022).
  41. Loan Fraud Users Detection in Online Lending Leveraging Multiple Data Views. Zhao, S., et al. Proc 37th AAAI Conf Artif Intell AAAI 2023, 37, 5428-5436 (2023).
  42. Fu, G., Sun, M., Xu, Q. An Alternative Credit Scoring System in China's Consumer Lending Market: A System Based on Digital Footprint Data. SSRN Electron J. , 1-51 (2020).
  43. Pang, S., Deng, C., Chen, S. System Dynamics Models of Online Lending Platform Based on Vensim Simulation Technology and Analysis of Interest Rate Evolution Trend. Comput Intell Neurosci. 2022, 9776138(2022).
  44. Tu, Y., Yan, X., Wang, H. Game Theory Analysis of Chinese DC/EP Loan and Internet Loan Models in the Context of Regulatory Goals. Sustain. 15 (9), 1-15 (2023).
  45. Von Solms, J. Integrating Regulatory Technology ( RegTech ) into the digital transformation of a bank Treasury. J Bank Regul. 22 (2), 152-168 (2021).
  46. Barupal, D. K., Fiehn, O. Generating the blood exposome database using a comprehensive text mining and database fusion approach. Environ Health Perspect. 127 (9), 2825-2830 (2019).
  47. Wolpert, D. H. Stacked generalization. Neur Netw. 5 (2), 2941-2259 (1992).
  48. Prokhorenkova, L., Gusev, G., Vorobev, A., Dorogush, A. V., Gulin, A. Catboost: Unbiased boosting with categorical features. Adv Neural Inf Process Syst. , 6638-6648 (2018).
  49. Freund, Y., Schapire, R. E. A Decision-Theoretic Generalization of On-Line Learning and an Application to Boosting. J Comput Syst Sci. 55 (1), 119-139 (1997).
  50. LightGBM: An effective decision tree gradient boosting method to predict customer loyalty in the finance industry. Machado, M. R., Karray, S., De Sousa, I. T. 14th Int Conf Comput Sci Educ ICCSE, , 1111-1116 (2019).
  51. Geurts, P., Ernst, D., Wehenkel, L. Extremely randomized trees. Mach Learn. 63 (1), 3-42 (2006).
  52. Sagi, O., Rokach, L. Ensemble learning: A survey. Wiley Interdiscip Rev Data Min Knowl Discov. 8 (4), 1-18 (2018).
  53. Khandani, A. E., Kim, A. J., Lo, A. W. Consumer credit-risk models via machine-learning algorithms. J Bank Financ. 34 (11), 2767-2787 (2010).
  54. Chen, Y. From Statistical Interpretations to Explainable AI in Machine Learning Enhancing Decision-Making in the Lending Industry. , Doctor of Philosophy, The University of Edinburgh. (2024).

Access restricted. Please log in or start a trial to view this content.

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

Stacking EnsembleLoan Approval PredictionMachine Learning TechniquesDigital LendingPeer To Peer LendingAlgorithmic Risk AssessmentGradient BoostingXGBoost ModelCredit ScoringFinancial Inclusion

Related Articles