Effects of Resampling Techniques on Imbalanced Data Classification: A New Under-resampling Method
Advances in Business and Management Forecasting
ISBN: 978-1-83982-091-5, eISBN: 978-1-83982-090-8
Publication date: 1 September 2021
Abstract
We study the performances of various predictive models including decision trees, random forests, neural networks, and linear discriminant analysis on an imbalanced data set of home loan applications. During the process, we propose our undersampling algorithm to cope with the issues created by the imbalance of the data. Our technique is shown to work competitively against popular resampling techniques such as random oversampling, undersampling, synthetic minority oversampling technique (SMOTE), and random oversampling examples (ROSE). We also investigate the relation between the true positive rate, true negative rate, and the imbalance of the data.
Keywords
Citation
Nguyen, S., Schumacher, P., Olinsky, A. and Quinn, J. (2021), "Effects of Resampling Techniques on Imbalanced Data Classification: A New Under-resampling Method", Lawrence, K.D. and Klimberg, R.K. (Ed.) Advances in Business and Management Forecasting (Advances in Business and Management Forecasting, Vol. 14), Emerald Publishing Limited, Leeds, pp. 51-70. https://doi.org/10.1108/S1477-407020210000014005
Publisher
:Emerald Publishing Limited
Copyright © 2021 by Emerald Publishing Limited