Nguyễn Thị Lan Anh * Nguyễn Thị Ngọc Hà

* Người chịu trách nhiệm về bài viết: (nguyenthilananh@dhsphue.edu.vn)

Abstract

Classifying the imbalanced data sets, especially the high-dimensional datasets is one of the important issues. In this paper, we present an undersampling algorithm with a subset of the available features to enhance the result of the high-dimensional imbalanced data sets classification.
Keywords: high-dimensional imbalanced data sets, undersampling methods

Tóm tắt

Phân lớp dữ liệu mất cân bằng đặc biệt với tập dữ liệu có số chiều lớn là một bài toán quan trọng trong thực tế. Trong bài báo này chúng tôi đề xuất một thuật toán làm giảm số lượng phần tử lớp đa số trên một tập con các đặc trưng để cải thiện hiệu suất phân lớp tập dữ liệu mất cân bằng và có số chiều lớn.
Từ khóa: Dữ liệu mất cân bằng có số chiều lớn, phương pháp làm giảm số lượng phần tử lớp đa số, phương pháp lựa chọn đặc trưng., feature selection methods.

Article Details

Tài liệu tham khảo

He H. and Garcia E. A. (2009). Learning from Imbalanced Data. IEEE Trans. Knowl. Data Eng. 21 (9): 1263–1284.

Chawla N. V, Bowyer K. W, Hall L. O., and Kegelmeyer W. P. (2002). SMOTE: Synthetic Minority Over-sampling Technique. Artif. Intell. Res. 16 (1): 321–357.

Han H., Wang W.Y., and Mao B.H. (2005). Borderline-SMOTE: A New Over-Sampling Method in Imbalanced Data Sets Learning. Advances in Intelligent Computing, ICIC 2005, Lecture Notes in Computer Science. 3644

Bunkhumpornpat C., Sinapiromsaran K., and Lursinsap C. (2009). Safe-Level-SMOTE: Safe-Level-Synthetic Minority Over-sampling Technique for handling the class imbalanced problem. 5476: 475–482

Maldonado S., López J., Vairetti C. (2018). An alternative SMOTE oversampling strategy for high-dimensional datasets. Applied Soft Computing Journal. 76: 380–389

Zhang Z. and Mani I. (2003). KNN Approach to Unbalanced Data Distribution: A Case Study involving Information Extraction. Workshop on Learning from Inmbalanced Datasets II, ICML, Washington DC

Sun Y., Wong A. K. C., and Kamel M. S. (2009). Classification of Imbalanced Data: A Review. J. Pattern Recognit. 23 (4): 687–719

Lichman M. (2013). UCI Machine Learning Repository, http://archive.ics.uci.edu/ml, Irvine, CA: University of California, School of Information and Computer Science.

Hall M., Frank E., Holmes G., Pfahringer B., Reutemann P., Written L. (2009). The WEKA Data Mining Software: An Update. ACM SIGKDD Explorations Newsletter. 11 (1): 10-18.

[10] Gower J.C. (1971). “A general coefficient of similarity and some of its properties”. Biometrics. 27 (1): 857–874