Turkish Journal of Electrical Engineering and Computer Sciences

Binary text classification using genetic programming with crossover-based oversampling for imbalanced datasets

DOI

10.55730/1300-0632.3978

Abstract

It is well known that classifiers trained using imbalanced datasets usually have a bias toward the majority class. In this context, classification models can present a high classification performance overall and for the majority class, even when the performance for the minority class is significantly lower. This paper presents a genetic programming (GP) model with a crossover-based oversampling technique for oversampling the imbalanced dataset for binary text classification. The aim of this study is to apply an oversampling technique to solve the imbalanced issue and improve the performance of the GP model that employed the proposed technique. The proposed technique employs a crossover operator for generating new samples for the minority class in an imbalanced text dataset. By using a combination of this crossover-based oversampling technique with GP, the performance was improved. It is shown that the proposed combination outperforms all GP applications that use the original dataset without resampling. Moreover, the performance of the proposed system surpassed GP approaches using the synthetic minority oversampling technique (SMOTE) and random oversampling. Further comparison with the state-of-the-art on five imbalanced text datasets in terms of F1-score shows the superior performance of the proposed approach.

Keywords

Imbalanced dataset, text classification, genetic programming, oversampling techniques, resampling

First Page

180

Last Page

192

Recommended Citation

ALJERO, MONA and DİMİLİLER, NAZİFE (2023) "Binary text classification using genetic programming with crossover-based oversampling for imbalanced datasets," Turkish Journal of Electrical Engineering and Computer Sciences: Vol. 31: No. 1, Article 12. https://doi.org/10.55730/1300-0632.3978
Available at: https://journals.tubitak.gov.tr/elektrik/vol31/iss1/12

Download

Included in

Computer Engineering Commons, Computer Sciences Commons, Electrical and Computer Engineering Commons

COinS

Turkish Journal of Electrical Engineering and Computer Sciences

Binary text classification using genetic programming with crossover-based oversampling for imbalanced datasets

DOI

Abstract

Keywords

First Page

Last Page

Recommended Citation

Included in

Issues by Year

Search

Turkish Journal of Electrical Engineering and Computer Sciences

Binary text classification using genetic programming with crossover-based oversampling for imbalanced datasets

Authors

DOI

Abstract

Keywords

First Page

Last Page

Recommended Citation

Included in

Share

Issues by Year

Search