SMOTE-text: A modified SMOTE for Turkish text classification

Küçük Resim Yok

Tarih

2021

Dergi Başlığı

Dergi ISSN

Cilt Başlığı

Yayıncı

Springer Science and Business Media Deutschland GmbH

Erişim Hakkı

info:eu-repo/semantics/closedAccess

Özet

One of the most common problems faced by large enterprise companies is the loss of knowhow after employee’s job replacements and quits. Creating a well-organized, indexed, connected, user friendly and sustainable digital enterprise memory can solve this problem and creates a practical knowhow transfer to new recruited personnel. In this regard, one of the problems that generated is the correct classification of documents that will be stored in the digital library. The most general meaning of text classification also known as text categorization is the process of categorizing text into labeled groups. A document can be related to one or more subjects and choosing the correct labels and classification is sometimes a challenging process. Information repository shows various distributions according to the company’s business areas. For a good and successful machine learning based text classification requires balanced datasets related with the business and previous samples. Due to the lack of documents from minor business creates imbalanced learning dataset. To overcome this problem synthetic data can be created with some methods but those methods are suitable for numerical inputs not proper for text classification. This article presents a modified version of Synthetic Minority Oversampling Technique SMOTE algorithm for text classification by integrating the Turkish dictionary for oversampling for text processing and classification.

Açıklama

Anahtar Kelimeler

Imbalanced Data Sets, Machine Learning, Oversampling, SMOTE-Text, Text Classification

Kaynak

Lecture Notes on Data Engineering and Communications Technologies

WoS Q Değeri

Scopus Q Değeri

Q3

Cilt

76

Sayı

Künye

Curukoglu, N., & Ozpinar, A. (2021). SMOTE-Text: A Modified SMOTE for Turkish Text Classification. In Lecture notes on data engineering and communications technologies, v.76, (pp. 82–92).