Yayın:
When does Synthetic Data Generation Work?

Yükleniyor...
Küçük Resim

Tarih

Kurum Yazarları

Danışman

item.page.editor

Editör

Bölüm / Program

Dergi Başlığı

Dergi ISSN

Cilt Başlığı

Yayıncı

IEEE

DOI

10.1109/siu53274.2021.9477956
View PlumX Details

Araştırma Projeleri

Akademik Birimler

Dergi Sayısı

Özet

Synthetic data generation is one of the methods used in machine learning to increase the performance of algorithms on datasets. However, these methods do not ensure success on each dataset. In this study, it has been investigated that which type of synthetic data generation algorithms are useful in which datasets by examining the effects of SMOTE, Borderline-SMOTE and Random data generation algorithms on 33 datasets. For this, each dataset has been fully balanced as a result of synthetic data generation. In order to evaluate the results, datasets are divided into three groups as balanced, partially balanced-unbalanced and unbalanced in accordance with the unbalance ratio. The datasets formed as a result of the data generation of the algorithms and the original datasets have been trained with an ANN models and their performance has been evaluated on the test set. Experimental results have shown that adding synthetic data to the datasets with the abovementioned algorithms generally increases the success in balanced and partially balanced-unbalanced datasets, but generally does not work in unbalanced datasets. Borderline-SMOTE, which produces border samples in balanced datasets, and SMOTE in partially balanced-unbalanced datasets have been more successful.

Tanım

Dergi veya Seri

29TH IEEE CONFERENCE ON SIGNAL PROCESSING AND COMMUNICATIONS APPLICATIONS (SIU 2021)

ISSN

ISBN

978-1-6654-3649-6

Haklar

Alıntı

Koleksiyonlar

Onay

Gözden geçir

Tamamlayıcı Bilgiler

Referans Gösteren

Related Patent

Related Goal

0

Views

0

Downloads