Yayın:
Data Feature Selection Methods on Distributed Big Data Processing Platforms

Yükleniyor...
Küçük Resim

Tarih

Kurum Yazarları

Danışman

item.page.editor

Editör

Bölüm / Program

Dergi Başlığı

Dergi ISSN

Cilt Başlığı

Yayıncı

IEEE

DOI

Araştırma Projeleri

Akademik Birimler

Dergi Sayısı

Özet

Nowadays, banking institutions keep their existing data in their relational databases and realize the necessary models using these data. The complex structure of enterprise data models, the large amount of data and the large number of features make it difficult to perform all kinds of analyzing (classification, clustering, regression, etc.) on the dataset. For these reasons, there is a need for software that can be integrated with existing tools that are easy to use for predicting high performance on large data sets, and that can produce the most accurate forecasts in comparison. Within the scope of this study, two feature selection categories which are commonly used in the literature have been studied. These are the Information Theory and Traditional Statistics categories. In the first category, Information Gain, Information Gain Ratio and Information Value methods have been implemented. In the second category, Chi-Square Test method is implemented. Our research examines how feature selection methods can be implemented on a large data processing platform. In addition, if different feature selection methods are used together, it is being investigated which combinations of feature selection can produce more successful results. In this research, a software architecture was developed that can determine the features of high estimating power in order to be able to respond to the requirements mentioned above, and the prototype application of this architecture has been produced. In this article, the methods used in the development process of the software, the algorithms and the details of the prototype application of the developed software are explained. The developed prototype has been applied on banking data, and the results obtained are analyzed comparatively. Our experiments show that the feature selection in the case of using together Information Gain, Information Value and Chi-Square Statistic methods gives the most successful results.

Tanım

Dergi veya Seri

2018 3RD INTERNATIONAL CONFERENCE ON COMPUTER SCIENCE AND ENGINEERING (UBMK)

ISSN

ISBN

978-1-5386-7893-0

Haklar

Alıntı

Koleksiyonlar

Onay

Gözden geçir

Tamamlayıcı Bilgiler

Referans Gösteren

Related Patent

Related Goal

0

Views

0

Downloads