Yayın: Data Feature Selection Methods on Distributed Big Data Processing Platforms
Yükleniyor...
Tarih
Danışman
item.page.editor
Editör
Bölüm / Program
Dergi Başlığı
Dergi ISSN
Cilt Başlığı
Yayıncı
IEEE
DOI
Özet
Nowadays, banking institutions keep their existing data in their relational databases and realize the necessary models using these data. The complex structure of enterprise data models, the large amount of data and the large number of features make it difficult to perform all kinds of analyzing (classification, clustering, regression, etc.) on the dataset. For these reasons, there is a need for software that can be integrated with existing tools that are easy to use for predicting high performance on large data sets, and that can produce the most accurate forecasts in comparison. Within the scope of this study, two feature selection categories which are commonly used in the literature have been studied. These are the Information Theory and Traditional Statistics categories. In the first category, Information Gain, Information Gain Ratio and Information Value methods have been implemented. In the second category, Chi-Square Test method is implemented. Our research examines how feature selection methods can be implemented on a large data processing platform. In addition, if different feature selection methods are used together, it is being investigated which combinations of feature selection can produce more successful results. In this research, a software architecture was developed that can determine the features of high estimating power in order to be able to respond to the requirements mentioned above, and the prototype application of this architecture has been produced. In this article, the methods used in the development process of the software, the algorithms and the details of the prototype application of the developed software are explained. The developed prototype has been applied on banking data, and the results obtained are analyzed comparatively. Our experiments show that the feature selection in the case of using together Information Gain, Information Value and Chi-Square Statistic methods gives the most successful results.
Tanım
Dergi veya Seri
2018 3RD INTERNATIONAL CONFERENCE ON COMPUTER SCIENCE AND ENGINEERING (UBMK)
ISSN
ISBN
978-1-5386-7893-0