Yayın:
Implementation of Data Preprocessing Techniques on Distributed Big Data Platforms

dc.contributor.authorCelik, Oguz
dc.contributor.authorHasanbasoglu, Muruvvet
dc.contributor.authorAktas, Mehmet S.
dc.contributor.authorKalipsiz, Oya
dc.contributor.authorKanli, Alper Nebi
dc.date.accessioned2026-06-27T14:21:46Z
dc.date.issued2019
dc.description.abstractWe are now in the era of Big Data, and the need for tools which can process and analyze such data is yet to be fulfilled. Big data mining aims to extract meaningful and valuable information from voluminous data that traditional data mining tools can not handle. One of the most vital steps of any data mining process is the preprocessing of the data. Our aim was to provide distributed implementation of some algorithms for two of the data preprocessing steps: outlier analysis and missing value imputation. The algorithms were implemented on Spark and this paper will focus on the details and performance of these algorithms on different distributed system setups.en
dc.description.sponsorshipCybersoft, RD Center
dc.description.urihttps://doi.org/10.1109/ubmk.2019.8907230
dc.identifier.doi10.1109/ubmk.2019.8907230
dc.identifier.endpage78
dc.identifier.isbn978-1-7281-3964-7
dc.identifier.startpage73
dc.identifier.urihttps://hdl.handle.net/20.500.14981/59720
dc.identifier.wos000609879900015
dc.language.isoeng
dc.publisherIEEE
dc.relation.conference4th International Conference on Computer Science and Engineering (UBMK)
dc.relation.ispartof2019 4TH INTERNATIONAL CONFERENCE ON COMPUTER SCIENCE AND ENGINEERING (UBMK)
dc.subjectBig Data
dc.subjectDistributed Computing
dc.subjectOutlier Analysis
dc.subjectMissing Value Imputation
dc.subjectComputer Science
dc.titleImplementation of Data Preprocessing Techniques on Distributed Big Data Platforms
dc.typeProceedings Paper
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar