Yayın:
On the big data processing algorithms for finding frequent sequences

dc.contributor.authorCan, Ali Burak
dc.contributor.authorZaval, Mounes
dc.contributor.authorUzun-Per, Meryem
dc.contributor.authorAktas, Mehmet S.
dc.date.accessioned2026-06-27T14:49:17Z
dc.date.issued2023
dc.description.abstractSequential pattern mining algorithms extract trendy sequence appearances inside ordered transactional datasets such as market basket datasets. There is a lack of research employing big data processing techniques to locate frequent sequences on large-scale datasets. Furthermore, there is a need for optimized sequential pattern mining algorithms that run on ordered one-dimensional sequences. We also observe a lack of sequential pattern search studies in the literature, where the focus is centered around multi-dimensional data sequences. Existing approaches that deal with ordered one-dimensional datasets suffer from scalability issues as the amount of data to be analyzed is enormous. This research investigates the big data processing techniques used to find frequent sequences in large-scale datasets. It also proposes a scalable sequence pattern mining algorithm called Sequential Pattern Acquisition by Reducing Search Space (SPARSS) designed for distributed data processing systems that efficiently handle large datasets containing sequential one-element data. It introduces a prototype implementation of SPARSS and provides information on the SPARSS's memory and time requirements, which were calculated as part of experimental studies on a real-world dataset. The results confirm our expectations and demonstrate SPARSS's superior scalability and run-time efficiency compared to other distributed algorithms.en
dc.description.urihttps://doi.org/10.1002/cpe.7660
dc.identifier.doi10.1002/cpe.7660
dc.identifier.eissn1532-0634
dc.identifier.issn1532-0626
dc.identifier.issue24
dc.identifier.urihttps://hdl.handle.net/20.500.14981/65184
dc.identifier.volume35
dc.identifier.wos000934843100001
dc.language.isoeng
dc.publisherWILEY
dc.relation.ispartofCONCURRENCY AND COMPUTATION-PRACTICE & EXPERIENCE
dc.subjectApache Spark
dc.subjectbig data
dc.subjectdistributed systems
dc.subjectDLA
dc.subjectGSP
dc.subjectPrefixSpan
dc.subjectsequential pattern mining
dc.subjectMINING SEQUENTIAL PATTERNS
dc.subjectPROGRAMMING-MODEL
dc.subjectPARALLEL
dc.subjectComputer Science
dc.titleOn the big data processing algorithms for finding frequent sequences
dc.typeArticle
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar