Yayın:
Multi-Stream Word-Based Compression Algorithm

dc.contributor.authorOzturk, Emir
dc.contributor.authorMesut, Altan
dc.contributor.authorDiri, Banu
dc.date.accessioned2026-06-27T14:10:52Z
dc.date.issued2017
dc.description.abstractIn this article, we present a novel word-based lossless compression algorithm for text files which uses a semi-static model. We named our algorithm as Multi-stream Word-based Compression Algorithm (MWCA), because it stores the compressed forms of the words in three individual streams depending on their frequencies in the text. It also stores two dictionaries and a bit vector as a side information. In our experiments MWCA obtains compression ratio over 3,23 bpc on average and 2,88 bpc on files larger than 50 MB. If a variable length encoder like Huffman Coding is used after MWCA, given ratios will reduce to 2,63 and 2,44 bpc respectively. With the advantage of its multi-stream structure MWCA could become a good solution especially for storing and searching big text data.en
dc.identifier.endpage37
dc.identifier.isbn978-1-5386-0930-9
dc.identifier.startpage34
dc.identifier.urihttps://hdl.handle.net/20.500.14981/57643
dc.identifier.wos000426856900007
dc.language.isotur
dc.publisherIEEE
dc.relation.conference2017 International Conference on Computer Science and Engineering (UBMK)
dc.relation.ispartof2017 INTERNATIONAL CONFERENCE ON COMPUTER SCIENCE AND ENGINEERING (UBMK)
dc.subjectData compression
dc.subjectText compression
dc.subjectNATURAL-LANGUAGE TEXT
dc.subjectComputer Science
dc.titleMulti-Stream Word-Based Compression Algorithm
dc.typeProceedings Paper
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar