Yayın:
Tokenization Standards and Evaluation in Natural Language Processing: A Comparative Analysis of Large Language Models on Turkish

dc.contributor.authorBayram, M. Ali
dc.contributor.authorFincan, Ali Arda
dc.contributor.authorGumus, Ahmet Semih
dc.contributor.authorKarakas, Sercan
dc.contributor.authorDiri, Banu
dc.contributor.authorYildirim, Savas
dc.date.accessioned2026-06-27T15:32:31Z
dc.date.issued2025
dc.description.abstractTokenization is a fundamental preprocessing step in Natural Language Processing (NLP), significantly impacting the capability of large language models (LLMs) to capture linguistic and semantic nuances. This study introduces a novel evaluation framework addressing tokenization challenges specific to morphologically-rich and low-resource languages such as Turkish. Utilizing the Turkish MMLU (TR-MMLU) dataset, comprising 6,200 multiple-choice questions from the Turkish education system, we assessed tokenizers based on vocabulary size, token count, processing time, language-specific token percentages (%TR), and token purity (%Pure). These newly proposed metrics measure how effectively tokenizers preserve linguistic structures. Our analysis reveals that language-specific token percentages exhibit a stronger correlation with downstream performance (e.g., MMLU scores) than token purity. Furthermore, increasing model parameters alone does not necessarily enhance linguistic performance, underscoring the importance of tailored, language-specific tokenization methods. The proposed framework establishes robust and practical tokenization standards for morphologically complex languages.en
dc.description.urihttps://doi.org/10.1109/siu66497.2025.11112220
dc.identifier.doi10.1109/siu66497.2025.11112220
dc.identifier.isbn979-8-3315-6656-2; 979-8-3315-6655-5
dc.identifier.issn2165-0608
dc.identifier.urihttps://hdl.handle.net/20.500.14981/71725
dc.identifier.wos001575462500250
dc.language.isotur
dc.publisherIEEE
dc.relation.conference33rd Conference on Signal Processing and Communications Applications-SIU-Annual
dc.relation.ispartof2025 33RD SIGNAL PROCESSING AND COMMUNICATIONS APPLICATIONS CONFERENCE, SIU
dc.rightsopenAccess
dc.subjectTokenization
dc.subjectLarge Language Models (LLM)
dc.subjectNatural Language Processing (NLP)
dc.subjectTurkish NLP
dc.subjectComputer Science
dc.subjectEngineering
dc.subjectTelecommunications
dc.titleTokenization Standards and Evaluation in Natural Language Processing: A Comparative Analysis of Large Language Models on Turkish
dc.typeProceedings Paper
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar