Yayın: TR-MMLU Benchmark for Large Language Models: Performance Evaluation, Challenges, and Opportunities for Improvement
| dc.contributor.author | Bayram, M. Ali | |
| dc.contributor.author | Fincan, Ali Arda | |
| dc.contributor.author | Gumus, Ahmet Semih | |
| dc.contributor.author | Diri, Banu | |
| dc.contributor.author | Yildirim, Savas | |
| dc.contributor.author | Aytas, Oner | |
| dc.date.accessioned | 2026-06-27T15:32:21Z | |
| dc.date.issued | 2025 | |
| dc.description.abstract | Language models have made significant advancements in understanding and generating human language, achieving remarkable success in various applications. However, evaluating these models remains a challenge, particularly for resource-limited languages like Turkish. To address this issue, we introduce the Turkish MMLU (TR-MMLU) benchmark, a comprehensive evaluation framework designed to assess the linguistic and conceptual capabilities of large language models (LLMs) in Turkish. TR-MMLU is based on a meticulously curated dataset comprising 6,200 multiple-choice questions across 62 sections within the Turkish education system. This benchmark provides a standard framework for Turkish NLP research, enabling detailed analyses of LLMs' capabilities in processing Turkish text. In this study, we evaluated state-of-the-art LLMs on TR-MMLU, highlighting areas for improvement in model design. TR-MMLU sets a new standard for advancing Turkish NLP research and inspiring future innovations. | en |
| dc.description.uri | https://doi.org/10.1109/siu66497.2025.11112154 | |
| dc.identifier.doi | 10.1109/siu66497.2025.11112154 | |
| dc.identifier.isbn | 979-8-3315-6656-2; 979-8-3315-6655-5 | |
| dc.identifier.issn | 2165-0608 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.14981/71694 | |
| dc.identifier.wos | 001575462500215 | |
| dc.language.iso | tur | |
| dc.publisher | IEEE | |
| dc.relation.conference | 33rd Conference on Signal Processing and Communications Applications-SIU-Annual | |
| dc.relation.ispartof | 2025 33RD SIGNAL PROCESSING AND COMMUNICATIONS APPLICATIONS CONFERENCE, SIU | |
| dc.rights | openAccess | |
| dc.subject | Large Language Models (LLM) | |
| dc.subject | Natural Language Processing (NLP) | |
| dc.subject | Artificial Intelligence | |
| dc.subject | Turkish NLP | |
| dc.subject | Computer Science | |
| dc.subject | Engineering | |
| dc.subject | Telecommunications | |
| dc.title | TR-MMLU Benchmark for Large Language Models: Performance Evaluation, Challenges, and Opportunities for Improvement | |
| dc.type | Proceedings Paper | |
| dspace.entity.type | Publication | |
| local.import.source | WOS |