Yayın:
TR-MMLU Benchmark for Large Language Models: Performance Evaluation, Challenges, and Opportunities for Improvement

Yükleniyor...
Küçük Resim

Tarih

Kurum Yazarları

Danışman

item.page.editor

Editör

Bölüm / Program

Dergi Başlığı

Dergi ISSN

Cilt Başlığı

Yayıncı

IEEE

DOI

10.1109/siu66497.2025.11112154
View PlumX Details

Araştırma Projeleri

Akademik Birimler

Dergi Sayısı

Özet

Language models have made significant advancements in understanding and generating human language, achieving remarkable success in various applications. However, evaluating these models remains a challenge, particularly for resource-limited languages like Turkish. To address this issue, we introduce the Turkish MMLU (TR-MMLU) benchmark, a comprehensive evaluation framework designed to assess the linguistic and conceptual capabilities of large language models (LLMs) in Turkish. TR-MMLU is based on a meticulously curated dataset comprising 6,200 multiple-choice questions across 62 sections within the Turkish education system. This benchmark provides a standard framework for Turkish NLP research, enabling detailed analyses of LLMs' capabilities in processing Turkish text. In this study, we evaluated state-of-the-art LLMs on TR-MMLU, highlighting areas for improvement in model design. TR-MMLU sets a new standard for advancing Turkish NLP research and inspiring future innovations.

Tanım

Dergi veya Seri

2025 33RD SIGNAL PROCESSING AND COMMUNICATIONS APPLICATIONS CONFERENCE, SIU

ISSN

2165-0608

ISBN

979-8-3315-6656-2; 979-8-3315-6655-5

Alıntı

Koleksiyonlar

Onay

Gözden geçir

Tamamlayıcı Bilgiler

Referans Gösteren

Related Patent

Related Goal

0

Views

0

Downloads