Yayın:
A Model-Based Evaluation Metric for Question Answering Systems

dc.contributor.authorBakir, Dilan
dc.contributor.authorAktas, Mehmet S.
dc.contributor.authorYildiz, Beytullah
dc.date.accessioned2026-06-27T15:12:56Z
dc.date.issued2025
dc.description.abstractThe paper addresses the limitations of traditional evaluation metrics for Question Answering (QA) systems that primarily focus on syntax and n-gram similarity. We propose a novel model-based evaluation metric, MQA-metric, and create a human-judgment-based dataset, squad-qametric and marco-qametric, to validate our approach. The research aims to solve several key problems: the objectivity in dataset labeling, the effectiveness of metrics when there is no syntax similarity, the impact of answer length on metric performance, and the influence of real answer quality on metric results. To tackle these challenges, we designed an interface for dataset labeling and conducted extensive experiments with human reviewers. Our analysis shows that the MQA-metric outperforms traditional metrics like BLEU, ROUGE and METEOR. Unlike existing metrics, MQA-metric leverages semantic comprehension through large language models (LLMs), enabling it to capture contextual nuances and synonymous expressions more effectively. This approach sets a standard for evaluating QA systems by prioritizing semantic accuracy over surface-level similarities. The proposed metric correlates better with human judgment, making it a more reliable tool for evaluating QA systems. Our contributions include the development of a robust evaluation workflow, creation of high-quality datasets, and an extensive comparison with existing evaluation methods. The results indicate that our model-based approach provides a significant improvement in assessing the quality of QA systems, which is crucial for their practical application and trustworthiness.en
dc.description.urihttps://doi.org/10.1142/s0218194025500032
dc.identifier.doi10.1142/s0218194025500032
dc.identifier.eissn1793-6403
dc.identifier.endpage262
dc.identifier.issn0218-1940
dc.identifier.issue2
dc.identifier.startpage243
dc.identifier.urihttps://hdl.handle.net/20.500.14981/69048
dc.identifier.volume35
dc.identifier.wos001405946200001
dc.language.isoeng
dc.publisherWORLD SCIENTIFIC PUBL CO PTE LTD
dc.relation.ispartofINTERNATIONAL JOURNAL OF SOFTWARE ENGINEERING AND KNOWLEDGE ENGINEERING
dc.subjectQuestion answering
dc.subjectgenerative model
dc.subjectevaluation metric
dc.subjectnatural language processing
dc.subjecttransformer models
dc.subjectlarge language model
dc.subjectComputer Science
dc.subjectEngineering
dc.titleA Model-Based Evaluation Metric for Question Answering Systems
dc.typeArticle
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar