Yayın:
A Novel Hybrid Large Language Model Approach for Reporting Panoramic Radiographs and Performance Comparison with Current Large Language Models

dc.contributor.authorBalel, Yunus
dc.contributor.authorSagtas, Kaan
dc.contributor.authorTeke, Fatih
dc.contributor.authorKurt, Mehmet Ali
dc.date.accessioned2026-06-27T15:31:51Z
dc.date.issued2026
dc.description.abstractLarge language models (LLMs) show potential in clinical reporting, yet current multimodal systems remain unreliable for interpreting panoramic radiographs due to limited visual diagnostic accuracy and high hallucination rates. This study introduces a hybrid framework that integrates a deep learning-based image analysis model with LLM-driven reporting to enhance reliability in dental radiology. A YOLOv12 model was trained on 30,954 panoramic radiographs (70% training, 15% validation, 15% testing) for tooth detection and 14-category segmentation. Detection and segmentation outputs were converted into structured JSON data and processed by locally hosted LLMs (DeepSeek R1, Mistral, Llama 3.2, Gemma 3, Qwen3, SmolLM3). Performance metrics included structural validity, consistency, response latency, and token length. Reporting accuracy was evaluated on 50 unseen radiographs, with expert assessment serving as the gold standard. The tooth-numbering model achieved Precision = 0.651, Recall = 0.699, and F1 = 0.674. The segmentation model achieved overall Precision = 0.816, Recall = 0.626, and F1 = 0.708, with highest F1-scores for ectopic/supernumerary teeth (0.994), impacted teeth (0.990), and implants (0.984). All hybrid LLMs produced structurally valid JSON outputs (100%). DeepSeek R1 showed the highest reporting accuracy (466 True findings), followed by Mistral (462), Llama 3.2 (442), and Gemma 3 (436). Hallucination counts were lowest in DeepSeek R1 (30) and highest in Gemma 3 (60). Commercial LLMs (ChatGPT-5, Gemini 2.5 Pro, DeepSeek R1-cloud) exhibited hallucinations in 100% of reports. Integrating structured image-derived findings with LLM reasoning markedly improves reporting accuracy and minimizes hallucinations. The hybrid framework outperforms commercial LLMs and represents a promising, reliable solution for AI-assisted dental radiographic interpretation.en
dc.description.urihttps://doi.org/10.1007/s10278-026-01880-9
dc.identifier.doi10.1007/s10278-026-01880-9
dc.identifier.eissn2948-2933
dc.identifier.issn2948-2925
dc.identifier.pubmed41741852
dc.identifier.urihttps://hdl.handle.net/20.500.14981/71594
dc.identifier.wos001699093000001
dc.language.isoeng
dc.publisherSPRINGER
dc.relation.ispartofJOURNAL OF IMAGING INFORMATICS IN MEDICINE
dc.subjectHybrid large language model
dc.subjectPanoramic dental radiography
dc.subjectYOLOv12 segmentation
dc.subjectRadiology report generation
dc.subjectDental artificial intelligence
dc.subjectRadiology, Nuclear Medicine & Medical Imaging
dc.titleA Novel Hybrid Large Language Model Approach for Reporting Panoramic Radiographs and Performance Comparison with Current Large Language Models
dc.typeArticle; Early Access
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar