Yayın:
A Multimodal Transformer-Based Framework for Emotion Analysis in Multilingual Video Content

dc.contributor.authorYakut, Sehmus
dc.contributor.authorTuten, Yusuf Taha
dc.contributor.authorCaglar, Eren
dc.contributor.authorAktas, Mehmet S.
dc.date.accessioned2026-06-27T15:30:31Z
dc.date.issued2026
dc.description.abstractThis research addresses the challenge of inferring complex psychological states, including stress, fatigue, anxiety, cognitive load, and boredom, from facial expressions. We propose an interpretable, literature-informed emotion-weighting methodology that transforms the eight-emotion probability outputs of facial emotion recognition models into continuous estimates of these five psychological states using weights derived from the Valence-Arousal framework, providing a principled bridge between discrete emotion predictions and higher-level affective constructs. The proposed formulation is evaluated across six representative deep learning architectures-a baseline CNN (ResNet-50), a modern CNN (ConvNeXt), a hybrid attention-based model (DDAMFN), and three Transformer-based models (ViT, BEiT, and Swin). Our results demonstrate that strong performance on discrete FER tasks does not directly translate to consistent behavior in complex state inference; instead, architectures capable of preserving subtle and distributed affective cues yield more stable and interpretable state estimates, with DDAMFN and Vision Transformer models exhibiting the most consistent performance across the evaluated psychological states. These findings highlight the central role of the proposed emotion-weighting formulation and the importance of architecture selection beyond categorical accuracy in complex affective state analysis.en
dc.description.urihttps://doi.org/10.3390/computers15020077
dc.identifier.doi10.3390/computers15020077
dc.identifier.issn2073-431X
dc.identifier.issue2
dc.identifier.urihttps://hdl.handle.net/20.500.14981/71327
dc.identifier.volume15
dc.identifier.wos001699931900001
dc.language.isoeng
dc.publisherMDPI
dc.relation.ispartofCOMPUTERS
dc.rightsopenAccess
dc.subjectfacial expression recognition (FER)
dc.subjectcomputer vision
dc.subjectdeep learning
dc.subjectCNN
dc.subjecttransformer
dc.subjectemotion recognition
dc.subjectstress
dc.subjectfatigue
dc.subjectboredom
dc.subjectanxiety
dc.subjectcognitive load
dc.subjectComputer Science
dc.titleA Multimodal Transformer-Based Framework for Emotion Analysis in Multilingual Video Content
dc.typeArticle
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar