Yayın:
Enhanced speech emotion recognition using averaged valence arousal dominance mapping and deep neural networks

dc.contributor.authorRizhinashvili, Davit
dc.contributor.authorSham, Abdallah Hussein
dc.contributor.authorAnbarjafari, Gholamreza
dc.date.accessioned2026-06-27T15:10:36Z
dc.date.issued2024
dc.description.abstractThis study delves into advancements in speech emotion recognition (SER) by establishing a novel approach for emotion mapping and prediction using the Valence-Arousal-Dominance (VAD) model. Central to this research is the creation of reliable emotion-to-VAD mappings, achieved by averaging outcomes from multiple pre-trained networks applied to the RAVDESS dataset. This approach adeptly resolves prior inconsistencies in emotion-to-VAD mappings and establishes a dependable framework for SER. The study also introduces a refined SER model, integrating the pre-trained Wave2Vec 2.0 with Long Short-Term Memory (LSTM) networks and linear layers, culminating in an output layer representing valence, arousal, and dominance. Notably, this model exhibits commendable accuracy across various datasets, such as RAVDESS, EMO-DB, CREMA-D, and TESS, thereby showcasing its robustness and adaptability, an improvement over earlier models susceptible to dataset-specific overfitting. The research further unveils a comprehensive speech analysis application, adept at denoising, segmenting, and profiling emotions in speech segments. This application features interactive emotion tracking and sentiment reports, illustrating its practicality in diverse applications. The study recognizes ongoing challenges in SER, especially in managing the subjective nature of emotion perception and integrating multimodal data. Although the research marks a progression in SER technology, it underscores the need for continuous research and careful consideration of ethical aspects in deploying such technologies. This work contributes to the SER domain by introducing a dependable method for emotion mapping, a robust model for emotion recognition, and a user-friendly application for practical implementations.en
dc.description.urihttps://doi.org/10.1007/s11760-024-03406-8
dc.identifier.doi10.1007/s11760-024-03406-8
dc.identifier.eissn1863-1711
dc.identifier.endpage7454
dc.identifier.issn1863-1703
dc.identifier.issue10
dc.identifier.startpage7445
dc.identifier.urihttps://hdl.handle.net/20.500.14981/68608
dc.identifier.volume18
dc.identifier.wos001268845200002
dc.language.isoeng
dc.publisherSPRINGER LONDON LTD
dc.relation.ispartofSIGNAL IMAGE AND VIDEO PROCESSING
dc.rightsopenAccess
dc.subjectSpeech emotion recognition
dc.subjectDeep neural networks
dc.subjectLSTM
dc.subjectSpeech analysis
dc.subjectValence arousal dominance
dc.subjectFEATURES
dc.subjectEngineering
dc.subjectImaging Science & Photographic Technology
dc.titleEnhanced speech emotion recognition using averaged valence arousal dominance mapping and deep neural networks
dc.typeArticle
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar