Yayın:
Can One Model Fit All? An Exploration of Wav2Lip's Lip-Syncing Generalizability Across Culturally Distinct Languages

dc.contributor.authorOskooei, Amirkia Rafiei
dc.contributor.authorYahsi, Ezgi
dc.contributor.authorSungur, Mehmet
dc.contributor.authorAktas, Mehmet S.
dc.date.accessioned2026-06-27T14:58:55Z
dc.date.issued2024
dc.description.abstractThis study explores the potential of Wav2Lip, a state-of-the-art lip-sync model, in multilingual environments. We assess its performance in generating lip-synchronized videos for Turkish, Persian, and Arabic languages. The evaluation results reveal promising language independence for Wav2Lip, achieving comparable accuracy to English. The research identifies the gap in research on lip-sync models for diverse languages and emphasizes the need for broader exploration. Additionally, we introduce a comprehensive Face-to-Face Translation workflow, outlining the fundamental elements for a seamless cross-lingual communication system. This work highlights the importance of Lip Sync models and the potential of Wav2Lip within such a system. By acknowledging current limitations and advocating for advancements in real-time models and high-resolution datasets, this study lays the groundwork for the development of revolutionary Face-to-Face Translation systems, fostering a future of barrier-free communication.en
dc.description.urihttps://doi.org/10.1007/978-3-031-65282-0_10
dc.identifier.doi10.1007/978-3-031-65282-0_10
dc.identifier.eissn1611-3349
dc.identifier.endpage164
dc.identifier.isbn978-3-031-65281-3; 978-3-031-65282-0
dc.identifier.issn0302-9743
dc.identifier.startpage149
dc.identifier.urihttps://hdl.handle.net/20.500.14981/66728
dc.identifier.volume14819
dc.identifier.wos001294370400010
dc.language.isoeng
dc.publisherSPRINGER INTERNATIONAL PUBLISHING AG
dc.relation.conference24th International Conference on Computational Science and Its Applications (ICCSA)
dc.relation.ispartofCOMPUTATIONAL SCIENCE AND ITS APPLICATIONS-ICCSA 2024 WORKSHOPS, PT V
dc.subjectTalking-face Generation
dc.subjectLip Sync
dc.subjectFace-to-Face Translation
dc.subjectGenerative AI
dc.subjectDeep Learning
dc.subjectComputer Vision
dc.subjectSERVICES
dc.subjectComputer Science
dc.titleCan One Model Fit All? An Exploration of Wav2Lip's Lip-Syncing Generalizability Across Culturally Distinct Languages
dc.typeProceedings Paper
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar