Yayın:
Whisper, Translate, Speak, Sync: Video Translation for Multilingual Video Conferencing Using Generative AI

dc.contributor.authorOskooei, Amirkia Rafiei
dc.contributor.authorCaglar, Eren
dc.contributor.authorSahini, Ibrahim
dc.contributor.authorKayabayi, Ayse
dc.contributor.authorAktas, Mehmet S.
dc.date.accessioned2026-06-27T15:23:38Z
dc.date.issued2026
dc.description.abstractThis paper addresses the growing need for seamless communication in multilingual video conferencing by presenting a novel, computationally efficient methodology for real-time video translation. While advancements in neural networks have enabled accurate speech translation and voice cloning, integrating these with lip synchronization for realistic talking head generation remains a challenge, particularly for real-time applications. This paper introduces a comprehensive video translation pipeline leveraging open-source deep learning models. We further propose a scalable system architecture incorporating a Token Ring mechanism to manage speaker turns and minimize computational load, addressing key challenges related to latency, scalability, and personalization in multilingual settings. A segmented batched processing protocol with inverse throughput thresholding and overlapping buffering is implemented to achieve near real-time performance. A simplified, universal prototype is developed to demonstrate the feasibility and efficacy of our approach, providing a foundation for building next-generation multilingual video conferencing systems. This work offers a practical framework for developers and businesses aiming to create inclusive and effective communication platforms.en
dc.description.urihttps://doi.org/10.1007/978-3-031-97606-3_15
dc.identifier.doi10.1007/978-3-031-97606-3_15
dc.identifier.eissn1611-3349
dc.identifier.endpage234
dc.identifier.isbn978-3-031-97605-6; 978-3-031-97606-3
dc.identifier.issn0302-9743
dc.identifier.startpage217
dc.identifier.urihttps://hdl.handle.net/20.500.14981/70442
dc.identifier.volume15890
dc.identifier.wos001564030700015
dc.language.isoeng
dc.publisherSPRINGER INTERNATIONAL PUBLISHING AG
dc.relation.conference25th International Conference on Computational Science and Applications-ICCSA-Annual
dc.relation.ispartofCOMPUTATIONAL SCIENCE AND ITS APPLICATIONS-ICCSA 2025 WORKSHOPS, PT V
dc.subjectVideo Translation
dc.subjectComputer Vision
dc.subjectDeep Learning
dc.subjectGenerative AI
dc.subjectHuman-AI Interaction
dc.subjectVideo Conferencing
dc.subjectSERVICES
dc.subjectComputer Science
dc.titleWhisper, Translate, Speak, Sync: Video Translation for Multilingual Video Conferencing Using Generative AI
dc.typeProceedings Paper
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar