Yayın: Whisper, Translate, Speak, Sync: Video Translation for Multilingual Video Conferencing Using Generative AI
| dc.contributor.author | Oskooei, Amirkia Rafiei | |
| dc.contributor.author | Caglar, Eren | |
| dc.contributor.author | Sahini, Ibrahim | |
| dc.contributor.author | Kayabayi, Ayse | |
| dc.contributor.author | Aktas, Mehmet S. | |
| dc.date.accessioned | 2026-06-27T15:23:38Z | |
| dc.date.issued | 2026 | |
| dc.description.abstract | This paper addresses the growing need for seamless communication in multilingual video conferencing by presenting a novel, computationally efficient methodology for real-time video translation. While advancements in neural networks have enabled accurate speech translation and voice cloning, integrating these with lip synchronization for realistic talking head generation remains a challenge, particularly for real-time applications. This paper introduces a comprehensive video translation pipeline leveraging open-source deep learning models. We further propose a scalable system architecture incorporating a Token Ring mechanism to manage speaker turns and minimize computational load, addressing key challenges related to latency, scalability, and personalization in multilingual settings. A segmented batched processing protocol with inverse throughput thresholding and overlapping buffering is implemented to achieve near real-time performance. A simplified, universal prototype is developed to demonstrate the feasibility and efficacy of our approach, providing a foundation for building next-generation multilingual video conferencing systems. This work offers a practical framework for developers and businesses aiming to create inclusive and effective communication platforms. | en |
| dc.description.uri | https://doi.org/10.1007/978-3-031-97606-3_15 | |
| dc.identifier.doi | 10.1007/978-3-031-97606-3_15 | |
| dc.identifier.eissn | 1611-3349 | |
| dc.identifier.endpage | 234 | |
| dc.identifier.isbn | 978-3-031-97605-6; 978-3-031-97606-3 | |
| dc.identifier.issn | 0302-9743 | |
| dc.identifier.startpage | 217 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.14981/70442 | |
| dc.identifier.volume | 15890 | |
| dc.identifier.wos | 001564030700015 | |
| dc.language.iso | eng | |
| dc.publisher | SPRINGER INTERNATIONAL PUBLISHING AG | |
| dc.relation.conference | 25th International Conference on Computational Science and Applications-ICCSA-Annual | |
| dc.relation.ispartof | COMPUTATIONAL SCIENCE AND ITS APPLICATIONS-ICCSA 2025 WORKSHOPS, PT V | |
| dc.subject | Video Translation | |
| dc.subject | Computer Vision | |
| dc.subject | Deep Learning | |
| dc.subject | Generative AI | |
| dc.subject | Human-AI Interaction | |
| dc.subject | Video Conferencing | |
| dc.subject | SERVICES | |
| dc.subject | Computer Science | |
| dc.title | Whisper, Translate, Speak, Sync: Video Translation for Multilingual Video Conferencing Using Generative AI | |
| dc.type | Proceedings Paper | |
| dspace.entity.type | Publication | |
| local.import.source | WOS |