Yayın: Generative AI for Video Translation: A Scalable Architecture for Multilingual Video Conferencing
| dc.contributor.author | Rafiei Oskooei, Amirkia | |
| dc.contributor.author | Caglar, Eren | |
| dc.contributor.author | Sahin, Ibrahim | |
| dc.contributor.author | Kayabay, Ayse | |
| dc.contributor.author | Aktas, Mehmet S. | |
| dc.date.accessioned | 2026-06-27T15:23:48Z | |
| dc.date.issued | 2025 | |
| dc.description.abstract | The real-time deployment of cascaded generative AI pipelines for applications like video translation is constrained by significant system-level challenges. These include the cumulative latency of sequential model inference and the quadratic (O(N-2)) computational complexity that renders multi-user video conferencing applications unscalable. This paper proposes and evaluates a practical system-level framework designed to mitigate these critical bottlenecks. The proposed architecture incorporates a turn-taking mechanism to reduce computational complexity from quadratic to linear in multi-user scenarios, and a segmented processing protocol to manage inference latency for a perceptually real-time experience. We implement a proof-of-concept pipeline and conduct a rigorous performance analysis across a multi-tiered hardware setup, including commodity (NVIDIA RTX 4060), cloud (NVIDIA T4), and enterprise (NVIDIA A100) GPUs. Our objective evaluation demonstrates that the system achieves real-time throughput (tau<1.0) on modern hardware. A subjective user study further validates the approach, showing that a predictable, initial processing delay is highly acceptable to users in exchange for a smooth, uninterrupted playback experience. The work presents a validated, end-to-end system design that offers a practical roadmap for deploying scalable, real-time generative AI applications in multilingual communication platforms. | en |
| dc.description.uri | https://doi.org/10.3390/app152312691 | |
| dc.identifier.doi | 10.3390/app152312691 | |
| dc.identifier.eissn | 2076-3417 | |
| dc.identifier.issue | 23 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.14981/70475 | |
| dc.identifier.volume | 15 | |
| dc.identifier.wos | 001635126500001 | |
| dc.language.iso | eng | |
| dc.publisher | MDPI | |
| dc.relation.ispartof | APPLIED SCIENCES-BASEL | |
| dc.rights | openAccess | |
| dc.subject | generative AI | |
| dc.subject | applied computer vision | |
| dc.subject | multimedia | |
| dc.subject | human-AI interaction | |
| dc.subject | deep learning | |
| dc.subject | Chemistry | |
| dc.subject | Engineering | |
| dc.subject | Materials Science | |
| dc.subject | Physics | |
| dc.title | Generative AI for Video Translation: A Scalable Architecture for Multilingual Video Conferencing | |
| dc.type | Article | |
| dspace.entity.type | Publication | |
| local.import.source | WOS |