Yayın:
Generative AI for Video Translation: A Scalable Architecture for Multilingual Video Conferencing

dc.contributor.authorRafiei Oskooei, Amirkia
dc.contributor.authorCaglar, Eren
dc.contributor.authorSahin, Ibrahim
dc.contributor.authorKayabay, Ayse
dc.contributor.authorAktas, Mehmet S.
dc.date.accessioned2026-06-27T15:23:48Z
dc.date.issued2025
dc.description.abstractThe real-time deployment of cascaded generative AI pipelines for applications like video translation is constrained by significant system-level challenges. These include the cumulative latency of sequential model inference and the quadratic (O(N-2)) computational complexity that renders multi-user video conferencing applications unscalable. This paper proposes and evaluates a practical system-level framework designed to mitigate these critical bottlenecks. The proposed architecture incorporates a turn-taking mechanism to reduce computational complexity from quadratic to linear in multi-user scenarios, and a segmented processing protocol to manage inference latency for a perceptually real-time experience. We implement a proof-of-concept pipeline and conduct a rigorous performance analysis across a multi-tiered hardware setup, including commodity (NVIDIA RTX 4060), cloud (NVIDIA T4), and enterprise (NVIDIA A100) GPUs. Our objective evaluation demonstrates that the system achieves real-time throughput (tau<1.0) on modern hardware. A subjective user study further validates the approach, showing that a predictable, initial processing delay is highly acceptable to users in exchange for a smooth, uninterrupted playback experience. The work presents a validated, end-to-end system design that offers a practical roadmap for deploying scalable, real-time generative AI applications in multilingual communication platforms.en
dc.description.urihttps://doi.org/10.3390/app152312691
dc.identifier.doi10.3390/app152312691
dc.identifier.eissn2076-3417
dc.identifier.issue23
dc.identifier.urihttps://hdl.handle.net/20.500.14981/70475
dc.identifier.volume15
dc.identifier.wos001635126500001
dc.language.isoeng
dc.publisherMDPI
dc.relation.ispartofAPPLIED SCIENCES-BASEL
dc.rightsopenAccess
dc.subjectgenerative AI
dc.subjectapplied computer vision
dc.subjectmultimedia
dc.subjecthuman-AI interaction
dc.subjectdeep learning
dc.subjectChemistry
dc.subjectEngineering
dc.subjectMaterials Science
dc.subjectPhysics
dc.titleGenerative AI for Video Translation: A Scalable Architecture for Multilingual Video Conferencing
dc.typeArticle
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar