Generative AI for Video Translation: A Scalable Architecture for Multilingual Video Conferencing
Fuente:
arXiv
Guardado en:
| Autores principales: | Oskooei, Amirkia Rafiei, Caglar, Eren, Sahin, Ibrahim, Kayabay, Ayse, Aktas, Mehmet S. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Asynchronous Pipeline Parallelism for Real-Time Multilingual Lip Synchronization in Video Communication Systems
por: Caglar, Eren, et al.
Publicado: (2025)
por: Caglar, Eren, et al.
Publicado: (2025)
BreakFun: Jailbreaking LLMs via Schema Exploitation
por: Oskooei, Amirkia Rafiei, et al.
Publicado: (2025)
por: Oskooei, Amirkia Rafiei, et al.
Publicado: (2025)
Natural Language Summarization Enables Multi-Repository Bug Localization by LLMs in Microservice Architectures
por: Oskooei, Amirkia Rafiei, et al.
Publicado: (2025)
por: Oskooei, Amirkia Rafiei, et al.
Publicado: (2025)
StyleMM: Stylized 3D Morphable Face Model via Text-Driven Aligned Image Translation
por: Lee, Seungmi, et al.
Publicado: (2025)
por: Lee, Seungmi, et al.
Publicado: (2025)
When Many-Shot Prompting Fails: An Empirical Study of LLM Code Translation
por: Oskooei, Amirkia Rafiei, et al.
Publicado: (2025)
por: Oskooei, Amirkia Rafiei, et al.
Publicado: (2025)
Cytoarchitecture in Words: Weakly Supervised Vision-Language Modeling for Human Brain Microscopy
por: Sutton, Matthew, et al.
Publicado: (2026)
por: Sutton, Matthew, et al.
Publicado: (2026)
Lightweight Complementary-Cue Fusion for Robust Video Face Forgery Detection
por: Baek, Sunghwan, et al.
Publicado: (2026)
por: Baek, Sunghwan, et al.
Publicado: (2026)
Detecting AI-Generated Videos with Spiking Neural Networks
por: Jang, Minsuk, et al.
Publicado: (2026)
por: Jang, Minsuk, et al.
Publicado: (2026)
Multimodal Ensemble with Conditional Feature Fusion for Dysgraphia Diagnosis in Children from Handwriting Samples
por: Kunhoth, Jayakanth, et al.
Publicado: (2024)
por: Kunhoth, Jayakanth, et al.
Publicado: (2024)
Relightable and Dynamic Gaussian Avatar Reconstruction from Monocular Video
por: Choi, Seonghwa, et al.
Publicado: (2025)
por: Choi, Seonghwa, et al.
Publicado: (2025)
GDDS: A Single Domain Generalized Defect Detection Frame of Open World Scenario using Gather and Distribute Domain-shift Suppression Network
por: Chen, Haiyong, et al.
Publicado: (2024)
por: Chen, Haiyong, et al.
Publicado: (2024)
Gaze-Guided Learning: Avoiding Shortcut Bias in Visual Classification
por: Li, Jiahang, et al.
Publicado: (2025)
por: Li, Jiahang, et al.
Publicado: (2025)
Understanding Identity Continuity in Thermal Video through Scene-Level Consistency
por: Sun, Wei-Chieh, et al.
Publicado: (2026)
por: Sun, Wei-Chieh, et al.
Publicado: (2026)
AVControl: Efficient Framework for Training Audio-Visual Controls
por: Ben-Yosef, Matan, et al.
Publicado: (2026)
por: Ben-Yosef, Matan, et al.
Publicado: (2026)
Cross-Domain Adversarial Augmentation: Stabilizing GANs for Medical and Handwriting Data Scarcity
por: Soad, Md. Sohanuzzaman, et al.
Publicado: (2026)
por: Soad, Md. Sohanuzzaman, et al.
Publicado: (2026)
Detecting Inpainted Video with Frequency Domain Insights
por: Tang, Quanhui, et al.
Publicado: (2024)
por: Tang, Quanhui, et al.
Publicado: (2024)
Human-Centric Anomaly Detection in Surveillance Videos Using YOLO-World and Spatio-Temporal Deep Learning
por: Naeen, Mohammad Ali Etemadi, et al.
Publicado: (2025)
por: Naeen, Mohammad Ali Etemadi, et al.
Publicado: (2025)
Motion Attribution for Video Generation
por: Wu, Xindi, et al.
Publicado: (2026)
por: Wu, Xindi, et al.
Publicado: (2026)
Text-Driven 3D Hand Motion Generation from Sign Language Data
por: Bensabath, Léore, et al.
Publicado: (2025)
por: Bensabath, Léore, et al.
Publicado: (2025)
Data Augmentation with Diffusion Models for Colon Polyp Localization on the Low Data Regime: How much real data is enough?
por: Tormos, Adrian, et al.
Publicado: (2024)
por: Tormos, Adrian, et al.
Publicado: (2024)
ADAT: Time-Series-Aware Adaptive Transformer Architecture for Sign Language Translation
por: Shahin, Nada, et al.
Publicado: (2025)
por: Shahin, Nada, et al.
Publicado: (2025)
Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models
por: Seo, Huichan, et al.
Publicado: (2025)
por: Seo, Huichan, et al.
Publicado: (2025)
TAG-WM: Tamper-Aware Generative Image Watermarking via Diffusion Inversion Sensitivity
por: Chen, Yuzhuo, et al.
Publicado: (2025)
por: Chen, Yuzhuo, et al.
Publicado: (2025)
Do Inpainting Yourself: Generative Facial Inpainting Guided by Exemplars
por: Lu, Wanglong, et al.
Publicado: (2022)
por: Lu, Wanglong, et al.
Publicado: (2022)
Explainable Classifier for Malignant Lymphoma Subtyping via Cell Graph and Image Fusion
por: Nishiyama, Daiki, et al.
Publicado: (2025)
por: Nishiyama, Daiki, et al.
Publicado: (2025)
Personalized QoE Prediction: A Demographic-Augmented Machine Learning Framework for 5G Video Streaming Networks
por: Ahmed, Syeda Zunaira, et al.
Publicado: (2025)
por: Ahmed, Syeda Zunaira, et al.
Publicado: (2025)
A Spatio-Temporal Deep Learning Approach For High-Resolution Gridded Monsoon Prediction
por: Borah, Parashjyoti, et al.
Publicado: (2026)
por: Borah, Parashjyoti, et al.
Publicado: (2026)
GLoT: A Novel Gated-Logarithmic Transformer for Efficient Sign Language Translation
por: Shahin, Nada, et al.
Publicado: (2025)
por: Shahin, Nada, et al.
Publicado: (2025)
GAN-based Content-Conditioned Generation of Handwritten Musical Symbols
por: Asbert, Gerard, et al.
Publicado: (2025)
por: Asbert, Gerard, et al.
Publicado: (2025)
Capacity Constraint Analysis Using Object Detection for Smart Manufacturing
por: Ahmad, Hafiz Mughees, et al.
Publicado: (2024)
por: Ahmad, Hafiz Mughees, et al.
Publicado: (2024)
SH17: A Dataset for Human Safety and Personal Protective Equipment Detection in Manufacturing Industry
por: Ahmad, Hafiz Mughees, et al.
Publicado: (2024)
por: Ahmad, Hafiz Mughees, et al.
Publicado: (2024)
AnyMoLe: Any Character Motion In-betweening Leveraging Video Diffusion Models
por: Yun, Kwan, et al.
Publicado: (2025)
por: Yun, Kwan, et al.
Publicado: (2025)
MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation
por: Bartkowiak, Patryk, et al.
Publicado: (2026)
por: Bartkowiak, Patryk, et al.
Publicado: (2026)
Vision transformers in domain adaptation and domain generalization: a study of robustness
por: Alijani, Shadi, et al.
Publicado: (2024)
por: Alijani, Shadi, et al.
Publicado: (2024)
Scene Detection Policies and Keyframe Extraction Strategies for Large-Scale Video Analysis
por: Korolkov, Vasilii
Publicado: (2025)
por: Korolkov, Vasilii
Publicado: (2025)
Visual Style Prompt Learning Using Diffusion Models for Blind Face Restoration
por: Lu, Wanglong, et al.
Publicado: (2024)
por: Lu, Wanglong, et al.
Publicado: (2024)
FACEMUG: A Multimodal Generative and Fusion Framework for Local Facial Editing
por: Lu, Wanglong, et al.
Publicado: (2024)
por: Lu, Wanglong, et al.
Publicado: (2024)
Interpreting Structured Perturbations in Image Protection Methods for Diffusion Models
por: Martin, Michael R., et al.
Publicado: (2025)
por: Martin, Michael R., et al.
Publicado: (2025)
SpectraNet: FFT-assisted Deep Learning Classifier for Deepfake Face Detection
por: Jayarathne, Nithira, et al.
Publicado: (2025)
por: Jayarathne, Nithira, et al.
Publicado: (2025)
Listen to the Unexpected: Self-Supervised Surprise Detection for Efficient Viewport Prediction
por: Khah, Arman Nik, et al.
Publicado: (2026)
por: Khah, Arman Nik, et al.
Publicado: (2026)
Ejemplares similares
-
Asynchronous Pipeline Parallelism for Real-Time Multilingual Lip Synchronization in Video Communication Systems
por: Caglar, Eren, et al.
Publicado: (2025) -
BreakFun: Jailbreaking LLMs via Schema Exploitation
por: Oskooei, Amirkia Rafiei, et al.
Publicado: (2025) -
Natural Language Summarization Enables Multi-Repository Bug Localization by LLMs in Microservice Architectures
por: Oskooei, Amirkia Rafiei, et al.
Publicado: (2025) -
StyleMM: Stylized 3D Morphable Face Model via Text-Driven Aligned Image Translation
por: Lee, Seungmi, et al.
Publicado: (2025) -
When Many-Shot Prompting Fails: An Empirical Study of LLM Code Translation
por: Oskooei, Amirkia Rafiei, et al.
Publicado: (2025)