Shared Latent Representation for Joint Text-to-Audio-Visual Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Yaman, Dogucan, Akti, Seymanur, Eyiokur, Fevziye Irem, Waibel, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Audio-Visual Speech Representation Expert for Enhanced Talking Face Video Generation and Evaluation
by: Yaman, Dogucan, et al.
Published: (2024)
by: Yaman, Dogucan, et al.
Published: (2024)
Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation
by: Yaman, Dogucan, et al.
Published: (2025)
by: Yaman, Dogucan, et al.
Published: (2025)
CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding
by: Eyiokur, Fevziye Irem, et al.
Published: (2025)
by: Eyiokur, Fevziye Irem, et al.
Published: (2025)
Audio-driven Talking Face Generation with Stabilized Synchronization Loss
by: Yaman, Dogucan, et al.
Published: (2023)
by: Yaman, Dogucan, et al.
Published: (2023)
A Multimodal Depth-Aware Method For Embodied Reference Understanding
by: Eyiokur, Fevziye Irem, et al.
Published: (2025)
by: Eyiokur, Fevziye Irem, et al.
Published: (2025)
Assessing Identity Leakage in Talking Face Generation: Metrics and Evaluation Framework
by: Yaman, Dogucan, et al.
Published: (2025)
by: Yaman, Dogucan, et al.
Published: (2025)
Lombard Speech Synthesis for Any Voice with Controllable Style Embeddings
by: Akti, Seymanur, et al.
Published: (2026)
by: Akti, Seymanur, et al.
Published: (2026)
Titanic Calling: Low Bandwidth Video Conference from the Titanic Wreck
by: Eyiokur, Fevziye Irem, et al.
Published: (2024)
by: Eyiokur, Fevziye Irem, et al.
Published: (2024)
AV-Flow: Transforming Text to Audio-Visual Human-like Interactions
by: Chatziagapi, Aggelina, et al.
Published: (2025)
by: Chatziagapi, Aggelina, et al.
Published: (2025)
AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
LTX-2: Efficient Joint Audio-Visual Foundation Model
by: HaCohen, Yoav, et al.
Published: (2026)
by: HaCohen, Yoav, et al.
Published: (2026)
JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
by: Chao, Jianghan, et al.
Published: (2025)
by: Chao, Jianghan, et al.
Published: (2025)
Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion
by: Akti, Seymanur, et al.
Published: (2025)
by: Akti, Seymanur, et al.
Published: (2025)
Joint Audio-Visual Idling Vehicle Detection with Streamlined Input Dependencies
by: Li, Xiwen, et al.
Published: (2024)
by: Li, Xiwen, et al.
Published: (2024)
TAViS: Text-bridged Audio-Visual Segmentation with Foundation Models
by: Luo, Ziyang, et al.
Published: (2025)
by: Luo, Ziyang, et al.
Published: (2025)
Unifying Latent and Lexicon Representations for Effective Video-Text Retrieval
by: Liu, Haowei, et al.
Published: (2024)
by: Liu, Haowei, et al.
Published: (2024)
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
by: Dai, Yifan, et al.
Published: (2026)
by: Dai, Yifan, et al.
Published: (2026)
Language-Guided Joint Audio-Visual Editing via One-Shot Adaptation
by: Liang, Susan, et al.
Published: (2024)
by: Liang, Susan, et al.
Published: (2024)
Latent Representations for Visual Proprioception in Inexpensive Robots
by: Sheikholeslami, Sahara, et al.
Published: (2025)
by: Sheikholeslami, Sahara, et al.
Published: (2025)
Exploiting Text-Image Latent Spaces for the Description of Visual Concepts
by: Schmalwasser, Laines, et al.
Published: (2024)
by: Schmalwasser, Laines, et al.
Published: (2024)
Text-Audio-Visual-conditioned Diffusion Model for Video Saliency Prediction
by: Yu, Li, et al.
Published: (2025)
by: Yu, Li, et al.
Published: (2025)
Foley Control: Aligning a Frozen Latent Text-to-Audio Model to Video
by: Rowles, Ciara, et al.
Published: (2025)
by: Rowles, Ciara, et al.
Published: (2025)
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval
by: Jeong, Boseung, et al.
Published: (2025)
by: Jeong, Boseung, et al.
Published: (2025)
JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion
by: Chen, Anthony, et al.
Published: (2026)
by: Chen, Anthony, et al.
Published: (2026)
CoVA: Text-Guided Composed Video Retrieval for Audio-Visual Content
by: Han, Gyuwon, et al.
Published: (2026)
by: Han, Gyuwon, et al.
Published: (2026)
Answering Diverse Questions via Text Attached with Key Audio-Visual Clues
by: Ye, Qilang, et al.
Published: (2024)
by: Ye, Qilang, et al.
Published: (2024)
Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation
by: Xu, Mingwang, et al.
Published: (2024)
by: Xu, Mingwang, et al.
Published: (2024)
SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery Detection
by: Liang, Yachao, et al.
Published: (2025)
by: Liang, Yachao, et al.
Published: (2025)
Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model
by: Seedance, Team, et al.
Published: (2025)
by: Seedance, Team, et al.
Published: (2025)
Enhancing Visual Representation for Text-based Person Searching
by: Shen, Wei, et al.
Published: (2024)
by: Shen, Wei, et al.
Published: (2024)
Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement
by: Nguyen, Tuan-Nam, et al.
Published: (2025)
by: Nguyen, Tuan-Nam, et al.
Published: (2025)
Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics
by: Shen, Ying, et al.
Published: (2026)
by: Shen, Ying, et al.
Published: (2026)
TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning
by: Kim, Seongah, et al.
Published: (2026)
by: Kim, Seongah, et al.
Published: (2026)
Latent Space Guided Scenario Sampling for Multimodal Segmentation Under Missing Modalities
by: Ulku, Irem, et al.
Published: (2026)
by: Ulku, Irem, et al.
Published: (2026)
VAEmo: Efficient Representation Learning for Visual-Audio Emotion with Knowledge Injection
by: Cheng, Hao, et al.
Published: (2025)
by: Cheng, Hao, et al.
Published: (2025)
OmniForcing: Unleashing Real-time Joint Audio-Visual Generation
by: Su, Yaofeng, et al.
Published: (2026)
by: Su, Yaofeng, et al.
Published: (2026)
Aligning Audio-Visual Joint Representations with an Agentic Workflow
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Audio Visual Segmentation Through Text Embeddings
by: Lee, Kyungbok, et al.
Published: (2025)
by: Lee, Kyungbok, et al.
Published: (2025)
SpikeGen: Decoupled "Rods and Cones" Visual Representation Processing with Latent Generative Framework
by: Dai, Gaole, et al.
Published: (2025)
by: Dai, Gaole, et al.
Published: (2025)
CLUE: Controllable Latent space of Unprompted Embeddings for Diversity Management in Text-to-Image Synthesis
by: Park, Keunwoo, et al.
Published: (2025)
by: Park, Keunwoo, et al.
Published: (2025)
Similar Items
-
Audio-Visual Speech Representation Expert for Enhanced Talking Face Video Generation and Evaluation
by: Yaman, Dogucan, et al.
Published: (2024) -
Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation
by: Yaman, Dogucan, et al.
Published: (2025) -
CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding
by: Eyiokur, Fevziye Irem, et al.
Published: (2025) -
Audio-driven Talking Face Generation with Stabilized Synchronization Loss
by: Yaman, Dogucan, et al.
Published: (2023) -
A Multimodal Depth-Aware Method For Embodied Reference Understanding
by: Eyiokur, Fevziye Irem, et al.
Published: (2025)