AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Choi, Jeongsoo, Kim, Ji-Hoon, Sung-Bin, Kim, Oh, Tae-Hyun, Chung, Joon Son |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment
von: Senocak, Arda, et al.
Veröffentlicht: (2024)
von: Senocak, Arda, et al.
Veröffentlicht: (2024)
DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization
von: Nguyen, Ngoc-Son, et al.
Veröffentlicht: (2026)
von: Nguyen, Ngoc-Son, et al.
Veröffentlicht: (2026)
Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment
von: Sung-Bin, Kim, et al.
Veröffentlicht: (2024)
von: Sung-Bin, Kim, et al.
Veröffentlicht: (2024)
CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2026)
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2026)
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2024)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2024)
AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech Representation
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2023)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2023)
PAVAS: Physics-Aware Video-to-Audio Synthesis
von: Hyun-Bin, Oh, et al.
Veröffentlicht: (2025)
von: Hyun-Bin, Oh, et al.
Veröffentlicht: (2025)
VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models
von: Sung-Bin, Kim, et al.
Veröffentlicht: (2025)
von: Sung-Bin, Kim, et al.
Veröffentlicht: (2025)
Multimodal Large Language Model is a Human-Aligned Annotator for Text-to-Image Generation
von: Wu, Xun, et al.
Veröffentlicht: (2024)
von: Wu, Xun, et al.
Veröffentlicht: (2024)
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
von: Wu, Yi, et al.
Veröffentlicht: (2025)
von: Wu, Yi, et al.
Veröffentlicht: (2025)
From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2025)
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2025)
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
von: Wang, Le, et al.
Veröffentlicht: (2025)
von: Wang, Le, et al.
Veröffentlicht: (2025)
Anisotropic Modality Align
von: Yu, Xiaomin, et al.
Veröffentlicht: (2026)
von: Yu, Xiaomin, et al.
Veröffentlicht: (2026)
GuardAlign: Test-time Safety Alignment in Multimodal Large Language Models
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
AKVSR: Audio Knowledge Empowered Visual Speech Recognition by Compressing Audio Knowledge of a Pretrained Model
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2023)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2023)
SD-DiT: Unleashing the Power of Self-supervised Discrimination in Diffusion Transformer
von: Zhu, Rui, et al.
Veröffentlicht: (2024)
von: Zhu, Rui, et al.
Veröffentlicht: (2024)
UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer
von: Wang, Xiang, et al.
Veröffentlicht: (2025)
von: Wang, Xiang, et al.
Veröffentlicht: (2025)
MAVFlow: Preserving Paralinguistic Elements with Conditional Flow Matching for Zero-Shot AV2AV Multilingual Translation
von: Cho, Sungwoo, et al.
Veröffentlicht: (2025)
von: Cho, Sungwoo, et al.
Veröffentlicht: (2025)
Improving Noise Robust Audio-Visual Speech Recognition via Router-Gated Cross-Modal Feature Fusion
von: Lim, DongHoon, et al.
Veröffentlicht: (2025)
von: Lim, DongHoon, et al.
Veröffentlicht: (2025)
MoLT: Mixture of Layer-Wise Tokens for Efficient Audio-Visual Learning
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
ComAlign: Compositional Alignment in Vision-Language Models
von: Abdollah, Ali, et al.
Veröffentlicht: (2024)
von: Abdollah, Ali, et al.
Veröffentlicht: (2024)
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
von: Ji, Yatai, et al.
Veröffentlicht: (2024)
von: Ji, Yatai, et al.
Veröffentlicht: (2024)
Disparity-based Stereo Image Compression with Aligned Cross-View Priors
von: Zhai, Yongqi, et al.
Veröffentlicht: (2022)
von: Zhai, Yongqi, et al.
Veröffentlicht: (2022)
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models
von: Yang, Jialiang, et al.
Veröffentlicht: (2026)
von: Yang, Jialiang, et al.
Veröffentlicht: (2026)
VerbDiff: Text-Only Diffusion Models with Enhanced Interaction Awareness
von: Cha, SeungJu, et al.
Veröffentlicht: (2025)
von: Cha, SeungJu, et al.
Veröffentlicht: (2025)
KAN-Based Fusion of Dual-Domain for Audio-Driven Facial Landmarks Generation
von: Vo-Thanh, Hoang-Son, et al.
Veröffentlicht: (2024)
von: Vo-Thanh, Hoang-Son, et al.
Veröffentlicht: (2024)
CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language Model
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
SAFIRE: Segment Any Forged Image Region
von: Kwon, Myung-Joon, et al.
Veröffentlicht: (2024)
von: Kwon, Myung-Joon, et al.
Veröffentlicht: (2024)
AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
HOP: Heterogeneous Topology-based Multimodal Entanglement for Co-Speech Gesture Generation
von: Cheng, Hongye, et al.
Veröffentlicht: (2025)
von: Cheng, Hongye, et al.
Veröffentlicht: (2025)
FoleyGRAM: Video-to-Audio Generation with GRAM-Aligned Multimodal Encoders
von: Gramaccioni, Riccardo Fosco, et al.
Veröffentlicht: (2025)
von: Gramaccioni, Riccardo Fosco, et al.
Veröffentlicht: (2025)
CLIP-PCQA: Exploring Subjective-Aligned Vision-Language Modeling for Point Cloud Quality Assessment
von: Liu, Yating, et al.
Veröffentlicht: (2025)
von: Liu, Yating, et al.
Veröffentlicht: (2025)
DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation
von: Cai, Minghong, et al.
Veröffentlicht: (2024)
von: Cai, Minghong, et al.
Veröffentlicht: (2024)
MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval
von: Ju, Yeong-Joon, et al.
Veröffentlicht: (2024)
von: Ju, Yeong-Joon, et al.
Veröffentlicht: (2024)
Multimodal Transformer With a Low-Computational-Cost Guarantee
von: Park, Sungjin, et al.
Veröffentlicht: (2024)
von: Park, Sungjin, et al.
Veröffentlicht: (2024)
A Novel FACS-Aligned Anatomical Text Description Paradigm for Fine-Grained Facial Behavior Synthesis
von: Wang, Jiahe, et al.
Veröffentlicht: (2026)
von: Wang, Jiahe, et al.
Veröffentlicht: (2026)
MHAD: Multimodal Home Activity Dataset with Multi-Angle Videos and Synchronized Physiological Signals
von: Yu, Lei, et al.
Veröffentlicht: (2024)
von: Yu, Lei, et al.
Veröffentlicht: (2024)
DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion Autoencoder
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
A Simple Baseline with Single-encoder for Referring Image Segmentation
von: Yu, Seonghoon, et al.
Veröffentlicht: (2024)
von: Yu, Seonghoon, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment
von: Senocak, Arda, et al.
Veröffentlicht: (2024) -
DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization
von: Nguyen, Ngoc-Son, et al.
Veröffentlicht: (2026) -
Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment
von: Sung-Bin, Kim, et al.
Veröffentlicht: (2024) -
CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2026) -
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2024)