Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liao, Junchao, Zhang, Zhenghao, Meng, Xiangyu, Li, Litao, Zhang, Ziying, Zhu, Siyu, Qin, Long, Wang, Weizhi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tora2: Motion and Appearance Customized Diffusion Transformer for Multi-Entity Video Generation
von: Zhang, Zhenghao, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenghao, et al.
Veröffentlicht: (2025)
Tora: Trajectory-oriented Diffusion Transformer for Video Generation
von: Zhang, Zhenghao, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenghao, et al.
Veröffentlicht: (2024)
TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation
von: Liu, Xiangyu, et al.
Veröffentlicht: (2026)
von: Liu, Xiangyu, et al.
Veröffentlicht: (2026)
Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation
von: Li, Chunyu, et al.
Veröffentlicht: (2026)
von: Li, Chunyu, et al.
Veröffentlicht: (2026)
ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling
von: Yang, Jianxuan, et al.
Veröffentlicht: (2026)
von: Yang, Jianxuan, et al.
Veröffentlicht: (2026)
PAVAS: Physics-Aware Video-to-Audio Synthesis
von: Hyun-Bin, Oh, et al.
Veröffentlicht: (2025)
von: Hyun-Bin, Oh, et al.
Veröffentlicht: (2025)
Omni2Sound: Towards Unified Video-Text-to-Audio Generation
von: Dai, Yusheng, et al.
Veröffentlicht: (2026)
von: Dai, Yusheng, et al.
Veröffentlicht: (2026)
Coherent Audio-Visual Editing via Conditional Audio Generation Following Video Edits
von: Ishii, Masato, et al.
Veröffentlicht: (2025)
von: Ishii, Masato, et al.
Veröffentlicht: (2025)
JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation
von: Liu, Kai, et al.
Veröffentlicht: (2026)
von: Liu, Kai, et al.
Veröffentlicht: (2026)
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
von: Chen, Yuheng, et al.
Veröffentlicht: (2026)
von: Chen, Yuheng, et al.
Veröffentlicht: (2026)
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
von: Cheng, Shihao, et al.
Veröffentlicht: (2026)
von: Cheng, Shihao, et al.
Veröffentlicht: (2026)
MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
von: Yang, Jianxuan, et al.
Veröffentlicht: (2025)
von: Yang, Jianxuan, et al.
Veröffentlicht: (2025)
Do Joint Audio-Video Generation Models Understand Physics?
von: Cui, Zijun, et al.
Veröffentlicht: (2026)
von: Cui, Zijun, et al.
Veröffentlicht: (2026)
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text
von: Pian, Weiguo, et al.
Veröffentlicht: (2026)
von: Pian, Weiguo, et al.
Veröffentlicht: (2026)
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
von: Li, Hebeizi, et al.
Veröffentlicht: (2026)
von: Li, Hebeizi, et al.
Veröffentlicht: (2026)
Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
von: Zhou, Yang-Hao, et al.
Veröffentlicht: (2026)
von: Zhou, Yang-Hao, et al.
Veröffentlicht: (2026)
DreamFoley: Scalable VLMs for High-Fidelity Video-to-Audio Generation
von: Li, Fu, et al.
Veröffentlicht: (2025)
von: Li, Fu, et al.
Veröffentlicht: (2025)
3MDiT: Unified Tri-Modal Diffusion Transformer for Text-Driven Synchronized Audio-Video Generation
von: Li, Yaoru, et al.
Veröffentlicht: (2025)
von: Li, Yaoru, et al.
Veröffentlicht: (2025)
Diffusion Models for Joint Audio-Video Generation
von: La Torre, Alejandro Paredes
Veröffentlicht: (2026)
von: La Torre, Alejandro Paredes
Veröffentlicht: (2026)
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
von: Wang, Le, et al.
Veröffentlicht: (2025)
von: Wang, Le, et al.
Veröffentlicht: (2025)
AudioStory: Generating Long-Form Narrative Audio with Large Language Models
von: Guo, Yuxin, et al.
Veröffentlicht: (2025)
von: Guo, Yuxin, et al.
Veröffentlicht: (2025)
STCTS: Generative Semantic Compression for Ultra-Low Bitrate Speech via Explicit Text-Prosody-Timbre Decomposition
von: Wang, Siyu, et al.
Veröffentlicht: (2025)
von: Wang, Siyu, et al.
Veröffentlicht: (2025)
Multifaceted Evaluation of Audio-Visual Capability for MLLMs: Effectiveness, Efficiency, Generalizability and Robustness
von: Zhao, Yusheng, et al.
Veröffentlicht: (2025)
von: Zhao, Yusheng, et al.
Veröffentlicht: (2025)
AccKV: Towards Efficient Audio-Video LLMs Inference via Adaptive-Focusing and Cross-Calibration KV Cache Optimization
von: Jiang, Zhonghua, et al.
Veröffentlicht: (2025)
von: Jiang, Zhonghua, et al.
Veröffentlicht: (2025)
Video-to-Audio Generation with Hidden Alignment
von: Xu, Manjie, et al.
Veröffentlicht: (2024)
von: Xu, Manjie, et al.
Veröffentlicht: (2024)
Audio-Guided Visual Perception for Audio-Visual Navigation
von: Wang, Yi, et al.
Veröffentlicht: (2025)
von: Wang, Yi, et al.
Veröffentlicht: (2025)
Apollo: Unified Multi-Task Audio-Video Joint Generation
von: Wang, Jun, et al.
Veröffentlicht: (2026)
von: Wang, Jun, et al.
Veröffentlicht: (2026)
AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control
von: Guo, Xinyue, et al.
Veröffentlicht: (2025)
von: Guo, Xinyue, et al.
Veröffentlicht: (2025)
Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation
von: Tong, Xinyi, et al.
Veröffentlicht: (2025)
von: Tong, Xinyi, et al.
Veröffentlicht: (2025)
Generative Audio Extension and Morphing
von: Seetharaman, Prem, et al.
Veröffentlicht: (2026)
von: Seetharaman, Prem, et al.
Veröffentlicht: (2026)
LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV
von: Liu, Tengfei, et al.
Veröffentlicht: (2026)
von: Liu, Tengfei, et al.
Veröffentlicht: (2026)
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video Generation
von: Tan, Weiting, et al.
Veröffentlicht: (2026)
von: Tan, Weiting, et al.
Veröffentlicht: (2026)
Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation
von: Liang, Yingshan, et al.
Veröffentlicht: (2025)
von: Liang, Yingshan, et al.
Veröffentlicht: (2025)
IsoSignVid2Aud: Sign Language Video to Audio Conversion without Text Intermediaries
von: Kavediya, Harsh, et al.
Veröffentlicht: (2025)
von: Kavediya, Harsh, et al.
Veröffentlicht: (2025)
Sound Sparks Motion: Audio and Text Tuning for Video Editing
von: Razlighi, AmirHossein Naghi, et al.
Veröffentlicht: (2026)
von: Razlighi, AmirHossein Naghi, et al.
Veröffentlicht: (2026)
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models
von: Yang, Jialiang, et al.
Veröffentlicht: (2026)
von: Yang, Jialiang, et al.
Veröffentlicht: (2026)
OmniForcing: Unleashing Real-time Joint Audio-Visual Generation
von: Su, Yaofeng, et al.
Veröffentlicht: (2026)
von: Su, Yaofeng, et al.
Veröffentlicht: (2026)
Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
von: Tian, Zeyue, et al.
Veröffentlicht: (2026)
von: Tian, Zeyue, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Tora2: Motion and Appearance Customized Diffusion Transformer for Multi-Entity Video Generation
von: Zhang, Zhenghao, et al.
Veröffentlicht: (2025) -
Tora: Trajectory-oriented Diffusion Transformer for Video Generation
von: Zhang, Zhenghao, et al.
Veröffentlicht: (2024) -
TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation
von: Liu, Xiangyu, et al.
Veröffentlicht: (2026) -
Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation
von: Li, Chunyu, et al.
Veröffentlicht: (2026) -
ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling
von: Yang, Jianxuan, et al.
Veröffentlicht: (2026)