Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Yupeng, Huang, Lianghua, Wu, Zhifan, Wang, Jiabao, Shi, Yupeng, Jiang, Biao, Zhou, Daquan, Liu, Yu, Cheng, Ming-Ming, Hou, Qibin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
von: Zhou, Yupeng, et al.
Veröffentlicht: (2024)
von: Zhou, Yupeng, et al.
Veröffentlicht: (2024)
MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation
von: Li, Haitian, et al.
Veröffentlicht: (2026)
von: Li, Haitian, et al.
Veröffentlicht: (2026)
FusID: Modality-Fused Semantic IDs for Generative Music Recommendation
von: Kim, Haven, et al.
Veröffentlicht: (2026)
von: Kim, Haven, et al.
Veröffentlicht: (2026)
MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
von: Zhou, Yang-Hao, et al.
Veröffentlicht: (2026)
von: Zhou, Yang-Hao, et al.
Veröffentlicht: (2026)
Efficient Video to Audio Mapper with Visual Scene Detection
von: Yi, Mingjing, et al.
Veröffentlicht: (2024)
von: Yi, Mingjing, et al.
Veröffentlicht: (2024)
Online Audio-Visual Autoregressive Speaker Extraction
von: Pan, Zexu, et al.
Veröffentlicht: (2025)
von: Pan, Zexu, et al.
Veröffentlicht: (2025)
AudioEval: Automatic Dual-Perspective and Multi-Dimensional Evaluation of Text-to-Audio-Generation
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
Temporally Aligned Audio for Video with Autoregression
von: Viertola, Ilpo, et al.
Veröffentlicht: (2024)
von: Viertola, Ilpo, et al.
Veröffentlicht: (2024)
AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing
von: Chen, William, et al.
Veröffentlicht: (2026)
von: Chen, William, et al.
Veröffentlicht: (2026)
Robust Distortion-Free Watermark for Autoregressive Audio Generation Models
von: Wu, Yihan, et al.
Veröffentlicht: (2025)
von: Wu, Yihan, et al.
Veröffentlicht: (2025)
Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based Audio-Visual Conformer
von: Wang, Haoxu, et al.
Veröffentlicht: (2024)
von: Wang, Haoxu, et al.
Veröffentlicht: (2024)
AR-1-to-3: Single Image to Consistent 3D Object Generation via Next-View Prediction
von: Zhang, Xuying, et al.
Veröffentlicht: (2025)
von: Zhang, Xuying, et al.
Veröffentlicht: (2025)
Sora Generates Videos with Stunning Geometrical Consistency
von: Li, Xuanyi, et al.
Veröffentlicht: (2024)
von: Li, Xuanyi, et al.
Veröffentlicht: (2024)
SRFormerV2: Taking a Closer Look at Permuted Self-Attention for Image Super-Resolution
von: Zhou, Yupeng, et al.
Veröffentlicht: (2023)
von: Zhou, Yupeng, et al.
Veröffentlicht: (2023)
FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
DreamFoley: Scalable VLMs for High-Fidelity Video-to-Audio Generation
von: Li, Fu, et al.
Veröffentlicht: (2025)
von: Li, Fu, et al.
Veröffentlicht: (2025)
LLM-ForcedAligner: A Non-Autoregressive and Accurate LLM-Based Forced Aligner for Multilingual and Long-Form Speech
von: Mu, Bingshen, et al.
Veröffentlicht: (2026)
von: Mu, Bingshen, et al.
Veröffentlicht: (2026)
OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
von: Zhou, Yupeng, et al.
Veröffentlicht: (2025)
von: Zhou, Yupeng, et al.
Veröffentlicht: (2025)
Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
von: Lei, Ke, et al.
Veröffentlicht: (2026)
von: Lei, Ke, et al.
Veröffentlicht: (2026)
Online Register for Dual-Mode Self-Supervised Speech Models: Mitigating The Lack of Future Context
von: Goto, Keita, et al.
Veröffentlicht: (2026)
von: Goto, Keita, et al.
Veröffentlicht: (2026)
Cross-attention and Self-attention for Audio-visual Speaker Diarization in MISP-Meeting Challenge
von: Li, Zhaoyang, et al.
Veröffentlicht: (2025)
von: Li, Zhaoyang, et al.
Veröffentlicht: (2025)
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models
von: Low, Chetwin, et al.
Veröffentlicht: (2025)
von: Low, Chetwin, et al.
Veröffentlicht: (2025)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
UniFlow-Audio: Unified Flow Matching for Audio Generation from Omni-Modalities
von: Xu, Xuenan, et al.
Veröffentlicht: (2025)
von: Xu, Xuenan, et al.
Veröffentlicht: (2025)
Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training
von: Sun, Houmin, et al.
Veröffentlicht: (2026)
von: Sun, Houmin, et al.
Veröffentlicht: (2026)
Waveform-Logmel Audio Neural Networks for Respiratory Sound Classification
von: Xie, Jiadong, et al.
Veröffentlicht: (2025)
von: Xie, Jiadong, et al.
Veröffentlicht: (2025)
Whisper-AuT: Domain-Adapted Audio Encoder for Efficient Audio-LLM Training
von: Qiu, Jielin, et al.
Veröffentlicht: (2026)
von: Qiu, Jielin, et al.
Veröffentlicht: (2026)
MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows
von: Li, Xiquan, et al.
Veröffentlicht: (2025)
von: Li, Xiquan, et al.
Veröffentlicht: (2025)
ChronosAudio: A Comprehensive Long-Audio Benchmark for Evaluating Audio-Large Language Models
von: Luo, Kaiwen, et al.
Veröffentlicht: (2026)
von: Luo, Kaiwen, et al.
Veröffentlicht: (2026)
AudioRole: An Audio Dataset for Character Role-Playing in Large Language Models
von: Li, Wenyu, et al.
Veröffentlicht: (2025)
von: Li, Wenyu, et al.
Veröffentlicht: (2025)
APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track
von: Miao, Deshui, et al.
Veröffentlicht: (2026)
von: Miao, Deshui, et al.
Veröffentlicht: (2026)
How to Label Resynthesized Audio: The Dual Role of Neural Audio Codecs in Audio Deepfake Detection
von: Xiao, Yixuan, et al.
Veröffentlicht: (2026)
von: Xiao, Yixuan, et al.
Veröffentlicht: (2026)
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
DualMark: Identifying Model and Training Data Origins in Generated Audio
von: Yang, Xuefeng, et al.
Veröffentlicht: (2025)
von: Yang, Xuefeng, et al.
Veröffentlicht: (2025)
Enhancing Speaker Verification with w2v-BERT 2.0 and Knowledge Distillation guided Structured Pruning
von: Li, Ze, et al.
Veröffentlicht: (2025)
von: Li, Ze, et al.
Veröffentlicht: (2025)
Dual Audio-Centric Modality Coupling for Talking Head Generation
von: Fu, Ao, et al.
Veröffentlicht: (2025)
von: Fu, Ao, et al.
Veröffentlicht: (2025)
MOVA: Towards Scalable and Synchronized Video-Audio Generation
von: OpenMOSS Team, et al.
Veröffentlicht: (2026)
von: OpenMOSS Team, et al.
Veröffentlicht: (2026)
BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech Synthesis
von: Xing, Jingyuan, et al.
Veröffentlicht: (2025)
von: Xing, Jingyuan, et al.
Veröffentlicht: (2025)
FAST: Fast Audio Spectrogram Transformer
von: Naman, Anugunj, et al.
Veröffentlicht: (2025)
von: Naman, Anugunj, et al.
Veröffentlicht: (2025)
AudioGS: Spectrogram-Based Audio Gaussian Splatting for Sound Field Reconstruction
von: Bi, Chunhao, et al.
Veröffentlicht: (2026)
von: Bi, Chunhao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
von: Zhou, Yupeng, et al.
Veröffentlicht: (2024) -
MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation
von: Li, Haitian, et al.
Veröffentlicht: (2026) -
FusID: Modality-Fused Semantic IDs for Generative Music Recommendation
von: Kim, Haven, et al.
Veröffentlicht: (2026) -
MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
von: Zhou, Yang-Hao, et al.
Veröffentlicht: (2026) -
Efficient Video to Audio Mapper with Visual Scene Detection
von: Yi, Mingjing, et al.
Veröffentlicht: (2024)