Error Analyses of Auto-Regressive Video Diffusion Models: A Unified Framework
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Jing, Zhang, Fengzhuo, Li, Xiaoli, Tan, Vincent Y. F., Pang, Tianyu, Du, Chao, Sun, Aixin, Yang, Zhuoran |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving Long-Text Alignment for Text-to-Image Diffusion Models
von: Liu, Luping, et al.
Veröffentlicht: (2024)
von: Liu, Luping, et al.
Veröffentlicht: (2024)
VIA: Unified Spatiotemporal Video Adaptation Framework for Global and Local Video Editing
von: Gu, Jing, et al.
Veröffentlicht: (2024)
von: Gu, Jing, et al.
Veröffentlicht: (2024)
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
von: Li, Hebeizi, et al.
Veröffentlicht: (2026)
von: Li, Hebeizi, et al.
Veröffentlicht: (2026)
DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion Autoencoder
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation
von: Zhou, S. Z., et al.
Veröffentlicht: (2025)
von: Zhou, S. Z., et al.
Veröffentlicht: (2025)
TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long Videos
von: Zhang, Chen-Lin, et al.
Veröffentlicht: (2025)
von: Zhang, Chen-Lin, et al.
Veröffentlicht: (2025)
M2ORT: Many-To-One Regression Transformer for Spatial Transcriptomics Prediction from Histopathology Images
von: Wang, Hongyi, et al.
Veröffentlicht: (2024)
von: Wang, Hongyi, et al.
Veröffentlicht: (2024)
Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective
von: Yuan, Hangjie, et al.
Veröffentlicht: (2025)
von: Yuan, Hangjie, et al.
Veröffentlicht: (2025)
Test-Time Backdoor Attacks on Multimodal Large Language Models
von: Lu, Dong, et al.
Veröffentlicht: (2024)
von: Lu, Dong, et al.
Veröffentlicht: (2024)
UBiSS: A Unified Framework for Bimodal Semantic Summarization of Videos
von: Mei, Yuting, et al.
Veröffentlicht: (2024)
von: Mei, Yuting, et al.
Veröffentlicht: (2024)
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
von: Lin, Yuanze, et al.
Veröffentlicht: (2025)
von: Lin, Yuanze, et al.
Veröffentlicht: (2025)
AutoAWG: Adverse Weather Generation with Adaptive Multi-Controls for Automotive Videos
von: Hu, Jiagao, et al.
Veröffentlicht: (2026)
von: Hu, Jiagao, et al.
Veröffentlicht: (2026)
Generalized Video Anomaly Event Detection: Systematic Taxonomy and Comparison of Deep Models
von: Liu, Yang, et al.
Veröffentlicht: (2023)
von: Liu, Yang, et al.
Veröffentlicht: (2023)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
One Framework to Rule Them All: Unifying Multimodal Tasks with LLM Neural-Tuning
von: Sun, Hao, et al.
Veröffentlicht: (2024)
von: Sun, Hao, et al.
Veröffentlicht: (2024)
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
von: Xu, Yifang, et al.
Veröffentlicht: (2025)
von: Xu, Yifang, et al.
Veröffentlicht: (2025)
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models
von: Yang, Haibo, et al.
Veröffentlicht: (2024)
von: Yang, Haibo, et al.
Veröffentlicht: (2024)
Bridging Your Imagination with Audio-Video Generation via a Unified Director
von: Zhang, Jiaxu, et al.
Veröffentlicht: (2025)
von: Zhang, Jiaxu, et al.
Veröffentlicht: (2025)
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation
von: Luo, Ziyang, et al.
Veröffentlicht: (2024)
von: Luo, Ziyang, et al.
Veröffentlicht: (2024)
TRIP: Temporal Residual Learning with Image Noise Prior for Image-to-Video Diffusion Models
von: Zhang, Zhongwei, et al.
Veröffentlicht: (2024)
von: Zhang, Zhongwei, et al.
Veröffentlicht: (2024)
Multi-proposal Collaboration and Multi-task Training for Weakly-supervised Video Moment Retrieval
von: Zhang, Bolin, et al.
Veröffentlicht: (2026)
von: Zhang, Bolin, et al.
Veröffentlicht: (2026)
Emotion-Qwen: A Unified Framework for Emotion and Vision Understanding
von: Huang, Dawei, et al.
Veröffentlicht: (2025)
von: Huang, Dawei, et al.
Veröffentlicht: (2025)
Memories are One-to-Many Mapping Alleviators in Talking Face Generation
von: Tang, Anni, et al.
Veröffentlicht: (2022)
von: Tang, Anni, et al.
Veröffentlicht: (2022)
Benchmarking Large Multimodal Models against Common Corruptions
von: Zhang, Jiawei, et al.
Veröffentlicht: (2024)
von: Zhang, Jiawei, et al.
Veröffentlicht: (2024)
Composing Concepts from Images and Videos via Concept-prompt Binding
von: Kong, Xianghao, et al.
Veröffentlicht: (2025)
von: Kong, Xianghao, et al.
Veröffentlicht: (2025)
StreamingEval: A Unified Evaluation Protocol towards Realistic Streaming Video Understanding
von: Tang, Guowei, et al.
Veröffentlicht: (2026)
von: Tang, Guowei, et al.
Veröffentlicht: (2026)
MCIHN: A Hybrid Network Model Based on Multi-path Cross-modal Interaction for Multimodal Emotion Recognition
von: Zhang, Haoyang, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyang, et al.
Veröffentlicht: (2025)
AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
von: Guan, Jiazhi, et al.
Veröffentlicht: (2025)
von: Guan, Jiazhi, et al.
Veröffentlicht: (2025)
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
von: Sun, Zeyi, et al.
Veröffentlicht: (2024)
von: Sun, Zeyi, et al.
Veröffentlicht: (2024)
VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and Understanding
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling
von: Yang, Jianxuan, et al.
Veröffentlicht: (2026)
von: Yang, Jianxuan, et al.
Veröffentlicht: (2026)
OneHOI: Unifying Human-Object Interaction Generation and Editing
von: Hoe, Jiun Tian, et al.
Veröffentlicht: (2026)
von: Hoe, Jiun Tian, et al.
Veröffentlicht: (2026)
JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
von: Chao, Jianghan, et al.
Veröffentlicht: (2025)
von: Chao, Jianghan, et al.
Veröffentlicht: (2025)
Bernini: Latent Semantic Planning for Video Diffusion
von: Bernini Team, et al.
Veröffentlicht: (2026)
von: Bernini Team, et al.
Veröffentlicht: (2026)
JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation
von: Liu, Kai, et al.
Veröffentlicht: (2026)
von: Liu, Kai, et al.
Veröffentlicht: (2026)
Omni2Sound: Towards Unified Video-Text-to-Audio Generation
von: Dai, Yusheng, et al.
Veröffentlicht: (2026)
von: Dai, Yusheng, et al.
Veröffentlicht: (2026)
StableMoFusion: Towards Robust and Efficient Diffusion-based Motion Generation Framework
von: Huang, Yiheng, et al.
Veröffentlicht: (2024)
von: Huang, Yiheng, et al.
Veröffentlicht: (2024)
AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video Hashing
von: Lian, Niu, et al.
Veröffentlicht: (2025)
von: Lian, Niu, et al.
Veröffentlicht: (2025)
Probabilistic Temporal Masked Attention for Cross-view Online Action Detection
von: Xie, Liping, et al.
Veröffentlicht: (2025)
von: Xie, Liping, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Improving Long-Text Alignment for Text-to-Image Diffusion Models
von: Liu, Luping, et al.
Veröffentlicht: (2024) -
VIA: Unified Spatiotemporal Video Adaptation Framework for Global and Local Video Editing
von: Gu, Jing, et al.
Veröffentlicht: (2024) -
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
von: Li, Hebeizi, et al.
Veröffentlicht: (2026) -
DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion Autoencoder
von: Du, Chenpeng, et al.
Veröffentlicht: (2023) -
A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation
von: Zhou, S. Z., et al.
Veröffentlicht: (2025)