KeyVID: Keyframe-Aware Video Diffusion for Audio-Synchronized Visual Animation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xingrui, Liu, Jiang, Wang, Ze, Yu, Xiaodong, Wu, Jialian, Sun, Ximeng, Su, Yusheng, Yuille, Alan, Liu, Zicheng, Barsoum, Emad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
MOVi: Training-free Text-conditioned Multi-Object Video Generation
von: Rahman, Aimon, et al.
Veröffentlicht: (2025)
von: Rahman, Aimon, et al.
Veröffentlicht: (2025)
Unleashing Hour-Scale Video Training for Long Video-Language Understanding
von: Lin, Jingyang, et al.
Veröffentlicht: (2025)
von: Lin, Jingyang, et al.
Veröffentlicht: (2025)
VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking
von: Lin, Jingyang, et al.
Veröffentlicht: (2026)
von: Lin, Jingyang, et al.
Veröffentlicht: (2026)
Agent Laboratory: Using LLM Agents as Research Assistants
von: Schmidgall, Samuel, et al.
Veröffentlicht: (2025)
von: Schmidgall, Samuel, et al.
Veröffentlicht: (2025)
Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation
von: Wang, Ze, et al.
Veröffentlicht: (2025)
von: Wang, Ze, et al.
Veröffentlicht: (2025)
Self-Taught Agentic Long Context Understanding
von: Zhuang, Yufan, et al.
Veröffentlicht: (2025)
von: Zhuang, Yufan, et al.
Veröffentlicht: (2025)
Latent Visual Reasoning
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
Learning from Online Videos at Inference Time for Computer-Use Agents
von: Liu, Yujian, et al.
Veröffentlicht: (2025)
von: Liu, Yujian, et al.
Veröffentlicht: (2025)
CD4LM: Consistency Distillation and aDaptive Decoding for Diffusion Language Models
von: Liang, Yihao, et al.
Veröffentlicht: (2026)
von: Liang, Yihao, et al.
Veröffentlicht: (2026)
ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning
von: Guo, Yuxiang, et al.
Veröffentlicht: (2025)
von: Guo, Yuxiang, et al.
Veröffentlicht: (2025)
DRIFT: Transferring Reasoning Priors for Efficient MLLM Fine-Tuning
von: Huang, Chao, et al.
Veröffentlicht: (2025)
von: Huang, Chao, et al.
Veröffentlicht: (2025)
TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games
von: Mishra, Prakamya, et al.
Veröffentlicht: (2025)
von: Mishra, Prakamya, et al.
Veröffentlicht: (2025)
Instella: Fully Open Language Models with Stellar Performance
von: Liu, Jiang, et al.
Veröffentlicht: (2025)
von: Liu, Jiang, et al.
Veröffentlicht: (2025)
Reliable Use of Lemmas via Eligibility Reasoning and Section$-$Aware Reinforcement Learning
von: Xu, Zhikun, et al.
Veröffentlicht: (2026)
von: Xu, Zhikun, et al.
Veröffentlicht: (2026)
Stabilizing Efficient Reasoning with Step-Level Advantage Selection
von: Wang, Han, et al.
Veröffentlicht: (2026)
von: Wang, Han, et al.
Veröffentlicht: (2026)
APRIL: Active Partial Rollouts in Reinforcement Learning to Tame Long-tail Generation
von: Zhou, Yuzhen, et al.
Veröffentlicht: (2025)
von: Zhou, Yuzhen, et al.
Veröffentlicht: (2025)
Audio-Synchronized Visual Animation
von: Zhang, Lin, et al.
Veröffentlicht: (2024)
von: Zhang, Lin, et al.
Veröffentlicht: (2024)
SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
KSDiff: Keyframe-Augmented Speech-Aware Dual-Path Diffusion for Facial Animation
von: Lyu, Tianle, et al.
Veröffentlicht: (2025)
von: Lyu, Tianle, et al.
Veröffentlicht: (2025)
Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion
von: Singh, Shivam, et al.
Veröffentlicht: (2026)
von: Singh, Shivam, et al.
Veröffentlicht: (2026)
CaptionQA: Is Your Caption as Useful as the Image Itself?
von: Yang, Shijia, et al.
Veröffentlicht: (2025)
von: Yang, Shijia, et al.
Veröffentlicht: (2025)
TermiGen: High-Fidelity Environment and Robust Trajectory Synthesis for Terminal Agents
von: Zhu, Kaijie, et al.
Veröffentlicht: (2026)
von: Zhu, Kaijie, et al.
Veröffentlicht: (2026)
KeyframeFace: Language-Driven Facial Animation via Semantic Keyframes
von: Wu, Jingchao, et al.
Veröffentlicht: (2025)
von: Wu, Jingchao, et al.
Veröffentlicht: (2025)
AdaptEvolve: Improving Efficiency of Evolutionary AI Agents through Adaptive Model Selection
von: Ray, Pretam, et al.
Veröffentlicht: (2026)
von: Ray, Pretam, et al.
Veröffentlicht: (2026)
TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering
von: Joshi, Vinay, et al.
Veröffentlicht: (2025)
von: Joshi, Vinay, et al.
Veröffentlicht: (2025)
Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
von: Wang, Xingrui, et al.
Veröffentlicht: (2024)
von: Wang, Xingrui, et al.
Veröffentlicht: (2024)
SAND-Math: Using LLMs to Generate Novel, Difficult and Useful Mathematics Questions and Answers
von: Manem, Chaitanya, et al.
Veröffentlicht: (2025)
von: Manem, Chaitanya, et al.
Veröffentlicht: (2025)
Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm
von: Zhang, Lin, et al.
Veröffentlicht: (2025)
von: Zhang, Lin, et al.
Veröffentlicht: (2025)
Taming Diffusion Prior for Image Super-Resolution with Domain Shift SDEs
von: Cui, Qinpeng, et al.
Veröffentlicht: (2024)
von: Cui, Qinpeng, et al.
Veröffentlicht: (2024)
Notational Animating: An Interactive Approach to Creating and Editing Animation Keyframes
von: Shi, Xinyu, et al.
Veröffentlicht: (2026)
von: Shi, Xinyu, et al.
Veröffentlicht: (2026)
KeyVideoLLM: Towards Large-scale Video Keyframe Selection
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
PR2: Predictive Routing Replay for MoE-Based LLM Reinforcement Learning
von: Dong, Daize, et al.
Veröffentlicht: (2026)
von: Dong, Daize, et al.
Veröffentlicht: (2026)
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
von: Singh, Aditya Kumar, et al.
Veröffentlicht: (2026)
von: Singh, Aditya Kumar, et al.
Veröffentlicht: (2026)
XS-VID: An Extremely Small Video Object Detection Dataset
von: Guo, Jiahao, et al.
Veröffentlicht: (2024)
von: Guo, Jiahao, et al.
Veröffentlicht: (2024)
Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation
von: Xu, Mingwang, et al.
Veröffentlicht: (2024)
von: Xu, Mingwang, et al.
Veröffentlicht: (2024)
A Comprehensive Review and Taxonomy of Audio-Visual Synchronization Techniques for Realistic Speech Animation
von: Fernandes, Jose Geraldo, et al.
Veröffentlicht: (2024)
von: Fernandes, Jose Geraldo, et al.
Veröffentlicht: (2024)
DiffSparse: Accelerating Diffusion Transformers with Learned Token Sparsity
von: Zhu, Haowei, et al.
Veröffentlicht: (2026)
von: Zhu, Haowei, et al.
Veröffentlicht: (2026)
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
von: He, Haibin, et al.
Veröffentlicht: (2026)
von: He, Haibin, et al.
Veröffentlicht: (2026)
Content and Style Aware Audio-Driven Facial Animation
von: Liu, Qingju, et al.
Veröffentlicht: (2024)
von: Liu, Qingju, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models
von: Wang, Xingrui, et al.
Veröffentlicht: (2025) -
MOVi: Training-free Text-conditioned Multi-Object Video Generation
von: Rahman, Aimon, et al.
Veröffentlicht: (2025) -
Unleashing Hour-Scale Video Training for Long Video-Language Understanding
von: Lin, Jingyang, et al.
Veröffentlicht: (2025) -
VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking
von: Lin, Jingyang, et al.
Veröffentlicht: (2026) -
Agent Laboratory: Using LLM Agents as Research Assistants
von: Schmidgall, Samuel, et al.
Veröffentlicht: (2025)