Saved in:
| Main Authors: | Li, Bingxuan, Cui, Yiming, He, Yicheng, Wang, Yiwei, Zhang, Shu, Wen, Longyin, Niu, Yulei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.24731 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound
by: Lee, Junwon, et al.
Published: (2024)
by: Lee, Junwon, et al.
Published: (2024)
Video-Guided Foley Sound Generation with Multimodal Controls
by: Chen, Ziyang, et al.
Published: (2024)
by: Chen, Ziyang, et al.
Published: (2024)
Referring Layer Decomposition
by: Chen, Fangyi, et al.
Published: (2026)
by: Chen, Fangyi, et al.
Published: (2026)
TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
by: Zhang, Guofeng, et al.
Published: (2025)
by: Zhang, Guofeng, et al.
Published: (2025)
HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation
by: Shan, Sizhe, et al.
Published: (2025)
by: Shan, Sizhe, et al.
Published: (2025)
ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling
by: Yang, Jianxuan, et al.
Published: (2026)
by: Yang, Jianxuan, et al.
Published: (2026)
EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use
by: Wen, Siwei, et al.
Published: (2026)
by: Wen, Siwei, et al.
Published: (2026)
Thinking With Bounding Boxes: Enhancing Spatio-Temporal Video Grounding via Reinforcement Fine-Tuning
by: Gu, Xin, et al.
Published: (2025)
by: Gu, Xin, et al.
Published: (2025)
Where do Large Vision-Language Models Look at when Answering Questions?
by: Xing, Xiaoying, et al.
Published: (2025)
by: Xing, Xiaoying, et al.
Published: (2025)
Structured Context Learning for Generic Event Boundary Detection
by: Gu, Xin, et al.
Published: (2025)
by: Gu, Xin, et al.
Published: (2025)
Chain of Event-Centric Causal Thought for Physically Plausible Video Generation
by: Wang, Zixuan, et al.
Published: (2026)
by: Wang, Zixuan, et al.
Published: (2026)
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts
by: Li, You, et al.
Published: (2026)
by: Li, You, et al.
Published: (2026)
FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips
by: Li, Mengtian, et al.
Published: (2026)
by: Li, Mengtian, et al.
Published: (2026)
Foley Control: Aligning a Frozen Latent Text-to-Audio Model to Video
by: Rowles, Ciara, et al.
Published: (2025)
by: Rowles, Ciara, et al.
Published: (2025)
PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment
by: Gramaccioni, Riccardo Fosco, et al.
Published: (2024)
by: Gramaccioni, Riccardo Fosco, et al.
Published: (2024)
Echo-Path: Pathology-Conditioned Echo Video Generation
by: Muhammad, Kabir Hamzah, et al.
Published: (2025)
by: Muhammad, Kabir Hamzah, et al.
Published: (2025)
TIE: Time Interval Encoding for Video Generation over Events
by: Shu, Zhilei, et al.
Published: (2026)
by: Shu, Zhilei, et al.
Published: (2026)
ArrowGEV: Grounding Events in Video via Learning the Arrow of Time
by: Yu, Fangxu, et al.
Published: (2026)
by: Yu, Fangxu, et al.
Published: (2026)
Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding
by: Zheng, Minghang, et al.
Published: (2025)
by: Zheng, Minghang, et al.
Published: (2025)
Accurate and Fast Compressed Video Captioning
by: Shen, Yaojie, et al.
Published: (2023)
by: Shen, Yaojie, et al.
Published: (2023)
CyberV: Cybernetics for Test-time Scaling in Video Understanding
by: Meng, Jiahao, et al.
Published: (2025)
by: Meng, Jiahao, et al.
Published: (2025)
Edit3K: Universal Representation Learning for Video Editing Components
by: Gu, Xin, et al.
Published: (2024)
by: Gu, Xin, et al.
Published: (2024)
FoleyGRAM: Video-to-Audio Generation with GRAM-Aligned Multimodal Encoders
by: Gramaccioni, Riccardo Fosco, et al.
Published: (2025)
by: Gramaccioni, Riccardo Fosco, et al.
Published: (2025)
METAL: A Multi-Agent Framework for Chart Generation with Test-Time Scaling
by: Li, Bingxuan, et al.
Published: (2025)
by: Li, Bingxuan, et al.
Published: (2025)
Beyond Raw Videos: Understanding Edited Videos with Large Multimodal Model
by: Xu, Lu, et al.
Published: (2024)
by: Xu, Lu, et al.
Published: (2024)
RAP: Retrieval-Augmented Planner for Adaptive Procedure Planning in Instructional Videos
by: Zare, Ali, et al.
Published: (2024)
by: Zare, Ali, et al.
Published: (2024)
E.M.Ground: A Temporal Grounding Vid-LLM with Holistic Event Perception and Matching
by: Nie, Jiahao, et al.
Published: (2026)
by: Nie, Jiahao, et al.
Published: (2026)
EchoShot: Multi-Shot Portrait Video Generation
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
by: Cheng, Shihao, et al.
Published: (2026)
by: Cheng, Shihao, et al.
Published: (2026)
D-Attn: Decomposed Attention for Large Vision-and-Language Models
by: Kuo, Chia-Wen, et al.
Published: (2025)
by: Kuo, Chia-Wen, et al.
Published: (2025)
DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation
by: Guo, Xu, et al.
Published: (2026)
by: Guo, Xu, et al.
Published: (2026)
AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer
by: Fang, Pengjun, et al.
Published: (2026)
by: Fang, Pengjun, et al.
Published: (2026)
VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing
by: Deng, Andong, et al.
Published: (2026)
by: Deng, Andong, et al.
Published: (2026)
EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion
by: Wei, Jiangchuan, et al.
Published: (2025)
by: Wei, Jiangchuan, et al.
Published: (2025)
SlotVTG: Object-Centric Adapter for Generalizable Video Temporal Grounding
by: Han, Jiwook, et al.
Published: (2026)
by: Han, Jiwook, et al.
Published: (2026)
SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation
by: Dou, Weijia, et al.
Published: (2026)
by: Dou, Weijia, et al.
Published: (2026)
TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation
by: Yan, Pengyu, et al.
Published: (2026)
by: Yan, Pengyu, et al.
Published: (2026)
CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognition
by: Yang, Hongji, et al.
Published: (2026)
by: Yang, Hongji, et al.
Published: (2026)
Similar Items
-
FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
by: Zhang, Yiming, et al.
Published: (2024) -
Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound
by: Lee, Junwon, et al.
Published: (2024) -
Video-Guided Foley Sound Generation with Multimodal Controls
by: Chen, Ziyang, et al.
Published: (2024) -
Referring Layer Decomposition
by: Chen, Fangyi, et al.
Published: (2026) -
TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
by: Zhang, Guofeng, et al.
Published: (2025)