InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Yiyuan, Kang, Yuhao, Zhang, Zhixin, Ding, Xiaohan, Zhao, Sanyuan, Yue, Xiangyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multimodal Long Video Modeling Based on Temporal Dynamic Context
von: Hao, Haoran, et al.
Veröffentlicht: (2025)
von: Hao, Haoran, et al.
Veröffentlicht: (2025)
Region-Constraint In-Context Generation for Instructional Video Editing
von: Zhang, Zhongwei, et al.
Veröffentlicht: (2025)
von: Zhang, Zhongwei, et al.
Veröffentlicht: (2025)
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
von: Wang, Yun, et al.
Veröffentlicht: (2025)
von: Wang, Yun, et al.
Veröffentlicht: (2025)
D-FCGS: Feedforward Compression of Dynamic Gaussian Splatting for Free-Viewpoint Videos
von: Zhang, Wenkang, et al.
Veröffentlicht: (2025)
von: Zhang, Wenkang, et al.
Veröffentlicht: (2025)
VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification
von: Meng, Jiahao, et al.
Veröffentlicht: (2026)
von: Meng, Jiahao, et al.
Veröffentlicht: (2026)
Edit As You Wish: Video Caption Editing with Multi-grained User Control
von: Yao, Linli, et al.
Veröffentlicht: (2023)
von: Yao, Linli, et al.
Veröffentlicht: (2023)
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
von: Cheng, Shihao, et al.
Veröffentlicht: (2026)
von: Cheng, Shihao, et al.
Veröffentlicht: (2026)
Viewport Prediction for Volumetric Video Streaming by Exploring Video Saliency and Trajectory Information
von: Li, Jie, et al.
Veröffentlicht: (2023)
von: Li, Jie, et al.
Veröffentlicht: (2023)
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
SSNVC: Single Stream Neural Video Compression with Implicit Temporal Information
von: Wang, Feng, et al.
Veröffentlicht: (2024)
von: Wang, Feng, et al.
Veröffentlicht: (2024)
DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation
von: Cai, Minghong, et al.
Veröffentlicht: (2024)
von: Cai, Minghong, et al.
Veröffentlicht: (2024)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
AIS 2024 Challenge on Video Quality Assessment of User-Generated Content: Methods and Results
von: Conde, Marcos V., et al.
Veröffentlicht: (2024)
von: Conde, Marcos V., et al.
Veröffentlicht: (2024)
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Learning Video Context as Interleaved Multimodal Sequences
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
von: Cai, Lingling, et al.
Veröffentlicht: (2024)
von: Cai, Lingling, et al.
Veröffentlicht: (2024)
MotionPro: A Precise Motion Controller for Image-to-Video Generation
von: Zhang, Zhongwei, et al.
Veröffentlicht: (2025)
von: Zhang, Zhongwei, et al.
Veröffentlicht: (2025)
Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence
von: Liao, Junchao, et al.
Veröffentlicht: (2026)
von: Liao, Junchao, et al.
Veröffentlicht: (2026)
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
MicroEmo: Time-Sensitive Multimodal Emotion Recognition with Micro-Expression Dynamics in Video Dialogues
von: Zhang, Liyun
Veröffentlicht: (2024)
von: Zhang, Liyun
Veröffentlicht: (2024)
Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing
von: Zhao, Pengcheng, et al.
Veröffentlicht: (2024)
von: Zhao, Pengcheng, et al.
Veröffentlicht: (2024)
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering
von: Meng, Yiran, et al.
Veröffentlicht: (2025)
von: Meng, Yiran, et al.
Veröffentlicht: (2025)
Image Conductor: Precision Control for Interactive Video Synthesis
von: Li, Yaowei, et al.
Veröffentlicht: (2024)
von: Li, Yaowei, et al.
Veröffentlicht: (2024)
VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
von: Lan, Xiaohan, et al.
Veröffentlicht: (2024)
von: Lan, Xiaohan, et al.
Veröffentlicht: (2024)
Hierarchical Action Recognition: A Contrastive Video-Language Approach with Hierarchical Interactions
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
MLLM-VADStory: Domain Knowledge-Driven Multimodal LLMs for Video Ad Storyline Insights
von: Yang, Jasmine, et al.
Veröffentlicht: (2026)
von: Yang, Jasmine, et al.
Veröffentlicht: (2026)
Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation
von: He, Liu, et al.
Veröffentlicht: (2024)
von: He, Liu, et al.
Veröffentlicht: (2024)
Explore the Limits of Omni-modal Pretraining at Scale
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
Human Motion Video Generation: A Survey
von: Xue, Haiwei, et al.
Veröffentlicht: (2025)
von: Xue, Haiwei, et al.
Veröffentlicht: (2025)
Generative Frame Sampler for Long Video Understanding
von: Yao, Linli, et al.
Veröffentlicht: (2025)
von: Yao, Linli, et al.
Veröffentlicht: (2025)
TAVGBench: Benchmarking Text to Audible-Video Generation
von: Mao, Yuxin, et al.
Veröffentlicht: (2024)
von: Mao, Yuxin, et al.
Veröffentlicht: (2024)
Follow-Your-MultiPose: Tuning-Free Multi-Character Text-to-Video Generation via Pose Guidance
von: Zhang, Beiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Beiyuan, et al.
Veröffentlicht: (2024)
Consistency-aware Fake Videos Detection on Short Video Platforms
von: Wang, Junxi, et al.
Veröffentlicht: (2025)
von: Wang, Junxi, et al.
Veröffentlicht: (2025)
Signal: Selective Interaction and Global-local Alignment for Multi-Modal Object Re-Identification
von: Liu, Yangyang, et al.
Veröffentlicht: (2025)
von: Liu, Yangyang, et al.
Veröffentlicht: (2025)
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
von: Pramanick, Shraman, et al.
Veröffentlicht: (2025)
von: Pramanick, Shraman, et al.
Veröffentlicht: (2025)
Multimodal Engagement Analysis from Facial Videos in the Classroom
von: Sümer, Ömer, et al.
Veröffentlicht: (2021)
von: Sümer, Ömer, et al.
Veröffentlicht: (2021)
Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose Interaction
von: Li, Dong, et al.
Veröffentlicht: (2025)
von: Li, Dong, et al.
Veröffentlicht: (2025)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines
von: Zhang, Zhixin, et al.
Veröffentlicht: (2024)
von: Zhang, Zhixin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Multimodal Long Video Modeling Based on Temporal Dynamic Context
von: Hao, Haoran, et al.
Veröffentlicht: (2025) -
Region-Constraint In-Context Generation for Instructional Video Editing
von: Zhang, Zhongwei, et al.
Veröffentlicht: (2025) -
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
von: Wang, Yun, et al.
Veröffentlicht: (2025) -
D-FCGS: Feedforward Compression of Dynamic Gaussian Splatting for Free-Viewpoint Videos
von: Zhang, Wenkang, et al.
Veröffentlicht: (2025) -
VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification
von: Meng, Jiahao, et al.
Veröffentlicht: (2026)