VIA: Unified Spatiotemporal Video Adaptation Framework for Global and Local Video Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Gu, Jing, Fang, Yuwei, Skorokhodov, Ivan, Wonka, Peter, Du, Xinya, Tulyakov, Sergey, Wang, Xin Eric |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Error Analyses of Auto-Regressive Video Diffusion Models: A Unified Framework
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
Taming Flow-based I2V Models for Creative Video Editing
by: Kong, Xianghao, et al.
Published: (2025)
by: Kong, Xianghao, et al.
Published: (2025)
TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long Videos
by: Zhang, Chen-Lin, et al.
Published: (2025)
by: Zhang, Chen-Lin, et al.
Published: (2025)
VCoME: Verbal Video Composition with Multimodal Editing Effects
by: Gong, Weibo, et al.
Published: (2024)
by: Gong, Weibo, et al.
Published: (2024)
Hybrid Local-Global Context Learning for Neural Video Compression
by: Zhai, Yongqi, et al.
Published: (2024)
by: Zhai, Yongqi, et al.
Published: (2024)
Compressed Deepfake Video Detection Based on 3D Spatiotemporal Trajectories
by: Chen, Zongmei, et al.
Published: (2024)
by: Chen, Zongmei, et al.
Published: (2024)
Region-Constraint In-Context Generation for Instructional Video Editing
by: Zhang, Zhongwei, et al.
Published: (2025)
by: Zhang, Zhongwei, et al.
Published: (2025)
AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
by: Ku, Max, et al.
Published: (2024)
by: Ku, Max, et al.
Published: (2024)
Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis
by: Menapace, Willi, et al.
Published: (2024)
by: Menapace, Willi, et al.
Published: (2024)
Bridging Your Imagination with Audio-Video Generation via a Unified Director
by: Zhang, Jiaxu, et al.
Published: (2025)
by: Zhang, Jiaxu, et al.
Published: (2025)
Hierarchical Patch Diffusion Models for High-Resolution Video Generation
by: Skorokhodov, Ivan, et al.
Published: (2024)
by: Skorokhodov, Ivan, et al.
Published: (2024)
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
by: Li, Hebeizi, et al.
Published: (2026)
by: Li, Hebeizi, et al.
Published: (2026)
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
by: Zhang, Pingping, et al.
Published: (2024)
by: Zhang, Pingping, et al.
Published: (2024)
Edit As You Wish: Video Caption Editing with Multi-grained User Control
by: Yao, Linli, et al.
Published: (2023)
by: Yao, Linli, et al.
Published: (2023)
FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
by: Cai, Lingling, et al.
Published: (2024)
by: Cai, Lingling, et al.
Published: (2024)
UBiSS: A Unified Framework for Bimodal Semantic Summarization of Videos
by: Mei, Yuting, et al.
Published: (2024)
by: Mei, Yuting, et al.
Published: (2024)
VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control
by: Bian, Yuxuan, et al.
Published: (2025)
by: Bian, Yuxuan, et al.
Published: (2025)
STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models
by: Fan, Linfeng, et al.
Published: (2026)
by: Fan, Linfeng, et al.
Published: (2026)
When and How to Cut Classical Concerts? A Multimodal Automated Video Editing Approach
by: Gonzálbez-Biosca, Daniel, et al.
Published: (2025)
by: Gonzálbez-Biosca, Daniel, et al.
Published: (2025)
Panonut360: A Head and Eye Tracking Dataset for Panoramic Video
by: Xu, Yutong, et al.
Published: (2024)
by: Xu, Yutong, et al.
Published: (2024)
GLANCE: A Global-Local Coordination Multi-Agent Framework for Music-Grounded Non-Linear Video Editing
by: Lin, Zihao, et al.
Published: (2026)
by: Lin, Zihao, et al.
Published: (2026)
LayerT2V: A Unified Multi-Layer Video Generation Framework
by: Li, Guangzhao, et al.
Published: (2025)
by: Li, Guangzhao, et al.
Published: (2025)
VideoMap: Supporting Video Editing Exploration, Brainstorming, and Prototyping in the Latent Space
by: Lin, David Chuan-En, et al.
Published: (2022)
by: Lin, David Chuan-En, et al.
Published: (2022)
Learning Spatial Adaptation and Temporal Coherence in Diffusion Models for Video Super-Resolution
by: Chen, Zhikai, et al.
Published: (2024)
by: Chen, Zhikai, et al.
Published: (2024)
Interactive Video Generation via Domain Adaptation
by: Rawal, Ishaan, et al.
Published: (2025)
by: Rawal, Ishaan, et al.
Published: (2025)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
by: Fang, Xinyu, et al.
Published: (2024)
by: Fang, Xinyu, et al.
Published: (2024)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
by: Zhang, Xueqiao, et al.
Published: (2025)
by: Zhang, Xueqiao, et al.
Published: (2025)
Sound Sparks Motion: Audio and Text Tuning for Video Editing
by: Razlighi, AmirHossein Naghi, et al.
Published: (2026)
by: Razlighi, AmirHossein Naghi, et al.
Published: (2026)
Editing Physiological Signals in Videos Using Latent Representations
by: Zhou, Tianwen, et al.
Published: (2025)
by: Zhou, Tianwen, et al.
Published: (2025)
StreamingEval: A Unified Evaluation Protocol towards Realistic Streaming Video Understanding
by: Tang, Guowei, et al.
Published: (2026)
by: Tang, Guowei, et al.
Published: (2026)
Mind the Time: Temporally-Controlled Multi-Event Video Generation
by: Wu, Ziyi, et al.
Published: (2024)
by: Wu, Ziyi, et al.
Published: (2024)
FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis
by: Liang, Feng, et al.
Published: (2023)
by: Liang, Feng, et al.
Published: (2023)
JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation
by: Liu, Kai, et al.
Published: (2026)
by: Liu, Kai, et al.
Published: (2026)
Omni2Sound: Towards Unified Video-Text-to-Audio Generation
by: Dai, Yusheng, et al.
Published: (2026)
by: Dai, Yusheng, et al.
Published: (2026)
OneHOI: Unifying Human-Object Interaction Generation and Editing
by: Hoe, Jiun Tian, et al.
Published: (2026)
by: Hoe, Jiun Tian, et al.
Published: (2026)
Quality-Aware Dynamic Resolution Adaptation Framework for Adaptive Video Streaming
by: Premkumar, Amritha, et al.
Published: (2024)
by: Premkumar, Amritha, et al.
Published: (2024)
A Tri-Dynamic Preprocessing Framework for UGC Video Compression
by: Zhao, Fei, et al.
Published: (2025)
by: Zhao, Fei, et al.
Published: (2025)
A Preprocessing Framework for Video Machine Vision under Compression
by: Zhao, Fei, et al.
Published: (2025)
by: Zhao, Fei, et al.
Published: (2025)
Composing Concepts from Images and Videos via Concept-prompt Binding
by: Kong, Xianghao, et al.
Published: (2025)
by: Kong, Xianghao, et al.
Published: (2025)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
by: Qin, Bosheng, et al.
Published: (2023)
by: Qin, Bosheng, et al.
Published: (2023)
Similar Items
-
Error Analyses of Auto-Regressive Video Diffusion Models: A Unified Framework
by: Wang, Jing, et al.
Published: (2025) -
Taming Flow-based I2V Models for Creative Video Editing
by: Kong, Xianghao, et al.
Published: (2025) -
TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long Videos
by: Zhang, Chen-Lin, et al.
Published: (2025) -
VCoME: Verbal Video Composition with Multimodal Editing Effects
by: Gong, Weibo, et al.
Published: (2024) -
Hybrid Local-Global Context Learning for Neural Video Compression
by: Zhai, Yongqi, et al.
Published: (2024)