SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Xinyao, Dong, Wenkai, Song, Yuxin, Fang, Bo, Zhang, Qi, Wang, Jing, Chen, Fan, Zhang, Hui, Feng, Haocheng, Lu, Yu, Zhou, Hang, Yuan, Chun, Wang, Jingdong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
by: Song, Yuxin, et al.
Published: (2025)
by: Song, Yuxin, et al.
Published: (2025)
RefAlign: Representation Alignment for Reference-to-Video Generation
by: Wang, Lei, et al.
Published: (2026)
by: Wang, Lei, et al.
Published: (2026)
MOSA: Motion-Guided Semantic Alignment for Dynamic Scene Graph Generation
by: Wang, Xuejiao, et al.
Published: (2026)
by: Wang, Xuejiao, et al.
Published: (2026)
MotionFollower: Editing Video Motion via Lightweight Score-Guided Diffusion
by: Tu, Shuyuan, et al.
Published: (2024)
by: Tu, Shuyuan, et al.
Published: (2024)
VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
by: Cong, Xiaoyan, et al.
Published: (2025)
by: Cong, Xiaoyan, et al.
Published: (2025)
SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models
by: Sun, Ye, et al.
Published: (2025)
by: Sun, Ye, et al.
Published: (2025)
Motion-Guided Semantic Alignment with Negative Prompts for Zero-Shot Video Action Recognition
by: Wang, Yiming, et al.
Published: (2026)
by: Wang, Yiming, et al.
Published: (2026)
MoCA-Video: Motion-Aware Concept Alignment for Consistent Video Editing
by: Zhang, Tong, et al.
Published: (2025)
by: Zhang, Tong, et al.
Published: (2025)
InstructVEdit: A Holistic Approach for Instructional Video Editing
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
In-Context Learning with Unpaired Clips for Instruction-based Video Editing
by: Liao, Xinyao, et al.
Published: (2025)
by: Liao, Xinyao, et al.
Published: (2025)
ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions
by: Chang, Di, et al.
Published: (2025)
by: Chang, Di, et al.
Published: (2025)
DISPLAY: Directable Human-Object Interaction Video Generation via Sparse Motion Guidance and Multi-Task Auxiliary
by: Guan, Jiazhi, et al.
Published: (2026)
by: Guan, Jiazhi, et al.
Published: (2026)
IVEBench: Modern Benchmark Suite for Instruction-Guided Video Editing Assessment
by: Chen, Yinan, et al.
Published: (2025)
by: Chen, Yinan, et al.
Published: (2025)
GaussianEditor: Editing 3D Gaussians Delicately with Text Instructions
by: Wang, Junjie, et al.
Published: (2023)
by: Wang, Junjie, et al.
Published: (2023)
MotionAgent: Fine-grained Controllable Video Generation via Motion Field Agent
by: Liao, Xinyao, et al.
Published: (2025)
by: Liao, Xinyao, et al.
Published: (2025)
Agent-SAMA: State-Aware Mobile Assistant
by: Guo, Linqiang, et al.
Published: (2025)
by: Guo, Linqiang, et al.
Published: (2025)
Re-HOLD: Video Hand Object Interaction Reenactment via adaptive Layout-instructed Diffusion Model
by: Fan, Yingying, et al.
Published: (2025)
by: Fan, Yingying, et al.
Published: (2025)
Disentangling Instruction Influence in Diffusion Transformers for Parallel Multi-Instruction-Guided Image Editing
by: Liu, Hui, et al.
Published: (2025)
by: Liu, Hui, et al.
Published: (2025)
Diffusion-Based Image Editing: An Unforeseen Adversary to Robust Invisible Watermarks
by: Fu, Wenkai, et al.
Published: (2025)
by: Fu, Wenkai, et al.
Published: (2025)
UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception
by: Song, Xinyang, et al.
Published: (2025)
by: Song, Xinyang, et al.
Published: (2025)
Routing to the Right Expertise: A Trustworthy Judge for Instruction-based Image Editing
by: Sun, Chenxi, et al.
Published: (2025)
by: Sun, Chenxi, et al.
Published: (2025)
What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing
by: Lin, Hangyu, et al.
Published: (2026)
by: Lin, Hangyu, et al.
Published: (2026)
LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing
by: Wang, Weicheng, et al.
Published: (2026)
by: Wang, Weicheng, et al.
Published: (2026)
EALG: Evolutionary Adversarial Generation of Language Model-Guided Generators for Combinatorial Optimization
by: Duan, Ruibo, et al.
Published: (2025)
by: Duan, Ruibo, et al.
Published: (2025)
InstructEngine: Instruction-driven Text-to-Image Alignment
by: Lu, Xingyu, et al.
Published: (2025)
by: Lu, Xingyu, et al.
Published: (2025)
Semantics-Aware Human Motion Generation from Audio Instructions
by: Wang, Zi-An, et al.
Published: (2025)
by: Wang, Zi-An, et al.
Published: (2025)
Semantics-aware Motion Retargeting with Vision-Language Models
by: Zhang, Haodong, et al.
Published: (2023)
by: Zhang, Haodong, et al.
Published: (2023)
Occlusion-Aware Physics-Semantic Keyframe Selection for Robust Video Editing
by: Liu, Lin, et al.
Published: (2026)
by: Liu, Lin, et al.
Published: (2026)
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
by: He, Haibin, et al.
Published: (2026)
by: He, Haibin, et al.
Published: (2026)
Shape-for-Motion: Precise and Consistent Video Editing with 3D Proxy
by: Liu, Yuhao, et al.
Published: (2025)
by: Liu, Yuhao, et al.
Published: (2025)
Causal Path Alignment: Anchoring the Optimization Trajectory for Controllable In-Parameter Knowledge Editing
by: Liu, Xiyu, et al.
Published: (2025)
by: Liu, Xiyu, et al.
Published: (2025)
OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing
by: He, Haoyang, et al.
Published: (2025)
by: He, Haoyang, et al.
Published: (2025)
RECOST: External Knowledge Guided Data-efficient Instruction Tuning
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrieval
by: Yang, Yuxin, et al.
Published: (2026)
by: Yang, Yuxin, et al.
Published: (2026)
STANCE: Motion Coherent Video Generation Via Sparse-to-Dense Anchored Encoding
by: Chen, Zhifei, et al.
Published: (2025)
by: Chen, Zhifei, et al.
Published: (2025)
Semantic-Guided Unsupervised Video Summarization
by: Liu, Haizhou, et al.
Published: (2026)
by: Liu, Haizhou, et al.
Published: (2026)
Unveiling Deep Semantic Uncertainty Perception for Language-Anchored Multi-modal Vision-Brain Alignment
by: Feng, Zehui, et al.
Published: (2025)
by: Feng, Zehui, et al.
Published: (2025)
MagDiff: Multi-Alignment Diffusion for High-Fidelity Video Generation and Editing
by: Zhao, Haoyu, et al.
Published: (2023)
by: Zhao, Haoyu, et al.
Published: (2023)
Visual Autoregressive Modeling for Instruction-Guided Image Editing
by: Mao, Qingyang, et al.
Published: (2025)
by: Mao, Qingyang, et al.
Published: (2025)
GeoBridge: A Semantic-Anchored Multi-View Foundation Model Bridging Images and Text for Geo-Localization
by: Song, Zixuan, et al.
Published: (2025)
by: Song, Zixuan, et al.
Published: (2025)
Similar Items
-
Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
by: Song, Yuxin, et al.
Published: (2025) -
RefAlign: Representation Alignment for Reference-to-Video Generation
by: Wang, Lei, et al.
Published: (2026) -
MOSA: Motion-Guided Semantic Alignment for Dynamic Scene Graph Generation
by: Wang, Xuejiao, et al.
Published: (2026) -
MotionFollower: Editing Video Motion via Lightweight Score-Guided Diffusion
by: Tu, Shuyuan, et al.
Published: (2024) -
VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
by: Cong, Xiaoyan, et al.
Published: (2025)