InstructVEdit: A Holistic Approach for Instructional Video Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Chi, Feng, Chengjian, Yan, Feng, Zhang, Qiming, Zhang, Mingjin, Zhong, Yujie, Zhang, Jing, Ma, Lin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unleashing the Power of Generic Segmentation Models: A Simple Baseline for Infrared Small Target Detection
by: Zhang, Mingjin, et al.
Published: (2024)
by: Zhang, Mingjin, et al.
Published: (2024)
UniMD: Towards Unifying Moment Retrieval and Temporal Action Detection
by: Zeng, Yingsen, et al.
Published: (2024)
by: Zeng, Yingsen, et al.
Published: (2024)
I2VEdit: First-Frame-Guided Video Editing via Image-to-Video Diffusion Models
by: Ouyang, Wenqi, et al.
Published: (2024)
by: Ouyang, Wenqi, et al.
Published: (2024)
Omni-I2C: A Holistic Benchmark for High-Fidelity Image-to-Code Generation
by: Zhou, Jiawei, et al.
Published: (2026)
by: Zhou, Jiawei, et al.
Published: (2026)
InstaGen: Enhancing Object Detection by Training on Synthetic Dataset
by: Feng, Chengjian, et al.
Published: (2024)
by: Feng, Chengjian, et al.
Published: (2024)
DisTime: Distribution-based Time Representation for Video Large Language Models
by: Zeng, Yingsen, et al.
Published: (2025)
by: Zeng, Yingsen, et al.
Published: (2025)
Na-IRSTD: Enhancing Infrared Small Target Detection via Native-Resolution Feature Selection and Fusion
by: Xu, Qian, et al.
Published: (2026)
by: Xu, Qian, et al.
Published: (2026)
RoboTron-Sim: Improving Real-World Driving via Simulated Hard-Case
by: Xiao, Baihui, et al.
Published: (2025)
by: Xiao, Baihui, et al.
Published: (2025)
InstructAV2AV: Instruction-Guided Audio-Video Joint Editing
by: Zheng, Haojie, et al.
Published: (2026)
by: Zheng, Haojie, et al.
Published: (2026)
SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing
by: Zhang, Xinyao, et al.
Published: (2026)
by: Zhang, Xinyao, et al.
Published: (2026)
DeepSketcher: Internalizing Visual Manipulation for Multimodal Reasoning
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
InstructRestore: Region-Customized Image Restoration with Human Instructions
by: Liu, Shuaizheng, et al.
Published: (2025)
by: Liu, Shuaizheng, et al.
Published: (2025)
ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
by: Fu, Haoyu, et al.
Published: (2025)
by: Fu, Haoyu, et al.
Published: (2025)
InstructDET: Diversifying Referring Object Detection with Generalized Instructions
by: Dang, Ronghao, et al.
Published: (2023)
by: Dang, Ronghao, et al.
Published: (2023)
RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving
by: Huang, Zhijian, et al.
Published: (2024)
by: Huang, Zhijian, et al.
Published: (2024)
InstructEngine: Instruction-driven Text-to-Image Alignment
by: Lu, Xingyu, et al.
Published: (2025)
by: Lu, Xingyu, et al.
Published: (2025)
InstructAttribute: Fine-grained Object Attributes editing with Instruction
by: Yin, Xingxi, et al.
Published: (2025)
by: Yin, Xingxi, et al.
Published: (2025)
Hierarchical Action Recognition: A Contrastive Video-Language Approach with Hierarchical Interactions
by: Zhang, Rui, et al.
Published: (2024)
by: Zhang, Rui, et al.
Published: (2024)
InstructSAM: Segment Any Instance with Any Instructions
by: Yuan, Yuqian, et al.
Published: (2026)
by: Yuan, Yuqian, et al.
Published: (2026)
AP-CAP: Advancing High-Quality Data Synthesis for Animal Pose Estimation via a Controllable Image Generation Pipeline
by: Wang, Lei, et al.
Published: (2025)
by: Wang, Lei, et al.
Published: (2025)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
by: Qin, Bosheng, et al.
Published: (2023)
by: Qin, Bosheng, et al.
Published: (2023)
VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
by: Cong, Xiaoyan, et al.
Published: (2025)
by: Cong, Xiaoyan, et al.
Published: (2025)
SketchVideo: Sketch-based Video Generation and Editing
by: Liu, Feng-Lin, et al.
Published: (2025)
by: Liu, Feng-Lin, et al.
Published: (2025)
Learning to Instruct for Visual Instruction Tuning
by: Zhou, Zhihan, et al.
Published: (2025)
by: Zhou, Zhihan, et al.
Published: (2025)
InstructLayout: Instruction-Driven 2D and 3D Layout Synthesis with Semantic Graph Prior
by: Lin, Chenguo, et al.
Published: (2024)
by: Lin, Chenguo, et al.
Published: (2024)
InstructUDrag: Joint Text Instructions and Object Dragging for Interactive Image Editing
by: Yu, Haoran, et al.
Published: (2025)
by: Yu, Haoran, et al.
Published: (2025)
RFSR: Improving ISR Diffusion Models via Reward Feedback Learning
by: Sun, Xiaopeng, et al.
Published: (2024)
by: Sun, Xiaopeng, et al.
Published: (2024)
LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing
by: Wang, Weicheng, et al.
Published: (2026)
by: Wang, Weicheng, et al.
Published: (2026)
InstructX: Towards Unified Visual Editing with MLLM Guidance
by: Mou, Chong, et al.
Published: (2025)
by: Mou, Chong, et al.
Published: (2025)
InstructHumans: Editing Animated 3D Human Textures with Instructions
by: Zhu, Jiayin, et al.
Published: (2024)
by: Zhu, Jiayin, et al.
Published: (2024)
InstructBrush: Learning Attention-based Instruction Optimization for Image Editing
by: Zhao, Ruoyu, et al.
Published: (2024)
by: Zhao, Ruoyu, et al.
Published: (2024)
TIGER: Text-Instructed 3D Gaussian Retrieval and Coherent Editing
by: Xu, Teng, et al.
Published: (2024)
by: Xu, Teng, et al.
Published: (2024)
Reasoning to Align: Implicit Reasoning in Diffusion Transformers for Video Editing
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
RSEdit: Text-Guided Image Editing for Remote Sensing
by: Zhenyuan, Chen, et al.
Published: (2026)
by: Zhenyuan, Chen, et al.
Published: (2026)
InstructAny2Pix: Flexible Visual Editing via Multimodal Instruction Following
by: Li, Shufan, et al.
Published: (2023)
by: Li, Shufan, et al.
Published: (2023)
IVEBench: Modern Benchmark Suite for Instruction-Guided Video Editing Assessment
by: Chen, Yinan, et al.
Published: (2025)
by: Chen, Yinan, et al.
Published: (2025)
InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models
by: Wei, Cong, et al.
Published: (2024)
by: Wei, Cong, et al.
Published: (2024)
ZONE: Zero-Shot Instruction-Guided Local Editing
by: Li, Shanglin, et al.
Published: (2023)
by: Li, Shanglin, et al.
Published: (2023)
RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
by: Liu, Fanfan, et al.
Published: (2024)
by: Liu, Fanfan, et al.
Published: (2024)
Region-Constraint In-Context Generation for Instructional Video Editing
by: Zhang, Zhongwei, et al.
Published: (2025)
by: Zhang, Zhongwei, et al.
Published: (2025)
Similar Items
-
Unleashing the Power of Generic Segmentation Models: A Simple Baseline for Infrared Small Target Detection
by: Zhang, Mingjin, et al.
Published: (2024) -
UniMD: Towards Unifying Moment Retrieval and Temporal Action Detection
by: Zeng, Yingsen, et al.
Published: (2024) -
I2VEdit: First-Frame-Guided Video Editing via Image-to-Video Diffusion Models
by: Ouyang, Wenqi, et al.
Published: (2024) -
Omni-I2C: A Holistic Benchmark for High-Fidelity Image-to-Code Generation
by: Zhou, Jiawei, et al.
Published: (2026) -
InstaGen: Enhancing Object Detection by Training on Synthetic Dataset
by: Feng, Chengjian, et al.
Published: (2024)