VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cong, Xiaoyan, Yang, Haotian, Wang, Angtian, Wang, Yizhi, Yang, Yiding, Zhang, Canyu, Ma, Chongyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ATI: Any Trajectory Instruction for Controllable Video Generation
von: Wang, Angtian, et al.
Veröffentlicht: (2025)
von: Wang, Angtian, et al.
Veröffentlicht: (2025)
CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance
von: Deng, Yufan, et al.
Veröffentlicht: (2025)
von: Deng, Yufan, et al.
Veröffentlicht: (2025)
HECTOR: Hybrid Editable Compositional Object References for Video Generation
von: Zhang, Guofeng, et al.
Veröffentlicht: (2026)
von: Zhang, Guofeng, et al.
Veröffentlicht: (2026)
TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
von: Zhang, Guofeng, et al.
Veröffentlicht: (2025)
von: Zhang, Guofeng, et al.
Veröffentlicht: (2025)
MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement
von: Deng, Yufan, et al.
Veröffentlicht: (2025)
von: Deng, Yufan, et al.
Veröffentlicht: (2025)
Envision: Embodied Visual Planning via Goal-Imagery Video Diffusion
von: Gu, Yuming, et al.
Veröffentlicht: (2025)
von: Gu, Yuming, et al.
Veröffentlicht: (2025)
InstructAV2AV: Instruction-Guided Audio-Video Joint Editing
von: Zheng, Haojie, et al.
Veröffentlicht: (2026)
von: Zheng, Haojie, et al.
Veröffentlicht: (2026)
BaseReward: A Strong Baseline for Multimodal Reward Model
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2025)
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2025)
SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing
von: Zhang, Xinyao, et al.
Veröffentlicht: (2026)
von: Zhang, Xinyao, et al.
Veröffentlicht: (2026)
IVEBench: Modern Benchmark Suite for Instruction-Guided Video Editing Assessment
von: Chen, Yinan, et al.
Veröffentlicht: (2025)
von: Chen, Yinan, et al.
Veröffentlicht: (2025)
LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing
von: Wang, Weicheng, et al.
Veröffentlicht: (2026)
von: Wang, Weicheng, et al.
Veröffentlicht: (2026)
OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing
von: He, Haoyang, et al.
Veröffentlicht: (2025)
von: He, Haoyang, et al.
Veröffentlicht: (2025)
Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset
von: Bai, Qingyan, et al.
Veröffentlicht: (2025)
von: Bai, Qingyan, et al.
Veröffentlicht: (2025)
Guiding Instruction-based Image Editing via Multimodal Large Language Models
von: Fu, Tsu-Jui, et al.
Veröffentlicht: (2023)
von: Fu, Tsu-Jui, et al.
Veröffentlicht: (2023)
FSVideo: Fast Speed Video Diffusion Model in a Highly-Compressed Latent Space
von: FSVideo Team, et al.
Veröffentlicht: (2026)
von: FSVideo Team, et al.
Veröffentlicht: (2026)
Multi-Reward as Condition for Instruction-based Image Editing
von: Gu, Xin, et al.
Veröffentlicht: (2024)
von: Gu, Xin, et al.
Veröffentlicht: (2024)
InstructVEdit: A Holistic Approach for Instructional Video Editing
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing
von: Wu, Keming, et al.
Veröffentlicht: (2025)
von: Wu, Keming, et al.
Veröffentlicht: (2025)
Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling
von: Liu, Gongye, et al.
Veröffentlicht: (2026)
von: Liu, Gongye, et al.
Veröffentlicht: (2026)
Instruction Guided Multi Object Image Editing with Quantity and Layout Consistency
von: Tan, Jiaqi, et al.
Veröffentlicht: (2025)
von: Tan, Jiaqi, et al.
Veröffentlicht: (2025)
ESA: Energy-Based Shot Assembly Optimization for Automatic Video Editing
von: Chen, Yaosen, et al.
Veröffentlicht: (2025)
von: Chen, Yaosen, et al.
Veröffentlicht: (2025)
What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing
von: Lin, Hangyu, et al.
Veröffentlicht: (2026)
von: Lin, Hangyu, et al.
Veröffentlicht: (2026)
AnchorSync: Global Consistency Optimization for Long Video Editing
von: Liu, Zichi, et al.
Veröffentlicht: (2025)
von: Liu, Zichi, et al.
Veröffentlicht: (2025)
MGIMM: Multi-Granularity Instruction Multimodal Model for Attribute-Guided Remote Sensing Image Detailed Description
von: Yang, Cong, et al.
Veröffentlicht: (2024)
von: Yang, Cong, et al.
Veröffentlicht: (2024)
Semantic Flow: Learning Semantic Field of Dynamic Scenes from Monocular Videos
von: Tian, Fengrui, et al.
Veröffentlicht: (2024)
von: Tian, Fengrui, et al.
Veröffentlicht: (2024)
DreamVE: Unified Instruction-based Image and Video Editing
von: Xia, Bin, et al.
Veröffentlicht: (2025)
von: Xia, Bin, et al.
Veröffentlicht: (2025)
Systematic Reward Gap Optimization for Mitigating VLM Hallucinations
von: He, Lehan, et al.
Veröffentlicht: (2024)
von: He, Lehan, et al.
Veröffentlicht: (2024)
Rationale Matters: Learning Transferable Rubrics via Proxy-Guided Critique for VLM Reward Models
von: Qiu, Weijie, et al.
Veröffentlicht: (2026)
von: Qiu, Weijie, et al.
Veröffentlicht: (2026)
Text-Driven Diverse Facial Texture Generation via Progressive Latent-Space Refinement
von: Wang, Chi, et al.
Veröffentlicht: (2024)
von: Wang, Chi, et al.
Veröffentlicht: (2024)
Neural Video Fields Editing
von: Yang, Shuzhou, et al.
Veröffentlicht: (2023)
von: Yang, Shuzhou, et al.
Veröffentlicht: (2023)
Chameleon: Benchmarking Detection and Backtracking on Commercial-Grade AI-Generated Videos
von: Liao, Xingming, et al.
Veröffentlicht: (2025)
von: Liao, Xingming, et al.
Veröffentlicht: (2025)
FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models
von: Cai, Kaitong, et al.
Veröffentlicht: (2025)
von: Cai, Kaitong, et al.
Veröffentlicht: (2025)
Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement
von: Liu, Yang, et al.
Veröffentlicht: (2025)
von: Liu, Yang, et al.
Veröffentlicht: (2025)
HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing
von: Bai, Jinbin, et al.
Veröffentlicht: (2024)
von: Bai, Jinbin, et al.
Veröffentlicht: (2024)
UniVideo: Unified Understanding, Generation, and Editing for Videos
von: Wei, Cong, et al.
Veröffentlicht: (2025)
von: Wei, Cong, et al.
Veröffentlicht: (2025)
VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction Editing Data and Long Captions
von: Wang, Ziteng, et al.
Veröffentlicht: (2025)
von: Wang, Ziteng, et al.
Veröffentlicht: (2025)
Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
von: Wang, Xingrui, et al.
Veröffentlicht: (2024)
von: Wang, Xingrui, et al.
Veröffentlicht: (2024)
Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling
von: Bai, Xuehai, et al.
Veröffentlicht: (2026)
von: Bai, Xuehai, et al.
Veröffentlicht: (2026)
ZONE: Zero-Shot Instruction-Guided Local Editing
von: Li, Shanglin, et al.
Veröffentlicht: (2023)
von: Li, Shanglin, et al.
Veröffentlicht: (2023)
Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing
von: Yeh, Chun-Hsiao, et al.
Veröffentlicht: (2025)
von: Yeh, Chun-Hsiao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ATI: Any Trajectory Instruction for Controllable Video Generation
von: Wang, Angtian, et al.
Veröffentlicht: (2025) -
CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance
von: Deng, Yufan, et al.
Veröffentlicht: (2025) -
HECTOR: Hybrid Editable Compositional Object References for Video Generation
von: Zhang, Guofeng, et al.
Veröffentlicht: (2026) -
TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
von: Zhang, Guofeng, et al.
Veröffentlicht: (2025) -
MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement
von: Deng, Yufan, et al.
Veröffentlicht: (2025)