VASE: Object-Centric Appearance and Shape Manipulation of Real Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Peruzzo, Elia, Goel, Vidit, Xu, Dejia, Xu, Xingqian, Jiang, Yifan, Wang, Zhangyang, Shi, Humphrey, Sebe, Nicu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PAIR-Diffusion: A Comprehensive Multimodal Object-Level Image Editor
by: Goel, Vidit, et al.
Published: (2023)
by: Goel, Vidit, et al.
Published: (2023)
RAGME: Retrieval Augmented Video Generation for Enhanced Motion Realism
by: Peruzzo, Elia, et al.
Published: (2025)
by: Peruzzo, Elia, et al.
Published: (2025)
OpenBias: Open-set Bias Detection in Text-to-Image Generative Models
by: D'Incà, Moreno, et al.
Published: (2024)
by: D'Incà, Moreno, et al.
Published: (2024)
Safe Vision-Language Models via Unsafe Weights Manipulation
by: D'Incà, Moreno, et al.
Published: (2025)
by: D'Incà, Moreno, et al.
Published: (2025)
GradBias: Unveiling Word Influence on Bias in Text-to-Image Generative Models
by: D'Incà, Moreno, et al.
Published: (2024)
by: D'Incà, Moreno, et al.
Published: (2024)
Versatile Diffusion: Text, Images and Variations All in One Diffusion Model
by: Xu, Xingqian, et al.
Published: (2022)
by: Xu, Xingqian, et al.
Published: (2022)
UVMap-ID: A Controllable and Personalized UV Map Generative Model
by: Wang, Weijie, et al.
Published: (2024)
by: Wang, Weijie, et al.
Published: (2024)
Efficient Image Generation with Variadic Attention Heads
by: Walton, Steven, et al.
Published: (2022)
by: Walton, Steven, et al.
Published: (2022)
Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention
by: Xu, Dejia, et al.
Published: (2024)
by: Xu, Dejia, et al.
Published: (2024)
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
by: Li, Jinlong, et al.
Published: (2026)
by: Li, Jinlong, et al.
Published: (2026)
CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation
by: Xu, Dejia, et al.
Published: (2024)
by: Xu, Dejia, et al.
Published: (2024)
4DGen: Grounded 4D Content Generation with Spatial-temporal Consistency
by: Yin, Yuyang, et al.
Published: (2023)
by: Yin, Yuyang, et al.
Published: (2023)
Appearance-Based Refinement for Object-Centric Motion Segmentation
by: Xie, Junyu, et al.
Published: (2023)
by: Xie, Junyu, et al.
Published: (2023)
IMG: Calibrating Diffusion Models via Implicit Multimodal Guidance
by: Guo, Jiayi, et al.
Published: (2025)
by: Guo, Jiayi, et al.
Published: (2025)
CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP
by: Xing, Songlong, et al.
Published: (2025)
by: Xing, Songlong, et al.
Published: (2025)
Transferable-guided Attention Is All You Need for Video Domain Adaptation
by: Sacilotti, André, et al.
Published: (2024)
by: Sacilotti, André, et al.
Published: (2024)
Cues3D: Unleashing the Power of Sole NeRF for Consistent and Unique Instances in Open-Vocabulary 3D Panoptic Segmentation
by: Xue, Feng, et al.
Published: (2025)
by: Xue, Feng, et al.
Published: (2025)
LightGaussian: Unbounded 3D Gaussian Compression with 15x Reduction and 200+ FPS
by: Fan, Zhiwen, et al.
Published: (2023)
by: Fan, Zhiwen, et al.
Published: (2023)
Detecting Text Manipulation in Images using Vision Language Models
by: Vidit, Vidit, et al.
Published: (2025)
by: Vidit, Vidit, et al.
Published: (2025)
Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
Reverse Personalization
by: Kung, Han-Wei, et al.
Published: (2025)
by: Kung, Han-Wei, et al.
Published: (2025)
Asymmetric GANs for Image-to-Image Translation
by: Tang, Hao, et al.
Published: (2019)
by: Tang, Hao, et al.
Published: (2019)
Object-Centric Instruction Augmentation for Robotic Manipulation
by: Wen, Junjie, et al.
Published: (2024)
by: Wen, Junjie, et al.
Published: (2024)
Diffusion4D: Fast Spatial-temporal Consistent 4D Generation via Video Diffusion Models
by: Liang, Hanwen, et al.
Published: (2024)
by: Liang, Hanwen, et al.
Published: (2024)
FSGS: Real-Time Few-shot View Synthesis using Gaussian Splatting
by: Zhu, Zehao, et al.
Published: (2023)
by: Zhu, Zehao, et al.
Published: (2023)
PVAFN: Point-Voxel Attention Fusion Network with Multi-Pooling Enhancing for 3D Object Detection
by: Li, Yidi, et al.
Published: (2024)
by: Li, Yidi, et al.
Published: (2024)
RankFeat&RankWeight: Rank-1 Feature/Weight Removal for Out-of-distribution Detection
by: Song, Yue, et al.
Published: (2023)
by: Song, Yue, et al.
Published: (2023)
FreeInsert: Disentangled Text-Guided Object Insertion in 3D Gaussian Scene without Spatial Priors
by: Li, Chenxi, et al.
Published: (2025)
by: Li, Chenxi, et al.
Published: (2025)
Uni4D: A Unified Self-Supervised Learning Framework for Point Cloud Videos
by: Zuo, Zhi, et al.
Published: (2025)
by: Zuo, Zhi, et al.
Published: (2025)
Vision+X: A Survey on Multimodal Learning in the Light of Data
by: Zhu, Ye, et al.
Published: (2022)
by: Zhu, Ye, et al.
Published: (2022)
Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding
by: Li, Jinlong, et al.
Published: (2025)
by: Li, Jinlong, et al.
Published: (2025)
Hierarchical Cross-Attention Network for Virtual Try-On
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
Rethinking the Learning Paradigm for Facial Expression Recognition
by: Wang, Weijie, et al.
Published: (2022)
by: Wang, Weijie, et al.
Published: (2022)
Multi-focal Conditioned Latent Diffusion for Person Image Synthesis
by: Liu, Jiaqi, et al.
Published: (2025)
by: Liu, Jiaqi, et al.
Published: (2025)
Social Reward: Evaluating and Enhancing Generative AI through Million-User Feedback from an Online Creative Community
by: Isajanyan, Arman, et al.
Published: (2024)
by: Isajanyan, Arman, et al.
Published: (2024)
Zero-Painter: Training-Free Layout Control for Text-to-Image Synthesis
by: Ohanyan, Marianna, et al.
Published: (2024)
by: Ohanyan, Marianna, et al.
Published: (2024)
Towards Physical Understanding in Video Generation: A 3D Point Regularization Approach
by: Chen, Yunuo, et al.
Published: (2025)
by: Chen, Yunuo, et al.
Published: (2025)
AGG: Amortized Generative 3D Gaussians for Single Image to 3D
by: Xu, Dejia, et al.
Published: (2024)
by: Xu, Dejia, et al.
Published: (2024)
ROSA: Reconstructing Object Shape and Appearance Textures by Adaptive Detail Transfer
by: Kaltheuner, Julian, et al.
Published: (2025)
by: Kaltheuner, Julian, et al.
Published: (2025)
LESS: Label-Efficient and Single-Stage Referring 3D Segmentation
by: Liu, Xuexun, et al.
Published: (2024)
by: Liu, Xuexun, et al.
Published: (2024)
Similar Items
-
PAIR-Diffusion: A Comprehensive Multimodal Object-Level Image Editor
by: Goel, Vidit, et al.
Published: (2023) -
RAGME: Retrieval Augmented Video Generation for Enhanced Motion Realism
by: Peruzzo, Elia, et al.
Published: (2025) -
OpenBias: Open-set Bias Detection in Text-to-Image Generative Models
by: D'Incà, Moreno, et al.
Published: (2024) -
Safe Vision-Language Models via Unsafe Weights Manipulation
by: D'Incà, Moreno, et al.
Published: (2025) -
GradBias: Unveiling Word Influence on Bias in Text-to-Image Generative Models
by: D'Incà, Moreno, et al.
Published: (2024)