VOID: Video Object and Interaction Deletion
Fuente:
arXiv
Salvato in:
| Autori principali: | Motamed, Saman, Harvey, William, Klein, Benjamin, Van Gool, Luc, Yuan, Zhuoning, Cheng, Ta-Ying |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility
di: Motamed, Saman, et al.
Pubblicazione: (2025)
di: Motamed, Saman, et al.
Pubblicazione: (2025)
Investigating the Effectiveness of Cross-Attention to Unlock Zero-Shot Editing of Text-to-Video Diffusion Models
di: Motamed, Saman, et al.
Pubblicazione: (2024)
di: Motamed, Saman, et al.
Pubblicazione: (2024)
Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models
di: Motamed, Saman, et al.
Pubblicazione: (2023)
di: Motamed, Saman, et al.
Pubblicazione: (2023)
Learning Generative Interactive Environments By Trained Agent Exploration
di: Kazemi, Naser, et al.
Pubblicazione: (2024)
di: Kazemi, Naser, et al.
Pubblicazione: (2024)
InTraGen: Trajectory-controlled Video Generation for Object Interactions
di: Liu, Zuhao, et al.
Pubblicazione: (2024)
di: Liu, Zuhao, et al.
Pubblicazione: (2024)
Learning Local and Global Temporal Contexts for Video Semantic Segmentation
di: Sun, Guolei, et al.
Pubblicazione: (2022)
di: Sun, Guolei, et al.
Pubblicazione: (2022)
Samba: Synchronized Set-of-Sequences Modeling for Multiple Object Tracking
di: Segu, Mattia, et al.
Pubblicazione: (2024)
di: Segu, Mattia, et al.
Pubblicazione: (2024)
A Unified and Interpretable Emotion Representation and Expression Generation
di: Paskaleva, Reni, et al.
Pubblicazione: (2024)
di: Paskaleva, Reni, et al.
Pubblicazione: (2024)
Language-Guided Instance-Aware Domain-Adaptive Panoptic Segmentation
di: Mansour, Elham Amin, et al.
Pubblicazione: (2024)
di: Mansour, Elham Amin, et al.
Pubblicazione: (2024)
ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives
di: Fu, Yuqian, et al.
Pubblicazione: (2024)
di: Fu, Yuqian, et al.
Pubblicazione: (2024)
Self-Explainable Affordance Learning with Embodied Caption
di: Zhang, Zhipeng, et al.
Pubblicazione: (2024)
di: Zhang, Zhipeng, et al.
Pubblicazione: (2024)
Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025
di: Fu, Yuqian, et al.
Pubblicazione: (2025)
di: Fu, Yuqian, et al.
Pubblicazione: (2025)
EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering
di: Li, Yanjun, et al.
Pubblicazione: (2025)
di: Li, Yanjun, et al.
Pubblicazione: (2025)
Do generative video models understand physical principles?
di: Motamed, Saman, et al.
Pubblicazione: (2025)
di: Motamed, Saman, et al.
Pubblicazione: (2025)
Probabilistic Sampling of Balanced K-Means using Adiabatic Quantum Computing
di: Zaech, Jan-Nico, et al.
Pubblicazione: (2023)
di: Zaech, Jan-Nico, et al.
Pubblicazione: (2023)
PhysDreamer: Physics-Based Interaction with 3D Objects via Video Generation
di: Zhang, Tianyuan, et al.
Pubblicazione: (2024)
di: Zhang, Tianyuan, et al.
Pubblicazione: (2024)
Hand-Object Interaction Pretraining from Videos
di: Singh, Himanshu Gaurav, et al.
Pubblicazione: (2024)
di: Singh, Himanshu Gaurav, et al.
Pubblicazione: (2024)
A Survey of Video Datasets for Grounded Event Understanding
di: Sanders, Kate, et al.
Pubblicazione: (2024)
di: Sanders, Kate, et al.
Pubblicazione: (2024)
VideoGameQA-Bench: Evaluating Vision-Language Models for Video Game Quality Assurance
di: Taesiri, Mohammad Reza, et al.
Pubblicazione: (2025)
di: Taesiri, Mohammad Reza, et al.
Pubblicazione: (2025)
Shallow Features Matter: Hierarchical Memory with Heterogeneous Interaction for Unsupervised Video Object Segmentation
di: Xiangyu, Zheng, et al.
Pubblicazione: (2025)
di: Xiangyu, Zheng, et al.
Pubblicazione: (2025)
Image Conductor: Precision Control for Interactive Video Synthesis
di: Li, Yaowei, et al.
Pubblicazione: (2024)
di: Li, Yaowei, et al.
Pubblicazione: (2024)
Rethinking Video Human-Object Interaction: Set Prediction over Time for Unified Detection and Anticipation
di: Luo, Yuanhao, et al.
Pubblicazione: (2026)
di: Luo, Yuanhao, et al.
Pubblicazione: (2026)
DynaHOI: Benchmarking Hand-Object Interaction for Dynamic Target
di: Hu, BoCheng, et al.
Pubblicazione: (2026)
di: Hu, BoCheng, et al.
Pubblicazione: (2026)
ESC: Erasing Space Concept for Knowledge Deletion
di: Lee, Tae-Young, et al.
Pubblicazione: (2025)
di: Lee, Tae-Young, et al.
Pubblicazione: (2025)
PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos
di: Zhou, Zhiyu, et al.
Pubblicazione: (2026)
di: Zhou, Zhiyu, et al.
Pubblicazione: (2026)
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
di: Goswami, Raktim Gautam, et al.
Pubblicazione: (2025)
di: Goswami, Raktim Gautam, et al.
Pubblicazione: (2025)
MultiDelete for Multimodal Machine Unlearning
di: Cheng, Jiali, et al.
Pubblicazione: (2023)
di: Cheng, Jiali, et al.
Pubblicazione: (2023)
Short-term Object Interaction Anticipation with Disentangled Object Detection @ Ego4D Short Term Object Interaction Anticipation Challenge
di: Cho, Hyunjin, et al.
Pubblicazione: (2024)
di: Cho, Hyunjin, et al.
Pubblicazione: (2024)
InteractMove: Text-Controlled Human-Object Interaction Generation in 3D Scenes with Movable Objects
di: Cai, Xinhao, et al.
Pubblicazione: (2025)
di: Cai, Xinhao, et al.
Pubblicazione: (2025)
GameGen-X: Interactive Open-world Game Video Generation
di: Che, Haoxuan, et al.
Pubblicazione: (2024)
di: Che, Haoxuan, et al.
Pubblicazione: (2024)
Open-World Object Counting in Videos
di: Amini-Naieni, Niki, et al.
Pubblicazione: (2025)
di: Amini-Naieni, Niki, et al.
Pubblicazione: (2025)
Interact-Custom: Customized Human Object Interaction Image Generation
di: Xu, Zhu, et al.
Pubblicazione: (2025)
di: Xu, Zhu, et al.
Pubblicazione: (2025)
Towards Universal Video Retrieval: Generalizing Video Embedding via Synthesized Multimodal Pyramid Curriculum
di: Guo, Zhuoning, et al.
Pubblicazione: (2025)
di: Guo, Zhuoning, et al.
Pubblicazione: (2025)
MADTempo: An Interactive System for Multi-Event Temporal Video Retrieval with Query Augmentation
di: Vu, Huu-An, et al.
Pubblicazione: (2025)
di: Vu, Huu-An, et al.
Pubblicazione: (2025)
EgoGrasp: World-Space Hand-Object Interaction Estimation from Egocentric Videos
di: Fu, Hongming, et al.
Pubblicazione: (2026)
di: Fu, Hongming, et al.
Pubblicazione: (2026)
Interact3D: Compositional 3D Generation of Interactive Objects
di: Shan, Hui, et al.
Pubblicazione: (2026)
di: Shan, Hui, et al.
Pubblicazione: (2026)
Temporal Object-Aware Vision Transformer for Few-Shot Video Object Detection
di: Kumar, Yogesh, et al.
Pubblicazione: (2025)
di: Kumar, Yogesh, et al.
Pubblicazione: (2025)
WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model
di: Li, Zongjian, et al.
Pubblicazione: (2024)
di: Li, Zongjian, et al.
Pubblicazione: (2024)
MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model
di: Tong, Jinguang, et al.
Pubblicazione: (2026)
di: Tong, Jinguang, et al.
Pubblicazione: (2026)
TempGlitch: Evaluating Vision-Language Models for Temporal Glitch Detection in Gameplay Videos
di: Yu, Yakun, et al.
Pubblicazione: (2026)
di: Yu, Yakun, et al.
Pubblicazione: (2026)
Documenti analoghi
-
TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility
di: Motamed, Saman, et al.
Pubblicazione: (2025) -
Investigating the Effectiveness of Cross-Attention to Unlock Zero-Shot Editing of Text-to-Video Diffusion Models
di: Motamed, Saman, et al.
Pubblicazione: (2024) -
Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models
di: Motamed, Saman, et al.
Pubblicazione: (2023) -
Learning Generative Interactive Environments By Trained Agent Exploration
di: Kazemi, Naser, et al.
Pubblicazione: (2024) -
InTraGen: Trajectory-controlled Video Generation for Object Interactions
di: Liu, Zuhao, et al.
Pubblicazione: (2024)