BVINet: Unlocking Blind Video Inpainting with Zero Annotations
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Zhiliang, Chen, Kerui, Li, Kun, Fan, Hehe, Yang, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Prompt-Aware Controllable Shadow Removal
by: Chen, Kerui, et al.
Published: (2025)
by: Chen, Kerui, et al.
Published: (2025)
EVA: Zero-shot Accurate Attributes and Multi-Object Video Editing
by: Yang, Xiangpeng, et al.
Published: (2024)
by: Yang, Xiangpeng, et al.
Published: (2024)
ZeroMamba: Exploring Visual State Space Model for Zero-Shot Learning
by: Hou, Wenjin, et al.
Published: (2024)
by: Hou, Wenjin, et al.
Published: (2024)
MMAD: Multi-label Micro-Action Detection in Videos
by: Li, Kun, et al.
Published: (2024)
by: Li, Kun, et al.
Published: (2024)
Scaling Video Understanding via Compact Latent Multi-Agent Collaboration
by: Chen, Kerui, et al.
Published: (2026)
by: Chen, Kerui, et al.
Published: (2026)
Prototype Learning for Micro-gesture Classification
by: Chen, Guoliang, et al.
Published: (2024)
by: Chen, Guoliang, et al.
Published: (2024)
ClusterStyle: Modeling Intra-Style Diversity with Prototypical Clustering for Stylized Motion Generation
by: Chen, Kerui, et al.
Published: (2025)
by: Chen, Kerui, et al.
Published: (2025)
MA-Bench: Towards Fine-grained Micro-Action Understanding
by: Li, Kun, et al.
Published: (2026)
by: Li, Kun, et al.
Published: (2026)
VideoGrain: Modulating Space-Time Attention for Multi-grained Video Editing
by: Yang, Xiangpeng, et al.
Published: (2025)
by: Yang, Xiangpeng, et al.
Published: (2025)
Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion
by: Zhou, Zhenglin, et al.
Published: (2025)
by: Zhou, Zhenglin, et al.
Published: (2025)
Motion Matters: Motion-guided Modulation Network for Skeleton-based Micro-Action Recognition
by: Gu, Jihao, et al.
Published: (2025)
by: Gu, Jihao, et al.
Published: (2025)
TransNormal: Dense Visual Semantics for Diffusion-based Transparent Object Normal Estimation
by: Li, Mingwei, et al.
Published: (2026)
by: Li, Mingwei, et al.
Published: (2026)
GraphTARIF: Linear Graph Transformer with Augmented Rank and Improved Focus
by: Hu, Zhaolin, et al.
Published: (2025)
by: Hu, Zhaolin, et al.
Published: (2025)
Prototypical Calibrating Ambiguous Samples for Micro-Action Recognition
by: Li, Kun, et al.
Published: (2024)
by: Li, Kun, et al.
Published: (2024)
TV-Dialogue: Crafting Theme-Aware Video Dialogues with Immersive Interaction
by: Wang, Sai, et al.
Published: (2025)
by: Wang, Sai, et al.
Published: (2025)
DeMoGen: Towards Decompositional Human Motion Generation with Energy-Based Diffusion Models
by: Zhang, Jianrong, et al.
Published: (2025)
by: Zhang, Jianrong, et al.
Published: (2025)
EnergyMoGen: Compositional Human Motion Generation with Energy-Based Diffusion Model in Latent Space
by: Zhang, Jianrong, et al.
Published: (2024)
by: Zhang, Jianrong, et al.
Published: (2024)
MTV-Inpaint: Multi-Task Long Video Inpainting
by: Yang, Shiyuan, et al.
Published: (2025)
by: Yang, Shiyuan, et al.
Published: (2025)
Incentivizing Generative Zero-Shot Learning via Outcome-Reward Reinforcement Learning with Visual Cues
by: Hou, Wenjin, et al.
Published: (2026)
by: Hou, Wenjin, et al.
Published: (2026)
VividDreamer: Invariant Score Distillation For Hyper-Realistic Text-to-3D Generation
by: Zhuo, Wenjie, et al.
Published: (2024)
by: Zhuo, Wenjie, et al.
Published: (2024)
Flow-Guided Diffusion for Video Inpainting
by: Gu, Bohai, et al.
Published: (2023)
by: Gu, Bohai, et al.
Published: (2023)
Let's Reward Step-by-Step: Step-Aware Contrastive Alignment for Vision-Language Navigation in Continuous Environments
by: Li, Haoyuan, et al.
Published: (2026)
by: Li, Haoyuan, et al.
Published: (2026)
HeadStudio: Text to Animatable Head Avatars with 3D Gaussian Splatting
by: Zhou, Zhenglin, et al.
Published: (2024)
by: Zhou, Zhenglin, et al.
Published: (2024)
TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
Prompt-Aware Adapter: Towards Learning Adaptive Visual Tokens for Multimodal Large Language Models
by: Zhang, Yue, et al.
Published: (2024)
by: Zhang, Yue, et al.
Published: (2024)
DVAR: Adversarial Multi-Agent Debate for Video Authenticity Detection
by: Qi, Hongyuan, et al.
Published: (2026)
by: Qi, Hongyuan, et al.
Published: (2026)
TSGS: Improving Gaussian Splatting for Transparent Surface Reconstruction via Normal and De-lighting Priors
by: Li, Mingwei, et al.
Published: (2025)
by: Li, Mingwei, et al.
Published: (2025)
Towards Language-Driven Video Inpainting via Multimodal Large Language Models
by: Wu, Jianzong, et al.
Published: (2024)
by: Wu, Jianzong, et al.
Published: (2024)
From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative Bootstrapping
by: He, Xu, et al.
Published: (2025)
by: He, Xu, et al.
Published: (2025)
DiTPainter: Efficient Video Inpainting with Diffusion Transformers
by: Wu, Xian, et al.
Published: (2025)
by: Wu, Xian, et al.
Published: (2025)
PM-VIS+: High-Performance Video Instance Segmentation without Video Annotation
by: Yang, Zhangjing, et al.
Published: (2024)
by: Yang, Zhangjing, et al.
Published: (2024)
Stroke3D: Lifting 2D strokes into rigged 3D model via latent diffusion models
by: Zhao, Ruisi, et al.
Published: (2026)
by: Zhao, Ruisi, et al.
Published: (2026)
InfiniDreamer: Arbitrarily Long Human Motion Generation via Segment Score Distillation
by: Zhuo, Wenjie, et al.
Published: (2024)
by: Zhuo, Wenjie, et al.
Published: (2024)
VIP: Video Inpainting Pipeline for Real World Human Removal
by: Sun, Huiming, et al.
Published: (2025)
by: Sun, Huiming, et al.
Published: (2025)
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts
by: Zhang, Yue, et al.
Published: (2025)
by: Zhang, Yue, et al.
Published: (2025)
Unlocking the Potential of Diffusion Priors in Blind Face Restoration
by: Miao, Yunqi, et al.
Published: (2025)
by: Miao, Yunqi, et al.
Published: (2025)
Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
by: Li, Haoyuan, et al.
Published: (2025)
by: Li, Haoyuan, et al.
Published: (2025)
RemoteZero: Geospatial Reasoning with Zero Human Annotations
by: Yao, Liang, et al.
Published: (2026)
by: Yao, Liang, et al.
Published: (2026)
AVID: Any-Length Video Inpainting with Diffusion Model
by: Zhang, Zhixing, et al.
Published: (2023)
by: Zhang, Zhixing, et al.
Published: (2023)
DiffuEraser: A Diffusion Model for Video Inpainting
by: Li, Xiaowen, et al.
Published: (2025)
by: Li, Xiaowen, et al.
Published: (2025)
Similar Items
-
Prompt-Aware Controllable Shadow Removal
by: Chen, Kerui, et al.
Published: (2025) -
EVA: Zero-shot Accurate Attributes and Multi-Object Video Editing
by: Yang, Xiangpeng, et al.
Published: (2024) -
ZeroMamba: Exploring Visual State Space Model for Zero-Shot Learning
by: Hou, Wenjin, et al.
Published: (2024) -
MMAD: Multi-label Micro-Action Detection in Videos
by: Li, Kun, et al.
Published: (2024) -
Scaling Video Understanding via Compact Latent Multi-Agent Collaboration
by: Chen, Kerui, et al.
Published: (2026)