Zero-Shot Video Semantic Segmentation based on Pre-Trained Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Qian, Eldesokey, Abdelrahman, Mendiratta, Mohit, Zhan, Fangneng, Kortylewski, Adam, Theobalt, Christian, Wonka, Peter |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PartEdit: Fine-Grained Image Editing using Pre-Trained Diffusion Models
by: Cvejic, Aleksandar, et al.
Published: (2025)
by: Cvejic, Aleksandar, et al.
Published: (2025)
FaceGPT: Self-supervised Learning to Chat about 3D Human Faces
by: Wang, Haoran, et al.
Published: (2024)
by: Wang, Haoran, et al.
Published: (2024)
LatentMan: Generating Consistent Animated Characters using Image Diffusion Models
by: Eldesokey, Abdelrahman, et al.
Published: (2023)
by: Eldesokey, Abdelrahman, et al.
Published: (2023)
Build-A-Scene: Interactive 3D Layout Control for Diffusion-Based Image Generation
by: Eldesokey, Abdelrahman, et al.
Published: (2024)
by: Eldesokey, Abdelrahman, et al.
Published: (2024)
General Neural Gauge Fields
by: Zhan, Fangneng, et al.
Published: (2023)
by: Zhan, Fangneng, et al.
Published: (2023)
TEDRA: Text-based Editing of Dynamic and Photoreal Actors
by: Sunagad, Basavaraj, et al.
Published: (2024)
by: Sunagad, Basavaraj, et al.
Published: (2024)
DatasetNeRF: Efficient 3D-aware Data Factory with Generative Radiance Fields
by: Chi, Yu, et al.
Published: (2023)
by: Chi, Yu, et al.
Published: (2023)
GRMM: Real-Time High-Fidelity Gaussian Morphable Head Model with Learned Residuals
by: Mendiratta, Mohit, et al.
Published: (2025)
by: Mendiratta, Mohit, et al.
Published: (2025)
EditCLIP: Representation Learning for Image Editing
by: Wang, Qian, et al.
Published: (2025)
by: Wang, Qian, et al.
Published: (2025)
ZeroKey: Point-Level Reasoning and Zero-Shot 3D Keypoint Detection from Large Language Models
by: Gong, Bingchen, et al.
Published: (2024)
by: Gong, Bingchen, et al.
Published: (2024)
DiffAge3D: Diffusion-based 3D-aware Face Aging
by: Wahid, Junaid, et al.
Published: (2024)
by: Wahid, Junaid, et al.
Published: (2024)
Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven Generation
by: Eldesokey, Abdelrahman, et al.
Published: (2025)
by: Eldesokey, Abdelrahman, et al.
Published: (2025)
Evolutive Rendering Models
by: Zhan, Fangneng, et al.
Published: (2024)
by: Zhan, Fangneng, et al.
Published: (2024)
NearID: Identity Representation Learning via Near-identity Distractors
by: Cvejic, Aleksandar, et al.
Published: (2026)
by: Cvejic, Aleksandar, et al.
Published: (2026)
Do It Yourself: Learning Semantic Correspondence from Pseudo-Labels
by: Dünkel, Olaf, et al.
Published: (2025)
by: Dünkel, Olaf, et al.
Published: (2025)
Common3D: Self-Supervised Learning of 3D Morphable Models for Common Objects in Neural Feature Space
by: Sommer, Leonhard, et al.
Published: (2025)
by: Sommer, Leonhard, et al.
Published: (2025)
SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models
by: Dünkel, Olaf, et al.
Published: (2026)
by: Dünkel, Olaf, et al.
Published: (2026)
Audio-Driven Universal Gaussian Head Avatars
by: Teotia, Kartik, et al.
Published: (2025)
by: Teotia, Kartik, et al.
Published: (2025)
DissolveStereo: Coarse Depth Injection for Zero-Shot Stereo Video Generation
by: Shi, Jian, et al.
Published: (2024)
by: Shi, Jian, et al.
Published: (2024)
AvatarMMC: 3D Head Avatar Generation and Editing with Multi-Modal Conditioning
by: Para, Wamiq Reyaz, et al.
Published: (2024)
by: Para, Wamiq Reyaz, et al.
Published: (2024)
ASH: Animatable Gaussian Splats for Efficient and Photoreal Human Rendering
by: Pang, Haokai, et al.
Published: (2023)
by: Pang, Haokai, et al.
Published: (2023)
Relightable Neural Actor with Intrinsic Decomposition and Pose Control
by: Luvizon, Diogo, et al.
Published: (2023)
by: Luvizon, Diogo, et al.
Published: (2023)
PocoLoco: A Point Cloud Diffusion Model of Human Shape in Loose Clothing
by: Seth, Siddharth, et al.
Published: (2024)
by: Seth, Siddharth, et al.
Published: (2024)
Out-of-Distribution Segmentation via Wasserstein-Based Evidential Uncertainty
by: Brosch, Arnold, et al.
Published: (2025)
by: Brosch, Arnold, et al.
Published: (2025)
Zero-Shot Video Restoration and Enhancement Using Pre-Trained Image Diffusion Model
by: Cao, Cong, et al.
Published: (2024)
by: Cao, Cong, et al.
Published: (2024)
StyleGaussian: Instant 3D Style Transfer with Gaussian Splatting
by: Liu, Kunhao, et al.
Published: (2024)
by: Liu, Kunhao, et al.
Published: (2024)
CNS-Bench: Benchmarking Image Classifier Robustness Under Continuous Nuisance Shifts
by: Dünkel, Olaf, et al.
Published: (2025)
by: Dünkel, Olaf, et al.
Published: (2025)
CRONOS: Benchmarking Counterfactual Physical Consistency in Video Models
by: Begiristain, León, et al.
Published: (2026)
by: Begiristain, León, et al.
Published: (2026)
Weakly Supervised 3D Open-vocabulary Segmentation
by: Liu, Kunhao, et al.
Published: (2023)
by: Liu, Kunhao, et al.
Published: (2023)
Zero-Shot Open-Vocabulary Tracking with Large Pre-Trained Models
by: Chu, Wen-Hsuan, et al.
Published: (2023)
by: Chu, Wen-Hsuan, et al.
Published: (2023)
Geometry Matters: 3D Foundation Priors for Learning Semantic Correspondence
by: Jesslen, Artur, et al.
Published: (2026)
by: Jesslen, Artur, et al.
Published: (2026)
Zero-Shot Video Deraining with Video Diffusion Models
by: Varanka, Tuomas, et al.
Published: (2025)
by: Varanka, Tuomas, et al.
Published: (2025)
Studying Image Diffusion Features for Zero-Shot Video Object Segmentation
by: Delatolas, Thanos, et al.
Published: (2025)
by: Delatolas, Thanos, et al.
Published: (2025)
Training-Free Semantic Video Composition via Pre-trained Diffusion Model
by: Guo, Jiaqi, et al.
Published: (2024)
by: Guo, Jiaqi, et al.
Published: (2024)
OmnimatteZero: Fast Training-free Omnimatte with Pre-trained Video Diffusion Models
by: Samuel, Dvir, et al.
Published: (2025)
by: Samuel, Dvir, et al.
Published: (2025)
LaGeM: A Large Geometry Model for 3D Representation Learning and Diffusion
by: Zhang, Biao, et al.
Published: (2024)
by: Zhang, Biao, et al.
Published: (2024)
LASPA: Latent Spatial Alignment for Fast Training-free Single Image Editing
by: Alharbi, Yazeed, et al.
Published: (2024)
by: Alharbi, Yazeed, et al.
Published: (2024)
Conditional Latent Diffusion Models for Zero-Shot Instance Segmentation
by: Ulmer, Maximilian, et al.
Published: (2025)
by: Ulmer, Maximilian, et al.
Published: (2025)
Zero-Shot Video Restoration and Enhancement with Assistance of Video Diffusion Models
by: Cao, Cong, et al.
Published: (2026)
by: Cao, Cong, et al.
Published: (2026)
PlaceIt3D: Language-Guided Object Placement in Real 3D Scenes
by: Abdelreheem, Ahmed, et al.
Published: (2025)
by: Abdelreheem, Ahmed, et al.
Published: (2025)
Similar Items
-
PartEdit: Fine-Grained Image Editing using Pre-Trained Diffusion Models
by: Cvejic, Aleksandar, et al.
Published: (2025) -
FaceGPT: Self-supervised Learning to Chat about 3D Human Faces
by: Wang, Haoran, et al.
Published: (2024) -
LatentMan: Generating Consistent Animated Characters using Image Diffusion Models
by: Eldesokey, Abdelrahman, et al.
Published: (2023) -
Build-A-Scene: Interactive 3D Layout Control for Diffusion-Based Image Generation
by: Eldesokey, Abdelrahman, et al.
Published: (2024) -
General Neural Gauge Fields
by: Zhan, Fangneng, et al.
Published: (2023)