Diffusion Model as a Generalist Segmentation Learner
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Haoxiao, Xiang, Antao, Sun, Haiyang, Sun, Peilin, Pan, Changhao, Chen, Yifu, Hong, Minjie, Wang, Weijie, Chen, Shuang, Chen, Yue, Zhao, Zhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TransDiff: Diffusion-Based Method for Manipulating Transparent Objects Using a Single RGB-D Image
von: Wang, Haoxiao, et al.
Veröffentlicht: (2025)
von: Wang, Haoxiao, et al.
Veröffentlicht: (2025)
Few-shot Learner Parameterization by Diffusion Time-steps
von: Yue, Zhongqi, et al.
Veröffentlicht: (2024)
von: Yue, Zhongqi, et al.
Veröffentlicht: (2024)
Image Generators are Generalist Vision Learners
von: Gabeur, Valentin, et al.
Veröffentlicht: (2026)
von: Gabeur, Valentin, et al.
Veröffentlicht: (2026)
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
von: Zhao, Canyu, et al.
Veröffentlicht: (2025)
von: Zhao, Canyu, et al.
Veröffentlicht: (2025)
Bayesian Diffusion Models for 3D Shape Reconstruction
von: Xu, Haiyang, et al.
Veröffentlicht: (2024)
von: Xu, Haiyang, et al.
Veröffentlicht: (2024)
Masked Diffusion as Self-supervised Representation Learner
von: Pan, Zixuan, et al.
Veröffentlicht: (2023)
von: Pan, Zixuan, et al.
Veröffentlicht: (2023)
Label-Efficient Data Augmentation with Video Diffusion Models for Guidewire Segmentation in Cardiac Fluoroscopy
von: Pan, Shaoyan, et al.
Veröffentlicht: (2024)
von: Pan, Shaoyan, et al.
Veröffentlicht: (2024)
Deep Learning for Inertial Positioning: A Survey
von: Chen, Changhao, et al.
Veröffentlicht: (2023)
von: Chen, Changhao, et al.
Veröffentlicht: (2023)
Matching Query Image Against Selected NeRF Feature for Efficient and Scalable Localization
von: Zhou, Huaiji, et al.
Veröffentlicht: (2024)
von: Zhou, Huaiji, et al.
Veröffentlicht: (2024)
UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing
von: Fu, Tsu-Jui, et al.
Veröffentlicht: (2025)
von: Fu, Tsu-Jui, et al.
Veröffentlicht: (2025)
VolSplat: Rethinking Feed-Forward 3D Gaussian Splatting with Voxel-Aligned Prediction
von: Wang, Weijie, et al.
Veröffentlicht: (2025)
von: Wang, Weijie, et al.
Veröffentlicht: (2025)
DriveGen3D: Boosting Feed-Forward Driving Scene Generation with Efficient Video Diffusion
von: Wang, Weijie, et al.
Veröffentlicht: (2025)
von: Wang, Weijie, et al.
Veröffentlicht: (2025)
DynaNav: Dynamic Feature and Layer Selection for Efficient Visual Navigation
von: Wang, Jiahui, et al.
Veröffentlicht: (2025)
von: Wang, Jiahui, et al.
Veröffentlicht: (2025)
Chimera: Improving Generalist Model with Domain-Specific Experts
von: Peng, Tianshuo, et al.
Veröffentlicht: (2024)
von: Peng, Tianshuo, et al.
Veröffentlicht: (2024)
Few to Big: Prototype Expansion Network via Diffusion Learner for Point Cloud Few-shot Semantic Segmentation
von: Zhao, Qianguang, et al.
Veröffentlicht: (2025)
von: Zhao, Qianguang, et al.
Veröffentlicht: (2025)
UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation
von: Chen, Dengbo, et al.
Veröffentlicht: (2025)
von: Chen, Dengbo, et al.
Veröffentlicht: (2025)
Embody4D: A Generalist 4D World Model for Embodied AI
von: Tu, Peiyan, et al.
Veröffentlicht: (2026)
von: Tu, Peiyan, et al.
Veröffentlicht: (2026)
Prompting Segment Anything Model with Domain-Adaptive Prototype for Generalizable Medical Image Segmentation
von: Wei, Zhikai, et al.
Veröffentlicht: (2024)
von: Wei, Zhikai, et al.
Veröffentlicht: (2024)
Bora: Biomedical Generalist Video Generation Model
von: Sun, Weixiang, et al.
Veröffentlicht: (2024)
von: Sun, Weixiang, et al.
Veröffentlicht: (2024)
Vision Generalist Model: A Survey
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
CBDiff:Conditional Bernoulli Diffusion Models for Image Forgery Localization
von: Lei, Zhou, et al.
Veröffentlicht: (2025)
von: Lei, Zhou, et al.
Veröffentlicht: (2025)
Semantic Localization Guiding Segment Anything Model For Reference Remote Sensing Image Segmentation
von: Li, Shuyang, et al.
Veröffentlicht: (2025)
von: Li, Shuyang, et al.
Veröffentlicht: (2025)
Group Diffusion Transformers are Unsupervised Multitask Learners
von: Huang, Lianghua, et al.
Veröffentlicht: (2024)
von: Huang, Lianghua, et al.
Veröffentlicht: (2024)
Task-Specific Zero-shot Quantization-Aware Training for Object Detection
von: Li, Changhao, et al.
Veröffentlicht: (2025)
von: Li, Changhao, et al.
Veröffentlicht: (2025)
GS: Generative Segmentation via Label Diffusion
von: Chen, Yuhao, et al.
Veröffentlicht: (2025)
von: Chen, Yuhao, et al.
Veröffentlicht: (2025)
PTQ4SAM: Post-Training Quantization for Segment Anything
von: Lv, Chengtao, et al.
Veröffentlicht: (2024)
von: Lv, Chengtao, et al.
Veröffentlicht: (2024)
3D Annotation-Free Learning by Distilling 2D Open-Vocabulary Segmentation Models for Autonomous Driving
von: Sun, Boyi, et al.
Veröffentlicht: (2024)
von: Sun, Boyi, et al.
Veröffentlicht: (2024)
Weakly-Supervised Semantic Segmentation with Image-Level Labels: from Traditional Models to Foundation Models
von: Chen, Zhaozheng, et al.
Veröffentlicht: (2023)
von: Chen, Zhaozheng, et al.
Veröffentlicht: (2023)
Retrieval-augmented Few-shot Medical Image Segmentation with Foundation Models
von: Zhao, Lin, et al.
Veröffentlicht: (2024)
von: Zhao, Lin, et al.
Veröffentlicht: (2024)
When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models
von: Sun, Zhengyang, et al.
Veröffentlicht: (2026)
von: Sun, Zhengyang, et al.
Veröffentlicht: (2026)
Generalizing Visual Geometry Priors to Sparse Gaussian Occupancy Prediction
von: Zhou, Changqing, et al.
Veröffentlicht: (2026)
von: Zhou, Changqing, et al.
Veröffentlicht: (2026)
AdaFSNet: Time Series Classification Based on Convolutional Network with a Adaptive and Effective Kernel Size Configuration
von: Wang, Haoxiao, et al.
Veröffentlicht: (2024)
von: Wang, Haoxiao, et al.
Veröffentlicht: (2024)
Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse
von: Zhang, Kuan, et al.
Veröffentlicht: (2026)
von: Zhang, Kuan, et al.
Veröffentlicht: (2026)
Mirage: One-Step Video Diffusion for Photorealistic and Coherent Asset Editing in Driving Scenes
von: Wang, Shuyun, et al.
Veröffentlicht: (2025)
von: Wang, Shuyun, et al.
Veröffentlicht: (2025)
StealthDiffusion: Towards Evading Diffusion Forensic Detection through Diffusion Model
von: Zhou, Ziyin, et al.
Veröffentlicht: (2024)
von: Zhou, Ziyin, et al.
Veröffentlicht: (2024)
GiT: Towards Generalist Vision Transformer through Universal Language Interface
von: Wang, Haiyang, et al.
Veröffentlicht: (2024)
von: Wang, Haiyang, et al.
Veröffentlicht: (2024)
Exploring Low-Dimensional Subspaces in Diffusion Models for Controllable Image Editing
von: Chen, Siyi, et al.
Veröffentlicht: (2024)
von: Chen, Siyi, et al.
Veröffentlicht: (2024)
Nonlinear Bipolar Compensation: Handling Outliers in Post-Training Quantization
von: Sun, Peilin, et al.
Veröffentlicht: (2026)
von: Sun, Peilin, et al.
Veröffentlicht: (2026)
Efficient Image Restoration through Low-Rank Adaptation and Stable Diffusion XL
von: Zhao, Haiyang
Veröffentlicht: (2024)
von: Zhao, Haiyang
Veröffentlicht: (2024)
Mora: Enabling Generalist Video Generation via A Multi-Agent Framework
von: Yuan, Zhengqing, et al.
Veröffentlicht: (2024)
von: Yuan, Zhengqing, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
TransDiff: Diffusion-Based Method for Manipulating Transparent Objects Using a Single RGB-D Image
von: Wang, Haoxiao, et al.
Veröffentlicht: (2025) -
Few-shot Learner Parameterization by Diffusion Time-steps
von: Yue, Zhongqi, et al.
Veröffentlicht: (2024) -
Image Generators are Generalist Vision Learners
von: Gabeur, Valentin, et al.
Veröffentlicht: (2026) -
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
von: Zhao, Canyu, et al.
Veröffentlicht: (2025) -
Bayesian Diffusion Models for 3D Shape Reconstruction
von: Xu, Haiyang, et al.
Veröffentlicht: (2024)