Toward a Diffusion-Based Generalist for Dense Vision Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fan, Yue, Xian, Yongqin, Zhai, Xiaohua, Kolesnikov, Alexander, Naeem, Muhammad Ferjad, Schiele, Bernt, Tombari, Federico |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
von: Kuzucu, Selim, et al.
Veröffentlicht: (2025)
von: Kuzucu, Selim, et al.
Veröffentlicht: (2025)
GiT: Towards Generalist Vision Transformer through Universal Language Interface
von: Wang, Haiyang, et al.
Veröffentlicht: (2024)
von: Wang, Haiyang, et al.
Veröffentlicht: (2024)
PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding
von: Kuzucu, Selim, et al.
Veröffentlicht: (2026)
von: Kuzucu, Selim, et al.
Veröffentlicht: (2026)
RefAM: Attention Magnets for Zero-Shot Referral Segmentation
von: Kukleva, Anna, et al.
Veröffentlicht: (2025)
von: Kukleva, Anna, et al.
Veröffentlicht: (2025)
Learning to Prompt with Text Only Supervision for Vision-Language Models
von: Khattak, Muhammad Uzair, et al.
Veröffentlicht: (2024)
von: Khattak, Muhammad Uzair, et al.
Veröffentlicht: (2024)
Human Pose Descriptions and Subject-Focused Attention for Improved Zero-Shot Transfer in Human-Centric Classification Tasks
von: Khan, Muhammad Saif Ullah, et al.
Veröffentlicht: (2024)
von: Khan, Muhammad Saif Ullah, et al.
Veröffentlicht: (2024)
Active Data Curation Effectively Distills Large-Scale Multimodal Models
von: Udandarao, Vishaal, et al.
Veröffentlicht: (2024)
von: Udandarao, Vishaal, et al.
Veröffentlicht: (2024)
LIME: Localized Image Editing via Attention Regularization in Diffusion Models
von: Simsar, Enis, et al.
Veröffentlicht: (2023)
von: Simsar, Enis, et al.
Veröffentlicht: (2023)
R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs
von: Xie, Jiahao, et al.
Veröffentlicht: (2026)
von: Xie, Jiahao, et al.
Veröffentlicht: (2026)
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models
von: Xie, Jiahao, et al.
Veröffentlicht: (2026)
von: Xie, Jiahao, et al.
Veröffentlicht: (2026)
How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs
von: Khattak, Muhammad Uzair, et al.
Veröffentlicht: (2024)
von: Khattak, Muhammad Uzair, et al.
Veröffentlicht: (2024)
Test-Time Visual In-Context Tuning
von: Xie, Jiahao, et al.
Veröffentlicht: (2025)
von: Xie, Jiahao, et al.
Veröffentlicht: (2025)
UIP2P: Unsupervised Instruction-based Image Editing via Edit Reversibility Constraint
von: Simsar, Enis, et al.
Veröffentlicht: (2024)
von: Simsar, Enis, et al.
Veröffentlicht: (2024)
Text-Conditioned Resampler For Long Form Video Understanding
von: Korbar, Bruno, et al.
Veröffentlicht: (2023)
von: Korbar, Bruno, et al.
Veröffentlicht: (2023)
Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
von: Plizzari, Chiara, et al.
Veröffentlicht: (2025)
von: Plizzari, Chiara, et al.
Veröffentlicht: (2025)
TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters
von: Wang, Haiyang, et al.
Veröffentlicht: (2024)
von: Wang, Haiyang, et al.
Veröffentlicht: (2024)
SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
von: Tschannen, Michael, et al.
Veröffentlicht: (2025)
von: Tschannen, Michael, et al.
Veröffentlicht: (2025)
OrCo: Towards Better Generalization via Orthogonality and Contrast for Few-Shot Class-Incremental Learning
von: Ahmed, Noor, et al.
Veröffentlicht: (2024)
von: Ahmed, Noor, et al.
Veröffentlicht: (2024)
MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
von: Segu, Mattia, et al.
Veröffentlicht: (2025)
von: Segu, Mattia, et al.
Veröffentlicht: (2025)
Towards Better Understanding Attribution Methods
von: Rao, Sukrut, et al.
Veröffentlicht: (2022)
von: Rao, Sukrut, et al.
Veröffentlicht: (2022)
B-cos Alignment for Inherently Interpretable CNNs and Vision Transformers
von: Böhle, Moritz, et al.
Veröffentlicht: (2023)
von: Böhle, Moritz, et al.
Veröffentlicht: (2023)
DWDN: Deep Wiener Deconvolution Network for Non-Blind Image Deblurring
von: Dong, Jiangxin, et al.
Veröffentlicht: (2021)
von: Dong, Jiangxin, et al.
Veröffentlicht: (2021)
VITAL: More Understandable Feature Visualization through Distribution Alignment and Relevant Information Flow
von: Gorgun, Ada, et al.
Veröffentlicht: (2025)
von: Gorgun, Ada, et al.
Veröffentlicht: (2025)
AnyUp: Universal Feature Upsampling
von: Wimmer, Thomas, et al.
Veröffentlicht: (2025)
von: Wimmer, Thomas, et al.
Veröffentlicht: (2025)
Sp2360: Sparse-view 360 Scene Reconstruction using Cascaded 2D Diffusion Priors
von: Paul, Soumava, et al.
Veröffentlicht: (2024)
von: Paul, Soumava, et al.
Veröffentlicht: (2024)
Temporal Concept Dynamics in Diffusion Models via Prompt-Conditioned Interventions
von: Gorgun, Ada, et al.
Veröffentlicht: (2025)
von: Gorgun, Ada, et al.
Veröffentlicht: (2025)
Discover-then-Name: Task-Agnostic Concept Bottlenecks via Automated Concept Discovery
von: Rao, Sukrut, et al.
Veröffentlicht: (2024)
von: Rao, Sukrut, et al.
Veröffentlicht: (2024)
Optimising for Interpretability: Convolutional Dynamic Alignment Networks
von: Böhle, Moritz, et al.
Veröffentlicht: (2021)
von: Böhle, Moritz, et al.
Veröffentlicht: (2021)
ClipTTT: CLIP-Guided Test-Time Training Helps LVLMs See Better
von: Nath, Mriganka, et al.
Veröffentlicht: (2026)
von: Nath, Mriganka, et al.
Veröffentlicht: (2026)
AIM: Amending Inherent Interpretability via Self-Supervised Masking
von: Alshami, Eyad, et al.
Veröffentlicht: (2025)
von: Alshami, Eyad, et al.
Veröffentlicht: (2025)
How to Probe: Simple Yet Effective Techniques for Improving Post-hoc Explanations
von: Gairola, Siddhartha, et al.
Veröffentlicht: (2025)
von: Gairola, Siddhartha, et al.
Veröffentlicht: (2025)
MTR++: Multi-Agent Motion Prediction with Symmetric Scene Modeling and Guided Intention Querying
von: Shi, Shaoshuai, et al.
Veröffentlicht: (2023)
von: Shi, Shaoshuai, et al.
Veröffentlicht: (2023)
MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment
von: Das, Anurag, et al.
Veröffentlicht: (2024)
von: Das, Anurag, et al.
Veröffentlicht: (2024)
OmniFashion: Towards Generalist Fashion Intelligence via Multi-Task Vision-Language Learning
von: Yang, Zhengwei, et al.
Veröffentlicht: (2026)
von: Yang, Zhengwei, et al.
Veröffentlicht: (2026)
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
von: Kim, Sanghwan, et al.
Veröffentlicht: (2025)
von: Kim, Sanghwan, et al.
Veröffentlicht: (2025)
InstructCV: Instruction-Tuned Text-to-Image Diffusion Models as Vision Generalists
von: Gan, Yulu, et al.
Veröffentlicht: (2023)
von: Gan, Yulu, et al.
Veröffentlicht: (2023)
Medical Vision Generalist: Unifying Medical Imaging Tasks in Context
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
Masked AutoDecoder is Effective Multi-Task Vision Generalist
von: Qiu, Han, et al.
Veröffentlicht: (2024)
von: Qiu, Han, et al.
Veröffentlicht: (2024)
SimNP: Learning Self-Similarity Priors Between Neural Points
von: Wewer, Christopher, et al.
Veröffentlicht: (2023)
von: Wewer, Christopher, et al.
Veröffentlicht: (2023)
PersonaHOI: Effortlessly Improving Personalized Face with Human-Object Interaction Generation
von: Hu, Xinting, et al.
Veröffentlicht: (2025)
von: Hu, Xinting, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
von: Kuzucu, Selim, et al.
Veröffentlicht: (2025) -
GiT: Towards Generalist Vision Transformer through Universal Language Interface
von: Wang, Haiyang, et al.
Veröffentlicht: (2024) -
PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding
von: Kuzucu, Selim, et al.
Veröffentlicht: (2026) -
RefAM: Attention Magnets for Zero-Shot Referral Segmentation
von: Kukleva, Anna, et al.
Veröffentlicht: (2025) -
Learning to Prompt with Text Only Supervision for Vision-Language Models
von: Khattak, Muhammad Uzair, et al.
Veröffentlicht: (2024)