Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Sanghwan, Xiao, Rui, Alaniz, Stephan, Xian, Yongqin, Akata, Zeynep |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FINER: MLLMs Hallucinate under Fine-grained Negative Queries
di: Xiao, Rui, et al.
Pubblicazione: (2026)
di: Xiao, Rui, et al.
Pubblicazione: (2026)
FLAIR: VLM with Fine-grained Language-informed Image Representations
di: Xiao, Rui, et al.
Pubblicazione: (2024)
di: Xiao, Rui, et al.
Pubblicazione: (2024)
LoFT: LoRA-fused Training Dataset Generation with Few-shot Guidance
di: Kim, Jae Myung, et al.
Pubblicazione: (2025)
di: Kim, Jae Myung, et al.
Pubblicazione: (2025)
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
di: Kim, Sanghwan, et al.
Pubblicazione: (2024)
di: Kim, Sanghwan, et al.
Pubblicazione: (2024)
From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs
di: Wu, Boyong, et al.
Pubblicazione: (2026)
di: Wu, Boyong, et al.
Pubblicazione: (2026)
A Large Scale Analysis of Gender Biases in Text-to-Image Generative Models
di: Girrbach, Leander, et al.
Pubblicazione: (2025)
di: Girrbach, Leander, et al.
Pubblicazione: (2025)
Explaining CLIP Zero-shot Predictions Through Concepts
di: Ozdemir, Onat, et al.
Pubblicazione: (2026)
di: Ozdemir, Onat, et al.
Pubblicazione: (2026)
SUB: Benchmarking CBM Generalization via Synthetic Attribute Substitutions
di: Bader, Jessica, et al.
Pubblicazione: (2025)
di: Bader, Jessica, et al.
Pubblicazione: (2025)
DataDream: Few-shot Guided Dataset Generation
di: Kim, Jae Myung, et al.
Pubblicazione: (2024)
di: Kim, Jae Myung, et al.
Pubblicazione: (2024)
Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
di: Girrbach, Leander, et al.
Pubblicazione: (2025)
di: Girrbach, Leander, et al.
Pubblicazione: (2025)
PALM: Predicting Actions through Language Models
di: Kim, Sanghwan, et al.
Pubblicazione: (2023)
di: Kim, Sanghwan, et al.
Pubblicazione: (2023)
Concept-Guided Interpretability via Neural Chunking
di: Wu, Shuchen, et al.
Pubblicazione: (2025)
di: Wu, Shuchen, et al.
Pubblicazione: (2025)
Audio-Visual Generalized Zero-Shot Learning using Pre-Trained Large Multi-Modal Models
di: Kurzendörfer, David, et al.
Pubblicazione: (2024)
di: Kurzendörfer, David, et al.
Pubblicazione: (2024)
Vision-by-Language for Training-Free Compositional Image Retrieval
di: Karthik, Shyamgopal, et al.
Pubblicazione: (2023)
di: Karthik, Shyamgopal, et al.
Pubblicazione: (2023)
Sparse Autoencoders are Topic Models
di: Girrbach, Leander, et al.
Pubblicazione: (2025)
di: Girrbach, Leander, et al.
Pubblicazione: (2025)
MIRAGE: Benchmarking and Aligning Multi-Instance Image Editing
di: Liu, Ziqian, et al.
Pubblicazione: (2026)
di: Liu, Ziqian, et al.
Pubblicazione: (2026)
Visual Jigsaw Post-Training Improves MLLMs
di: Wu, Penghao, et al.
Pubblicazione: (2025)
di: Wu, Penghao, et al.
Pubblicazione: (2025)
Reflecting on the State of Rehearsal-free Continual Learning with Pretrained Models
di: Thede, Lukas, et al.
Pubblicazione: (2024)
di: Thede, Lukas, et al.
Pubblicazione: (2024)
DeLoRA: Decoupling Angles and Strength in Low-rank Adaptation
di: Bini, Massimo, et al.
Pubblicazione: (2025)
di: Bini, Massimo, et al.
Pubblicazione: (2025)
Distilling ODE Solvers of Diffusion Models into Smaller Steps
di: Kim, Sanghwan, et al.
Pubblicazione: (2023)
di: Kim, Sanghwan, et al.
Pubblicazione: (2023)
Pick-and-Draw: Training-free Semantic Guidance for Text-to-Image Personalization
di: Lv, Henglei, et al.
Pubblicazione: (2024)
di: Lv, Henglei, et al.
Pubblicazione: (2024)
EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval
di: Hummel, Thomas, et al.
Pubblicazione: (2024)
di: Hummel, Thomas, et al.
Pubblicazione: (2024)
[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs
di: Wang, Ao, et al.
Pubblicazione: (2024)
di: Wang, Ao, et al.
Pubblicazione: (2024)
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
di: Xiao, Kelaiti, et al.
Pubblicazione: (2025)
di: Xiao, Kelaiti, et al.
Pubblicazione: (2025)
ReNO: Enhancing One-step Text-to-Image Models through Reward-based Noise Optimization
di: Eyring, Luca, et al.
Pubblicazione: (2024)
di: Eyring, Luca, et al.
Pubblicazione: (2024)
Rethinking Concept Bottleneck Models: From Pitfalls to Solutions
di: Tapli, Merve, et al.
Pubblicazione: (2026)
di: Tapli, Merve, et al.
Pubblicazione: (2026)
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
di: Zhang, Jiarui, et al.
Pubblicazione: (2025)
di: Zhang, Jiarui, et al.
Pubblicazione: (2025)
Toward a Diffusion-Based Generalist for Dense Vision Tasks
di: Fan, Yue, et al.
Pubblicazione: (2024)
di: Fan, Yue, et al.
Pubblicazione: (2024)
DiffuseHigh: Training-free Progressive High-Resolution Image Synthesis through Structure Guidance
di: Kim, Younghyun, et al.
Pubblicazione: (2024)
di: Kim, Younghyun, et al.
Pubblicazione: (2024)
Stitch: Training-Free Position Control in Multimodal Diffusion Transformers
di: Bader, Jessica, et al.
Pubblicazione: (2025)
di: Bader, Jessica, et al.
Pubblicazione: (2025)
Improving Intervention Efficacy via Concept Realignment in Concept Bottleneck Models
di: Singhi, Nishad, et al.
Pubblicazione: (2024)
di: Singhi, Nishad, et al.
Pubblicazione: (2024)
Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study
di: Huang, Yiran, et al.
Pubblicazione: (2025)
di: Huang, Yiran, et al.
Pubblicazione: (2025)
Growing Visual Generative Capacity for Pre-Trained MLLMs
di: Wang, Hanyu, et al.
Pubblicazione: (2025)
di: Wang, Hanyu, et al.
Pubblicazione: (2025)
ETHER: Efficient Finetuning of Large-Scale Models with Hyperplane Reflections
di: Bini, Massimo, et al.
Pubblicazione: (2024)
di: Bini, Massimo, et al.
Pubblicazione: (2024)
Task-Adaptive Saliency Guidance for Exemplar-free Class Incremental Learning
di: Liu, Xialei, et al.
Pubblicazione: (2022)
di: Liu, Xialei, et al.
Pubblicazione: (2022)
Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding
di: Kim, Jiwan, et al.
Pubblicazione: (2026)
di: Kim, Jiwan, et al.
Pubblicazione: (2026)
The Manifold Hypothesis for Gradient-Based Explanations
di: Bordt, Sebastian, et al.
Pubblicazione: (2022)
di: Bordt, Sebastian, et al.
Pubblicazione: (2022)
SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model
di: Xu, Dongli, et al.
Pubblicazione: (2025)
di: Xu, Dongli, et al.
Pubblicazione: (2025)
Scalable Ranked Preference Optimization for Text-to-Image Generation
di: Karthik, Shyamgopal, et al.
Pubblicazione: (2024)
di: Karthik, Shyamgopal, et al.
Pubblicazione: (2024)
Training-Free Reasoning and Reflection in MLLMs
di: Wei, Hongchen, et al.
Pubblicazione: (2025)
di: Wei, Hongchen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
FINER: MLLMs Hallucinate under Fine-grained Negative Queries
di: Xiao, Rui, et al.
Pubblicazione: (2026) -
FLAIR: VLM with Fine-grained Language-informed Image Representations
di: Xiao, Rui, et al.
Pubblicazione: (2024) -
LoFT: LoRA-fused Training Dataset Generation with Few-shot Guidance
di: Kim, Jae Myung, et al.
Pubblicazione: (2025) -
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
di: Kim, Sanghwan, et al.
Pubblicazione: (2024) -
From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs
di: Wu, Boyong, et al.
Pubblicazione: (2026)