ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Irene, Lin, Wei, Mirza, M. Jehanzeb, Hansen, Jacob A., Doveh, Sivan, Butoi, Victor Ion, Herzig, Roei, Arbelle, Assaf, Kuehne, Hilde, Darrell, Trevor, Gan, Chuang, Oliva, Aude, Feris, Rogerio, Karlinsky, Leonid |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Teaching VLMs to Localize Specific Objects from In-context Examples
di: Doveh, Sivan, et al.
Pubblicazione: (2024)
di: Doveh, Sivan, et al.
Pubblicazione: (2024)
Navigating the Labyrinth: Evaluating LLMs' Ability to Reason About Search Problems
di: Borazjanizadeh, Nasim, et al.
Pubblicazione: (2024)
di: Borazjanizadeh, Nasim, et al.
Pubblicazione: (2024)
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
di: Huang, Brandon, et al.
Pubblicazione: (2024)
di: Huang, Brandon, et al.
Pubblicazione: (2024)
Visualizing Thought: Conceptual Diagrams Enable Robust Planning in LMMs
di: Borazjanizadeh, Nasim, et al.
Pubblicazione: (2025)
di: Borazjanizadeh, Nasim, et al.
Pubblicazione: (2025)
Latent Implicit Visual Reasoning
di: Li, Kelvin, et al.
Pubblicazione: (2025)
di: Li, Kelvin, et al.
Pubblicazione: (2025)
Comparison Visual Instruction Tuning
di: Lin, Wei, et al.
Pubblicazione: (2024)
di: Lin, Wei, et al.
Pubblicazione: (2024)
Towards Multimodal In-Context Learning for Vision & Language Models
di: Doveh, Sivan, et al.
Pubblicazione: (2024)
di: Doveh, Sivan, et al.
Pubblicazione: (2024)
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
di: Mitra, Chancharik, et al.
Pubblicazione: (2024)
di: Mitra, Chancharik, et al.
Pubblicazione: (2024)
Meta-Prompting for Automating Zero-shot Visual Recognition with LLMs
di: Mirza, M. Jehanzeb, et al.
Pubblicazione: (2024)
di: Mirza, M. Jehanzeb, et al.
Pubblicazione: (2024)
NumeroLogic: Number Encoding for Enhanced LLMs' Numerical Reasoning
di: Schwartz, Eli, et al.
Pubblicazione: (2024)
di: Schwartz, Eli, et al.
Pubblicazione: (2024)
MAEDAY: MAE for few and zero shot AnomalY-Detection
di: Schwartz, Eli, et al.
Pubblicazione: (2022)
di: Schwartz, Eli, et al.
Pubblicazione: (2022)
Activation Reward Models for Few-Shot Model Alignment
di: Chai, Tianning, et al.
Pubblicazione: (2025)
di: Chai, Tianning, et al.
Pubblicazione: (2025)
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
di: Gavrikov, Paul, et al.
Pubblicazione: (2025)
di: Gavrikov, Paul, et al.
Pubblicazione: (2025)
TTRV: Test-Time Reinforcement Learning for Vision Language Models
di: Singh, Akshit, et al.
Pubblicazione: (2025)
di: Singh, Akshit, et al.
Pubblicazione: (2025)
State-Space Large Audio Language Models
di: Bhati, Saurabhchand, et al.
Pubblicazione: (2024)
di: Bhati, Saurabhchand, et al.
Pubblicazione: (2024)
DASS: Distilled Audio State Space Models Are Stronger and More Duration-Scalable Learners
di: Bhati, Saurabhchand, et al.
Pubblicazione: (2024)
di: Bhati, Saurabhchand, et al.
Pubblicazione: (2024)
LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content
di: Shabtay, Nimrod, et al.
Pubblicazione: (2024)
di: Shabtay, Nimrod, et al.
Pubblicazione: (2024)
PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies
di: Selch, Lukas, et al.
Pubblicazione: (2025)
di: Selch, Lukas, et al.
Pubblicazione: (2025)
CALM: Class-Conditional Sparse Attention Vectors for Large Audio-Language Models
di: Mehta, Videet, et al.
Pubblicazione: (2026)
di: Mehta, Videet, et al.
Pubblicazione: (2026)
DAVE: A VLM Vision Encoder for Document Understanding and Web Agents
di: Huang, Brandon, et al.
Pubblicazione: (2025)
di: Huang, Brandon, et al.
Pubblicazione: (2025)
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
di: Rouditchenko, Andrew, et al.
Pubblicazione: (2024)
di: Rouditchenko, Andrew, et al.
Pubblicazione: (2024)
AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers
di: Araujo, Edson, et al.
Pubblicazione: (2026)
di: Araujo, Edson, et al.
Pubblicazione: (2026)
$\textit{Trans-LoRA}$: towards data-free Transferable Parameter Efficient Finetuning
di: Wang, Runqian, et al.
Pubblicazione: (2024)
di: Wang, Runqian, et al.
Pubblicazione: (2024)
Augmenting In-Context-Learning in LLMs via Automatic Data Labeling and Refinement
di: Shtok, Joseph, et al.
Pubblicazione: (2024)
di: Shtok, Joseph, et al.
Pubblicazione: (2024)
GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models
di: Mirza, M. Jehanzeb, et al.
Pubblicazione: (2024)
di: Mirza, M. Jehanzeb, et al.
Pubblicazione: (2024)
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion
di: Hansen, Jacob, et al.
Pubblicazione: (2025)
di: Hansen, Jacob, et al.
Pubblicazione: (2025)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
di: Mitra, Chancharik, et al.
Pubblicazione: (2023)
di: Mitra, Chancharik, et al.
Pubblicazione: (2023)
Exploring Modality Guidance to Enhance VFM-based Feature Fusion for UDA in 3D Semantic Segmentation
di: Spoecklberger, Johannes, et al.
Pubblicazione: (2025)
di: Spoecklberger, Johannes, et al.
Pubblicazione: (2025)
Towards Audio Token Compression in Large Audio Language Models
di: Bhati, Saurabhchand, et al.
Pubblicazione: (2025)
di: Bhati, Saurabhchand, et al.
Pubblicazione: (2025)
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition
di: Rouditchenko, Andrew, et al.
Pubblicazione: (2025)
di: Rouditchenko, Andrew, et al.
Pubblicazione: (2025)
ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation
di: Kondic, Jovana, et al.
Pubblicazione: (2025)
di: Kondic, Jovana, et al.
Pubblicazione: (2025)
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
di: Araujo, Edson, et al.
Pubblicazione: (2025)
di: Araujo, Edson, et al.
Pubblicazione: (2025)
Overflow Prevention Enhances Long-Context Recurrent LLMs
di: Ben-Kish, Assaf, et al.
Pubblicazione: (2025)
di: Ben-Kish, Assaf, et al.
Pubblicazione: (2025)
TTA-Vid: Generalized Test-Time Adaptation for Video Reasoning
di: Jahagirdar, Soumya Shamarao, et al.
Pubblicazione: (2026)
di: Jahagirdar, Soumya Shamarao, et al.
Pubblicazione: (2026)
TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering
di: Shang, Chuyi, et al.
Pubblicazione: (2024)
di: Shang, Chuyi, et al.
Pubblicazione: (2024)
Recursive Visual Programming
di: Ge, Jiaxin, et al.
Pubblicazione: (2023)
di: Ge, Jiaxin, et al.
Pubblicazione: (2023)
$\texttt{BATCLIP}$: Bimodal Online Test-Time Adaptation for CLIP
di: Maharana, Sarthak Kumar, et al.
Pubblicazione: (2024)
di: Maharana, Sarthak Kumar, et al.
Pubblicazione: (2024)
In-Context Learning Enables Robot Action Prediction in LLMs
di: Yin, Yida, et al.
Pubblicazione: (2024)
di: Yin, Yida, et al.
Pubblicazione: (2024)
Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?
di: Rouditchenko, Andrew, et al.
Pubblicazione: (2025)
di: Rouditchenko, Andrew, et al.
Pubblicazione: (2025)
CAMELoT: Towards Large Language Models with Training-Free Consolidated Associative Memory
di: He, Zexue, et al.
Pubblicazione: (2024)
di: He, Zexue, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Teaching VLMs to Localize Specific Objects from In-context Examples
di: Doveh, Sivan, et al.
Pubblicazione: (2024) -
Navigating the Labyrinth: Evaluating LLMs' Ability to Reason About Search Problems
di: Borazjanizadeh, Nasim, et al.
Pubblicazione: (2024) -
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
di: Huang, Brandon, et al.
Pubblicazione: (2024) -
Visualizing Thought: Conceptual Diagrams Enable Robust Planning in LMMs
di: Borazjanizadeh, Nasim, et al.
Pubblicazione: (2025) -
Latent Implicit Visual Reasoning
di: Li, Kelvin, et al.
Pubblicazione: (2025)