Uncovering Grounding IDs: How External Cues Shape Multimodal Binding
Fuente:
arXiv
Guardado en:
| Autores principales: | Hasani, Hosein, Izadi, Amirmohammad, Askari, Fatemeh, Bagherian, Mobin, Mohammadian, Sadegh, Izadi, Mohammad, Baghshah, Mahdieh Soleymani |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Understanding Counting Mechanisms in Large Language and Vision-Language Models
por: Hasani, Hosein, et al.
Publicado: (2025)
por: Hasani, Hosein, et al.
Publicado: (2025)
Causal Attribution via Activation Patching
por: Izadi, Amirmohammad, et al.
Publicado: (2026)
por: Izadi, Amirmohammad, et al.
Publicado: (2026)
Mechanistic Interpretability of Large-Scale Counting in LLMs through a System-2 Strategy
por: Hasani, Hosein, et al.
Publicado: (2026)
por: Hasani, Hosein, et al.
Publicado: (2026)
Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
por: Izadi, Amirmohammad, et al.
Publicado: (2025)
por: Izadi, Amirmohammad, et al.
Publicado: (2025)
Improving 3D Few-Shot Segmentation with Inference-Time Pseudo-Labeling
por: Mozafari, Mohammad, et al.
Publicado: (2024)
por: Mozafari, Mohammad, et al.
Publicado: (2024)
ComAlign: Compositional Alignment in Vision-Language Models
por: Abdollah, Ali, et al.
Publicado: (2024)
por: Abdollah, Ali, et al.
Publicado: (2024)
T2I-FineEval: Fine-Grained Compositional Metric for Text-to-Image Evaluation
por: Hosseini, Seyed Mohammad Hadi, et al.
Publicado: (2025)
por: Hosseini, Seyed Mohammad Hadi, et al.
Publicado: (2025)
Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection
por: Zohrabi, Reihaneh, et al.
Publicado: (2025)
por: Zohrabi, Reihaneh, et al.
Publicado: (2025)
HaloProbe: Bayesian Detection and Mitigation of Object Hallucinations in Vision-Language Models
por: Zohrabi, Reihaneh, et al.
Publicado: (2026)
por: Zohrabi, Reihaneh, et al.
Publicado: (2026)
CAREL: Instruction-guided reinforcement learning with cross-modal auxiliary objectives
por: Saghafian, Armin, et al.
Publicado: (2024)
por: Saghafian, Armin, et al.
Publicado: (2024)
Fine-Grained Alignment and Noise Refinement for Compositional Text-to-Image Generation
por: Izadi, Amir Mohammad, et al.
Publicado: (2025)
por: Izadi, Amir Mohammad, et al.
Publicado: (2025)
Deciphering the Role of Representation Disentanglement: Investigating Compositional Generalization in CLIP Models
por: Abbasi, Reza, et al.
Publicado: (2024)
por: Abbasi, Reza, et al.
Publicado: (2024)
Trained Models Tell Us How to Make Them Robust to Spurious Correlation without Group Annotation
por: Ghaznavi, Mahdi, et al.
Publicado: (2024)
por: Ghaznavi, Mahdi, et al.
Publicado: (2024)
LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions
por: Mehri, Faridoun, et al.
Publicado: (2024)
por: Mehri, Faridoun, et al.
Publicado: (2024)
Dilated Balanced Cross Entropy Loss for Medical Image Segmentation
por: Hosseini, Seyed Mohsen, et al.
Publicado: (2024)
por: Hosseini, Seyed Mohsen, et al.
Publicado: (2024)
Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models
por: Rezaei, Parham, et al.
Publicado: (2025)
por: Rezaei, Parham, et al.
Publicado: (2025)
Diffusion Beats Autoregressive: An Evaluation of Compositional Generation in Text-to-Image Models
por: Marioriyad, Arash, et al.
Publicado: (2024)
por: Marioriyad, Arash, et al.
Publicado: (2024)
Limits and Gains of Test-Time Scaling in Vision-Language Reasoning
por: Ahmadpour, Mohammadjavad, et al.
Publicado: (2025)
por: Ahmadpour, Mohammadjavad, et al.
Publicado: (2025)
Language Plays a Pivotal Role in the Object-Attribute Compositional Generalization of CLIP
por: Abbasi, Reza, et al.
Publicado: (2024)
por: Abbasi, Reza, et al.
Publicado: (2024)
Classification of Breast Cancer Histopathology Images using a Modified Supervised Contrastive Learning Method
por: Sani, Matina Mahdizadeh, et al.
Publicado: (2024)
por: Sani, Matina Mahdizadeh, et al.
Publicado: (2024)
Eye-Q: A Multilingual Benchmark for Visual Word Puzzle Solving and Image-to-Phrase Reasoning
por: Najar, Ali, et al.
Publicado: (2026)
por: Najar, Ali, et al.
Publicado: (2026)
Attention Overlap Is Responsible for The Entity Missing Problem in Text-to-image Diffusion Models!
por: Marioriyad, Arash, et al.
Publicado: (2024)
por: Marioriyad, Arash, et al.
Publicado: (2024)
Analyzing CLIP's Performance Limitations in Multi-Object Scenarios: A Controlled High-Resolution Study
por: Abbasi, Reza, et al.
Publicado: (2025)
por: Abbasi, Reza, et al.
Publicado: (2025)
CLIP Under the Microscope: A Fine-Grained Analysis of Multi-Object Representation
por: Abbasi, Reza, et al.
Publicado: (2025)
por: Abbasi, Reza, et al.
Publicado: (2025)
Infinity and Beyond: Compositional Alignment in VAR and Diffusion T2I Models
por: Shahabadi, Hossein, et al.
Publicado: (2025)
por: Shahabadi, Hossein, et al.
Publicado: (2025)
Hidden Meanings in Plain Sight: RebusBench for Evaluating Cognitive Visual Reasoning
por: Kasaei, Seyed Amir, et al.
Publicado: (2026)
por: Kasaei, Seyed Amir, et al.
Publicado: (2026)
GABInsight: Exploring Gender-Activity Binding Bias in Vision-Language Models
por: Abdollahi, Ali, et al.
Publicado: (2024)
por: Abdollahi, Ali, et al.
Publicado: (2024)
No Concept Left Behind: Test-Time Optimization for Compositional Text-to-Image Generation
por: Sameti, Mohammad Hossein, et al.
Publicado: (2025)
por: Sameti, Mohammad Hossein, et al.
Publicado: (2025)
Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation
por: Kasaei, Seyed Amir, et al.
Publicado: (2025)
por: Kasaei, Seyed Amir, et al.
Publicado: (2025)
Decompose-and-Compose: A Compositional Approach to Mitigating Spurious Correlation
por: Noohdani, Fahimeh Hosseini, et al.
Publicado: (2024)
por: Noohdani, Fahimeh Hosseini, et al.
Publicado: (2024)
CARINOX: Inference-time Scaling with Category-Aware Reward-based Initial Noise Optimization and Exploration
por: Kasaei, Seyed Amir, et al.
Publicado: (2025)
por: Kasaei, Seyed Amir, et al.
Publicado: (2025)
HyCoVAD: A Hybrid SSL-LLM Model for Complex Video Anomaly Detection
por: Hemmatyar, Mohammad Mahdi, et al.
Publicado: (2025)
por: Hemmatyar, Mohammad Mahdi, et al.
Publicado: (2025)
RODEO: Robust Outlier Detection via Exposing Adaptive Out-of-Distribution Samples
por: Mirzaei, Hossein, et al.
Publicado: (2025)
por: Mirzaei, Hossein, et al.
Publicado: (2025)
Enhancing Few-Shot Image Classification through Learnable Multi-Scale Embedding and Attention Mechanisms
por: Askari, Fatemeh, et al.
Publicado: (2024)
por: Askari, Fatemeh, et al.
Publicado: (2024)
Grounding Task Assistance with Multimodal Cues from a Single Demonstration
por: Sarch, Gabriel, et al.
Publicado: (2025)
por: Sarch, Gabriel, et al.
Publicado: (2025)
SUSD: Structured Unsupervised Skill Discovery through State Factorization
por: Hosseini, Seyed Mohammad Hadi, et al.
Publicado: (2026)
por: Hosseini, Seyed Mohammad Hadi, et al.
Publicado: (2026)
DL-EWF: Deep Learning Empowering Women's Fashion with Grounded-Segment-Anything Segmentation for Body Shape Classification
por: Asghari, Fatemeh, et al.
Publicado: (2024)
por: Asghari, Fatemeh, et al.
Publicado: (2024)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
por: Marioriyad, Arash, et al.
Publicado: (2025)
por: Marioriyad, Arash, et al.
Publicado: (2025)
Social-MAE: A Transformer-Based Multimodal Autoencoder for Face and Voice
por: Bohy, Hugo, et al.
Publicado: (2025)
por: Bohy, Hugo, et al.
Publicado: (2025)
A Contrastive Teacher-Student Framework for Novelty Detection under Style Shifts
por: Mirzaei, Hossein, et al.
Publicado: (2025)
por: Mirzaei, Hossein, et al.
Publicado: (2025)
Ejemplares similares
-
Understanding Counting Mechanisms in Large Language and Vision-Language Models
por: Hasani, Hosein, et al.
Publicado: (2025) -
Causal Attribution via Activation Patching
por: Izadi, Amirmohammad, et al.
Publicado: (2026) -
Mechanistic Interpretability of Large-Scale Counting in LLMs through a System-2 Strategy
por: Hasani, Hosein, et al.
Publicado: (2026) -
Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
por: Izadi, Amirmohammad, et al.
Publicado: (2025) -
Improving 3D Few-Shot Segmentation with Inference-Time Pseudo-Labeling
por: Mozafari, Mohammad, et al.
Publicado: (2024)