The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Anvekar, Tejas, Bardoliya, Fenil, Turaga, Pavan K., Baral, Chitta, Gupta, Vivek |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Map&Make: Schema Guided Text to Table Generation
by: Ahuja, Naman, et al.
Published: (2025)
by: Ahuja, Naman, et al.
Published: (2025)
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
by: Yilmaz, Nilay, et al.
Published: (2025)
by: Yilmaz, Nilay, et al.
Published: (2025)
On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation
by: Chatterjee, Agneet, et al.
Published: (2024)
by: Chatterjee, Agneet, et al.
Published: (2024)
Mahalanobis k-NN: A Statistical Lens for Robust Point-Cloud Registrations
by: Anvekar, Tejas, et al.
Published: (2024)
by: Anvekar, Tejas, et al.
Published: (2024)
GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
by: Siingh, Shikhhar, et al.
Published: (2025)
by: Siingh, Shikhhar, et al.
Published: (2025)
MMTABREAL: Real-World Benchmark for Multimodal Table Understanding
by: Titiya, Prasham, et al.
Published: (2025)
by: Titiya, Prasham, et al.
Published: (2025)
Grounding Stylistic Domain Generalization with Quantitative Domain Shift Measures and Synthetic Scene Images
by: Luo, Yiran, et al.
Published: (2024)
by: Luo, Yiran, et al.
Published: (2024)
REVISION: Rendering Tools Enable Spatial Fidelity in Vision-Language Models
by: Chatterjee, Agneet, et al.
Published: (2024)
by: Chatterjee, Agneet, et al.
Published: (2024)
ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models
by: Patel, Maitreya, et al.
Published: (2023)
by: Patel, Maitreya, et al.
Published: (2023)
Rethinking Information Synthesis in Multimodal Question Answering A Multi-Agent Perspective
by: Rajput, Krishna Singh, et al.
Published: (2025)
by: Rajput, Krishna Singh, et al.
Published: (2025)
Automatic Temporal Segmentation for Post-Stroke Rehabilitation: A Keypoint Detection and Temporal Segmentation Approach for Small Datasets
by: Lee, Jisoo, et al.
Published: (2025)
by: Lee, Jisoo, et al.
Published: (2025)
ActionCOMET: A Zero-shot Approach to Learn Image-specific Commonsense Concepts about Actions
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
A Benchmark Grocery Dataset of Realworld Point Clouds From Single View
by: Sheshappanavar, Shivanand Venkanna, et al.
Published: (2024)
by: Sheshappanavar, Shivanand Venkanna, et al.
Published: (2024)
Intra-class Patch Swap for Self-Distillation
by: Choi, Hongjun, et al.
Published: (2025)
by: Choi, Hongjun, et al.
Published: (2025)
Leveraging Topological Guidance for Improved Knowledge Distillation
by: Jeon, Eun Som, et al.
Published: (2024)
by: Jeon, Eun Som, et al.
Published: (2024)
Help Me Identify: Is an LLM+VQA System All We Need to Identify Visual Concepts?
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems
by: Chen, Shuhang, et al.
Published: (2025)
by: Chen, Shuhang, et al.
Published: (2025)
AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models
by: Malaviya, Vatsal, et al.
Published: (2025)
by: Malaviya, Vatsal, et al.
Published: (2025)
RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions
by: Pathiraja, Bimsara, et al.
Published: (2025)
by: Pathiraja, Bimsara, et al.
Published: (2025)
Do MLLMs Exhibit Human-like Perceptual Behaviors? HVSBench: A Benchmark for MLLM Alignment with Human Perceptual Behavior
by: Lin, Jiaying, et al.
Published: (2024)
by: Lin, Jiaying, et al.
Published: (2024)
ViTaB-A: Evaluating Multimodal Large Language Models on Visual Table Attribution
by: Alqurnawi, Yahia, et al.
Published: (2026)
by: Alqurnawi, Yahia, et al.
Published: (2026)
Harnessing Synthetic Preference Data for Enhancing Temporal Understanding of Video-LLMs
by: Vani, Sameep, et al.
Published: (2025)
by: Vani, Sameep, et al.
Published: (2025)
Ground Reaction Force Estimation via Time-aware Knowledge Distillation
by: Jeon, Eun Som, et al.
Published: (2025)
by: Jeon, Eun Som, et al.
Published: (2025)
Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs
by: Wu, Jingze, et al.
Published: (2026)
by: Wu, Jingze, et al.
Published: (2026)
CMAG: Concept-Scaffolded Retrieval for Marketplace Avatar Generation
by: Goel, Rajeev, et al.
Published: (2026)
by: Goel, Rajeev, et al.
Published: (2026)
Getting it Right: Improving Spatial Consistency in Text-to-Image Models
by: Chatterjee, Agneet, et al.
Published: (2024)
by: Chatterjee, Agneet, et al.
Published: (2024)
GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation
by: Li, Rang, et al.
Published: (2025)
by: Li, Rang, et al.
Published: (2025)
Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions
by: Parmar, Mihir, et al.
Published: (2022)
by: Parmar, Mihir, et al.
Published: (2022)
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs
by: Wu, Yixuan, et al.
Published: (2025)
by: Wu, Yixuan, et al.
Published: (2025)
CodePercept: Code-Grounded Visual STEM Perception for MLLMs
by: Guan, Tongkun, et al.
Published: (2026)
by: Guan, Tongkun, et al.
Published: (2026)
Dual Caption Preference Optimization for Diffusion Models
by: Saeidi, Amir, et al.
Published: (2025)
by: Saeidi, Amir, et al.
Published: (2025)
Chimera: Compositional Image Generation using Part-based Concepting
by: Singh, Shivam, et al.
Published: (2025)
by: Singh, Shivam, et al.
Published: (2025)
Guiding Diffusion with Deep Geometric Moments: Balancing Fidelity and Variation
by: Jung, Sangmin, et al.
Published: (2025)
by: Jung, Sangmin, et al.
Published: (2025)
DecompDreamer: A Composition-Aware Curriculum for Structured 3D Asset Generation
by: Nath, Utkarsh, et al.
Published: (2025)
by: Nath, Utkarsh, et al.
Published: (2025)
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
by: Fallah, Forouzan, et al.
Published: (2025)
by: Fallah, Forouzan, et al.
Published: (2025)
Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning
by: Zhang, Bob, et al.
Published: (2025)
by: Zhang, Bob, et al.
Published: (2025)
RobustMerge: Parameter-Efficient Model Merging for MLLMs with Direction Robustness
by: Zeng, Fanhu, et al.
Published: (2025)
by: Zeng, Fanhu, et al.
Published: (2025)
Deep Geometric Moments Promote Shape Consistency in Text-to-3D Generation
by: Nath, Utkarsh, et al.
Published: (2024)
by: Nath, Utkarsh, et al.
Published: (2024)
Similar Items
-
Map&Make: Schema Guided Text to Table Generation
by: Ahuja, Naman, et al.
Published: (2025) -
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
by: Yilmaz, Nilay, et al.
Published: (2025) -
On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation
by: Chatterjee, Agneet, et al.
Published: (2024) -
Mahalanobis k-NN: A Statistical Lens for Robust Point-Cloud Registrations
by: Anvekar, Tejas, et al.
Published: (2024) -
GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
by: Siingh, Shikhhar, et al.
Published: (2025)