Hyperphantasia: A Benchmark for Evaluating the Mental Visualization Capabilities of Multimodal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Sepehri, Mohammad Shahab, Tinaz, Berk, Fabian, Zalan, Soltanolkotabi, Mahdi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ConceptMix++: Leveling the Playing Field in Text-to-Image Benchmarking via Iterative Prompt Optimization
by: Gan, Haosheng, et al.
Published: (2025)
by: Gan, Haosheng, et al.
Published: (2025)
MediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation models
by: Sepehri, Mohammad Shahab, et al.
Published: (2024)
by: Sepehri, Mohammad Shahab, et al.
Published: (2024)
MosaicMRI: A Diverse Dataset and Benchmark for Raw Musculoskeletal MRI
by: Arguello, Paula, et al.
Published: (2026)
by: Arguello, Paula, et al.
Published: (2026)
Emergence and Evolution of Interpretable Concepts in Diffusion Models
by: Tinaz, Berk, et al.
Published: (2025)
by: Tinaz, Berk, et al.
Published: (2025)
DiracDiffusion: Denoising and Incremental Reconstruction with Assured Data-Consistency
by: Fabian, Zalan, et al.
Published: (2023)
by: Fabian, Zalan, et al.
Published: (2023)
Serpent: Scalable and Efficient Image Restoration via Multi-scale Structured State Space Models
by: Sepehri, Mohammad Shahab, et al.
Published: (2024)
by: Sepehri, Mohammad Shahab, et al.
Published: (2024)
ATHENA: Adaptive Test-Time Steering for Improving Count Fidelity in Diffusion Models
by: Sepehri, Mohammad Shahab, et al.
Published: (2026)
by: Sepehri, Mohammad Shahab, et al.
Published: (2026)
Adapt and Diffuse: Sample-adaptive Reconstruction via Latent Diffusion Models
by: Fabian, Zalan, et al.
Published: (2023)
by: Fabian, Zalan, et al.
Published: (2023)
HARMONY: Hidden Activation Representations and Model Output-Aware Uncertainty Estimation for Vision-Language Models
by: Mushtaq, Erum, et al.
Published: (2025)
by: Mushtaq, Erum, et al.
Published: (2025)
Do You See Me : A Multidimensional Benchmark for Evaluating Visual Perception in Multimodal LLMs
by: Kanade, Aditya, et al.
Published: (2025)
by: Kanade, Aditya, et al.
Published: (2025)
Evaluating Linguistic Capabilities of Multimodal LLMs in the Lens of Few-Shot Learning
by: Dogan, Mustafa, et al.
Published: (2024)
by: Dogan, Mustafa, et al.
Published: (2024)
MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
by: Shi, Yang, et al.
Published: (2025)
by: Shi, Yang, et al.
Published: (2025)
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
Segmentation as A Plug-and-Play Capability for Frozen Multimodal LLMs
by: Liu, Jiazhen, et al.
Published: (2025)
by: Liu, Jiazhen, et al.
Published: (2025)
ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction
by: Guo, Zichun, et al.
Published: (2026)
by: Guo, Zichun, et al.
Published: (2026)
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs
by: Jiang, Tianxiang, et al.
Published: (2025)
by: Jiang, Tianxiang, et al.
Published: (2025)
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs
by: Jian, Ai, et al.
Published: (2025)
by: Jian, Ai, et al.
Published: (2025)
Localizing Before Answering: A Hallucination Evaluation Benchmark for Grounded Medical Multimodal LLMs
by: Nguyen, Dung, et al.
Published: (2025)
by: Nguyen, Dung, et al.
Published: (2025)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
by: Xu, Zelin, et al.
Published: (2026)
by: Xu, Zelin, et al.
Published: (2026)
DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal Models
by: Wang, Jiarui, et al.
Published: (2025)
by: Wang, Jiarui, et al.
Published: (2025)
Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs
by: Sinha, Rohit, et al.
Published: (2026)
by: Sinha, Rohit, et al.
Published: (2026)
Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
by: Wang, Youze, et al.
Published: (2025)
by: Wang, Youze, et al.
Published: (2025)
OCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities
by: Zhang, Peirong, et al.
Published: (2025)
by: Zhang, Peirong, et al.
Published: (2025)
GC-KBVQA: A New Four-Stage Framework for Enhancing Knowledge Based Visual Question Answering Performance
by: Moradi, Mohammad Mahdi, et al.
Published: (2025)
by: Moradi, Mohammad Mahdi, et al.
Published: (2025)
Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
by: Jain, Jitesh, et al.
Published: (2024)
by: Jain, Jitesh, et al.
Published: (2024)
RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
by: Li, Jiaang, et al.
Published: (2025)
by: Li, Jiaang, et al.
Published: (2025)
Customized Visual Storytelling with Unified Multimodal LLMs
by: Li, Wei-Hua, et al.
Published: (2026)
by: Li, Wei-Hua, et al.
Published: (2026)
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
by: Törtei, Brigitta Malagurski, et al.
Published: (2025)
by: Törtei, Brigitta Malagurski, et al.
Published: (2025)
MMRA: A Benchmark for Evaluating Multi-Granularity and Multi-Image Relational Association Capabilities in Large Visual Language Models
by: Wu, Siwei, et al.
Published: (2024)
by: Wu, Siwei, et al.
Published: (2024)
Evaluating Graphical Perception with Multimodal LLMs
by: Nguyen, Rami Huu, et al.
Published: (2025)
by: Nguyen, Rami Huu, et al.
Published: (2025)
VideoAesBench: Benchmarking the Video Aesthetics Perception Capabilities of Large Multimodal Models
by: Li, Yunhao, et al.
Published: (2026)
by: Li, Yunhao, et al.
Published: (2026)
DuwatBench: Bridging Language and Visual Heritage through an Arabic Calligraphy Benchmark for Multimodal Understanding
by: Patle, Shubham, et al.
Published: (2026)
by: Patle, Shubham, et al.
Published: (2026)
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs
by: Zhou, Yikang, et al.
Published: (2025)
by: Zhou, Yikang, et al.
Published: (2025)
WebVR: Benchmarking Multimodal LLMs for WebPage Recreation from Videos via Human-Aligned Visual Rubrics
by: Dai, Yuhong, et al.
Published: (2026)
by: Dai, Yuhong, et al.
Published: (2026)
Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency
by: Wang, Zhikai, et al.
Published: (2025)
by: Wang, Zhikai, et al.
Published: (2025)
CryptoMamba: Leveraging State Space Models for Accurate Bitcoin Price Prediction
by: Sepehri, Mohammad Shahab, et al.
Published: (2025)
by: Sepehri, Mohammad Shahab, et al.
Published: (2025)
GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing
by: Fang, Rongyao, et al.
Published: (2025)
by: Fang, Rongyao, et al.
Published: (2025)
SpineBench: Benchmarking Multimodal LLMs for Spinal Pathology Analysis
by: Zhang, Chenghanyu, et al.
Published: (2025)
by: Zhang, Chenghanyu, et al.
Published: (2025)
HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding
by: Li, Keliang, et al.
Published: (2024)
by: Li, Keliang, et al.
Published: (2024)
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
by: Roberts, Jonathan, et al.
Published: (2025)
by: Roberts, Jonathan, et al.
Published: (2025)
Similar Items
-
ConceptMix++: Leveling the Playing Field in Text-to-Image Benchmarking via Iterative Prompt Optimization
by: Gan, Haosheng, et al.
Published: (2025) -
MediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation models
by: Sepehri, Mohammad Shahab, et al.
Published: (2024) -
MosaicMRI: A Diverse Dataset and Benchmark for Raw Musculoskeletal MRI
by: Arguello, Paula, et al.
Published: (2026) -
Emergence and Evolution of Interpretable Concepts in Diffusion Models
by: Tinaz, Berk, et al.
Published: (2025) -
DiracDiffusion: Denoising and Incremental Reconstruction with Assured Data-Consistency
by: Fabian, Zalan, et al.
Published: (2023)