MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Shengyu, Ye, Tongrui, Zhang, Jianbo, Zhang, Zicheng, Li, Chunyi, Zhai, Guangtao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MirrorBench: A Benchmark to Evaluate Conversational User-Proxy Agents for Human-Likeness
by: Hathidara, Ashutosh, et al.
Published: (2026)
by: Hathidara, Ashutosh, et al.
Published: (2026)
Affordance Benchmark for MLLMs
by: Wang, Junying, et al.
Published: (2025)
by: Wang, Junying, et al.
Published: (2025)
Redundancy Principles for MLLMs Benchmarks
by: Zhang, Zicheng, et al.
Published: (2025)
by: Zhang, Zicheng, et al.
Published: (2025)
Q-Mirror: Unlocking the Multi-Modal Potential of Scientific Text-Only QA Pairs
by: Wang, Junying, et al.
Published: (2025)
by: Wang, Junying, et al.
Published: (2025)
A-Bench: Are LMMs Masters at Evaluating AI-generated Images?
by: Zhang, Zicheng, et al.
Published: (2024)
by: Zhang, Zicheng, et al.
Published: (2024)
PuzzleBench: A Fully Dynamic Evaluation Framework for Large Multimodal Models on Puzzle Solving
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
GeoX-Bench: Benchmarking Cross-View Geo-Localization and Pose Estimation Capabilities of Large Multimodal Models
by: Zheng, Yushuo, et al.
Published: (2025)
by: Zheng, Yushuo, et al.
Published: (2025)
Quality Assessment in the Era of Large Models: A Survey
by: Zhang, Zicheng, et al.
Published: (2024)
by: Zhang, Zicheng, et al.
Published: (2024)
Market-Bench: Benchmarking Large Language Models on Economic and Trade Competition
by: Zheng, Yushuo, et al.
Published: (2026)
by: Zheng, Yushuo, et al.
Published: (2026)
Explore the Hallucination on Low-level Perception for MLLMs
by: Sun, Yinan, et al.
Published: (2024)
by: Sun, Yinan, et al.
Published: (2024)
StableToolBench-MirrorAPI: Modeling Tool Environments as Mirrors of 7,000+ Real-World APIs
by: Guo, Zhicheng, et al.
Published: (2025)
by: Guo, Zhicheng, et al.
Published: (2025)
A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation
by: Shen, Ye, et al.
Published: (2025)
by: Shen, Ye, et al.
Published: (2025)
Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision
by: Wu, Haoning, et al.
Published: (2023)
by: Wu, Haoning, et al.
Published: (2023)
MirrorShield: Towards Universal Defense Against Jailbreaks via Entropy-Guided Mirror Crafting
by: Pu, Rui, et al.
Published: (2025)
by: Pu, Rui, et al.
Published: (2025)
Embodied Representation Alignment with Mirror Neurons
by: Zhu, Wentao, et al.
Published: (2025)
by: Zhu, Wentao, et al.
Published: (2025)
LM Fight Arena: Benchmarking Large Multimodal Models via Game Competition
by: Zheng, Yushuo, et al.
Published: (2025)
by: Zheng, Yushuo, et al.
Published: (2025)
Self-Emotion-Mediated Exploration in Artificial Intelligence Mirrors: Findings from Cognitive Psychology
by: Assunção, Gustavo, et al.
Published: (2023)
by: Assunção, Gustavo, et al.
Published: (2023)
MISC: Ultra-low Bitrate Image Semantic Compression Driven by Large Multimodal Model
by: Li, Chunyi, et al.
Published: (2024)
by: Li, Chunyi, et al.
Published: (2024)
MirrorMark: Generalizable Mirrored Sampling for Multi-bit LLM Watermarking
by: Jiang, Ya, et al.
Published: (2026)
by: Jiang, Ya, et al.
Published: (2026)
Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All?
by: Ziliotto, Filippo, et al.
Published: (2026)
by: Ziliotto, Filippo, et al.
Published: (2026)
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
by: Roytburg, Dani, et al.
Published: (2025)
by: Roytburg, Dani, et al.
Published: (2025)
The Ever-Evolving Science Exam
by: Wang, Junying, et al.
Published: (2025)
by: Wang, Junying, et al.
Published: (2025)
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
by: Li, Caorui, et al.
Published: (2025)
by: Li, Caorui, et al.
Published: (2025)
Plasticity as the Mirror of Empowerment
by: Abel, David, et al.
Published: (2025)
by: Abel, David, et al.
Published: (2025)
MirrorMamba: Towards Scalable and Robust Mirror Detection in Videos
by: Song, Rui, et al.
Published: (2025)
by: Song, Rui, et al.
Published: (2025)
Mirror Mirror on the Wall, Have I Forgotten it All? A New Framework for Evaluating Machine Unlearning
by: Brimhall, Brennon, et al.
Published: (2025)
by: Brimhall, Brennon, et al.
Published: (2025)
The AI in the Mirror: LLM Self-Recognition in an Iterated Public Goods Game
by: Long, Olivia, et al.
Published: (2025)
by: Long, Olivia, et al.
Published: (2025)
Find Them All: Unveiling MLLMs for Versatile Person Re-identification
by: Li, Jinhao, et al.
Published: (2025)
by: Li, Jinhao, et al.
Published: (2025)
OBI-Bench: Can LMMs Aid in Study of Ancient Script on Oracle Bones?
by: Chen, Zijian, et al.
Published: (2024)
by: Chen, Zijian, et al.
Published: (2024)
Evaluating from Benign to Dynamic Adversarial: A Squid Game for Large Language Models
by: Chen, Zijian, et al.
Published: (2025)
by: Chen, Zijian, et al.
Published: (2025)
KidsArtBench: Multi-Dimensional Children's Art Evaluation with Attribute-Aware MLLMs
by: Ye, Mingrui, et al.
Published: (2025)
by: Ye, Mingrui, et al.
Published: (2025)
MEMO-Bench: A Multiple Benchmark for Text-to-Image and Multimodal Large Language Models on Human Emotion Analysis
by: Zhou, Yingjie, et al.
Published: (2024)
by: Zhou, Yingjie, et al.
Published: (2024)
Policy Mirror Descent with Lookahead
by: Protopapas, Kimon, et al.
Published: (2024)
by: Protopapas, Kimon, et al.
Published: (2024)
RaPA: Enhancing Transferable Targeted Attacks via Random Parameter Pruning
by: Su, Tongrui, et al.
Published: (2025)
by: Su, Tongrui, et al.
Published: (2025)
STAR : Bridging Statistical and Agentic Reasoning for Large Model Performance Prediction
by: Wang, Xiaoxiao, et al.
Published: (2026)
by: Wang, Xiaoxiao, et al.
Published: (2026)
Beyond State-Wise Mirror Descent: Offline Policy Optimization with Parametric Policies
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
Stochastic MeanFlow Policies: One-Step Generative Control with Entropic Mirror Descent
by: Wang, Zeyuan, et al.
Published: (2026)
by: Wang, Zeyuan, et al.
Published: (2026)
Do Large Language Models Mirror Cognitive Language Processing?
by: Ren, Yuqi, et al.
Published: (2024)
by: Ren, Yuqi, et al.
Published: (2024)
LMM-PCQA: Assisting Point Cloud Quality Assessment with LMM
by: Zhang, Zicheng, et al.
Published: (2024)
by: Zhang, Zicheng, et al.
Published: (2024)
VQA$^2$: Visual Question Answering for Video Quality Assessment
by: Jia, Ziheng, et al.
Published: (2024)
by: Jia, Ziheng, et al.
Published: (2024)
Similar Items
-
MirrorBench: A Benchmark to Evaluate Conversational User-Proxy Agents for Human-Likeness
by: Hathidara, Ashutosh, et al.
Published: (2026) -
Affordance Benchmark for MLLMs
by: Wang, Junying, et al.
Published: (2025) -
Redundancy Principles for MLLMs Benchmarks
by: Zhang, Zicheng, et al.
Published: (2025) -
Q-Mirror: Unlocking the Multi-Modal Potential of Scientific Text-Only QA Pairs
by: Wang, Junying, et al.
Published: (2025) -
A-Bench: Are LMMs Masters at Evaluating AI-generated Images?
by: Zhang, Zicheng, et al.
Published: (2024)