Affordance Benchmark for MLLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Junying, Li, Wenzhe, Wu, Yalun, Liang, Yingji, Guo, Yijin, Li, Chunyi, Duan, Haodong, Zhang, Zicheng, Zhai, Guangtao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Ever-Evolving Science Exam
von: Wang, Junying, et al.
Veröffentlicht: (2025)
von: Wang, Junying, et al.
Veröffentlicht: (2025)
Q-Mirror: Unlocking the Multi-Modal Potential of Scientific Text-Only QA Pairs
von: Wang, Junying, et al.
Veröffentlicht: (2025)
von: Wang, Junying, et al.
Veröffentlicht: (2025)
Redundancy Principles for MLLMs Benchmarks
von: Zhang, Zicheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zicheng, et al.
Veröffentlicht: (2025)
A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation
von: Shen, Ye, et al.
Veröffentlicht: (2025)
von: Shen, Ye, et al.
Veröffentlicht: (2025)
MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror
von: Guo, Shengyu, et al.
Veröffentlicht: (2026)
von: Guo, Shengyu, et al.
Veröffentlicht: (2026)
Human-Centric Evaluation for Foundation Models
von: Guo, Yijin, et al.
Veröffentlicht: (2025)
von: Guo, Yijin, et al.
Veröffentlicht: (2025)
H2HTalk: Evaluating Large Language Models as Emotional Companion
von: Wang, Boyang, et al.
Veröffentlicht: (2025)
von: Wang, Boyang, et al.
Veröffentlicht: (2025)
EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory
von: Shen, Ye, et al.
Veröffentlicht: (2026)
von: Shen, Ye, et al.
Veröffentlicht: (2026)
QoNext: Towards Next-generation QoE for Foundation Models
von: Guo, Yijin, et al.
Veröffentlicht: (2025)
von: Guo, Yijin, et al.
Veröffentlicht: (2025)
GeoX-Bench: Benchmarking Cross-View Geo-Localization and Pose Estimation Capabilities of Large Multimodal Models
von: Zheng, Yushuo, et al.
Veröffentlicht: (2025)
von: Zheng, Yushuo, et al.
Veröffentlicht: (2025)
STAR : Bridging Statistical and Agentic Reasoning for Large Model Performance Prediction
von: Wang, Xiaoxiao, et al.
Veröffentlicht: (2026)
von: Wang, Xiaoxiao, et al.
Veröffentlicht: (2026)
MEMO-Bench: A Multiple Benchmark for Text-to-Image and Multimodal Large Language Models on Human Emotion Analysis
von: Zhou, Yingjie, et al.
Veröffentlicht: (2024)
von: Zhou, Yingjie, et al.
Veröffentlicht: (2024)
Improve MLLM Benchmark Efficiency through Interview
von: Wen, Farong, et al.
Veröffentlicht: (2025)
von: Wen, Farong, et al.
Veröffentlicht: (2025)
Information Density Principle for MLLM Benchmarks
von: Li, Chunyi, et al.
Veröffentlicht: (2025)
von: Li, Chunyi, et al.
Veröffentlicht: (2025)
Quality Assessment in the Era of Large Models: A Survey
von: Zhang, Zicheng, et al.
Veröffentlicht: (2024)
von: Zhang, Zicheng, et al.
Veröffentlicht: (2024)
Explore the Hallucination on Low-level Perception for MLLMs
von: Sun, Yinan, et al.
Veröffentlicht: (2024)
von: Sun, Yinan, et al.
Veröffentlicht: (2024)
Evaluating from Benign to Dynamic Adversarial: A Squid Game for Large Language Models
von: Chen, Zijian, et al.
Veröffentlicht: (2025)
von: Chen, Zijian, et al.
Veröffentlicht: (2025)
Children's Intelligence Tests Pose Challenges for MLLMs? KidGym: A 2D Grid-Based Reasoning Benchmark for MLLMs
von: Ye, Hengwei, et al.
Veröffentlicht: (2026)
von: Ye, Hengwei, et al.
Veröffentlicht: (2026)
Quasi-symbolic Semantic Geometry over Transformer-based Variational AutoEncoder
von: Zhang, Yingji, et al.
Veröffentlicht: (2022)
von: Zhang, Yingji, et al.
Veröffentlicht: (2022)
Learning Disentangled Semantic Spaces of Explanations via Invertible Neural Networks
von: Zhang, Yingji, et al.
Veröffentlicht: (2023)
von: Zhang, Yingji, et al.
Veröffentlicht: (2023)
Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision
von: Wu, Haoning, et al.
Veröffentlicht: (2023)
von: Wu, Haoning, et al.
Veröffentlicht: (2023)
A Survey on Fairness in Large Language Models
von: Li, Yingji, et al.
Veröffentlicht: (2023)
von: Li, Yingji, et al.
Veröffentlicht: (2023)
AdaCodec: A Predictive Visual Code for Video MLLMs
von: Hou, Haowen, et al.
Veröffentlicht: (2026)
von: Hou, Haowen, et al.
Veröffentlicht: (2026)
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities
von: Jiang, Shixin, et al.
Veröffentlicht: (2024)
von: Jiang, Shixin, et al.
Veröffentlicht: (2024)
LM Fight Arena: Benchmarking Large Multimodal Models via Game Competition
von: Zheng, Yushuo, et al.
Veröffentlicht: (2025)
von: Zheng, Yushuo, et al.
Veröffentlicht: (2025)
LangVAE and LangSpace: Building and Probing for Language Model VAEs
von: Carvalho, Danilo S., et al.
Veröffentlicht: (2025)
von: Carvalho, Danilo S., et al.
Veröffentlicht: (2025)
Judge Before Answer: Can MLLM Discern the False Premise in Question?
von: Li, Jidong, et al.
Veröffentlicht: (2025)
von: Li, Jidong, et al.
Veröffentlicht: (2025)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
von: Wang, Chonghua, et al.
Veröffentlicht: (2024)
von: Wang, Chonghua, et al.
Veröffentlicht: (2024)
COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMs
von: Guo, Peizheng, et al.
Veröffentlicht: (2025)
von: Guo, Peizheng, et al.
Veröffentlicht: (2025)
ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints
von: Chen, Pei-An, et al.
Veröffentlicht: (2026)
von: Chen, Pei-An, et al.
Veröffentlicht: (2026)
Shortcut Learning in In-Context Learning: A Survey
von: Song, Rui, et al.
Veröffentlicht: (2024)
von: Song, Rui, et al.
Veröffentlicht: (2024)
Toward Robust In-Context Learning: Leveraging Out-of-distribution Proxies for Target Inaccessible Demonstration Retrieval
von: Xu, Hao, et al.
Veröffentlicht: (2026)
von: Xu, Hao, et al.
Veröffentlicht: (2026)
Enhancing Multimodal Large Language Models for Ancient Chinese Character Evolution Analysis via Glyph-Driven Fine-Tuning
von: Song, Rui, et al.
Veröffentlicht: (2026)
von: Song, Rui, et al.
Veröffentlicht: (2026)
Find Them All: Unveiling MLLMs for Versatile Person Re-identification
von: Li, Jinhao, et al.
Veröffentlicht: (2025)
von: Li, Jinhao, et al.
Veröffentlicht: (2025)
Scaling-up Perceptual Video Quality Assessment
von: Jia, Ziheng, et al.
Veröffentlicht: (2025)
von: Jia, Ziheng, et al.
Veröffentlicht: (2025)
SIQA: Toward Reliable Scientific Image Quality Assessment
von: Li, Wenzhe, et al.
Veröffentlicht: (2026)
von: Li, Wenzhe, et al.
Veröffentlicht: (2026)
Up to 36x Speedup: Mask-based Parallel Inference Paradigm for Key Information Extraction in MLLMs
von: Wang, Xinzhong, et al.
Veröffentlicht: (2026)
von: Wang, Xinzhong, et al.
Veröffentlicht: (2026)
Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification?
von: Lin, Fei, et al.
Veröffentlicht: (2025)
von: Lin, Fei, et al.
Veröffentlicht: (2025)
CNSL-bench: Benchmarking the Sign Language Understanding Capabilities of MLLMs on Chinese National Sign Language
von: Zhao, Rui, et al.
Veröffentlicht: (2026)
von: Zhao, Rui, et al.
Veröffentlicht: (2026)
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement
von: Cao, Maosong, et al.
Veröffentlicht: (2025)
von: Cao, Maosong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The Ever-Evolving Science Exam
von: Wang, Junying, et al.
Veröffentlicht: (2025) -
Q-Mirror: Unlocking the Multi-Modal Potential of Scientific Text-Only QA Pairs
von: Wang, Junying, et al.
Veröffentlicht: (2025) -
Redundancy Principles for MLLMs Benchmarks
von: Zhang, Zicheng, et al.
Veröffentlicht: (2025) -
A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation
von: Shen, Ye, et al.
Veröffentlicht: (2025) -
MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror
von: Guo, Shengyu, et al.
Veröffentlicht: (2026)