InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models via Human Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Henry Hengyuan, Pei, Wenqi, Tao, Yifei, Mei, Haiyang, Shou, Mike Zheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912692241956864
author Zhao, Henry Hengyuan
Pei, Wenqi
Tao, Yifei
Mei, Haiyang
Shou, Mike Zheng
author_facet Zhao, Henry Hengyuan
Pei, Wenqi
Tao, Yifei
Mei, Haiyang
Shou, Mike Zheng
contents Existing benchmarks do not test Large Multimodal Models (LMMs) on their interactive intelligence with human users, which is vital for developing general-purpose AI assistants. We design InterFeedback, an interactive framework, which can be applied to any LMM and dataset to assess this ability autonomously. On top of this, we introduce InterFeedback-Bench which evaluates interactive intelligence using two representative datasets, MMMU-Pro and MathVerse, to test 10 different open-source LMMs. Additionally, we present InterFeedback-Human, a newly collected dataset of 120 cases designed for manually testing interactive performance in leading models such as OpenAI-o1 and Claude-Sonnet-4. Our evaluation results indicate that even the state-of-the-art LMM, OpenAI-o1, struggles to refine its responses based on human feedback, achieving an average score of less than 50%. Our findings point to the need for methods that can enhance LMMs' capabilities to interpret and benefit from feedback.
format Preprint
id arxiv_https___arxiv_org_abs_2502_15027
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models via Human Feedback
Zhao, Henry Hengyuan
Pei, Wenqi
Tao, Yifei
Mei, Haiyang
Shou, Mike Zheng
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Human-Computer Interaction
Existing benchmarks do not test Large Multimodal Models (LMMs) on their interactive intelligence with human users, which is vital for developing general-purpose AI assistants. We design InterFeedback, an interactive framework, which can be applied to any LMM and dataset to assess this ability autonomously. On top of this, we introduce InterFeedback-Bench which evaluates interactive intelligence using two representative datasets, MMMU-Pro and MathVerse, to test 10 different open-source LMMs. Additionally, we present InterFeedback-Human, a newly collected dataset of 120 cases designed for manually testing interactive performance in leading models such as OpenAI-o1 and Claude-Sonnet-4. Our evaluation results indicate that even the state-of-the-art LMM, OpenAI-o1, struggles to refine its responses based on human feedback, achieving an average score of less than 50%. Our findings point to the need for methods that can enhance LMMs' capabilities to interpret and benefit from feedback.
title InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models via Human Feedback
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Human-Computer Interaction
url https://arxiv.org/abs/2502.15027