Saved in:
Bibliographic Details
Main Authors: Wang, Meiping, Zhong, Jian, Han, Rongduo, Kang, Liming, Shi, Zhengkun, Liang, Xiao, Lin, Xing, Gao, Nan, Zhang, Haining
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.09507
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917030140051456
author Wang, Meiping
Zhong, Jian
Han, Rongduo
Kang, Liming
Shi, Zhengkun
Liang, Xiao
Lin, Xing
Gao, Nan
Zhang, Haining
author_facet Wang, Meiping
Zhong, Jian
Han, Rongduo
Kang, Liming
Shi, Zhengkun
Liang, Xiao
Lin, Xing
Gao, Nan
Zhang, Haining
contents With the rapid development of mobile intelligent assistant technologies, multi-modal AI assistants have become essential interfaces for daily user interactions. However, current evaluation methods face challenges including high manual costs, inconsistent standards, and subjective bias. This paper proposes an automated multi-modal evaluation framework based on large language models and multi-agent collaboration. The framework employs a three-tier agent architecture consisting of interaction evaluation agents, semantic verification agents, and experience decision agents. Through supervised fine-tuning on the Qwen3-8B model, we achieve a significant evaluation matching accuracy with human experts. Experimental results on eight major intelligent agents demonstrate the framework's effectiveness in predicting users' satisfaction and identifying generation defects.
format Preprint
id arxiv_https___arxiv_org_abs_2508_09507
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Automated Multi-modal Evaluation Framework for Mobile Intelligent Assistants Based on Large Language Models and Multi-Agent Collaboration
Wang, Meiping
Zhong, Jian
Han, Rongduo
Kang, Liming
Shi, Zhengkun
Liang, Xiao
Lin, Xing
Gao, Nan
Zhang, Haining
Artificial Intelligence
With the rapid development of mobile intelligent assistant technologies, multi-modal AI assistants have become essential interfaces for daily user interactions. However, current evaluation methods face challenges including high manual costs, inconsistent standards, and subjective bias. This paper proposes an automated multi-modal evaluation framework based on large language models and multi-agent collaboration. The framework employs a three-tier agent architecture consisting of interaction evaluation agents, semantic verification agents, and experience decision agents. Through supervised fine-tuning on the Qwen3-8B model, we achieve a significant evaluation matching accuracy with human experts. Experimental results on eight major intelligent agents demonstrate the framework's effectiveness in predicting users' satisfaction and identifying generation defects.
title An Automated Multi-modal Evaluation Framework for Mobile Intelligent Assistants Based on Large Language Models and Multi-Agent Collaboration
topic Artificial Intelligence
url https://arxiv.org/abs/2508.09507