Saved in:
Bibliographic Details
Main Authors: Zheng, Yushuo, Zhang, Zicheng, Min, Xiongkuo, Duan, Huiyu, Zhai, Guangtao
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.08928
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912640399310848
author Zheng, Yushuo
Zhang, Zicheng
Min, Xiongkuo
Duan, Huiyu
Zhai, Guangtao
author_facet Zheng, Yushuo
Zhang, Zicheng
Min, Xiongkuo
Duan, Huiyu
Zhai, Guangtao
contents Existing benchmarks for large multimodal models (LMMs) often fail to capture their performance in real-time, adversarial environments. We introduce LM Fight Arena (Large Model Fight Arena), a novel framework that evaluates LMMs by pitting them against each other in the classic fighting game Mortal Kombat II, a task requiring rapid visual understanding and tactical, sequential decision-making. In a controlled tournament, we test six leading open- and closed-source models, where each agent operates controlling the same character to ensure a fair comparison. The models are prompted to interpret game frames and state data to select their next actions. Unlike static evaluations, LM Fight Arena provides a fully automated, reproducible, and objective assessment of an LMM's strategic reasoning capabilities in a dynamic setting. This work introduces a challenging and engaging benchmark that bridges the gap between AI evaluation and interactive entertainment.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08928
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LM Fight Arena: Benchmarking Large Multimodal Models via Game Competition
Zheng, Yushuo
Zhang, Zicheng
Min, Xiongkuo
Duan, Huiyu
Zhai, Guangtao
Artificial Intelligence
Existing benchmarks for large multimodal models (LMMs) often fail to capture their performance in real-time, adversarial environments. We introduce LM Fight Arena (Large Model Fight Arena), a novel framework that evaluates LMMs by pitting them against each other in the classic fighting game Mortal Kombat II, a task requiring rapid visual understanding and tactical, sequential decision-making. In a controlled tournament, we test six leading open- and closed-source models, where each agent operates controlling the same character to ensure a fair comparison. The models are prompted to interpret game frames and state data to select their next actions. Unlike static evaluations, LM Fight Arena provides a fully automated, reproducible, and objective assessment of an LMM's strategic reasoning capabilities in a dynamic setting. This work introduces a challenging and engaging benchmark that bridges the gap between AI evaluation and interactive entertainment.
title LM Fight Arena: Benchmarking Large Multimodal Models via Game Competition
topic Artificial Intelligence
url https://arxiv.org/abs/2510.08928