MMTB: Evaluating Terminal Agents on Multimedia-File Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Heo, Chiyeong, Kim, Jaechang, Kwon, Junhyuk, Kim, Hoyoung, Park, Dongmin, Lee, Jonghyun, Ok, Jungseul
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913113253609472
author Heo, Chiyeong
Kim, Jaechang
Kwon, Junhyuk
Kim, Hoyoung
Park, Dongmin
Lee, Jonghyun
Ok, Jungseul
author_facet Heo, Chiyeong
Kim, Jaechang
Kwon, Junhyuk
Kim, Hoyoung
Park, Dongmin
Lee, Jonghyun
Ok, Jungseul
contents Terminals provide a powerful interface for AI agents by exposing diverse tools for automating complex workflows, yet existing terminal-agent benchmarks largely focus on tasks grounded in text, code, and structured files. However, many real-world workflows require practitioners to work directly with audio and video files. Working with such multimedia files calls for terminal agents not only to understand multimedia content, but also to convert auditory and visual evidence across related files into appropriate actions. To evaluate terminal agents on multimedia-file tasks, we introduce MultiMedia-TerminalBench (MMTB), a benchmark of 105 tasks across 5 meta-categories where terminal agents directly operate with audio and video files. Alongside MMTB, we propose Terminus-MM, a multimedia harness that extends Terminus-KIRA with audio and video perception for terminal agents. Together, MMTB and Terminus-MM support a controlled study of multimedia terminal agents, revealing how different forms of multimedia access shape task outcomes and determine which evidence agents rely on to construct executable terminal workflows. MMTB media and metadata are released at https://huggingface.co/datasets/mm-tbench/mmtb-media
format Preprint
id arxiv_https___arxiv_org_abs_2605_10966
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MMTB: Evaluating Terminal Agents on Multimedia-File Tasks
Heo, Chiyeong
Kim, Jaechang
Kwon, Junhyuk
Kim, Hoyoung
Park, Dongmin
Lee, Jonghyun
Ok, Jungseul
Multimedia
Artificial Intelligence
Terminals provide a powerful interface for AI agents by exposing diverse tools for automating complex workflows, yet existing terminal-agent benchmarks largely focus on tasks grounded in text, code, and structured files. However, many real-world workflows require practitioners to work directly with audio and video files. Working with such multimedia files calls for terminal agents not only to understand multimedia content, but also to convert auditory and visual evidence across related files into appropriate actions. To evaluate terminal agents on multimedia-file tasks, we introduce MultiMedia-TerminalBench (MMTB), a benchmark of 105 tasks across 5 meta-categories where terminal agents directly operate with audio and video files. Alongside MMTB, we propose Terminus-MM, a multimedia harness that extends Terminus-KIRA with audio and video perception for terminal agents. Together, MMTB and Terminus-MM support a controlled study of multimedia terminal agents, revealing how different forms of multimedia access shape task outcomes and determine which evidence agents rely on to construct executable terminal workflows. MMTB media and metadata are released at https://huggingface.co/datasets/mm-tbench/mmtb-media
title MMTB: Evaluating Terminal Agents on Multimedia-File Tasks
topic Multimedia
Artificial Intelligence
url https://arxiv.org/abs/2605.10966