Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Almeida, Thales Sales, Santos, João Guilherme Alves, Laitz, Thiago, Bonás, Giovana Kerche
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909794504278016
author Almeida, Thales Sales
Santos, João Guilherme Alves
Laitz, Thiago
Bonás, Giovana Kerche
author_facet Almeida, Thales Sales
Santos, João Guilherme Alves
Laitz, Thiago
Bonás, Giovana Kerche
contents Large language models (LLMs) are increasingly deployed as task-oriented agents, where success depends on their ability to generate accurate function calls under realistic, multilingual conditions. However, existing agent evaluations largely overlook cultural and linguistic diversity, often relying on monolingual or naively translated benchmarks. We introduce Ticket-Bench, a benchmark for multilingual agent evaluation in task-oriented scenarios. Ticket-Bench simulates the domain of soccer ticket purchases across six major languages: Portuguese, English, Spanish, German, Italian, and French. Using localized teams, cities, and user profiles to provide a higher level of realism. We evaluate a wide range of commercial and open-source LLMs, measuring function-calling accuracy and consistency across languages. Results show that reasoning-oriented models (e.g., GPT-5, Qwen3-235B) dominate performance but still exhibit notable cross-lingual disparities. These findings underscore the need for culturally aware, multilingual benchmarks to guide the development of robust LLM agents.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14477
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
Almeida, Thales Sales
Santos, João Guilherme Alves
Laitz, Thiago
Bonás, Giovana Kerche
Computation and Language
Large language models (LLMs) are increasingly deployed as task-oriented agents, where success depends on their ability to generate accurate function calls under realistic, multilingual conditions. However, existing agent evaluations largely overlook cultural and linguistic diversity, often relying on monolingual or naively translated benchmarks. We introduce Ticket-Bench, a benchmark for multilingual agent evaluation in task-oriented scenarios. Ticket-Bench simulates the domain of soccer ticket purchases across six major languages: Portuguese, English, Spanish, German, Italian, and French. Using localized teams, cities, and user profiles to provide a higher level of realism. We evaluate a wide range of commercial and open-source LLMs, measuring function-calling accuracy and consistency across languages. Results show that reasoning-oriented models (e.g., GPT-5, Qwen3-235B) dominate performance but still exhibit notable cross-lingual disparities. These findings underscore the need for culturally aware, multilingual benchmarks to guide the development of robust LLM agents.
title Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
topic Computation and Language
url https://arxiv.org/abs/2509.14477