TeleMath: A Benchmark for Large Language Models in Telecom Mathematical Problem Solving

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Colle, Vincenzo, Sana, Mohamed, Piovesan, Nicola, De Domenico, Antonio, Ayed, Fadhel, Debbah, Merouane
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918056964390912
author Colle, Vincenzo
Sana, Mohamed
Piovesan, Nicola
De Domenico, Antonio
Ayed, Fadhel
Debbah, Merouane
author_facet Colle, Vincenzo
Sana, Mohamed
Piovesan, Nicola
De Domenico, Antonio
Ayed, Fadhel
Debbah, Merouane
contents The increasing adoption of artificial intelligence in telecommunications has raised interest in the capability of Large Language Models (LLMs) to address domain-specific, mathematically intensive tasks. Although recent advancements have improved the performance of LLMs in general mathematical reasoning, their effectiveness within specialized domains, such as signal processing, network optimization, and performance analysis, remains largely unexplored. To address this gap, we introduce TeleMath, the first benchmark dataset specifically designed to evaluate LLM performance in solving mathematical problems with numerical solutions in the telecommunications domain. Comprising 500 question-answer (QnA) pairs, TeleMath covers a wide spectrum of topics in the telecommunications field. This paper outlines the proposed QnAs generation pipeline, starting from a selected seed of problems crafted by Subject Matter Experts. The evaluation of a wide range of open-source LLMs reveals that best performance on TeleMath is achieved by recent models explicitly designed for mathematical or logical reasoning. In contrast, general-purpose models, even those with a large number of parameters, often struggle with these challenges. We have released the dataset and the evaluation code to ease result reproducibility and support future research.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10674
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TeleMath: A Benchmark for Large Language Models in Telecom Mathematical Problem Solving
Colle, Vincenzo
Sana, Mohamed
Piovesan, Nicola
De Domenico, Antonio
Ayed, Fadhel
Debbah, Merouane
Artificial Intelligence
Computation and Language
The increasing adoption of artificial intelligence in telecommunications has raised interest in the capability of Large Language Models (LLMs) to address domain-specific, mathematically intensive tasks. Although recent advancements have improved the performance of LLMs in general mathematical reasoning, their effectiveness within specialized domains, such as signal processing, network optimization, and performance analysis, remains largely unexplored. To address this gap, we introduce TeleMath, the first benchmark dataset specifically designed to evaluate LLM performance in solving mathematical problems with numerical solutions in the telecommunications domain. Comprising 500 question-answer (QnA) pairs, TeleMath covers a wide spectrum of topics in the telecommunications field. This paper outlines the proposed QnAs generation pipeline, starting from a selected seed of problems crafted by Subject Matter Experts. The evaluation of a wide range of open-source LLMs reveals that best performance on TeleMath is achieved by recent models explicitly designed for mathematical or logical reasoning. In contrast, general-purpose models, even those with a large number of parameters, often struggle with these challenges. We have released the dataset and the evaluation code to ease result reproducibility and support future research.
title TeleMath: A Benchmark for Large Language Models in Telecom Mathematical Problem Solving
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.10674