Evaluating Open-Source Large Language Models for Technical Telecom Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Caraus, Arina, Buscemi, Alessio, Kumar, Sumit, Turcanu, Ion
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911177924149248
author Caraus, Arina
Buscemi, Alessio
Kumar, Sumit
Turcanu, Ion
author_facet Caraus, Arina
Buscemi, Alessio
Kumar, Sumit
Turcanu, Ion
contents Large Language Models (LLMs) have shown remarkable capabilities across various fields. However, their performance in technical domains such as telecommunications remains underexplored. This paper evaluates two open-source LLMs, Gemma 3 27B and DeepSeek R1 32B, on factual and reasoning-based questions derived from advanced wireless communications material. We construct a benchmark of 105 question-answer pairs and assess performance using lexical metrics, semantic similarity, and LLM-as-a-judge scoring. We also analyze consistency, judgment reliability, and hallucination through source attribution and score variance. Results show that Gemma excels in semantic fidelity and LLM-rated correctness, while DeepSeek demonstrates slightly higher lexical consistency. Additional findings highlight current limitations in telecom applications and the need for domain-adapted models to support trustworthy Artificial Intelligence (AI) assistants in engineering.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21949
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Open-Source Large Language Models for Technical Telecom Question Answering
Caraus, Arina
Buscemi, Alessio
Kumar, Sumit
Turcanu, Ion
Networking and Internet Architecture
Computation and Language
Large Language Models (LLMs) have shown remarkable capabilities across various fields. However, their performance in technical domains such as telecommunications remains underexplored. This paper evaluates two open-source LLMs, Gemma 3 27B and DeepSeek R1 32B, on factual and reasoning-based questions derived from advanced wireless communications material. We construct a benchmark of 105 question-answer pairs and assess performance using lexical metrics, semantic similarity, and LLM-as-a-judge scoring. We also analyze consistency, judgment reliability, and hallucination through source attribution and score variance. Results show that Gemma excels in semantic fidelity and LLM-rated correctness, while DeepSeek demonstrates slightly higher lexical consistency. Additional findings highlight current limitations in telecom applications and the need for domain-adapted models to support trustworthy Artificial Intelligence (AI) assistants in engineering.
title Evaluating Open-Source Large Language Models for Technical Telecom Question Answering
topic Networking and Internet Architecture
Computation and Language
url https://arxiv.org/abs/2509.21949