SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Djuhera, Aladin, Kadhe, Swanand Ravindra, Ahmed, Farhan, Zawad, Syed, Koch, Fernando, Saad, Walid, Boche, Holger
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917253022220288
author Djuhera, Aladin
Kadhe, Swanand Ravindra
Ahmed, Farhan
Zawad, Syed
Koch, Fernando
Saad, Walid
Boche, Holger
author_facet Djuhera, Aladin
Kadhe, Swanand Ravindra
Ahmed, Farhan
Zawad, Syed
Koch, Fernando
Saad, Walid
Boche, Holger
contents Fine-tuning large language models (LLMs) on telecom datasets is a common practice to adapt general-purpose models to the telecom domain. However, little attention has been paid to how this process may compromise model safety. Recent research has shown that even benign fine-tuning can degrade the safety alignment of LLMs, causing them to respond to harmful or unethical user queries. In this paper, we investigate this issue by fine-tuning LLMs on three representative telecom datasets and show that safety degrades even for light telecom domain adaptation. To this end, we introduce TeleHarm, the first telecom-specific red-teaming benchmark, which we use alongside established DirectHarm and HexPhi datasets to systematically assess harmful behavior. We further extend our analysis to publicly available TeleLLMs that were continually pre-trained on large telecom corpora, revealing that safety alignment is severely lacking, primarily due to the omission of safety-focused instruction tuning. To address these issues, we evaluate three realignment defenses: SafeInstruct, SafeLoRA, SafeMERGE. We show that, across all settings, the proposed defenses can effectively restore safety without compromising telecom task performance, leading to Safe teleCOMMunication (SafeCOMM) models. Our work serves as both a diagnostic study and practical guide for safety realignment in telecom-tuned LLMs, underscoring the need for safety-aware instruction and fine-tuning in the telecom domain.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00062
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models
Djuhera, Aladin
Kadhe, Swanand Ravindra
Ahmed, Farhan
Zawad, Syed
Koch, Fernando
Saad, Walid
Boche, Holger
Computers and Society
Computation and Language
Cryptography and Security
Machine Learning
Fine-tuning large language models (LLMs) on telecom datasets is a common practice to adapt general-purpose models to the telecom domain. However, little attention has been paid to how this process may compromise model safety. Recent research has shown that even benign fine-tuning can degrade the safety alignment of LLMs, causing them to respond to harmful or unethical user queries. In this paper, we investigate this issue by fine-tuning LLMs on three representative telecom datasets and show that safety degrades even for light telecom domain adaptation. To this end, we introduce TeleHarm, the first telecom-specific red-teaming benchmark, which we use alongside established DirectHarm and HexPhi datasets to systematically assess harmful behavior. We further extend our analysis to publicly available TeleLLMs that were continually pre-trained on large telecom corpora, revealing that safety alignment is severely lacking, primarily due to the omission of safety-focused instruction tuning. To address these issues, we evaluate three realignment defenses: SafeInstruct, SafeLoRA, SafeMERGE. We show that, across all settings, the proposed defenses can effectively restore safety without compromising telecom task performance, leading to Safe teleCOMMunication (SafeCOMM) models. Our work serves as both a diagnostic study and practical guide for safety realignment in telecom-tuned LLMs, underscoring the need for safety-aware instruction and fine-tuning in the telecom domain.
title SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models
topic Computers and Society
Computation and Language
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2506.00062