SALMAN: Stability Analysis of Language Models Through the Maps Between Graph-based Manifolds

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Cheng, Wuxinlin, Cao, Yupeng, Wu, Jinwen, Subbalakshmi, Koduvayur, Han, Tian, Feng, Zhuo
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912554489479168
author Cheng, Wuxinlin
Cao, Yupeng
Wu, Jinwen
Subbalakshmi, Koduvayur
Han, Tian
Feng, Zhuo
author_facet Cheng, Wuxinlin
Cao, Yupeng
Wu, Jinwen
Subbalakshmi, Koduvayur
Han, Tian
Feng, Zhuo
contents Recent strides in pretrained transformer-based language models have propelled state-of-the-art performance in numerous NLP tasks. Yet, as these models grow in size and deployment, their robustness under input perturbations becomes an increasingly urgent question. Existing robustness methods often diverge between small-parameter and large-scale models (LLMs), and they typically rely on labor-intensive, sample-specific adversarial designs. In this paper, we propose a unified, local (sample-level) robustness framework (SALMAN) that evaluates model stability without modifying internal parameters or resorting to complex perturbation heuristics. Central to our approach is a novel Distance Mapping Distortion (DMD) measure, which ranks each sample's susceptibility by comparing input-to-output distance mappings in a near-linear complexity manner. By demonstrating significant gains in attack efficiency and robust training, we position our framework as a practical, model-agnostic tool for advancing the reliability of transformer-based NLP systems.
format Preprint
id arxiv_https___arxiv_org_abs_2508_18306
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SALMAN: Stability Analysis of Language Models Through the Maps Between Graph-based Manifolds
Cheng, Wuxinlin
Cao, Yupeng
Wu, Jinwen
Subbalakshmi, Koduvayur
Han, Tian
Feng, Zhuo
Machine Learning
Artificial Intelligence
Computation and Language
Recent strides in pretrained transformer-based language models have propelled state-of-the-art performance in numerous NLP tasks. Yet, as these models grow in size and deployment, their robustness under input perturbations becomes an increasingly urgent question. Existing robustness methods often diverge between small-parameter and large-scale models (LLMs), and they typically rely on labor-intensive, sample-specific adversarial designs. In this paper, we propose a unified, local (sample-level) robustness framework (SALMAN) that evaluates model stability without modifying internal parameters or resorting to complex perturbation heuristics. Central to our approach is a novel Distance Mapping Distortion (DMD) measure, which ranks each sample's susceptibility by comparing input-to-output distance mappings in a near-linear complexity manner. By demonstrating significant gains in attack efficiency and robust training, we position our framework as a practical, model-agnostic tool for advancing the reliability of transformer-based NLP systems.
title SALMAN: Stability Analysis of Language Models Through the Maps Between Graph-based Manifolds
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.18306