Exposing Vulnerabilities in Explanation for Time Series Classifiers via Dual-Target Attacks

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Bohan, Liu, Zewen, Lin, Lu, Liu, Hui, Xiong, Li, Jin, Ming, Jin, Wei
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914620027961344
author Wang, Bohan
Liu, Zewen
Lin, Lu
Liu, Hui
Xiong, Li
Jin, Ming
Jin, Wei
author_facet Wang, Bohan
Liu, Zewen
Lin, Lu
Liu, Hui
Xiong, Li
Jin, Ming
Jin, Wei
contents Interpretable time series deep learning systems are often assessed by checking temporal consistency on explanations, implicitly treating this as evidence of robustness. We show that this assumption can fail: Predictions and explanations can be adversarially decoupled, enabling targeted misclassification while the explanation remains plausible and consistent with a chosen reference rationale. We propose TSEF (Time Series Explanation Fooler), a dual-target attack that jointly manipulates the classifier and explainer outputs. In contrast to single-objective misclassification attacks that disrupt explanation and spread attribution mass broadly, TSEF achieves targeted prediction changes while keeping explanations consistent with the reference. Across multiple datasets and explainer backbones, our results consistently reveal that explanation stability is a misleading proxy for decision robustness and motivate coupling-aware robustness evaluations for trustworthy time series tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02763
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Exposing Vulnerabilities in Explanation for Time Series Classifiers via Dual-Target Attacks
Wang, Bohan
Liu, Zewen
Lin, Lu
Liu, Hui
Xiong, Li
Jin, Ming
Jin, Wei
Machine Learning
Interpretable time series deep learning systems are often assessed by checking temporal consistency on explanations, implicitly treating this as evidence of robustness. We show that this assumption can fail: Predictions and explanations can be adversarially decoupled, enabling targeted misclassification while the explanation remains plausible and consistent with a chosen reference rationale. We propose TSEF (Time Series Explanation Fooler), a dual-target attack that jointly manipulates the classifier and explainer outputs. In contrast to single-objective misclassification attacks that disrupt explanation and spread attribution mass broadly, TSEF achieves targeted prediction changes while keeping explanations consistent with the reference. Across multiple datasets and explainer backbones, our results consistently reveal that explanation stability is a misleading proxy for decision robustness and motivate coupling-aware robustness evaluations for trustworthy time series tasks.
title Exposing Vulnerabilities in Explanation for Time Series Classifiers via Dual-Target Attacks
topic Machine Learning
url https://arxiv.org/abs/2602.02763