Interpreting Emergent Features in Deep Learning-based Side-channel Analysis

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Karayalçin, Sengim, Krček, Marina, Picek, Stjepan
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917094516326400
author Karayalçin, Sengim
Krček, Marina
Picek, Stjepan
author_facet Karayalçin, Sengim
Krček, Marina
Picek, Stjepan
contents Side-channel analysis (SCA) poses a real-world threat by exploiting unintentional physical signals to extract secret information from secure devices. Evaluation labs also use the same techniques to certify device security. In recent years, deep learning has emerged as a prominent method for SCA, achieving state-of-the-art attack performance at the cost of interpretability. Understanding how neural networks extract secrets is crucial for security evaluators aiming to defend against such attacks, as only by understanding the attack can one propose better countermeasures. In this work, we apply mechanistic interpretability to neural networks trained for SCA, revealing \textit{how} models exploit \textit{what} leakage in side-channel traces. We focus on sudden jumps in performance to reverse engineer learned representations, ultimately recovering secret masks and moving the evaluation process from black-box to white-box. Our results show that mechanistic interpretability can scale to realistic SCA settings, even when relevant inputs are sparse, model accuracies are low, and side-channel protections prevent standard input interventions.
format Preprint
id arxiv_https___arxiv_org_abs_2502_00384
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Interpreting Emergent Features in Deep Learning-based Side-channel Analysis
Karayalçin, Sengim
Krček, Marina
Picek, Stjepan
Cryptography and Security
Machine Learning
Side-channel analysis (SCA) poses a real-world threat by exploiting unintentional physical signals to extract secret information from secure devices. Evaluation labs also use the same techniques to certify device security. In recent years, deep learning has emerged as a prominent method for SCA, achieving state-of-the-art attack performance at the cost of interpretability. Understanding how neural networks extract secrets is crucial for security evaluators aiming to defend against such attacks, as only by understanding the attack can one propose better countermeasures. In this work, we apply mechanistic interpretability to neural networks trained for SCA, revealing \textit{how} models exploit \textit{what} leakage in side-channel traces. We focus on sudden jumps in performance to reverse engineer learned representations, ultimately recovering secret masks and moving the evaluation process from black-box to white-box. Our results show that mechanistic interpretability can scale to realistic SCA settings, even when relevant inputs are sparse, model accuracies are low, and side-channel protections prevent standard input interventions.
title Interpreting Emergent Features in Deep Learning-based Side-channel Analysis
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2502.00384