Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lasnier, Théo, Antoun, Wissam, Kulumba, Francis, Seddah, Djamé |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Language-Switching Triggers Take a Latent Detour Through Language Models
von: Kulumba, Francis, et al.
Veröffentlicht: (2026)
von: Kulumba, Francis, et al.
Veröffentlicht: (2026)
From Text to Source: Results in Detecting Large Language Model-Generated Content
von: Antoun, Wissam, et al.
Veröffentlicht: (2023)
von: Antoun, Wissam, et al.
Veröffentlicht: (2023)
CamemBERT 2.0: A Smarter French Language Model Aged to Perfection
von: Antoun, Wissam, et al.
Veröffentlicht: (2024)
von: Antoun, Wissam, et al.
Veröffentlicht: (2024)
ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance
von: Antoun, Wissam, et al.
Veröffentlicht: (2025)
von: Antoun, Wissam, et al.
Veröffentlicht: (2025)
When Tables Go Crazy: Evaluating Multimodal Models on French Financial Documents
von: Mouilleron, Virginie, et al.
Veröffentlicht: (2026)
von: Mouilleron, Virginie, et al.
Veröffentlicht: (2026)
Beyond Dataset Creation: Critical View of Annotation Variation and Bias Probing of a Dataset for Online Radical Content Detection
von: Riabi, Arij, et al.
Veröffentlicht: (2024)
von: Riabi, Arij, et al.
Veröffentlicht: (2024)
Disentangling meaning from language in LLM-based machine translation
von: Lasnier, Théo, et al.
Veröffentlicht: (2026)
von: Lasnier, Théo, et al.
Veröffentlicht: (2026)
Gaperon: A Peppered English-French Generative Language Model Suite
von: Godey, Nathan, et al.
Veröffentlicht: (2025)
von: Godey, Nathan, et al.
Veröffentlicht: (2025)
HALvest-Contrastive: Retrieval-Like Authorship Attribution with Patch-Level Late Interaction
von: Kulumba, Francis, et al.
Veröffentlicht: (2024)
von: Kulumba, Francis, et al.
Veröffentlicht: (2024)
Can Character-based Language Models Improve Downstream Task Performance in Low-Resource and Noisy Language Scenarios?
von: Riabi, Arij, et al.
Veröffentlicht: (2021)
von: Riabi, Arij, et al.
Veröffentlicht: (2021)
Enriching the NArabizi Treebank: A Multifaceted Approach to Supporting an Under-Resourced Language
von: Riabi, Arij, et al.
Veröffentlicht: (2023)
von: Riabi, Arij, et al.
Veröffentlicht: (2023)
Common Ground, Diverse Roots: The Difficulty of Classifying Common Examples in Spanish Varieties
von: Lopetegui, Javier A., et al.
Veröffentlicht: (2024)
von: Lopetegui, Javier A., et al.
Veröffentlicht: (2024)
Where Does Authorship Signal Emerge in Encoder-Based Language Models?
von: Kulumba, Francis, et al.
Veröffentlicht: (2026)
von: Kulumba, Francis, et al.
Veröffentlicht: (2026)
Cloaked Classifiers: Pseudonymization Strategies on Sensitive Classification Tasks
von: Riabi, Arij, et al.
Veröffentlicht: (2024)
von: Riabi, Arij, et al.
Veröffentlicht: (2024)
Rethinking the Multilingual Reasoning Gap with Layer Swap
von: Lasbordes, Maxence, et al.
Veröffentlicht: (2026)
von: Lasbordes, Maxence, et al.
Veröffentlicht: (2026)
Mechanistic Exploration of Backdoored Large Language Model Attention Patterns
von: Baker, Mohammed Abu, et al.
Veröffentlicht: (2025)
von: Baker, Mohammed Abu, et al.
Veröffentlicht: (2025)
Mechanistic Circuit-Based Knowledge Editing in Large Language Models
von: Zhao, Tianyi, et al.
Veröffentlicht: (2026)
von: Zhao, Tianyi, et al.
Veröffentlicht: (2026)
Multilingual, Multimodal Pipeline for Creating Authentic and Structured Fact-Checked Claim Dataset
von: Hüsünbeyi, Z. Melce, et al.
Veröffentlicht: (2026)
von: Hüsünbeyi, Z. Melce, et al.
Veröffentlicht: (2026)
Hidden Ads: Behavior Triggered Semantic Backdoors for Advertisement Injection in Vision Language Models
von: Yao, Duanyi, et al.
Veröffentlicht: (2026)
von: Yao, Duanyi, et al.
Veröffentlicht: (2026)
Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
Mechanistic Behavior Editing of Language Models
von: Singh, Joykirat, et al.
Veröffentlicht: (2024)
von: Singh, Joykirat, et al.
Veröffentlicht: (2024)
Goal Hijacking Attack on Large Language Models via Pseudo-Conversation Injection
von: Chen, Zheng, et al.
Veröffentlicht: (2024)
von: Chen, Zheng, et al.
Veröffentlicht: (2024)
Hijacking Large Language Models via Adversarial In-Context Learning
von: Zhou, Xiangyu, et al.
Veröffentlicht: (2023)
von: Zhou, Xiangyu, et al.
Veröffentlicht: (2023)
When Emotion Becomes Trigger: Emotion-style dynamic Backdoor Attack Parasitising Large Language Models
von: Liu, Ziyu, et al.
Veröffentlicht: (2026)
von: Liu, Ziyu, et al.
Veröffentlicht: (2026)
Merging Triggers, Breaking Backdoors: Defensive Poisoning for Instruction-Tuned Language Models
von: Kim, San, et al.
Veröffentlicht: (2026)
von: Kim, San, et al.
Veröffentlicht: (2026)
ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs
von: Zhao, Gejian, et al.
Veröffentlicht: (2025)
von: Zhao, Gejian, et al.
Veröffentlicht: (2025)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
von: Kim, Geonhee, et al.
Veröffentlicht: (2024)
von: Kim, Geonhee, et al.
Veröffentlicht: (2024)
Mechanistic Indicators of Steering Effectiveness in Large Language Models
von: Jafari, Mehdi, et al.
Veröffentlicht: (2026)
von: Jafari, Mehdi, et al.
Veröffentlicht: (2026)
Empirical Analysis of Large Vision-Language Models against Goal Hijacking via Visual Prompt Injection
von: Kimura, Subaru, et al.
Veröffentlicht: (2024)
von: Kimura, Subaru, et al.
Veröffentlicht: (2024)
Lethe: Purifying Backdoored Large Language Models with Knowledge Dilution
von: Chen, Chen, et al.
Veröffentlicht: (2025)
von: Chen, Chen, et al.
Veröffentlicht: (2025)
Neutralizing Backdoors through Information Conflicts for Large Language Models
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
Behavioral Analysis of Information Salience in Large Language Models
von: Trienes, Jan, et al.
Veröffentlicht: (2025)
von: Trienes, Jan, et al.
Veröffentlicht: (2025)
Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Models
von: Nie, Ercong, et al.
Veröffentlicht: (2025)
von: Nie, Ercong, et al.
Veröffentlicht: (2025)
Mechanistic Indicators of Understanding in Large Language Models
von: Beckmann, Pierre, et al.
Veröffentlicht: (2025)
von: Beckmann, Pierre, et al.
Veröffentlicht: (2025)
Investigating Adversarial Trigger Transfer in Large Language Models
von: Meade, Nicholas, et al.
Veröffentlicht: (2024)
von: Meade, Nicholas, et al.
Veröffentlicht: (2024)
Seeing No Evil: Blinding Large Vision-Language Models to Safety Instructions via Adversarial Attention Hijacking
von: Li, Jingru, et al.
Veröffentlicht: (2026)
von: Li, Jingru, et al.
Veröffentlicht: (2026)
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
von: Tong, Terry, et al.
Veröffentlicht: (2024)
von: Tong, Terry, et al.
Veröffentlicht: (2024)
Mechanistic Decoding of Cognitive Constructs in Large Language Models
von: Shou, Yitong, et al.
Veröffentlicht: (2026)
von: Shou, Yitong, et al.
Veröffentlicht: (2026)
Mechanistic Interpretability of Emotion Inference in Large Language Models
von: Tak, Ala N., et al.
Veröffentlicht: (2025)
von: Tak, Ala N., et al.
Veröffentlicht: (2025)
Hijacking Context in Large Multi-modal Models
von: Jeong, Joonhyun
Veröffentlicht: (2023)
von: Jeong, Joonhyun
Veröffentlicht: (2023)
Ähnliche Einträge
-
Language-Switching Triggers Take a Latent Detour Through Language Models
von: Kulumba, Francis, et al.
Veröffentlicht: (2026) -
From Text to Source: Results in Detecting Large Language Model-Generated Content
von: Antoun, Wissam, et al.
Veröffentlicht: (2023) -
CamemBERT 2.0: A Smarter French Language Model Aged to Perfection
von: Antoun, Wissam, et al.
Veröffentlicht: (2024) -
ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance
von: Antoun, Wissam, et al.
Veröffentlicht: (2025) -
When Tables Go Crazy: Evaluating Multimodal Models on French Financial Documents
von: Mouilleron, Virginie, et al.
Veröffentlicht: (2026)