Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models
Fuente:
arXiv
Saved in:
| Main Authors: | Braun, Tobias, Grebe, Jonas Henry, Shakibania, Hossein, Rohrbach, Anna, Rohrbach, Marcus |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
by: Braun, Tobias, et al.
Published: (2025)
by: Braun, Tobias, et al.
Published: (2025)
Geometric Erasure by Contrastive Velocity Matching in Rectified Flows
by: Grebe, Jonas Henry, et al.
Published: (2026)
by: Grebe, Jonas Henry, et al.
Published: (2026)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
by: Jiang, Peihai, et al.
Published: (2025)
by: Jiang, Peihai, et al.
Published: (2025)
Compromising Embodied Agents with Contextual Backdoor Attacks
by: Liu, Aishan, et al.
Published: (2024)
by: Liu, Aishan, et al.
Published: (2024)
Exploring Backdoor Vulnerabilities of Chat Models
by: Hao, Yunzhuo, et al.
Published: (2024)
by: Hao, Yunzhuo, et al.
Published: (2024)
Membership Inference Attacks on Tokenizers of Large Language Models
by: Tong, Meng, et al.
Published: (2025)
by: Tong, Meng, et al.
Published: (2025)
Selection-Based Vulnerabilities: Clean-Label Backdoor Attacks in Active Learning
by: Zhi, Yuhan, et al.
Published: (2025)
by: Zhi, Yuhan, et al.
Published: (2025)
Token-Efficient Prompt Injection Attack: Provoking Cessation in LLM Reasoning via Adaptive Token Compression
by: Cui, Yu, et al.
Published: (2025)
by: Cui, Yu, et al.
Published: (2025)
Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models
by: Yu, Miao, et al.
Published: (2025)
by: Yu, Miao, et al.
Published: (2025)
Reliable Model Watermarking: Defending Against Theft without Compromising on Evasion
by: Zhu, Hongyu, et al.
Published: (2024)
by: Zhu, Hongyu, et al.
Published: (2024)
Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
by: Xu, Jiashu, et al.
Published: (2023)
by: Xu, Jiashu, et al.
Published: (2023)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
by: Tong, Terry, et al.
Published: (2025)
by: Tong, Terry, et al.
Published: (2025)
MTRE: Multi-Token Reliability Estimation for Hallucination Detection in VLMs
by: Zollicoffer, Geigh, et al.
Published: (2025)
by: Zollicoffer, Geigh, et al.
Published: (2025)
Exploiting the Vulnerability of Large Language Models via Defense-Aware Architectural Backdoor
by: Miah, Abdullah Arafat, et al.
Published: (2024)
by: Miah, Abdullah Arafat, et al.
Published: (2024)
GUARD-SLM: Token Activation-Based Defense Against Jailbreak Attacks for Small Language Models
by: Mia, Md Jueal, et al.
Published: (2026)
by: Mia, Md Jueal, et al.
Published: (2026)
Unified Neural Backdoor Removal with Only Few Clean Samples through Unlearning and Relearning
by: Min, Nay Myat, et al.
Published: (2024)
by: Min, Nay Myat, et al.
Published: (2024)
Backdoor Sentinel: Detecting and Detoxifying Backdoors in Diffusion Models via Temporal Noise Consistency
by: Wang, Bingzheng, et al.
Published: (2026)
by: Wang, Bingzheng, et al.
Published: (2026)
When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations
by: Ge, Huaizhi, et al.
Published: (2024)
by: Ge, Huaizhi, et al.
Published: (2024)
Understanding Secret Leakage Risks in Code LLMs: A Tokenization Perspective
by: Chen, Meifang, et al.
Published: (2026)
by: Chen, Meifang, et al.
Published: (2026)
Privacy Guard & Token Parsimony by Prompt and Context Handling and LLM Routing
by: Langiu, Alessio
Published: (2026)
by: Langiu, Alessio
Published: (2026)
AI-Governed Agent Architecture for Web-Trustworthy Tokenization of Alternative Assets
by: Borjigin, Ailiya, et al.
Published: (2025)
by: Borjigin, Ailiya, et al.
Published: (2025)
UTF:Undertrained Tokens as Fingerprints A Novel Approach to LLM Identification
by: Cai, Jiacheng, et al.
Published: (2024)
by: Cai, Jiacheng, et al.
Published: (2024)
TokenMark: A Modality-Agnostic Watermark for Pre-trained Transformers
by: Xu, Hengyuan, et al.
Published: (2024)
by: Xu, Hengyuan, et al.
Published: (2024)
Less Is More -- Until It Breaks: Security Pitfalls of Vision Token Compression in Large Vision-Language Models
by: Zhang, Xiaomei, et al.
Published: (2026)
by: Zhang, Xiaomei, et al.
Published: (2026)
Exploiting Layer-Specific Vulnerabilities to Backdoor Attack in Federated Learning
by: Foroughi, Mohammad Hadi, et al.
Published: (2026)
by: Foroughi, Mohammad Hadi, et al.
Published: (2026)
Backdooring Bias in Large Language Models
by: Das, Anudeep, et al.
Published: (2026)
by: Das, Anudeep, et al.
Published: (2026)
Lightweight and Fast Backdoor Model Detection
by: Yu, Yinbo, et al.
Published: (2026)
by: Yu, Yinbo, et al.
Published: (2026)
DeBackdoor: A Deductive Framework for Detecting Backdoor Attacks on Deep Models with Limited Data
by: Popovic, Dorde, et al.
Published: (2025)
by: Popovic, Dorde, et al.
Published: (2025)
BackdoorMBTI: A Backdoor Learning Multimodal Benchmark Tool Kit for Backdoor Defense Evaluation
by: Yu, Haiyang, et al.
Published: (2024)
by: Yu, Haiyang, et al.
Published: (2024)
MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation
by: Zhu, Wentian, et al.
Published: (2025)
by: Zhu, Wentian, et al.
Published: (2025)
Auditing Pay-Per-Token in Large Language Models
by: Velasco, Ander Artola, et al.
Published: (2025)
by: Velasco, Ander Artola, et al.
Published: (2025)
TERD: A Unified Framework for Safeguarding Diffusion Models Against Backdoors
by: Mo, Yichuan, et al.
Published: (2024)
by: Mo, Yichuan, et al.
Published: (2024)
Towards Backdoor Stealthiness in Model Parameter Space
by: Xu, Xiaoyun, et al.
Published: (2025)
by: Xu, Xiaoyun, et al.
Published: (2025)
Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers
by: Zhao, Andrew, et al.
Published: (2025)
by: Zhao, Andrew, et al.
Published: (2025)
Backdoor4Good: Benchmarking Beneficial Uses of Backdoors in LLMs
by: Li, Yige, et al.
Published: (2026)
by: Li, Yige, et al.
Published: (2026)
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
by: Guo, Weiyang, et al.
Published: (2026)
by: Guo, Weiyang, et al.
Published: (2026)
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
by: Li, Yige, et al.
Published: (2025)
by: Li, Yige, et al.
Published: (2025)
Safety Alignment Should Be Made More Than Just a Few Tokens Deep
by: Qi, Xiangyu, et al.
Published: (2024)
by: Qi, Xiangyu, et al.
Published: (2024)
Kill Two Birds with One Stone! Trajectory enabled Unified Online Detection of Adversarial Examples and Backdoor Attacks
by: Fu, Anmin, et al.
Published: (2025)
by: Fu, Anmin, et al.
Published: (2025)
Similar Items
-
Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
by: Braun, Tobias, et al.
Published: (2025) -
Geometric Erasure by Contrastive Velocity Matching in Rectified Flows
by: Grebe, Jonas Henry, et al.
Published: (2026) -
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
by: Jiang, Peihai, et al.
Published: (2025) -
Compromising Embodied Agents with Contextual Backdoor Attacks
by: Liu, Aishan, et al.
Published: (2024) -
Exploring Backdoor Vulnerabilities of Chat Models
by: Hao, Yunzhuo, et al.
Published: (2024)