Guardado en:
| Autores principales: | Baker, Mohammed Abu, Babu-Saheer, Lakshmi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2508.15847 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Composite Backdoor Attacks Against Large Language Models
por: Huang, Hai, et al.
Publicado: (2023)
por: Huang, Hai, et al.
Publicado: (2023)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
por: Xu, Jiashu, et al.
Publicado: (2023)
por: Xu, Jiashu, et al.
Publicado: (2023)
Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language Models
por: Das, Anindya Sundar, et al.
Publicado: (2025)
por: Das, Anindya Sundar, et al.
Publicado: (2025)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
por: Cho, Hakaze, et al.
Publicado: (2025)
por: Cho, Hakaze, et al.
Publicado: (2025)
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
por: Yan, Jun, et al.
Publicado: (2023)
por: Yan, Jun, et al.
Publicado: (2023)
SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
por: Xue, Eric, et al.
Publicado: (2025)
por: Xue, Eric, et al.
Publicado: (2025)
Mechanistic Anomaly Detection for "Quirky" Language Models
por: Johnston, David O., et al.
Publicado: (2025)
por: Johnston, David O., et al.
Publicado: (2025)
Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
por: Shang, Bingqi, et al.
Publicado: (2025)
por: Shang, Bingqi, et al.
Publicado: (2025)
CROW: Eliminating Backdoors from Large Language Models via Internal Consistency Regularization
por: Min, Nay Myat, et al.
Publicado: (2024)
por: Min, Nay Myat, et al.
Publicado: (2024)
Disentangling Exploration of Large Language Models by Optimal Exploitation
por: Grams, Tim, et al.
Publicado: (2025)
por: Grams, Tim, et al.
Publicado: (2025)
Eliminating Position Bias of Language Models: A Mechanistic Approach
por: Wang, Ziqi, et al.
Publicado: (2024)
por: Wang, Ziqi, et al.
Publicado: (2024)
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models
por: Huo, Yifu, et al.
Publicado: (2026)
por: Huo, Yifu, et al.
Publicado: (2026)
Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models
por: Binkowski, Jakub, et al.
Publicado: (2026)
por: Binkowski, Jakub, et al.
Publicado: (2026)
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
por: Peng, Runyu, et al.
Publicado: (2026)
por: Peng, Runyu, et al.
Publicado: (2026)
Test-Time Backdoor Attacks on Multimodal Large Language Models
por: Lu, Dong, et al.
Publicado: (2024)
por: Lu, Dong, et al.
Publicado: (2024)
How do Large Language Models Understand Relevance? A Mechanistic Interpretability Perspective
por: Liu, Qi, et al.
Publicado: (2025)
por: Liu, Qi, et al.
Publicado: (2025)
Analyzing And Editing Inner Mechanisms Of Backdoored Language Models
por: Lamparth, Max, et al.
Publicado: (2023)
por: Lamparth, Max, et al.
Publicado: (2023)
Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models
por: Liu, Runze, et al.
Publicado: (2025)
por: Liu, Runze, et al.
Publicado: (2025)
Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention Calibration
por: Yu, Zhongzhi, et al.
Publicado: (2024)
por: Yu, Zhongzhi, et al.
Publicado: (2024)
Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning Models
por: Tan, Wenhui, et al.
Publicado: (2026)
por: Tan, Wenhui, et al.
Publicado: (2026)
Z-Pruner: Post-Training Pruning of Large Language Models for Efficiency without Retraining
por: Bhuiyan, Samiul Basir, et al.
Publicado: (2025)
por: Bhuiyan, Samiul Basir, et al.
Publicado: (2025)
Watch Out for Your Guidance on Generation! Exploring Conditional Backdoor Attacks against Large Language Models
por: He, Jiaming, et al.
Publicado: (2024)
por: He, Jiaming, et al.
Publicado: (2024)
On the Impact of Language Nuances on Sentiment Analysis with Large Language Models: Paraphrasing, Sarcasm, and Emojis
por: Bhargava, Naman, et al.
Publicado: (2025)
por: Bhargava, Naman, et al.
Publicado: (2025)
Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges
por: Liu, Qin, et al.
Publicado: (2024)
por: Liu, Qin, et al.
Publicado: (2024)
Deep Learning Detection Method for Large Language Models-Generated Scientific Content
por: Alhijawi, Bushra, et al.
Publicado: (2024)
por: Alhijawi, Bushra, et al.
Publicado: (2024)
Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization
por: Son, Seungwoo, et al.
Publicado: (2024)
por: Son, Seungwoo, et al.
Publicado: (2024)
Entropic-Time Inference: Self-Organizing Large Language Model Decoding Beyond Attention
por: Kiruluta, Andrew
Publicado: (2026)
por: Kiruluta, Andrew
Publicado: (2026)
Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention
por: Gao, Bin, et al.
Publicado: (2024)
por: Gao, Bin, et al.
Publicado: (2024)
SEEKR: Selective Attention-Guided Knowledge Retention for Continual Learning of Large Language Models
por: He, Jinghan, et al.
Publicado: (2024)
por: He, Jinghan, et al.
Publicado: (2024)
LLM-PS: Empowering Large Language Models for Time Series Forecasting with Temporal Patterns and Semantics
por: Tang, Jialiang, et al.
Publicado: (2025)
por: Tang, Jialiang, et al.
Publicado: (2025)
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
por: Dai, Runpeng, et al.
Publicado: (2025)
por: Dai, Runpeng, et al.
Publicado: (2025)
Sliding Window Attention Training for Efficient Large Language Models
por: Fu, Zichuan, et al.
Publicado: (2025)
por: Fu, Zichuan, et al.
Publicado: (2025)
Instruction Following by Principled Boosting Attention of Large Language Models
por: Guardieiro, Vitoria, et al.
Publicado: (2025)
por: Guardieiro, Vitoria, et al.
Publicado: (2025)
Explaining Large Language Models with gSMILE
por: Dehghani, Zeinab, et al.
Publicado: (2025)
por: Dehghani, Zeinab, et al.
Publicado: (2025)
Dynamic Attention-Guided Context Decoding for Mitigating Context Faithfulness Hallucinations in Large Language Models
por: Huang, Yanwen, et al.
Publicado: (2025)
por: Huang, Yanwen, et al.
Publicado: (2025)
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
por: Tyukin, Georgy, et al.
Publicado: (2024)
por: Tyukin, Georgy, et al.
Publicado: (2024)
Unmasking Hallucinations: A Causal Graph-Attention Perspective on Factual Reliability in Large Language Models
por: kurra, Sailesh kiran, et al.
Publicado: (2026)
por: kurra, Sailesh kiran, et al.
Publicado: (2026)
Exploring Activation Patterns of Parameters in Language Models
por: Wang, Yudong, et al.
Publicado: (2024)
por: Wang, Yudong, et al.
Publicado: (2024)
UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models
por: Lin, Huawei, et al.
Publicado: (2025)
por: Lin, Huawei, et al.
Publicado: (2025)
Detecting and Understanding Vulnerabilities in Language Models via Mechanistic Interpretability
por: García-Carrasco, Jorge, et al.
Publicado: (2024)
por: García-Carrasco, Jorge, et al.
Publicado: (2024)
Ejemplares similares
-
Composite Backdoor Attacks Against Large Language Models
por: Huang, Hai, et al.
Publicado: (2023) -
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
por: Xu, Jiashu, et al.
Publicado: (2023) -
Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language Models
por: Das, Anindya Sundar, et al.
Publicado: (2025) -
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
por: Cho, Hakaze, et al.
Publicado: (2025) -
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
por: Yan, Jun, et al.
Publicado: (2023)