Analyzing Multi-Head Attention on Trojan BERT Models
Fuente:
arXiv
Saved in:
| Main Author: | Wang, Jingwei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Latent Multi-Head Attention for Small Language Models
by: Mehta, Sushant, et al.
Published: (2025)
by: Mehta, Sushant, et al.
Published: (2025)
Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models
by: Mąka, Paweł, et al.
Published: (2024)
by: Mąka, Paweł, et al.
Published: (2024)
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
by: Lu, Yi, et al.
Published: (2024)
by: Lu, Yi, et al.
Published: (2024)
Analyzing Narrative Processing in Large Language Models (LLMs): Using GPT4 to test BERT
by: Krauss, Patrick, et al.
Published: (2024)
by: Krauss, Patrick, et al.
Published: (2024)
Analyzing Gender Polarity in Short Social Media Texts with BERT: The Role of Emojis and Emoticons
by: Jazi, Saba Yousefian, et al.
Published: (2024)
by: Jazi, Saba Yousefian, et al.
Published: (2024)
KliniskVestBERT: BERT Model Specialised to Norwegian Clinical Texts
by: Autenried, Christian, et al.
Published: (2026)
by: Autenried, Christian, et al.
Published: (2026)
Unitary Multi-Margin BERT for Robust Natural Language Processing
by: Chang, Hao-Yuan, et al.
Published: (2024)
by: Chang, Hao-Yuan, et al.
Published: (2024)
KuBERT: Central Kurdish BERT Model and Its Application for Sentiment Analysis
by: Awlla, Kozhin muhealddin, et al.
Published: (2025)
by: Awlla, Kozhin muhealddin, et al.
Published: (2025)
BERT-LSH: Reducing Absolute Compute For Attention
by: Li, Zezheng, et al.
Published: (2024)
by: Li, Zezheng, et al.
Published: (2024)
Advancing Pancreatic Cancer Prediction with a Next Visit Token Prediction Head on top of Med-BERT
by: He, Jianping, et al.
Published: (2025)
by: He, Jianping, et al.
Published: (2025)
Sycophancy Hides Linearly in the Attention Heads
by: Genadi, Rifo, et al.
Published: (2026)
by: Genadi, Rifo, et al.
Published: (2026)
NeoBERT: A Next-Generation BERT
by: Breton, Lola Le, et al.
Published: (2025)
by: Breton, Lola Le, et al.
Published: (2025)
Debiasing LLMs by Masking Unfairness-Driving Attention Heads
by: Han, Tingxu, et al.
Published: (2025)
by: Han, Tingxu, et al.
Published: (2025)
BERT-VBD: Vietnamese Multi-Document Summarization Framework
by: Vuong, Tuan-Cuong, et al.
Published: (2024)
by: Vuong, Tuan-Cuong, et al.
Published: (2024)
In-Context Learning in Speech Language Models: Analyzing the Role of Acoustic Features, Linguistic Structure, and Induction Heads
by: Pouw, Charlotte, et al.
Published: (2026)
by: Pouw, Charlotte, et al.
Published: (2026)
CA-BERT: Leveraging Context Awareness for Enhanced Multi-Turn Chat Interaction
by: Liu, Minghao, et al.
Published: (2024)
by: Liu, Minghao, et al.
Published: (2024)
NoVo: Norm Voting off Hallucinations with Attention Heads in Large Language Models
by: Ho, Zheng Yi, et al.
Published: (2024)
by: Ho, Zheng Yi, et al.
Published: (2024)
Focusing on Language: Revealing and Exploiting Language Attention Heads in Multilingual Large Language Models
by: Liu, Xin, et al.
Published: (2025)
by: Liu, Xin, et al.
Published: (2025)
IM-BERT: Enhancing Robustness of BERT through the Implicit Euler Method
by: Kim, Mihyeon, et al.
Published: (2025)
by: Kim, Mihyeon, et al.
Published: (2025)
Multi-BERT: Leveraging Adapters and Prompt Tuning for Low-Resource Multi-Domain Adaptation
by: Azad, Parham Abed, et al.
Published: (2024)
by: Azad, Parham Abed, et al.
Published: (2024)
Optimized Biomedical Question-Answering Services with LLM and Multi-BERT Integration
by: Qian, Cheng, et al.
Published: (2024)
by: Qian, Cheng, et al.
Published: (2024)
Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs
by: Ji, Tao, et al.
Published: (2025)
by: Ji, Tao, et al.
Published: (2025)
Falcon Mamba: The First Competitive Attention-free 7B Language Model
by: Zuo, Jingwei, et al.
Published: (2024)
by: Zuo, Jingwei, et al.
Published: (2024)
Multi-Head Attention Is a Multi-Player Game
by: Chakrabarti, Kushal, et al.
Published: (2026)
by: Chakrabarti, Kushal, et al.
Published: (2026)
Ensemble BERT: A student social network text sentiment classification model based on ensemble learning and BERT architecture
by: Jiang, Kai, et al.
Published: (2024)
by: Jiang, Kai, et al.
Published: (2024)
Attention Mechanism and Heuristic Approach: Context-Aware File Ranking Using Multi-Head Self-Attention
by: Sharma, Pradeep Kumar, et al.
Published: (2026)
by: Sharma, Pradeep Kumar, et al.
Published: (2026)
Chinese ModernBERT with Whole-Word Masking
by: Zhao, Zeyu, et al.
Published: (2025)
by: Zhao, Zeyu, et al.
Published: (2025)
RooseBERT: A New Deal For Political Language Modelling
by: Dore, Deborah, et al.
Published: (2025)
by: Dore, Deborah, et al.
Published: (2025)
Multi-RADS Synthetic Radiology Report Dataset and Head-to-Head Benchmarking of 41 Open-Weight and Proprietary Language Models
by: Bose, Kartik, et al.
Published: (2026)
by: Bose, Kartik, et al.
Published: (2026)
Bangla MedER: Multi-BERT Ensemble Approach for the Recognition of Bangla Medical Entity
by: Aurpa, Tanjim Taharat, et al.
Published: (2025)
by: Aurpa, Tanjim Taharat, et al.
Published: (2025)
Enhancing Text Classification with a Novel Multi-Agent Collaboration Framework Leveraging BERT
by: Baban, Hediyeh, et al.
Published: (2025)
by: Baban, Hediyeh, et al.
Published: (2025)
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
by: Chen, Yilong, et al.
Published: (2024)
by: Chen, Yilong, et al.
Published: (2024)
BPDec: Unveiling the Potential of Masked Language Modeling Decoder in BERT pretraining
by: Liang, Wen, et al.
Published: (2024)
by: Liang, Wen, et al.
Published: (2024)
Exploring Variability in Fine-Tuned Models for Text Classification with DistilBERT
by: Lorenzoni, Giuliano, et al.
Published: (2024)
by: Lorenzoni, Giuliano, et al.
Published: (2024)
MEDBERT.de: A Comprehensive German BERT Model for the Medical Domain
by: Bressem, Keno K., et al.
Published: (2023)
by: Bressem, Keno K., et al.
Published: (2023)
Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse Attention
by: Yang, Lijie, et al.
Published: (2025)
by: Yang, Lijie, et al.
Published: (2025)
Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
by: Wang, Haoran, et al.
Published: (2023)
by: Wang, Haoran, et al.
Published: (2023)
On the Role of Attention Heads in Large Language Model Safety
by: Zhou, Zhenhong, et al.
Published: (2024)
by: Zhou, Zhenhong, et al.
Published: (2024)
Which Attention Heads Matter for In-Context Learning?
by: Yin, Kayo, et al.
Published: (2025)
by: Yin, Kayo, et al.
Published: (2025)
AraPoemBERT: A Pretrained Language Model for Arabic Poetry Analysis
by: Qarah, Faisal
Published: (2024)
by: Qarah, Faisal
Published: (2024)
Similar Items
-
Latent Multi-Head Attention for Small Language Models
by: Mehta, Sushant, et al.
Published: (2025) -
Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models
by: Mąka, Paweł, et al.
Published: (2024) -
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
by: Lu, Yi, et al.
Published: (2024) -
Analyzing Narrative Processing in Large Language Models (LLMs): Using GPT4 to test BERT
by: Krauss, Patrick, et al.
Published: (2024) -
Analyzing Gender Polarity in Short Social Media Texts with BERT: The Role of Emojis and Emoticons
by: Jazi, Saba Yousefian, et al.
Published: (2024)