Does RoBERTa Perform Better than BERT in Continual Learning: An Attention Sink Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Bai, Xueying, Sun, Yifan, Balasubramanian, Niranjan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RoBERTurk: Adjusting RoBERTa for Turkish
by: Tas, Nuri
Published: (2024)
by: Tas, Nuri
Published: (2024)
Continual Learning with Global Alignment
by: Bai, Xueying, et al.
Published: (2022)
by: Bai, Xueying, et al.
Published: (2022)
Multilingual Hope Speech Detection: A Comparative Study of Logistic Regression, mBERT, and XLM-RoBERTa with Active Learning
by: Abiola, T. O., et al.
Published: (2025)
by: Abiola, T. O., et al.
Published: (2025)
Performance Evaluation of Emotion Classification in Japanese Using RoBERTa and DeBERTa
by: Takenaka, Yoichi
Published: (2025)
by: Takenaka, Yoichi
Published: (2025)
NER- RoBERTa: Fine-Tuning RoBERTa for Named Entity Recognition (NER) within low-resource languages
by: Abdullah, Abdulhady Abas, et al.
Published: (2024)
by: Abdullah, Abdulhady Abas, et al.
Published: (2024)
The Role of Model Architecture and Scale in Predicting Molecular Properties: Insights from Fine-Tuning RoBERTa, BART, and LLaMA
by: Youngmin, Lee, et al.
Published: (2024)
by: Youngmin, Lee, et al.
Published: (2024)
SG-UniBuc-NLP at SemEval-2026 Task 6: Multi-Head RoBERTa with Chunking for Long-Context Evasion Detection
by: Stefan, Gabriel, et al.
Published: (2026)
by: Stefan, Gabriel, et al.
Published: (2026)
The Large Language Model GreekLegalRoBERTa
by: Saketos, Vasileios, et al.
Published: (2024)
by: Saketos, Vasileios, et al.
Published: (2024)
Antibody Foundational Model : Ab-RoBERTa
by: Huh, Eunna, et al.
Published: (2025)
by: Huh, Eunna, et al.
Published: (2025)
Multiclass Hate Speech Detection with RoBERTa-OTA: Integrating Transformer Attention and Graph Convolutional Networks
by: Abusaqer, Mahmoud, et al.
Published: (2026)
by: Abusaqer, Mahmoud, et al.
Published: (2026)
Evaluating Simple Debiasing Techniques in RoBERTa-based Hate Speech Detection Models
by: Iftimie, Diana, et al.
Published: (2025)
by: Iftimie, Diana, et al.
Published: (2025)
A RoBERTa-Based Functional Syntax Annotation Model for Chinese Texts
by: Xiaohui, Han, et al.
Published: (2025)
by: Xiaohui, Han, et al.
Published: (2025)
Logits-Constrained Framework with RoBERTa for Ancient Chinese NER
by: Hua, Wenjie, et al.
Published: (2025)
by: Hua, Wenjie, et al.
Published: (2025)
Comparing Pre-trained Human Language Models: Is it Better with Human Context as Groups, Individual Traits, or Both?
by: Soni, Nikita, et al.
Published: (2024)
by: Soni, Nikita, et al.
Published: (2024)
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
by: Peng, Runyu, et al.
Published: (2026)
by: Peng, Runyu, et al.
Published: (2026)
Data Quality Matters: Suicide Intention Detection on Social Media Posts Using RoBERTa-CNN
by: Lin, Emily, et al.
Published: (2024)
by: Lin, Emily, et al.
Published: (2024)
Bilingual Sexism Classification: Fine-Tuned XLM-RoBERTa and GPT-3.5 Few-Shot Learning
by: Azadi, AmirMohammad, et al.
Published: (2024)
by: Azadi, AmirMohammad, et al.
Published: (2024)
EfficientQA : a RoBERTa Based Phrase-Indexed Question-Answering System
by: Chaybouti, Sofian, et al.
Published: (2021)
by: Chaybouti, Sofian, et al.
Published: (2021)
RoBERTa-BiLSTM: A Context-Aware Hybrid Model for Sentiment Analysis
by: Rahman, Md. Mostafizer, et al.
Published: (2024)
by: Rahman, Md. Mostafizer, et al.
Published: (2024)
TartuNLP @ SIGTYP 2024 Shared Task: Adapting XLM-RoBERTa for Ancient and Historical Languages
by: Dorkin, Aleksei, et al.
Published: (2024)
by: Dorkin, Aleksei, et al.
Published: (2024)
Sentiment Analysis Based on RoBERTa for Amazon Review: An Empirical Study on Decision Making
by: Guo, Xinli
Published: (2024)
by: Guo, Xinli
Published: (2024)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
by: Fu, Zizhuo, et al.
Published: (2026)
by: Fu, Zizhuo, et al.
Published: (2026)
Why Antiwork: A RoBERTa-Based System for Work-Related Stress Identification and Leading Factor Analysis
by: Lu, Tao, et al.
Published: (2024)
by: Lu, Tao, et al.
Published: (2024)
On the Existence and Behavior of Secondary Attention Sinks
by: Wong, Jeffrey T. H., et al.
Published: (2025)
by: Wong, Jeffrey T. H., et al.
Published: (2025)
"AGI" team at SHROOM-CAP: Data-Centric Approach to Multilingual Hallucination Detection using XLM-RoBERTa
by: Rathva, Harsh, et al.
Published: (2025)
by: Rathva, Harsh, et al.
Published: (2025)
SemEval-2024 Task 8: Weighted Layer Averaging RoBERTa for Black-Box Machine-Generated Text Detection
by: Datta, Ayan, et al.
Published: (2024)
by: Datta, Ayan, et al.
Published: (2024)
NCL-BU at SemEval-2026 Task 3: Fine-tuning XLM-RoBERTa for Multilingual Dimensional Sentiment Regression
by: Wu, Tong, et al.
Published: (2026)
by: Wu, Tong, et al.
Published: (2026)
Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
by: Shang, Bingqi, et al.
Published: (2025)
by: Shang, Bingqi, et al.
Published: (2025)
Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models
by: Binkowski, Jakub, et al.
Published: (2026)
by: Binkowski, Jakub, et al.
Published: (2026)
Demystifying the Slash Pattern in Attention: The Role of RoPE
by: Cheng, Yuan, et al.
Published: (2026)
by: Cheng, Yuan, et al.
Published: (2026)
Does Biomedical Training Lead to Better Medical Performance?
by: Dada, Amin, et al.
Published: (2024)
by: Dada, Amin, et al.
Published: (2024)
Attention Sinks in Massively Multilingual Neural Machine Translation:Discovery, Analysis, and Mitigation
by: Mutisya, Hillary, et al.
Published: (2026)
by: Mutisya, Hillary, et al.
Published: (2026)
Frayed RoPE and Long Inputs: A Geometric Perspective
by: Wertheimer, Davis, et al.
Published: (2026)
by: Wertheimer, Davis, et al.
Published: (2026)
Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention Calibration
by: Yu, Zhongzhi, et al.
Published: (2024)
by: Yu, Zhongzhi, et al.
Published: (2024)
EDU-NER-2025: Named Entity Recognition in Urdu Educational Texts using XLM-RoBERTa with X (formerly Twitter)
by: Ullah, Fida, et al.
Published: (2025)
by: Ullah, Fida, et al.
Published: (2025)
Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization
by: Son, Seungwoo, et al.
Published: (2024)
by: Son, Seungwoo, et al.
Published: (2024)
QuadAI at SemEval-2026 Task 3: Ensemble Learning of Hybrid RoBERTa and LLMs for Dimensional Aspect-Based Sentiment Analysis
by: de Vink, A. J. W., et al.
Published: (2026)
by: de Vink, A. J. W., et al.
Published: (2026)
Large Human Language Models: A Need and the Challenges
by: Soni, Nikita, et al.
Published: (2023)
by: Soni, Nikita, et al.
Published: (2023)
Can We Use Probing to Better Understand Fine-tuning and Knowledge Distillation of the BERT NLU?
by: Hościłowicz, Jakub, et al.
Published: (2023)
by: Hościłowicz, Jakub, et al.
Published: (2023)
When Attention Sink Emerges in Language Models: An Empirical View
by: Gu, Xiangming, et al.
Published: (2024)
by: Gu, Xiangming, et al.
Published: (2024)
Similar Items
-
RoBERTurk: Adjusting RoBERTa for Turkish
by: Tas, Nuri
Published: (2024) -
Continual Learning with Global Alignment
by: Bai, Xueying, et al.
Published: (2022) -
Multilingual Hope Speech Detection: A Comparative Study of Logistic Regression, mBERT, and XLM-RoBERTa with Active Learning
by: Abiola, T. O., et al.
Published: (2025) -
Performance Evaluation of Emotion Classification in Japanese Using RoBERTa and DeBERTa
by: Takenaka, Yoichi
Published: (2025) -
NER- RoBERTa: Fine-Tuning RoBERTa for Named Entity Recognition (NER) within low-resource languages
by: Abdullah, Abdulhady Abas, et al.
Published: (2024)