ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance
Fuente:
arXiv
Saved in:
| Main Authors: | Antoun, Wissam, Sagot, Benoît, Seddah, Djamé |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Text to Source: Results in Detecting Large Language Model-Generated Content
by: Antoun, Wissam, et al.
Published: (2023)
by: Antoun, Wissam, et al.
Published: (2023)
Language-Switching Triggers Take a Latent Detour Through Language Models
by: Kulumba, Francis, et al.
Published: (2026)
by: Kulumba, Francis, et al.
Published: (2026)
CamemBERT 2.0: A Smarter French Language Model Aged to Perfection
by: Antoun, Wissam, et al.
Published: (2024)
by: Antoun, Wissam, et al.
Published: (2024)
Can Character-based Language Models Improve Downstream Task Performance in Low-Resource and Noisy Language Scenarios?
by: Riabi, Arij, et al.
Published: (2021)
by: Riabi, Arij, et al.
Published: (2021)
Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models
by: Lasnier, Théo, et al.
Published: (2026)
by: Lasnier, Théo, et al.
Published: (2026)
Gaperon: A Peppered English-French Generative Language Model Suite
by: Godey, Nathan, et al.
Published: (2025)
by: Godey, Nathan, et al.
Published: (2025)
Beyond Dataset Creation: Critical View of Annotation Variation and Bias Probing of a Dataset for Online Radical Content Detection
by: Riabi, Arij, et al.
Published: (2024)
by: Riabi, Arij, et al.
Published: (2024)
Disentangling meaning from language in LLM-based machine translation
by: Lasnier, Théo, et al.
Published: (2026)
by: Lasnier, Théo, et al.
Published: (2026)
Patent Language Model Pretraining with ModernBERT
by: Yousefiramandi, Amirhossein, et al.
Published: (2025)
by: Yousefiramandi, Amirhossein, et al.
Published: (2025)
Chinese ModernBERT with Whole-Word Masking
by: Zhao, Zeyu, et al.
Published: (2025)
by: Zhao, Zeyu, et al.
Published: (2025)
Enriching the NArabizi Treebank: A Multifaceted Approach to Supporting an Under-Resourced Language
by: Riabi, Arij, et al.
Published: (2023)
by: Riabi, Arij, et al.
Published: (2023)
BioClinical ModernBERT: A State-of-the-Art Long-Context Encoder for Biomedical and Clinical NLP
by: Sounack, Thomas, et al.
Published: (2025)
by: Sounack, Thomas, et al.
Published: (2025)
TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish
by: Türker, Melikşah, et al.
Published: (2025)
by: Türker, Melikşah, et al.
Published: (2025)
When Tables Go Crazy: Evaluating Multimodal Models on French Financial Documents
by: Mouilleron, Virginie, et al.
Published: (2026)
by: Mouilleron, Virginie, et al.
Published: (2026)
Common Ground, Diverse Roots: The Difficulty of Classifying Common Examples in Spanish Varieties
by: Lopetegui, Javier A., et al.
Published: (2024)
by: Lopetegui, Javier A., et al.
Published: (2024)
Spatial ModernBERT: Spatial-Aware Transformer for Table and Key-Value Extraction in Financial Documents at Scale
by: Javis AI Team, et al.
Published: (2025)
by: Javis AI Team, et al.
Published: (2025)
ModernBERT + ColBERT: Enhancing biomedical RAG through an advanced re-ranking retriever
by: Rivera, Eduardo Martínez, et al.
Published: (2025)
by: Rivera, Eduardo Martínez, et al.
Published: (2025)
A Diversity Diet for a Healthier Model: A Case Study of French ModernBERT
by: Estève, Louis, et al.
Published: (2026)
by: Estève, Louis, et al.
Published: (2026)
ModernBERT is More Efficient than Conventional BERT for Chest CT Findings Classification in Japanese Radiology Reports
by: Yamagishi, Yosuke, et al.
Published: (2025)
by: Yamagishi, Yosuke, et al.
Published: (2025)
Clinical ModernBERT: An efficient and long context encoder for biomedical text
by: Lee, Simon A., et al.
Published: (2025)
by: Lee, Simon A., et al.
Published: (2025)
Cloaked Classifiers: Pseudonymization Strategies on Sensitive Classification Tasks
by: Riabi, Arij, et al.
Published: (2024)
by: Riabi, Arij, et al.
Published: (2024)
NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus
by: Silva, Enzo S. N., et al.
Published: (2026)
by: Silva, Enzo S. N., et al.
Published: (2026)
Pretraining Finnish ModernBERTs
by: Reunamo, Akseli, et al.
Published: (2025)
by: Reunamo, Akseli, et al.
Published: (2025)
Rethinking the Multilingual Reasoning Gap with Layer Swap
by: Lasbordes, Maxence, et al.
Published: (2026)
by: Lasbordes, Maxence, et al.
Published: (2026)
llm-jp-modernbert: A ModernBERT Model Trained on a Large-Scale Japanese Corpus with Long Context Length
by: Sugiura, Issa, et al.
Published: (2025)
by: Sugiura, Issa, et al.
Published: (2025)
m3BERT: A Modern, Multi-lingual, Matryoshka Bidirectional Encoder
by: Wang, Yaoxiang, et al.
Published: (2026)
by: Wang, Yaoxiang, et al.
Published: (2026)
Performance Evaluation of Emotion Classification in Japanese Using RoBERTa and DeBERTa
by: Takenaka, Yoichi
Published: (2025)
by: Takenaka, Yoichi
Published: (2025)
Multilingual, Multimodal Pipeline for Creating Authentic and Structured Fact-Checked Claim Dataset
by: Hüsünbeyi, Z. Melce, et al.
Published: (2026)
by: Hüsünbeyi, Z. Melce, et al.
Published: (2026)
LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens
by: Zebaze, Armel, et al.
Published: (2025)
by: Zebaze, Armel, et al.
Published: (2025)
TopXGen: Topic-Diverse Parallel Data Generation for Low-Resource Machine Translation
by: Zebaze, Armel, et al.
Published: (2025)
by: Zebaze, Armel, et al.
Published: (2025)
AraModernBERT: Transtokenized Initialization and Long-Context Encoder Modeling for Arabic
by: Elshehy, Omar, et al.
Published: (2026)
by: Elshehy, Omar, et al.
Published: (2026)
Anisotropy Is Inherent to Self-Attention in Transformers
by: Godey, Nathan, et al.
Published: (2024)
by: Godey, Nathan, et al.
Published: (2024)
Does RoBERTa Perform Better than BERT in Continual Learning: An Attention Sink Perspective
by: Bai, Xueying, et al.
Published: (2024)
by: Bai, Xueying, et al.
Published: (2024)
mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
by: Marone, Marc, et al.
Published: (2025)
by: Marone, Marc, et al.
Published: (2025)
HALvest-Contrastive: Retrieval-Like Authorship Attribution with Patch-Level Late Interaction
by: Kulumba, Francis, et al.
Published: (2024)
by: Kulumba, Francis, et al.
Published: (2024)
A French Version of the OLDI Seed Corpus
by: Marmonier, Malik, et al.
Published: (2025)
by: Marmonier, Malik, et al.
Published: (2025)
Explicit Learning and the LLM in Machine Translation
by: Marmonier, Malik, et al.
Published: (2025)
by: Marmonier, Malik, et al.
Published: (2025)
Compositional Translation: A Novel LLM-based Approach for Low-resource Machine Translation
by: Zebaze, Armel, et al.
Published: (2025)
by: Zebaze, Armel, et al.
Published: (2025)
Tree of Problems: Improving structured problem solving with compositionality
by: Zebaze, Armel, et al.
Published: (2024)
by: Zebaze, Armel, et al.
Published: (2024)
Testing the Deliteralization Hypothesis in Human and Machine Translation
by: Marmonier, Malik, et al.
Published: (2026)
by: Marmonier, Malik, et al.
Published: (2026)
Similar Items
-
From Text to Source: Results in Detecting Large Language Model-Generated Content
by: Antoun, Wissam, et al.
Published: (2023) -
Language-Switching Triggers Take a Latent Detour Through Language Models
by: Kulumba, Francis, et al.
Published: (2026) -
CamemBERT 2.0: A Smarter French Language Model Aged to Perfection
by: Antoun, Wissam, et al.
Published: (2024) -
Can Character-based Language Models Improve Downstream Task Performance in Low-Resource and Noisy Language Scenarios?
by: Riabi, Arij, et al.
Published: (2021) -
Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models
by: Lasnier, Théo, et al.
Published: (2026)