Guardado en:
| Autores principales: | Estève, Louis, Servan, Christophe, Lavergne, Thomas, Savary, Agata |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2602.22014 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Formalising lexical and syntactic diversity for data sampling in French
por: Estève, Louis, et al.
Publicado: (2025)
por: Estève, Louis, et al.
Publicado: (2025)
Patent Language Model Pretraining with ModernBERT
por: Yousefiramandi, Amirhossein, et al.
Publicado: (2025)
por: Yousefiramandi, Amirhossein, et al.
Publicado: (2025)
Chinese ModernBERT with Whole-Word Masking
por: Zhao, Zeyu, et al.
Publicado: (2025)
por: Zhao, Zeyu, et al.
Publicado: (2025)
TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish
por: Türker, Melikşah, et al.
Publicado: (2025)
por: Türker, Melikşah, et al.
Publicado: (2025)
A survey of diversity quantification in natural language processing: The why, what, where and how
por: Estève, Louis, et al.
Publicado: (2025)
por: Estève, Louis, et al.
Publicado: (2025)
ModernBERT + ColBERT: Enhancing biomedical RAG through an advanced re-ranking retriever
por: Rivera, Eduardo Martínez, et al.
Publicado: (2025)
por: Rivera, Eduardo Martínez, et al.
Publicado: (2025)
NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus
por: Silva, Enzo S. N., et al.
Publicado: (2026)
por: Silva, Enzo S. N., et al.
Publicado: (2026)
Clinical ModernBERT: An efficient and long context encoder for biomedical text
por: Lee, Simon A., et al.
Publicado: (2025)
por: Lee, Simon A., et al.
Publicado: (2025)
ModernBERT is More Efficient than Conventional BERT for Chest CT Findings Classification in Japanese Radiology Reports
por: Yamagishi, Yosuke, et al.
Publicado: (2025)
por: Yamagishi, Yosuke, et al.
Publicado: (2025)
BioClinical ModernBERT: A State-of-the-Art Long-Context Encoder for Biomedical and Clinical NLP
por: Sounack, Thomas, et al.
Publicado: (2025)
por: Sounack, Thomas, et al.
Publicado: (2025)
Pretraining Finnish ModernBERTs
por: Reunamo, Akseli, et al.
Publicado: (2025)
por: Reunamo, Akseli, et al.
Publicado: (2025)
llm-jp-modernbert: A ModernBERT Model Trained on a Large-Scale Japanese Corpus with Long Context Length
por: Sugiura, Issa, et al.
Publicado: (2025)
por: Sugiura, Issa, et al.
Publicado: (2025)
A Benchmark Evaluation of Clinical Named Entity Recognition in French
por: Bannour, Nesrine, et al.
Publicado: (2024)
por: Bannour, Nesrine, et al.
Publicado: (2024)
ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance
por: Antoun, Wissam, et al.
Publicado: (2025)
por: Antoun, Wissam, et al.
Publicado: (2025)
Spatial ModernBERT: Spatial-Aware Transformer for Table and Key-Value Extraction in Financial Documents at Scale
por: Javis AI Team, et al.
Publicado: (2025)
por: Javis AI Team, et al.
Publicado: (2025)
New Semantic Task for the French Spoken Language Understanding MEDIA Benchmark
por: Alavoine, Nadège, et al.
Publicado: (2024)
por: Alavoine, Nadège, et al.
Publicado: (2024)
LLM-based Atomic Propositions help weak extractors: Evaluation of a Propositioner for triplet extraction
por: Pommeret, Luc, et al.
Publicado: (2026)
por: Pommeret, Luc, et al.
Publicado: (2026)
Leveraging Information Retrieval to Enhance Spoken Language Understanding Prompts in Few-Shot Learning
por: Lepagnol, Pierre, et al.
Publicado: (2025)
por: Lepagnol, Pierre, et al.
Publicado: (2025)
CamemBERT 2.0: A Smarter French Language Model Aged to Perfection
por: Antoun, Wissam, et al.
Publicado: (2024)
por: Antoun, Wissam, et al.
Publicado: (2024)
FRASIMED: a Clinical French Annotated Resource Produced through Crosslingual BERT-Based Annotation Projection
por: Zaghir, Jamil, et al.
Publicado: (2023)
por: Zaghir, Jamil, et al.
Publicado: (2023)
mALBERT: Is a Compact Multilingual BERT Model Still Worth It?
por: Servan, Christophe, et al.
Publicado: (2024)
por: Servan, Christophe, et al.
Publicado: (2024)
An investigation of structures responsible for gender bias in BERT and DistilBERT
por: Leteno, Thibaud, et al.
Publicado: (2024)
por: Leteno, Thibaud, et al.
Publicado: (2024)
How Gender Interacts with Political Values: A Case Study on Czech BERT Models
por: Ali, Adnan Al, et al.
Publicado: (2024)
por: Ali, Adnan Al, et al.
Publicado: (2024)
Breaking MLPerf Training: A Case Study on Optimizing BERT
por: Kim, Yongdeok, et al.
Publicado: (2024)
por: Kim, Yongdeok, et al.
Publicado: (2024)
m3BERT: A Modern, Multi-lingual, Matryoshka Bidirectional Encoder
por: Wang, Yaoxiang, et al.
Publicado: (2026)
por: Wang, Yaoxiang, et al.
Publicado: (2026)
Understanding the Interplay of Scale, Data, and Bias in Language Models: A Case Study with BERT
por: Ali, Muhammad, et al.
Publicado: (2024)
por: Ali, Muhammad, et al.
Publicado: (2024)
Construction Identification and Disambiguation Using BERT: A Case Study of NPN
por: Scivetti, Wesley, et al.
Publicado: (2025)
por: Scivetti, Wesley, et al.
Publicado: (2025)
A Dataset for Pharmacovigilance in German, French, and Japanese: Annotating Adverse Drug Reactions across Languages
por: Raithel, Lisa, et al.
Publicado: (2024)
por: Raithel, Lisa, et al.
Publicado: (2024)
AraModernBERT: Transtokenized Initialization and Long-Context Encoder Modeling for Arabic
por: Elshehy, Omar, et al.
Publicado: (2026)
por: Elshehy, Omar, et al.
Publicado: (2026)
CamemBERT-bio: Leveraging Continual Pre-training for Cost-Effective Models on French Biomedical Data
por: Touchent, Rian, et al.
Publicado: (2023)
por: Touchent, Rian, et al.
Publicado: (2023)
mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
por: Marone, Marc, et al.
Publicado: (2025)
por: Marone, Marc, et al.
Publicado: (2025)
Low-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-Study
por: Khan, Eeham, et al.
Publicado: (2025)
por: Khan, Eeham, et al.
Publicado: (2025)
NeoBERT: A Next-Generation BERT
por: Breton, Lola Le, et al.
Publicado: (2025)
por: Breton, Lola Le, et al.
Publicado: (2025)
Emotionally Aware Moderation: The Potential of Emotion Monitoring in Shaping Healthier Social Media Conversations
por: Su, Xiaotian, et al.
Publicado: (2025)
por: Su, Xiaotian, et al.
Publicado: (2025)
Modeling the Construction of a Literary Archetype: The Case of the Detective Figure in French Literature
por: Barré, Jean, et al.
Publicado: (2025)
por: Barré, Jean, et al.
Publicado: (2025)
From BERT to T5: A Study of Named Entity Recognition
por: Jia, Mei
Publicado: (2026)
por: Jia, Mei
Publicado: (2026)
ConfliBERT: A Language Model for Political Conflict
por: Brandt, Patrick T., et al.
Publicado: (2024)
por: Brandt, Patrick T., et al.
Publicado: (2024)
Histoires Morales: A French Dataset for Assessing Moral Alignment
por: Leteno, Thibaud, et al.
Publicado: (2025)
por: Leteno, Thibaud, et al.
Publicado: (2025)
SpikeBERT: A Language Spikformer Learned from BERT with Knowledge Distillation
por: Lv, Changze, et al.
Publicado: (2023)
por: Lv, Changze, et al.
Publicado: (2023)
ColBERT-XM: A Modular Multi-Vector Representation Model for Zero-Shot Multilingual Information Retrieval
por: Louis, Antoine, et al.
Publicado: (2024)
por: Louis, Antoine, et al.
Publicado: (2024)
Ejemplares similares
-
Formalising lexical and syntactic diversity for data sampling in French
por: Estève, Louis, et al.
Publicado: (2025) -
Patent Language Model Pretraining with ModernBERT
por: Yousefiramandi, Amirhossein, et al.
Publicado: (2025) -
Chinese ModernBERT with Whole-Word Masking
por: Zhao, Zeyu, et al.
Publicado: (2025) -
TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish
por: Türker, Melikşah, et al.
Publicado: (2025) -
A survey of diversity quantification in natural language processing: The why, what, where and how
por: Estève, Louis, et al.
Publicado: (2025)