Patent Language Model Pretraining with ModernBERT
Fuente:
arXiv
Saved in:
| Main Authors: | Yousefiramandi, Amirhossein, Cooney, Ciaran |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fine-Tuning Causal LLMs for Text Classification: Embedding-Based vs. Instruction-Based Approaches
by: Yousefiramandi, Amirhossein, et al.
Published: (2025)
by: Yousefiramandi, Amirhossein, et al.
Published: (2025)
Clinical ModernBERT: An efficient and long context encoder for biomedical text
by: Lee, Simon A., et al.
Published: (2025)
by: Lee, Simon A., et al.
Published: (2025)
Cumulative-Goodness Free-Riding in Forward-Forward Networks: Real, Repairable, but Not Accuracy-Dominant
by: Yousefiramandi, Amirhossein
Published: (2026)
by: Yousefiramandi, Amirhossein
Published: (2026)
Chinese ModernBERT with Whole-Word Masking
by: Zhao, Zeyu, et al.
Published: (2025)
by: Zhao, Zeyu, et al.
Published: (2025)
NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus
by: Silva, Enzo S. N., et al.
Published: (2026)
by: Silva, Enzo S. N., et al.
Published: (2026)
BioClinical ModernBERT: A State-of-the-Art Long-Context Encoder for Biomedical and Clinical NLP
by: Sounack, Thomas, et al.
Published: (2025)
by: Sounack, Thomas, et al.
Published: (2025)
MrBERT: Modern Multilingual Encoders via Vocabulary, Domain, and Dimensional Adaptation
by: Tamayo, Daniel, et al.
Published: (2026)
by: Tamayo, Daniel, et al.
Published: (2026)
Pretraining Large Language Models with NVFP4
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
$\texttt{PatentAgent}$: Intelligent Agent for Automated Pharmaceutical Patent Analysis
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Revisiting Multilingual Data Mixtures in Language Model Pretraining
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
In-context Pretraining: Language Modeling Beyond Document Boundaries
by: Shi, Weijia, et al.
Published: (2023)
by: Shi, Weijia, et al.
Published: (2023)
Discovering Knowledge-Critical Subnetworks in Pretrained Language Models
by: Bayazit, Deniz, et al.
Published: (2023)
by: Bayazit, Deniz, et al.
Published: (2023)
Harnessing Large Language Models: Fine-tuned BERT for Detecting Charismatic Leadership Tactics in Natural Language
by: Saeid, Yasser, et al.
Published: (2024)
by: Saeid, Yasser, et al.
Published: (2024)
Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
by: McLeish, Sean, et al.
Published: (2025)
by: McLeish, Sean, et al.
Published: (2025)
BERT-JEPA: Reorganizing CLS Embeddings for Language-Invariant Semantics
by: Gillin, Taj, et al.
Published: (2026)
by: Gillin, Taj, et al.
Published: (2026)
Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models
by: Huang, Yukun, et al.
Published: (2025)
by: Huang, Yukun, et al.
Published: (2025)
Tracing the Representation Geometry of Language Models from Pretraining to Post-training
by: Li, Melody Zixuan, et al.
Published: (2025)
by: Li, Melody Zixuan, et al.
Published: (2025)
Generalizable and Stable Finetuning of Pretrained Language Models on Low-Resource Texts
by: Somayajula, Sai Ashish, et al.
Published: (2024)
by: Somayajula, Sai Ashish, et al.
Published: (2024)
Exploring the Reversal Curse and Other Deductive Logical Reasoning in BERT and GPT-Based Large Language Models
by: Wu, Da, et al.
Published: (2023)
by: Wu, Da, et al.
Published: (2023)
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
by: Ali, Mehdi, et al.
Published: (2025)
by: Ali, Mehdi, et al.
Published: (2025)
MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining
by: Xiaomi, LLM-Core, et al.
Published: (2025)
by: Xiaomi, LLM-Core, et al.
Published: (2025)
IDEA Prune: An Integrated Enlarge-and-Prune Pipeline in Generative Language Model Pretraining
by: Li, Yixiao, et al.
Published: (2025)
by: Li, Yixiao, et al.
Published: (2025)
Pretrained Generative Language Models as General Learning Frameworks for Sequence-Based Tasks
by: Fauber, Ben
Published: (2024)
by: Fauber, Ben
Published: (2024)
Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data
by: Wang, Xinyi, et al.
Published: (2024)
by: Wang, Xinyi, et al.
Published: (2024)
Instruct-Tuning Pretrained Causal Language Models for Ancient Greek Papyrology and Epigraphy
by: Cullhed, Eric
Published: (2024)
by: Cullhed, Eric
Published: (2024)
BiMix: A Bivariate Data Mixing Law for Language Model Pretraining
by: Ge, Ce, et al.
Published: (2024)
by: Ge, Ce, et al.
Published: (2024)
Scaling BERT Models for Turkish Automatic Punctuation and Capitalization Correction
by: Saoud, Abdulkader, et al.
Published: (2024)
by: Saoud, Abdulkader, et al.
Published: (2024)
Cancer Diagnosis Categorization in Electronic Health Records Using Large Language Models and BioBERT: Model Performance Evaluation Study
by: Hashtarkhani, Soheil, et al.
Published: (2025)
by: Hashtarkhani, Soheil, et al.
Published: (2025)
MedicalBERT: enhancing biomedical natural language processing using pretrained BERT-based model
by: Reddy, K. Sahit, et al.
Published: (2025)
by: Reddy, K. Sahit, et al.
Published: (2025)
HMI: Hierarchical Knowledge Management for Efficient Multi-Tenant Inference in Pretrained Language Models
by: Zhang, Jun, et al.
Published: (2025)
by: Zhang, Jun, et al.
Published: (2025)
Injecting Structured Biomedical Knowledge into Language Models: Continual Pretraining vs. GraphRAG
by: Klila, Jaafer, et al.
Published: (2026)
by: Klila, Jaafer, et al.
Published: (2026)
Fine-tuning can Help Detect Pretraining Data from Large Language Models
by: Zhang, Hengxiang, et al.
Published: (2024)
by: Zhang, Hengxiang, et al.
Published: (2024)
Toward Understanding BERT-Like Pre-Training for DNA Foundation Models
by: Liang, Chaoqi, et al.
Published: (2023)
by: Liang, Chaoqi, et al.
Published: (2023)
BioPars: A Pretrained Biomedical Large Language Model for Persian Biomedical Text Mining
by: Merzah, Baqer M., et al.
Published: (2025)
by: Merzah, Baqer M., et al.
Published: (2025)
BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains
by: Labrak, Yanis, et al.
Published: (2024)
by: Labrak, Yanis, et al.
Published: (2024)
Ayn: A Tiny yet Competitive Indian Legal Language Model Pretrained from Scratch
by: Niyogi, Mitodru, et al.
Published: (2024)
by: Niyogi, Mitodru, et al.
Published: (2024)
Improving VTE Identification through Language Models from Radiology Reports: A Comparative Study of Mamba, Phi-3 Mini, and BERT
by: Deng, Jamie, et al.
Published: (2024)
by: Deng, Jamie, et al.
Published: (2024)
Pretraining Finnish ModernBERTs
by: Reunamo, Akseli, et al.
Published: (2025)
by: Reunamo, Akseli, et al.
Published: (2025)
BERT-LSH: Reducing Absolute Compute For Attention
by: Li, Zezheng, et al.
Published: (2024)
by: Li, Zezheng, et al.
Published: (2024)
MuRating: A High Quality Data Selecting Approach to Multilingual Large Language Model Pretraining
by: Chen, Zhixun, et al.
Published: (2025)
by: Chen, Zhixun, et al.
Published: (2025)
Similar Items
-
Fine-Tuning Causal LLMs for Text Classification: Embedding-Based vs. Instruction-Based Approaches
by: Yousefiramandi, Amirhossein, et al.
Published: (2025) -
Clinical ModernBERT: An efficient and long context encoder for biomedical text
by: Lee, Simon A., et al.
Published: (2025) -
Cumulative-Goodness Free-Riding in Forward-Forward Networks: Real, Repairable, but Not Accuracy-Dominant
by: Yousefiramandi, Amirhossein
Published: (2026) -
Chinese ModernBERT with Whole-Word Masking
by: Zhao, Zeyu, et al.
Published: (2025) -
NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus
by: Silva, Enzo S. N., et al.
Published: (2026)