Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content
Fuente:
arXiv
Saved in:
| Main Authors: | Touchent, Rian, Godey, Nathan, de la Clergerie, Eric |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Causal Language Modeling Detour Improves Encoder Continued Pretraining
by: Touchent, Rian, et al.
Published: (2026)
by: Touchent, Rian, et al.
Published: (2026)
CamemBERT-bio: Leveraging Continual Pre-training for Cost-Effective Models on French Biomedical Data
by: Touchent, Rian, et al.
Published: (2023)
by: Touchent, Rian, et al.
Published: (2023)
Gaperon: A Peppered English-French Generative Language Model Suite
by: Godey, Nathan, et al.
Published: (2025)
by: Godey, Nathan, et al.
Published: (2025)
Anisotropy Is Inherent to Self-Attention in Transformers
by: Godey, Nathan, et al.
Published: (2024)
by: Godey, Nathan, et al.
Published: (2024)
Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck
by: Godey, Nathan, et al.
Published: (2024)
by: Godey, Nathan, et al.
Published: (2024)
On the Scaling Laws of Geographical Representation in Language Models
by: Godey, Nathan, et al.
Published: (2024)
by: Godey, Nathan, et al.
Published: (2024)
CamemBERT 2.0: A Smarter French Language Model Aged to Perfection
by: Antoun, Wissam, et al.
Published: (2024)
by: Antoun, Wissam, et al.
Published: (2024)
BiomedSQL: Text-to-SQL for Scientific Reasoning on Biomedical Knowledge Bases
by: Koretsky, Mathew J., et al.
Published: (2025)
by: Koretsky, Mathew J., et al.
Published: (2025)
Patent Representation Learning via Self-supervision
by: Zuo, You, et al.
Published: (2025)
by: Zuo, You, et al.
Published: (2025)
Towards Enriched Controllability for Educational Question Generation
by: Leite, Bernardo, et al.
Published: (2023)
by: Leite, Bernardo, et al.
Published: (2023)
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
by: Mendu, Sai Krishna, et al.
Published: (2025)
by: Mendu, Sai Krishna, et al.
Published: (2025)
Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression
by: Godey, Nathan, et al.
Published: (2025)
by: Godey, Nathan, et al.
Published: (2025)
SWE2: SubWord Enriched and Significant Word Emphasized Framework for Hate Speech Detection
by: Mou, Guanyi, et al.
Published: (2024)
by: Mou, Guanyi, et al.
Published: (2024)
AF Adapter: Continual Pretraining for Building Chinese Biomedical Language Model
by: Yan, Yongyu, et al.
Published: (2022)
by: Yan, Yongyu, et al.
Published: (2022)
FinGPT: Enhancing Sentiment-Based Stock Movement Prediction with Dissemination-Aware and Context-Enriched LLMs
by: Liang, Yixuan, et al.
Published: (2024)
by: Liang, Yixuan, et al.
Published: (2024)
BioPars: A Pretrained Biomedical Large Language Model for Persian Biomedical Text Mining
by: Merzah, Baqer M., et al.
Published: (2025)
by: Merzah, Baqer M., et al.
Published: (2025)
BioCoref: Benchmarking Biomedical Coreference Resolution with LLMs
by: Salem, Nourah M, et al.
Published: (2025)
by: Salem, Nourah M, et al.
Published: (2025)
Empowering Small-Scale Knowledge Graphs: A Strategy of Leveraging General-Purpose Knowledge Graphs for Enriched Embeddings
by: Sawczyn, Albert, et al.
Published: (2024)
by: Sawczyn, Albert, et al.
Published: (2024)
Mining Hidden Thoughts from Texts: Evaluating Continual Pretraining with Synthetic Data for LLM Reasoning
by: Ishibashi, Yoichi, et al.
Published: (2025)
by: Ishibashi, Yoichi, et al.
Published: (2025)
Enriching language models with graph-based context information to better understand textual data
by: Roethel, Albert, et al.
Published: (2023)
by: Roethel, Albert, et al.
Published: (2023)
Decompose, Enrich, and Extract! Schema-aware Event Extraction using LLMs
by: Shiri, Fatemeh, et al.
Published: (2024)
by: Shiri, Fatemeh, et al.
Published: (2024)
BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language Associations
by: Pei, Qizhi, et al.
Published: (2023)
by: Pei, Qizhi, et al.
Published: (2023)
Lost in Backpropagation: The LM Head is a Gradient Bottleneck
by: Godey, Nathan, et al.
Published: (2026)
by: Godey, Nathan, et al.
Published: (2026)
Unknown Unknowns: Why Hidden Intentions in LLMs Evade Detection
by: Srivastav, Devansh, et al.
Published: (2026)
by: Srivastav, Devansh, et al.
Published: (2026)
Compressing LLMs: The Truth is Rarely Pure and Never Simple
by: Jaiswal, Ajay, et al.
Published: (2023)
by: Jaiswal, Ajay, et al.
Published: (2023)
Injecting Structured Biomedical Knowledge into Language Models: Continual Pretraining vs. GraphRAG
by: Klila, Jaafer, et al.
Published: (2026)
by: Klila, Jaafer, et al.
Published: (2026)
MedDec: A Dataset for Extracting Medical Decisions from Discharge Summaries
by: Elgaar, Mohamed, et al.
Published: (2024)
by: Elgaar, Mohamed, et al.
Published: (2024)
RichSpace: Enriching Text-to-Video Prompt Space via Text Embedding Interpolation
by: Cao, Yuefan, et al.
Published: (2025)
by: Cao, Yuefan, et al.
Published: (2025)
Enriching GNNs with Text Contextual Representations for Detecting Disinformation Campaigns on Social Media
by: da Silva, Bruno Croso Cunha, et al.
Published: (2024)
by: da Silva, Bruno Croso Cunha, et al.
Published: (2024)
ProxSparse: Regularized Learning of Semi-Structured Sparsity Masks for Pretrained LLMs
by: Liu, Hongyi, et al.
Published: (2025)
by: Liu, Hongyi, et al.
Published: (2025)
Implicit Identity Technologies for LLMs: Fingerprinting and Watermarking across Datasets, Models, and Generated Content
by: Liu, Bing, et al.
Published: (2026)
by: Liu, Bing, et al.
Published: (2026)
EnrichIndex: Using LLMs to Enrich Retrieval Indices Offline
by: Chen, Peter Baile, et al.
Published: (2025)
by: Chen, Peter Baile, et al.
Published: (2025)
LLMs are not Zero-Shot Reasoners for Biomedical Information Extraction
by: Nagar, Aishik, et al.
Published: (2024)
by: Nagar, Aishik, et al.
Published: (2024)
Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs
by: Hu, Zhiyuan, et al.
Published: (2026)
by: Hu, Zhiyuan, et al.
Published: (2026)
Extracting Protein-Protein Interactions (PPIs) from Biomedical Literature using Attention-based Relational Context Information
by: Park, Gilchan, et al.
Published: (2024)
by: Park, Gilchan, et al.
Published: (2024)
A Large-Scale Vision-Language Dataset Derived from Open Scientific Literature to Advance Biomedical Generalist AI
by: Lozano, Alejandro, et al.
Published: (2025)
by: Lozano, Alejandro, et al.
Published: (2025)
Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs
by: Mathew, Yohan, et al.
Published: (2024)
by: Mathew, Yohan, et al.
Published: (2024)
Can GRPO Help LLMs Transcend Their Pretraining Origin?
by: Ni, Kangqi, et al.
Published: (2025)
by: Ni, Kangqi, et al.
Published: (2025)
Instruct-Tuning Pretrained Causal Language Models for Ancient Greek Papyrology and Epigraphy
by: Cullhed, Eric
Published: (2024)
by: Cullhed, Eric
Published: (2024)
Extracting Unlearned Information from LLMs with Activation Steering
by: Seyitoğlu, Atakan, et al.
Published: (2024)
by: Seyitoğlu, Atakan, et al.
Published: (2024)
Similar Items
-
A Causal Language Modeling Detour Improves Encoder Continued Pretraining
by: Touchent, Rian, et al.
Published: (2026) -
CamemBERT-bio: Leveraging Continual Pre-training for Cost-Effective Models on French Biomedical Data
by: Touchent, Rian, et al.
Published: (2023) -
Gaperon: A Peppered English-French Generative Language Model Suite
by: Godey, Nathan, et al.
Published: (2025) -
Anisotropy Is Inherent to Self-Attention in Transformers
by: Godey, Nathan, et al.
Published: (2024) -
Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck
by: Godey, Nathan, et al.
Published: (2024)