On the Importance of Pretraining Data Alignment for Atomic Property Prediction
Fuente:
arXiv
Salvato in:
| Autori principali: | Ghunaim, Yasir, Hammoud, Hasan Abed Al Kader, Ghanem, Bernard |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DiffCLIP: Differential Attention Meets CLIP
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2025)
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2025)
Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2025)
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2025)
Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2025)
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2025)
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection
di: Shairah, Harethah Abu, et al.
Pubblicazione: (2025)
di: Shairah, Harethah Abu, et al.
Pubblicazione: (2025)
On Pretraining Data Diversity for Self-Supervised Learning
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2024)
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2024)
Train Long, Think Short: Curriculum Learning for Efficient Reasoning
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2025)
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2025)
An Embarrassingly Simple Defense Against LLM Abliteration Attacks
di: Shairah, Harethah Abu, et al.
Pubblicazione: (2025)
di: Shairah, Harethah Abu, et al.
Pubblicazione: (2025)
Model Merging and Safety Alignment: One Bad Model Spoils the Bunch
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2024)
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2024)
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2024)
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2024)
AraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language Models
di: Zbeeb, Mohammad, et al.
Pubblicazione: (2025)
di: Zbeeb, Mohammad, et al.
Pubblicazione: (2025)
Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning
di: Alssum, Lama, et al.
Pubblicazione: (2025)
di: Alssum, Lama, et al.
Pubblicazione: (2025)
QuanBench+: A Unified Multi-Framework Benchmark for LLM-Based Quantum Code Generation
di: Slim, Ali, et al.
Pubblicazione: (2026)
di: Slim, Ali, et al.
Pubblicazione: (2026)
TAPS: Task Aware Proposal Distributions for Speculative Sampling
di: Zbib, Mohamad, et al.
Pubblicazione: (2026)
di: Zbib, Mohamad, et al.
Pubblicazione: (2026)
Towards Faster and More Compact Foundation Models for Molecular Property Prediction
di: Ghunaim, Yasir, et al.
Pubblicazione: (2025)
di: Ghunaim, Yasir, et al.
Pubblicazione: (2025)
Large-Scale Knowledge Integration for Enhanced Molecular Property Prediction
di: Ghunaim, Yasir, et al.
Pubblicazione: (2024)
di: Ghunaim, Yasir, et al.
Pubblicazione: (2024)
Forget Less, Retain More: A Lightweight Regularizer for Rehearsal-Based Continual Learning
di: Alssum, Lama, et al.
Pubblicazione: (2025)
di: Alssum, Lama, et al.
Pubblicazione: (2025)
From Categories to Classifiers: Name-Only Continual Learning by Exploring the Web
di: Prabhu, Ameya, et al.
Pubblicazione: (2023)
di: Prabhu, Ameya, et al.
Pubblicazione: (2023)
Reasoning Vectors: Transferring Chain-of-Thought Capabilities via Task Arithmetic
di: Zbeeb, Mohammad, et al.
Pubblicazione: (2025)
di: Zbeeb, Mohammad, et al.
Pubblicazione: (2025)
Supervised Pretraining for Material Property Prediction
di: Rahman, Chowdhury Mohammad Abid, et al.
Pubblicazione: (2025)
di: Rahman, Chowdhury Mohammad Abid, et al.
Pubblicazione: (2025)
Data Assimilation in Chaotic Systems Using Deep Reinforcement Learning
di: Hammoud, Mohamad Abed El Rahman, et al.
Pubblicazione: (2024)
di: Hammoud, Mohamad Abed El Rahman, et al.
Pubblicazione: (2024)
Multi-level Self-supervised Pretraining on Compositional Hierarchical Graph for Molecular Property Prediction
di: Liu, Xiayu, et al.
Pubblicazione: (2026)
di: Liu, Xiayu, et al.
Pubblicazione: (2026)
Two-Stage Pretraining for Molecular Property Prediction in the Wild
di: Wijaya, Kevin Tirta, et al.
Pubblicazione: (2024)
di: Wijaya, Kevin Tirta, et al.
Pubblicazione: (2024)
Benchmarking the Medical Understanding and Reasoning of Large Language Models in Arabic Healthcare Tasks
di: AlDahoul, Nouar, et al.
Pubblicazione: (2025)
di: AlDahoul, Nouar, et al.
Pubblicazione: (2025)
Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models
di: AlDahoul, Nouar, et al.
Pubblicazione: (2025)
di: AlDahoul, Nouar, et al.
Pubblicazione: (2025)
Feature Importance Depends on Properties of the Data: Towards Choosing the Correct Explanations for Your Data and Decision Trees based Models
di: Ayad, Célia Wafa, et al.
Pubblicazione: (2025)
di: Ayad, Célia Wafa, et al.
Pubblicazione: (2025)
Quantifying the Importance of Data Alignment in Downstream Model Performance
di: Chawla, Krrish, et al.
Pubblicazione: (2025)
di: Chawla, Krrish, et al.
Pubblicazione: (2025)
Hierarchical Structure-Property Alignment for Data-Efficient Molecular Generation and Editing
di: Fan, Ziyu, et al.
Pubblicazione: (2025)
di: Fan, Ziyu, et al.
Pubblicazione: (2025)
CANDID DAC: Leveraging Coupled Action Dimensions with Importance Differences in DAC
di: Bordne, Philipp, et al.
Pubblicazione: (2024)
di: Bordne, Philipp, et al.
Pubblicazione: (2024)
Unlock the Potential of Large Language Models for Predictive Tabular Tasks in Data Science with Table-Specific Pretraining
di: Yang, Yazheng, et al.
Pubblicazione: (2024)
di: Yang, Yazheng, et al.
Pubblicazione: (2024)
MMPolymer: A Multimodal Multitask Pretraining Framework for Polymer Property Prediction
di: Wang, Fanmeng, et al.
Pubblicazione: (2024)
di: Wang, Fanmeng, et al.
Pubblicazione: (2024)
Entropy-Regularized Adjoint Matching for Offline Reinforcement Learning
di: Ghanem, Abdelghani, et al.
Pubblicazione: (2026)
di: Ghanem, Abdelghani, et al.
Pubblicazione: (2026)
Label Informed Contrastive Pretraining for Node Importance Estimation on Knowledge Graphs
di: Zhang, Tianyu, et al.
Pubblicazione: (2024)
di: Zhang, Tianyu, et al.
Pubblicazione: (2024)
Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries
di: Mahajan, Divyat, et al.
Pubblicazione: (2025)
di: Mahajan, Divyat, et al.
Pubblicazione: (2025)
A Scalable Pretraining Framework for Link Prediction with Efficient Adaptation
di: Song, Yu, et al.
Pubblicazione: (2025)
di: Song, Yu, et al.
Pubblicazione: (2025)
Clin-JEPA: A Multi-Phase Co-Training Framework for Joint-Embedding Predictive Pretraining on EHR Patient Trajectories
di: Yang, Yixuan, et al.
Pubblicazione: (2026)
di: Yang, Yixuan, et al.
Pubblicazione: (2026)
Gap-K%: Measuring Top-1 Prediction Gap for Detecting Pretraining Data
di: Kwak, Minseo, et al.
Pubblicazione: (2026)
di: Kwak, Minseo, et al.
Pubblicazione: (2026)
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
di: AlDahoul, Nouar, et al.
Pubblicazione: (2025)
di: AlDahoul, Nouar, et al.
Pubblicazione: (2025)
Structural Alignment in Link Prediction
di: Sardina, Jeffrey Seathrún
Pubblicazione: (2025)
di: Sardina, Jeffrey Seathrún
Pubblicazione: (2025)
Cross-Modal Reconstruction Pretraining for Ramp Flow Prediction at Highway Interchanges
di: Li, Yongchao, et al.
Pubblicazione: (2025)
di: Li, Yongchao, et al.
Pubblicazione: (2025)
A Pretrained Probabilistic Transformer for City-Scale Traffic Volume Prediction
di: Shen, Shiyu, et al.
Pubblicazione: (2025)
di: Shen, Shiyu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
DiffCLIP: Differential Attention Meets CLIP
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2025) -
Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2025) -
Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2025) -
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection
di: Shairah, Harethah Abu, et al.
Pubblicazione: (2025) -
On Pretraining Data Diversity for Self-Supervised Learning
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2024)