Influence-driven Curriculum Learning for Pre-training on Limited Data
Fuente:
arXiv
Saved in:
| Main Authors: | Schoenegger, Loris, Thoma, Lukas, Blevins, Terra, Roth, Benjamin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Compact Example-Based Explanations for Language Models
by: Schoenegger, Loris, et al.
Published: (2026)
by: Schoenegger, Loris, et al.
Published: (2026)
Select or Project? Evaluating Lower-dimensional Vectors for LLM Training Data Explanations
by: Hinterleitner, Lukas, et al.
Published: (2026)
by: Hinterleitner, Lukas, et al.
Published: (2026)
WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training
by: Tian, Changxin, et al.
Published: (2025)
by: Tian, Changxin, et al.
Published: (2025)
Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment
by: Zhang, Yizhuo, et al.
Published: (2025)
by: Zhang, Yizhuo, et al.
Published: (2025)
Pre-trained Models Perform the Best When Token Distributions Follow Zipf's Law
by: He, Yanjin, et al.
Published: (2025)
by: He, Yanjin, et al.
Published: (2025)
BabyLlama-2: Ensemble-Distilled Models Consistently Outperform Teachers With Limited Data
by: Tastet, Jean-Loup, et al.
Published: (2024)
by: Tastet, Jean-Loup, et al.
Published: (2024)
Your Pretrained Model Tells the Difficulty Itself: A Self-Adaptive Curriculum Learning Paradigm for Natural Language Understanding
by: Feng, Qi, et al.
Published: (2025)
by: Feng, Qi, et al.
Published: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
PowLU: An Activation Function for Stable Pre-Training of LLMs
by: Jiang, Peijie, et al.
Published: (2026)
by: Jiang, Peijie, et al.
Published: (2026)
Geometric Deviation as an Unsupervised Pre-Generation Reliability Signal: Probing LLM Representations for Answerability
by: Du, Yucheng
Published: (2026)
by: Du, Yucheng
Published: (2026)
Overcoming Long-Context Limitations of State-Space Models via Context-Dependent Sparse Attention
by: Zhan, Zhihao, et al.
Published: (2025)
by: Zhan, Zhihao, et al.
Published: (2025)
$FastDoc$: Domain-Specific Fast Continual Pre-training Technique using Document-Level Metadata and Taxonomy
by: Nandy, Abhilash, et al.
Published: (2023)
by: Nandy, Abhilash, et al.
Published: (2023)
Improving Influence-based Instruction Tuning Data Selection for Balanced Learning of Diverse Capabilities
by: Dai, Qirun, et al.
Published: (2025)
by: Dai, Qirun, et al.
Published: (2025)
On Initializing Transformers with Pre-trained Embeddings
by: Kim, Ha Young, et al.
Published: (2024)
by: Kim, Ha Young, et al.
Published: (2024)
Reinforcement Learning for Latent-Space Thinking in LLMs
by: Özeren, Enes, et al.
Published: (2025)
by: Özeren, Enes, et al.
Published: (2025)
MALoRA: Mixture of Asymmetric Low-Rank Adaptation for Enhanced Multi-Task Learning
by: Wang, Xujia, et al.
Published: (2024)
by: Wang, Xujia, et al.
Published: (2024)
Large Language Model (LLM) Bias Index -- LLMBI
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
Engineering A Large Language Model From Scratch
by: Oketunji, Abiodun Finbarrs
Published: (2024)
by: Oketunji, Abiodun Finbarrs
Published: (2024)
The Data Efficiency Frontier of Financial Foundation Models: Scaling Laws from Continued Pretraining
by: Ponnock, Jesse
Published: (2025)
by: Ponnock, Jesse
Published: (2025)
Continuous-Depth Transformers with Learned Control Dynamics
by: Jemley, Peter
Published: (2026)
by: Jemley, Peter
Published: (2026)
COA-GPT: Generative Pre-trained Transformers for Accelerated Course of Action Development in Military Operations
by: Goecks, Vinicius G., et al.
Published: (2024)
by: Goecks, Vinicius G., et al.
Published: (2024)
Arabic Hate Speech Identification and Masking in Social Media using Deep Learning Models and Pre-trained Models Fine-tuning
by: Doghmash, Salam Thabet, et al.
Published: (2025)
by: Doghmash, Salam Thabet, et al.
Published: (2025)
Pre-training data selection for biomedical domain adaptation using journal impact metrics
by: Laï-king, Mathieu, et al.
Published: (2024)
by: Laï-king, Mathieu, et al.
Published: (2024)
Annotation Entropy Predicts Per-Example Learning Dynamics in LoRA Fine-Tuning
by: Steele, Brady
Published: (2026)
by: Steele, Brady
Published: (2026)
Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility
by: Szilvasy, Gergely, et al.
Published: (2026)
by: Szilvasy, Gergely, et al.
Published: (2026)
MMSciBench: Benchmarking Language Models on Chinese Multimodal Scientific Problems
by: Ye, Xinwu, et al.
Published: (2025)
by: Ye, Xinwu, et al.
Published: (2025)
Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMs
by: Zhang, Zhaowei, et al.
Published: (2025)
by: Zhang, Zhaowei, et al.
Published: (2025)
Interpreto: An Explainability Library for Transformers
by: Poché, Antonin, et al.
Published: (2025)
by: Poché, Antonin, et al.
Published: (2025)
GenAI Content Detection Task 3: Cross-Domain Machine-Generated Text Detection Challenge
by: Dugan, Liam, et al.
Published: (2025)
by: Dugan, Liam, et al.
Published: (2025)
Reconstructing Syllable Sequences in Abugida Scripts with Incomplete Inputs
by: Thu, Ye Kyaw, et al.
Published: (2025)
by: Thu, Ye Kyaw, et al.
Published: (2025)
Ambiguity in LLMs is a concept missing problem
by: Hu, Zhibo, et al.
Published: (2025)
by: Hu, Zhibo, et al.
Published: (2025)
Constraint-Driven Small Language Models Based on Agent and OpenAlex Knowledge Graph: Mining Conceptual Pathways and Discovering Innovation Points in Academic Papers
by: Xia, Ziye, et al.
Published: (2025)
by: Xia, Ziye, et al.
Published: (2025)
Digital Guardians: Can GPT-4, Perspective API, and Moderation API reliably detect hate speech in reader comments of German online newspapers?
by: Weber, Manuel, et al.
Published: (2025)
by: Weber, Manuel, et al.
Published: (2025)
QiMeng-Attention: SOTA Attention Operator is generated by SOTA Attention Algorithm
by: Zhou, Qirui, et al.
Published: (2025)
by: Zhou, Qirui, et al.
Published: (2025)
HYPEROFA: Expanding LLM Vocabulary to New Languages via Hypernetwork-Based Embedding Initialization
by: Özeren, Enes, et al.
Published: (2025)
by: Özeren, Enes, et al.
Published: (2025)
ZERA: Zero-init Instruction Evolving Refinement Agent -- From Zero Instructions to Structured Prompts via Principle-based Optimization
by: Yi, Seungyoun, et al.
Published: (2025)
by: Yi, Seungyoun, et al.
Published: (2025)
QUAD: Quantization and Parameter-Efficient Tuning of LLM with Activation Decomposition
by: Hu, Yuxuan, et al.
Published: (2025)
by: Hu, Yuxuan, et al.
Published: (2025)
Enhancing Traffic Accident Classifications: Application of NLP Methods for City Safety
by: Özeren, Enes, et al.
Published: (2025)
by: Özeren, Enes, et al.
Published: (2025)
In-Context Algebra
by: Todd, Eric, et al.
Published: (2025)
by: Todd, Eric, et al.
Published: (2025)
From Facts to Folklore: Evaluating Large Language Models on Bengali Cultural Knowledge
by: Chowdhury, Nafis, et al.
Published: (2025)
by: Chowdhury, Nafis, et al.
Published: (2025)
Similar Items
-
Compact Example-Based Explanations for Language Models
by: Schoenegger, Loris, et al.
Published: (2026) -
Select or Project? Evaluating Lower-dimensional Vectors for LLM Training Data Explanations
by: Hinterleitner, Lukas, et al.
Published: (2026) -
WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training
by: Tian, Changxin, et al.
Published: (2025) -
Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment
by: Zhang, Yizhuo, et al.
Published: (2025) -
Pre-trained Models Perform the Best When Token Distributions Follow Zipf's Law
by: He, Yanjin, et al.
Published: (2025)