Influence-driven Curriculum Learning for Pre-training on Limited Data
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Schoenegger, Loris, Thoma, Lukas, Blevins, Terra, Roth, Benjamin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Compact Example-Based Explanations for Language Models
von: Schoenegger, Loris, et al.
Veröffentlicht: (2026)
von: Schoenegger, Loris, et al.
Veröffentlicht: (2026)
Select or Project? Evaluating Lower-dimensional Vectors for LLM Training Data Explanations
von: Hinterleitner, Lukas, et al.
Veröffentlicht: (2026)
von: Hinterleitner, Lukas, et al.
Veröffentlicht: (2026)
WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training
von: Tian, Changxin, et al.
Veröffentlicht: (2025)
von: Tian, Changxin, et al.
Veröffentlicht: (2025)
Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment
von: Zhang, Yizhuo, et al.
Veröffentlicht: (2025)
von: Zhang, Yizhuo, et al.
Veröffentlicht: (2025)
Pre-trained Models Perform the Best When Token Distributions Follow Zipf's Law
von: He, Yanjin, et al.
Veröffentlicht: (2025)
von: He, Yanjin, et al.
Veröffentlicht: (2025)
BabyLlama-2: Ensemble-Distilled Models Consistently Outperform Teachers With Limited Data
von: Tastet, Jean-Loup, et al.
Veröffentlicht: (2024)
von: Tastet, Jean-Loup, et al.
Veröffentlicht: (2024)
Your Pretrained Model Tells the Difficulty Itself: A Self-Adaptive Curriculum Learning Paradigm for Natural Language Understanding
von: Feng, Qi, et al.
Veröffentlicht: (2025)
von: Feng, Qi, et al.
Veröffentlicht: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
PowLU: An Activation Function for Stable Pre-Training of LLMs
von: Jiang, Peijie, et al.
Veröffentlicht: (2026)
von: Jiang, Peijie, et al.
Veröffentlicht: (2026)
Geometric Deviation as an Unsupervised Pre-Generation Reliability Signal: Probing LLM Representations for Answerability
von: Du, Yucheng
Veröffentlicht: (2026)
von: Du, Yucheng
Veröffentlicht: (2026)
Overcoming Long-Context Limitations of State-Space Models via Context-Dependent Sparse Attention
von: Zhan, Zhihao, et al.
Veröffentlicht: (2025)
von: Zhan, Zhihao, et al.
Veröffentlicht: (2025)
$FastDoc$: Domain-Specific Fast Continual Pre-training Technique using Document-Level Metadata and Taxonomy
von: Nandy, Abhilash, et al.
Veröffentlicht: (2023)
von: Nandy, Abhilash, et al.
Veröffentlicht: (2023)
Improving Influence-based Instruction Tuning Data Selection for Balanced Learning of Diverse Capabilities
von: Dai, Qirun, et al.
Veröffentlicht: (2025)
von: Dai, Qirun, et al.
Veröffentlicht: (2025)
On Initializing Transformers with Pre-trained Embeddings
von: Kim, Ha Young, et al.
Veröffentlicht: (2024)
von: Kim, Ha Young, et al.
Veröffentlicht: (2024)
Reinforcement Learning for Latent-Space Thinking in LLMs
von: Özeren, Enes, et al.
Veröffentlicht: (2025)
von: Özeren, Enes, et al.
Veröffentlicht: (2025)
MALoRA: Mixture of Asymmetric Low-Rank Adaptation for Enhanced Multi-Task Learning
von: Wang, Xujia, et al.
Veröffentlicht: (2024)
von: Wang, Xujia, et al.
Veröffentlicht: (2024)
Large Language Model (LLM) Bias Index -- LLMBI
von: Oketunji, Abiodun Finbarrs, et al.
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs, et al.
Veröffentlicht: (2023)
Engineering A Large Language Model From Scratch
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2024)
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2024)
The Data Efficiency Frontier of Financial Foundation Models: Scaling Laws from Continued Pretraining
von: Ponnock, Jesse
Veröffentlicht: (2025)
von: Ponnock, Jesse
Veröffentlicht: (2025)
Continuous-Depth Transformers with Learned Control Dynamics
von: Jemley, Peter
Veröffentlicht: (2026)
von: Jemley, Peter
Veröffentlicht: (2026)
COA-GPT: Generative Pre-trained Transformers for Accelerated Course of Action Development in Military Operations
von: Goecks, Vinicius G., et al.
Veröffentlicht: (2024)
von: Goecks, Vinicius G., et al.
Veröffentlicht: (2024)
Arabic Hate Speech Identification and Masking in Social Media using Deep Learning Models and Pre-trained Models Fine-tuning
von: Doghmash, Salam Thabet, et al.
Veröffentlicht: (2025)
von: Doghmash, Salam Thabet, et al.
Veröffentlicht: (2025)
Pre-training data selection for biomedical domain adaptation using journal impact metrics
von: Laï-king, Mathieu, et al.
Veröffentlicht: (2024)
von: Laï-king, Mathieu, et al.
Veröffentlicht: (2024)
Annotation Entropy Predicts Per-Example Learning Dynamics in LoRA Fine-Tuning
von: Steele, Brady
Veröffentlicht: (2026)
von: Steele, Brady
Veröffentlicht: (2026)
Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility
von: Szilvasy, Gergely, et al.
Veröffentlicht: (2026)
von: Szilvasy, Gergely, et al.
Veröffentlicht: (2026)
MMSciBench: Benchmarking Language Models on Chinese Multimodal Scientific Problems
von: Ye, Xinwu, et al.
Veröffentlicht: (2025)
von: Ye, Xinwu, et al.
Veröffentlicht: (2025)
Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMs
von: Zhang, Zhaowei, et al.
Veröffentlicht: (2025)
von: Zhang, Zhaowei, et al.
Veröffentlicht: (2025)
Interpreto: An Explainability Library for Transformers
von: Poché, Antonin, et al.
Veröffentlicht: (2025)
von: Poché, Antonin, et al.
Veröffentlicht: (2025)
GenAI Content Detection Task 3: Cross-Domain Machine-Generated Text Detection Challenge
von: Dugan, Liam, et al.
Veröffentlicht: (2025)
von: Dugan, Liam, et al.
Veröffentlicht: (2025)
Reconstructing Syllable Sequences in Abugida Scripts with Incomplete Inputs
von: Thu, Ye Kyaw, et al.
Veröffentlicht: (2025)
von: Thu, Ye Kyaw, et al.
Veröffentlicht: (2025)
Ambiguity in LLMs is a concept missing problem
von: Hu, Zhibo, et al.
Veröffentlicht: (2025)
von: Hu, Zhibo, et al.
Veröffentlicht: (2025)
Constraint-Driven Small Language Models Based on Agent and OpenAlex Knowledge Graph: Mining Conceptual Pathways and Discovering Innovation Points in Academic Papers
von: Xia, Ziye, et al.
Veröffentlicht: (2025)
von: Xia, Ziye, et al.
Veröffentlicht: (2025)
Digital Guardians: Can GPT-4, Perspective API, and Moderation API reliably detect hate speech in reader comments of German online newspapers?
von: Weber, Manuel, et al.
Veröffentlicht: (2025)
von: Weber, Manuel, et al.
Veröffentlicht: (2025)
QiMeng-Attention: SOTA Attention Operator is generated by SOTA Attention Algorithm
von: Zhou, Qirui, et al.
Veröffentlicht: (2025)
von: Zhou, Qirui, et al.
Veröffentlicht: (2025)
HYPEROFA: Expanding LLM Vocabulary to New Languages via Hypernetwork-Based Embedding Initialization
von: Özeren, Enes, et al.
Veröffentlicht: (2025)
von: Özeren, Enes, et al.
Veröffentlicht: (2025)
ZERA: Zero-init Instruction Evolving Refinement Agent -- From Zero Instructions to Structured Prompts via Principle-based Optimization
von: Yi, Seungyoun, et al.
Veröffentlicht: (2025)
von: Yi, Seungyoun, et al.
Veröffentlicht: (2025)
QUAD: Quantization and Parameter-Efficient Tuning of LLM with Activation Decomposition
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
Enhancing Traffic Accident Classifications: Application of NLP Methods for City Safety
von: Özeren, Enes, et al.
Veröffentlicht: (2025)
von: Özeren, Enes, et al.
Veröffentlicht: (2025)
In-Context Algebra
von: Todd, Eric, et al.
Veröffentlicht: (2025)
von: Todd, Eric, et al.
Veröffentlicht: (2025)
From Facts to Folklore: Evaluating Large Language Models on Bengali Cultural Knowledge
von: Chowdhury, Nafis, et al.
Veröffentlicht: (2025)
von: Chowdhury, Nafis, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Compact Example-Based Explanations for Language Models
von: Schoenegger, Loris, et al.
Veröffentlicht: (2026) -
Select or Project? Evaluating Lower-dimensional Vectors for LLM Training Data Explanations
von: Hinterleitner, Lukas, et al.
Veröffentlicht: (2026) -
WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training
von: Tian, Changxin, et al.
Veröffentlicht: (2025) -
Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment
von: Zhang, Yizhuo, et al.
Veröffentlicht: (2025) -
Pre-trained Models Perform the Best When Token Distributions Follow Zipf's Law
von: He, Yanjin, et al.
Veröffentlicht: (2025)