Saved in:
| Main Authors: | Rybakov, Oleg, Chrzanowski, Mike, Dykas, Peter, Xue, Jinze, Lanir, Ben |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2410.16682 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MoR: Mixture Of Representations For Mixed-Precision Training
by: Su, Bor-Yiing, et al.
Published: (2025)
by: Su, Bor-Yiing, et al.
Published: (2025)
Replaying pre-training data improves fine-tuning
by: Kotha, Suhas, et al.
Published: (2026)
by: Kotha, Suhas, et al.
Published: (2026)
Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training
by: Zhong, Zexuan, et al.
Published: (2024)
by: Zhong, Zexuan, et al.
Published: (2024)
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
by: McKinzie, Brandon, et al.
Published: (2024)
by: McKinzie, Brandon, et al.
Published: (2024)
SimulTron: On-Device Simultaneous Speech to Speech Translation
by: Agranovich, Alex, et al.
Published: (2024)
by: Agranovich, Alex, et al.
Published: (2024)
Variance Control via Weight Rescaling in LLM Pre-training
by: Owen, Louis, et al.
Published: (2025)
by: Owen, Louis, et al.
Published: (2025)
Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning
by: Ding, Bowen, et al.
Published: (2025)
by: Ding, Bowen, et al.
Published: (2025)
ECG-LLM -- training and evaluation of domain-specific large language models for electrocardiography
by: Ahrens, Lara, et al.
Published: (2025)
by: Ahrens, Lara, et al.
Published: (2025)
Pre-training LLM without Learning Rate Decay Enhances Supervised Fine-Tuning
by: Yano, Kazuki, et al.
Published: (2026)
by: Yano, Kazuki, et al.
Published: (2026)
DUMP: Automated Distribution-Level Curriculum Learning for RL-based LLM Post-training
by: Wang, Zhenting, et al.
Published: (2025)
by: Wang, Zhenting, et al.
Published: (2025)
Thinking with Knowledge Graphs: Enhancing LLM Reasoning Through Structured Data
by: Wu, Xue, et al.
Published: (2024)
by: Wu, Xue, et al.
Published: (2024)
The benefits of query-based KGQA systems for complex and temporal questions in LLM era
by: Alekseev, Artem, et al.
Published: (2025)
by: Alekseev, Artem, et al.
Published: (2025)
SALSA: Single-pass Autoregressive LLM Structured Classification
by: Berdichevsky, Ruslan, et al.
Published: (2025)
by: Berdichevsky, Ruslan, et al.
Published: (2025)
Hadamard Adapter: An Extreme Parameter-Efficient Adapter Tuning Method for Pre-trained Language Models
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
Cookbook: A framework for improving LLM generative abilities via programmatic data generating templates
by: Narayan, Avanika, et al.
Published: (2024)
by: Narayan, Avanika, et al.
Published: (2024)
PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding
by: Nakka, Krishna Kanth, et al.
Published: (2024)
by: Nakka, Krishna Kanth, et al.
Published: (2024)
RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains
by: Jiang, Haoxiang, et al.
Published: (2026)
by: Jiang, Haoxiang, et al.
Published: (2026)
nanoLM: an Affordable LLM Pre-training Benchmark via Accurate Loss Prediction across Scales
by: Yao, Yiqun, et al.
Published: (2023)
by: Yao, Yiqun, et al.
Published: (2023)
A Lightweight Method to Disrupt Memorized Sequences in LLM
by: Prashant, Parjanya Prajakta, et al.
Published: (2025)
by: Prashant, Parjanya Prajakta, et al.
Published: (2025)
LLM Dataset Inference: Did you train on my dataset?
by: Maini, Pratyush, et al.
Published: (2024)
by: Maini, Pratyush, et al.
Published: (2024)
Robust LLM safeguarding via refusal feature adversarial training
by: Yu, Lei, et al.
Published: (2024)
by: Yu, Lei, et al.
Published: (2024)
Post-training an LLM for RAG? Train on Self-Generated Demonstrations
by: Finlayson, Matthew, et al.
Published: (2025)
by: Finlayson, Matthew, et al.
Published: (2025)
Pretraining Large Language Models with NVFP4
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
BTS: Harmonizing Specialized Experts into a Generalist LLM
by: Zhang, Qizhen, et al.
Published: (2025)
by: Zhang, Qizhen, et al.
Published: (2025)
Behavior Structformer: Learning Players Representations with Structured Tokenization
by: Smirnov, Oleg, et al.
Published: (2024)
by: Smirnov, Oleg, et al.
Published: (2024)
Confidence Estimation for Error Detection in Text-to-SQL Systems
by: Somov, Oleg, et al.
Published: (2025)
by: Somov, Oleg, et al.
Published: (2025)
Probabilistic Soundness Guarantees in LLM Reasoning Chains
by: You, Weiqiu, et al.
Published: (2025)
by: You, Weiqiu, et al.
Published: (2025)
When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method
by: Zhang, Biao, et al.
Published: (2024)
by: Zhang, Biao, et al.
Published: (2024)
LLM Unlearning Under the Microscope: A Full-Stack View on Methods and Metrics
by: Fan, Chongyu, et al.
Published: (2025)
by: Fan, Chongyu, et al.
Published: (2025)
Deep Learning-based Method for Expressing Knowledge Boundary of Black-Box LLM
by: Sheng, Haotian, et al.
Published: (2026)
by: Sheng, Haotian, et al.
Published: (2026)
Thinking Augmented Pre-training
by: Wang, Liang, et al.
Published: (2025)
by: Wang, Liang, et al.
Published: (2025)
NuNER: Entity Recognition Encoder Pre-training via LLM-Annotated Data
by: Bogdanov, Sergei, et al.
Published: (2024)
by: Bogdanov, Sergei, et al.
Published: (2024)
Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
by: Bergsma, Shane, et al.
Published: (2025)
by: Bergsma, Shane, et al.
Published: (2025)
Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
by: Jing, Yi, et al.
Published: (2026)
by: Jing, Yi, et al.
Published: (2026)
Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods
by: Doshi, Jai, et al.
Published: (2024)
by: Doshi, Jai, et al.
Published: (2024)
Domaino1s: Guiding LLM Reasoning for Explainable Answers in High-Stakes Domains
by: Chu, Xu, et al.
Published: (2025)
by: Chu, Xu, et al.
Published: (2025)
Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment
by: Zhang, Yizhuo, et al.
Published: (2025)
by: Zhang, Yizhuo, et al.
Published: (2025)
LLM Probing with Contrastive Eigenproblems: Improving Understanding and Applicability of CCS
by: Schouten, Stefan F., et al.
Published: (2025)
by: Schouten, Stefan F., et al.
Published: (2025)
Health Insurance Coverage Rule Interpretation Corpus: Law, Policy, and Medical Guidance for Health Insurance Coverage Understanding
by: Gartner, Mike
Published: (2025)
by: Gartner, Mike
Published: (2025)
WISCA: A Lightweight Model Transition Method to Improve LLM Training via Weight Scaling
by: Li, Jiacheng, et al.
Published: (2025)
by: Li, Jiacheng, et al.
Published: (2025)
Similar Items
-
MoR: Mixture Of Representations For Mixed-Precision Training
by: Su, Bor-Yiing, et al.
Published: (2025) -
Replaying pre-training data improves fine-tuning
by: Kotha, Suhas, et al.
Published: (2026) -
Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training
by: Zhong, Zexuan, et al.
Published: (2024) -
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
by: McKinzie, Brandon, et al.
Published: (2024) -
SimulTron: On-Device Simultaneous Speech to Speech Translation
by: Agranovich, Alex, et al.
Published: (2024)