Scalable Parameter and Memory Efficient Pretraining for LLM: Recent Algorithmic Advances and Benchmarking
Fuente:
arXiv
Saved in:
| Main Authors: | Glentis, Athanasios, Li, Jiaxiang, Shang, Qiulin, Han, Andi, Tsaknakis, Ioannis, Wei, Quan, Hong, Mingyi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
by: Glentis, Athanasios, et al.
Published: (2025)
by: Glentis, Athanasios, et al.
Published: (2025)
Do LLMs Recognize Your Latent Preferences? A Benchmark for Latent Information Discovery in Personalized Interaction
by: Tsaknakis, Ioannis, et al.
Published: (2025)
by: Tsaknakis, Ioannis, et al.
Published: (2025)
Getting More Juice Out of the SFT Data: Reward Learning from Human Demonstration Improves SFT for LLM Alignment
by: Li, Jiaxiang, et al.
Published: (2024)
by: Li, Jiaxiang, et al.
Published: (2024)
Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates
by: Glentis, Athanasios, et al.
Published: (2026)
by: Glentis, Athanasios, et al.
Published: (2026)
A Doubly Stochastically Perturbed Algorithm for Linearly Constrained Bilevel Optimization
by: Khanduri, Prashant, et al.
Published: (2025)
by: Khanduri, Prashant, et al.
Published: (2025)
A Discretization Approach for Bilevel Optimization with Low-Dimensional and Non-Convex Lower-Level
by: Jiang, Xiaotian, et al.
Published: (2025)
by: Jiang, Xiaotian, et al.
Published: (2025)
A Framework for Quantifying How Pre-Training and Context Benefit In-Context Learning
by: Song, Bingqing, et al.
Published: (2025)
by: Song, Bingqing, et al.
Published: (2025)
KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
by: Cheng, Rongxin, et al.
Published: (2024)
by: Cheng, Rongxin, et al.
Published: (2024)
Systematic Analysis for Pretrained Language Model Priming for Parameter-Efficient Fine-tuning
by: Huang, Shih-Cheng, et al.
Published: (2022)
by: Huang, Shih-Cheng, et al.
Published: (2022)
Muon Outperforms Adam in Tail-End Associative Memory Learning
by: Wang, Shuche, et al.
Published: (2025)
by: Wang, Shuche, et al.
Published: (2025)
A Scalable Pretraining Framework for Link Prediction with Efficient Adaptation
by: Song, Yu, et al.
Published: (2025)
by: Song, Yu, et al.
Published: (2025)
RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models
by: Wei, Quan, et al.
Published: (2025)
by: Wei, Quan, et al.
Published: (2025)
SimpleMem: Efficient Lifelong Memory for LLM Agents
by: Liu, Jiaqi, et al.
Published: (2026)
by: Liu, Jiaqi, et al.
Published: (2026)
Hippocampus: An Efficient and Scalable Memory Module for Agentic AI
by: Li, Yi, et al.
Published: (2026)
by: Li, Yi, et al.
Published: (2026)
Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language Models
by: Cao, Jiaqi, et al.
Published: (2025)
by: Cao, Jiaqi, et al.
Published: (2025)
Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management
by: Yang, Xinjun, et al.
Published: (2025)
by: Yang, Xinjun, et al.
Published: (2025)
Learning Reward and Policy Jointly from Demonstration and Preference Improves Alignment
by: Li, Chenliang, et al.
Published: (2024)
by: Li, Chenliang, et al.
Published: (2024)
ConstraintBench: Benchmarking LLM Constraint Reasoning on Direct Optimization
by: Tso, Joseph, et al.
Published: (2026)
by: Tso, Joseph, et al.
Published: (2026)
Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight
by: Ye, Junze, et al.
Published: (2025)
by: Ye, Junze, et al.
Published: (2025)
Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
by: Wei, Tianxin, et al.
Published: (2025)
by: Wei, Tianxin, et al.
Published: (2025)
BEAM: Bi-level Memory-adaptive Algorithmic Evolution for LLM-Powered Heuristic Design
by: Xiang, Chuyang, et al.
Published: (2026)
by: Xiang, Chuyang, et al.
Published: (2026)
CloneMem: Benchmarking Long-Term Memory for AI Clones
by: Hu, Sen, et al.
Published: (2026)
by: Hu, Sen, et al.
Published: (2026)
Beyond Random Sampling: Efficient Language Model Pretraining via Curriculum Learning
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
BIPEFT: Budget-Guided Iterative Search for Parameter Efficient Fine-Tuning of Large Pretrained Language Models
by: Chang, Aofei, et al.
Published: (2024)
by: Chang, Aofei, et al.
Published: (2024)
ScaleNet: Scaling up Pretrained Neural Networks with Incremental Parameters
by: Hao, Zhiwei, et al.
Published: (2025)
by: Hao, Zhiwei, et al.
Published: (2025)
A Meta-Level Learning Algorithm for Sequential Hyper-Parameter Space Reduction in AutoML
by: Borboudakis, Giorgos, et al.
Published: (2023)
by: Borboudakis, Giorgos, et al.
Published: (2023)
Effectively Steer LLM To Follow Preference via Building Confident Directions
by: Song, Bingqing, et al.
Published: (2025)
by: Song, Bingqing, et al.
Published: (2025)
LLM Inference Serving: Survey of Recent Advances and Opportunities
by: Li, Baolin, et al.
Published: (2024)
by: Li, Baolin, et al.
Published: (2024)
Efficient Detection of Bad Benchmark Items with Novel Scalability Coefficients
by: Hardy, Michael, et al.
Published: (2026)
by: Hardy, Michael, et al.
Published: (2026)
Towards Agentic OS: An LLM Agent Framework for Linux Schedulers
by: Zheng, Yusheng, et al.
Published: (2025)
by: Zheng, Yusheng, et al.
Published: (2025)
Parameter-Efficient Fine-Tuning of Large Pretrained Models for Instance Segmentation Tasks
by: Baker, Nermeen Abou, et al.
Published: (2026)
by: Baker, Nermeen Abou, et al.
Published: (2026)
LLM-Inspired Pretrain-Then-Finetune for Small-Data, Large-Scale Optimization
by: Zhang, Zishi, et al.
Published: (2026)
by: Zhang, Zishi, et al.
Published: (2026)
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
[Re] Benchmarking LLM Capabilities in Negotiation through Scoreable Games
by: Pollo, Jorge Carrasco, et al.
Published: (2026)
by: Pollo, Jorge Carrasco, et al.
Published: (2026)
Tuning LLMs by RAG Principles: Towards LLM-native Memory
by: Wei, Jiale, et al.
Published: (2025)
by: Wei, Jiale, et al.
Published: (2025)
Secondary Structure-Guided Novel Protein Sequence Generation with Latent Graph Diffusion
by: Hu, Yutong, et al.
Published: (2024)
by: Hu, Yutong, et al.
Published: (2024)
A Survey on Recent Advances in LLM-Based Multi-turn Dialogue Systems
by: Yi, Zihao, et al.
Published: (2024)
by: Yi, Zihao, et al.
Published: (2024)
Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents
by: Wang, Zihao, et al.
Published: (2025)
by: Wang, Zihao, et al.
Published: (2025)
On Entropy Control in LLM-RL Algorithms
by: Shen, Han
Published: (2025)
by: Shen, Han
Published: (2025)
FrontierOR: Benchmarking LLMs' Capacity for Efficient Algorithm Design in Large-Scale Optimization
by: Kong, Minwei, et al.
Published: (2026)
by: Kong, Minwei, et al.
Published: (2026)
Similar Items
-
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
by: Glentis, Athanasios, et al.
Published: (2025) -
Do LLMs Recognize Your Latent Preferences? A Benchmark for Latent Information Discovery in Personalized Interaction
by: Tsaknakis, Ioannis, et al.
Published: (2025) -
Getting More Juice Out of the SFT Data: Reward Learning from Human Demonstration Improves SFT for LLM Alignment
by: Li, Jiaxiang, et al.
Published: (2024) -
Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates
by: Glentis, Athanasios, et al.
Published: (2026) -
A Doubly Stochastically Perturbed Algorithm for Linearly Constrained Bilevel Optimization
by: Khanduri, Prashant, et al.
Published: (2025)