MLKV: Efficiently Scaling up Large Embedding Model Training with Disk-based Key-Value Storage
Fuente:
arXiv
Saved in:
| Main Authors: | He, Yongjun, Waleffe, Roger, Han, Zhichao, George, Johnu, Yuan, Binhang, Zhang, Zitao, Shan, Yinan, Zhao, Yang, Dutta, Debojyoti, Rekatsinas, Theodoros, Zhang, Ce |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Predictive Scaling Laws for Efficient GRPO Training of Large Reasoning Models
by: Nimmaturi, Datta, et al.
Published: (2025)
by: Nimmaturi, Datta, et al.
Published: (2025)
Armada: Memory-Efficient Distributed Training of Large-Scale Graph Neural Networks
by: Waleffe, Roger, et al.
Published: (2025)
by: Waleffe, Roger, et al.
Published: (2025)
MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding
by: Zuhri, Zayd Muhammad Kawakibi, et al.
Published: (2024)
by: Zuhri, Zayd Muhammad Kawakibi, et al.
Published: (2024)
TSDS: Data Selection for Task-Specific Model Finetuning
by: Liu, Zifan, et al.
Published: (2024)
by: Liu, Zifan, et al.
Published: (2024)
Action Shapley: A Training Data Selection Metric for World Model in Reinforcement Learning
by: Ghosh, Rajat, et al.
Published: (2026)
by: Ghosh, Rajat, et al.
Published: (2026)
Fundamental Challenges in Evaluating Text2SQL Solutions and Detecting Their Limitations
by: Renggli, Cedric, et al.
Published: (2025)
by: Renggli, Cedric, et al.
Published: (2025)
GraphFramEx: Towards Systematic Evaluation of Explainability Methods for Graph Neural Networks
by: Amara, Kenza, et al.
Published: (2022)
by: Amara, Kenza, et al.
Published: (2022)
Efficient Alignment of Large Language Models via Data Sampling
by: Khera, Amrit, et al.
Published: (2024)
by: Khera, Amrit, et al.
Published: (2024)
CPP-UT-Bench: Can LLMs Write Complex Unit Tests in C++?
by: Bhargava, Vaishnavi, et al.
Published: (2024)
by: Bhargava, Vaishnavi, et al.
Published: (2024)
BVLSM: Write-Efficient LSM-Tree Storage via WAL-Time Key-Value Separation
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories
by: Soni, Aditya Bharat, et al.
Published: (2026)
by: Soni, Aditya Bharat, et al.
Published: (2026)
ΩSFormer: Dual-Modal Ω-like Super-Resolution Transformer Network for Cross-scale and High-accuracy Terraced Field Vectorization Extraction
by: Li, Chang, et al.
Published: (2024)
by: Li, Chang, et al.
Published: (2024)
GraphSnapShot: Caching Local Structure for Fast Graph Learning
by: Liu, Dong, et al.
Published: (2024)
by: Liu, Dong, et al.
Published: (2024)
Robust Carbon‐Dot Optical Disks for Orthogonal Amplitude‐Polarization Encryption Storage
by: Jingying Miao, et al.
Published: (2024)
by: Jingying Miao, et al.
Published: (2024)
A posteriori certification for neural network approximations to PDEs
by: Ernst, Lewin, et al.
Published: (2025)
by: Ernst, Lewin, et al.
Published: (2025)
A Physics Constrained Machine Learning Pipeline for Young's Modulus Prediction in Multimaterial Hyperelastic Cylinders Guided by Contact Mechanics
by: Christoforos Rekatsinas, et al.
Published: (2025)
by: Christoforos Rekatsinas, et al.
Published: (2025)
RANGER -- Repository-Level Agent for Graph-Enhanced Retrieval
by: Shah, Pratik, et al.
Published: (2025)
by: Shah, Pratik, et al.
Published: (2025)
CR-Bench: Evaluating the Real-World Utility of AI Code Review Agents
by: Pereira, Kristen, et al.
Published: (2026)
by: Pereira, Kristen, et al.
Published: (2026)
A Multi-Agent Framework for Stateful Inference-Time Search
by: Lalan, Arshika, et al.
Published: (2025)
by: Lalan, Arshika, et al.
Published: (2025)
Predictive Value of Combined CRP and INR for Intracranial Hypertension in Cerebral Venous Thrombosis
by: Jiahui Yan, et al.
Published: (2025)
by: Jiahui Yan, et al.
Published: (2025)
On triviality of $\mathbb{A}^2$-forms admitting a nontrivial $\mathbb{G}_a$-action
by: Saha, Debojyoti
Published: (2026)
by: Saha, Debojyoti
Published: (2026)
On Generalised Danielewski Surfaces over fields of arbitrary characteristic
by: Saha, Debojyoti
Published: (2025)
by: Saha, Debojyoti
Published: (2025)
HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
by: Yan, Ran, et al.
Published: (2024)
by: Yan, Ran, et al.
Published: (2024)
Chameleon: a Heterogeneous and Disaggregated Accelerator System for Retrieval-Augmented Language Models
by: Jiang, Wenqi, et al.
Published: (2023)
by: Jiang, Wenqi, et al.
Published: (2023)
From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint Training
by: Liu, Tianqiao, et al.
Published: (2025)
by: Liu, Tianqiao, et al.
Published: (2025)
Go-UT-Bench: A Fine-Tuning Dataset for LLM-Based Unit Test Generation in Go
by: Pipalani, Yashshi, et al.
Published: (2025)
by: Pipalani, Yashshi, et al.
Published: (2025)
BAR Conjecture: the Feasibility of Inference Budget-Constrained LLM Services with Authenticity and Reasoning
by: Zhou, Jinan, et al.
Published: (2025)
by: Zhou, Jinan, et al.
Published: (2025)
Deep Uzawa for Kinetic Transport with Lagrange-Enforced Boundaries
by: Makridakis, Charalambos, et al.
Published: (2025)
by: Makridakis, Charalambos, et al.
Published: (2025)
COLE: A Column-based Learned Storage for Blockchain Systems
by: Zhang, Ce, et al.
Published: (2023)
by: Zhang, Ce, et al.
Published: (2023)
Graph Embedding in the Graph Fractional Fourier Transform Domain
by: Sheng, Changjie, et al.
Published: (2025)
by: Sheng, Changjie, et al.
Published: (2025)
GRUHP : An Adaptive Feature Selection Model for Hard Disk Drive Failure Prediction in Large‐Scale Storage Systems
by: Qinlu He, et al.
Published: (2025)
by: Qinlu He, et al.
Published: (2025)
Symmetry Results for Cyclotomic Multiple Hurwitz Zeta Values via Contour Integrals
by: Xu, Ce
Published: (2026)
by: Xu, Ce
Published: (2026)
COLE$^+$: Towards Practical Column-based Learned Storage for Blockchain Systems
by: Zhang, Ce, et al.
Published: (2026)
by: Zhang, Ce, et al.
Published: (2026)
KeyDroid: A Large-Scale Analysis of Secure Key Storage in Android Apps
by: Blessing, Jenny, et al.
Published: (2025)
by: Blessing, Jenny, et al.
Published: (2025)
Quake: Adaptive Indexing for Vector Search
by: Mohoney, Jason, et al.
Published: (2025)
by: Mohoney, Jason, et al.
Published: (2025)
Explicit Evaluation of Euler-Apéry Type Multiple Zeta Star Values and Multiple $t$-Star Values
by: Xu, Ce, et al.
Published: (2022)
by: Xu, Ce, et al.
Published: (2022)
EdgePrompt: A Distributed Key-Value Inference Framework for LLMs in 6G Networks
by: Ning, Jiahong, et al.
Published: (2025)
by: Ning, Jiahong, et al.
Published: (2025)
Generation of Heterogeneous PET Images from Uniform Organ Activity Maps Using a Pretrained Domain-Adapted Diffusion Model
by: Li, Suya, et al.
Published: (2026)
by: Li, Suya, et al.
Published: (2026)
DeepZero: Scaling up Zeroth-Order Optimization for Deep Model Training
by: Chen, Aochuan, et al.
Published: (2023)
by: Chen, Aochuan, et al.
Published: (2023)
Liar's vertex-edge domination in unit disk graph
by: Bhattacharya, Debojyoti, et al.
Published: (2025)
by: Bhattacharya, Debojyoti, et al.
Published: (2025)
Similar Items
-
Predictive Scaling Laws for Efficient GRPO Training of Large Reasoning Models
by: Nimmaturi, Datta, et al.
Published: (2025) -
Armada: Memory-Efficient Distributed Training of Large-Scale Graph Neural Networks
by: Waleffe, Roger, et al.
Published: (2025) -
MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding
by: Zuhri, Zayd Muhammad Kawakibi, et al.
Published: (2024) -
TSDS: Data Selection for Task-Specific Model Finetuning
by: Liu, Zifan, et al.
Published: (2024) -
Action Shapley: A Training Data Selection Metric for World Model in Reinforcement Learning
by: Ghosh, Rajat, et al.
Published: (2026)