DistZO2: High-Throughput and Memory-Efficient Zeroth-Order Fine-tuning LLMs with Distributed Parallel Computing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Liangyu, Xie, Huanyi, Wang, Di |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ZO2: Scalable Zeroth-Order Fine-Tuning for Extremely Large Language Models with Limited GPU Memory
von: Wang, Liangyu, et al.
Veröffentlicht: (2025)
von: Wang, Liangyu, et al.
Veröffentlicht: (2025)
TeZO: Empowering the Low-Rankness on the Temporal Dimension in the Zeroth-Order Optimization for Fine-tuning LLMs
von: Sun, Yan, et al.
Veröffentlicht: (2025)
von: Sun, Yan, et al.
Veröffentlicht: (2025)
A High-Throughput Compute-Efficient POMDP Hide-And-Seek-Engine (HASE) for Multi-Agent Operations
von: Flavin, Timothy, et al.
Veröffentlicht: (2026)
von: Flavin, Timothy, et al.
Veröffentlicht: (2026)
CurvZO: Adaptive Curvature-Guided Sparse Zeroth-Order Optimization for Efficient LLM Fine-Tuning
von: Wang, Shuo, et al.
Veröffentlicht: (2026)
von: Wang, Shuo, et al.
Veröffentlicht: (2026)
ElasticZO: A Memory-Efficient On-Device Learning with Combined Zeroth- and First-Order Optimization
von: Sugiura, Keisuke, et al.
Veröffentlicht: (2025)
von: Sugiura, Keisuke, et al.
Veröffentlicht: (2025)
AdaMeZO: Adam-style Zeroth-Order Optimizer for LLM Fine-tuning Without Maintaining the Moments
von: Cai, Zhijie, et al.
Veröffentlicht: (2026)
von: Cai, Zhijie, et al.
Veröffentlicht: (2026)
DistMLIP: A Distributed Inference Platform for Machine Learning Interatomic Potentials
von: Han, Kevin, et al.
Veröffentlicht: (2025)
von: Han, Kevin, et al.
Veröffentlicht: (2025)
QuZO: Quantized Zeroth-Order Fine-Tuning for Large Language Models
von: Zhou, Jiajun, et al.
Veröffentlicht: (2025)
von: Zhou, Jiajun, et al.
Veröffentlicht: (2025)
Reducing Compute Waste in LLMs through Kernel-Level DVFS
von: Spaan, Jeffrey, et al.
Veröffentlicht: (2026)
von: Spaan, Jeffrey, et al.
Veröffentlicht: (2026)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
Throughput Optimization as a Strategic Lever in Large-Scale AI Systems: Evidence from Dataloader and Memory Profiling Innovations
von: Jha, Mayank
Veröffentlicht: (2026)
von: Jha, Mayank
Veröffentlicht: (2026)
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
von: Shao, Zishan, et al.
Veröffentlicht: (2025)
von: Shao, Zishan, et al.
Veröffentlicht: (2025)
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
von: Zhou, Zikai, et al.
Veröffentlicht: (2025)
von: Zhou, Zikai, et al.
Veröffentlicht: (2025)
Cloud Computing Energy Consumption Prediction Based on Kernel Extreme Learning Machine Algorithm Improved by Vector Weighted Average Algorithm
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
von: Holmes, Connor, et al.
Veröffentlicht: (2024)
von: Holmes, Connor, et al.
Veröffentlicht: (2024)
Parallel Implementations Assessment of a Spatial-Spectral Classifier for Hyperspectral Clinical Applications
von: Lazcano, Raquel, et al.
Veröffentlicht: (2024)
von: Lazcano, Raquel, et al.
Veröffentlicht: (2024)
Large-Scale Data Parallelization of Product Quantization and Inverted Indexing Using Dask
von: Abraham, Ashley N., et al.
Veröffentlicht: (2026)
von: Abraham, Ashley N., et al.
Veröffentlicht: (2026)
Accelerating Diffusion LLMs via Adaptive Parallel Decoding
von: Israel, Daniel, et al.
Veröffentlicht: (2025)
von: Israel, Daniel, et al.
Veröffentlicht: (2025)
Sparse MeZO: Less Parameters for Better Performance in Zeroth-Order LLM Fine-Tuning
von: Liu, Yong, et al.
Veröffentlicht: (2024)
von: Liu, Yong, et al.
Veröffentlicht: (2024)
MaZO: Masked Zeroth-Order Optimization for Multi-Task Fine-Tuning of Large Language Models
von: Zhang, Zhen, et al.
Veröffentlicht: (2025)
von: Zhang, Zhen, et al.
Veröffentlicht: (2025)
Efficient Reinforcement Learning for Routing Jobs in Heterogeneous Queueing Systems
von: Jali, Neharika, et al.
Veröffentlicht: (2024)
von: Jali, Neharika, et al.
Veröffentlicht: (2024)
gDist: Efficient Distance Computation between 3D Meshes on GPU
von: Fang, Peng, et al.
Veröffentlicht: (2024)
von: Fang, Peng, et al.
Veröffentlicht: (2024)
Simultaneous Computation and Memory Efficient Zeroth-Order Optimizer for Fine-Tuning Large Language Models
von: Wang, Fei, et al.
Veröffentlicht: (2024)
von: Wang, Fei, et al.
Veröffentlicht: (2024)
Conformer-Based Speech Recognition On Extreme Edge-Computing Devices
von: Xu, Mingbin, et al.
Veröffentlicht: (2023)
von: Xu, Mingbin, et al.
Veröffentlicht: (2023)
Memory Analysis on the Training Course of DeepSeek Models
von: Zhang, Ping, et al.
Veröffentlicht: (2025)
von: Zhang, Ping, et al.
Veröffentlicht: (2025)
Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models
von: Wang, Liangyu, et al.
Veröffentlicht: (2025)
von: Wang, Liangyu, et al.
Veröffentlicht: (2025)
AR1-ZO: Topology-Aware Rank-1 Zeroth-Order Queries for High-Rank LoRA Fine-Tuning
von: Chen, Ziye, et al.
Veröffentlicht: (2026)
von: Chen, Ziye, et al.
Veröffentlicht: (2026)
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models
von: Ding, Shiwei, et al.
Veröffentlicht: (2025)
von: Ding, Shiwei, et al.
Veröffentlicht: (2025)
OMPILOT: Harnessing Transformer Models for Auto Parallelization to Shared Memory Computing Paradigms
von: Bhattacharjee, Arijit, et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Arijit, et al.
Veröffentlicht: (2025)
A Structure-Aware Framework for Learning Device Placements on Computation Graphs
von: Duan, Shukai, et al.
Veröffentlicht: (2024)
von: Duan, Shukai, et al.
Veröffentlicht: (2024)
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
von: Zhao, Youpeng, et al.
Veröffentlicht: (2024)
von: Zhao, Youpeng, et al.
Veröffentlicht: (2024)
APOLLO: SGD-like Memory, AdamW-level Performance
von: Zhu, Hanqing, et al.
Veröffentlicht: (2024)
von: Zhu, Hanqing, et al.
Veröffentlicht: (2024)
Efficient GPU implementation of randomized SVD and its applications
von: Struski, Łukasz, et al.
Veröffentlicht: (2021)
von: Struski, Łukasz, et al.
Veröffentlicht: (2021)
Zeroth-Order Fine-Tuning of LLMs in Random Subspaces
von: Yu, Ziming, et al.
Veröffentlicht: (2024)
von: Yu, Ziming, et al.
Veröffentlicht: (2024)
Ariel-ML: Computing Parallelization with Embedded Rust for Neural Networks on Heterogeneous Multi-core Microcontrollers
von: Huang, Zhaolan, et al.
Veröffentlicht: (2025)
von: Huang, Zhaolan, et al.
Veröffentlicht: (2025)
Accuracy and Consumption analysis from a compressed model by CompactifAI from Multiverse Computing
von: Fovet, Damien, et al.
Veröffentlicht: (2025)
von: Fovet, Damien, et al.
Veröffentlicht: (2025)
Towards Computational Performance Engineering for Unsupervised Concept Drift Detection -- Complexities, Benchmarking, Performance Analysis
von: Werner, Elias, et al.
Veröffentlicht: (2023)
von: Werner, Elias, et al.
Veröffentlicht: (2023)
ZO-DARTS++: An Efficient and Size-Variable Zeroth-Order Neural Architecture Search Algorithm
von: Xie, Lunchen, et al.
Veröffentlicht: (2025)
von: Xie, Lunchen, et al.
Veröffentlicht: (2025)
AdaGradSelect: An adaptive gradient-guided layer selection method for efficient fine-tuning of SLMs
von: Kumar, Anshul, et al.
Veröffentlicht: (2025)
von: Kumar, Anshul, et al.
Veröffentlicht: (2025)
Integration of a systolic array based hardware accelerator into a DNN operator auto-tuning framework
von: Peccia, F. N., et al.
Veröffentlicht: (2022)
von: Peccia, F. N., et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
ZO2: Scalable Zeroth-Order Fine-Tuning for Extremely Large Language Models with Limited GPU Memory
von: Wang, Liangyu, et al.
Veröffentlicht: (2025) -
TeZO: Empowering the Low-Rankness on the Temporal Dimension in the Zeroth-Order Optimization for Fine-tuning LLMs
von: Sun, Yan, et al.
Veröffentlicht: (2025) -
A High-Throughput Compute-Efficient POMDP Hide-And-Seek-Engine (HASE) for Multi-Agent Operations
von: Flavin, Timothy, et al.
Veröffentlicht: (2026) -
CurvZO: Adaptive Curvature-Guided Sparse Zeroth-Order Optimization for Efficient LLM Fine-Tuning
von: Wang, Shuo, et al.
Veröffentlicht: (2026) -
ElasticZO: A Memory-Efficient On-Device Learning with Combined Zeroth- and First-Order Optimization
von: Sugiura, Keisuke, et al.
Veröffentlicht: (2025)