Speeding up Model Loading with fastsafetensors
Fuente:
arXiv
Saved in:
| Main Authors: | Yoshimura, Takeshi, Chiba, Tatsuhiro, Sethi, Manish, Waddington, Daniel, Sundararaman, Swaminathan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accuracy Is Speed: Towards Long-Context-Aware Routing for Distributed LLM Serving
by: Yoshimura, Takeshi, et al.
Published: (2026)
by: Yoshimura, Takeshi, et al.
Published: (2026)
Rorqual: Speeding up Narwhal with TEEs
by: Freitas, Luciano, et al.
Published: (2024)
by: Freitas, Luciano, et al.
Published: (2024)
A Robust Power Model Training Framework for Cloud Native Runtime Energy Metric Exporter
by: Choochotkaew, Sunyanan, et al.
Published: (2024)
by: Choochotkaew, Sunyanan, et al.
Published: (2024)
A Communication- and Memory-Aware Model for Load Balancing Tasks
by: Lifflander, Jonathan, et al.
Published: (2024)
by: Lifflander, Jonathan, et al.
Published: (2024)
Scalable and Performant Data Loading
by: Hira, Moto, et al.
Published: (2025)
by: Hira, Moto, et al.
Published: (2025)
Optimizing Robot Dispersion on Grids: with and without Fault Tolerance
by: Banerjee, Rik, et al.
Published: (2024)
by: Banerjee, Rik, et al.
Published: (2024)
Optimal Fault-Tolerant Dispersion on Oriented Grids
by: Banerjee, Rik, et al.
Published: (2024)
by: Banerjee, Rik, et al.
Published: (2024)
Speeding up Local Optimization in Vehicle Routing with Tensor-based GPU Acceleration
by: Lei, Zhenyu, et al.
Published: (2025)
by: Lei, Zhenyu, et al.
Published: (2025)
Accelerating Loading WebGraphs in ParaGrapher
by: Esfahani, Mohsen Koohi
Published: (2025)
by: Esfahani, Mohsen Koohi
Published: (2025)
ExClique: An Express Consensus Algorithm for High-Speed Transaction Process in Blockchains
by: Zhao, Chonghe, et al.
Published: (2025)
by: Zhao, Chonghe, et al.
Published: (2025)
Chasing the Speed of Light: Low-Latency Planetary-Scale Adaptive Byzantine Consensus
by: Berger, Christian, et al.
Published: (2023)
by: Berger, Christian, et al.
Published: (2023)
Everywhere & Nowhere: Envisioning a Computing Continuum for Science
by: Parashar, Manish
Published: (2024)
by: Parashar, Manish
Published: (2024)
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
by: Tran, Phuong, et al.
Published: (2025)
by: Tran, Phuong, et al.
Published: (2025)
Distributed Load Balancing with Workload-Dependent Service Rates
by: Zhang, Wenxin, et al.
Published: (2024)
by: Zhang, Wenxin, et al.
Published: (2024)
Inference Load-Aware Orchestration for Hierarchical Federated Learning
by: Lackinger, Anna, et al.
Published: (2024)
by: Lackinger, Anna, et al.
Published: (2024)
Memory Efficient and Staleness Free Pipeline Parallel DNN Training Framework with Improved Convergence Speed
by: Dutta, Ankita, et al.
Published: (2025)
by: Dutta, Ankita, et al.
Published: (2025)
Wait or Not to Wait: Evaluating Trade-Offs between Speed and Precision in Blockchain-based Federated Aggregation
by: Nguyen, Huong, et al.
Published: (2024)
by: Nguyen, Huong, et al.
Published: (2024)
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models
by: Wang, Wei, et al.
Published: (2024)
by: Wang, Wei, et al.
Published: (2024)
Morpheus: Lightweight RTT Prediction for Performance-Aware Load Balancing
by: Giannakopoulos, Panagiotis, et al.
Published: (2025)
by: Giannakopoulos, Panagiotis, et al.
Published: (2025)
Fine-grained MoE Load Balancing with Linear Programming
by: Zhao, Chenqi, et al.
Published: (2025)
by: Zhao, Chenqi, et al.
Published: (2025)
Hiding Latencies in Network-Based Image Loading for Deep Learning
by: Versaci, Francesco, et al.
Published: (2025)
by: Versaci, Francesco, et al.
Published: (2025)
Review of Hybrid Load Balancing Algorithms in Cloud Computing Environment
by: Ijeoma, Chukwuneke Chiamaka, et al.
Published: (2022)
by: Ijeoma, Chukwuneke Chiamaka, et al.
Published: (2022)
Load Balanced Parallel Node Generation for Meshless Numerical Methods
by: Vehovar, Jon, et al.
Published: (2026)
by: Vehovar, Jon, et al.
Published: (2026)
Flare: Leveraging Serverless Elasticity to Absorb Microservice Load Spikes
by: Dehigama, Dilina, et al.
Published: (2026)
by: Dehigama, Dilina, et al.
Published: (2026)
Supervised Distributed Computing: Efficiency and Robustness under a Majority of Adversarial Workers
by: Augustine, John, et al.
Published: (2026)
by: Augustine, John, et al.
Published: (2026)
Leveraging Public Cloud Infrastructure for Real-time Connected Vehicle Speed Advisory at a Signalized Corridor
by: Deng, Hsien-Wen, et al.
Published: (2024)
by: Deng, Hsien-Wen, et al.
Published: (2024)
Load Balancing in Strongly Inhomogeneous Simulations -- a Vlasiator Case Study
by: Kotipalo, Leo, et al.
Published: (2025)
by: Kotipalo, Leo, et al.
Published: (2025)
Distributed Load Orchestration for Vision Computing in Multi-Access Edge Computing
by: Boing, Ricardo N., et al.
Published: (2022)
by: Boing, Ricardo N., et al.
Published: (2022)
KnapsackLB: Enabling Performance-Aware Layer-4 Load Balancing
by: Gandhi, Rohan, et al.
Published: (2024)
by: Gandhi, Rohan, et al.
Published: (2024)
Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving
by: Cheng, Ke, et al.
Published: (2024)
by: Cheng, Ke, et al.
Published: (2024)
AES-SpMM: Balancing Accuracy and Speed by Adaptive Edge Sampling Strategy to Accelerate SpMM in GNNs
by: Song, Yingchen, et al.
Published: (2025)
by: Song, Yingchen, et al.
Published: (2025)
TD-Orch: Scalable Load-Balancing for Distributed Systems with Applications to Graph Processing
by: Zhao, Yiwei, et al.
Published: (2025)
by: Zhao, Yiwei, et al.
Published: (2025)
ReaLB: Real-Time Load Balancing for Multimodal MoE Inference
by: Wang, Yingping, et al.
Published: (2026)
by: Wang, Yingping, et al.
Published: (2026)
LB4OMP: A Dynamic Load Balancing Library for Multithreaded Applications
by: Korndörfer, Jonas H. Müller, et al.
Published: (2021)
by: Korndörfer, Jonas H. Müller, et al.
Published: (2021)
QEdgeProxy: QoS-Aware Load Balancing for IoT Services in the Computing Continuum
by: Čilić, Ivan, et al.
Published: (2024)
by: Čilić, Ivan, et al.
Published: (2024)
Online Load and Graph Balancing for Random Order Inputs
by: Im, Sungjin, et al.
Published: (2024)
by: Im, Sungjin, et al.
Published: (2024)
Distributed Download from an External Data Source in Asynchronous Faulty Settings
by: Augustine, John, et al.
Published: (2025)
by: Augustine, John, et al.
Published: (2025)
SkyWalker: A Locality-Aware Cross-Region Load Balancer for LLM Inference
by: Xia, Tian, et al.
Published: (2025)
by: Xia, Tian, et al.
Published: (2025)
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
by: Yuan, Yitao, et al.
Published: (2025)
by: Yuan, Yitao, et al.
Published: (2025)
An Analytical Overview Of Virtual Machine Load Balancing Scheduling Algorithms with their Comparative Case Study
by: Vaidya, Priyank, et al.
Published: (2025)
by: Vaidya, Priyank, et al.
Published: (2025)
Similar Items
-
Accuracy Is Speed: Towards Long-Context-Aware Routing for Distributed LLM Serving
by: Yoshimura, Takeshi, et al.
Published: (2026) -
Rorqual: Speeding up Narwhal with TEEs
by: Freitas, Luciano, et al.
Published: (2024) -
A Robust Power Model Training Framework for Cloud Native Runtime Energy Metric Exporter
by: Choochotkaew, Sunyanan, et al.
Published: (2024) -
A Communication- and Memory-Aware Model for Load Balancing Tasks
by: Lifflander, Jonathan, et al.
Published: (2024) -
Scalable and Performant Data Loading
by: Hira, Moto, et al.
Published: (2025)