You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ding, Shiwei, Zhang, Lan, Wang, Zhenlin, Ateniese, Giuseppe, Yuan, Xiaoyong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tuning the Tuner: Introducing Hyperparameter Optimization for Auto-Tuning
von: Willemsen, Floris-Jan, et al.
Veröffentlicht: (2025)
von: Willemsen, Floris-Jan, et al.
Veröffentlicht: (2025)
Accelerating Mobile Inference through Fine-Grained CPU-GPU Co-Execution
von: Li, Zhuojin, et al.
Veröffentlicht: (2025)
von: Li, Zhuojin, et al.
Veröffentlicht: (2025)
CloudFormer: An Attention-based Performance Prediction for Public Clouds with Unknown Workload
von: Shahbazinia, Amirhossein, et al.
Veröffentlicht: (2025)
von: Shahbazinia, Amirhossein, et al.
Veröffentlicht: (2025)
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
von: Sun, Tingyang, et al.
Veröffentlicht: (2026)
von: Sun, Tingyang, et al.
Veröffentlicht: (2026)
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
von: Titopoulos, Vasileios, et al.
Veröffentlicht: (2025)
von: Titopoulos, Vasileios, et al.
Veröffentlicht: (2025)
KVDirect: Distributed Disaggregated LLM Inference
von: Chen, Shiyang, et al.
Veröffentlicht: (2024)
von: Chen, Shiyang, et al.
Veröffentlicht: (2024)
Ridgeline: A 2D Roofline Model for Distributed Systems
von: Checconi, Fabio, et al.
Veröffentlicht: (2022)
von: Checconi, Fabio, et al.
Veröffentlicht: (2022)
Modeling and Characterizing Service Interference in Dynamic Infrastructures
von: Medel, VÍctor, et al.
Veröffentlicht: (2024)
von: Medel, VÍctor, et al.
Veröffentlicht: (2024)
Distributed Matrix-Based Sampling for Graph Neural Network Training
von: Tripathy, Alok, et al.
Veröffentlicht: (2023)
von: Tripathy, Alok, et al.
Veröffentlicht: (2023)
FalconFS: Distributed File System for Large-Scale Deep Learning Pipeline
von: Xu, Jingwei, et al.
Veröffentlicht: (2025)
von: Xu, Jingwei, et al.
Veröffentlicht: (2025)
QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
When Less is More: Achieving Faster Convergence in Distributed Edge Machine Learning
von: Basani, Advik Raj, et al.
Veröffentlicht: (2024)
von: Basani, Advik Raj, et al.
Veröffentlicht: (2024)
MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs
von: Sarkar, Aishwarya, et al.
Veröffentlicht: (2024)
von: Sarkar, Aishwarya, et al.
Veröffentlicht: (2024)
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
von: Dutt, Anurag, et al.
Veröffentlicht: (2025)
von: Dutt, Anurag, et al.
Veröffentlicht: (2025)
Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI Training
von: Papavasileiou, Ioannis, et al.
Veröffentlicht: (2026)
von: Papavasileiou, Ioannis, et al.
Veröffentlicht: (2026)
Bringing Auto-tuning to HIP: Analysis of Tuning Impact and Difficulty on AMD and Nvidia GPUs
von: Lurati, Milo, et al.
Veröffentlicht: (2024)
von: Lurati, Milo, et al.
Veröffentlicht: (2024)
Performance Optimization in Stream Processing Systems: Experiment-Driven Configuration Tuning for Kafka Streams
von: Chen, David, et al.
Veröffentlicht: (2026)
von: Chen, David, et al.
Veröffentlicht: (2026)
CARAT: Client-Side Adaptive RPC and Cache Co-Tuning for Parallel File Systems
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
A Multi-Port Concurrent Communication Model for handling Compute Intensive Tasks on Distributed Satellite System Constellations
von: Veeravalli, Bharadwaj
Veröffentlicht: (2026)
von: Veeravalli, Bharadwaj
Veröffentlicht: (2026)
On Orchestrating Parallel Broadcasts for Distributed Ledgers
von: Sheng, Peiyao, et al.
Veröffentlicht: (2024)
von: Sheng, Peiyao, et al.
Veröffentlicht: (2024)
An Online Probabilistic Distributed Tracing System
von: Toslali, M., et al.
Veröffentlicht: (2024)
von: Toslali, M., et al.
Veröffentlicht: (2024)
Towards a Peer-to-Peer Data Distribution Layer for Efficient and Collaborative Resource Optimization of Distributed Dataflow Applications
von: Scheinert, Dominik, et al.
Veröffentlicht: (2023)
von: Scheinert, Dominik, et al.
Veröffentlicht: (2023)
A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
von: Apanasevich, L., et al.
Veröffentlicht: (2024)
von: Apanasevich, L., et al.
Veröffentlicht: (2024)
DIAL: Decentralized I/O AutoTuning via Learned Client-side Local Metrics for Parallel File System
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs
von: Cornelius, Melanie, et al.
Veröffentlicht: (2025)
von: Cornelius, Melanie, et al.
Veröffentlicht: (2025)
Kubernetes in Action: Exploring the Performance of Kubernetes Distributions in the Cloud
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024)
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024)
LLMPerf: GPU Performance Modeling meets Large Language Models
von: Nguyen, Khoi N. M., et al.
Veröffentlicht: (2025)
von: Nguyen, Khoi N. M., et al.
Veröffentlicht: (2025)
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
von: Lacey, Dane C., et al.
Veröffentlicht: (2024)
von: Lacey, Dane C., et al.
Veröffentlicht: (2024)
BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
von: Hu, Xiannan, et al.
Veröffentlicht: (2025)
von: Hu, Xiannan, et al.
Veröffentlicht: (2025)
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
von: Zhuang, Chen, et al.
Veröffentlicht: (2025)
von: Zhuang, Chen, et al.
Veröffentlicht: (2025)
CGSim: A Simulation Framework for Large Scale Distributed Computing Environment
von: Vatsavai, Sairam Sri, et al.
Veröffentlicht: (2025)
von: Vatsavai, Sairam Sri, et al.
Veröffentlicht: (2025)
Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study
von: McDonald, Jesse, et al.
Veröffentlicht: (2024)
von: McDonald, Jesse, et al.
Veröffentlicht: (2024)
Towards a Scalable and Efficient PGAS-based Distributed OpenMP
von: Shan, Baodi, et al.
Veröffentlicht: (2024)
von: Shan, Baodi, et al.
Veröffentlicht: (2024)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
von: Karfakis, George, et al.
Veröffentlicht: (2025)
von: Karfakis, George, et al.
Veröffentlicht: (2025)
Multi-DNN Inference of Sparse Models on Edge SoCs
von: Luo, Jiawei, et al.
Veröffentlicht: (2026)
von: Luo, Jiawei, et al.
Veröffentlicht: (2026)
Towards Universal Performance Modeling for Machine Learning Training on Multi-GPU Platforms
von: Lin, Zhongyi, et al.
Veröffentlicht: (2024)
von: Lin, Zhongyi, et al.
Veröffentlicht: (2024)
Phantora: Maximizing Code Reuse in Simulation-based Machine Learning System Performance Estimation
von: Qin, Jianxing, et al.
Veröffentlicht: (2025)
von: Qin, Jianxing, et al.
Veröffentlicht: (2025)
Beyond Thread States: Diagnosing Performance Degradation with eBPF and Thread Dynamics
von: Landau, Diogo, et al.
Veröffentlicht: (2026)
von: Landau, Diogo, et al.
Veröffentlicht: (2026)
Cyclic Data Streaming on GPUs for Short Range Stencils Applied to Molecular Dynamics
von: Rose, Martin, et al.
Veröffentlicht: (2025)
von: Rose, Martin, et al.
Veröffentlicht: (2025)
Matryoshka: Optimization of Dynamic Diverse Quantum Chemistry Systems via Elastic Parallelism Transformation
von: Wang, Tuowei, et al.
Veröffentlicht: (2024)
von: Wang, Tuowei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Tuning the Tuner: Introducing Hyperparameter Optimization for Auto-Tuning
von: Willemsen, Floris-Jan, et al.
Veröffentlicht: (2025) -
Accelerating Mobile Inference through Fine-Grained CPU-GPU Co-Execution
von: Li, Zhuojin, et al.
Veröffentlicht: (2025) -
CloudFormer: An Attention-based Performance Prediction for Public Clouds with Unknown Workload
von: Shahbazinia, Amirhossein, et al.
Veröffentlicht: (2025) -
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
von: Sun, Tingyang, et al.
Veröffentlicht: (2026) -
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
von: Titopoulos, Vasileios, et al.
Veröffentlicht: (2025)