A Sparsity Predicting Approach for Large Language Models via Activation Pattern Clustering
Fuente:
arXiv
Saved in:
| Main Authors: | Dhar, Nobel, Deng, Bobin, Islam, Md Romyull, Zhang, Xinyue, Nasif, Kazi Fahim Ahmad, Suo, Kun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Characterizing and Understanding Energy Footprint and Efficiency of Small Language Model on Edges
by: Islam, Md Romyull, et al.
Published: (2025)
by: Islam, Md Romyull, et al.
Published: (2025)
Optimizing CDN Architectures: Multi-Metric Algorithmic Breakthroughs for Edge and Distributed Performance
by: Absur, Md Nurul, et al.
Published: (2024)
by: Absur, Md Nurul, et al.
Published: (2024)
Activation Sparsity Opportunities for Compressing General Large Language Models
by: Dhar, Nobel, et al.
Published: (2024)
by: Dhar, Nobel, et al.
Published: (2024)
Performance Characterization of Distributed Deep Learning Strategies: A Quantitative Evaluation of DDP, FSDP, and Parameter Server Architectures on GPU Clusters
by: Ovi, Md Sultanul Islam
Published: (2025)
by: Ovi, Md Sultanul Islam
Published: (2025)
Enabling Dynamic Sparsity in Quantized LLM Inference
by: Wang, Rongxiang, et al.
Published: (2025)
by: Wang, Rongxiang, et al.
Published: (2025)
Predictable LLM Serving on GPU Clusters
by: Darzi, Erfan, et al.
Published: (2025)
by: Darzi, Erfan, et al.
Published: (2025)
Improving SpGEMM Performance Through Matrix Reordering and Cluster-wise Computation
by: Islam, Abdullah Al Raqibul, et al.
Published: (2025)
by: Islam, Abdullah Al Raqibul, et al.
Published: (2025)
Sparsity-Aware Roofline Models for Sparse Matrix-Matrix Multiplication
by: Qian, Matthew, et al.
Published: (2026)
by: Qian, Matthew, et al.
Published: (2026)
Optimizing Memory Allocation in Distributed Clusters with Predictive Modeling
by: Bader, Jonathan, et al.
Published: (2026)
by: Bader, Jonathan, et al.
Published: (2026)
Sparsity-Preserving Encodings for Straggler-Optimal Distributed Matrix Computations at the Edge
by: Das, Anindya Bijoy, et al.
Published: (2024)
by: Das, Anindya Bijoy, et al.
Published: (2024)
ConChain: A Scheme for Contention-free and Attack Resilient BlockChain
by: Bappy, Faisal Haque, et al.
Published: (2023)
by: Bappy, Faisal Haque, et al.
Published: (2023)
Characterizing Communication Patterns in Distributed Large Language Model Inference
by: Xu, Lang, et al.
Published: (2025)
by: Xu, Lang, et al.
Published: (2025)
A Deep Dive into the Google Cluster Workload Traces: Analyzing the Application Failure Characteristics and User Behaviors
by: Bappy, Faisal Haque, et al.
Published: (2023)
by: Bappy, Faisal Haque, et al.
Published: (2023)
Augur: Pre-Execution Energy Prediction for Workflow Tasks in Heterogeneous Clusters
by: West, Kathleen, et al.
Published: (2026)
by: West, Kathleen, et al.
Published: (2026)
Adaptive Parallel Downloader for Large Genomic Datasets
by: Swargo, Rasman Mubtasim, et al.
Published: (2025)
by: Swargo, Rasman Mubtasim, et al.
Published: (2025)
Mitigating Interference of Microservices with a Scoring Mechanism in Large-scale Clusters
by: Yang, Dingyu, et al.
Published: (2024)
by: Yang, Dingyu, et al.
Published: (2024)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
by: Liu, Di, et al.
Published: (2026)
by: Liu, Di, et al.
Published: (2026)
Evolving Topics in Federated Learning: Trends, and Emerging Directions for IS
by: Uddin, Md Raihan, et al.
Published: (2024)
by: Uddin, Md Raihan, et al.
Published: (2024)
Predicting the Performance of Scientific Workflow Tasks for Cluster Resource Management: An Overview of the State of the Art
by: Bader, Jonathan, et al.
Published: (2025)
by: Bader, Jonathan, et al.
Published: (2025)
DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-based Clusters
by: Bai, Haoyu, et al.
Published: (2024)
by: Bai, Haoyu, et al.
Published: (2024)
Fantasy: Efficient Large-scale Vector Search on GPU Clusters with GPUDirect Async
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
PRISM: Dynamic Primitive-Based Forecasting for Large-Scale GPU Cluster Workloads
by: Wu, Xin, et al.
Published: (2026)
by: Wu, Xin, et al.
Published: (2026)
MSAO: Adaptive Modality Sparsity-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
by: Yang, Zheming, et al.
Published: (2026)
by: Yang, Zheming, et al.
Published: (2026)
Looking for (Genomic) Needles in a Haystack: Sparsity-Driven Search for Identifying Correlated Genetic Mutations in Cancer
by: Prabhu, Ritvik, et al.
Published: (2026)
by: Prabhu, Ritvik, et al.
Published: (2026)
GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
by: Duan, Jiaang, et al.
Published: (2025)
by: Duan, Jiaang, et al.
Published: (2025)
Selection Guidelines for Geo-Replicated SMR Protocols: A Communication Pattern-based Latency Modeling Approach
by: Shiozaki, Kohya, et al.
Published: (2024)
by: Shiozaki, Kohya, et al.
Published: (2024)
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
by: Zhang, Mingjun, et al.
Published: (2025)
by: Zhang, Mingjun, et al.
Published: (2025)
Understanding Large-Scale HPC System Behavior Through Cluster-Based Visual Analytics
by: Austin, Allison, et al.
Published: (2026)
by: Austin, Allison, et al.
Published: (2026)
C-Koordinator: Interference-aware Management for Large-scale and Co-located Microservice Clusters
by: Song, Shengye, et al.
Published: (2025)
by: Song, Shengye, et al.
Published: (2025)
FASTEN: Towards a FAult-tolerant and STorage EfficieNt Cloud: Balancing Between Replication and Deduplication
by: Ahmed, Sabbir, et al.
Published: (2023)
by: Ahmed, Sabbir, et al.
Published: (2023)
Resilient Packet Forwarding: A Reinforcement Learning Approach to Routing in Gaussian Interconnected Networks with Clustered Faults
by: Charrwi, Mohammad Walid, et al.
Published: (2025)
by: Charrwi, Mohammad Walid, et al.
Published: (2025)
AIReSim: A Discrete Event Simulator for Large-scale AI Cluster Reliability Modeling
by: Pattabiraman, Karthik, et al.
Published: (2026)
by: Pattabiraman, Karthik, et al.
Published: (2026)
DSV: Exploiting Dynamic Sparsity to Accelerate Large-Scale Video DiT Training
by: Tan, Xin, et al.
Published: (2025)
by: Tan, Xin, et al.
Published: (2025)
Exploring Novel Data Storage Approaches for Large-Scale Numerical Weather Prediction
by: Gil, Nicolau Manubens
Published: (2026)
by: Gil, Nicolau Manubens
Published: (2026)
Efficiently Reproducing Distributed Workflows in Notebook-based Systems
by: Azaz, Talha, et al.
Published: (2026)
by: Azaz, Talha, et al.
Published: (2026)
The Robustness of Spiking Neural Networks in Federated Learning with Compression Against Non-omniscient Byzantine Attacks
by: Nguyen, Manh V., et al.
Published: (2025)
by: Nguyen, Manh V., et al.
Published: (2025)
Local problems in trees across a wide range of distributed models
by: Dhar, Anubhav, et al.
Published: (2024)
by: Dhar, Anubhav, et al.
Published: (2024)
ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
by: Xiang, Yuxing, et al.
Published: (2025)
by: Xiang, Yuxing, et al.
Published: (2025)
CrashEventLLM: Predicting System Crashes with Large Language Models
by: Mudgal, Priyanka, et al.
Published: (2024)
by: Mudgal, Priyanka, et al.
Published: (2024)
H2:Towards Efficient Large-Scale LLM Training on Hyper-Heterogeneous Cluster over 1,000 Chips
by: Tang, Ding, et al.
Published: (2025)
by: Tang, Ding, et al.
Published: (2025)
Similar Items
-
Characterizing and Understanding Energy Footprint and Efficiency of Small Language Model on Edges
by: Islam, Md Romyull, et al.
Published: (2025) -
Optimizing CDN Architectures: Multi-Metric Algorithmic Breakthroughs for Edge and Distributed Performance
by: Absur, Md Nurul, et al.
Published: (2024) -
Activation Sparsity Opportunities for Compressing General Large Language Models
by: Dhar, Nobel, et al.
Published: (2024) -
Performance Characterization of Distributed Deep Learning Strategies: A Quantitative Evaluation of DDP, FSDP, and Parameter Server Architectures on GPU Clusters
by: Ovi, Md Sultanul Islam
Published: (2025) -
Enabling Dynamic Sparsity in Quantized LLM Inference
by: Wang, Rongxiang, et al.
Published: (2025)