Modeling Utilization to Identify Shared-Memory Atomic Bottlenecks
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Dong, Rongcui, Pai, Sreepathi |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Characterizing VLA Models: Identifying the Action Generation Bottleneck for Edge AI Architectures
par: Vishwanathan, Manoj, et autres
Publié: (2026)
par: Vishwanathan, Manoj, et autres
Publié: (2026)
Tuning Fast Memory Size based on Modeling of Page Migration for Tiered Memory
par: Chen, Shangye, et autres
Publié: (2024)
par: Chen, Shangye, et autres
Publié: (2024)
Noise Injection for__Performance Bottleneck Analysis
par: Delval, Aurélien, et autres
Publié: (2025)
par: Delval, Aurélien, et autres
Publié: (2025)
An Analysis of Performance Bottlenecks in MRI Pre-Processing
par: Dugré, Mathieu, et autres
Publié: (2024)
par: Dugré, Mathieu, et autres
Publié: (2024)
Memshare: Memory Sharing for Multicore Computation in R with an Application to Feature Selection by Mutual Information using PDE
par: Thrun, Michael C., et autres
Publié: (2025)
par: Thrun, Michael C., et autres
Publié: (2025)
Prefill vs. Decode Bottlenecks: SRAM-Frequency Tradeoffs and the Memory-Bandwidth Ceiling
par: Atmer, Hannah, et autres
Publié: (2025)
par: Atmer, Hannah, et autres
Publié: (2025)
Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered Memory
par: Ren, Jie, et autres
Publié: (2025)
par: Ren, Jie, et autres
Publié: (2025)
LLload: An Easy-to-Use HPC Utilization Tool
par: Byun, Chansup, et autres
Publié: (2024)
par: Byun, Chansup, et autres
Publié: (2024)
gigiProfiler: Diagnosing Performance Issues by Uncovering Application Resource Bottlenecks
par: Hu, Yigong, et autres
Publié: (2025)
par: Hu, Yigong, et autres
Publié: (2025)
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services
par: Tang, Lingfeng, et autres
Publié: (2025)
par: Tang, Lingfeng, et autres
Publié: (2025)
Heterogeneous Memory Pool Tuning
par: Vaverka, Filip, et autres
Publié: (2025)
par: Vaverka, Filip, et autres
Publié: (2025)
Analysis and Mitigation of Shared Resource Contention on Heterogeneous Multicore: An Industrial Case Study
par: Bechtel, Michael, et autres
Publié: (2023)
par: Bechtel, Michael, et autres
Publié: (2023)
Updates on the Low-Level Abstraction of Memory Access
par: Gruber, Bernhard Manfred
Publié: (2023)
par: Gruber, Bernhard Manfred
Publié: (2023)
SparseX: Efficient Segment-Level KV Cache Sharing for Interleaved LLM Serving
par: Zhang, Quqing, et autres
Publié: (2026)
par: Zhang, Quqing, et autres
Publié: (2026)
Evaluating the Performance of the DeepSeek Model in Confidential Computing Environment
par: Dong, Ben, et autres
Publié: (2025)
par: Dong, Ben, et autres
Publié: (2025)
A Few Fit Most: Improving Performance Portability of SGEMM on GPUs using Multi-Versioning
par: Hochgraf, Robert, et autres
Publié: (2025)
par: Hochgraf, Robert, et autres
Publié: (2025)
Atys: An Efficient Profiling Framework for Identifying Hotspot Functions in Large-scale Cloud Microservices
par: Sun, Jiaqi, et autres
Publié: (2025)
par: Sun, Jiaqi, et autres
Publié: (2025)
Memory Analysis on the Training Course of DeepSeek Models
par: Zhang, Ping, et autres
Publié: (2025)
par: Zhang, Ping, et autres
Publié: (2025)
EDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPC
par: Shen, Siyuan, et autres
Publié: (2025)
par: Shen, Siyuan, et autres
Publié: (2025)
gpu tracker: Python Package for Tracking and Profiling GPU and Other Hardware Utilization in Both Desktop and High-Performance Computing Environments
par: Huckvale, Erik D., et autres
Publié: (2024)
par: Huckvale, Erik D., et autres
Publié: (2024)
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
par: Arif, Moiz, et autres
Publié: (2026)
par: Arif, Moiz, et autres
Publié: (2026)
WritePolicyBench: Benchmarking Memory Write Policies under Byte Budgets
par: Cham, Edgard El
Publié: (2026)
par: Cham, Edgard El
Publié: (2026)
Analysis and Evaluation of Using Microsecond-Latency Memory for In-Memory Indices and Caches in SSD-Based Key-Value Stores
par: Bando, Yosuke, et autres
Publié: (2025)
par: Bando, Yosuke, et autres
Publié: (2025)
Examem: Low-Overhead Memory Instrumentation for Intelligent Memory Systems
par: Poduval, Ashwin, et autres
Publié: (2024)
par: Poduval, Ashwin, et autres
Publié: (2024)
Shared Memory-contention-aware Concurrent DNN Execution for Diversely Heterogeneous System-on-Chips
par: Dagli, Ismet, et autres
Publié: (2023)
par: Dagli, Ismet, et autres
Publié: (2023)
Mitigating GIL Bottlenecks in Edge AI Systems
par: Mandal, Mridankan, et autres
Publié: (2026)
par: Mandal, Mridankan, et autres
Publié: (2026)
OMPILOT: Harnessing Transformer Models for Auto Parallelization to Shared Memory Computing Paradigms
par: Bhattacharjee, Arijit, et autres
Publié: (2025)
par: Bhattacharjee, Arijit, et autres
Publié: (2025)
HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing
par: Huang, Haochen, et autres
Publié: (2025)
par: Huang, Haochen, et autres
Publié: (2025)
Heterogeneous Memory Benchmarking Toolkit
par: Ghaemi, Golsana, et autres
Publié: (2025)
par: Ghaemi, Golsana, et autres
Publié: (2025)
Optimizing System Memory Bandwidth with Micron CXL Memory Expansion Modules on Intel Xeon 6 Processors
par: Sehgal, Rohit, et autres
Publié: (2024)
par: Sehgal, Rohit, et autres
Publié: (2024)
ZO2: Scalable Zeroth-Order Fine-Tuning for Extremely Large Language Models with Limited GPU Memory
par: Wang, Liangyu, et autres
Publié: (2025)
par: Wang, Liangyu, et autres
Publié: (2025)
A Controlled Study of Memory Hierarchy Transitions in Quantum Circuit Simulation on Apple M4 Pro Unified Memory Architecture
par: Pratipat, Gyan
Publié: (2026)
par: Pratipat, Gyan
Publié: (2026)
Virtual-Memory Powersort
par: Moltmann, Finn, et autres
Publié: (2026)
par: Moltmann, Finn, et autres
Publié: (2026)
Impact of Data-Oriented and Object-Oriented Design on Performance and Cache Utilization with Artificial Intelligence Algorithms in Multi-Threaded CPUs
par: Arantes, Gabriel M., et autres
Publié: (2025)
par: Arantes, Gabriel M., et autres
Publié: (2025)
Performance Characterization of AutoNUMA Memory Tiering on Graph Analytics
par: Moura, Diego, et autres
Publié: (2022)
par: Moura, Diego, et autres
Publié: (2022)
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
par: Shao, Zishan, et autres
Publié: (2025)
par: Shao, Zishan, et autres
Publié: (2025)
ETM2: Empowering Traditional Memory Bandwidth Regulation using ETM
par: Zuepke, Alexander, et autres
Publié: (2026)
par: Zuepke, Alexander, et autres
Publié: (2026)
Delegation with Trust<T>: A Scalable, Type- and Memory-Safe Alternative to Locks
par: Ahmad, Noaman, et autres
Publié: (2024)
par: Ahmad, Noaman, et autres
Publié: (2024)
A$^3$PIM: An Automated, Analytic and Accurate Processing-in-Memory Offloader
par: Jiang, Qingcai, et autres
Publié: (2024)
par: Jiang, Qingcai, et autres
Publié: (2024)
Modeling Interfering Sources in Shared Queues for Timely Computations in Edge Computing Systems
par: Akar, Nail, et autres
Publié: (2024)
par: Akar, Nail, et autres
Publié: (2024)
Documents similaires
-
Characterizing VLA Models: Identifying the Action Generation Bottleneck for Edge AI Architectures
par: Vishwanathan, Manoj, et autres
Publié: (2026) -
Tuning Fast Memory Size based on Modeling of Page Migration for Tiered Memory
par: Chen, Shangye, et autres
Publié: (2024) -
Noise Injection for__Performance Bottleneck Analysis
par: Delval, Aurélien, et autres
Publié: (2025) -
An Analysis of Performance Bottlenecks in MRI Pre-Processing
par: Dugré, Mathieu, et autres
Publié: (2024) -
Memshare: Memory Sharing for Multicore Computation in R with an Application to Feature Selection by Mutual Information using PDE
par: Thrun, Michael C., et autres
Publié: (2025)