Long-term Monitoring of Kernel and Hardware Events to Understand Latency Variance
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Fang, Huang, Yuyang, Yu, Miao, Ma, Sixiang, Liu, Tongping, Wang, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaler: Efficient and Effective Cross Flow Analysis
by: Steven, et al.
Published: (2024)
by: Steven, et al.
Published: (2024)
EDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPC
by: Shen, Siyuan, et al.
Published: (2025)
by: Shen, Siyuan, et al.
Published: (2025)
PANDA: Noise-Resilient Antagonist Identification in Production Datacenters
by: Zhou, Sixiang, et al.
Published: (2025)
by: Zhou, Sixiang, et al.
Published: (2025)
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
by: Chen, Feiyang, et al.
Published: (2025)
by: Chen, Feiyang, et al.
Published: (2025)
ProTrain: Efficient LLM Training via Memory-Aware Techniques
by: Yang, Hanmei, et al.
Published: (2024)
by: Yang, Hanmei, et al.
Published: (2024)
Data-Driven Power Modeling and Monitoring via Hardware Performance Counter Tracking
by: Mazzola, Sergio, et al.
Published: (2025)
by: Mazzola, Sergio, et al.
Published: (2025)
Data-Driven Power Modeling and Monitoring via Hardware Performance Counters Tracking
by: Mazzola, Sergio, et al.
Published: (2024)
by: Mazzola, Sergio, et al.
Published: (2024)
Latency and Privacy-Aware Resource Allocation in Vehicular Edge Computing
by: Ahmadvand, Hossein, et al.
Published: (2025)
by: Ahmadvand, Hossein, et al.
Published: (2025)
Reducing Latency of LLM Search Agent via Speculation-based Algorithm-System Co-Design
by: Huang, Zixiao, et al.
Published: (2025)
by: Huang, Zixiao, et al.
Published: (2025)
Mosaic: Cross-Modal Clustering for Efficient Video Understanding
by: Wang, Tuowei, et al.
Published: (2026)
by: Wang, Tuowei, et al.
Published: (2026)
PM2Lat: Highly Accurate and Generalized Prediction of DNN Execution Latency on GPUs
by: Le, Truong-Thanh, et al.
Published: (2026)
by: Le, Truong-Thanh, et al.
Published: (2026)
A Latency-Constrained, Gated Recurrent Unit (GRU) Implementation in the Versal AI Engine
by: Sapkas, M., et al.
Published: (2025)
by: Sapkas, M., et al.
Published: (2025)
Multi-Strided Access Patterns to Boost Hardware Prefetching
by: Blom, Miguel O., et al.
Published: (2024)
by: Blom, Miguel O., et al.
Published: (2024)
KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
by: Wang, Han, et al.
Published: (2026)
by: Wang, Han, et al.
Published: (2026)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
by: Islam, Tanzima Z., et al.
Published: (2024)
by: Islam, Tanzima Z., et al.
Published: (2024)
Layer Importance and Hallucination Analysis in Large Language Models via Enhanced Activation Variance-Sparsity
by: Song, Zichen, et al.
Published: (2024)
by: Song, Zichen, et al.
Published: (2024)
WaveTune: Wave-aware Bilinear Modeling for Efficient GPU Kernel Auto-tuning
by: Zhang, Kaixuan, et al.
Published: (2026)
by: Zhang, Kaixuan, et al.
Published: (2026)
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
by: Wang, Haoxin, et al.
Published: (2025)
by: Wang, Haoxin, et al.
Published: (2025)
Motion-to-Motion Latency Measurement Framework for Connected and Autonomous Vehicle Teleoperation
by: Provost, François, et al.
Published: (2025)
by: Provost, François, et al.
Published: (2025)
An Experimental Study of Low-Latency Video Streaming over 5G
by: Khan, Imran, et al.
Published: (2024)
by: Khan, Imran, et al.
Published: (2024)
CounterPoint: Using Hardware Event Counters to Refute and Refine Microarchitectural Assumptions (Extended Version)
by: Lindsay, Nick, et al.
Published: (2026)
by: Lindsay, Nick, et al.
Published: (2026)
Latency Based Tiling
by: Cashman, Jack
Published: (2025)
by: Cashman, Jack
Published: (2025)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
by: Kong, Linghao, et al.
Published: (2026)
by: Kong, Linghao, et al.
Published: (2026)
Analysis and Evaluation of Using Microsecond-Latency Memory for In-Memory Indices and Caches in SSD-Based Key-Value Stores
by: Bando, Yosuke, et al.
Published: (2025)
by: Bando, Yosuke, et al.
Published: (2025)
Cloud Computing Energy Consumption Prediction Based on Kernel Extreme Learning Machine Algorithm Improved by Vector Weighted Average Algorithm
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
Numerical Kernels on a Spatial Accelerator: A Study of Tenstorrent Wormhole
by: Taylor, Maya, et al.
Published: (2026)
by: Taylor, Maya, et al.
Published: (2026)
Personalized Model-Based Design of Human Centric AI enabled CPS for Long term usage
by: Ngabonziza, Bernard, et al.
Published: (2026)
by: Ngabonziza, Bernard, et al.
Published: (2026)
AutoKernel: Autonomous GPU Kernel Optimization via Iterative Agent-Driven Search
by: Jaber, Jaber, et al.
Published: (2026)
by: Jaber, Jaber, et al.
Published: (2026)
Dissecting RISC-V Performance: Practical PMU Profiling and Hardware-Agnostic Roofline Analysis on Emerging Platforms
by: Batashev, Alexander
Published: (2025)
by: Batashev, Alexander
Published: (2025)
eHashPipe: Lightweight Top-K and Per-PID Resource Monitoring with eBPF
by: Dai, Yuanjun, et al.
Published: (2025)
by: Dai, Yuanjun, et al.
Published: (2025)
Can Increasing the Hit Ratio Hurt Cache Throughput? (Long Version)
by: Qiu, Ziyue, et al.
Published: (2024)
by: Qiu, Ziyue, et al.
Published: (2024)
Hardware optimization on Android for inference of AI models
by: Gherasim, Iulius, et al.
Published: (2025)
by: Gherasim, Iulius, et al.
Published: (2025)
Lightweight Software Kernels and Hardware Extensions for Efficient Sparse Deep Neural Networks on Microcontrollers
by: Daghero, Francesco, et al.
Published: (2025)
by: Daghero, Francesco, et al.
Published: (2025)
gpu tracker: Python Package for Tracking and Profiling GPU and Other Hardware Utilization in Both Desktop and High-Performance Computing Environments
by: Huckvale, Erik D., et al.
Published: (2024)
by: Huckvale, Erik D., et al.
Published: (2024)
FALCON: Feedback-driven Adaptive Long/short-term memory reinforced Coding Optimization system
by: Li, Zeyuan, et al.
Published: (2024)
by: Li, Zeyuan, et al.
Published: (2024)
Exploring Topologies in Quantum Annealing: A Hardware-Aware Perspective
by: Bifulco, Mario, et al.
Published: (2025)
by: Bifulco, Mario, et al.
Published: (2025)
Hardware Acceleration for Knowledge Graph Processing: Challenges & Recent Developments
by: Besta, Maciej, et al.
Published: (2024)
by: Besta, Maciej, et al.
Published: (2024)
From 8 Seconds to 370ms: Kernel-Fused SAR Imaging on Apple Silicon via Single-Dispatch FFT Pipelines
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
by: Zhang, Lingqi, et al.
Published: (2025)
by: Zhang, Lingqi, et al.
Published: (2025)
Heuristic-Based Merging of HPC Traces to Extend Hardware Counter Coverage
by: Aubach, Júlia Orteu, et al.
Published: (2026)
by: Aubach, Júlia Orteu, et al.
Published: (2026)
Similar Items
-
Scaler: Efficient and Effective Cross Flow Analysis
by: Steven, et al.
Published: (2024) -
EDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPC
by: Shen, Siyuan, et al.
Published: (2025) -
PANDA: Noise-Resilient Antagonist Identification in Production Datacenters
by: Zhou, Sixiang, et al.
Published: (2025) -
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
by: Chen, Feiyang, et al.
Published: (2025) -
ProTrain: Efficient LLM Training via Memory-Aware Techniques
by: Yang, Hanmei, et al.
Published: (2024)