Heuristic-Based Merging of HPC Traces to Extend Hardware Counter Coverage
Fuente:
arXiv
Salvato in:
| Autori principali: | Aubach, Júlia Orteu, Banchelli, Fabio, Ramírez, Marc Clascà, Garcia-Gasulla, Marta |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Hardware-Agnostic and Insightful Efficiency Metrics for Accelerated Systems: Definition and Implementation within TALP
di: Rahimi, Ghazal, et al.
Pubblicazione: (2026)
di: Rahimi, Ghazal, et al.
Pubblicazione: (2026)
Hardware optimization on Android for inference of AI models
di: Gherasim, Iulius, et al.
Pubblicazione: (2025)
di: Gherasim, Iulius, et al.
Pubblicazione: (2025)
Hardware-efficient tractable probabilistic inference for TinyML Neurosymbolic AI applications
di: Leslin, Jelin, et al.
Pubblicazione: (2025)
di: Leslin, Jelin, et al.
Pubblicazione: (2025)
HPC Application Parameter Autotuning on Edge Devices: A Bandit Learning Approach
di: Hossain, Abrar, et al.
Pubblicazione: (2025)
di: Hossain, Abrar, et al.
Pubblicazione: (2025)
TESSERACT: Eliminating Experimental Bias in Malware Classification across Space and Time (Extended Version)
di: Kan, Zeliang, et al.
Pubblicazione: (2024)
di: Kan, Zeliang, et al.
Pubblicazione: (2024)
CounterPoint: Using Hardware Event Counters to Refute and Refine Microarchitectural Assumptions (Extended Version)
di: Lindsay, Nick, et al.
Pubblicazione: (2026)
di: Lindsay, Nick, et al.
Pubblicazione: (2026)
Introducing MareNostrum5: A European pre-exascale energy-efficient system designed to serve a broad spectrum of scientific workloads
di: Banchelli, Fabio, et al.
Pubblicazione: (2025)
di: Banchelli, Fabio, et al.
Pubblicazione: (2025)
Toward A Formalized Approach for Spike Sorting Algorithms and Hardware Evaluation
di: Zhang, Tim, et al.
Pubblicazione: (2022)
di: Zhang, Tim, et al.
Pubblicazione: (2022)
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
di: Chen, Feiyang, et al.
Pubblicazione: (2025)
di: Chen, Feiyang, et al.
Pubblicazione: (2025)
A Practical Two-Stage Framework for GPU Resource and Power Prediction in Heterogeneous HPC Systems
di: Oztop, Beste, et al.
Pubblicazione: (2026)
di: Oztop, Beste, et al.
Pubblicazione: (2026)
REAM: Merging Improves Pruning of Experts in LLMs
di: Jha, Saurav, et al.
Pubblicazione: (2026)
di: Jha, Saurav, et al.
Pubblicazione: (2026)
TALP-Pages: An easy-to-integrate continuous performance monitoring framework
di: Seitz, Valentin, et al.
Pubblicazione: (2025)
di: Seitz, Valentin, et al.
Pubblicazione: (2025)
Conformer-Based Speech Recognition On Extreme Edge-Computing Devices
di: Xu, Mingbin, et al.
Pubblicazione: (2023)
di: Xu, Mingbin, et al.
Pubblicazione: (2023)
A Kernel-Based Approach for Accurate Steady-State Detection in Performance Time Series
di: Beseda, Martin, et al.
Pubblicazione: (2025)
di: Beseda, Martin, et al.
Pubblicazione: (2025)
Automating Energy-Efficient GPU Kernel Generation: A Fast Search-Based Compilation Approach
di: Zhang, Yijia, et al.
Pubblicazione: (2024)
di: Zhang, Yijia, et al.
Pubblicazione: (2024)
Cloud Computing Energy Consumption Prediction Based on Kernel Extreme Learning Machine Algorithm Improved by Vector Weighted Average Algorithm
di: Wang, Yuqing, et al.
Pubblicazione: (2025)
di: Wang, Yuqing, et al.
Pubblicazione: (2025)
Batched DGEMMs for scientific codes running on long vector architectures
di: Banchelli, Fabio, et al.
Pubblicazione: (2025)
di: Banchelli, Fabio, et al.
Pubblicazione: (2025)
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
di: Patwari, Rajeev, et al.
Pubblicazione: (2025)
di: Patwari, Rajeev, et al.
Pubblicazione: (2025)
Exploiting long vectors with a CFD code: a co-design show case
di: Blancafort, Marc, et al.
Pubblicazione: (2024)
di: Blancafort, Marc, et al.
Pubblicazione: (2024)
Optimizing Methane Detection On Board Satellites: Speed, Accuracy, and Low-Power Solutions for Resource-Constrained Hardware
di: Herec, Jonáš, et al.
Pubblicazione: (2025)
di: Herec, Jonáš, et al.
Pubblicazione: (2025)
Foundations of the Theory of Performance-Based Ranking
di: Piérard, Sébastien, et al.
Pubblicazione: (2024)
di: Piérard, Sébastien, et al.
Pubblicazione: (2024)
MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces
di: Sridharan, Srinivas, et al.
Pubblicazione: (2026)
di: Sridharan, Srinivas, et al.
Pubblicazione: (2026)
GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models
di: Taneja, Maanas, et al.
Pubblicazione: (2026)
di: Taneja, Maanas, et al.
Pubblicazione: (2026)
MarginGate: Sparse Margin-Triggered Verification for Batch-Invariant LLM Inference
di: Chu, Kexin, et al.
Pubblicazione: (2026)
di: Chu, Kexin, et al.
Pubblicazione: (2026)
Reducing Compute Waste in LLMs through Kernel-Level DVFS
di: Spaan, Jeffrey, et al.
Pubblicazione: (2026)
di: Spaan, Jeffrey, et al.
Pubblicazione: (2026)
KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
di: Wang, Han, et al.
Pubblicazione: (2026)
di: Wang, Han, et al.
Pubblicazione: (2026)
Single-Thread JPEG Decoder Benchmarks Mis-Evaluate ML Data Loaders
di: Iglovikov, Vladimir, et al.
Pubblicazione: (2026)
di: Iglovikov, Vladimir, et al.
Pubblicazione: (2026)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
di: Kong, Linghao, et al.
Pubblicazione: (2026)
di: Kong, Linghao, et al.
Pubblicazione: (2026)
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
di: Ziller, Thomas, et al.
Pubblicazione: (2026)
di: Ziller, Thomas, et al.
Pubblicazione: (2026)
Approximating Uniform Random Rotations by Two-Block Structured Hadamard Rotations in High Dimensions
di: Zilca, Tomer, et al.
Pubblicazione: (2026)
di: Zilca, Tomer, et al.
Pubblicazione: (2026)
EARL: Energy-Aware Optimization of Liquid State Machines for Pervasive AI
di: Iqbal, Zain, et al.
Pubblicazione: (2026)
di: Iqbal, Zain, et al.
Pubblicazione: (2026)
Large-Scale Data Parallelization of Product Quantization and Inverted Indexing Using Dask
di: Abraham, Ashley N., et al.
Pubblicazione: (2026)
di: Abraham, Ashley N., et al.
Pubblicazione: (2026)
AutoKernel: Autonomous GPU Kernel Optimization via Iterative Agent-Driven Search
di: Jaber, Jaber, et al.
Pubblicazione: (2026)
di: Jaber, Jaber, et al.
Pubblicazione: (2026)
A Scalable k-Medoids Clustering via Whale Optimization Algorithm
di: Chenan, Huang, et al.
Pubblicazione: (2024)
di: Chenan, Huang, et al.
Pubblicazione: (2024)
CPINN-ABPI: Physics-Informed Neural Networks for Accurate Power Estimation in MPSoCs
di: Elshamy, Mohamed R., et al.
Pubblicazione: (2025)
di: Elshamy, Mohamed R., et al.
Pubblicazione: (2025)
DistZO2: High-Throughput and Memory-Efficient Zeroth-Order Fine-tuning LLMs with Distributed Parallel Computing
di: Wang, Liangyu, et al.
Pubblicazione: (2025)
di: Wang, Liangyu, et al.
Pubblicazione: (2025)
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
di: An, Zihao, et al.
Pubblicazione: (2025)
di: An, Zihao, et al.
Pubblicazione: (2025)
Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants
di: You, Bozhi, et al.
Pubblicazione: (2025)
di: You, Bozhi, et al.
Pubblicazione: (2025)
Parallel Implementations Assessment of a Spatial-Spectral Classifier for Hyperspectral Clinical Applications
di: Lazcano, Raquel, et al.
Pubblicazione: (2024)
di: Lazcano, Raquel, et al.
Pubblicazione: (2024)
A Structure-Aware Framework for Learning Device Placements on Computation Graphs
di: Duan, Shukai, et al.
Pubblicazione: (2024)
di: Duan, Shukai, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Hardware-Agnostic and Insightful Efficiency Metrics for Accelerated Systems: Definition and Implementation within TALP
di: Rahimi, Ghazal, et al.
Pubblicazione: (2026) -
Hardware optimization on Android for inference of AI models
di: Gherasim, Iulius, et al.
Pubblicazione: (2025) -
Hardware-efficient tractable probabilistic inference for TinyML Neurosymbolic AI applications
di: Leslin, Jelin, et al.
Pubblicazione: (2025) -
HPC Application Parameter Autotuning on Edge Devices: A Bandit Learning Approach
di: Hossain, Abrar, et al.
Pubblicazione: (2025) -
TESSERACT: Eliminating Experimental Bias in Malware Classification across Space and Time (Extended Version)
di: Kan, Zeliang, et al.
Pubblicazione: (2024)