Guardado en:
| Autores principales: | Wu, Xingfu, Oli, Tupendra, Qian, Justin H., Taylor, Valerie, Hersam, Mark C., Sangwan, Vinod K. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2406.18445 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Co-Design of 2D Heterojunctions for Data Filtering in Tracking Systems
por: Oli, Tupendra, et al.
Publicado: (2024)
por: Oli, Tupendra, et al.
Publicado: (2024)
Integrating ytopt and libEnsemble to Autotune OpenMC
por: Wu, Xingfu, et al.
Publicado: (2024)
por: Wu, Xingfu, et al.
Publicado: (2024)
Generalizing Scaling Laws for Dense and Sparse Large Language Models
por: Hossain, Md Arafat, et al.
Publicado: (2025)
por: Hossain, Md Arafat, et al.
Publicado: (2025)
HPC Application Parameter Autotuning on Edge Devices: A Bandit Learning Approach
por: Hossain, Abrar, et al.
Publicado: (2025)
por: Hossain, Abrar, et al.
Publicado: (2025)
Machine Learning-driven Autotuning of Graphics Processing Unit Accelerated Computational Fluid Dynamics for Enhanced Performance
por: Xue, Weicheng, et al.
Publicado: (2023)
por: Xue, Weicheng, et al.
Publicado: (2023)
PixelBrax: Learning Continuous Control from Pixels End-to-End on the GPU
por: McInroe, Trevor, et al.
Publicado: (2025)
por: McInroe, Trevor, et al.
Publicado: (2025)
XLB: A High Performance Layer-7 Load Balancer for Microservices using eBPF-based In-kernel Interposition
por: Wang, Yuejie, et al.
Publicado: (2026)
por: Wang, Yuejie, et al.
Publicado: (2026)
SLO-Guard: Crash-Aware, Budget-Consistent Autotuning for SLO-Constrained LLM Serving
por: Lysenstøen, Christian
Publicado: (2026)
por: Lysenstøen, Christian
Publicado: (2026)
FlipFlop: A Static Analysis-based Energy Optimization Framework for GPU Kernels
por: Rajput, Saurabhsingh, et al.
Publicado: (2026)
por: Rajput, Saurabhsingh, et al.
Publicado: (2026)
Opal: A Modular Framework for Optimizing Performance using Analytics and LLMs
por: Zaeed, Mohammad, et al.
Publicado: (2025)
por: Zaeed, Mohammad, et al.
Publicado: (2025)
Atys: An Efficient Profiling Framework for Identifying Hotspot Functions in Large-scale Cloud Microservices
por: Sun, Jiaqi, et al.
Publicado: (2025)
por: Sun, Jiaqi, et al.
Publicado: (2025)
A Modular Graph-Native Query Optimization Framework
por: Lyu, Bingqing, et al.
Publicado: (2024)
por: Lyu, Bingqing, et al.
Publicado: (2024)
Accelerating Transistor-Level Simulation of Integrated Circuits via Equivalence of RC Long-Chain Structures
por: Tang, Ruibai, et al.
Publicado: (2025)
por: Tang, Ruibai, et al.
Publicado: (2025)
Optimizing Cloud-native Services with SAGA: A Service Affinity Graph-based Approach
por: Dinh-Tuan, Hai, et al.
Publicado: (2025)
por: Dinh-Tuan, Hai, et al.
Publicado: (2025)
Optimas: An Intelligent Analytics-Informed Generative AI Framework for Performance Optimization
por: Zaeed, Mohammad, et al.
Publicado: (2026)
por: Zaeed, Mohammad, et al.
Publicado: (2026)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
por: Zhang, Li, et al.
Publicado: (2025)
por: Zhang, Li, et al.
Publicado: (2025)
Evaluating the Performance of the DeepSeek Model in Confidential Computing Environment
por: Dong, Ben, et al.
Publicado: (2025)
por: Dong, Ben, et al.
Publicado: (2025)
From Profiling to Optimization: Unveiling the Profile Guided Optimization
por: Liu, Bingxin, et al.
Publicado: (2025)
por: Liu, Bingxin, et al.
Publicado: (2025)
A Dataset of Performance Measurements and Alerts from Mozilla (Data Artifact)
por: Besbes, Mohamed Bilel, et al.
Publicado: (2025)
por: Besbes, Mohamed Bilel, et al.
Publicado: (2025)
Energy-Efficient Transformer Inference: Optimization Strategies for Time Series Classification
por: Kermani, Arshia, et al.
Publicado: (2025)
por: Kermani, Arshia, et al.
Publicado: (2025)
GreedySnake: Accelerating SSD-Offloaded LLM Training with Efficient Scheduling and Optimizer Step Overlapping
por: Yin, Yishu, et al.
Publicado: (2025)
por: Yin, Yishu, et al.
Publicado: (2025)
A Tale of Three Location Trackers: AirTag, SmartTag, and Tile
por: Jang, HyunSeok Daniel, et al.
Publicado: (2025)
por: Jang, HyunSeok Daniel, et al.
Publicado: (2025)
ytopt: Autotuning Scientific Applications for Energy Efficiency at Large Scales
por: Xingfu Wu, et al.
Publicado: (2024)
por: Xingfu Wu, et al.
Publicado: (2024)
Numerical Kernels on a Spatial Accelerator: A Study of Tenstorrent Wormhole
por: Taylor, Maya, et al.
Publicado: (2026)
por: Taylor, Maya, et al.
Publicado: (2026)
LOOPerSet: A Large-Scale Dataset for Data-Driven Polyhedral Compiler Optimization
por: Merouani, Massinissa, et al.
Publicado: (2025)
por: Merouani, Massinissa, et al.
Publicado: (2025)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
por: Wang, Yuxin, et al.
Publicado: (2024)
por: Wang, Yuxin, et al.
Publicado: (2024)
PORTAL: Controllable Landscape Generator for Continuous Optimization-Part I: Framework
por: Yazdani, Danial, et al.
Publicado: (2025)
por: Yazdani, Danial, et al.
Publicado: (2025)
Two Criteria for Performance Analysis of Optimization Algorithms
por: Jing, Yunpeng, et al.
Publicado: (2024)
por: Jing, Yunpeng, et al.
Publicado: (2024)
Performance Characterization and Optimizations of Traditional ML Applications
por: Kumar, Harsh, et al.
Publicado: (2024)
por: Kumar, Harsh, et al.
Publicado: (2024)
Optimizing Winograd Convolution on ARMv8 processors
por: Gui, Haoyuan, et al.
Publicado: (2024)
por: Gui, Haoyuan, et al.
Publicado: (2024)
AR-PPF: Advanced Resolution-Based Pixel Preemption Data Filtering for Efficient Time-Series Data Analysis
por: Kim, Taewoong, et al.
Publicado: (2024)
por: Kim, Taewoong, et al.
Publicado: (2024)
A Zoned Storage Optimized Flash Cache on ZNS SSDs
por: Yang, Chongzhuo, et al.
Publicado: (2024)
por: Yang, Chongzhuo, et al.
Publicado: (2024)
Obfuscation as an Effective Signal for Prioritizing Cross-Chain Smart Contract Audits: Large-Scale Measurement and Risk Profiling
por: Zhao, Yao, et al.
Publicado: (2026)
por: Zhao, Yao, et al.
Publicado: (2026)
Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered Memory
por: Ren, Jie, et al.
Publicado: (2025)
por: Ren, Jie, et al.
Publicado: (2025)
Systematic Performance Evaluation Framework for LEO Mega-Constellation Satellite Networks
por: Wang, Yu, et al.
Publicado: (2024)
por: Wang, Yu, et al.
Publicado: (2024)
A Microbenchmark Framework for Performance Evaluation of OpenMP Target Offloading
por: Atif, Mohammad, et al.
Publicado: (2025)
por: Atif, Mohammad, et al.
Publicado: (2025)
Performance Optimization of 3D Stencil Computation on ARM Scalable Vector Extension
por: Chen, Hongguang
Publicado: (2025)
por: Chen, Hongguang
Publicado: (2025)
DSO: A GPU Energy Efficiency Optimizer by Fusing Dynamic and Static Information
por: Wang, Qiang, et al.
Publicado: (2024)
por: Wang, Qiang, et al.
Publicado: (2024)
Toward Smart Scheduling in Tapis
por: Stubbs, Joe, et al.
Publicado: (2024)
por: Stubbs, Joe, et al.
Publicado: (2024)
Effects of the Auto-Correlation of Delays on the Age of Information: A Gaussian Process Framework
por: Inoie, Atsushi, et al.
Publicado: (2025)
por: Inoie, Atsushi, et al.
Publicado: (2025)
Ejemplares similares
-
Co-Design of 2D Heterojunctions for Data Filtering in Tracking Systems
por: Oli, Tupendra, et al.
Publicado: (2024) -
Integrating ytopt and libEnsemble to Autotune OpenMC
por: Wu, Xingfu, et al.
Publicado: (2024) -
Generalizing Scaling Laws for Dense and Sparse Large Language Models
por: Hossain, Md Arafat, et al.
Publicado: (2025) -
HPC Application Parameter Autotuning on Edge Devices: A Bandit Learning Approach
por: Hossain, Abrar, et al.
Publicado: (2025) -
Machine Learning-driven Autotuning of Graphics Processing Unit Accelerated Computational Fluid Dynamics for Enhanced Performance
por: Xue, Weicheng, et al.
Publicado: (2023)