Privacy-Preserving Performance Profiling of In-The-Wild GPUs
Fuente:
arXiv
Saved in:
| Main Authors: | McDougall, Ian, Davies, Michael, Chatterjee, Rahul, Jha, Somesh, Sankaralingam, Karthikeyan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pedagogically Motivated and Composable Open-Source RISC-V Processors for Computer Science Education
by: McDougall, Ian, et al.
Published: (2025)
by: McDougall, Ian, et al.
Published: (2025)
Beyond Static Policies: Exploring Dynamic Policy Selection for Single-Thread Performance Optimization
by: Zhang, Yanxin, et al.
Published: (2026)
by: Zhang, Yanxin, et al.
Published: (2026)
IPU: Flexible Hardware Introspection Units
by: McDougall, Ian, et al.
Published: (2023)
by: McDougall, Ian, et al.
Published: (2023)
LIMINAL: Exploring The Frontiers of LLM Decode Performance
by: Davies, Michael, et al.
Published: (2025)
by: Davies, Michael, et al.
Published: (2025)
Kitsune: Enabling Dataflow Execution on GPUs
by: Davies, Michael, et al.
Published: (2025)
by: Davies, Michael, et al.
Published: (2025)
SAHM: State-Aware Heterogeneous Multicore for Single-Thread Performance
by: Wadle, Shayne, et al.
Published: (2025)
by: Wadle, Shayne, et al.
Published: (2025)
NeuroScalar: A Deep Learning Framework for Fast, Accurate, and In-the-Wild Cycle-Level Performance Prediction
by: Wadle, Shayne, et al.
Published: (2025)
by: Wadle, Shayne, et al.
Published: (2025)
Computer Architecture's AlphaZero Moment: Automated Discovery in an Encircled World
by: Sankaralingam, Karthikeyan
Published: (2026)
by: Sankaralingam, Karthikeyan
Published: (2026)
The Impact Market to Save Conference Peer Review: Decoupling Dissemination and Credentialing
by: Sankaralingam, Karthikeyan
Published: (2025)
by: Sankaralingam, Karthikeyan
Published: (2025)
Control Flow Management in Modern GPUs
by: Shoushtary, Mojtaba Abaie, et al.
Published: (2024)
by: Shoushtary, Mojtaba Abaie, et al.
Published: (2024)
Study on the Particle Sorting Performance for Reactor Monte Carlo Neutron Transport on Apple Unified Memory GPUs
by: Liu, Changyuan
Published: (2024)
by: Liu, Changyuan
Published: (2024)
A Systematic Characterization of LLM Inference on GPUs
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
Ironman: Accelerating Oblivious Transfer Extension for Privacy-Preserving AI with Near-Memory Processing
by: Lin, Chenqi, et al.
Published: (2025)
by: Lin, Chenqi, et al.
Published: (2025)
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
by: Hong, Jeongmin, et al.
Published: (2024)
by: Hong, Jeongmin, et al.
Published: (2024)
Fleet: Hierarchical Task-based Abstraction for Megakernels on Multi-Die GPUs
by: Chowdhary, Sangeeta, et al.
Published: (2026)
by: Chowdhary, Sangeeta, et al.
Published: (2026)
CarbonSet: A Dataset to Analyze Trends and Benchmark the Sustainability of CPUs and GPUs
by: Hu, Jiajun, et al.
Published: (2025)
by: Hu, Jiajun, et al.
Published: (2025)
Virgo: Cluster-level Matrix Unit Integration in GPUs for Scalability and Energy Efficiency
by: Kim, Hansung, et al.
Published: (2024)
by: Kim, Hansung, et al.
Published: (2024)
RealProbe: An Automated and Lightweight Performance Profiler for In-FPGA Execution of High-Level Synthesis Designs
by: Kim, Jiho, et al.
Published: (2025)
by: Kim, Jiho, et al.
Published: (2025)
GPIR: Enabling Practical Private Information Retrieval with GPUs
by: Ji, Hyesung, et al.
Published: (2026)
by: Ji, Hyesung, et al.
Published: (2026)
Hidden Risks of Unmonitored GPUs in Intelligent Transportation Systems
by: Puspa, Sefatun-Noor, et al.
Published: (2026)
by: Puspa, Sefatun-Noor, et al.
Published: (2026)
Profile-Guided Temporal Prefetching
by: Li, Mengming, et al.
Published: (2025)
by: Li, Mengming, et al.
Published: (2025)
How Much Progress Has There Been in NVIDIA Datacenter GPUs?
by: Del Sozzo, Emanuele, et al.
Published: (2026)
by: Del Sozzo, Emanuele, et al.
Published: (2026)
Constable: Improving Performance and Power Efficiency by Safely Eliminating Load Instruction Execution
by: Bera, Rahul, et al.
Published: (2024)
by: Bera, Rahul, et al.
Published: (2024)
Sectored DRAM: A Practical Energy-Efficient and High-Performance Fine-Grained DRAM Architecture
by: Olgun, Ataberk, et al.
Published: (2022)
by: Olgun, Ataberk, et al.
Published: (2022)
Best Practices for Large Load Interconnections: A North American Perspective on Data Centers
by: Zahedi, Rafi, et al.
Published: (2026)
by: Zahedi, Rafi, et al.
Published: (2026)
APINT: A Full-Stack Framework for Acceleration of Privacy-Preserving Inference of Transformers based on Garbled Circuits
by: Cho, Hyunjun, et al.
Published: (2025)
by: Cho, Hyunjun, et al.
Published: (2025)
Lightweight Congruence Profiling for Early Design Exploration of Heterogeneous FPGAs
by: Boston, Allen, et al.
Published: (2025)
by: Boston, Allen, et al.
Published: (2025)
A Mess of Memory System Benchmarking, Simulation and Application Profiling
by: Esmaili-Dokht, Pouya, et al.
Published: (2024)
by: Esmaili-Dokht, Pouya, et al.
Published: (2024)
Architectural Design and Performance Analysis of FPGA based AI Accelerators: A Comprehensive Review
by: Chatterjee, Soumita, et al.
Published: (2026)
by: Chatterjee, Soumita, et al.
Published: (2026)
SPRING: Systematic Profiling of Randomly Interconnected Neural Networks Generated by HLS
by: Shi, Rui, et al.
Published: (2025)
by: Shi, Rui, et al.
Published: (2025)
Understanding Simulated Architecture via gem5 Call-Stack Profiling
by: Söderström, Johan, et al.
Published: (2026)
by: Söderström, Johan, et al.
Published: (2026)
Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs
by: Soi, Rupanshu, et al.
Published: (2025)
by: Soi, Rupanshu, et al.
Published: (2025)
NeCTAr: A Heterogeneous RISC-V SoC for Language Model Inference in Intel 16
by: Schmulbach, Viansa, et al.
Published: (2025)
by: Schmulbach, Viansa, et al.
Published: (2025)
Evaluating Computing Platforms for Sustainability: A Comparative Analysis of FPGAs against ASICs, GPUs, and CPUs
by: Sudarshan, Chetan Choppali, et al.
Published: (2026)
by: Sudarshan, Chetan Choppali, et al.
Published: (2026)
Spec2Assertion: Automatic Pre-RTL Assertion Generation using Large Language Models with Progressive Regularization
by: Wu, Fenghua, et al.
Published: (2025)
by: Wu, Fenghua, et al.
Published: (2025)
In-Storage Domain-Specific Acceleration for Serverless Computing
by: Mahapatra, Rohan, et al.
Published: (2023)
by: Mahapatra, Rohan, et al.
Published: (2023)
Performance and Energy Benefits of MRDIMMs
by: Díaz, Pau, et al.
Published: (2026)
by: Díaz, Pau, et al.
Published: (2026)
Look-Up Table based Neural Network Hardware
by: Sen, Ovishake, et al.
Published: (2024)
by: Sen, Ovishake, et al.
Published: (2024)
Latch Based Design for Fast Voltage Droop Response
by: Srinivas, Shreyas, et al.
Published: (2025)
by: Srinivas, Shreyas, et al.
Published: (2025)
Demystifying FPGA Hard NoC Performance
by: Liu, Sihao, et al.
Published: (2025)
by: Liu, Sihao, et al.
Published: (2025)
Similar Items
-
Pedagogically Motivated and Composable Open-Source RISC-V Processors for Computer Science Education
by: McDougall, Ian, et al.
Published: (2025) -
Beyond Static Policies: Exploring Dynamic Policy Selection for Single-Thread Performance Optimization
by: Zhang, Yanxin, et al.
Published: (2026) -
IPU: Flexible Hardware Introspection Units
by: McDougall, Ian, et al.
Published: (2023) -
LIMINAL: Exploring The Frontiers of LLM Decode Performance
by: Davies, Michael, et al.
Published: (2025) -
Kitsune: Enabling Dataflow Execution on GPUs
by: Davies, Michael, et al.
Published: (2025)