GPUArmor: A Hardware-Software Co-design for Efficient and Scalable Memory Safety on GPUs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ziad, Mohamed Tarek Ibn, Damani, Sana, Stephenson, Mark, Keckler, Stephen W., Jaleel, Aamer |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Kitsune: Enabling Dataflow Execution on GPUs
von: Davies, Michael, et al.
Veröffentlicht: (2025)
von: Davies, Michael, et al.
Veröffentlicht: (2025)
Benchmarking Compound AI Applications for Hardware-Software Co-Design
von: Samuthrsindh, Paramuth, et al.
Veröffentlicht: (2026)
von: Samuthrsindh, Paramuth, et al.
Veröffentlicht: (2026)
xNVMe: Unleashing Storage Hardware-Software Co-design
von: Lund, Simon A. F., et al.
Veröffentlicht: (2024)
von: Lund, Simon A. F., et al.
Veröffentlicht: (2024)
Optimizing Hardware Resource Partitioning and Job Allocations on Modern GPUs under Power Caps
von: Arima, Eishi, et al.
Veröffentlicht: (2024)
von: Arima, Eishi, et al.
Veröffentlicht: (2024)
Improving Multi-Instance GPU Efficiency via Sub-Entry Sharing TLB Design
von: Li, Bingyao, et al.
Veröffentlicht: (2024)
von: Li, Bingyao, et al.
Veröffentlicht: (2024)
Utility-Driven Speculative Decoding for Mixture-of-Experts
von: Saxena, Anish, et al.
Veröffentlicht: (2025)
von: Saxena, Anish, et al.
Veröffentlicht: (2025)
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
von: Dwaraknath, Rajat Vadiraj, et al.
Veröffentlicht: (2026)
von: Dwaraknath, Rajat Vadiraj, et al.
Veröffentlicht: (2026)
Combining Performance and Productivity: Accelerating the Network Sensing Graph Challenge with GPUs and Commodity Data Science Software
von: Samsi, Siddharth, et al.
Veröffentlicht: (2025)
von: Samsi, Siddharth, et al.
Veröffentlicht: (2025)
Scalable Graph Indexing using GPUs for Approximate Nearest Neighbor Search
von: Li, Zhonggen, et al.
Veröffentlicht: (2025)
von: Li, Zhonggen, et al.
Veröffentlicht: (2025)
DiFache: Efficient and Scalable Caching on Disaggregated Memory using Decentralized Coherence
von: Zhang, Hanze, et al.
Veröffentlicht: (2025)
von: Zhang, Hanze, et al.
Veröffentlicht: (2025)
FREESH: Fair, Resource- and Energy-Efficient Scheduling for LLM Serving on Heterogeneous GPUs
von: He, Xuan, et al.
Veröffentlicht: (2025)
von: He, Xuan, et al.
Veröffentlicht: (2025)
TurboFFT: Co-Designed High-Performance and Fault-Tolerant Fast Fourier Transform on GPUs
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
ESG: Pipeline-Conscious Efficient Scheduling of DNN Workflows on Serverless Platforms with Shareable GPUs
von: Hui, Xinning, et al.
Veröffentlicht: (2024)
von: Hui, Xinning, et al.
Veröffentlicht: (2024)
ParamSpMM: Adaptive and Efficient Sparse Matrix-Matrix Multiplication on GPUs for GNNs
von: Zhang, Lixing, et al.
Veröffentlicht: (2026)
von: Zhang, Lixing, et al.
Veröffentlicht: (2026)
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
von: Patke, Archit, et al.
Veröffentlicht: (2025)
von: Patke, Archit, et al.
Veröffentlicht: (2025)
An Adaptive Distributed Stencil Abstraction for GPUs
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
Accelerating Maximal Biclique Enumeration on GPUs
von: Hsieh, Chou-Ying, et al.
Veröffentlicht: (2024)
von: Hsieh, Chou-Ying, et al.
Veröffentlicht: (2024)
Parallelizing Maximal Clique Enumeration on GPUs
von: Almasri, Mohammad, et al.
Veröffentlicht: (2022)
von: Almasri, Mohammad, et al.
Veröffentlicht: (2022)
Optimizing sDTW for AMD GPUs
von: Latta-Lin, Daniel, et al.
Veröffentlicht: (2024)
von: Latta-Lin, Daniel, et al.
Veröffentlicht: (2024)
TokenSim: Enabling Hardware and Software Exploration for Large Language Model Inference Systems
von: Wu, Feiyang, et al.
Veröffentlicht: (2025)
von: Wu, Feiyang, et al.
Veröffentlicht: (2025)
HMTRace: Hardware-Assisted Memory-Tagging based Dynamic Data Race Detection
von: Shastri, Jaidev, et al.
Veröffentlicht: (2024)
von: Shastri, Jaidev, et al.
Veröffentlicht: (2024)
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs
von: He, Guoliang, et al.
Veröffentlicht: (2025)
von: He, Guoliang, et al.
Veröffentlicht: (2025)
Thread and Data Mapping in Software Transactional Memory: An Overview
von: Pasqualin, Douglas Pereira, et al.
Veröffentlicht: (2022)
von: Pasqualin, Douglas Pereira, et al.
Veröffentlicht: (2022)
Serving Compound Inference Systems on Datacenter GPUs
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
von: Jangda, Abhinav, et al.
Veröffentlicht: (2024)
von: Jangda, Abhinav, et al.
Veröffentlicht: (2024)
Optimal Workload Placement on Multi-Instance GPUs
von: Turkkan, Bekir, et al.
Veröffentlicht: (2024)
von: Turkkan, Bekir, et al.
Veröffentlicht: (2024)
Towards Energy-Efficient Serverless Computing with Hardware Isolation
von: Carl, Natalie, et al.
Veröffentlicht: (2025)
von: Carl, Natalie, et al.
Veröffentlicht: (2025)
FLARE: A Dataflow-Aware and Scalable Hardware Architecture for Neural-Hybrid Scientific Lossy Compression
von: Jia, Wenqi, et al.
Veröffentlicht: (2025)
von: Jia, Wenqi, et al.
Veröffentlicht: (2025)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
von: Brock, Benjamin, et al.
Veröffentlicht: (2023)
von: Brock, Benjamin, et al.
Veröffentlicht: (2023)
Accurate Computation of the Logarithm of Modified Bessel Functions on GPUs
von: Plesner, Andreas, et al.
Veröffentlicht: (2024)
von: Plesner, Andreas, et al.
Veröffentlicht: (2024)
Self-Evolving Distributed Memory Architecture for Scalable AI Systems
von: Li, Zixuan, et al.
Veröffentlicht: (2026)
von: Li, Zixuan, et al.
Veröffentlicht: (2026)
FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion
von: Chang, Li-Wen, et al.
Veröffentlicht: (2024)
von: Chang, Li-Wen, et al.
Veröffentlicht: (2024)
Managing Multi Instance GPUs for High Throughput and Energy Savings
von: Saraha, Abhijeet, et al.
Veröffentlicht: (2025)
von: Saraha, Abhijeet, et al.
Veröffentlicht: (2025)
Analytical Performance Estimation during Code Generation on Modern GPUs
von: Ernst, Dominik, et al.
Veröffentlicht: (2022)
von: Ernst, Dominik, et al.
Veröffentlicht: (2022)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep Learning
von: An, Wei, et al.
Veröffentlicht: (2024)
von: An, Wei, et al.
Veröffentlicht: (2024)
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
von: Gao, Wei, et al.
Veröffentlicht: (2026)
von: Gao, Wei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Kitsune: Enabling Dataflow Execution on GPUs
von: Davies, Michael, et al.
Veröffentlicht: (2025) -
Benchmarking Compound AI Applications for Hardware-Software Co-Design
von: Samuthrsindh, Paramuth, et al.
Veröffentlicht: (2026) -
xNVMe: Unleashing Storage Hardware-Software Co-design
von: Lund, Simon A. F., et al.
Veröffentlicht: (2024) -
Optimizing Hardware Resource Partitioning and Job Allocations on Modern GPUs under Power Caps
von: Arima, Eishi, et al.
Veröffentlicht: (2024) -
Improving Multi-Instance GPU Efficiency via Sub-Entry Sharing TLB Design
von: Li, Bingyao, et al.
Veröffentlicht: (2024)