Hardware-Assisted Virtualization of Neural Processing Units for Cloud Platforms
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xue, Yuqi, Liu, Yiqi, Nai, Lifeng, Huang, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs
von: Cai, Tianhao, et al.
Veröffentlicht: (2025)
von: Cai, Tianhao, et al.
Veröffentlicht: (2025)
Hardware Memory Management for Future Mobile Hybrid Memory Systems
von: Wen, Fei, et al.
Veröffentlicht: (2020)
von: Wen, Fei, et al.
Veröffentlicht: (2020)
Arcus: SLO Management for Accelerators in the Cloud with Traffic Shaping
von: Zhao, Jiechen, et al.
Veröffentlicht: (2024)
von: Zhao, Jiechen, et al.
Veröffentlicht: (2024)
Virtuoso: Enabling Fast and Accurate Virtual Memory Research via an Imitation-based Operating System Simulation Methodology
von: Kanellopoulos, Konstantinos, et al.
Veröffentlicht: (2024)
von: Kanellopoulos, Konstantinos, et al.
Veröffentlicht: (2024)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
von: Huang, Wei, et al.
Veröffentlicht: (2023)
von: Huang, Wei, et al.
Veröffentlicht: (2023)
Attention, Distillation, and Tabularization: Towards Practical Neural Network-Based Prefetching
von: Zhang, Pengmiao, et al.
Veröffentlicht: (2023)
von: Zhang, Pengmiao, et al.
Veröffentlicht: (2023)
CounterPoint: Using Hardware Event Counters to Refute and Refine Microarchitectural Assumptions (Extended Version)
von: Lindsay, Nick, et al.
Veröffentlicht: (2026)
von: Lindsay, Nick, et al.
Veröffentlicht: (2026)
GraNNite: Enabling High-Performance Execution of Graph Neural Networks on Resource-Constrained Neural Processing Units
von: Das, Arghadip, et al.
Veröffentlicht: (2025)
von: Das, Arghadip, et al.
Veröffentlicht: (2025)
Ariadne: A Hotness-Aware and Size-Adaptive Compressed Swap Technique for Fast Application Relaunch and Reduced CPU Usage on Mobile Devices
von: Liang, Yu, et al.
Veröffentlicht: (2025)
von: Liang, Yu, et al.
Veröffentlicht: (2025)
ConZone+: Practical Zoned Flash Storage Emulation for Consumer Devices
von: Yu, Dingcui, et al.
Veröffentlicht: (2025)
von: Yu, Dingcui, et al.
Veröffentlicht: (2025)
Brain-inspired AI for Edge Intelligence: a systematic review
von: Cheng, Yingchao, et al.
Veröffentlicht: (2026)
von: Cheng, Yingchao, et al.
Veröffentlicht: (2026)
Extending Straight-Through Estimation for Robust Neural Networks on Analog CIM Hardware
von: Feng, Yuannuo, et al.
Veröffentlicht: (2025)
von: Feng, Yuannuo, et al.
Veröffentlicht: (2025)
ReGate: Enabling Power Gating in Neural Processing Units
von: Xue, Yuqi, et al.
Veröffentlicht: (2025)
von: Xue, Yuqi, et al.
Veröffentlicht: (2025)
Dynamic Voltage and Frequency Scaling for Intermittent Computing
von: Maioli, Andrea, et al.
Veröffentlicht: (2024)
von: Maioli, Andrea, et al.
Veröffentlicht: (2024)
Rethinking Programmed I/O for Fast Devices, Cheap Cores, and Coherent Interconnects
von: Ruzhanskaia, Anastasiia, et al.
Veröffentlicht: (2024)
von: Ruzhanskaia, Anastasiia, et al.
Veröffentlicht: (2024)
eBPF-mm: Userspace-guided memory management in Linux with eBPF
von: Mores, Konstantinos, et al.
Veröffentlicht: (2024)
von: Mores, Konstantinos, et al.
Veröffentlicht: (2024)
Where Linux Breaks Under Radiation: A Cross-Architecture Kernel-Level Characterization of Proton-Induced Failures in COTS SoCs
von: Memon, Saad, et al.
Veröffentlicht: (2025)
von: Memon, Saad, et al.
Veröffentlicht: (2025)
Preemption-Enhanced Benchmark Suite for FPGAs
von: Malik, Arsalan Ali, et al.
Veröffentlicht: (2025)
von: Malik, Arsalan Ali, et al.
Veröffentlicht: (2025)
Enabling Syscall Intercept for RISC-V
von: Andrić, Petar, et al.
Veröffentlicht: (2025)
von: Andrić, Petar, et al.
Veröffentlicht: (2025)
WebAssembly on Resource-Constrained IoT Devices: Performance, Efficiency, and Portability
von: Has, Mislav, et al.
Veröffentlicht: (2025)
von: Has, Mislav, et al.
Veröffentlicht: (2025)
AERO: Adaptive and Efficient Runtime-Aware OTA Updates for Energy-Harvesting IoT
von: Wei, Wei, et al.
Veröffentlicht: (2026)
von: Wei, Wei, et al.
Veröffentlicht: (2026)
ASIC-based Compression Accelerators for Storage Systems: Design, Placement, and Profiling Insights
von: Lu, Tao, et al.
Veröffentlicht: (2025)
von: Lu, Tao, et al.
Veröffentlicht: (2025)
LearnedFTL: A Learning-Based Page-Level FTL for Reducing Double Reads in Flash-Based SSDs
von: Wang, Shengzhe, et al.
Veröffentlicht: (2023)
von: Wang, Shengzhe, et al.
Veröffentlicht: (2023)
An Introduction to the Compute Express Link (CXL) Interconnect
von: Sharma, Debendra Das, et al.
Veröffentlicht: (2023)
von: Sharma, Debendra Das, et al.
Veröffentlicht: (2023)
Neural Network Quantization for Microcontrollers: A Comprehensive Survey of Methods, Platforms, and Applications
von: Abushahla, Hamza A., et al.
Veröffentlicht: (2025)
von: Abushahla, Hamza A., et al.
Veröffentlicht: (2025)
ConSmax: Hardware-Friendly Alternative Softmax with Learnable Parameters
von: Liu, Shiwei, et al.
Veröffentlicht: (2024)
von: Liu, Shiwei, et al.
Veröffentlicht: (2024)
HPD: Hybrid Projection Decomposition for Robust State Space Models on Analog CIM Hardware
von: Feng, Yuannuo, et al.
Veröffentlicht: (2025)
von: Feng, Yuannuo, et al.
Veröffentlicht: (2025)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
Architect in the Loop Agentic Hardware Design and Verification
von: Mohammed, Mubarek
Veröffentlicht: (2025)
von: Mohammed, Mubarek
Veröffentlicht: (2025)
ML For Hardware Design Interpretability: Challenges and Opportunities
von: Baartmans, Raymond, et al.
Veröffentlicht: (2025)
von: Baartmans, Raymond, et al.
Veröffentlicht: (2025)
GRAU: Generic Reconfigurable Activation Unit Design for Neural Network Hardware Accelerators
von: Liu, Yuhao, et al.
Veröffentlicht: (2026)
von: Liu, Yuhao, et al.
Veröffentlicht: (2026)
Learning to Compare Hardware Designs for High-Level Synthesis
von: Bai, Yunsheng, et al.
Veröffentlicht: (2024)
von: Bai, Yunsheng, et al.
Veröffentlicht: (2024)
CXLMemSim: A pure software simulated CXL.mem for performance characterization
von: Yang, Yiwei, et al.
Veröffentlicht: (2023)
von: Yang, Yiwei, et al.
Veröffentlicht: (2023)
Putting the Context back into Memory
von: Roberts, David A.
Veröffentlicht: (2025)
von: Roberts, David A.
Veröffentlicht: (2025)
A Limits Study of Memory-side Tiering Telemetry
von: Petrucci, Vinicius, et al.
Veröffentlicht: (2025)
von: Petrucci, Vinicius, et al.
Veröffentlicht: (2025)
HDReason: Algorithm-Hardware Codesign for Hyperdimensional Knowledge Graph Reasoning
von: Chen, Hanning, et al.
Veröffentlicht: (2024)
von: Chen, Hanning, et al.
Veröffentlicht: (2024)
Challenges and Research Directions for Large Language Model Inference Hardware
von: Ma, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Ma, Xiaoyu, et al.
Veröffentlicht: (2026)
Hardware Efficient Approximate Convolution with Tunable Error Tolerance for CNNs
von: Shashidhar, Vishal, et al.
Veröffentlicht: (2026)
von: Shashidhar, Vishal, et al.
Veröffentlicht: (2026)
Enabling Physical AI at the Edge: Hardware-Accelerated Recovery of System Dynamics
von: Xu, Bin, et al.
Veröffentlicht: (2025)
von: Xu, Bin, et al.
Veröffentlicht: (2025)
The NIC should be part of the OS
von: Xu, Pengcheng, et al.
Veröffentlicht: (2025)
von: Xu, Pengcheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs
von: Cai, Tianhao, et al.
Veröffentlicht: (2025) -
Hardware Memory Management for Future Mobile Hybrid Memory Systems
von: Wen, Fei, et al.
Veröffentlicht: (2020) -
Arcus: SLO Management for Accelerators in the Cloud with Traffic Shaping
von: Zhao, Jiechen, et al.
Veröffentlicht: (2024) -
Virtuoso: Enabling Fast and Accurate Virtual Memory Research via an Imitation-based Operating System Simulation Methodology
von: Kanellopoulos, Konstantinos, et al.
Veröffentlicht: (2024) -
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
von: Huang, Wei, et al.
Veröffentlicht: (2023)