Good things come in small packages: Should we build AI clusters with Lite-GPUs?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Canakci, Burcu, Liu, Junyi, Wu, Xingbo, Cheriere, Nathanaël, Costa, Paolo, Legtchenko, Sergey, Narayanan, Dushyanth, Rowstron, Ant |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Managed-Retention Memory: A New Class of Memory for the AI Era
von: Legtchenko, Sergey, et al.
Veröffentlicht: (2025)
von: Legtchenko, Sergey, et al.
Veröffentlicht: (2025)
Towards Memory Specialization: A Case for Long-Term and Short-Term RAM
von: Li, Peijing, et al.
Veröffentlicht: (2025)
von: Li, Peijing, et al.
Veröffentlicht: (2025)
Control Flow Management in Modern GPUs
von: Shoushtary, Mojtaba Abaie, et al.
Veröffentlicht: (2024)
von: Shoushtary, Mojtaba Abaie, et al.
Veröffentlicht: (2024)
Privacy-Preserving Performance Profiling of In-The-Wild GPUs
von: McDougall, Ian, et al.
Veröffentlicht: (2025)
von: McDougall, Ian, et al.
Veröffentlicht: (2025)
A Systematic Characterization of LLM Inference on GPUs
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
Fault Injection in On-Chip Interconnects: A Comparative Study of Wishbone, AXI-Lite, and AXI
von: Zhao, Hongwei, et al.
Veröffentlicht: (2025)
von: Zhao, Hongwei, et al.
Veröffentlicht: (2025)
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
von: Hong, Jeongmin, et al.
Veröffentlicht: (2024)
von: Hong, Jeongmin, et al.
Veröffentlicht: (2024)
Fleet: Hierarchical Task-based Abstraction for Megakernels on Multi-Die GPUs
von: Chowdhary, Sangeeta, et al.
Veröffentlicht: (2026)
von: Chowdhary, Sangeeta, et al.
Veröffentlicht: (2026)
Virgo: Cluster-level Matrix Unit Integration in GPUs for Scalability and Energy Efficiency
von: Kim, Hansung, et al.
Veröffentlicht: (2024)
von: Kim, Hansung, et al.
Veröffentlicht: (2024)
CarbonSet: A Dataset to Analyze Trends and Benchmark the Sustainability of CPUs and GPUs
von: Hu, Jiajun, et al.
Veröffentlicht: (2025)
von: Hu, Jiajun, et al.
Veröffentlicht: (2025)
RulePlanner: All-in-One Reinforcement Learner for Unifying Design Rules in 3D Floorplanning
von: Zhong, Ruizhe, et al.
Veröffentlicht: (2026)
von: Zhong, Ruizhe, et al.
Veröffentlicht: (2026)
Study on the Particle Sorting Performance for Reactor Monte Carlo Neutron Transport on Apple Unified Memory GPUs
von: Liu, Changyuan
Veröffentlicht: (2024)
von: Liu, Changyuan
Veröffentlicht: (2024)
GPIR: Enabling Practical Private Information Retrieval with GPUs
von: Ji, Hyesung, et al.
Veröffentlicht: (2026)
von: Ji, Hyesung, et al.
Veröffentlicht: (2026)
Hidden Risks of Unmonitored GPUs in Intelligent Transportation Systems
von: Puspa, Sefatun-Noor, et al.
Veröffentlicht: (2026)
von: Puspa, Sefatun-Noor, et al.
Veröffentlicht: (2026)
Hardware and software build flow with SoCMake
von: Pejašinović, Risto, et al.
Veröffentlicht: (2025)
von: Pejašinović, Risto, et al.
Veröffentlicht: (2025)
How Much Progress Has There Been in NVIDIA Datacenter GPUs?
von: Del Sozzo, Emanuele, et al.
Veröffentlicht: (2026)
von: Del Sozzo, Emanuele, et al.
Veröffentlicht: (2026)
SnipSnap: A Joint Compression Format and Dataflow Co-Optimization Framework for Efficient Sparse LLM Accelerator Design
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
Kitsune: Enabling Dataflow Execution on GPUs
von: Davies, Michael, et al.
Veröffentlicht: (2025)
von: Davies, Michael, et al.
Veröffentlicht: (2025)
Characterizing and Understanding HGNN Training on GPUs
von: Han, Dengke, et al.
Veröffentlicht: (2024)
von: Han, Dengke, et al.
Veröffentlicht: (2024)
Evaluating Computing Platforms for Sustainability: A Comparative Analysis of FPGAs against ASICs, GPUs, and CPUs
von: Sudarshan, Chetan Choppali, et al.
Veröffentlicht: (2026)
von: Sudarshan, Chetan Choppali, et al.
Veröffentlicht: (2026)
An integrated design of energy and indoor environmental quality monitoring system for effective building performance management
von: Zakka, Vincent Gbouna, et al.
Veröffentlicht: (2025)
von: Zakka, Vincent Gbouna, et al.
Veröffentlicht: (2025)
A Reconfigurable Computing In-Memory Macro with Charge-sharing-based Weighted Accumulator
von: Yang, Junyi, et al.
Veröffentlicht: (2026)
von: Yang, Junyi, et al.
Veröffentlicht: (2026)
CADC: Crossbar-Aware Dendritic Convolution for Efficient In-memory Computing
von: Dong, Shuai, et al.
Veröffentlicht: (2025)
von: Dong, Shuai, et al.
Veröffentlicht: (2025)
ControlPULPlet: A Flexible Real-time Multi-core RISC-V Controller for 2.5D Systems-in-package
von: Ottaviano, Alessandro, et al.
Veröffentlicht: (2024)
von: Ottaviano, Alessandro, et al.
Veröffentlicht: (2024)
Acore-CIM: build accurate and reliable mixed-signal CIM cores with RISC-V controlled self-calibration
von: Numan, Omar, et al.
Veröffentlicht: (2025)
von: Numan, Omar, et al.
Veröffentlicht: (2025)
HetGPU: The pursuit of making binary compatibility towards GPUs
von: Yang, Yiwei, et al.
Veröffentlicht: (2025)
von: Yang, Yiwei, et al.
Veröffentlicht: (2025)
Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs
von: Soi, Rupanshu, et al.
Veröffentlicht: (2025)
von: Soi, Rupanshu, et al.
Veröffentlicht: (2025)
Tawa: Automatic Warp Specialization for Modern GPUs with Asynchronous References
von: Chen, Hongzheng, et al.
Veröffentlicht: (2025)
von: Chen, Hongzheng, et al.
Veröffentlicht: (2025)
Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs
von: Yadav, Divakar Kumar, et al.
Veröffentlicht: (2026)
von: Yadav, Divakar Kumar, et al.
Veröffentlicht: (2026)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
Near-Memory Architecture for Threshold-Ordinal Surface-Based Corner Detection of Event Cameras
von: Shang, Hongyang, et al.
Veröffentlicht: (2025)
von: Shang, Hongyang, et al.
Veröffentlicht: (2025)
smallNet: Implementation of a convolutional layer in tiny FPGAs
von: Bascuñán, Fernanda Zapata, et al.
Veröffentlicht: (2025)
von: Bascuñán, Fernanda Zapata, et al.
Veröffentlicht: (2025)
In-Memory ADC-Based Nonlinear Activation Quantization for Efficient In-Memory Computing
von: Dong, Shuai, et al.
Veröffentlicht: (2026)
von: Dong, Shuai, et al.
Veröffentlicht: (2026)
Efficient stereo matching on embedded GPUs with zero-means cross correlation
von: Chang, Qiong, et al.
Veröffentlicht: (2022)
von: Chang, Qiong, et al.
Veröffentlicht: (2022)
Characterizing the Behavior of Training Mamba-based State Space Models on GPUs
von: Baruah, Trinayan, et al.
Veröffentlicht: (2025)
von: Baruah, Trinayan, et al.
Veröffentlicht: (2025)
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
von: Qararyah, Fareed, et al.
Veröffentlicht: (2024)
von: Qararyah, Fareed, et al.
Veröffentlicht: (2024)
Are LLMs Any Good for High-Level Synthesis?
von: Liao, Yuchao, et al.
Veröffentlicht: (2024)
von: Liao, Yuchao, et al.
Veröffentlicht: (2024)
Evaluation of Hardware-based Video Encoders on Modern GPUs for UHD Live-Streaming
von: Arunruangsirilert, Kasidis, et al.
Veröffentlicht: (2025)
von: Arunruangsirilert, Kasidis, et al.
Veröffentlicht: (2025)
Topkima-Former: Low-energy, Low-Latency Inference for Transformers using top-k In-memory ADC
von: Dong, Shuai, et al.
Veröffentlicht: (2024)
von: Dong, Shuai, et al.
Veröffentlicht: (2024)
Bombyx: OpenCilk Compilation for FPGA Hardware Acceleration
von: Shahawy, Mohamed, et al.
Veröffentlicht: (2025)
von: Shahawy, Mohamed, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Managed-Retention Memory: A New Class of Memory for the AI Era
von: Legtchenko, Sergey, et al.
Veröffentlicht: (2025) -
Towards Memory Specialization: A Case for Long-Term and Short-Term RAM
von: Li, Peijing, et al.
Veröffentlicht: (2025) -
Control Flow Management in Modern GPUs
von: Shoushtary, Mojtaba Abaie, et al.
Veröffentlicht: (2024) -
Privacy-Preserving Performance Profiling of In-The-Wild GPUs
von: McDougall, Ian, et al.
Veröffentlicht: (2025) -
A Systematic Characterization of LLM Inference on GPUs
von: Wang, Haonan, et al.
Veröffentlicht: (2025)