How Much Progress Has There Been in NVIDIA Datacenter GPUs?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Del Sozzo, Emanuele, Fleming, Martin, Flamm, Kenneth, Thompson, Neil |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Power Stabilization for AI Training Datacenters
von: Choukse, Esha, et al.
Veröffentlicht: (2025)
von: Choukse, Esha, et al.
Veröffentlicht: (2025)
Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs
von: Yadav, Divakar Kumar, et al.
Veröffentlicht: (2026)
von: Yadav, Divakar Kumar, et al.
Veröffentlicht: (2026)
Characterizing and Understanding HGNN Training on GPUs
von: Han, Dengke, et al.
Veröffentlicht: (2024)
von: Han, Dengke, et al.
Veröffentlicht: (2024)
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
LLM4EDA: Emerging Progress in Large Language Models for Electronic Design Automation
von: Zhong, Ruizhe, et al.
Veröffentlicht: (2023)
von: Zhong, Ruizhe, et al.
Veröffentlicht: (2023)
RuC: HDL-Agnostic Rule Completion Benchmark Generation
von: Domingo, Arnau Ayguadé, et al.
Veröffentlicht: (2026)
von: Domingo, Arnau Ayguadé, et al.
Veröffentlicht: (2026)
Analyzing Modern NVIDIA GPU cores
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
IMAGINE: An 8-to-1b 22nm FD-SOI Compute-In-Memory CNN Accelerator With an End-to-End Analog Charge-Based 0.15-8POPS/W Macro Featuring Distribution-Aware Data Reshaping
von: Kneip, Adrian, et al.
Veröffentlicht: (2024)
von: Kneip, Adrian, et al.
Veröffentlicht: (2024)
HiAER-Spike Software-Hardware Reconfigurable Platform for Event-Driven Neuromorphic Computing at Scale
von: Frank, Gwenevere, et al.
Veröffentlicht: (2026)
von: Frank, Gwenevere, et al.
Veröffentlicht: (2026)
EdgeReasoning: Characterizing Reasoning LLM Deployment on Edge GPUs
von: Kubwimana, Benjamin, et al.
Veröffentlicht: (2025)
von: Kubwimana, Benjamin, et al.
Veröffentlicht: (2025)
DDC: A Vision for a Disaggregated Datacenter
von: Ewais, Mohammad, et al.
Veröffentlicht: (2024)
von: Ewais, Mohammad, et al.
Veröffentlicht: (2024)
Microbenchmarking NVIDIA's Blackwell Architecture: An in-depth Architectural Analysis
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2025)
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2025)
Thermal Analysis for NVIDIA GTX480 Fermi GPU Architecture
von: Nagendra, Savinay
Veröffentlicht: (2024)
von: Nagendra, Savinay
Veröffentlicht: (2024)
Good things come in small packages: Should we build AI clusters with Lite-GPUs?
von: Canakci, Burcu, et al.
Veröffentlicht: (2025)
von: Canakci, Burcu, et al.
Veröffentlicht: (2025)
Design Conductor: An agent autonomously builds a 1.5 GHz Linux-capable RISC-V CPU
von: The Verkor Team, et al.
Veröffentlicht: (2026)
von: The Verkor Team, et al.
Veröffentlicht: (2026)
Diagnosing FP4 inference: a layer-wise and block-wise sensitivity analysis of NVFP4 and MXFP4
von: Cim, Musa, et al.
Veröffentlicht: (2026)
von: Cim, Musa, et al.
Veröffentlicht: (2026)
POET: Power-Oriented Evolutionary Tuning for LLM-Based RTL PPA Optimization
von: Ping, Heng, et al.
Veröffentlicht: (2026)
von: Ping, Heng, et al.
Veröffentlicht: (2026)
Enhancing LUT-based Deep Neural Networks Inference through Architecture and Connectivity Optimization
von: Lou, Binglei, et al.
Veröffentlicht: (2026)
von: Lou, Binglei, et al.
Veröffentlicht: (2026)
HYPERHEURIST: A Simulated Annealing-Based Control Framework for LLM-Driven Code Generation in Optimized Hardware Design
von: Ahir, Shiva, et al.
Veröffentlicht: (2026)
von: Ahir, Shiva, et al.
Veröffentlicht: (2026)
Configuration Over Selection: Hyperparameter Sensitivity Exceeds Model Differences in Open-Source LLMs for RTL Generation
von: Shao, Minghao, et al.
Veröffentlicht: (2026)
von: Shao, Minghao, et al.
Veröffentlicht: (2026)
AMSnet-q: Unsupervised Circuit Identification and Performance Labeling for AMS Circuits
von: Zhang, Ze, et al.
Veröffentlicht: (2026)
von: Zhang, Ze, et al.
Veröffentlicht: (2026)
Resource Utilization of Differentiable Logic Gate Networks Deployed on FPGAs
von: Wormald, Stephen, et al.
Veröffentlicht: (2026)
von: Wormald, Stephen, et al.
Veröffentlicht: (2026)
EdgeCIM: A Hardware-Software Co-Design for CIM-Based Acceleration of Small Language Models
von: Bazzi, Jinane, et al.
Veröffentlicht: (2026)
von: Bazzi, Jinane, et al.
Veröffentlicht: (2026)
Dr. RTL: Autonomous Agentic RTL Optimization through Tool-Grounded Self-Improvement
von: Fang, Wenji, et al.
Veröffentlicht: (2026)
von: Fang, Wenji, et al.
Veröffentlicht: (2026)
Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding
von: Fan, Wang, et al.
Veröffentlicht: (2026)
von: Fan, Wang, et al.
Veröffentlicht: (2026)
AHASD: Asynchronous Heterogeneous Architecture for LLM Adaptive Drafting Speculative Decoding on Mobile Devices
von: Zirui, Ma, et al.
Veröffentlicht: (2026)
von: Zirui, Ma, et al.
Veröffentlicht: (2026)
RAGnaroX: A Secure, Local-Hosted ChatOps Assistant Using Small Language Models
von: Dornauer, Benedikt, et al.
Veröffentlicht: (2026)
von: Dornauer, Benedikt, et al.
Veröffentlicht: (2026)
A3D: Agentic AI flow for autonomous Accelerator Design
von: Nallathambi, Abinand, et al.
Veröffentlicht: (2026)
von: Nallathambi, Abinand, et al.
Veröffentlicht: (2026)
Building Reliable Arithmetic Multipliers Under NBTI Aging and Process Variations
von: Heidary, Masoud, et al.
Veröffentlicht: (2026)
von: Heidary, Masoud, et al.
Veröffentlicht: (2026)
AceleradorSNN: A Neuromorphic Cognitive System Integrating Spiking Neural Networks and DynamicImage Signal Processing on FPGA
von: Gutierrez, Daniel, et al.
Veröffentlicht: (2026)
von: Gutierrez, Daniel, et al.
Veröffentlicht: (2026)
RPU -- A Reasoning Processing Unit
von: Adiletta, Matthew, et al.
Veröffentlicht: (2026)
von: Adiletta, Matthew, et al.
Veröffentlicht: (2026)
SA-Kura: An Energy-Efficient Systolic Array Accelerator for Locally-Coupled Kuramoto Drift in Diffusion Sampling
von: Jin, Jeongmin, et al.
Veröffentlicht: (2026)
von: Jin, Jeongmin, et al.
Veröffentlicht: (2026)
MX-SAFE: Versatile Inference- and Training-Proof Microscaling Format with On-the-Fly Exponent and Mantissa Bit Allocation
von: Park, Dahoon, et al.
Veröffentlicht: (2026)
von: Park, Dahoon, et al.
Veröffentlicht: (2026)
RulePlanner: All-in-One Reinforcement Learner for Unifying Design Rules in 3D Floorplanning
von: Zhong, Ruizhe, et al.
Veröffentlicht: (2026)
von: Zhong, Ruizhe, et al.
Veröffentlicht: (2026)
AssertLLM2: A Comprehensive LLM Benchmark for Assertion Generation from Design Specifications
von: Wu, Yuchao, et al.
Veröffentlicht: (2026)
von: Wu, Yuchao, et al.
Veröffentlicht: (2026)
HAVEN: Hybrid Automated Verification ENgine for UVM Testbench Synthesis with LLMs
von: Meng, Chang-Chih, et al.
Veröffentlicht: (2026)
von: Meng, Chang-Chih, et al.
Veröffentlicht: (2026)
LP5X-PIM Sim: A High-Fidelity HW/SW Integrated Simulator for LPDDR5X-PIM
von: Cha, SangHoon, et al.
Veröffentlicht: (2026)
von: Cha, SangHoon, et al.
Veröffentlicht: (2026)
GRAU: Generic Reconfigurable Activation Unit Design for Neural Network Hardware Accelerators
von: Liu, Yuhao, et al.
Veröffentlicht: (2026)
von: Liu, Yuhao, et al.
Veröffentlicht: (2026)
Neuromorphic Computing for Low-Power Artificial Intelligence
von: Katti, Keshava, et al.
Veröffentlicht: (2026)
von: Katti, Keshava, et al.
Veröffentlicht: (2026)
Comparative Characterization of KV Cache Management Strategies for LLM Inference
von: Mamo, Oteo, et al.
Veröffentlicht: (2026)
von: Mamo, Oteo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Power Stabilization for AI Training Datacenters
von: Choukse, Esha, et al.
Veröffentlicht: (2025) -
Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs
von: Yadav, Divakar Kumar, et al.
Veröffentlicht: (2026) -
Characterizing and Understanding HGNN Training on GPUs
von: Han, Dengke, et al.
Veröffentlicht: (2024) -
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025) -
LLM4EDA: Emerging Progress in Large Language Models for Electronic Design Automation
von: Zhong, Ruizhe, et al.
Veröffentlicht: (2023)