Lincoln AI Computing Survey (LAICS) and Trends
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Reuther, Albert, Michaleas, Peter, Jones, Michael, Gadepally, Vijay, Kepner, Jeremy |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Toward a Universal GPU Instruction Set Architecture: A Cross-Vendor Analysis of Hardware-Invariant Computational Primitives in Parallel Processors
von: Abraham, Ojima, et al.
Veröffentlicht: (2026)
von: Abraham, Ojima, et al.
Veröffentlicht: (2026)
Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System Modeling
von: Li, Yinrong, et al.
Veröffentlicht: (2026)
von: Li, Yinrong, et al.
Veröffentlicht: (2026)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
von: Jo, Myeong Jun
Veröffentlicht: (2026)
von: Jo, Myeong Jun
Veröffentlicht: (2026)
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
von: Pati, Suchita, et al.
Veröffentlicht: (2024)
von: Pati, Suchita, et al.
Veröffentlicht: (2024)
Wattchmen: Watching the Wattchers -- High Fidelity, Flexible GPU Energy Modeling
von: Tran, Brandon, et al.
Veröffentlicht: (2026)
von: Tran, Brandon, et al.
Veröffentlicht: (2026)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
von: Cheng, Long, et al.
Veröffentlicht: (2026)
von: Cheng, Long, et al.
Veröffentlicht: (2026)
Mitigating the Memory Bottleneck with Machine Learning-Driven and Data-Aware Microarchitectural Techniques
von: Bera, Rahul
Veröffentlicht: (2026)
von: Bera, Rahul
Veröffentlicht: (2026)
Compute Can't Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure
von: Jung, Myoungsoo
Veröffentlicht: (2025)
von: Jung, Myoungsoo
Veröffentlicht: (2025)
Exploring the Design Space for Message-Driven Systems for Dynamic Graph Processing using CCA
von: Chandio, Bibrak Qamar, et al.
Veröffentlicht: (2024)
von: Chandio, Bibrak Qamar, et al.
Veröffentlicht: (2024)
LLaMCAT: Optimizing Large Language Model Inference with Cache Arbitration and Throttling
von: Zhou, Zhongchun, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongchun, et al.
Veröffentlicht: (2025)
Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects
von: De Sensi, Daniele, et al.
Veröffentlicht: (2024)
von: De Sensi, Daniele, et al.
Veröffentlicht: (2024)
Inside VOLT: Designing an Open-Source GPU Compiler
von: Jeong, Shinnung, et al.
Veröffentlicht: (2025)
von: Jeong, Shinnung, et al.
Veröffentlicht: (2025)
Splitwise: Efficient generative LLM inference using phase splitting
von: Patel, Pratyush, et al.
Veröffentlicht: (2023)
von: Patel, Pratyush, et al.
Veröffentlicht: (2023)
Optimizing Foundation Model Inference on a Many-tiny-core Open-source RISC-V Platform
von: Potocnik, Viviane, et al.
Veröffentlicht: (2024)
von: Potocnik, Viviane, et al.
Veröffentlicht: (2024)
Global Optimizations & Lightweight Dynamic Logic for Concurrency
von: Pati, Suchita, et al.
Veröffentlicht: (2024)
von: Pati, Suchita, et al.
Veröffentlicht: (2024)
Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
von: Iliakopoulou, Nikoleta, et al.
Veröffentlicht: (2024)
von: Iliakopoulou, Nikoleta, et al.
Veröffentlicht: (2024)
Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
von: Georgiou, Athos
Veröffentlicht: (2026)
von: Georgiou, Athos
Veröffentlicht: (2026)
Sustainable Supercomputing for AI: GPU Power Capping at HPC Scale
von: Zhao, Dan, et al.
Veröffentlicht: (2024)
von: Zhao, Dan, et al.
Veröffentlicht: (2024)
TriADA: Massively Parallel Trilinear Matrix-by-Tensor Multiply-Add Algorithm and Device Architecture for the Acceleration of 3D Discrete Transformations
von: Sedukhin, Stanislav, et al.
Veröffentlicht: (2025)
von: Sedukhin, Stanislav, et al.
Veröffentlicht: (2025)
Factor Machine: Mixed-signal Architecture for Fine-Grained Graph-Based Computing
von: Dudek, Piotr
Veröffentlicht: (2024)
von: Dudek, Piotr
Veröffentlicht: (2024)
A Survey on Heterogeneous Computing Using SmartNICs and Emerging Data Processing Units
von: Tibbetts, Nathan, et al.
Veröffentlicht: (2025)
von: Tibbetts, Nathan, et al.
Veröffentlicht: (2025)
Athena: Synergizing Data Prefetching and Off-Chip Prediction via Online Reinforcement Learning
von: Bera, Rahul, et al.
Veröffentlicht: (2026)
von: Bera, Rahul, et al.
Veröffentlicht: (2026)
Flex-MIG: Enabling Distributed Execution on MIG
von: Kim, Myeongsu, et al.
Veröffentlicht: (2025)
von: Kim, Myeongsu, et al.
Veröffentlicht: (2025)
Application-Driven Exascale: The JUPITER Benchmark Suite
von: Herten, Andreas, et al.
Veröffentlicht: (2024)
von: Herten, Andreas, et al.
Veröffentlicht: (2024)
FPGA-Accelerated Lock Management and Transaction Processing: Architecture, Optimization, and Design Space Exploration
von: Zhu, Shien, et al.
Veröffentlicht: (2026)
von: Zhu, Shien, et al.
Veröffentlicht: (2026)
pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup Tables
von: Ferreira, João Dinis, et al.
Veröffentlicht: (2021)
von: Ferreira, João Dinis, et al.
Veröffentlicht: (2021)
Coordinated Reinforcement Learning Prefetching Architecture for Multicore Systems
von: Siddiqui, Mohammed Humaid, et al.
Veröffentlicht: (2025)
von: Siddiqui, Mohammed Humaid, et al.
Veröffentlicht: (2025)
Scheduler-Driven Job Atomization
von: Konopa, Michal, et al.
Veröffentlicht: (2025)
von: Konopa, Michal, et al.
Veröffentlicht: (2025)
JASDA: Introducing Job-Aware Scheduling in Scheduler-Driven Job Atomization
von: Konopa, Michal, et al.
Veröffentlicht: (2025)
von: Konopa, Michal, et al.
Veröffentlicht: (2025)
Ten-Four: An Open-Source Fused Dot Product Unit for Mixed-Precision GPGPU Tensor Cores
von: Rout, Nikhil, et al.
Veröffentlicht: (2025)
von: Rout, Nikhil, et al.
Veröffentlicht: (2025)
Accelerating Frontier MoE Training with 3D Integrated Optics
von: Bernadskiy, Mikhail, et al.
Veröffentlicht: (2025)
von: Bernadskiy, Mikhail, et al.
Veröffentlicht: (2025)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
von: Ganjihal, Sanjeev Rao
Veröffentlicht: (2026)
von: Ganjihal, Sanjeev Rao
Veröffentlicht: (2026)
Directives for Function Offloading in 5G Networks Based on a Performance Characteristics Analysis
von: Dettinger, Falk, et al.
Veröffentlicht: (2025)
von: Dettinger, Falk, et al.
Veröffentlicht: (2025)
NCCL EP: Towards a Unified Expert Parallel Communication API for NCCL
von: Goldman, Amos, et al.
Veröffentlicht: (2026)
von: Goldman, Amos, et al.
Veröffentlicht: (2026)
A comprehensive evaluation of spatial co-execution on GPUs using MPS and MIG technologies
von: Villarrubia, Jorge, et al.
Veröffentlicht: (2026)
von: Villarrubia, Jorge, et al.
Veröffentlicht: (2026)
GPU-Augmented OLAP Execution Engine: GPU Offloading
von: Chang, Ilsun
Veröffentlicht: (2025)
von: Chang, Ilsun
Veröffentlicht: (2025)
Efficient and Scalable Architecture for Multiple-chip Implementation of Simulated Bifurcation Machines
von: Kashimata, Tomoya, et al.
Veröffentlicht: (2023)
von: Kashimata, Tomoya, et al.
Veröffentlicht: (2023)
Sky$^ε$-Tree: Embracing the Batch Updates of B$^ε$-trees through Access Port Parallelism on Skyrmion Racetrack Memory
von: Tsai, Yu-Shiang, et al.
Veröffentlicht: (2024)
von: Tsai, Yu-Shiang, et al.
Veröffentlicht: (2024)
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures
von: Zhao, Chenggang, et al.
Veröffentlicht: (2025)
von: Zhao, Chenggang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Toward a Universal GPU Instruction Set Architecture: A Cross-Vendor Analysis of Hardware-Invariant Computational Primitives in Parallel Processors
von: Abraham, Ojima, et al.
Veröffentlicht: (2026) -
Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System Modeling
von: Li, Yinrong, et al.
Veröffentlicht: (2026) -
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
von: Jo, Myeong Jun
Veröffentlicht: (2026) -
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
von: Pati, Suchita, et al.
Veröffentlicht: (2024) -
Wattchmen: Watching the Wattchers -- High Fidelity, Flexible GPU Energy Modeling
von: Tran, Brandon, et al.
Veröffentlicht: (2026)