Lincoln AI Computing Survey (LAICS) and Trends
Fuente:
arXiv
Saved in:
| Main Authors: | Reuther, Albert, Michaleas, Peter, Jones, Michael, Gadepally, Vijay, Kepner, Jeremy |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Toward a Universal GPU Instruction Set Architecture: A Cross-Vendor Analysis of Hardware-Invariant Computational Primitives in Parallel Processors
by: Abraham, Ojima, et al.
Published: (2026)
by: Abraham, Ojima, et al.
Published: (2026)
Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System Modeling
by: Li, Yinrong, et al.
Published: (2026)
by: Li, Yinrong, et al.
Published: (2026)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
by: Jo, Myeong Jun
Published: (2026)
by: Jo, Myeong Jun
Published: (2026)
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
by: Pati, Suchita, et al.
Published: (2024)
by: Pati, Suchita, et al.
Published: (2024)
Wattchmen: Watching the Wattchers -- High Fidelity, Flexible GPU Energy Modeling
by: Tran, Brandon, et al.
Published: (2026)
by: Tran, Brandon, et al.
Published: (2026)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
by: Peng, Hongwu, et al.
Published: (2023)
by: Peng, Hongwu, et al.
Published: (2023)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
by: Cheng, Long, et al.
Published: (2026)
by: Cheng, Long, et al.
Published: (2026)
Mitigating the Memory Bottleneck with Machine Learning-Driven and Data-Aware Microarchitectural Techniques
by: Bera, Rahul
Published: (2026)
by: Bera, Rahul
Published: (2026)
Compute Can't Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure
by: Jung, Myoungsoo
Published: (2025)
by: Jung, Myoungsoo
Published: (2025)
Exploring the Design Space for Message-Driven Systems for Dynamic Graph Processing using CCA
by: Chandio, Bibrak Qamar, et al.
Published: (2024)
by: Chandio, Bibrak Qamar, et al.
Published: (2024)
LLaMCAT: Optimizing Large Language Model Inference with Cache Arbitration and Throttling
by: Zhou, Zhongchun, et al.
Published: (2025)
by: Zhou, Zhongchun, et al.
Published: (2025)
Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects
by: De Sensi, Daniele, et al.
Published: (2024)
by: De Sensi, Daniele, et al.
Published: (2024)
Inside VOLT: Designing an Open-Source GPU Compiler
by: Jeong, Shinnung, et al.
Published: (2025)
by: Jeong, Shinnung, et al.
Published: (2025)
Splitwise: Efficient generative LLM inference using phase splitting
by: Patel, Pratyush, et al.
Published: (2023)
by: Patel, Pratyush, et al.
Published: (2023)
Optimizing Foundation Model Inference on a Many-tiny-core Open-source RISC-V Platform
by: Potocnik, Viviane, et al.
Published: (2024)
by: Potocnik, Viviane, et al.
Published: (2024)
Global Optimizations & Lightweight Dynamic Logic for Concurrency
by: Pati, Suchita, et al.
Published: (2024)
by: Pati, Suchita, et al.
Published: (2024)
Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
by: Iliakopoulou, Nikoleta, et al.
Published: (2024)
by: Iliakopoulou, Nikoleta, et al.
Published: (2024)
Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
by: Georgiou, Athos
Published: (2026)
by: Georgiou, Athos
Published: (2026)
Sustainable Supercomputing for AI: GPU Power Capping at HPC Scale
by: Zhao, Dan, et al.
Published: (2024)
by: Zhao, Dan, et al.
Published: (2024)
TriADA: Massively Parallel Trilinear Matrix-by-Tensor Multiply-Add Algorithm and Device Architecture for the Acceleration of 3D Discrete Transformations
by: Sedukhin, Stanislav, et al.
Published: (2025)
by: Sedukhin, Stanislav, et al.
Published: (2025)
Factor Machine: Mixed-signal Architecture for Fine-Grained Graph-Based Computing
by: Dudek, Piotr
Published: (2024)
by: Dudek, Piotr
Published: (2024)
A Survey on Heterogeneous Computing Using SmartNICs and Emerging Data Processing Units
by: Tibbetts, Nathan, et al.
Published: (2025)
by: Tibbetts, Nathan, et al.
Published: (2025)
Athena: Synergizing Data Prefetching and Off-Chip Prediction via Online Reinforcement Learning
by: Bera, Rahul, et al.
Published: (2026)
by: Bera, Rahul, et al.
Published: (2026)
Flex-MIG: Enabling Distributed Execution on MIG
by: Kim, Myeongsu, et al.
Published: (2025)
by: Kim, Myeongsu, et al.
Published: (2025)
Application-Driven Exascale: The JUPITER Benchmark Suite
by: Herten, Andreas, et al.
Published: (2024)
by: Herten, Andreas, et al.
Published: (2024)
FPGA-Accelerated Lock Management and Transaction Processing: Architecture, Optimization, and Design Space Exploration
by: Zhu, Shien, et al.
Published: (2026)
by: Zhu, Shien, et al.
Published: (2026)
pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup Tables
by: Ferreira, João Dinis, et al.
Published: (2021)
by: Ferreira, João Dinis, et al.
Published: (2021)
Coordinated Reinforcement Learning Prefetching Architecture for Multicore Systems
by: Siddiqui, Mohammed Humaid, et al.
Published: (2025)
by: Siddiqui, Mohammed Humaid, et al.
Published: (2025)
Scheduler-Driven Job Atomization
by: Konopa, Michal, et al.
Published: (2025)
by: Konopa, Michal, et al.
Published: (2025)
JASDA: Introducing Job-Aware Scheduling in Scheduler-Driven Job Atomization
by: Konopa, Michal, et al.
Published: (2025)
by: Konopa, Michal, et al.
Published: (2025)
Ten-Four: An Open-Source Fused Dot Product Unit for Mixed-Precision GPGPU Tensor Cores
by: Rout, Nikhil, et al.
Published: (2025)
by: Rout, Nikhil, et al.
Published: (2025)
Accelerating Frontier MoE Training with 3D Integrated Optics
by: Bernadskiy, Mikhail, et al.
Published: (2025)
by: Bernadskiy, Mikhail, et al.
Published: (2025)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
by: Ganjihal, Sanjeev Rao
Published: (2026)
by: Ganjihal, Sanjeev Rao
Published: (2026)
Directives for Function Offloading in 5G Networks Based on a Performance Characteristics Analysis
by: Dettinger, Falk, et al.
Published: (2025)
by: Dettinger, Falk, et al.
Published: (2025)
NCCL EP: Towards a Unified Expert Parallel Communication API for NCCL
by: Goldman, Amos, et al.
Published: (2026)
by: Goldman, Amos, et al.
Published: (2026)
A comprehensive evaluation of spatial co-execution on GPUs using MPS and MIG technologies
by: Villarrubia, Jorge, et al.
Published: (2026)
by: Villarrubia, Jorge, et al.
Published: (2026)
GPU-Augmented OLAP Execution Engine: GPU Offloading
by: Chang, Ilsun
Published: (2025)
by: Chang, Ilsun
Published: (2025)
Efficient and Scalable Architecture for Multiple-chip Implementation of Simulated Bifurcation Machines
by: Kashimata, Tomoya, et al.
Published: (2023)
by: Kashimata, Tomoya, et al.
Published: (2023)
Sky$^ε$-Tree: Embracing the Batch Updates of B$^ε$-trees through Access Port Parallelism on Skyrmion Racetrack Memory
by: Tsai, Yu-Shiang, et al.
Published: (2024)
by: Tsai, Yu-Shiang, et al.
Published: (2024)
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures
by: Zhao, Chenggang, et al.
Published: (2025)
by: Zhao, Chenggang, et al.
Published: (2025)
Similar Items
-
Toward a Universal GPU Instruction Set Architecture: A Cross-Vendor Analysis of Hardware-Invariant Computational Primitives in Parallel Processors
by: Abraham, Ojima, et al.
Published: (2026) -
Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System Modeling
by: Li, Yinrong, et al.
Published: (2026) -
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
by: Jo, Myeong Jun
Published: (2026) -
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
by: Pati, Suchita, et al.
Published: (2024) -
Wattchmen: Watching the Wattchers -- High Fidelity, Flexible GPU Energy Modeling
by: Tran, Brandon, et al.
Published: (2026)