Toward a Universal GPU Instruction Set Architecture: A Cross-Vendor Analysis of Hardware-Invariant Computational Primitives in Parallel Processors
Fuente:
arXiv
Saved in:
| Main Authors: | Abraham, Ojima, Okoli, Onyinye |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lincoln AI Computing Survey (LAICS) and Trends
by: Reuther, Albert, et al.
Published: (2025)
by: Reuther, Albert, et al.
Published: (2025)
Inside VOLT: Designing an Open-Source GPU Compiler
by: Jeong, Shinnung, et al.
Published: (2025)
by: Jeong, Shinnung, et al.
Published: (2025)
Efficient and Scalable Architecture for Multiple-chip Implementation of Simulated Bifurcation Machines
by: Kashimata, Tomoya, et al.
Published: (2023)
by: Kashimata, Tomoya, et al.
Published: (2023)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
by: Jo, Myeong Jun
Published: (2026)
by: Jo, Myeong Jun
Published: (2026)
Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects
by: De Sensi, Daniele, et al.
Published: (2024)
by: De Sensi, Daniele, et al.
Published: (2024)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
by: Peng, Hongwu, et al.
Published: (2023)
by: Peng, Hongwu, et al.
Published: (2023)
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
by: Pati, Suchita, et al.
Published: (2024)
by: Pati, Suchita, et al.
Published: (2024)
Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System Modeling
by: Li, Yinrong, et al.
Published: (2026)
by: Li, Yinrong, et al.
Published: (2026)
GPU-Augmented OLAP Execution Engine: GPU Offloading
by: Chang, Ilsun
Published: (2025)
by: Chang, Ilsun
Published: (2025)
Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
by: Georgiou, Athos
Published: (2026)
by: Georgiou, Athos
Published: (2026)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
by: Ganjihal, Sanjeev Rao
Published: (2026)
by: Ganjihal, Sanjeev Rao
Published: (2026)
Optimizing Foundation Model Inference on a Many-tiny-core Open-source RISC-V Platform
by: Potocnik, Viviane, et al.
Published: (2024)
by: Potocnik, Viviane, et al.
Published: (2024)
Mitigating the Memory Bottleneck with Machine Learning-Driven and Data-Aware Microarchitectural Techniques
by: Bera, Rahul
Published: (2026)
by: Bera, Rahul
Published: (2026)
Sky$^ε$-Tree: Embracing the Batch Updates of B$^ε$-trees through Access Port Parallelism on Skyrmion Racetrack Memory
by: Tsai, Yu-Shiang, et al.
Published: (2024)
by: Tsai, Yu-Shiang, et al.
Published: (2024)
Wattchmen: Watching the Wattchers -- High Fidelity, Flexible GPU Energy Modeling
by: Tran, Brandon, et al.
Published: (2026)
by: Tran, Brandon, et al.
Published: (2026)
Deep Recommender Models Inference: Automatic Asymmetric Data Flow Optimization
by: Ruggeri, Giuseppe, et al.
Published: (2025)
by: Ruggeri, Giuseppe, et al.
Published: (2025)
Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
by: Iliakopoulou, Nikoleta, et al.
Published: (2024)
by: Iliakopoulou, Nikoleta, et al.
Published: (2024)
Compute Can't Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure
by: Jung, Myoungsoo
Published: (2025)
by: Jung, Myoungsoo
Published: (2025)
Exploring the Design Space for Message-Driven Systems for Dynamic Graph Processing using CCA
by: Chandio, Bibrak Qamar, et al.
Published: (2024)
by: Chandio, Bibrak Qamar, et al.
Published: (2024)
Unlocking Python's Cores: Hardware Usage and Energy Implications of Removing the GIL
by: Salazar, José Daniel Montoya
Published: (2026)
by: Salazar, José Daniel Montoya
Published: (2026)
Splitwise: Efficient generative LLM inference using phase splitting
by: Patel, Pratyush, et al.
Published: (2023)
by: Patel, Pratyush, et al.
Published: (2023)
Directives for Function Offloading in 5G Networks Based on a Performance Characteristics Analysis
by: Dettinger, Falk, et al.
Published: (2025)
by: Dettinger, Falk, et al.
Published: (2025)
FCDP: Fully Cached Data Parallel for Communication-Avoiding Large-Scale Training
by: Park, Gyeongseo, et al.
Published: (2026)
by: Park, Gyeongseo, et al.
Published: (2026)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
by: Zhang, Chen, et al.
Published: (2026)
by: Zhang, Chen, et al.
Published: (2026)
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures
by: Zhao, Chenggang, et al.
Published: (2025)
by: Zhao, Chenggang, et al.
Published: (2025)
TriADA: Massively Parallel Trilinear Matrix-by-Tensor Multiply-Add Algorithm and Device Architecture for the Acceleration of 3D Discrete Transformations
by: Sedukhin, Stanislav, et al.
Published: (2025)
by: Sedukhin, Stanislav, et al.
Published: (2025)
Serverless GPU Architecture for Enterprise HR Analytics: A Production-Scale BDaaS Implementation
by: Zhang, Guilin, et al.
Published: (2025)
by: Zhang, Guilin, et al.
Published: (2025)
How Fast Can Graph Computations Go on Fine-grained Parallel Architectures
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
by: Cheng, Long, et al.
Published: (2026)
by: Cheng, Long, et al.
Published: (2026)
Kunlun Anomaly Troubleshooter: Enabling Kernel-Level Anomaly Detection and Causal Reasoning for Large Model Distributed Inference
by: Liu, Yuyang, et al.
Published: (2025)
by: Liu, Yuyang, et al.
Published: (2025)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
by: Qiu, Tong Dong, et al.
Published: (2023)
by: Qiu, Tong Dong, et al.
Published: (2023)
GigaAPI for GPU Parallelization
by: Suvarna, M., et al.
Published: (2025)
by: Suvarna, M., et al.
Published: (2025)
Cognitive Infrastructure: A Unified DCIM Framework for AI Data Centers
by: Sunkara, Krishna Chaitanya
Published: (2026)
by: Sunkara, Krishna Chaitanya
Published: (2026)
Parallelizing a modern GPU simulator
by: Huerta, Rodrigo, et al.
Published: (2025)
by: Huerta, Rodrigo, et al.
Published: (2025)
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
by: He, Jiaao, et al.
Published: (2024)
by: He, Jiaao, et al.
Published: (2024)
Nezha: Deployable and High-Performance Consensus Using Synchronized Clocks
by: Geng, Jinkun, et al.
Published: (2022)
by: Geng, Jinkun, et al.
Published: (2022)
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
by: Jarmusch, Aaron, et al.
Published: (2026)
by: Jarmusch, Aaron, et al.
Published: (2026)
LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear Programming
by: Shen, Siyuan, et al.
Published: (2024)
by: Shen, Siyuan, et al.
Published: (2024)
Moonshot: Optimizing Chain-Based Rotating Leader BFT via Optimistic Proposals
by: Doidge, Isaac, et al.
Published: (2024)
by: Doidge, Isaac, et al.
Published: (2024)
Exploiting Spot Instances for Time-Critical Cloud Workloads Using Optimal Randomized Strategies
by: Bhuyan, Neelkamal, et al.
Published: (2026)
by: Bhuyan, Neelkamal, et al.
Published: (2026)
Similar Items
-
Lincoln AI Computing Survey (LAICS) and Trends
by: Reuther, Albert, et al.
Published: (2025) -
Inside VOLT: Designing an Open-Source GPU Compiler
by: Jeong, Shinnung, et al.
Published: (2025) -
Efficient and Scalable Architecture for Multiple-chip Implementation of Simulated Bifurcation Machines
by: Kashimata, Tomoya, et al.
Published: (2023) -
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
by: Jo, Myeong Jun
Published: (2026) -
Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects
by: De Sensi, Daniele, et al.
Published: (2024)