Coordinated Reinforcement Learning Prefetching Architecture for Multicore Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Siddiqui, Mohammed Humaid, Guzman, Fernando, Wu, Yufei, Ann, Ruishu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The $qs$ Inequality: Quantifying the Double Penalty of Mixture-of-Experts at Inference
von: Adhinarayanan, Vignesh, et al.
Veröffentlicht: (2026)
von: Adhinarayanan, Vignesh, et al.
Veröffentlicht: (2026)
Accelerating State-Vector Quantum Simulation on Integrated GPUs via Cache Locality Optimization: A Cross-Architecture Evaluation
von: Thomaz, Gabriel Fernandes, et al.
Veröffentlicht: (2026)
von: Thomaz, Gabriel Fernandes, et al.
Veröffentlicht: (2026)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
von: Jo, Myeong Jun
Veröffentlicht: (2026)
von: Jo, Myeong Jun
Veröffentlicht: (2026)
Global Optimizations & Lightweight Dynamic Logic for Concurrency
von: Pati, Suchita, et al.
Veröffentlicht: (2024)
von: Pati, Suchita, et al.
Veröffentlicht: (2024)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
von: Cheng, Long, et al.
Veröffentlicht: (2026)
von: Cheng, Long, et al.
Veröffentlicht: (2026)
Cross-Platform Fused MoE Dispatch in Triton: Portable Expert Routing Without CUDA
von: Mitra, Subhadip
Veröffentlicht: (2026)
von: Mitra, Subhadip
Veröffentlicht: (2026)
Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System Modeling
von: Li, Yinrong, et al.
Veröffentlicht: (2026)
von: Li, Yinrong, et al.
Veröffentlicht: (2026)
NCCL EP: Towards a Unified Expert Parallel Communication API for NCCL
von: Goldman, Amos, et al.
Veröffentlicht: (2026)
von: Goldman, Amos, et al.
Veröffentlicht: (2026)
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
von: Pati, Suchita, et al.
Veröffentlicht: (2024)
von: Pati, Suchita, et al.
Veröffentlicht: (2024)
Efficient and Scalable Architecture for Multiple-chip Implementation of Simulated Bifurcation Machines
von: Kashimata, Tomoya, et al.
Veröffentlicht: (2023)
von: Kashimata, Tomoya, et al.
Veröffentlicht: (2023)
GPU-Initiated Networking for NCCL
von: Hamidouche, Khaled, et al.
Veröffentlicht: (2025)
von: Hamidouche, Khaled, et al.
Veröffentlicht: (2025)
Athena: Synergizing Data Prefetching and Off-Chip Prediction via Online Reinforcement Learning
von: Bera, Rahul, et al.
Veröffentlicht: (2026)
von: Bera, Rahul, et al.
Veröffentlicht: (2026)
SLED: A Speculative LLM Decoding Framework for Efficient Edge Serving
von: Li, Xiangchen, et al.
Veröffentlicht: (2025)
von: Li, Xiangchen, et al.
Veröffentlicht: (2025)
Splitwise: Efficient generative LLM inference using phase splitting
von: Patel, Pratyush, et al.
Veröffentlicht: (2023)
von: Patel, Pratyush, et al.
Veröffentlicht: (2023)
FPGA-Accelerated Lock Management and Transaction Processing: Architecture, Optimization, and Design Space Exploration
von: Zhu, Shien, et al.
Veröffentlicht: (2026)
von: Zhu, Shien, et al.
Veröffentlicht: (2026)
Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects
von: De Sensi, Daniele, et al.
Veröffentlicht: (2024)
von: De Sensi, Daniele, et al.
Veröffentlicht: (2024)
Next-Generation Event-Driven Architectures: Performance, Scalability, and Intelligent Orchestration Across Messaging Frameworks
von: Arafat, Jahidul, et al.
Veröffentlicht: (2025)
von: Arafat, Jahidul, et al.
Veröffentlicht: (2025)
Inside VOLT: Designing an Open-Source GPU Compiler
von: Jeong, Shinnung, et al.
Veröffentlicht: (2025)
von: Jeong, Shinnung, et al.
Veröffentlicht: (2025)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
von: Ganjihal, Sanjeev Rao
Veröffentlicht: (2026)
von: Ganjihal, Sanjeev Rao
Veröffentlicht: (2026)
Application-Driven Exascale: The JUPITER Benchmark Suite
von: Herten, Andreas, et al.
Veröffentlicht: (2024)
von: Herten, Andreas, et al.
Veröffentlicht: (2024)
Accelerating Frontier MoE Training with 3D Integrated Optics
von: Bernadskiy, Mikhail, et al.
Veröffentlicht: (2025)
von: Bernadskiy, Mikhail, et al.
Veröffentlicht: (2025)
LLaMCAT: Optimizing Large Language Model Inference with Cache Arbitration and Throttling
von: Zhou, Zhongchun, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongchun, et al.
Veröffentlicht: (2025)
Optimizing Multi-DNN Inference on Mobile Devices through Heterogeneous Processor Co-Execution
von: Gao, Yunquan, et al.
Veröffentlicht: (2025)
von: Gao, Yunquan, et al.
Veröffentlicht: (2025)
Wattchmen: Watching the Wattchers -- High Fidelity, Flexible GPU Energy Modeling
von: Tran, Brandon, et al.
Veröffentlicht: (2026)
von: Tran, Brandon, et al.
Veröffentlicht: (2026)
Efficiently Scheduling Parallel DAG Tasks on Identical Multiprocessors
von: Lendve, Shardul, et al.
Veröffentlicht: (2024)
von: Lendve, Shardul, et al.
Veröffentlicht: (2024)
Readout-Side Bypass for Residual Hybrid Quantum-Classical Models
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
von: Georgiou, Athos
Veröffentlicht: (2026)
von: Georgiou, Athos
Veröffentlicht: (2026)
LAMMPS-KOKKOS: Performance Portable Molecular Dynamics Across Exascale Architectures
von: Johansson, Anders, et al.
Veröffentlicht: (2025)
von: Johansson, Anders, et al.
Veröffentlicht: (2025)
Compiler Support for Speculation in Decoupled Access/Execute Architectures
von: Szafarczyk, Robert, et al.
Veröffentlicht: (2025)
von: Szafarczyk, Robert, et al.
Veröffentlicht: (2025)
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
von: Luo, Weile, et al.
Veröffentlicht: (2025)
von: Luo, Weile, et al.
Veröffentlicht: (2025)
General-Purpose Multicore Architectures
von: Ghose, Saugata
Veröffentlicht: (2024)
von: Ghose, Saugata
Veröffentlicht: (2024)
SGDRC: Software-Defined Dynamic Resource Control for Concurrent DNN Inference on NVIDIA GPUs
von: Zhang, Yongkang, et al.
Veröffentlicht: (2024)
von: Zhang, Yongkang, et al.
Veröffentlicht: (2024)
FlashSpread: IO-Aware GPU Simulation of Non-Markovian Epidemic Dynamics via Kernel Fusion
von: Shakeri, Heman, et al.
Veröffentlicht: (2026)
von: Shakeri, Heman, et al.
Veröffentlicht: (2026)
RapidStream IR: Infrastructure for FPGA High-Level Physical Synthesis
von: Lau, Jason, et al.
Veröffentlicht: (2024)
von: Lau, Jason, et al.
Veröffentlicht: (2024)
Optimizing Foundation Model Inference on a Many-tiny-core Open-source RISC-V Platform
von: Potocnik, Viviane, et al.
Veröffentlicht: (2024)
von: Potocnik, Viviane, et al.
Veröffentlicht: (2024)
Deep Learning Model Deployment in Multiple Cloud Providers: an Exploratory Study Using Low Computing Power Environments
von: Lemos, Elayne, et al.
Veröffentlicht: (2025)
von: Lemos, Elayne, et al.
Veröffentlicht: (2025)
Compute Can't Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure
von: Jung, Myoungsoo
Veröffentlicht: (2025)
von: Jung, Myoungsoo
Veröffentlicht: (2025)
Method for determining the acceleration of a parallel specialised computer system based on Amdahl's law
von: Filipchenko, Aleksandr S.
Veröffentlicht: (2024)
von: Filipchenko, Aleksandr S.
Veröffentlicht: (2024)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
The $qs$ Inequality: Quantifying the Double Penalty of Mixture-of-Experts at Inference
von: Adhinarayanan, Vignesh, et al.
Veröffentlicht: (2026) -
Accelerating State-Vector Quantum Simulation on Integrated GPUs via Cache Locality Optimization: A Cross-Architecture Evaluation
von: Thomaz, Gabriel Fernandes, et al.
Veröffentlicht: (2026) -
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
von: Jo, Myeong Jun
Veröffentlicht: (2026) -
Global Optimizations & Lightweight Dynamic Logic for Concurrency
von: Pati, Suchita, et al.
Veröffentlicht: (2024) -
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
von: Cheng, Long, et al.
Veröffentlicht: (2026)