MSCCL++: Rethinking GPU Communication Abstractions for AI Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hwang, Changho, Cheng, Peng, Dathathri, Roshan, Jangda, Abhinav, Maleki, Saeed, Musuvathi, Madan, Saarikivi, Olli, Shah, Aashaka, Yang, Ziyue, Li, Binyang, Rocha, Caio, Zhou, Qinghua, Ghazimirsaeed, Mahdieh, Anantharamu, Sreevatsa, Jose, Jithin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Framework for Fine-Grained Synchronization of Dependent GPU Kernels
von: Jangda, Abhinav, et al.
Veröffentlicht: (2023)
von: Jangda, Abhinav, et al.
Veröffentlicht: (2023)
PEAK: A Performance Engineering AI-Assistant for GPU Kernels Powered by Natural Language Transformations
von: Tariq, Muhammad Usman, et al.
Veröffentlicht: (2025)
von: Tariq, Muhammad Usman, et al.
Veröffentlicht: (2025)
A Non‐Dissipative, Energy‐Conserving, Arbitrary High‐Order Numerical Method and Its Efficient Implementation for Incompressible Flow Simulation in Complex Geometries
von: Sreevatsa Anantharamu, et al.
Veröffentlicht: (2024)
von: Sreevatsa Anantharamu, et al.
Veröffentlicht: (2024)
Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants
von: You, Bozhi, et al.
Veröffentlicht: (2025)
von: You, Bozhi, et al.
Veröffentlicht: (2025)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
von: Jangda, Abhinav, et al.
Veröffentlicht: (2024)
von: Jangda, Abhinav, et al.
Veröffentlicht: (2024)
How Many Parameters Does Your Task Really Need? Task Specific Pruning with LLM-Sieve
von: Reda, Waleed, et al.
Veröffentlicht: (2025)
von: Reda, Waleed, et al.
Veröffentlicht: (2025)
Splitwise: Efficient generative LLM inference using phase splitting
von: Patel, Pratyush, et al.
Veröffentlicht: (2023)
von: Patel, Pratyush, et al.
Veröffentlicht: (2023)
GPU-Virt-Bench: A Comprehensive Benchmarking Framework for Software-Based GPU Virtualization Systems
von: VG, Jithin, et al.
Veröffentlicht: (2025)
von: VG, Jithin, et al.
Veröffentlicht: (2025)
LLM-Vectorizer: LLM-based Verified Loop Vectorizer
von: Taneja, Jubi, et al.
Veröffentlicht: (2024)
von: Taneja, Jubi, et al.
Veröffentlicht: (2024)
Abstraction-based Control of Unknown Continuous-Space Models with Just Two Trajectories
von: Samari, Behrad, et al.
Veröffentlicht: (2024)
von: Samari, Behrad, et al.
Veröffentlicht: (2024)
CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning
von: Zhu, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xinyu, et al.
Veröffentlicht: (2026)
Impact of Landslide on Tourism in Wayanad: A Study on Customer Satisfaction and Recovery Efforts
von: Jithin Scaria
Veröffentlicht: (2025)
von: Jithin Scaria
Veröffentlicht: (2025)
Beyond Component Strength: Synergistic Integration and Adaptive Calibration in Multi-Agent RAG Systems
von: Krishnan, Jithin
Veröffentlicht: (2025)
von: Krishnan, Jithin
Veröffentlicht: (2025)
Basic Legibility Protocols Improve Trusted Monitoring
von: Sreevatsa, Ashwin, et al.
Veröffentlicht: (2026)
von: Sreevatsa, Ashwin, et al.
Veröffentlicht: (2026)
Information Theory for Data Science
von: Suh, Changho
Veröffentlicht: (2024)
von: Suh, Changho
Veröffentlicht: (2024)
Hilbert scheme of smooth curves of degree sixteen in $\mathbb{P}^5$
von: Keem, Changho
Veröffentlicht: (2025)
von: Keem, Changho
Veröffentlicht: (2025)
Convex Optimization for Machine Learning
von: Suh, Changho
Veröffentlicht: (2023)
von: Suh, Changho
Veröffentlicht: (2023)
Hilbert scheme of linearly normal curves in $\mathbb{P}^r$ with index of speciality five and beyond
von: Keem, Changho
Veröffentlicht: (2023)
von: Keem, Changho
Veröffentlicht: (2023)
Hilbert scheme of smooth projective curves of unexpected dimension \& existence of a component with less than the expected number of moduli
von: Keem, Changho
Veröffentlicht: (2026)
von: Keem, Changho
Veröffentlicht: (2026)
Watermarking Needs Input Repetition Masking
von: Khachaturov, David, et al.
Veröffentlicht: (2025)
von: Khachaturov, David, et al.
Veröffentlicht: (2025)
SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication
von: Khalilov, Mikhail, et al.
Veröffentlicht: (2025)
von: Khalilov, Mikhail, et al.
Veröffentlicht: (2025)
ForestColl: Throughput-Optimal Collective Communications on Heterogeneous Network Fabrics
von: Zhao, Liangyu, et al.
Veröffentlicht: (2024)
von: Zhao, Liangyu, et al.
Veröffentlicht: (2024)
SYNTHESIS AND EVOLUTION: CULTURAL SYNCRETISM AND THE HISTORICAL FORMATION OF THE CONTEMPORARY HINDU SPIRITUAL LANDSCAPE
von: Jithin Sankar N
Veröffentlicht: (2025)
von: Jithin Sankar N
Veröffentlicht: (2025)
Cl+ and HCl+ in Reaction with H2 and Isotopologues: A Glance into H Abstraction and Indirect Exchange at Astrophysical Conditions
von: Jiménez-Redondo, Miguel, et al.
Veröffentlicht: (2025)
von: Jiménez-Redondo, Miguel, et al.
Veröffentlicht: (2025)
Rethinking Analytical Processing in the GPU Era
von: Yogatama, Bobbi, et al.
Veröffentlicht: (2025)
von: Yogatama, Bobbi, et al.
Veröffentlicht: (2025)
Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods
von: Jang, Yeonwoo, et al.
Veröffentlicht: (2025)
von: Jang, Yeonwoo, et al.
Veröffentlicht: (2025)
ForensiBlock: A Provenance-Driven Blockchain Framework for Data Forensics and Auditability
von: Akbarfam, Asma Jodeiri, et al.
Veröffentlicht: (2023)
von: Akbarfam, Asma Jodeiri, et al.
Veröffentlicht: (2023)
Stochastic Barnes-Hut Approximation for Fast Summation on the GPU
von: Madan, Abhishek, et al.
Veröffentlicht: (2025)
von: Madan, Abhishek, et al.
Veröffentlicht: (2025)
When Abstraction Breaks Physics: Rethinking Modular Design in Quantum Software
von: Zhao, Jianjun
Veröffentlicht: (2025)
von: Zhao, Jianjun
Veröffentlicht: (2025)
The epigenetic scar: How reproductive violence will shape generations of health in Gaza
von: Kanza Farhan, et al.
Veröffentlicht: (2025)
von: Kanza Farhan, et al.
Veröffentlicht: (2025)
Co-Design and Evaluation of a CPU-Free MPI GPU Communication Abstraction and Implementation
von: Bridges, Patrick G., et al.
Veröffentlicht: (2026)
von: Bridges, Patrick G., et al.
Veröffentlicht: (2026)
GPU-Native Multi-Area State Estimation via SIMD Abstraction and Boundary Condensation
von: Xu, Yifei, et al.
Veröffentlicht: (2026)
von: Xu, Yifei, et al.
Veröffentlicht: (2026)
Hilbert scheme and Hilbert functions of smooth curves of degrees at most $15$ in $\mathbb{P}^5$
von: Ballico, Edoardo, et al.
Veröffentlicht: (2025)
von: Ballico, Edoardo, et al.
Veröffentlicht: (2025)
Secant rank and syzygies of projections of elliptic normal curves
von: Han, Changho, et al.
Veröffentlicht: (2026)
von: Han, Changho, et al.
Veröffentlicht: (2026)
The stacky Batyrev-Manin conjecture and modular curves
von: Darda, Ratko, et al.
Veröffentlicht: (2026)
von: Darda, Ratko, et al.
Veröffentlicht: (2026)
On the Hilbert scheme of smooth curves of degree $d=15$ in $\mathbb{P}^5$
von: Ballico, Edoardo, et al.
Veröffentlicht: (2023)
von: Ballico, Edoardo, et al.
Veröffentlicht: (2023)
Multi-modal Machine Learning for Vehicle Rating Predictions Using Image, Text, and Parametric Data
von: Su, Hanqi, et al.
Veröffentlicht: (2023)
von: Su, Hanqi, et al.
Veröffentlicht: (2023)
Reservoir Sampling over Joins
von: Dai, Binyang, et al.
Veröffentlicht: (2024)
von: Dai, Binyang, et al.
Veröffentlicht: (2024)
Polyzwitterionic Organohydrogel and Soft Composite with Tunable Sol–Gel Properties Enabling On‐Demand Functionalization with Colloids
von: Ziyue Miao, et al.
Veröffentlicht: (2025)
von: Ziyue Miao, et al.
Veröffentlicht: (2025)
Thinking with Visual Abstract: Enhancing Multimodal Reasoning via Visual Abstraction
von: Liu, Dairu, et al.
Veröffentlicht: (2025)
von: Liu, Dairu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Framework for Fine-Grained Synchronization of Dependent GPU Kernels
von: Jangda, Abhinav, et al.
Veröffentlicht: (2023) -
PEAK: A Performance Engineering AI-Assistant for GPU Kernels Powered by Natural Language Transformations
von: Tariq, Muhammad Usman, et al.
Veröffentlicht: (2025) -
A Non‐Dissipative, Energy‐Conserving, Arbitrary High‐Order Numerical Method and Its Efficient Implementation for Incompressible Flow Simulation in Complex Geometries
von: Sreevatsa Anantharamu, et al.
Veröffentlicht: (2024) -
Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants
von: You, Bozhi, et al.
Veröffentlicht: (2025) -
Fast Kronecker Matrix-Matrix Multiplication on GPUs
von: Jangda, Abhinav, et al.
Veröffentlicht: (2024)