Cross-Platform Fused MoE Dispatch in Triton: Portable Expert Routing Without CUDA
Fuente:
arXiv
Guardado en:
| Autor principal: | Mitra, Subhadip |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Active Inference-Based Adaptive Routing for Heterogeneous Edge AI Services
por: Wang, Zihang, et al.
Publicado: (2026)
por: Wang, Zihang, et al.
Publicado: (2026)
The $qs$ Inequality: Quantifying the Double Penalty of Mixture-of-Experts at Inference
por: Adhinarayanan, Vignesh, et al.
Publicado: (2026)
por: Adhinarayanan, Vignesh, et al.
Publicado: (2026)
ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM Inference
por: Zhu, Xiongwei, et al.
Publicado: (2026)
por: Zhu, Xiongwei, et al.
Publicado: (2026)
Seamless acceleration of Fortran intrinsics via AMD AI engines
por: Brown, Nick, et al.
Publicado: (2025)
por: Brown, Nick, et al.
Publicado: (2025)
The Landscape of GPU-Centric Communication
por: Unat, Didem, et al.
Publicado: (2024)
por: Unat, Didem, et al.
Publicado: (2024)
On the Performance of Cloud-based ARM SVE for Zero-Knowledge Proving Systems
por: Loghin, Dumitrel, et al.
Publicado: (2025)
por: Loghin, Dumitrel, et al.
Publicado: (2025)
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
por: Afzal, Ayesha, et al.
Publicado: (2026)
por: Afzal, Ayesha, et al.
Publicado: (2026)
Coordinated Reinforcement Learning Prefetching Architecture for Multicore Systems
por: Siddiqui, Mohammed Humaid, et al.
Publicado: (2025)
por: Siddiqui, Mohammed Humaid, et al.
Publicado: (2025)
Simulating Cloud Environments of Connected Vehicles for Anomaly Detection
por: Weiß, M., et al.
Publicado: (2024)
por: Weiß, M., et al.
Publicado: (2024)
A Hybrid Heuristic Framework for Resource-Efficient Querying of Scientific Experiments Data
por: Patel, Mayank, et al.
Publicado: (2025)
por: Patel, Mayank, et al.
Publicado: (2025)
Efficient Construction of Large Search Spaces for Auto-Tuning
por: Willemsen, Floris-Jan, et al.
Publicado: (2025)
por: Willemsen, Floris-Jan, et al.
Publicado: (2025)
AMP4EC: Adaptive Model Partitioning Framework for Efficient Deep Learning Inference in Edge Computing Environments
por: Zhang, Guilin, et al.
Publicado: (2025)
por: Zhang, Guilin, et al.
Publicado: (2025)
Experimentally Evaluating the Resource Efficiency of Big Data Autoscaling
por: Will, Jonathan, et al.
Publicado: (2025)
por: Will, Jonathan, et al.
Publicado: (2025)
Creation of AI-driven Smart Spaces for Enhanced Indoor Environments -- A Survey
por: Varol, Aygün, et al.
Publicado: (2024)
por: Varol, Aygün, et al.
Publicado: (2024)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
por: Jo, Myeong Jun
Publicado: (2026)
por: Jo, Myeong Jun
Publicado: (2026)
Intelligent Cloud Orchestration: A Hybrid Predictive and Heuristic Framework for Cost Optimization
por: Nagoriya, Heet, et al.
Publicado: (2026)
por: Nagoriya, Heet, et al.
Publicado: (2026)
KPI2KVI: A Multi Agent Workflow for Calculating Key Value Indicators from Service Descriptions
por: Shokrnezhad, Masoud, et al.
Publicado: (2026)
por: Shokrnezhad, Masoud, et al.
Publicado: (2026)
E-QUARTIC: Energy Efficient Edge Ensemble of Convolutional Neural Networks for Resource-Optimized Learning
por: Zhang, Le, et al.
Publicado: (2024)
por: Zhang, Le, et al.
Publicado: (2024)
Hierarchical Recursive Precision for Accelerating Symmetric Linear Solves on MXUs
por: Carrica, Vicki, et al.
Publicado: (2026)
por: Carrica, Vicki, et al.
Publicado: (2026)
Evaluating Fault Tolerance and Scalability in Distributed File Systems: A Case Study of GFS, HDFS, and MinIO
por: Malhotra, Shubham, et al.
Publicado: (2025)
por: Malhotra, Shubham, et al.
Publicado: (2025)
The Sunk Carbon Fallacy: Rethinking Carbon Footprint Metrics for Effective Carbon-Aware Scheduling
por: Bashir, Noman, et al.
Publicado: (2024)
por: Bashir, Noman, et al.
Publicado: (2024)
EWSJF: An Adaptive Scheduler with Hybrid Partitioning for Mixed-Workload LLM Inference
por: Sidik, Bronislav, et al.
Publicado: (2026)
por: Sidik, Bronislav, et al.
Publicado: (2026)
ACME: Adaptive Customization of Large Models via Distributed Systems
por: Dai, Ziming, et al.
Publicado: (2025)
por: Dai, Ziming, et al.
Publicado: (2025)
An Empirical Study of the Impact of Federated Learning on Machine Learning Model Accuracy
por: Yang, Haotian, et al.
Publicado: (2025)
por: Yang, Haotian, et al.
Publicado: (2025)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
por: Kolluru, Saicharan
Publicado: (2025)
por: Kolluru, Saicharan
Publicado: (2025)
Efficient and Scalable Architecture for Multiple-chip Implementation of Simulated Bifurcation Machines
por: Kashimata, Tomoya, et al.
Publicado: (2023)
por: Kashimata, Tomoya, et al.
Publicado: (2023)
Parameter-Efficient and Personalized Federated Training of Generative Models at the Edge
por: Khan, Kabir, et al.
Publicado: (2025)
por: Khan, Kabir, et al.
Publicado: (2025)
TAGC: Optimizing Gradient Communication in Distributed Transformer Training
por: Polyakov, Igor, et al.
Publicado: (2025)
por: Polyakov, Igor, et al.
Publicado: (2025)
Is RISC-V Ready for Machine Learning? Portable Gaussian Processes Using Asynchronous Tasks
por: Strack, Alexander, et al.
Publicado: (2026)
por: Strack, Alexander, et al.
Publicado: (2026)
Libra: Unleashing GPU Heterogeneity for High-Performance Sparse Matrix Multiplication
por: Shi, Jinliang, et al.
Publicado: (2025)
por: Shi, Jinliang, et al.
Publicado: (2025)
Vectorized Adaptive Histograms for Sparse Oblique Forests
por: Lubonja, Ariel, et al.
Publicado: (2026)
por: Lubonja, Ariel, et al.
Publicado: (2026)
SLO-Guard: Crash-Aware, Budget-Consistent Autotuning for SLO-Constrained LLM Serving
por: Lysenstøen, Christian
Publicado: (2026)
por: Lysenstøen, Christian
Publicado: (2026)
Intent-driven scheduling of backup jobs
por: Dutta, Souvik, et al.
Publicado: (2024)
por: Dutta, Souvik, et al.
Publicado: (2024)
Federated Fine-Tuning of LLMs on the Very Edge: The Good, the Bad, the Ugly
por: Woisetschläger, Herbert, et al.
Publicado: (2023)
por: Woisetschläger, Herbert, et al.
Publicado: (2023)
Feature-Aware Task-to-Core Allocation in Embedded Multi-core Platforms via Statistical Learning
por: Pivezhandi, Mohammad, et al.
Publicado: (2025)
por: Pivezhandi, Mohammad, et al.
Publicado: (2025)
Cognitive Infrastructure: A Unified DCIM Framework for AI Data Centers
por: Sunkara, Krishna Chaitanya
Publicado: (2026)
por: Sunkara, Krishna Chaitanya
Publicado: (2026)
Secure Federated XGBoost with CUDA-accelerated Homomorphic Encryption via NVIDIA FLARE
por: Xu, Ziyue, et al.
Publicado: (2025)
por: Xu, Ziyue, et al.
Publicado: (2025)
Safety-Critical Edge Robotics Architecture with Bounded End-to-End Latency
por: Gala, Gautam, et al.
Publicado: (2024)
por: Gala, Gautam, et al.
Publicado: (2024)
Prediction based computation offloading and resource allocation for multi-access ISAC enabled IoT system
por: Le, Duc-Thuan
Publicado: (2024)
por: Le, Duc-Thuan
Publicado: (2024)
A Selective Homomorphic Encryption Approach for Faster Privacy-Preserving Federated Learning
por: Korkmaz, Abdulkadir, et al.
Publicado: (2025)
por: Korkmaz, Abdulkadir, et al.
Publicado: (2025)
Ejemplares similares
-
Active Inference-Based Adaptive Routing for Heterogeneous Edge AI Services
por: Wang, Zihang, et al.
Publicado: (2026) -
The $qs$ Inequality: Quantifying the Double Penalty of Mixture-of-Experts at Inference
por: Adhinarayanan, Vignesh, et al.
Publicado: (2026) -
ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM Inference
por: Zhu, Xiongwei, et al.
Publicado: (2026) -
Seamless acceleration of Fortran intrinsics via AMD AI engines
por: Brown, Nick, et al.
Publicado: (2025) -
The Landscape of GPU-Centric Communication
por: Unat, Didem, et al.
Publicado: (2024)