Context-Driven Performance Modeling for Causal Inference Operators on Neural Processing Units
Fuente:
arXiv
Saved in:
| Main Authors: | Gupta, Neelesh, Jayanth, Rakshith, Parikh, Dhruv, Prasanna, Viktor |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ScalableHD: Scalable and High-Throughput Hyperdimensional Computing Inference on Multi-Core CPUs
by: Parikh, Dhruv, et al.
Published: (2025)
by: Parikh, Dhruv, et al.
Published: (2025)
Benchmarking the Performance of Large Language Models on the Cerebras Wafer Scale Engine
by: Zhang, Zuoning, et al.
Published: (2024)
by: Zhang, Zuoning, et al.
Published: (2024)
Accelerating Dynamic Image Graph Construction on FPGA for Vision GNNs
by: Ramachandran, Anvitha, et al.
Published: (2025)
by: Ramachandran, Anvitha, et al.
Published: (2025)
GraphLeap: Decoupling Graph Construction and Convolution for Vision GNN Acceleration on FPGA
by: Ramachandran, Anvitha, et al.
Published: (2026)
by: Ramachandran, Anvitha, et al.
Published: (2026)
Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference
by: Deshmukh, Dhruv, et al.
Published: (2025)
by: Deshmukh, Dhruv, et al.
Published: (2025)
A Distributed Framework for Causal Modeling of Performance Variability in GPU Traces
by: Lahiry, Ankur, et al.
Published: (2025)
by: Lahiry, Ankur, et al.
Published: (2025)
FedGMI: Generative Model-Driven Federated Learning for Probabilistic Mixture Inference
by: Hou, Qijun, et al.
Published: (2026)
by: Hou, Qijun, et al.
Published: (2026)
Practical Performance Guarantees for Pipelined DNN Inference
by: Archer, Aaron, et al.
Published: (2023)
by: Archer, Aaron, et al.
Published: (2023)
No Request Left Behind: Tackling Heterogeneity in Long-Context LLM Inference with Medha
by: Agrawal, Amey, et al.
Published: (2024)
by: Agrawal, Amey, et al.
Published: (2024)
Adaptive Stream Processing on Edge Devices through Active Inference
by: Sedlak, Boris, et al.
Published: (2024)
by: Sedlak, Boris, et al.
Published: (2024)
Understanding and Improving Communication Performance in Multi-node LLM Inference
by: Singhania, Prajwal, et al.
Published: (2025)
by: Singhania, Prajwal, et al.
Published: (2025)
Accelerating ViT Inference on FPGA through Static and Dynamic Pruning
by: Parikh, Dhruv, et al.
Published: (2024)
by: Parikh, Dhruv, et al.
Published: (2024)
Scalable and Cost-Efficient ML Inference: Parallel Batch Processing with Serverless Functions
by: Barrak, Amine, et al.
Published: (2025)
by: Barrak, Amine, et al.
Published: (2025)
ClusterViG: Efficient Globally Aware Vision GNNs via Image Partitioning
by: Parikh, Dhruv, et al.
Published: (2025)
by: Parikh, Dhruv, et al.
Published: (2025)
What Operations can be Performed Directly on Compressed Arrays, and with What Error?
by: Agarwal, Tripti, et al.
Published: (2024)
by: Agarwal, Tripti, et al.
Published: (2024)
Performance Modeling and Workload Analysis of Distributed Large Language Model Training and Inference
by: Kundu, Joyjit, et al.
Published: (2024)
by: Kundu, Joyjit, et al.
Published: (2024)
GPU-Accelerated Optimization of Transformer-Based Neural Networks for Real-Time Inference
by: Mukherjee, Soutrik, et al.
Published: (2026)
by: Mukherjee, Soutrik, et al.
Published: (2026)
BLaST: High Performance Inference and Pretraining using BLock Sparse Transformers
by: Okanovic, Patrik, et al.
Published: (2025)
by: Okanovic, Patrik, et al.
Published: (2025)
Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection
by: Gupta, Vima, et al.
Published: (2024)
by: Gupta, Vima, et al.
Published: (2024)
Context-Aware Inference via Performance Forecasting in Decentralized Learning Networks
by: Pfeffer, Joel, et al.
Published: (2025)
by: Pfeffer, Joel, et al.
Published: (2025)
SuperFedNAS: Cost-Efficient Federated Neural Architecture Search for On-Device Inference
by: Khare, Alind, et al.
Published: (2023)
by: Khare, Alind, et al.
Published: (2023)
From Models to Operators: Rethinking Autoscaling Granularity for Large Generative Models
by: Cui, Xingqi, et al.
Published: (2025)
by: Cui, Xingqi, et al.
Published: (2025)
LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load
by: Tummalapalli, Pranay, et al.
Published: (2026)
by: Tummalapalli, Pranay, et al.
Published: (2026)
Fast Distributed Inference Serving for Large Language Models
by: Wu, Bingyang, et al.
Published: (2023)
by: Wu, Bingyang, et al.
Published: (2023)
Priority-Aware Model-Distributed Inference at Edge Networks
by: Li, Teng, et al.
Published: (2024)
by: Li, Teng, et al.
Published: (2024)
CascadeServe: Unlocking Model Cascades for Inference Serving
by: Kossmann, Ferdi, et al.
Published: (2024)
by: Kossmann, Ferdi, et al.
Published: (2024)
Intelligent Orchestration of Distributed Large Foundation Model Inference at the Edge
by: Koch, Fernando, et al.
Published: (2025)
by: Koch, Fernando, et al.
Published: (2025)
Collaborative Speculative Inference for Efficient LLM Inference Serving
by: Gao, Luyao, et al.
Published: (2025)
by: Gao, Luyao, et al.
Published: (2025)
Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
by: Liu, Mengfan, et al.
Published: (2025)
by: Liu, Mengfan, et al.
Published: (2025)
Designing Large Foundation Models for Efficient Training and Inference: A Survey
by: Liu, Dong, et al.
Published: (2024)
by: Liu, Dong, et al.
Published: (2024)
ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
by: Fu, Yao, et al.
Published: (2024)
by: Fu, Yao, et al.
Published: (2024)
Knowledge-Driven Federated Graph Learning on Model Heterogeneity
by: Wu, Zhengyu, et al.
Published: (2025)
by: Wu, Zhengyu, et al.
Published: (2025)
Enabling Efficient Serverless Inference Serving for LLM (Large Language Model) in the Cloud
by: Ghosh, Himel
Published: (2024)
by: Ghosh, Himel
Published: (2024)
Graph Neural Network Training Systems: A Performance Comparison of Full-Graph and Mini-Batch
by: Bajaj, Saurabh, et al.
Published: (2024)
by: Bajaj, Saurabh, et al.
Published: (2024)
MatchNAS: Optimizing Edge AI in Sparse-Label Data Contexts via Automating Deep Neural Network Porting for Mobile Deployment
by: Huang, Hongtao, et al.
Published: (2024)
by: Huang, Hongtao, et al.
Published: (2024)
Comprehensive Performance Modeling and System Design Insights for Foundation Models
by: Subramanian, Shashank, et al.
Published: (2024)
by: Subramanian, Shashank, et al.
Published: (2024)
Near-Zero-Overhead Freshness for Recommendation Systems via Inference-Side Model Updates
by: Yu, Wenjun, et al.
Published: (2025)
by: Yu, Wenjun, et al.
Published: (2025)
FedECA: Federated External Control Arms for Causal Inference with Time-To-Event Data in Distributed Settings
by: Terrail, Jean Ogier du, et al.
Published: (2023)
by: Terrail, Jean Ogier du, et al.
Published: (2023)
Salted Inference: Enhancing Privacy while Maintaining Efficiency of Split Inference in Mobile Computing
by: Malekzadeh, Mohammad, et al.
Published: (2023)
by: Malekzadeh, Mohammad, et al.
Published: (2023)
DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism
by: Jiang, Chenyu, et al.
Published: (2025)
by: Jiang, Chenyu, et al.
Published: (2025)
Similar Items
-
ScalableHD: Scalable and High-Throughput Hyperdimensional Computing Inference on Multi-Core CPUs
by: Parikh, Dhruv, et al.
Published: (2025) -
Benchmarking the Performance of Large Language Models on the Cerebras Wafer Scale Engine
by: Zhang, Zuoning, et al.
Published: (2024) -
Accelerating Dynamic Image Graph Construction on FPGA for Vision GNNs
by: Ramachandran, Anvitha, et al.
Published: (2025) -
GraphLeap: Decoupling Graph Construction and Convolution for Vision GNN Acceleration on FPGA
by: Ramachandran, Anvitha, et al.
Published: (2026) -
Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference
by: Deshmukh, Dhruv, et al.
Published: (2025)