Saved in:
| Main Authors: | Kalakanti, Arun Kumar, Rao, Shrisha |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2408.14169 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cooperative Solutions to Exploration Tasks Under Speed and Budget Constraints
by: Karishma, et al.
Published: (2022)
by: Karishma, et al.
Published: (2022)
Speeding up Local Optimization in Vehicle Routing with Tensor-based GPU Acceleration
by: Lei, Zhenyu, et al.
Published: (2025)
by: Lei, Zhenyu, et al.
Published: (2025)
FedRAV: Hierarchically Federated Region-Learning for Traffic Object Classification of Autonomous Vehicles
by: Zhai, Yijun, et al.
Published: (2024)
by: Zhai, Yijun, et al.
Published: (2024)
FedPAW: Federated Learning with Personalized Aggregation Weights for Urban Vehicle Speed Prediction
by: He, Yuepeng, et al.
Published: (2024)
by: He, Yuepeng, et al.
Published: (2024)
Electricity Cost Minimization for Multi-Workflow Allocation in Geo-Distributed Data Centers
by: Wang, Shuang, et al.
Published: (2025)
by: Wang, Shuang, et al.
Published: (2025)
AI Inference as Relocatable Electricity Demand: A Latency-Constrained Energy-Geography Framework
by: Luo, Xubin, et al.
Published: (2026)
by: Luo, Xubin, et al.
Published: (2026)
Striking the Right Balance between Compute and Copy: Improving LLM Inferencing Under Speculative Decoding
by: Ramachandran, Arun, et al.
Published: (2025)
by: Ramachandran, Arun, et al.
Published: (2025)
AI-Driven Cloud Resource Optimization for Multi-Cluster Environments
by: Punniyamoorthy, Vinoth, et al.
Published: (2025)
by: Punniyamoorthy, Vinoth, et al.
Published: (2025)
Distributed Inference on Mobile Edge and Cloud: A Data-Cartography based Clustering Approach
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
Dynamic Scheduling Strategies for Resource Optimization in Computing Environments
by: Wang, Xiaoye
Published: (2024)
by: Wang, Xiaoye
Published: (2024)
Rethinking Dynamic Networks and Heterogeneous Computing with Automatic Parallelization
by: Wu, Ruilong, et al.
Published: (2025)
by: Wu, Ruilong, et al.
Published: (2025)
Balanced and Elastic End-to-end Training of Dynamic LLMs
by: Wahib, Mohamed, et al.
Published: (2025)
by: Wahib, Mohamed, et al.
Published: (2025)
Towards an Introspective Dynamic Model of Globally Distributed Computing Infrastructures
by: Kilic, Ozgur O., et al.
Published: (2025)
by: Kilic, Ozgur O., et al.
Published: (2025)
Janus: Collaborative Vision Transformer Under Dynamic Network Environment
by: Jiang, Linyi, et al.
Published: (2025)
by: Jiang, Linyi, et al.
Published: (2025)
Cooperative Cognitive Dynamic System in UAV Swarms: Reconfigurable Mechanism and Framework
by: Jia, Ziye, et al.
Published: (2024)
by: Jia, Ziye, et al.
Published: (2024)
DeF-DReL: Systematic Deployment of Serverless Functions in Fog and Cloud environments using Deep Reinforcement Learning
by: Dehury, Chinmaya Kumar, et al.
Published: (2021)
by: Dehury, Chinmaya Kumar, et al.
Published: (2021)
Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
DIP: Efficient Large Multimodal Model Training with Dynamic Interleaved Pipeline
by: Xue, Zhenliang, et al.
Published: (2025)
by: Xue, Zhenliang, et al.
Published: (2025)
Dynamic Resource Allocation for Virtual Machine Migration Optimization using Machine Learning
by: Gong, Yulu, et al.
Published: (2024)
by: Gong, Yulu, et al.
Published: (2024)
Accurate GPU Memory Prediction for Deep Learning Jobs through Dynamic Analysis
by: Shi, Jiabo, et al.
Published: (2025)
by: Shi, Jiabo, et al.
Published: (2025)
Reducing Fragmentation and Starvation in GPU Clusters through Dynamic Multi-Objective Scheduling
by: Mamirov, Akhmadillo
Published: (2025)
by: Mamirov, Akhmadillo
Published: (2025)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
by: Kumar, Satyam, et al.
Published: (2026)
by: Kumar, Satyam, et al.
Published: (2026)
FedDCT: A Dynamic Cross-Tier Federated Learning Framework in Wireless Networks
by: Xian, Youquan, et al.
Published: (2023)
by: Xian, Youquan, et al.
Published: (2023)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
by: Zheng, Wanyi, et al.
Published: (2025)
by: Zheng, Wanyi, et al.
Published: (2025)
Dynamic Scheduling for Vehicle-to-Vehicle Communications Enhanced Federated Learning
by: Yan, Jintao, et al.
Published: (2024)
by: Yan, Jintao, et al.
Published: (2024)
Autonomous Systems Dependability in the era of AI: Design Challenges in Safety, Security, Reliability and Certification
by: Ranjbar, Behnaz, et al.
Published: (2026)
by: Ranjbar, Behnaz, et al.
Published: (2026)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
by: Wang, Haodong, et al.
Published: (2025)
by: Wang, Haodong, et al.
Published: (2025)
Scalable Cloud-Native Architectures for Intelligent PMU Data Processing
by: Chockalingam, Nachiappan, et al.
Published: (2025)
by: Chockalingam, Nachiappan, et al.
Published: (2025)
Scaling Multi Agent Reinforcement Learning for Underwater Acoustic Tracking via Autonomous Vehicles
by: Gallici, Matteo, et al.
Published: (2025)
by: Gallici, Matteo, et al.
Published: (2025)
The infrastructure powering IBM's Gen AI model development
by: Gershon, Talia, et al.
Published: (2024)
by: Gershon, Talia, et al.
Published: (2024)
GEM: GPU-Variability-Aware Expert to GPU Mapping for MoE Systems
by: Wawdhane, Sourish, et al.
Published: (2026)
by: Wawdhane, Sourish, et al.
Published: (2026)
FreeRide: Harvesting Bubbles in Pipeline Parallelism
by: Zhang, Jiashu, et al.
Published: (2024)
by: Zhang, Jiashu, et al.
Published: (2024)
KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
by: Cheng, Rongxin, et al.
Published: (2024)
by: Cheng, Rongxin, et al.
Published: (2024)
Ensemble Method for System Failure Detection Using Large-Scale Telemetry Data
by: Mudgal, Priyanka, et al.
Published: (2024)
by: Mudgal, Priyanka, et al.
Published: (2024)
Topology-aware Preemptive Scheduling for Co-located LLM Workloads
by: Zhang, Ping, et al.
Published: (2024)
by: Zhang, Ping, et al.
Published: (2024)
Can Large Language Models Write Parallel Code?
by: Nichols, Daniel, et al.
Published: (2024)
by: Nichols, Daniel, et al.
Published: (2024)
LLM as HPC Expert: Extending RAG Architecture for HPC Data
by: Miyashita, Yusuke, et al.
Published: (2024)
by: Miyashita, Yusuke, et al.
Published: (2024)
Boosting Asynchronous Decentralized Learning with Model Fragmentation
by: Biswas, Sayan, et al.
Published: (2024)
by: Biswas, Sayan, et al.
Published: (2024)
FedFT: Improving Communication Performance for Federated Learning with Frequency Space Transformation
by: Palihawadana, Chamath, et al.
Published: (2024)
by: Palihawadana, Chamath, et al.
Published: (2024)
Isambard-AI: a leadership class supercomputer optimised specifically for Artificial Intelligence
by: McIntosh-Smith, Simon, et al.
Published: (2024)
by: McIntosh-Smith, Simon, et al.
Published: (2024)
Similar Items
-
Cooperative Solutions to Exploration Tasks Under Speed and Budget Constraints
by: Karishma, et al.
Published: (2022) -
Speeding up Local Optimization in Vehicle Routing with Tensor-based GPU Acceleration
by: Lei, Zhenyu, et al.
Published: (2025) -
FedRAV: Hierarchically Federated Region-Learning for Traffic Object Classification of Autonomous Vehicles
by: Zhai, Yijun, et al.
Published: (2024) -
FedPAW: Federated Learning with Personalized Aggregation Weights for Urban Vehicle Speed Prediction
by: He, Yuepeng, et al.
Published: (2024) -
Electricity Cost Minimization for Multi-Workflow Allocation in Geo-Distributed Data Centers
by: Wang, Shuang, et al.
Published: (2025)