Efficient and Scalable Agentic AI with Heterogeneous Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Asgar, Zain, Nguyen, Michelle, Katti, Sachin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HiDP: Hierarchical DNN Partitioning for Distributed Inference on Heterogeneous Edge Platforms
von: Taufique, Zain, et al.
Veröffentlicht: (2024)
von: Taufique, Zain, et al.
Veröffentlicht: (2024)
Adaptive and Resource-efficient Agentic AI Systems for Mobile and Embedded Devices: A Survey
von: Liu, Sicong, et al.
Veröffentlicht: (2025)
von: Liu, Sicong, et al.
Veröffentlicht: (2025)
Efficient Federated Learning with Heterogeneous Data and Adaptive Dropout
von: Liu, Ji, et al.
Veröffentlicht: (2025)
von: Liu, Ji, et al.
Veröffentlicht: (2025)
Mind the Gap: Revealing Inconsistencies Across Heterogeneous AI Accelerators
von: Wen, Elliott, et al.
Veröffentlicht: (2025)
von: Wen, Elliott, et al.
Veröffentlicht: (2025)
Hubs and Spokes Learning: Efficient and Scalable Collaborative Machine Learning
von: Sharma, Atul, et al.
Veröffentlicht: (2025)
von: Sharma, Atul, et al.
Veröffentlicht: (2025)
LCFed: An Efficient Clustered Federated Learning Framework for Heterogeneous Data
von: Zhang, Yuxin, et al.
Veröffentlicht: (2025)
von: Zhang, Yuxin, et al.
Veröffentlicht: (2025)
Towards Robust and Efficient Federated Low-Rank Adaptation with Heterogeneous Clients
von: Koo, Jabin, et al.
Veröffentlicht: (2024)
von: Koo, Jabin, et al.
Veröffentlicht: (2024)
Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers
von: Singh, Siddharth, et al.
Veröffentlicht: (2025)
von: Singh, Siddharth, et al.
Veröffentlicht: (2025)
Effective Heterogeneous Federated Learning via Efficient Hypernetwork-based Weight Generation
von: Shin, Yujin, et al.
Veröffentlicht: (2024)
von: Shin, Yujin, et al.
Veröffentlicht: (2024)
Cost-Efficient Multimodal LLM Inference via Cross-Tier GPU Heterogeneity
von: Yu, Donglin
Veröffentlicht: (2026)
von: Yu, Donglin
Veröffentlicht: (2026)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
SatFed: A Resource-Efficient LEO Satellite-Assisted Heterogeneous Federated Learning Framework
von: Zhang, Yuxin, et al.
Veröffentlicht: (2024)
von: Zhang, Yuxin, et al.
Veröffentlicht: (2024)
HASFL: Heterogeneity-aware Split Federated Learning over Edge Computing Systems
von: Lin, Zheng, et al.
Veröffentlicht: (2025)
von: Lin, Zheng, et al.
Veröffentlicht: (2025)
Loop Improvement: An Efficient Approach for Extracting Shared Features from Heterogeneous Data without Central Server
von: Li, Fei, et al.
Veröffentlicht: (2024)
von: Li, Fei, et al.
Veröffentlicht: (2024)
HSplitLoRA: A Heterogeneous Split Parameter-Efficient Fine-Tuning Framework for Large Language Models
von: Lin, Zheng, et al.
Veröffentlicht: (2025)
von: Lin, Zheng, et al.
Veröffentlicht: (2025)
Leyline: KV Cache Directives for Agentic Inference
von: Ma, Bole, et al.
Veröffentlicht: (2026)
von: Ma, Bole, et al.
Veröffentlicht: (2026)
PubSub-VFL: Towards Efficient Two-Party Split Learning in Heterogeneous Environments via Publisher/Subscriber Architecture
von: Liu, Yi, et al.
Veröffentlicht: (2025)
von: Liu, Yi, et al.
Veröffentlicht: (2025)
ARL-Tangram: Unleash the Resource Efficiency in Agentic Reinforcement Learning
von: Xiao, Bangjun, et al.
Veröffentlicht: (2026)
von: Xiao, Bangjun, et al.
Veröffentlicht: (2026)
Scalable Explainability-as-a-Service (XaaS) for Edge AI Systems
von: Singh, Samaresh Kumar, et al.
Veröffentlicht: (2026)
von: Singh, Samaresh Kumar, et al.
Veröffentlicht: (2026)
RollArt: Scaling Agentic RL Training via Disaggregated Infrastructure
von: Gao, Wei, et al.
Veröffentlicht: (2025)
von: Gao, Wei, et al.
Veröffentlicht: (2025)
Irminsul: MLA-Native Position-Independent Caching for Agentic LLM Serving
von: Ma, Bole, et al.
Veröffentlicht: (2026)
von: Ma, Bole, et al.
Veröffentlicht: (2026)
Context Parallelism for Scalable Million-Token Inference
von: Yang, Amy, et al.
Veröffentlicht: (2024)
von: Yang, Amy, et al.
Veröffentlicht: (2024)
Galvatron: An Automatic Distributed System for Efficient Foundation Model Training
von: Liu, Xinyi, et al.
Veröffentlicht: (2025)
von: Liu, Xinyi, et al.
Veröffentlicht: (2025)
Ravnest: Decentralized Asynchronous Training on Heterogeneous Devices
von: Menon, Anirudh Rajiv, et al.
Veröffentlicht: (2024)
von: Menon, Anirudh Rajiv, et al.
Veröffentlicht: (2024)
SwapNet: Efficient Swapping for DNN Inference on Edge AI Devices Beyond the Memory Budget
von: Wang, Kun, et al.
Veröffentlicht: (2024)
von: Wang, Kun, et al.
Veröffentlicht: (2024)
ELIS: Efficient LLM Iterative Scheduling System with Response Length Predictor
von: Choi, Seungbeom, et al.
Veröffentlicht: (2025)
von: Choi, Seungbeom, et al.
Veröffentlicht: (2025)
Scalable Artificial Intelligence for Science: Perspectives, Methods and Exemplars
von: Brewer, Wesley, et al.
Veröffentlicht: (2024)
von: Brewer, Wesley, et al.
Veröffentlicht: (2024)
Measuring Heterogeneity in Machine Learning with Distributed Energy Distance
von: Fan, Mengchen, et al.
Veröffentlicht: (2025)
von: Fan, Mengchen, et al.
Veröffentlicht: (2025)
Beyond Aggregation: Guiding Clients in Heterogeneous Federated Learning
von: Wang, Zijian, et al.
Veröffentlicht: (2025)
von: Wang, Zijian, et al.
Veröffentlicht: (2025)
FedUV: Uniformity and Variance for Heterogeneous Federated Learning
von: Son, Ha Min, et al.
Veröffentlicht: (2024)
von: Son, Ha Min, et al.
Veröffentlicht: (2024)
PipeOffload: Improving Scalability of Pipeline Parallelism with Memory Optimization
von: Wan, Xinyi, et al.
Veröffentlicht: (2025)
von: Wan, Xinyi, et al.
Veröffentlicht: (2025)
Laminar: A Scalable Asynchronous RL Post-Training Framework
von: Sheng, Guangming, et al.
Veröffentlicht: (2025)
von: Sheng, Guangming, et al.
Veröffentlicht: (2025)
The Big Send-off: Scalable and Performant Collectives for Deep Learning
von: Singh, Siddharth, et al.
Veröffentlicht: (2025)
von: Singh, Siddharth, et al.
Veröffentlicht: (2025)
DistShap: Scalable GNN Explanations with Distributed Shapley Values
von: Akkas, Selahattin, et al.
Veröffentlicht: (2025)
von: Akkas, Selahattin, et al.
Veröffentlicht: (2025)
Scalable and Adaptive Parallel Training of Graph Transformer on Large Graphs
von: Lin, Jun-Liang, et al.
Veröffentlicht: (2026)
von: Lin, Jun-Liang, et al.
Veröffentlicht: (2026)
Characterizing Mobile SoC for Accelerating Heterogeneous LLM Inference
von: Chen, Le, et al.
Veröffentlicht: (2025)
von: Chen, Le, et al.
Veröffentlicht: (2025)
Adaptive Active Inference Agents for Heterogeneous and Lifelong Federated Learning
von: Danilenka, Anastasiya, et al.
Veröffentlicht: (2024)
von: Danilenka, Anastasiya, et al.
Veröffentlicht: (2024)
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
von: Shen, Zheyu, et al.
Veröffentlicht: (2025)
von: Shen, Zheyu, et al.
Veröffentlicht: (2025)
DeepHYDRA: Resource-Efficient Time-Series Anomaly Detection in Dynamically-Configured Systems
von: Stehle, Franz Kevin, et al.
Veröffentlicht: (2024)
von: Stehle, Franz Kevin, et al.
Veröffentlicht: (2024)
FSD-Inference: Fully Serverless Distributed Inference with Scalable Cloud Communication
von: Oakley, Joe, et al.
Veröffentlicht: (2024)
von: Oakley, Joe, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HiDP: Hierarchical DNN Partitioning for Distributed Inference on Heterogeneous Edge Platforms
von: Taufique, Zain, et al.
Veröffentlicht: (2024) -
Adaptive and Resource-efficient Agentic AI Systems for Mobile and Embedded Devices: A Survey
von: Liu, Sicong, et al.
Veröffentlicht: (2025) -
Efficient Federated Learning with Heterogeneous Data and Adaptive Dropout
von: Liu, Ji, et al.
Veröffentlicht: (2025) -
Mind the Gap: Revealing Inconsistencies Across Heterogeneous AI Accelerators
von: Wen, Elliott, et al.
Veröffentlicht: (2025) -
Hubs and Spokes Learning: Efficient and Scalable Collaborative Machine Learning
von: Sharma, Atul, et al.
Veröffentlicht: (2025)