Gespeichert in:
| Hauptverfasser: | Jiang, Linyi, Fu, Silvery D., Zhu, Yifei, Li, Bo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2502.10047 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hyperion: Low-Latency Ultra-HD Video Analytics via Collaborative Vision Transformer Inference
von: Jiang, Linyi, et al.
Veröffentlicht: (2025)
von: Jiang, Linyi, et al.
Veröffentlicht: (2025)
Dynamic Scheduling Strategies for Resource Optimization in Computing Environments
von: Wang, Xiaoye
Veröffentlicht: (2024)
von: Wang, Xiaoye
Veröffentlicht: (2024)
Rethinking Dynamic Networks and Heterogeneous Computing with Automatic Parallelization
von: Wu, Ruilong, et al.
Veröffentlicht: (2025)
von: Wu, Ruilong, et al.
Veröffentlicht: (2025)
DIP: Efficient Large Multimodal Model Training with Dynamic Interleaved Pipeline
von: Xue, Zhenliang, et al.
Veröffentlicht: (2025)
von: Xue, Zhenliang, et al.
Veröffentlicht: (2025)
Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
Striking the Right Balance between Compute and Copy: Improving LLM Inferencing Under Speculative Decoding
von: Ramachandran, Arun, et al.
Veröffentlicht: (2025)
von: Ramachandran, Arun, et al.
Veröffentlicht: (2025)
FedDCT: A Dynamic Cross-Tier Federated Learning Framework in Wireless Networks
von: Xian, Youquan, et al.
Veröffentlicht: (2023)
von: Xian, Youquan, et al.
Veröffentlicht: (2023)
A Resource-Adaptive Approach for Federated Learning under Resource-Constrained Environments
von: Zhang, Ruirui, et al.
Veröffentlicht: (2024)
von: Zhang, Ruirui, et al.
Veröffentlicht: (2024)
Dynamic Resource Allocation for Virtual Machine Migration Optimization using Machine Learning
von: Gong, Yulu, et al.
Veröffentlicht: (2024)
von: Gong, Yulu, et al.
Veröffentlicht: (2024)
Collaborative Split Federated Learning with Parallel Training and Aggregation
von: Papageorgiou, Yiannis, et al.
Veröffentlicht: (2025)
von: Papageorgiou, Yiannis, et al.
Veröffentlicht: (2025)
SFPrompt: Communication-Efficient Split Federated Fine-Tuning for Large Pre-Trained Models over Resource-Limited Devices
von: Cao, Linxiao, et al.
Veröffentlicht: (2024)
von: Cao, Linxiao, et al.
Veröffentlicht: (2024)
PackKV: Reducing KV Cache Memory Footprint through LLM-Aware Lossy Compression
von: Jiang, Bo, et al.
Veröffentlicht: (2025)
von: Jiang, Bo, et al.
Veröffentlicht: (2025)
KVComp: A High-Performance, LLM-Aware, Lossy Compression Framework for KV Cache
von: Jiang, Bo, et al.
Veröffentlicht: (2025)
von: Jiang, Bo, et al.
Veröffentlicht: (2025)
PolyKAN: Efficient Fused GPU Operators for Polynomial Kolmogorov-Arnold Network Variants
von: Yu, Mingkun, et al.
Veröffentlicht: (2025)
von: Yu, Mingkun, et al.
Veröffentlicht: (2025)
AI-Driven Cloud Resource Optimization for Multi-Cluster Environments
von: Punniyamoorthy, Vinoth, et al.
Veröffentlicht: (2025)
von: Punniyamoorthy, Vinoth, et al.
Veröffentlicht: (2025)
ECCENTRIC: Edge-Cloud Collaboration Framework for Distributed Inference Using Knowledge Adaptation
von: Kamani, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
von: Kamani, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
Cooperative Cognitive Dynamic System in UAV Swarms: Reconfigurable Mechanism and Framework
von: Jia, Ziye, et al.
Veröffentlicht: (2024)
von: Jia, Ziye, et al.
Veröffentlicht: (2024)
xDiT: an Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism
von: Fang, Jiarui, et al.
Veröffentlicht: (2024)
von: Fang, Jiarui, et al.
Veröffentlicht: (2024)
Transforming Future Data Center Operations and Management via Physical AI
von: Cao, Zhiwei, et al.
Veröffentlicht: (2025)
von: Cao, Zhiwei, et al.
Veröffentlicht: (2025)
Hardware Utilization and Inference Performance of Edge Object Detection Under Fault Injection
von: Pasandideh, Faezeh, et al.
Veröffentlicht: (2026)
von: Pasandideh, Faezeh, et al.
Veröffentlicht: (2026)
Deploying Graph Neural Networks in Wireless Networks: A Link Stability Viewpoint
von: Li, Jun, et al.
Veröffentlicht: (2024)
von: Li, Jun, et al.
Veröffentlicht: (2024)
Adaptive Fault Tolerance Mechanisms of Large Language Models in Cloud Computing Environments
von: Jin, Yihong, et al.
Veröffentlicht: (2025)
von: Jin, Yihong, et al.
Veröffentlicht: (2025)
Astra: Efficient and Money-saving Automatic Parallel Strategies Search on Heterogeneous GPUs
von: Wang, Peiran, et al.
Veröffentlicht: (2025)
von: Wang, Peiran, et al.
Veröffentlicht: (2025)
Towards Real-Time Neural Volumetric Rendering on Mobile Devices: A Measurement Study
von: Wang, Zhe, et al.
Veröffentlicht: (2024)
von: Wang, Zhe, et al.
Veröffentlicht: (2024)
Benchmarking Federated Learning in Edge Computing Environments: A Systematic Review and Performance Evaluation
von: Aribe Jr., Sales, et al.
Veröffentlicht: (2026)
von: Aribe Jr., Sales, et al.
Veröffentlicht: (2026)
Mesh-Attention: A New Communication-Efficient Distributed Attention with Improved Data Locality
von: Chen, Sirui, et al.
Veröffentlicht: (2025)
von: Chen, Sirui, et al.
Veröffentlicht: (2025)
Seesaw: High-throughput LLM Inference via Model Re-sharding
von: Su, Qidong, et al.
Veröffentlicht: (2025)
von: Su, Qidong, et al.
Veröffentlicht: (2025)
InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training
von: Wang, Shiju, et al.
Veröffentlicht: (2025)
von: Wang, Shiju, et al.
Veröffentlicht: (2025)
Decentralized AI: Permissionless LLM Inference on POKT Network
von: Olshansky, Daniel, et al.
Veröffentlicht: (2024)
von: Olshansky, Daniel, et al.
Veröffentlicht: (2024)
Demystifying the Communication Characteristics for Distributed Transformer Models
von: Anthony, Quentin, et al.
Veröffentlicht: (2024)
von: Anthony, Quentin, et al.
Veröffentlicht: (2024)
KAITIAN: A Unified Communication Framework for Enabling Efficient Collaboration Across Heterogeneous Accelerators in Embodied AI Systems
von: Lin, Jieke, et al.
Veröffentlicht: (2025)
von: Lin, Jieke, et al.
Veröffentlicht: (2025)
Mosaic: Data-Free Knowledge Distillation via Mixture-of-Experts for Heterogeneous Distributed Environments
von: Liu, Junming, et al.
Veröffentlicht: (2025)
von: Liu, Junming, et al.
Veröffentlicht: (2025)
High-Dimensional Data Processing: Benchmarking Machine Learning and Deep Learning Architectures in Local and Distributed Environments
von: Rodriguez, Julian, et al.
Veröffentlicht: (2025)
von: Rodriguez, Julian, et al.
Veröffentlicht: (2025)
ParaGAN: A Scalable Distributed Training Framework for Generative Adversarial Networks
von: Shi, Ziji, et al.
Veröffentlicht: (2024)
von: Shi, Ziji, et al.
Veröffentlicht: (2024)
LLM Inference Serving: Survey of Recent Advances and Opportunities
von: Li, Baolin, et al.
Veröffentlicht: (2024)
von: Li, Baolin, et al.
Veröffentlicht: (2024)
Transformer-Based Model for Cold Start Mitigation in FaaS Architecture
von: Mouen, Alexandre Savi Fayam Mbala, et al.
Veröffentlicht: (2025)
von: Mouen, Alexandre Savi Fayam Mbala, et al.
Veröffentlicht: (2025)
HadaCore: Tensor Core Accelerated Hadamard Transform Kernel
von: Agarwal, Krish, et al.
Veröffentlicht: (2024)
von: Agarwal, Krish, et al.
Veröffentlicht: (2024)
MoE-Lens: Towards the Hardware Limit of High-Throughput MoE LLM Serving Under Resource Constraints
von: Yuan, Yichao, et al.
Veröffentlicht: (2025)
von: Yuan, Yichao, et al.
Veröffentlicht: (2025)
Profiling-Driven Adaptive Distributed Transformer Inference on Embedded Edge Deployment
von: Qazi, Muhammad Azlan, et al.
Veröffentlicht: (2026)
von: Qazi, Muhammad Azlan, et al.
Veröffentlicht: (2026)
FedSAC: Dynamic Submodel Allocation for Collaborative Fairness in Federated Learning
von: Wang, Zihui, et al.
Veröffentlicht: (2024)
von: Wang, Zihui, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Hyperion: Low-Latency Ultra-HD Video Analytics via Collaborative Vision Transformer Inference
von: Jiang, Linyi, et al.
Veröffentlicht: (2025) -
Dynamic Scheduling Strategies for Resource Optimization in Computing Environments
von: Wang, Xiaoye
Veröffentlicht: (2024) -
Rethinking Dynamic Networks and Heterogeneous Computing with Automatic Parallelization
von: Wu, Ruilong, et al.
Veröffentlicht: (2025) -
DIP: Efficient Large Multimodal Model Training with Dynamic Interleaved Pipeline
von: Xue, Zhenliang, et al.
Veröffentlicht: (2025) -
Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving
von: Li, Rui, et al.
Veröffentlicht: (2025)