Communication-Efficient Hybrid Language Model via Uncertainty-Aware Opportunistic and Compressed Transmission
Fuente:
arXiv
Salvato in:
| Autori principali: | Oh, Seungeun, Kim, Jinhyuk, Park, Jihong, Ko, Seung-Woo, Choi, Jinho, Quek, Tony Q. S., Kim, Seong-Lyun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models
di: Oh, Seungeun, et al.
Pubblicazione: (2024)
di: Oh, Seungeun, et al.
Pubblicazione: (2024)
Action Deviation-Aware Inference for Low-Latency Wireless Robots
di: Park, Jeyoung, et al.
Pubblicazione: (2025)
di: Park, Jeyoung, et al.
Pubblicazione: (2025)
Breaking the Capacity Bottleneck in Model-Heterogeneous Federated Learning via Gradual Model Restoration
di: Ma, Chengjie, et al.
Pubblicazione: (2025)
di: Ma, Chengjie, et al.
Pubblicazione: (2025)
Federated Inference for Heterogeneous LLM Communication and Collaboration
di: Chen, Zihan, et al.
Pubblicazione: (2026)
di: Chen, Zihan, et al.
Pubblicazione: (2026)
Privacy-Preserving Split Learning with Vision Transformers using Patch-Wise Random and Noisy CutMix
di: Oh, Seungeun, et al.
Pubblicazione: (2024)
di: Oh, Seungeun, et al.
Pubblicazione: (2024)
Integrated user scheduling and beam steering in over-the-air federated learning for mobile IoT
di: Liu, Shengheng, et al.
Pubblicazione: (2025)
di: Liu, Shengheng, et al.
Pubblicazione: (2025)
AGAThA: Fast and Efficient GPU Acceleration of Guided Sequence Alignment for Long Read Mapping
di: Park, Seongyeon, et al.
Pubblicazione: (2024)
di: Park, Seongyeon, et al.
Pubblicazione: (2024)
Hyperion: Hierarchical Scheduling for Parallel LLM Acceleration in Multi-tier Networks
di: Ma, Mulei, et al.
Pubblicazione: (2025)
di: Ma, Mulei, et al.
Pubblicazione: (2025)
Accelerating Wireless Distributed Learning via Hybrid Split and Federated Learning Optimization
di: Guo, Kun, et al.
Pubblicazione: (2025)
di: Guo, Kun, et al.
Pubblicazione: (2025)
ScalePool: Hybrid XLink-CXL Fabric for Composable Resource Disaggregation in Unified Scale-up Domains
di: Woo, Hyein, et al.
Pubblicazione: (2025)
di: Woo, Hyein, et al.
Pubblicazione: (2025)
FlexiWalker: Extensible GPU Framework for Efficient Dynamic Random Walks with Runtime Adaptation
di: Park, Seongyeon, et al.
Pubblicazione: (2025)
di: Park, Seongyeon, et al.
Pubblicazione: (2025)
Robust Federated Fine-Tuning in Heterogeneous Networks with Unreliable Connections: An Aggregation View
di: Wang, Yanmeng, et al.
Pubblicazione: (2025)
di: Wang, Yanmeng, et al.
Pubblicazione: (2025)
PathWeaver: A High-Throughput Multi-GPU System for Graph-Based Approximate Nearest Neighbor Search
di: Kim, Sukjin, et al.
Pubblicazione: (2025)
di: Kim, Sukjin, et al.
Pubblicazione: (2025)
Communication-Efficient Federated Learning by Quantized Variance Reduction for Heterogeneous Wireless Edge Networks
di: Wang, Shuai, et al.
Pubblicazione: (2025)
di: Wang, Shuai, et al.
Pubblicazione: (2025)
Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation
di: Kim, Joon Ha, et al.
Pubblicazione: (2026)
di: Kim, Joon Ha, et al.
Pubblicazione: (2026)
FLARE: A Dataflow-Aware and Scalable Hardware Architecture for Neural-Hybrid Scientific Lossy Compression
di: Jia, Wenqi, et al.
Pubblicazione: (2025)
di: Jia, Wenqi, et al.
Pubblicazione: (2025)
Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony
di: Wang, Shaoyu, et al.
Pubblicazione: (2025)
di: Wang, Shaoyu, et al.
Pubblicazione: (2025)
Characterizing Compute-Communication Overlap in GPU-Accelerated Distributed Deep Learning: Performance and Power Implications
di: Lee, Seonho, et al.
Pubblicazione: (2025)
di: Lee, Seonho, et al.
Pubblicazione: (2025)
SpotVista: Availability-Aware Recommendation System for Reliable and Cost-Efficient Multi-Node Spot Instances
di: Kim, Taeyoon, et al.
Pubblicazione: (2026)
di: Kim, Taeyoon, et al.
Pubblicazione: (2026)
Wireless Distributed Matrix-Vector Multiplication using Over-the-Air Computation and Analog Coding
di: Choi, Jinho
Pubblicazione: (2024)
di: Choi, Jinho
Pubblicazione: (2024)
PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM Devices
di: Noh, Si Ung, et al.
Pubblicazione: (2024)
di: Noh, Si Ung, et al.
Pubblicazione: (2024)
ExeGPT: Constraint-Aware Resource Scheduling for LLM Inference
di: Oh, Hyungjun, et al.
Pubblicazione: (2024)
di: Oh, Hyungjun, et al.
Pubblicazione: (2024)
FourierCompress: Layer-Aware Spectral Activation Compression for Efficient and Accurate Collaborative LLM Inference
di: Ma, Jian, et al.
Pubblicazione: (2025)
di: Ma, Jian, et al.
Pubblicazione: (2025)
Secure Communication in the Presence of an RIS-Enhanced Eavesdropper in MIMO Networks
di: Zhang, Gaoyuan, et al.
Pubblicazione: (2025)
di: Zhang, Gaoyuan, et al.
Pubblicazione: (2025)
NAVIS: Concurrent Search and Update with Low Position-Seeking Overhead in On-SSD Graph-Based Vector Search
di: Song, Jaeyong, et al.
Pubblicazione: (2026)
di: Song, Jaeyong, et al.
Pubblicazione: (2026)
Scaling Up Throughput-oriented LLM Inference Applications on Heterogeneous Opportunistic GPU Clusters with Pervasive Context Management
di: Phung, Thanh Son, et al.
Pubblicazione: (2025)
di: Phung, Thanh Son, et al.
Pubblicazione: (2025)
FedCostAware: Enabling Cost-Aware Federated Learning on the Cloud
di: Sinha, Aditya, et al.
Pubblicazione: (2025)
di: Sinha, Aditya, et al.
Pubblicazione: (2025)
LLMSched: Uncertainty-Aware Workload Scheduling for Compound LLM Applications
di: Zhu, Botao, et al.
Pubblicazione: (2025)
di: Zhu, Botao, et al.
Pubblicazione: (2025)
GriNNder: Breaking the Memory Capacity Wall in Full-Graph GNN Training with Storage Offloading
di: Song, Jaeyong, et al.
Pubblicazione: (2026)
di: Song, Jaeyong, et al.
Pubblicazione: (2026)
SageSched: Efficient LLM Scheduling Confronting Demand Uncertainty and Hybridity
di: Gan, Zhenghao, et al.
Pubblicazione: (2026)
di: Gan, Zhenghao, et al.
Pubblicazione: (2026)
Efficiently Executing High-throughput Lightweight LLM Inference Applications on Heterogeneous Opportunistic GPU Clusters with Pervasive Context Management
di: Phung, Thanh Son, et al.
Pubblicazione: (2025)
di: Phung, Thanh Son, et al.
Pubblicazione: (2025)
A Survey on Resource Management in Joint Communication and Computing-Embedded SAGIN
di: Chen, Qian, et al.
Pubblicazione: (2024)
di: Chen, Qian, et al.
Pubblicazione: (2024)
Byzantine Attacks in RIS-Enhanced Cooperative Spectrum Sensing: A Decision Fusion Perspective
di: Zhang, Gaoyuan, et al.
Pubblicazione: (2025)
di: Zhang, Gaoyuan, et al.
Pubblicazione: (2025)
TopoSZp: Lightweight Topology-Aware Error-controlled Compression for Scientific Data
di: Agarwal, Tripti, et al.
Pubblicazione: (2026)
di: Agarwal, Tripti, et al.
Pubblicazione: (2026)
GCAPS: GPU Context-Aware Preemptive Priority-based Scheduling for Real-Time Tasks
di: Wang, Yidi, et al.
Pubblicazione: (2024)
di: Wang, Yidi, et al.
Pubblicazione: (2024)
Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
di: Lee, Sanghyeon, et al.
Pubblicazione: (2025)
di: Lee, Sanghyeon, et al.
Pubblicazione: (2025)
Deadline-Aware Bandwidth Allocation for Semantic Generative Communication with Diffusion Models
di: Choi, Jinhyuk, et al.
Pubblicazione: (2025)
di: Choi, Jinhyuk, et al.
Pubblicazione: (2025)
AAPA: An Archetype-Aware Predictive Autoscaler with Uncertainty Quantification for Serverless Workloads on Kubernetes
di: Zhang, Guilin, et al.
Pubblicazione: (2025)
di: Zhang, Guilin, et al.
Pubblicazione: (2025)
FedSZ: Leveraging Error-Bounded Lossy Compression for Federated Learning Communications
di: Wilkins, Grant, et al.
Pubblicazione: (2023)
di: Wilkins, Grant, et al.
Pubblicazione: (2023)
Efficient Time-Aware Partitioning of Quantum Circuits for Distributed Quantum Computing
di: Wu, Raymond P. H., et al.
Pubblicazione: (2026)
di: Wu, Raymond P. H., et al.
Pubblicazione: (2026)
Documenti analoghi
-
Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models
di: Oh, Seungeun, et al.
Pubblicazione: (2024) -
Action Deviation-Aware Inference for Low-Latency Wireless Robots
di: Park, Jeyoung, et al.
Pubblicazione: (2025) -
Breaking the Capacity Bottleneck in Model-Heterogeneous Federated Learning via Gradual Model Restoration
di: Ma, Chengjie, et al.
Pubblicazione: (2025) -
Federated Inference for Heterogeneous LLM Communication and Collaboration
di: Chen, Zihan, et al.
Pubblicazione: (2026) -
Privacy-Preserving Split Learning with Vision Transformers using Patch-Wise Random and Noisy CutMix
di: Oh, Seungeun, et al.
Pubblicazione: (2024)