Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Oh, Seungeun, Kim, Jinhyuk, Park, Jihong, Ko, Seung-Woo, Quek, Tony Q. S., Kim, Seong-Lyun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Communication-Efficient Hybrid Language Model via Uncertainty-Aware Opportunistic and Compressed Transmission
by: Oh, Seungeun, et al.
Published: (2025)
by: Oh, Seungeun, et al.
Published: (2025)
Breaking the Capacity Bottleneck in Model-Heterogeneous Federated Learning via Gradual Model Restoration
by: Ma, Chengjie, et al.
Published: (2025)
by: Ma, Chengjie, et al.
Published: (2025)
Action Deviation-Aware Inference for Low-Latency Wireless Robots
by: Park, Jeyoung, et al.
Published: (2025)
by: Park, Jeyoung, et al.
Published: (2025)
Federated Inference for Heterogeneous LLM Communication and Collaboration
by: Chen, Zihan, et al.
Published: (2026)
by: Chen, Zihan, et al.
Published: (2026)
PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM Devices
by: Noh, Si Ung, et al.
Published: (2024)
by: Noh, Si Ung, et al.
Published: (2024)
A Survey on Resource Management in Joint Communication and Computing-Embedded SAGIN
by: Chen, Qian, et al.
Published: (2024)
by: Chen, Qian, et al.
Published: (2024)
Legible Consensus: Topology-Aware Quorum Geometry for Asymmetric Networks
by: Mason, Tony
Published: (2026)
by: Mason, Tony
Published: (2026)
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
by: Chung, Euijun, et al.
Published: (2026)
by: Chung, Euijun, et al.
Published: (2026)
Privacy-Preserving Split Learning with Vision Transformers using Patch-Wise Random and Noisy CutMix
by: Oh, Seungeun, et al.
Published: (2024)
by: Oh, Seungeun, et al.
Published: (2024)
DUET: Disaggregated Hybrid Mamba-Transformer LLMs with Prefill and Decode-Specific Packages
by: Kanani, Alish, et al.
Published: (2026)
by: Kanani, Alish, et al.
Published: (2026)
IOMMU Support for Virtual-Address Remote DMA in an ARMv8 environment
by: Psistakis, Antonis
Published: (2025)
by: Psistakis, Antonis
Published: (2025)
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
by: Kim, Hyeseong, et al.
Published: (2026)
by: Kim, Hyeseong, et al.
Published: (2026)
Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
by: Du, Chengze, et al.
Published: (2025)
by: Du, Chengze, et al.
Published: (2025)
Joint wireless and computing resource management with optimal slice selection in in-network-edge metaverse system
by: Rashid, Sulaiman Muhammad, et al.
Published: (2024)
by: Rashid, Sulaiman Muhammad, et al.
Published: (2024)
FengHuang: Next-Generation Memory Orchestration for AI Inferencing
by: Li, Jiamin, et al.
Published: (2025)
by: Li, Jiamin, et al.
Published: (2025)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
by: Meng, William, et al.
Published: (2025)
by: Meng, William, et al.
Published: (2025)
A Hybrid Approach to Monitor Context Parameters for Optimising Caching for Context-Aware IoT Applications
by: Manchanda, Ashish, et al.
Published: (2024)
by: Manchanda, Ashish, et al.
Published: (2024)
Automated Deep Neural Network Inference Partitioning for Distributed Embedded Systems
by: Kreß, Fabian, et al.
Published: (2024)
by: Kreß, Fabian, et al.
Published: (2024)
Contextual Chain: Single-State Ledger Design for Mobile/IoT Networks with Frequent Partitions
by: Kim, Song-Ju
Published: (2026)
by: Kim, Song-Ju
Published: (2026)
Datapath Combinational Equivalence Checking With Hybrid Sweeping Engines and Parallelization
by: Chen, Zhihan, et al.
Published: (2024)
by: Chen, Zhihan, et al.
Published: (2024)
Optimizing Task Scheduling in Fog Computing with Deadline Awareness
by: Sirjani, Mohammad Sadegh, et al.
Published: (2025)
by: Sirjani, Mohammad Sadegh, et al.
Published: (2025)
Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with Voxel
by: Liu, Yiqi, et al.
Published: (2026)
by: Liu, Yiqi, et al.
Published: (2026)
RevaMp3D: Architecting the Processor Core and Cache Hierarchy for Systems with Monolithically-Integrated Logic and Memory
by: Ghiasi, Nika Mansouri, et al.
Published: (2022)
by: Ghiasi, Nika Mansouri, et al.
Published: (2022)
iHAC: A Hybrid Cluster Architecture for Enhanced Performance and Resilience
by: Muntaka, Siddique Abubakr, et al.
Published: (2026)
by: Muntaka, Siddique Abubakr, et al.
Published: (2026)
HgPCN: A Heterogeneous Architecture for E2E Embedded Point Cloud Inference
by: Gao, Yiming, et al.
Published: (2025)
by: Gao, Yiming, et al.
Published: (2025)
RapidOMS: FPGA-based Open Modification Spectral Library Searching with HD Computing
by: Pinge, Sumukh, et al.
Published: (2024)
by: Pinge, Sumukh, et al.
Published: (2024)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
by: Xu, Weihong, et al.
Published: (2025)
by: Xu, Weihong, et al.
Published: (2025)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
by: Adnan, Muhammad, et al.
Published: (2024)
by: Adnan, Muhammad, et al.
Published: (2024)
CELLO: Co-designing Schedule and Hybrid Implicit/Explicit Buffer for Complex Tensor Reuse
by: Garg, Raveesh, et al.
Published: (2023)
by: Garg, Raveesh, et al.
Published: (2023)
Enabling Time-Aware Priority Traffic Management over Distributed FPGA Nodes
by: Scionti, Alberto, et al.
Published: (2025)
by: Scionti, Alberto, et al.
Published: (2025)
Topology-Aware Virtualization over Inter-Core Connected Neural Processing Units
by: Feng, Dahu, et al.
Published: (2025)
by: Feng, Dahu, et al.
Published: (2025)
cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node Communications
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
ELMoE-3D: Leveraging Intrinsic Elasticity of MoE for Hybrid-Bonding-Enabled Self-Speculative Decoding in On-Premises Serving
by: Choi, Yuseon, et al.
Published: (2026)
by: Choi, Yuseon, et al.
Published: (2026)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
by: Zhang, Chen, et al.
Published: (2026)
by: Zhang, Chen, et al.
Published: (2026)
Cache Your Prompt When It's Green: Carbon-Aware Caching for Large Language Model Serving
by: Tian, Yuyang, et al.
Published: (2025)
by: Tian, Yuyang, et al.
Published: (2025)
Sequence-Aware Split Heuristic to Mitigate SM Underutilization in FlashAttention-3 Low-Head-Count Decoding
by: Font, Martí Llopart, et al.
Published: (2026)
by: Font, Martí Llopart, et al.
Published: (2026)
To Stream or Not to Stream: Towards A Quantitative Model for Remote HPC Processing Decisions
by: Castro, Flavio, et al.
Published: (2025)
by: Castro, Flavio, et al.
Published: (2025)
SCENIC: Stream Computation-Enhanced SmartNIC
by: Ramhorst, Benjamin, et al.
Published: (2026)
by: Ramhorst, Benjamin, et al.
Published: (2026)
The Role of Federated Learning in a Wireless World with Foundation Models
by: Chen, Zihan, et al.
Published: (2023)
by: Chen, Zihan, et al.
Published: (2023)
Strategic Server Deployment under Uncertainty in Mobile Edge Computing
by: Tran, Duc A., et al.
Published: (2025)
by: Tran, Duc A., et al.
Published: (2025)
Similar Items
-
Communication-Efficient Hybrid Language Model via Uncertainty-Aware Opportunistic and Compressed Transmission
by: Oh, Seungeun, et al.
Published: (2025) -
Breaking the Capacity Bottleneck in Model-Heterogeneous Federated Learning via Gradual Model Restoration
by: Ma, Chengjie, et al.
Published: (2025) -
Action Deviation-Aware Inference for Low-Latency Wireless Robots
by: Park, Jeyoung, et al.
Published: (2025) -
Federated Inference for Heterogeneous LLM Communication and Collaboration
by: Chen, Zihan, et al.
Published: (2026) -
PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM Devices
by: Noh, Si Ung, et al.
Published: (2024)