GKT: A Novel Guidance-Based Knowledge Transfer Framework For Efficient Cloud-edge Collaboration LLM Deployment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yao, Yao, Li, Zuchao, Zhao, Hai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SirLLM: Streaming Infinite Retentive LLM
von: Yao, Yao, et al.
Veröffentlicht: (2024)
von: Yao, Yao, et al.
Veröffentlicht: (2024)
Beyond Chain-of-Thought, Effective Graph-of-Thought Reasoning in Language Models
von: Yao, Yao, et al.
Veröffentlicht: (2023)
von: Yao, Yao, et al.
Veröffentlicht: (2023)
Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
von: Shi, Luohe, et al.
Veröffentlicht: (2024)
von: Shi, Luohe, et al.
Veröffentlicht: (2024)
Reference Trustable Decoding: A Training-Free Augmentation Paradigm for Large Language Models
von: Shi, Luohe, et al.
Veröffentlicht: (2024)
von: Shi, Luohe, et al.
Veröffentlicht: (2024)
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
von: Yang, Haoqi, et al.
Veröffentlicht: (2025)
von: Yang, Haoqi, et al.
Veröffentlicht: (2025)
A Coin Has Two Sides: A Novel Detector-Corrector Framework for Chinese Spelling Correction
von: Zeng, Xiangke, et al.
Veröffentlicht: (2024)
von: Zeng, Xiangke, et al.
Veröffentlicht: (2024)
Dual Reasoning: A GNN-LLM Collaborative Framework for Knowledge Graph Question Answering
von: Liu, Guangyi, et al.
Veröffentlicht: (2024)
von: Liu, Guangyi, et al.
Veröffentlicht: (2024)
How Deep is Love in LLMs' Hearts? Exploring Semantic Size in Human-like Cognition
von: Yao, Yao, et al.
Veröffentlicht: (2025)
von: Yao, Yao, et al.
Veröffentlicht: (2025)
Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios
von: Shi, Luohe, et al.
Veröffentlicht: (2025)
von: Shi, Luohe, et al.
Veröffentlicht: (2025)
DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
Venturing into Uncharted Waters: The Navigation Compass from Transformer to Mamba
von: Zou, Yuchen, et al.
Veröffentlicht: (2024)
von: Zou, Yuchen, et al.
Veröffentlicht: (2024)
Faster MoE LLM Inference for Extremely Large Models
von: Yang, Haoqi, et al.
Veröffentlicht: (2025)
von: Yang, Haoqi, et al.
Veröffentlicht: (2025)
RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding
von: Zhang, Zihong, et al.
Veröffentlicht: (2026)
von: Zhang, Zihong, et al.
Veröffentlicht: (2026)
Efficient and Effective Internal Memory Retrieval for LLM-Based Healthcare Prediction
von: Li, Mingchen, et al.
Veröffentlicht: (2026)
von: Li, Mingchen, et al.
Veröffentlicht: (2026)
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding
von: Shi, Luohe, et al.
Veröffentlicht: (2025)
von: Shi, Luohe, et al.
Veröffentlicht: (2025)
Collaborative Distillation Strategies for Parameter-Efficient Language Model Deployment
von: Meng, Xiandong, et al.
Veröffentlicht: (2025)
von: Meng, Xiandong, et al.
Veröffentlicht: (2025)
Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions
von: Ma, Xinbei, et al.
Veröffentlicht: (2024)
von: Ma, Xinbei, et al.
Veröffentlicht: (2024)
LLM-Based Insight Extraction for Contact Center Analytics and Cost-Efficient Deployment
von: Embar, Varsha, et al.
Veröffentlicht: (2025)
von: Embar, Varsha, et al.
Veröffentlicht: (2025)
Sparse is Enough in Fine-tuning Pre-trained Large Language Models
von: Song, Weixi, et al.
Veröffentlicht: (2023)
von: Song, Weixi, et al.
Veröffentlicht: (2023)
From Isolated Scoring to Collaborative Ranking: A Comparison-Native Framework for LLM-Based Paper Evaluation
von: Zheng, Pujun, et al.
Veröffentlicht: (2026)
von: Zheng, Pujun, et al.
Veröffentlicht: (2026)
Semantics-Preserved Distortion for Personal Privacy Protection in Information Management
von: Li, Jiajia, et al.
Veröffentlicht: (2022)
von: Li, Jiajia, et al.
Veröffentlicht: (2022)
Incorporating External Knowledge and Goal Guidance for LLM-based Conversational Recommender Systems
von: Li, Chuang, et al.
Veröffentlicht: (2024)
von: Li, Chuang, et al.
Veröffentlicht: (2024)
Variation is the Key: A Variation-Based Framework for LLM-Generated Text Detection
von: Li, Xuecong, et al.
Veröffentlicht: (2026)
von: Li, Xuecong, et al.
Veröffentlicht: (2026)
Streamlining Cloud-Native Application Development and Deployment with Robust Encapsulation
von: Lertpongrujikorn, Pawissanutt, et al.
Veröffentlicht: (2024)
von: Lertpongrujikorn, Pawissanutt, et al.
Veröffentlicht: (2024)
Vocabulary Customization for Efficient Domain-Specific LLM Deployment
von: Herold, Christian, et al.
Veröffentlicht: (2025)
von: Herold, Christian, et al.
Veröffentlicht: (2025)
Segment First or Comprehend First? Explore the Limit of Unsupervised Word Segmentation with Large Language Models
von: Zhang, Zihong, et al.
Veröffentlicht: (2025)
von: Zhang, Zihong, et al.
Veröffentlicht: (2025)
When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration
von: Shi, Quan, et al.
Veröffentlicht: (2025)
von: Shi, Quan, et al.
Veröffentlicht: (2025)
VHASR: A Multimodal Speech Recognition System With Vision Hotwords
von: Hu, Jiliang, et al.
Veröffentlicht: (2024)
von: Hu, Jiliang, et al.
Veröffentlicht: (2024)
From AR to Diffusion: Efficiently Adapting Large Language Models with Strictly Causal and Elastic Horizons
von: Ma, Xiangyu, et al.
Veröffentlicht: (2026)
von: Ma, Xiangyu, et al.
Veröffentlicht: (2026)
Enhancing LLM-based Hatred and Toxicity Detection with Meta-Toxic Knowledge Graph
von: Zhao, Yibo, et al.
Veröffentlicht: (2024)
von: Zhao, Yibo, et al.
Veröffentlicht: (2024)
DIY-MKG: An LLM-Based Polyglot Language Learning System
von: Tang, Kenan, et al.
Veröffentlicht: (2025)
von: Tang, Kenan, et al.
Veröffentlicht: (2025)
ToM: Leveraging Tree-oriented MapReduce for Long-Context Reasoning in Large Language Models
von: Guo, Jiani, et al.
Veröffentlicht: (2025)
von: Guo, Jiani, et al.
Veröffentlicht: (2025)
Steering LLM Thinking with Budget Guidance
von: Li, Junyan, et al.
Veröffentlicht: (2025)
von: Li, Junyan, et al.
Veröffentlicht: (2025)
Alignment Drift in Long-Term Human-LLM Interaction: A Mechanism-Oriented Framework
von: Yao, Xintong
Veröffentlicht: (2026)
von: Yao, Xintong
Veröffentlicht: (2026)
SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers
von: Tang, Zicong, et al.
Veröffentlicht: (2025)
von: Tang, Zicong, et al.
Veröffentlicht: (2025)
Mitigating Catastrophic Forgetting in Multi-domain Chinese Spelling Correction by Multi-stage Knowledge Transfer Framework
von: Xing, Peng, et al.
Veröffentlicht: (2024)
von: Xing, Peng, et al.
Veröffentlicht: (2024)
Reliable Reasoning Path: Distilling Effective Guidance for LLM Reasoning with Knowledge Graphs
von: Xiao, Yilin, et al.
Veröffentlicht: (2025)
von: Xiao, Yilin, et al.
Veröffentlicht: (2025)
An Enhanced Prompt-Based LLM Reasoning Scheme via Knowledge Graph-Integrated Collaboration
von: Li, Yihao, et al.
Veröffentlicht: (2024)
von: Li, Yihao, et al.
Veröffentlicht: (2024)
Challenges in Deploying Long-Context Transformers: A Theoretical Peak Performance Analysis
von: Fu, Yao
Veröffentlicht: (2024)
von: Fu, Yao
Veröffentlicht: (2024)
Ähnliche Einträge
-
SirLLM: Streaming Infinite Retentive LLM
von: Yao, Yao, et al.
Veröffentlicht: (2024) -
Beyond Chain-of-Thought, Effective Graph-of-Thought Reasoning in Language Models
von: Yao, Yao, et al.
Veröffentlicht: (2023) -
Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
von: Shi, Luohe, et al.
Veröffentlicht: (2024) -
Reference Trustable Decoding: A Training-Free Augmentation Paradigm for Large Language Models
von: Shi, Luohe, et al.
Veröffentlicht: (2024) -
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs
von: Zhao, Yi, et al.
Veröffentlicht: (2025)