Krul: Efficient State Restoration for Multi-turn Conversations with Dynamic Cross-layer KV Sharing
Fuente:
arXiv
Salvato in:
| Autori principali: | Wen, Junyi, Liang, Junyuan, Hong, Zicong, Chen, Wuhui, Cai, Ting, Zheng, Zibin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline
di: Fang, Zhiyuan, et al.
Pubblicazione: (2025)
di: Fang, Zhiyuan, et al.
Pubblicazione: (2025)
Fate: Fast Edge Inference of Mixture-of-Experts Models via Cross-Layer Gate
di: Fang, Zhiyuan, et al.
Pubblicazione: (2025)
di: Fang, Zhiyuan, et al.
Pubblicazione: (2025)
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
di: Gu, Yifeng, et al.
Pubblicazione: (2025)
di: Gu, Yifeng, et al.
Pubblicazione: (2025)
A Systematic Study of Cross-Layer KV Sharing for Efficient LLM Inference
di: Wu, You, et al.
Pubblicazione: (2024)
di: Wu, You, et al.
Pubblicazione: (2024)
Reconstructing KV Caches with Cross-layer Fusion For Enhanced Transformers
di: Lin, Hongzhan, et al.
Pubblicazione: (2025)
di: Lin, Hongzhan, et al.
Pubblicazione: (2025)
KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference
di: Lin, Jian, et al.
Pubblicazione: (2026)
di: Lin, Jian, et al.
Pubblicazione: (2026)
Broaden your SCOPE! Efficient Multi-turn Conversation Planning for LLMs with Semantic Space
di: Chen, Zhiliang, et al.
Pubblicazione: (2025)
di: Chen, Zhiliang, et al.
Pubblicazione: (2025)
SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers
di: Tang, Zicong, et al.
Pubblicazione: (2025)
di: Tang, Zicong, et al.
Pubblicazione: (2025)
CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation Generation
di: Cheng, Yiruo, et al.
Pubblicazione: (2024)
di: Cheng, Yiruo, et al.
Pubblicazione: (2024)
Pixel-Level Reasoning Segmentation via Multi-turn Conversations
di: Cai, Dexian, et al.
Pubblicazione: (2025)
di: Cai, Dexian, et al.
Pubblicazione: (2025)
FlowKV: Enhancing Multi-Turn Conversational Coherence in LLMs via Isolated Key-Value Cache Management
di: Liu, Xiang, et al.
Pubblicazione: (2025)
di: Liu, Xiang, et al.
Pubblicazione: (2025)
AM$^3$Safety: Towards Data Efficient Alignment of Multi-modal Multi-turn Safety for MLLMs
di: Zhu, Han, et al.
Pubblicazione: (2026)
di: Zhu, Han, et al.
Pubblicazione: (2026)
Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention
di: Gao, Bin, et al.
Pubblicazione: (2024)
di: Gao, Bin, et al.
Pubblicazione: (2024)
FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMs
di: Fan, Zhiting, et al.
Pubblicazione: (2024)
di: Fan, Zhiting, et al.
Pubblicazione: (2024)
On the Multi-turn Instruction Following for Conversational Web Agents
di: Deng, Yang, et al.
Pubblicazione: (2024)
di: Deng, Yang, et al.
Pubblicazione: (2024)
Beyond KV Caching: Shared Attention for Efficient LLMs
di: Liao, Bingli, et al.
Pubblicazione: (2024)
di: Liao, Bingli, et al.
Pubblicazione: (2024)
DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving
di: Liu, Yuhan, et al.
Pubblicazione: (2024)
di: Liu, Yuhan, et al.
Pubblicazione: (2024)
GriDB: Scaling Blockchain Database via Sharding and Off-Chain Cross-Shard Mechanism
di: Hong, Zicong, et al.
Pubblicazione: (2024)
di: Hong, Zicong, et al.
Pubblicazione: (2024)
KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
di: Yang, Yifei, et al.
Pubblicazione: (2024)
di: Yang, Yifei, et al.
Pubblicazione: (2024)
Learning Multi-Indicator Weights for Data Selection: A Joint Task-Model Adaptation Framework with Efficient Proxies
di: Song, Jingze, et al.
Pubblicazione: (2026)
di: Song, Jingze, et al.
Pubblicazione: (2026)
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
di: Cai, Zefan, et al.
Pubblicazione: (2024)
di: Cai, Zefan, et al.
Pubblicazione: (2024)
DyMoE: Dynamic Expert Orchestration with Mixed-Precision Quantization for Efficient MoE Inference on Edge
di: Huang, Yuegui, et al.
Pubblicazione: (2026)
di: Huang, Yuegui, et al.
Pubblicazione: (2026)
Rethinking Text-to-SQL: Dynamic Multi-turn SQL Interaction for Real-world Database Exploration
di: Sun, Linzhuang, et al.
Pubblicazione: (2025)
di: Sun, Linzhuang, et al.
Pubblicazione: (2025)
A State-Update Prompting Strategy for Efficient and Robust Multi-turn Dialogue
di: Liu, Ziyi
Pubblicazione: (2025)
di: Liu, Ziyi
Pubblicazione: (2025)
Training and Serving System of Foundation Models: A Comprehensive Survey
di: Zhou, Jiahang, et al.
Pubblicazione: (2024)
di: Zhou, Jiahang, et al.
Pubblicazione: (2024)
NaturalConv: A Chinese Dialogue Dataset Towards Multi-turn Topic-driven Conversation
di: Wang, Xiaoyang, et al.
Pubblicazione: (2021)
di: Wang, Xiaoyang, et al.
Pubblicazione: (2021)
Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective Model
di: Niu, Fuqiang, et al.
Pubblicazione: (2024)
di: Niu, Fuqiang, et al.
Pubblicazione: (2024)
LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation
di: Chen, Han, et al.
Pubblicazione: (2025)
di: Chen, Han, et al.
Pubblicazione: (2025)
Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training
di: Chen, Maximillian, et al.
Pubblicazione: (2024)
di: Chen, Maximillian, et al.
Pubblicazione: (2024)
Token Statistics Reveal Conversational Drift in Multi-turn LLM Interaction
di: Hafez, Wael, et al.
Pubblicazione: (2026)
di: Hafez, Wael, et al.
Pubblicazione: (2026)
SemShareKV: Efficient KVCache Sharing for Semantically Similar Prompts via Token-Level LSH Matching
di: Zhao, Xinye, et al.
Pubblicazione: (2025)
di: Zhao, Xinye, et al.
Pubblicazione: (2025)
ForesightKV: Optimizing KV Cache Eviction for Reasoning Models by Learning Long-Term Contribution
di: Dong, Zican, et al.
Pubblicazione: (2026)
di: Dong, Zican, et al.
Pubblicazione: (2026)
MAC: A Multi-Agent Framework for Interactive User Clarification in Multi-turn Conversations
di: Acikgoz, Emre Can, et al.
Pubblicazione: (2025)
di: Acikgoz, Emre Can, et al.
Pubblicazione: (2025)
SafeMT: Multi-turn Safety for Multimodal Language Models
di: Zhu, Han, et al.
Pubblicazione: (2025)
di: Zhu, Han, et al.
Pubblicazione: (2025)
Source-primed Multi-turn Conversation Helps Large Language Models Translate Documents
di: Hu, Hanxu, et al.
Pubblicazione: (2025)
di: Hu, Hanxu, et al.
Pubblicazione: (2025)
PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference
di: Patel, Ishan, et al.
Pubblicazione: (2026)
di: Patel, Ishan, et al.
Pubblicazione: (2026)
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
di: Zhou, Xiabin, et al.
Pubblicazione: (2024)
di: Zhou, Xiabin, et al.
Pubblicazione: (2024)
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
di: Tong, Terry, et al.
Pubblicazione: (2024)
di: Tong, Terry, et al.
Pubblicazione: (2024)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
di: Wang, Yixuan, et al.
Pubblicazione: (2025)
di: Wang, Yixuan, et al.
Pubblicazione: (2025)
LabelCoRank: Revolutionizing Long Tail Multi-Label Classification with Co-Occurrence Reranking
di: Yan, Yan, et al.
Pubblicazione: (2025)
di: Yan, Yan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline
di: Fang, Zhiyuan, et al.
Pubblicazione: (2025) -
Fate: Fast Edge Inference of Mixture-of-Experts Models via Cross-Layer Gate
di: Fang, Zhiyuan, et al.
Pubblicazione: (2025) -
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
di: Gu, Yifeng, et al.
Pubblicazione: (2025) -
A Systematic Study of Cross-Layer KV Sharing for Efficient LLM Inference
di: Wu, You, et al.
Pubblicazione: (2024) -
Reconstructing KV Caches with Cross-layer Fusion For Enhanced Transformers
di: Lin, Hongzhan, et al.
Pubblicazione: (2025)