AdapShot: Adaptive Many-Shot In-Context Learning with Semantic-Aware KV Cache Reuse
Fuente:
arXiv
Salvato in:
| Autori principali: | Ou, Jie, Guo, Jinyu, Guo, Shiyao, Li, Yuang, Wu, Ruiqi, Wang, Zhaokun, Li, Wenyi, Tian, Wenhong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
di: li, Fei, et al.
Pubblicazione: (2026)
di: li, Fei, et al.
Pubblicazione: (2026)
Accelerating Adaptive Retrieval Augmented Generation via Instruction-Driven Representation Reduction of Retrieval Overlaps
di: Ou, Jie, et al.
Pubblicazione: (2025)
di: Ou, Jie, et al.
Pubblicazione: (2025)
ALTER: Asymmetric LoRA for Token-Entropy-Guided Unlearning of LLMs
di: Chen, Xunlei, et al.
Pubblicazione: (2026)
di: Chen, Xunlei, et al.
Pubblicazione: (2026)
CAP: Controllable Alignment Prompting for Unlearning in LLMs
di: Wang, Zhaokun, et al.
Pubblicazione: (2026)
di: Wang, Zhaokun, et al.
Pubblicazione: (2026)
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
di: Zhou, Xiabin, et al.
Pubblicazione: (2024)
di: Zhou, Xiabin, et al.
Pubblicazione: (2024)
HASH-RAG: Bridging Deep Hashing with Retriever for Efficient, Fine Retrieval and Augmented Generation
di: Guo, Jinyu, et al.
Pubblicazione: (2025)
di: Guo, Jinyu, et al.
Pubblicazione: (2025)
Noise-Robustness Through Noise: A Framework combining Asymmetric LoRA with Poisoning MoE
di: Wang, Zhaokun, et al.
Pubblicazione: (2025)
di: Wang, Zhaokun, et al.
Pubblicazione: (2025)
MAPLE: Many-Shot Adaptive Pseudo-Labeling for In-Context Learning
di: Chen, Zihan, et al.
Pubblicazione: (2025)
di: Chen, Zihan, et al.
Pubblicazione: (2025)
Many-Shot In-Context Learning
di: Agarwal, Rishabh, et al.
Pubblicazione: (2024)
di: Agarwal, Rishabh, et al.
Pubblicazione: (2024)
CacheClip: Accelerating RAG with Effective KV Cache Reuse
di: Yang, Bin, et al.
Pubblicazione: (2025)
di: Yang, Bin, et al.
Pubblicazione: (2025)
DASH-KV: Accelerating Long-Context LLM Inference via Asymmetric KV Cache Hashing
di: Guo, Jinyu, et al.
Pubblicazione: (2026)
di: Guo, Jinyu, et al.
Pubblicazione: (2026)
Towards More Economical Context-Augmented LLM Generation by Reusing Stored KV Cache
di: Li, Hanchen, et al.
Pubblicazione: (2025)
di: Li, Hanchen, et al.
Pubblicazione: (2025)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
di: Liu, Guangda, et al.
Pubblicazione: (2024)
di: Liu, Guangda, et al.
Pubblicazione: (2024)
MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference
di: Li, Yu, et al.
Pubblicazione: (2026)
di: Li, Yu, et al.
Pubblicazione: (2026)
ZSMerge: Zero-Shot KV Cache Compression for Memory-Efficient Long-Context LLMs
di: Liu, Xin, et al.
Pubblicazione: (2025)
di: Liu, Xin, et al.
Pubblicazione: (2025)
Compressing Many-Shots in In-Context Learning
di: Khatri, Devvrit, et al.
Pubblicazione: (2025)
di: Khatri, Devvrit, et al.
Pubblicazione: (2025)
KVmix: Gradient-Based Layer Importance-Aware Mixed-Precision Quantization for KV Cache
di: Li, Fei, et al.
Pubblicazione: (2025)
di: Li, Fei, et al.
Pubblicazione: (2025)
CSKV: Training-Efficient Channel Shrinking for KV Cache in Long-Context Scenarios
di: Wang, Luning, et al.
Pubblicazione: (2024)
di: Wang, Luning, et al.
Pubblicazione: (2024)
AFANet: Adaptive Frequency-Aware Network for Weakly-Supervised Few-Shot Semantic Segmentation
di: Ma, Jiaqi, et al.
Pubblicazione: (2024)
di: Ma, Jiaqi, et al.
Pubblicazione: (2024)
Lossless Acceleration of Large Language Model via Adaptive N-gram Parallel Decoding
di: Ou, Jie, et al.
Pubblicazione: (2024)
di: Ou, Jie, et al.
Pubblicazione: (2024)
RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache Reuse
di: Geng, Yingsheng, et al.
Pubblicazione: (2026)
di: Geng, Yingsheng, et al.
Pubblicazione: (2026)
On Many-Shot In-Context Learning for Long-Context Evaluation
di: Zou, Kaijian, et al.
Pubblicazione: (2024)
di: Zou, Kaijian, et al.
Pubblicazione: (2024)
EFIM: Efficient Serving of LLMs for Infilling Tasks with Improved KV Cache Reuse
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
di: Liu, Xiang, et al.
Pubblicazione: (2025)
di: Liu, Xiang, et al.
Pubblicazione: (2025)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
di: Li, Kunxi, et al.
Pubblicazione: (2025)
di: Li, Kunxi, et al.
Pubblicazione: (2025)
Many-Shot In-Context Learning for Molecular Inverse Design
di: Moayedpour, Saeed, et al.
Pubblicazione: (2024)
di: Moayedpour, Saeed, et al.
Pubblicazione: (2024)
Selecting Demonstrations for Many-Shot In-Context Learning via Gradient Matching
di: Zhang, Jianfei, et al.
Pubblicazione: (2025)
di: Zhang, Jianfei, et al.
Pubblicazione: (2025)
Enabling Small Models for Zero-Shot Selection and Reuse through Model Label Learning
di: Zhang, Jia, et al.
Pubblicazione: (2024)
di: Zhang, Jia, et al.
Pubblicazione: (2024)
SABlock: Semantic-Aware KV Cache Eviction with Adaptive Compression Block Size
di: Chen, Jinhan, et al.
Pubblicazione: (2025)
di: Chen, Jinhan, et al.
Pubblicazione: (2025)
KV-CAR: KV Cache Compression using Autoencoders and KV Reuse in Large Language Models
di: Roy, Sourjya, et al.
Pubblicazione: (2025)
di: Roy, Sourjya, et al.
Pubblicazione: (2025)
Neuron: Learning Context-Aware Evolving Representations for Zero-Shot Skeleton Action Recognition
di: Chen, Yang, et al.
Pubblicazione: (2024)
di: Chen, Yang, et al.
Pubblicazione: (2024)
Efficient Many-Shot In-Context Learning with Dynamic Block-Sparse Attention
di: Xiao, Emily, et al.
Pubblicazione: (2025)
di: Xiao, Emily, et al.
Pubblicazione: (2025)
Many-Shot In-Context Learning in Multimodal Foundation Models
di: Jiang, Yixing, et al.
Pubblicazione: (2024)
di: Jiang, Yixing, et al.
Pubblicazione: (2024)
Towards Compute-Optimal Many-Shot In-Context Learning
di: Golchin, Shahriar, et al.
Pubblicazione: (2025)
di: Golchin, Shahriar, et al.
Pubblicazione: (2025)
OjaKV: Context-Aware Online Low-Rank KV Cache Compression
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025)
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025)
KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs
di: Chen, Chuangtao, et al.
Pubblicazione: (2026)
di: Chen, Chuangtao, et al.
Pubblicazione: (2026)
Distilling Many-Shot In-Context Learning into a Cheat Sheet
di: Honda, Ukyo, et al.
Pubblicazione: (2025)
di: Honda, Ukyo, et al.
Pubblicazione: (2025)
Where and What Matters: Sensitivity-Aware Task Vectors for Many-Shot Multimodal In-Context Learning
di: Ma, Ziyu, et al.
Pubblicazione: (2025)
di: Ma, Ziyu, et al.
Pubblicazione: (2025)
Meta-Semantics Augmented Few-Shot Relational Learning
di: Wu, Han, et al.
Pubblicazione: (2025)
di: Wu, Han, et al.
Pubblicazione: (2025)
Learning by Neighbor-Aware Semantics, Deciding by Open-form Flows: Towards Robust Zero-Shot Skeleton Action Recognition
di: Chen, Yang, et al.
Pubblicazione: (2025)
di: Chen, Yang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
di: li, Fei, et al.
Pubblicazione: (2026) -
Accelerating Adaptive Retrieval Augmented Generation via Instruction-Driven Representation Reduction of Retrieval Overlaps
di: Ou, Jie, et al.
Pubblicazione: (2025) -
ALTER: Asymmetric LoRA for Token-Entropy-Guided Unlearning of LLMs
di: Chen, Xunlei, et al.
Pubblicazione: (2026) -
CAP: Controllable Alignment Prompting for Unlearning in LLMs
di: Wang, Zhaokun, et al.
Pubblicazione: (2026) -
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
di: Zhou, Xiabin, et al.
Pubblicazione: (2024)