An experimental study of KV cache reuse strategies in chunk-level caching systems
Fuente:
arXiv
Guardado en:
| Autores principales: | Cestola, Samuel, Xia, Tianxiang, Weiyan, Zheng, Pengfei, Zheng, Didona, Diego |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Accelerating Suffix Jailbreak attacks with Prefix-Shared KV-cache
por: Wang, Xinhai, et al.
Publicado: (2026)
por: Wang, Xinhai, et al.
Publicado: (2026)
KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference
por: Nadali, Alireza, et al.
Publicado: (2026)
por: Nadali, Alireza, et al.
Publicado: (2026)
Constraint-Driven Small Language Models Based on Agent and OpenAlex Knowledge Graph: Mining Conceptual Pathways and Discovering Innovation Points in Academic Papers
por: Xia, Ziye, et al.
Publicado: (2025)
por: Xia, Ziye, et al.
Publicado: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMs
por: Zhang, Zhaowei, et al.
Publicado: (2025)
por: Zhang, Zhaowei, et al.
Publicado: (2025)
Enhancing Automated Essay Scoring with Three Techniques: Two-Stage Fine-Tuning, Score Alignment, and Self-Training
por: Choi, Hongseok, et al.
Publicado: (2026)
por: Choi, Hongseok, et al.
Publicado: (2026)
SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference
por: Chauhan, Anay, et al.
Publicado: (2026)
por: Chauhan, Anay, et al.
Publicado: (2026)
Transactional Attention: Semantic Sponsorship for KV-Cache Retention
por: Basu, Abhinaba
Publicado: (2026)
por: Basu, Abhinaba
Publicado: (2026)
Large Language Model (LLM) Bias Index -- LLMBI
por: Oketunji, Abiodun Finbarrs, et al.
Publicado: (2023)
por: Oketunji, Abiodun Finbarrs, et al.
Publicado: (2023)
Engineering A Large Language Model From Scratch
por: Oketunji, Abiodun Finbarrs
Publicado: (2024)
por: Oketunji, Abiodun Finbarrs
Publicado: (2024)
Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction
por: Garcia, Gabriel
Publicado: (2026)
por: Garcia, Gabriel
Publicado: (2026)
Dual-Phase Federated Deep Unlearning via Weight-Aware Rollback and Reconstruction
por: Zhou, Changjun, et al.
Publicado: (2025)
por: Zhou, Changjun, et al.
Publicado: (2025)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
por: Fadli, Samih
Publicado: (2025)
por: Fadli, Samih
Publicado: (2025)
Generalizing Test-time Compute-optimal Scaling as an Optimizable Graph
por: Wang, Fali, et al.
Publicado: (2025)
por: Wang, Fali, et al.
Publicado: (2025)
Reasoning Over the Glyphs: Evaluation of LLM's Decipherment of Rare Scripts
por: Shih, Yu-Fei, et al.
Publicado: (2025)
por: Shih, Yu-Fei, et al.
Publicado: (2025)
MMSciBench: Benchmarking Language Models on Chinese Multimodal Scientific Problems
por: Ye, Xinwu, et al.
Publicado: (2025)
por: Ye, Xinwu, et al.
Publicado: (2025)
Influence-driven Curriculum Learning for Pre-training on Limited Data
por: Schoenegger, Loris, et al.
Publicado: (2025)
por: Schoenegger, Loris, et al.
Publicado: (2025)
Towards Latent Diffusion Suitable For Text
por: Midavaine, Nesta, et al.
Publicado: (2026)
por: Midavaine, Nesta, et al.
Publicado: (2026)
SpectralLoRA: Is Low-Frequency Structure Sufficient for LoRA Adaptation? A Spectral Analysis of Weight Updates
por: Singh, Rajveer
Publicado: (2026)
por: Singh, Rajveer
Publicado: (2026)
BLP-2023 Task 2: Sentiment Analysis
por: Hasan, Md. Arid, et al.
Publicado: (2023)
por: Hasan, Md. Arid, et al.
Publicado: (2023)
FIM-LoRA: Task-Informative Rank Allocation for LoRA via Calibration-Time Gradient-Variance Estimation
por: Sathyavageeswaran, Ramakrishnan
Publicado: (2026)
por: Sathyavageeswaran, Ramakrishnan
Publicado: (2026)
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
por: Hanna, Michael, et al.
Publicado: (2024)
por: Hanna, Michael, et al.
Publicado: (2024)
Integrating Expert Labels into LLM-based Emission Goal Detection: Example Selection vs Automatic Prompt Design
por: Wrzalik, Marco, et al.
Publicado: (2024)
por: Wrzalik, Marco, et al.
Publicado: (2024)
Where Should LoRA Go? Component-Type Placement in Hybrid Language Models
por: Borobia, Hector, et al.
Publicado: (2026)
por: Borobia, Hector, et al.
Publicado: (2026)
Interpreto: An Explainability Library for Transformers
por: Poché, Antonin, et al.
Publicado: (2025)
por: Poché, Antonin, et al.
Publicado: (2025)
RightNow-Arabic-0.5B-Turbo: An Open Sub-1B Arabic Language Model via Vocabulary Injection and Edge-First Deployment
por: Jaber, Jaber, et al.
Publicado: (2026)
por: Jaber, Jaber, et al.
Publicado: (2026)
Structured-Sparse Attention for Entity Tracking with Subquadratic Sequence Complexity
por: Zhao, Hangyue, et al.
Publicado: (2026)
por: Zhao, Hangyue, et al.
Publicado: (2026)
Efficacy of ByT5 in Multilingual Translation of Biblical Texts for Underrepresented Languages
por: Aars, Corinne, et al.
Publicado: (2024)
por: Aars, Corinne, et al.
Publicado: (2024)
GenAI Content Detection Task 3: Cross-Domain Machine-Generated Text Detection Challenge
por: Dugan, Liam, et al.
Publicado: (2025)
por: Dugan, Liam, et al.
Publicado: (2025)
WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training
por: Tian, Changxin, et al.
Publicado: (2025)
por: Tian, Changxin, et al.
Publicado: (2025)
PowLU: An Activation Function for Stable Pre-Training of LLMs
por: Jiang, Peijie, et al.
Publicado: (2026)
por: Jiang, Peijie, et al.
Publicado: (2026)
Resonant Context Anchoring: Decoupling Attention Routing and Signal Gain at Inference Time
por: Zhao, Mingkuan, et al.
Publicado: (2026)
por: Zhao, Mingkuan, et al.
Publicado: (2026)
Reconstructing Syllable Sequences in Abugida Scripts with Incomplete Inputs
por: Thu, Ye Kyaw, et al.
Publicado: (2025)
por: Thu, Ye Kyaw, et al.
Publicado: (2025)
Human-interpretable clustering of short-text using large language models
por: Miller, Justin K., et al.
Publicado: (2024)
por: Miller, Justin K., et al.
Publicado: (2024)
Scaling Laws for Forgetting When Fine-Tuning Large Language Models
por: Kalajdzievski, Damjan
Publicado: (2024)
por: Kalajdzievski, Damjan
Publicado: (2024)
Structure-Guided Entity Resolution: Fine-Tuning LLMs for Robust Name Matching in Complex Linguistic Contexts
por: Chourasia, Shivam, et al.
Publicado: (2026)
por: Chourasia, Shivam, et al.
Publicado: (2026)
Lon-ea at SemEval-2023 Task 11: A Comparison of Activation Functions for Soft and Hard Label Prediction
por: Hosseini, Peyman, et al.
Publicado: (2023)
por: Hosseini, Peyman, et al.
Publicado: (2023)
A Framework for Fine-Tuning LLMs using Heterogeneous Feedback
por: Aponte, Ryan, et al.
Publicado: (2024)
por: Aponte, Ryan, et al.
Publicado: (2024)
Ambiguity in LLMs is a concept missing problem
por: Hu, Zhibo, et al.
Publicado: (2025)
por: Hu, Zhibo, et al.
Publicado: (2025)
Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment
por: Zhang, Yizhuo, et al.
Publicado: (2025)
por: Zhang, Yizhuo, et al.
Publicado: (2025)
Ejemplares similares
-
Accelerating Suffix Jailbreak attacks with Prefix-Shared KV-cache
por: Wang, Xinhai, et al.
Publicado: (2026) -
KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference
por: Nadali, Alireza, et al.
Publicado: (2026) -
Constraint-Driven Small Language Models Based on Agent and OpenAlex Knowledge Graph: Mining Conceptual Pathways and Discovering Innovation Points in Academic Papers
por: Xia, Ziye, et al.
Publicado: (2025) -
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
por: Oketunji, Abiodun Finbarrs
Publicado: (2023) -
Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMs
por: Zhang, Zhaowei, et al.
Publicado: (2025)