Salvato in:
| Autori principali: | Biton, Dvir David, Friedman, Roy |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2603.03301 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025)
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025)
dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching
di: Liu, Zhiyuan, et al.
Pubblicazione: (2025)
di: Liu, Zhiyuan, et al.
Pubblicazione: (2025)
CaRT: Teaching LLM Agents to Know When They Know Enough
di: Liu, Grace, et al.
Pubblicazione: (2025)
di: Liu, Grace, et al.
Pubblicazione: (2025)
Screening Is Enough
di: Nakanishi, Ken M.
Pubblicazione: (2026)
di: Nakanishi, Ken M.
Pubblicazione: (2026)
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
di: Saxena, Utkarsh, et al.
Pubblicazione: (2024)
di: Saxena, Utkarsh, et al.
Pubblicazione: (2024)
KV Cache Transform Coding for Compact Storage in LLM Inference
di: Staniszewski, Konrad, et al.
Pubblicazione: (2025)
di: Staniszewski, Konrad, et al.
Pubblicazione: (2025)
PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents
di: Gu, Zhuohan, et al.
Pubblicazione: (2026)
di: Gu, Zhuohan, et al.
Pubblicazione: (2026)
TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference
di: Dzikanyanga, Gradwell, et al.
Pubblicazione: (2026)
di: Dzikanyanga, Gradwell, et al.
Pubblicazione: (2026)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
di: Liu, Guangda, et al.
Pubblicazione: (2025)
di: Liu, Guangda, et al.
Pubblicazione: (2025)
Enough Coin Flips Can Make LLMs Act Bayesian
di: Gupta, Ritwik, et al.
Pubblicazione: (2025)
di: Gupta, Ritwik, et al.
Pubblicazione: (2025)
Understanding LLM Embeddings for Regression
di: Tang, Eric, et al.
Pubblicazione: (2024)
di: Tang, Eric, et al.
Pubblicazione: (2024)
Static Word Embeddings for Sentence Semantic Representation
di: Wada, Takashi, et al.
Pubblicazione: (2025)
di: Wada, Takashi, et al.
Pubblicazione: (2025)
MeanCache: User-Centric Semantic Caching for LLM Web Services
di: Gill, Waris, et al.
Pubblicazione: (2024)
di: Gill, Waris, et al.
Pubblicazione: (2024)
User-LLM: Efficient LLM Contextualization with User Embeddings
di: Ning, Lin, et al.
Pubblicazione: (2024)
di: Ning, Lin, et al.
Pubblicazione: (2024)
Reward Is Enough: LLMs Are In-Context Reinforcement Learners
di: Song, Kefan, et al.
Pubblicazione: (2025)
di: Song, Kefan, et al.
Pubblicazione: (2025)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
di: Tian, Yuxuan, et al.
Pubblicazione: (2025)
di: Tian, Yuxuan, et al.
Pubblicazione: (2025)
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
di: Kang, Hao, et al.
Pubblicazione: (2024)
di: Kang, Hao, et al.
Pubblicazione: (2024)
On the Detectability of LLM-Generated Text: What Exactly Is LLM-Generated Text?
di: Geng, Mingmeng, et al.
Pubblicazione: (2025)
di: Geng, Mingmeng, et al.
Pubblicazione: (2025)
BPO: Staying Close to the Behavior LLM Creates Better Online LLM Alignment
di: Xu, Wenda, et al.
Pubblicazione: (2024)
di: Xu, Wenda, et al.
Pubblicazione: (2024)
Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation
di: Tang, Pingzhi, et al.
Pubblicazione: (2026)
di: Tang, Pingzhi, et al.
Pubblicazione: (2026)
AgenticCache: Cache-Driven Asynchronous Planning for Embodied AI Agents
di: Kim, Hojoon, et al.
Pubblicazione: (2026)
di: Kim, Hojoon, et al.
Pubblicazione: (2026)
Generalist Foundation Models Are Not Clinical Enough for Hospital Operations
di: Jiang, Lavender Y., et al.
Pubblicazione: (2025)
di: Jiang, Lavender Y., et al.
Pubblicazione: (2025)
Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation
di: He, Zhenyu, et al.
Pubblicazione: (2024)
di: He, Zhenyu, et al.
Pubblicazione: (2024)
BERT-JEPA: Reorganizing CLS Embeddings for Language-Invariant Semantics
di: Gillin, Taj, et al.
Pubblicazione: (2026)
di: Gillin, Taj, et al.
Pubblicazione: (2026)
A General Framework for Producing Interpretable Semantic Text Embeddings
di: Sun, Yiqun, et al.
Pubblicazione: (2024)
di: Sun, Yiqun, et al.
Pubblicazione: (2024)
Confidence-aware Self-Semantic Distillation on Knowledge Graph Embedding
di: Liu, Yichen, et al.
Pubblicazione: (2022)
di: Liu, Yichen, et al.
Pubblicazione: (2022)
Output Embedding Centering for Stable LLM Pretraining
di: Stollenwerk, Felix, et al.
Pubblicazione: (2026)
di: Stollenwerk, Felix, et al.
Pubblicazione: (2026)
Aligned at the Start: Conceptual Groupings in LLM Embeddings
di: Khatir, Mehrdad, et al.
Pubblicazione: (2024)
di: Khatir, Mehrdad, et al.
Pubblicazione: (2024)
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
di: Liu, Akide, et al.
Pubblicazione: (2024)
di: Liu, Akide, et al.
Pubblicazione: (2024)
OccamLLM: Fast and Exact Language Model Arithmetic in a Single Step
di: Dugan, Owen, et al.
Pubblicazione: (2024)
di: Dugan, Owen, et al.
Pubblicazione: (2024)
LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
di: Yang, Junxiao, et al.
Pubblicazione: (2026)
di: Yang, Junxiao, et al.
Pubblicazione: (2026)
When Less is Enough: Efficient Inference via Collaborative Reasoning
di: Chen, Yilei, et al.
Pubblicazione: (2026)
di: Chen, Yilei, et al.
Pubblicazione: (2026)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
di: Zhuang, Haomin, et al.
Pubblicazione: (2024)
di: Zhuang, Haomin, et al.
Pubblicazione: (2024)
Improving Uncertainty Quantification in Large Language Models via Semantic Embeddings
di: Grewal, Yashvir S., et al.
Pubblicazione: (2024)
di: Grewal, Yashvir S., et al.
Pubblicazione: (2024)
Representing Rule-based Chatbots with Transformers
di: Friedman, Dan, et al.
Pubblicazione: (2024)
di: Friedman, Dan, et al.
Pubblicazione: (2024)
DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving
di: Liu, Yuhan, et al.
Pubblicazione: (2024)
di: Liu, Yuhan, et al.
Pubblicazione: (2024)
Sparse is Enough in Fine-tuning Pre-trained Large Language Models
di: Song, Weixi, et al.
Pubblicazione: (2023)
di: Song, Weixi, et al.
Pubblicazione: (2023)
Open or Closed LLM for Lesser-Resourced Languages? Lessons from Greek
di: Pavlopoulos, John, et al.
Pubblicazione: (2025)
di: Pavlopoulos, John, et al.
Pubblicazione: (2025)
The Semantic Illusion: Certified Limits of Embedding-Based Hallucination Detection in RAG Systems
di: Sinha, Debu
Pubblicazione: (2025)
di: Sinha, Debu
Pubblicazione: (2025)
LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models
di: Shi, Dachuan, et al.
Pubblicazione: (2025)
di: Shi, Dachuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025) -
dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching
di: Liu, Zhiyuan, et al.
Pubblicazione: (2025) -
CaRT: Teaching LLM Agents to Know When They Know Enough
di: Liu, Grace, et al.
Pubblicazione: (2025) -
Screening Is Enough
di: Nakanishi, Ken M.
Pubblicazione: (2026) -
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
di: Saxena, Utkarsh, et al.
Pubblicazione: (2024)