KV-Embedding: Training-free Text Embedding via Internal KV Re-routing in Decoder-only LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Tang, Yixuan, Yang, Yi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FinMTEB: Finance Massive Text Embedding Benchmark
por: Tang, Yixuan, et al.
Publicado: (2025)
por: Tang, Yixuan, et al.
Publicado: (2025)
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding
por: Shi, Luohe, et al.
Publicado: (2025)
por: Shi, Luohe, et al.
Publicado: (2025)
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
por: Ji, Shiyu, et al.
Publicado: (2026)
por: Ji, Shiyu, et al.
Publicado: (2026)
TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering
por: Joshi, Vinay, et al.
Publicado: (2025)
por: Joshi, Vinay, et al.
Publicado: (2025)
GAPrune: Gradient-Alignment Pruning for Domain-Aware Embeddings
por: Tang, Yixuan, et al.
Publicado: (2025)
por: Tang, Yixuan, et al.
Publicado: (2025)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
por: Liao, Mengqi, et al.
Publicado: (2025)
por: Liao, Mengqi, et al.
Publicado: (2025)
Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
por: Wu, Chengyue, et al.
Publicado: (2025)
por: Wu, Chengyue, et al.
Publicado: (2025)
ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal Smoothing
por: An, Yongqi, et al.
Publicado: (2026)
por: An, Yongqi, et al.
Publicado: (2026)
Revealing the Numeracy Gap: An Empirical Investigation of Text Embedding Models
por: Deng, Ningyuan, et al.
Publicado: (2025)
por: Deng, Ningyuan, et al.
Publicado: (2025)
Crystal-KV: Efficient KV Cache Management for Chain-of-Thought LLMs via Answer-First Principle
por: Wang, Zihan, et al.
Publicado: (2026)
por: Wang, Zihan, et al.
Publicado: (2026)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
por: Jo, Dongwon, et al.
Publicado: (2025)
por: Jo, Dongwon, et al.
Publicado: (2025)
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
por: Zhou, Xiabin, et al.
Publicado: (2024)
por: Zhou, Xiabin, et al.
Publicado: (2024)
Pooling And Attention: What Are Effective Designs For LLM-Based Embedding Models?
por: Tang, Yixuan, et al.
Publicado: (2024)
por: Tang, Yixuan, et al.
Publicado: (2024)
Do We Need Domain-Specific Embedding Models? An Empirical Investigation
por: Tang, Yixuan, et al.
Publicado: (2024)
por: Tang, Yixuan, et al.
Publicado: (2024)
ManifoldKV: Training-Free KV Cache Compression via Euclidean Outlier Detection
por: Datta, Debajyoti, et al.
Publicado: (2026)
por: Datta, Debajyoti, et al.
Publicado: (2026)
KV-CoRE: Benchmarking Data-Dependent Low-Rank Compressibility of KV-Caches in LLMs
por: Chen, Jian, et al.
Publicado: (2026)
por: Chen, Jian, et al.
Publicado: (2026)
DASH-KV: Accelerating Long-Context LLM Inference via Asymmetric KV Cache Hashing
por: Guo, Jinyu, et al.
Publicado: (2026)
por: Guo, Jinyu, et al.
Publicado: (2026)
Residual-Mass Accounting for Partial-KV Decoding
por: Hoshi, Yasuto, et al.
Publicado: (2026)
por: Hoshi, Yasuto, et al.
Publicado: (2026)
ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs
por: Qi, Yanlin, et al.
Publicado: (2026)
por: Qi, Yanlin, et al.
Publicado: (2026)
GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
por: Huang, Kung-Hsiang, et al.
Publicado: (2025)
por: Huang, Kung-Hsiang, et al.
Publicado: (2025)
Value-Guided KV Compression for LLMs via Approximated CUR Decomposition
por: Sengupta, Ayan, et al.
Publicado: (2025)
por: Sengupta, Ayan, et al.
Publicado: (2025)
Conan-Embedding-v2: Training an LLM from Scratch for Text Embeddings
por: Li, Shiyu, et al.
Publicado: (2025)
por: Li, Shiyu, et al.
Publicado: (2025)
GRKV: Global Regression for Training-Free KV Cache Compression in Long-Context LLMs
por: Peng, Junjie, et al.
Publicado: (2026)
por: Peng, Junjie, et al.
Publicado: (2026)
EntmaxKV: Support-Aware Decoding for Entmax Attention
por: Duarte, Gonçalo, et al.
Publicado: (2026)
por: Duarte, Gonçalo, et al.
Publicado: (2026)
In-context KV-Cache Eviction for LLMs via Attention-Gate
por: Zeng, Zihao, et al.
Publicado: (2024)
por: Zeng, Zihao, et al.
Publicado: (2024)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
por: Liu, Xiang, et al.
Publicado: (2025)
por: Liu, Xiang, et al.
Publicado: (2025)
Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads
por: He, Xingyang, et al.
Publicado: (2025)
por: He, Xingyang, et al.
Publicado: (2025)
Causal2Vec: Improving Decoder-only LLMs as Embedding Models through a Contextual Token
por: Lin, Ailiang, et al.
Publicado: (2025)
por: Lin, Ailiang, et al.
Publicado: (2025)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
por: Zhu, Yuxuan, et al.
Publicado: (2025)
por: Zhu, Yuxuan, et al.
Publicado: (2025)
KV Cache Steering for Controlling Frozen LLMs
por: Belitsky, Max, et al.
Publicado: (2025)
por: Belitsky, Max, et al.
Publicado: (2025)
SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers
por: Tang, Zicong, et al.
Publicado: (2025)
por: Tang, Zicong, et al.
Publicado: (2025)
CTkvr: KV Cache Retrieval for Long-Context LLMs via Centroid then Token Indexing
por: Lu, Kuan, et al.
Publicado: (2025)
por: Lu, Kuan, et al.
Publicado: (2025)
Where Matters More Than What: Decoding-aligned KV Cache Compression via Position-aware Pseudo Queries
por: Tian, Zhenxu, et al.
Publicado: (2026)
por: Tian, Zhenxu, et al.
Publicado: (2026)
MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection
por: Lin, Bokai, et al.
Publicado: (2024)
por: Lin, Bokai, et al.
Publicado: (2024)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
por: Su, Zunhai, et al.
Publicado: (2025)
por: Su, Zunhai, et al.
Publicado: (2025)
TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization
por: Yao, Dingyu, et al.
Publicado: (2025)
por: Yao, Dingyu, et al.
Publicado: (2025)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
por: Du, Dayou, et al.
Publicado: (2025)
por: Du, Dayou, et al.
Publicado: (2025)
DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity
por: Hao, Jitai, et al.
Publicado: (2026)
por: Hao, Jitai, et al.
Publicado: (2026)
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
por: Yang, Dongquan, et al.
Publicado: (2025)
por: Yang, Dongquan, et al.
Publicado: (2025)
Embedding-based In-Context Prompt Training for Enhancing LLMs as Text Encoders
por: Lin, Ailiang, et al.
Publicado: (2026)
por: Lin, Ailiang, et al.
Publicado: (2026)
Ejemplares similares
-
FinMTEB: Finance Massive Text Embedding Benchmark
por: Tang, Yixuan, et al.
Publicado: (2025) -
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding
por: Shi, Luohe, et al.
Publicado: (2025) -
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
por: Ji, Shiyu, et al.
Publicado: (2026) -
TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering
por: Joshi, Vinay, et al.
Publicado: (2025) -
GAPrune: Gradient-Alignment Pruning for Domain-Aware Embeddings
por: Tang, Yixuan, et al.
Publicado: (2025)