Saved in:
| Main Authors: | Ren, Siyu, Zhu, Kenny Q. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2402.06262 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
by: Gu, Yifeng, et al.
Published: (2025)
by: Gu, Yifeng, et al.
Published: (2025)
Stepwise Alignment for Constrained Language Model Policy Optimization
by: Wachi, Akifumi, et al.
Published: (2024)
by: Wachi, Akifumi, et al.
Published: (2024)
Reformulating KV Cache Eviction Problem for Long-Context LLM Inference
by: Mai, Tho, et al.
Published: (2026)
by: Mai, Tho, et al.
Published: (2026)
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
by: Yang, Xintong, et al.
Published: (2026)
by: Yang, Xintong, et al.
Published: (2026)
EMO: Earth Mover Distance Optimization for Auto-Regressive Language Modeling
by: Ren, Siyu, et al.
Published: (2023)
by: Ren, Siyu, et al.
Published: (2023)
SynthEHR-Eviction: Enhancing Eviction SDoH Detection with LLM-Augmented Synthetic EHR Data
by: Yao, Zonghai, et al.
Published: (2025)
by: Yao, Zonghai, et al.
Published: (2025)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
by: Feng, Yuan, et al.
Published: (2024)
by: Feng, Yuan, et al.
Published: (2024)
A Cognitive Evaluation Benchmark of Image Reasoning and Description for Large Vision-Language Models
by: Song, Xiujie, et al.
Published: (2024)
by: Song, Xiujie, et al.
Published: (2024)
Eliciting Trustworthiness Priors of Large Language Models via Economic Games
by: Yan, Siyu, et al.
Published: (2026)
by: Yan, Siyu, et al.
Published: (2026)
Judge Q: Trainable Queries for Optimized Information Retention in KV Cache Eviction
by: Liu, Yijun, et al.
Published: (2025)
by: Liu, Yijun, et al.
Published: (2025)
Audit-of-Understanding: Posterior-Constrained Inference for Mathematical Reasoning in Language Models
by: Abdaljalil, Samir, et al.
Published: (2025)
by: Abdaljalil, Samir, et al.
Published: (2025)
Measuring Human and AI Values Based on Generative Psychometrics with Large Language Models
by: Ye, Haoran, et al.
Published: (2024)
by: Ye, Haoran, et al.
Published: (2024)
Generative Psycho-Lexical Approach for Constructing Value Systems in Large Language Models
by: Ye, Haoran, et al.
Published: (2025)
by: Ye, Haoran, et al.
Published: (2025)
KV-Runahead: Scalable Causal LLM Inference by Parallel Key-Value Cache Generation
by: Cho, Minsik, et al.
Published: (2024)
by: Cho, Minsik, et al.
Published: (2024)
Inference-Time Language Model Alignment via Integrated Value Guidance
by: Liu, Zhixuan, et al.
Published: (2024)
by: Liu, Zhixuan, et al.
Published: (2024)
RobustKV: Defending Large Language Models against Jailbreak Attacks via KV Eviction
by: Jiang, Tanqiu, et al.
Published: (2024)
by: Jiang, Tanqiu, et al.
Published: (2024)
Topic Over Source: The Key to Effective Data Mixing for Language Models Pre-training
by: Peng, Jiahui, et al.
Published: (2025)
by: Peng, Jiahui, et al.
Published: (2025)
Representations as Language: An Information-Theoretic Framework for Interpretability
by: Conklin, Henry, et al.
Published: (2024)
by: Conklin, Henry, et al.
Published: (2024)
Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
by: Li, Kunxi, et al.
Published: (2025)
by: Li, Kunxi, et al.
Published: (2025)
Activation-aware Probe-Query: Effective Key-Value Retrieval for Long-Context LLMs Inference
by: Xiao, Qingfa, et al.
Published: (2025)
by: Xiao, Qingfa, et al.
Published: (2025)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
by: Liao, Mengqi, et al.
Published: (2025)
by: Liao, Mengqi, et al.
Published: (2025)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
by: Wang, Guangtao, et al.
Published: (2025)
by: Wang, Guangtao, et al.
Published: (2025)
Schema Key Wording as an Instruction Channel in Structured Generation under Constrained Decoding
by: Le, Yifan
Published: (2026)
by: Le, Yifan
Published: (2026)
Key-Point-Driven Mathematical Reasoning Distillation of Large Language Model
by: Zhu, Xunyu, et al.
Published: (2024)
by: Zhu, Xunyu, et al.
Published: (2024)
Translate-and-Revise: Boosting Large Language Models for Constrained Translation
by: Huang, Pengcheng, et al.
Published: (2024)
by: Huang, Pengcheng, et al.
Published: (2024)
WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More
by: Yue, Yuxuan, et al.
Published: (2024)
by: Yue, Yuxuan, et al.
Published: (2024)
Is Your Image a Good Storyteller?
by: Song, Xiujie, et al.
Published: (2024)
by: Song, Xiujie, et al.
Published: (2024)
Large Language Model for Patent Concept Generation
by: Ren, Runtao, et al.
Published: (2024)
by: Ren, Runtao, et al.
Published: (2024)
Correlated Errors in Large Language Models
by: Kim, Elliot, et al.
Published: (2025)
by: Kim, Elliot, et al.
Published: (2025)
Depression Diagnosis Dialogue Simulation: Self-improving Psychiatrist with Tertiary Memory
by: Lan, Kunyao, et al.
Published: (2024)
by: Lan, Kunyao, et al.
Published: (2024)
TaskBench: Benchmarking Large Language Models for Task Automation
by: Shen, Yongliang, et al.
Published: (2023)
by: Shen, Yongliang, et al.
Published: (2023)
Conversational Control with Ontologies for Large Language Models: A Lightweight Framework for Constrained Generation
by: Gendron, Barbara, et al.
Published: (2026)
by: Gendron, Barbara, et al.
Published: (2026)
Unlocking Anticipatory Text Generation: A Constrained Approach for Large Language Models Decoding
by: Tu, Lifu, et al.
Published: (2023)
by: Tu, Lifu, et al.
Published: (2023)
Benchmarking Multi-National Value Alignment for Large Language Models
by: Shi, Weijie, et al.
Published: (2025)
by: Shi, Weijie, et al.
Published: (2025)
Boosting Self-Efficacy and Performance of Large Language Models via Verbal Efficacy Stimulations
by: Chen, Rui, et al.
Published: (2025)
by: Chen, Rui, et al.
Published: (2025)
Symbol-LLM: Towards Foundational Symbol-centric Interface For Large Language Models
by: Xu, Fangzhi, et al.
Published: (2023)
by: Xu, Fangzhi, et al.
Published: (2023)
CoLLEGe: Concept Embedding Generation for Large Language Models
by: Teehan, Ryan, et al.
Published: (2024)
by: Teehan, Ryan, et al.
Published: (2024)
Parallel Key-Value Cache Fusion for Position Invariant RAG
by: Oh, Philhoon, et al.
Published: (2025)
by: Oh, Philhoon, et al.
Published: (2025)
Abductive Inference in Retrieval-Augmented Language Models: Generating and Validating Missing Premises
by: Lin, Shiyin
Published: (2025)
by: Lin, Shiyin
Published: (2025)
Similar Items
-
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
by: Gu, Yifeng, et al.
Published: (2025) -
Stepwise Alignment for Constrained Language Model Policy Optimization
by: Wachi, Akifumi, et al.
Published: (2024) -
Reformulating KV Cache Eviction Problem for Long-Context LLM Inference
by: Mai, Tho, et al.
Published: (2026) -
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
by: Yang, Xintong, et al.
Published: (2026) -
EMO: Earth Mover Distance Optimization for Auto-Regressive Language Modeling
by: Ren, Siyu, et al.
Published: (2023)