SnapKV: LLM Knows What You are Looking for Before Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yuhong, Huang, Yingbing, Yang, Bowen, Venkitesh, Bharat, Locatelli, Acyr, Ye, Hanchen, Cai, Tianle, Lewis, Patrick, Chen, Deming |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SnapKV
by: Anonymous
Published: (2025)
by: Anonymous
Published: (2025)
Rope to Nope and Back Again: A New Hybrid Attention Strategy
by: Yang, Bowen, et al.
Published: (2025)
by: Yang, Bowen, et al.
Published: (2025)
New Solutions on LLM Acceleration, Optimization, and Application
by: Huang, Yingbing, et al.
Published: (2024)
by: Huang, Yingbing, et al.
Published: (2024)
KVCapsule: Efficient Sequential KV Cache Compression for Vision-Language Models with Asymmetric Redundancy
by: Huang, Yingbing, et al.
Published: (2026)
by: Huang, Yingbing, et al.
Published: (2026)
Look Before You Leap: Autonomous Exploration for LLM Agents
by: Ye, Ziang, et al.
Published: (2026)
by: Ye, Ziang, et al.
Published: (2026)
JAM: Controllable and Responsible Text Generation via Causal Reasoning and Latent Vector Manipulation
by: Huang, Yingbing, et al.
Published: (2025)
by: Huang, Yingbing, et al.
Published: (2025)
CompressKV: Semantic Retrieval Heads Know What Tokens are Not Important Before Generation
by: Lin, Xiaolin, et al.
Published: (2025)
by: Lin, Xiaolin, et al.
Published: (2025)
StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMs
by: Ye, Hanchen, et al.
Published: (2025)
by: Ye, Hanchen, et al.
Published: (2025)
BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts
by: Zhang, Qizhen, et al.
Published: (2024)
by: Zhang, Qizhen, et al.
Published: (2024)
Efficient Benchmarking Is Just Feature Selection and Multiple Regression
by: Bowyer, Sam, et al.
Published: (2026)
by: Bowyer, Sam, et al.
Published: (2026)
Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
by: Cai, Tianle, et al.
Published: (2024)
by: Cai, Tianle, et al.
Published: (2024)
Polynomials, Divided Differences, and Codes
by: Venkitesh, S.
Published: (2024)
by: Venkitesh, S.
Published: (2024)
Know You Before You Speak: User-State Modeling for LLM Personalization in Multi-Turn Conversation
by: Luo, Jiani, et al.
Published: (2026)
by: Luo, Jiani, et al.
Published: (2026)
When Magnetic Field Lines Stretch, Snap, and Expand: A New Look at Solar Flares with L-maps
by: Kazachenko, Maria D., et al.
Published: (2025)
by: Kazachenko, Maria D., et al.
Published: (2025)
Visually Dehallucinative Instruction Generation: Know What You Don't Know
by: Cha, Sungguk, et al.
Published: (2024)
by: Cha, Sungguk, et al.
Published: (2024)
Towards More Economical Context-Augmented LLM Generation by Reusing Stored KV Cache
by: Li, Hanchen, et al.
Published: (2025)
by: Li, Hanchen, et al.
Published: (2025)
Look Before You Leap: Enhancing Attention and Vigilance Regarding Harmful Content with GuidelineLLM
by: Zhang, Shaoqing, et al.
Published: (2024)
by: Zhang, Shaoqing, et al.
Published: (2024)
Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts
by: Gritsch, Nikolas, et al.
Published: (2024)
by: Gritsch, Nikolas, et al.
Published: (2024)
Understanding Likelihood Over-optimisation in Direct Alignment Algorithms
by: Shi, Zhengyan, et al.
Published: (2024)
by: Shi, Zhengyan, et al.
Published: (2024)
"Don't Look, But I Know You Do": Norms and Observer Effects in Shared LLM Accounts
by: Song, Ji Eun, et al.
Published: (2026)
by: Song, Ji Eun, et al.
Published: (2026)
O que os analistas pensam sobre a homossexualidade?
by: Acyr Maya
Published: (2007)
by: Acyr Maya
Published: (2007)
MolSnap: Snap-Fast Molecular Generation with Latent Variational Mean Flow
by: Ahamed, Md Atik, et al.
Published: (2025)
by: Ahamed, Md Atik, et al.
Published: (2025)
Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models
by: Huang, Yuheng, et al.
Published: (2023)
by: Huang, Yuheng, et al.
Published: (2023)
Do You Know What Your Mission Is?
by: Balas, Janet L.
Published: (2007)
by: Balas, Janet L.
Published: (2007)
Random Reed-Solomon Codes are List Recoverable with Optimal List Size
by: Doron, Dean, et al.
Published: (2024)
by: Doron, Dean, et al.
Published: (2024)
List Recoverable Codes: The Good, the Bad, and the Unknown (hopefully not Ugly)
by: Resch, Nicolas, et al.
Published: (2025)
by: Resch, Nicolas, et al.
Published: (2025)
Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive Reasoning
by: Li, Yian, et al.
Published: (2024)
by: Li, Yian, et al.
Published: (2024)
Know What You Know: Metacognitive Entropy Calibration for Verifiable RL Reasoning
by: Zhao, Qiannian, et al.
Published: (2026)
by: Zhao, Qiannian, et al.
Published: (2026)
Know What You Don't Know: Uncertainty Calibration of Process Reward Models
by: Park, Young-Jin, et al.
Published: (2025)
by: Park, Young-Jin, et al.
Published: (2025)
Know What You Don't Know: Selective Prediction for Early Exit DNNs
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
What You Shouldn't Know About Quantum Computers
by: Ferrie, Chris
Published: (2024)
by: Ferrie, Chris
Published: (2024)
Open Access: What You Need to Know Now
by: Crawford, Walt
Published: (2011)
by: Crawford, Walt
Published: (2011)
Getting Ready for RDA: What You Need to Know
by: Hart, Amy
Published: (2010)
by: Hart, Amy
Published: (2010)
Knowing Before Saying: LLM Representations Encode Information About Chain-of-Thought Success Before Completion
by: Afzal, Anum, et al.
Published: (2025)
by: Afzal, Anum, et al.
Published: (2025)
Look-Around Before You Leap: High-Frequency Injected Transformer for Image Restoration
by: Zhou, Shihao, et al.
Published: (2024)
by: Zhou, Shihao, et al.
Published: (2024)
Subgraph Extraction-based Feedback-guided Iterative Scheduling for HLS
by: Ye, Hanchen, et al.
Published: (2024)
by: Ye, Hanchen, et al.
Published: (2024)
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement
by: Ding, Peng, et al.
Published: (2025)
by: Ding, Peng, et al.
Published: (2025)
Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation
by: Wanyan, Yuyang, et al.
Published: (2025)
by: Wanyan, Yuyang, et al.
Published: (2025)
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
by: Cai, Will, et al.
Published: (2025)
by: Cai, Will, et al.
Published: (2025)
What Are You Doing? A Closer Look at Controllable Human Video Generation
by: Bugliarello, Emanuele, et al.
Published: (2025)
by: Bugliarello, Emanuele, et al.
Published: (2025)
Similar Items
-
SnapKV
by: Anonymous
Published: (2025) -
Rope to Nope and Back Again: A New Hybrid Attention Strategy
by: Yang, Bowen, et al.
Published: (2025) -
New Solutions on LLM Acceleration, Optimization, and Application
by: Huang, Yingbing, et al.
Published: (2024) -
KVCapsule: Efficient Sequential KV Cache Compression for Vision-Language Models with Asymmetric Redundancy
by: Huang, Yingbing, et al.
Published: (2026) -
Look Before You Leap: Autonomous Exploration for LLM Agents
by: Ye, Ziang, et al.
Published: (2026)