InfiniPot: Infinite Context Processing on Memory-Constrained LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Minsoo, Shim, Kyuhong, Choi, Jungwook, Chang, Simyung |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding
by: Kim, Minsoo, et al.
Published: (2025)
by: Kim, Minsoo, et al.
Published: (2025)
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention
by: Munkhdalai, Tsendsuren, et al.
Published: (2024)
by: Munkhdalai, Tsendsuren, et al.
Published: (2024)
Preserving Pre-trained Representation Space: On Effectiveness of Prefix-tuning for Large Multi-modal Models
by: Kim, Donghoon, et al.
Published: (2024)
by: Kim, Donghoon, et al.
Published: (2024)
Feedback Adaptation for Retrieval-Augmented Generation
by: Bang, Jihwan, et al.
Published: (2026)
by: Bang, Jihwan, et al.
Published: (2026)
Crayon: Customized On-Device LLM via Instant Adapter Blending and Edge-Server Hybrid Inference
by: Bang, Jihwan, et al.
Published: (2024)
by: Bang, Jihwan, et al.
Published: (2024)
Chain-of-Rank: Enhancing Large Language Models for Domain-Specific RAG in Edge Device
by: Lee, Juntae, et al.
Published: (2025)
by: Lee, Juntae, et al.
Published: (2025)
Human-inspired Episodic Memory for Infinite Context LLMs
by: Fountas, Zafeirios, et al.
Published: (2024)
by: Fountas, Zafeirios, et al.
Published: (2024)
Probing the Limits of Compressive Memory: A Study of Infini-Attention in Small-Scale Pretraining
by: Huang, Ruizhe, et al.
Published: (2025)
by: Huang, Ruizhe, et al.
Published: (2025)
Unlocking Transfer Learning for Open-World Few-Shot Recognition
by: Kim, Byeonggeun, et al.
Published: (2024)
by: Kim, Byeonggeun, et al.
Published: (2024)
General-Purpose Retrieval-Enhanced Medical Prediction Model Using Near-Infinite History
by: Kim, Junu, et al.
Published: (2023)
by: Kim, Junu, et al.
Published: (2023)
Learning When to Attend: Conditional Memory Access for Long-Context LLMs
by: Choudhary, Sakshi, et al.
Published: (2026)
by: Choudhary, Sakshi, et al.
Published: (2026)
HMT: Hierarchical Memory Transformer for Efficient Long Context Language Processing
by: He, Zifan, et al.
Published: (2024)
by: He, Zifan, et al.
Published: (2024)
Learning Contextual Retrieval for Robust Conversational Search
by: Yang, Seunghan, et al.
Published: (2025)
by: Yang, Seunghan, et al.
Published: (2025)
Compressed Context Memory For Online Language Model Interaction
by: Kim, Jang-Hyun, et al.
Published: (2023)
by: Kim, Jang-Hyun, et al.
Published: (2023)
Improving Conversational Abilities of Quantized Large Language Models via Direct Preference Alignment
by: Lee, Janghwan, et al.
Published: (2024)
by: Lee, Janghwan, et al.
Published: (2024)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
Online Adaptation of Language Models with a Memory of Amortized Contexts
by: Tack, Jihoon, et al.
Published: (2024)
by: Tack, Jihoon, et al.
Published: (2024)
Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models
by: Kim, Donghoon, et al.
Published: (2025)
by: Kim, Donghoon, et al.
Published: (2025)
Adapting LLMs for Efficient Context Processing through Soft Prompt Compression
by: Wang, Cangqing, et al.
Published: (2024)
by: Wang, Cangqing, et al.
Published: (2024)
Do MLLMs Capture How Interfaces Guide User Behavior? A Benchmark for Multimodal UI/UX Design Understanding
by: Jeon, Jaehyun, et al.
Published: (2025)
by: Jeon, Jaehyun, et al.
Published: (2025)
Intrinsic Fingerprint of LLMs: Continue Training is NOT All You Need to Steal A Model!
by: Yoon, Do-hyeon, et al.
Published: (2025)
by: Yoon, Do-hyeon, et al.
Published: (2025)
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU
by: Lee, Heejun, et al.
Published: (2025)
by: Lee, Heejun, et al.
Published: (2025)
Intrinsic Entropy of Context Length Scaling in LLMs
by: Shi, Jingzhe, et al.
Published: (2025)
by: Shi, Jingzhe, et al.
Published: (2025)
Enhancing Computation Efficiency in Large Language Models through Weight and Activation Quantization
by: Lee, Janghwan, et al.
Published: (2023)
by: Lee, Janghwan, et al.
Published: (2023)
Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
by: Bansal, Rachit, et al.
Published: (2025)
by: Bansal, Rachit, et al.
Published: (2025)
Guiding LLMs The Right Way: Fast, Non-Invasive Constrained Generation
by: Beurer-Kellner, Luca, et al.
Published: (2024)
by: Beurer-Kellner, Luca, et al.
Published: (2024)
EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration
by: Nie, Allen, et al.
Published: (2024)
by: Nie, Allen, et al.
Published: (2024)
LLMem: Estimating GPU Memory Usage for Fine-Tuning Pre-Trained LLMs
by: Kim, Taeho, et al.
Published: (2024)
by: Kim, Taeho, et al.
Published: (2024)
PoeTone: A Framework for Constrained Generation of Structured Chinese Songci with LLMs
by: Qu, Zhan, et al.
Published: (2025)
by: Qu, Zhan, et al.
Published: (2025)
Toward Conversational Agents with Context and Time Sensitive Long-term Memory
by: Alonso, Nick, et al.
Published: (2024)
by: Alonso, Nick, et al.
Published: (2024)
Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts
by: Sivtsov, Danil, et al.
Published: (2025)
by: Sivtsov, Danil, et al.
Published: (2025)
Controllable Generation via Locally Constrained Resampling
by: Ahmed, Kareem, et al.
Published: (2024)
by: Ahmed, Kareem, et al.
Published: (2024)
A Controlled Study on Long Context Extension and Generalization in LLMs
by: Lu, Yi, et al.
Published: (2024)
by: Lu, Yi, et al.
Published: (2024)
Long-Short Alignment for Effective Long-Context Modeling in LLMs
by: Du, Tianqi, et al.
Published: (2025)
by: Du, Tianqi, et al.
Published: (2025)
Training-Inference Consistent Segmented Execution for Long-Context LLMs
by: Shang, Xianpeng, et al.
Published: (2026)
by: Shang, Xianpeng, et al.
Published: (2026)
Training LLMs for Multi-Step Tool Orchestration with Constrained Data Synthesis and Graduated Rewards
by: Jiayang, Cheng, et al.
Published: (2026)
by: Jiayang, Cheng, et al.
Published: (2026)
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference
by: Li, Weizhuo, et al.
Published: (2024)
by: Li, Weizhuo, et al.
Published: (2024)
ZERA: Zero-init Instruction Evolving Refinement Agent -- From Zero Instructions to Structured Prompts via Principle-based Optimization
by: Yi, Seungyoun, et al.
Published: (2025)
by: Yi, Seungyoun, et al.
Published: (2025)
CrEst: Credibility Estimation for Contexts in LLMs via Weak Supervision
by: Adila, Dyah, et al.
Published: (2025)
by: Adila, Dyah, et al.
Published: (2025)
Similar Items
-
InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding
by: Kim, Minsoo, et al.
Published: (2025) -
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
by: Kim, Eunji, et al.
Published: (2024) -
Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention
by: Munkhdalai, Tsendsuren, et al.
Published: (2024) -
Preserving Pre-trained Representation Space: On Effectiveness of Prefix-tuning for Large Multi-modal Models
by: Kim, Donghoon, et al.
Published: (2024) -
Feedback Adaptation for Retrieval-Augmented Generation
by: Bang, Jihwan, et al.
Published: (2026)