CItruS: Chunked Instruction-aware State Eviction for Long Sequence Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Bai, Yu, Zou, Xiyuan, Huang, Heyan, Chen, Sanxing, Rondeau, Marc-Antoine, Gao, Yang, Cheung, Jackie Chi Kit |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Identifying and Analyzing Performance-Critical Tokens in Large Language Models
by: Bai, Yu, et al.
Published: (2024)
by: Bai, Yu, et al.
Published: (2024)
Stochastic Chameleons: Irrelevant Context Hallucinations Reveal Class-Based (Mis)Generalization in LLMs
by: Cheng, Ziling, et al.
Published: (2025)
by: Cheng, Ziling, et al.
Published: (2025)
A Controlled Reevaluation of Coreference Resolution Models
by: Porada, Ian, et al.
Published: (2024)
by: Porada, Ian, et al.
Published: (2024)
Solving the Challenge Set without Solving the Task: On Winograd Schemas as a Test of Pronominal Coreference Resolution
by: Porada, Ian, et al.
Published: (2024)
by: Porada, Ian, et al.
Published: (2024)
PreSumm: Predicting Summarization Performance Without Summarizing
by: Koniaev, Steven, et al.
Published: (2025)
by: Koniaev, Steven, et al.
Published: (2025)
Mechanistic Understanding and Mitigation of Language Model Non-Factual Hallucinations
by: Yu, Lei, et al.
Published: (2024)
by: Yu, Lei, et al.
Published: (2024)
Improving the Calibration of Confidence Scores in Text Generation Using the Output Distribution's Characteristics
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2025)
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2025)
Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2026)
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2026)
$\texttt{COSMIC}$: Mutual Information for Task-Agnostic Summarization Evaluation
by: Darrin, Maxime, et al.
Published: (2024)
by: Darrin, Maxime, et al.
Published: (2024)
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
by: Chehbouni, Khaoula, et al.
Published: (2025)
by: Chehbouni, Khaoula, et al.
Published: (2025)
How Far Can In-Context Alignment Go? Exploring the State of In-Context Alignment
by: Huang, Heyan, et al.
Published: (2024)
by: Huang, Heyan, et al.
Published: (2024)
Challenges to Evaluating the Generalization of Coreference Resolution Models: A Measurement Modeling Perspective
by: Porada, Ian, et al.
Published: (2023)
by: Porada, Ian, et al.
Published: (2023)
Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation
by: Cheng, Ziling, et al.
Published: (2025)
by: Cheng, Ziling, et al.
Published: (2025)
Can Vision Language Models Be Adaptive in Mathematics Education? A Learner Model-based Rubric Study
by: Gao, Jie, et al.
Published: (2026)
by: Gao, Jie, et al.
Published: (2026)
Subtopic-aware View Sampling and Temporal Aggregation for Long-form Document Matching
by: Zhou, Youchao, et al.
Published: (2024)
by: Zhou, Youchao, et al.
Published: (2024)
$(RSA)^2$: A Rhetorical-Strategy-Aware Rational Speech Act Framework for Figurative Language Understanding
by: Piano, Cesare Spinoso-Di, et al.
Published: (2025)
by: Piano, Cesare Spinoso-Di, et al.
Published: (2025)
FlashBack:Efficient Retrieval-Augmented Language Modeling for Long Context Inference
by: Liu, Runheng, et al.
Published: (2024)
by: Liu, Runheng, et al.
Published: (2024)
LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long Reasoning
by: Zhang, Haoyue, et al.
Published: (2025)
by: Zhang, Haoyue, et al.
Published: (2025)
Real-time Factuality Assessment from Adversarial Feedback
by: Chen, Sanxing, et al.
Published: (2024)
by: Chen, Sanxing, et al.
Published: (2024)
ECBD: Evidence-Centered Benchmark Design for NLP
by: Liu, Yu Lu, et al.
Published: (2024)
by: Liu, Yu Lu, et al.
Published: (2024)
Testing the Assumptions of Active Learning for Translation Tasks with Few Samples
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2026)
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2026)
Error Diversity Matters: An Error-Resistant Ensemble Method for Unsupervised Dependency Parsing
by: Shayegh, Behzad, et al.
Published: (2024)
by: Shayegh, Behzad, et al.
Published: (2024)
ChatShop: Interactive Information Seeking with Language Agents
by: Chen, Sanxing, et al.
Published: (2024)
by: Chen, Sanxing, et al.
Published: (2024)
To Trust or Not to Trust? Enhancing Large Language Models' Situated Faithfulness to External Contexts
by: Huang, Yukun, et al.
Published: (2024)
by: Huang, Yukun, et al.
Published: (2024)
Locret: Enhancing Eviction in Long-Context LLM Inference with Trained Retaining Heads on Consumer-Grade Devices
by: Huang, Yuxiang, et al.
Published: (2024)
by: Huang, Yuxiang, et al.
Published: (2024)
Building Knowledge-Grounded Dialogue Systems with Graph-Based Semantic Modeling
by: Yang, Yizhe, et al.
Published: (2022)
by: Yang, Yizhe, et al.
Published: (2022)
Word Matters: What Influences Domain Adaptation in Summarization?
by: Li, Yinghao, et al.
Published: (2024)
by: Li, Yinghao, et al.
Published: (2024)
Beyond Chunk-Local Extraction: Cross-Chunk Graph Augmentation for GraphRAG
by: Zhang, Jiaming, et al.
Published: (2026)
by: Zhang, Jiaming, et al.
Published: (2026)
Reformulating KV Cache Eviction Problem for Long-Context LLM Inference
by: Mai, Tho, et al.
Published: (2026)
by: Mai, Tho, et al.
Published: (2026)
SynthEHR-Eviction: Enhancing Eviction SDoH Detection with LLM-Augmented Synthetic EHR Data
by: Yao, Zonghai, et al.
Published: (2025)
by: Yao, Zonghai, et al.
Published: (2025)
R1-Compress: Long Chain-of-Thought Compression via Chunk Compression and Search
by: Wang, Yibo, et al.
Published: (2025)
by: Wang, Yibo, et al.
Published: (2025)
Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization
by: Duan, Shaohua, et al.
Published: (2025)
by: Duan, Shaohua, et al.
Published: (2025)
Biological Sequence with Language Model Prompting: A Survey
by: Jiang, Jiyue, et al.
Published: (2025)
by: Jiang, Jiyue, et al.
Published: (2025)
Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models
by: Günther, Michael, et al.
Published: (2024)
by: Günther, Michael, et al.
Published: (2024)
CAMI: A Counselor Agent Supporting Motivational Interviewing through State Inference and Topic Exploration
by: Yang, Yizhe, et al.
Published: (2025)
by: Yang, Yizhe, et al.
Published: (2025)
Does This Summary Answer My Question? Modeling Query-Focused Summary Readers with Rational Speech Acts
by: Piano, Cesare Spinoso-Di, et al.
Published: (2024)
by: Piano, Cesare Spinoso-Di, et al.
Published: (2024)
On the Loss of Context-awareness in General Instruction Fine-tuning
by: Wang, Yihan, et al.
Published: (2024)
by: Wang, Yihan, et al.
Published: (2024)
RetentiveKV: State-Space Memory for Uncertainty-Aware Multimodal KV Cache Eviction
by: Liu, Sihao, et al.
Published: (2026)
by: Liu, Sihao, et al.
Published: (2026)
RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence Modeling
by: Wei, Xiuying, et al.
Published: (2025)
by: Wei, Xiuying, et al.
Published: (2025)
SABlock: Semantic-Aware KV Cache Eviction with Adaptive Compression Block Size
by: Chen, Jinhan, et al.
Published: (2025)
by: Chen, Jinhan, et al.
Published: (2025)
Similar Items
-
Identifying and Analyzing Performance-Critical Tokens in Large Language Models
by: Bai, Yu, et al.
Published: (2024) -
Stochastic Chameleons: Irrelevant Context Hallucinations Reveal Class-Based (Mis)Generalization in LLMs
by: Cheng, Ziling, et al.
Published: (2025) -
A Controlled Reevaluation of Coreference Resolution Models
by: Porada, Ian, et al.
Published: (2024) -
Solving the Challenge Set without Solving the Task: On Winograd Schemas as a Test of Pronominal Coreference Resolution
by: Porada, Ian, et al.
Published: (2024) -
PreSumm: Predicting Summarization Performance Without Summarizing
by: Koniaev, Steven, et al.
Published: (2025)