Found in the Middle: How Language Models Use Long Contexts Better via Plug-and-Play Positional Encoding
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zhenyu, Chen, Runjin, Liu, Shiwei, Yao, Zhewei, Ruwase, Olatunji, Chen, Beidi, Wu, Xiaoxia, Wang, Zhangyang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LoCoCo: Dropping In Convolutions for Long Context Compression
by: Cai, Ruisi, et al.
Published: (2024)
by: Cai, Ruisi, et al.
Published: (2024)
Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer
by: Yao, Jinghan, et al.
Published: (2024)
by: Yao, Jinghan, et al.
Published: (2024)
SEAL: Steerable Reasoning Calibration of Large Language Models for Free
by: Chen, Runjin, et al.
Published: (2025)
by: Chen, Runjin, et al.
Published: (2025)
Found in the Middle: Calibrating Positional Attention Bias Improves Long Context Utilization
by: Hsieh, Cheng-Yu, et al.
Published: (2024)
by: Hsieh, Cheng-Yu, et al.
Published: (2024)
APE: Faster and Longer Context-Augmented Generation via Adaptive Parallel Encoding
by: Yang, Xinyu, et al.
Published: (2025)
by: Yang, Xinyu, et al.
Published: (2025)
An Efficient Recipe for Long Context Extension via Middle-Focused Positional Encoding
by: Wu, Tong, et al.
Published: (2024)
by: Wu, Tong, et al.
Published: (2024)
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
by: Gupta, Ahan, et al.
Published: (2026)
by: Gupta, Ahan, et al.
Published: (2026)
FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
by: Xia, Haojun, et al.
Published: (2024)
by: Xia, Haojun, et al.
Published: (2024)
LLaGA: Large Language and Graph Assistant
by: Chen, Runjin, et al.
Published: (2024)
by: Chen, Runjin, et al.
Published: (2024)
FastPersist: Accelerating Model Checkpointing in Deep Learning
by: Wang, Guanhua, et al.
Published: (2024)
by: Wang, Guanhua, et al.
Published: (2024)
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
by: Lian, Xinyu, et al.
Published: (2025)
by: Lian, Xinyu, et al.
Published: (2025)
ARM: A Learnable, Plug-and-Play Module for CLIP-based Open-vocabulary Semantic Segmentation
by: Liu, Ziquan, et al.
Published: (2025)
by: Liu, Ziquan, et al.
Published: (2025)
Domino: Eliminating Communication in LLM Training via Generic Tensor Slicing and Overlapping
by: Wang, Guanhua, et al.
Published: (2024)
by: Wang, Guanhua, et al.
Published: (2024)
Plug-and-Play Logit Fusion for Heterogeneous Pathology Foundation Models
by: Huang, Gexin, et al.
Published: (2026)
by: Huang, Gexin, et al.
Published: (2026)
Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference
by: Dong, Harry, et al.
Published: (2024)
by: Dong, Harry, et al.
Published: (2024)
GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
by: Zhao, Jiawei, et al.
Published: (2024)
by: Zhao, Jiawei, et al.
Published: (2024)
Long-Context Language Modeling with Parallel Context Encoding
by: Yen, Howard, et al.
Published: (2024)
by: Yen, Howard, et al.
Published: (2024)
Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding
by: Zhu, Jiajun, et al.
Published: (2025)
by: Zhu, Jiajun, et al.
Published: (2025)
Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences
by: Bekman, Stas, et al.
Published: (2025)
by: Bekman, Stas, et al.
Published: (2025)
LLMs Can Get "Brain Rot": A Pilot Study on Twitter/X
by: Xing, Shuo, et al.
Published: (2025)
by: Xing, Shuo, et al.
Published: (2025)
From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications
by: Jaiswal, Ajay, et al.
Published: (2024)
by: Jaiswal, Ajay, et al.
Published: (2024)
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning
by: Perin, Gabriel J., et al.
Published: (2025)
by: Perin, Gabriel J., et al.
Published: (2025)
GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?
by: Zhou, Yang, et al.
Published: (2025)
by: Zhou, Yang, et al.
Published: (2025)
Multiverse: Your Language Models Secretly Decide How to Parallelize and Merge Generation
by: Yang, Xinyu, et al.
Published: (2025)
by: Yang, Xinyu, et al.
Published: (2025)
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
by: Tanaka, Masahiro, et al.
Published: (2025)
by: Tanaka, Masahiro, et al.
Published: (2025)
Mojito: Motion Trajectory and Intensity Control for Video Generation
by: He, Xuehai, et al.
Published: (2024)
by: He, Xuehai, et al.
Published: (2024)
Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models
by: Chen, Jiaxing, et al.
Published: (2024)
by: Chen, Jiaxing, et al.
Published: (2024)
Plug-and-Play AMC: Context Is King in Training-Free, Open-Set Modulation with LLMs
by: Rostami, Mohammad, et al.
Published: (2025)
by: Rostami, Mohammad, et al.
Published: (2025)
Plug-and-Play Context Feature Reuse for Efficient Masked Generation
by: Liu, Xuejie, et al.
Published: (2025)
by: Liu, Xuejie, et al.
Published: (2025)
Position: The Turing-Completeness of Autoregressive Transformers Relies Heavily on Context Management
by: Cui, Guanyu, et al.
Published: (2026)
by: Cui, Guanyu, et al.
Published: (2026)
TempoFit: Plug-and-Play Layer-Wise Temporal KV Memory for Long-Horizon Vision-Language-Action Manipulation
by: Sun, Jun, et al.
Published: (2026)
by: Sun, Jun, et al.
Published: (2026)
Unrolling Plug-and-Play Network for Hyperspectral Unmixing
by: Zhao, Min, et al.
Published: (2024)
by: Zhao, Min, et al.
Published: (2024)
CIP: A Plug-and-Play Causal Prompting Framework for Mitigating Hallucinations under Long-Context Noise
by: Ma, Qingsen, et al.
Published: (2025)
by: Ma, Qingsen, et al.
Published: (2025)
Generation-Augmented Generation: A Plug-and-Play Framework for Private Knowledge Injection in Large Language Models
by: Li, Rongji, et al.
Published: (2026)
by: Li, Rongji, et al.
Published: (2026)
How do Language Models Generate Slang: A Systematic Comparison between Human and Machine-Generated Slang Usages
by: Wu, Siyang, et al.
Published: (2025)
by: Wu, Siyang, et al.
Published: (2025)
Chasing Better Deep Image Priors between Over- and Under-parameterization
by: Wu, Qiming, et al.
Published: (2024)
by: Wu, Qiming, et al.
Published: (2024)
V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding
by: Ge, Junqi, et al.
Published: (2024)
by: Ge, Junqi, et al.
Published: (2024)
PDR: A Plug-and-Play Positional Decay Framework for LLM Pre-training Data Detection
by: Liu, Jinhan, et al.
Published: (2026)
by: Liu, Jinhan, et al.
Published: (2026)
BetterDepth: Plug-and-Play Diffusion Refiner for Zero-Shot Monocular Depth Estimation
by: Zhang, Xiang, et al.
Published: (2024)
by: Zhang, Xiang, et al.
Published: (2024)
Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation
by: He, Zhenyu, et al.
Published: (2024)
by: He, Zhenyu, et al.
Published: (2024)
Similar Items
-
LoCoCo: Dropping In Convolutions for Long Context Compression
by: Cai, Ruisi, et al.
Published: (2024) -
Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer
by: Yao, Jinghan, et al.
Published: (2024) -
SEAL: Steerable Reasoning Calibration of Large Language Models for Free
by: Chen, Runjin, et al.
Published: (2025) -
Found in the Middle: Calibrating Positional Attention Bias Improves Long Context Utilization
by: Hsieh, Cheng-Yu, et al.
Published: (2024) -
APE: Faster and Longer Context-Augmented Generation via Adaptive Parallel Encoding
by: Yang, Xinyu, et al.
Published: (2025)