Structured Packing in LLM Training Improves Long Context Utilization
Fuente:
arXiv
Saved in:
| Main Authors: | Staniszewski, Konrad, Tworkowski, Szymon, Jaszczur, Sebastian, Zhao, Yu, Michalewski, Henryk, Kuciński, Łukasz, Miłoś, Piotr |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Analysing The Impact of Sequence Composition on Language Model Pre-Training
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
Catalytic Role Of Noise And Necessity Of Inductive Biases In The Emergence Of Compositional Communication
by: Kuciński, Łukasz, et al.
Published: (2021)
by: Kuciński, Łukasz, et al.
Published: (2021)
There and Back Again: On the relation between Noise and Image Inversions in Diffusion Models
by: Staniszewski, Łukasz, et al.
Published: (2024)
by: Staniszewski, Łukasz, et al.
Published: (2024)
Beyond Lines and Circles: Unveiling the Geometric Reasoning Gap in Large Language Models
by: Mouselinos, Spyridon, et al.
Published: (2024)
by: Mouselinos, Spyridon, et al.
Published: (2024)
KV Cache Transform Coding for Compact Storage in LLM Inference
by: Staniszewski, Konrad, et al.
Published: (2025)
by: Staniszewski, Konrad, et al.
Published: (2025)
Magnushammer: A Transformer-Based Approach to Premise Selection
by: Mikuła, Maciej, et al.
Published: (2023)
by: Mikuła, Maciej, et al.
Published: (2023)
Inference-Time Hyper-Scaling with KV Cache Compression
by: Łańcucki, Adrian, et al.
Published: (2025)
by: Łańcucki, Adrian, et al.
Published: (2025)
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
by: Pióro, Maciej, et al.
Published: (2024)
by: Pióro, Maciej, et al.
Published: (2024)
Off-Policy Correction For Multi-Agent Reinforcement Learning
by: Zawalski, Michał, et al.
Published: (2021)
by: Zawalski, Michał, et al.
Published: (2021)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
by: Ludziejewski, Jan, et al.
Published: (2025)
by: Ludziejewski, Jan, et al.
Published: (2025)
Found in the Middle: Calibrating Positional Attention Bias Improves Long Context Utilization
by: Hsieh, Cheng-Yu, et al.
Published: (2024)
by: Hsieh, Cheng-Yu, et al.
Published: (2024)
Squeezed Attention: Accelerating Long Context Length LLM Inference
by: Hooper, Coleman, et al.
Published: (2024)
by: Hooper, Coleman, et al.
Published: (2024)
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
by: Gu, Zhuohan, et al.
Published: (2024)
by: Gu, Zhuohan, et al.
Published: (2024)
Lightweight Latent Verifiers for Efficient Meta-Generation Strategies
by: Piotrowski, Bartosz, et al.
Published: (2025)
by: Piotrowski, Bartosz, et al.
Published: (2025)
Mixture of Tokens: Continuous MoE through Cross-Example Aggregation
by: Antoniak, Szymon, et al.
Published: (2023)
by: Antoniak, Szymon, et al.
Published: (2023)
Scaling Laws for Fine-Grained Mixture of Experts
by: Krajewski, Jakub, et al.
Published: (2024)
by: Krajewski, Jakub, et al.
Published: (2024)
ACC: Compiling Agent Trajectories for Long-Context Training
by: Su, Qisheng, et al.
Published: (2026)
by: Su, Qisheng, et al.
Published: (2026)
Make Your LLM Fully Utilize the Context
by: An, Shengnan, et al.
Published: (2024)
by: An, Shengnan, et al.
Published: (2024)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
Long Context Pre-Training with Lighthouse Attention
by: Peng, Bowen, et al.
Published: (2026)
by: Peng, Bowen, et al.
Published: (2026)
Training-free Context-adaptive Attention for Efficient Long Context Modeling
by: You, Zeng, et al.
Published: (2025)
by: You, Zeng, et al.
Published: (2025)
Grounding Data Science Code Generation with Input-Output Specifications
by: Wen, Yeming, et al.
Published: (2024)
by: Wen, Yeming, et al.
Published: (2024)
Locret: Enhancing Eviction in Long-Context LLM Inference with Trained Retaining Heads on Consumer-Grade Devices
by: Huang, Yuxiang, et al.
Published: (2024)
by: Huang, Yuxiang, et al.
Published: (2024)
DoubleDipper: Improving Long-Context LLMs via Context Recycling
by: Cattan, Arie, et al.
Published: (2024)
by: Cattan, Arie, et al.
Published: (2024)
SCORE: Specificity, Context Utilization, Robustness, and Relevance for Reference-Free LLM Evaluation
by: Shomee, Homaira Huda, et al.
Published: (2026)
by: Shomee, Homaira Huda, et al.
Published: (2026)
A Little Goes a Long Way: Efficient Long Context Training and Inference with Partial Contexts
by: Ge, Suyu, et al.
Published: (2024)
by: Ge, Suyu, et al.
Published: (2024)
Sentinel: Decoding Context Utilization via Attention Probing for Efficient LLM Context Compression
by: Zhang, Yong, et al.
Published: (2025)
by: Zhang, Yong, et al.
Published: (2025)
Mindscape-Aware Retrieval Augmented Generation for Improved Long Context Understanding
by: Li, Yuqing, et al.
Published: (2025)
by: Li, Yuqing, et al.
Published: (2025)
Shopping Companion: Benchmarking and Training LLM Agents for Long-Horizon Preference-Grounded E-Commerce Tasks
by: Yu, Zijian, et al.
Published: (2026)
by: Yu, Zijian, et al.
Published: (2026)
Lag-Relative Sparse Attention In Long Context Training
by: Liang, Manlai, et al.
Published: (2025)
by: Liang, Manlai, et al.
Published: (2025)
MAML-en-LLM: Model Agnostic Meta-Training of LLMs for Improved In-Context Learning
by: Sinha, Sanchit, et al.
Published: (2024)
by: Sinha, Sanchit, et al.
Published: (2024)
Insights into LLM Long-Context Failures: When Transformers Know but Don't Tell
by: Lu, Taiming, et al.
Published: (2024)
by: Lu, Taiming, et al.
Published: (2024)
LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization
by: Liu, Jiarui, et al.
Published: (2025)
by: Liu, Jiarui, et al.
Published: (2025)
CompLLM: Compression for Long Context Q&A
by: Berton, Gabriele, et al.
Published: (2025)
by: Berton, Gabriele, et al.
Published: (2025)
Perception Compressor: A Training-Free Prompt Compression Framework in Long Context Scenarios
by: Tang, Jiwei, et al.
Published: (2024)
by: Tang, Jiwei, et al.
Published: (2024)
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference
by: Li, Weizhuo, et al.
Published: (2024)
by: Li, Weizhuo, et al.
Published: (2024)
Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing
by: Long, Lingkun, et al.
Published: (2026)
by: Long, Lingkun, et al.
Published: (2026)
Training-Free Long-Context Scaling of Large Language Models
by: An, Chenxin, et al.
Published: (2024)
by: An, Chenxin, et al.
Published: (2024)
Threshold Filtering Packing for Supervised Fine-Tuning: Training Related Samples within Packs
by: Dong, Jiancheng, et al.
Published: (2024)
by: Dong, Jiancheng, et al.
Published: (2024)
From Similarity to Structure: Training-free LLM Context Compression with Hybrid Graph Priors
by: Zhou, Yitian, et al.
Published: (2026)
by: Zhou, Yitian, et al.
Published: (2026)
Similar Items
-
Analysing The Impact of Sequence Composition on Language Model Pre-Training
by: Zhao, Yu, et al.
Published: (2024) -
Catalytic Role Of Noise And Necessity Of Inductive Biases In The Emergence Of Compositional Communication
by: Kuciński, Łukasz, et al.
Published: (2021) -
There and Back Again: On the relation between Noise and Image Inversions in Diffusion Models
by: Staniszewski, Łukasz, et al.
Published: (2024) -
Beyond Lines and Circles: Unveiling the Geometric Reasoning Gap in Large Language Models
by: Mouselinos, Spyridon, et al.
Published: (2024) -
KV Cache Transform Coding for Compact Storage in LLM Inference
by: Staniszewski, Konrad, et al.
Published: (2025)