Why Does the Effective Context Length of LLMs Fall Short?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | An, Chenxin, Zhang, Jun, Zhong, Ming, Li, Lei, Gong, Shansan, Luo, Yao, Xu, Jingjing, Kong, Lingpeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Training-Free Long-Context Scaling of Large Language Models
von: An, Chenxin, et al.
Veröffentlicht: (2024)
von: An, Chenxin, et al.
Veröffentlicht: (2024)
GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models
von: Li, Mukai, et al.
Veröffentlicht: (2024)
von: Li, Mukai, et al.
Veröffentlicht: (2024)
DiffuCoder: Understanding and Improving Masked Diffusion Models for Code Generation
von: Gong, Shansan, et al.
Veröffentlicht: (2025)
von: Gong, Shansan, et al.
Veröffentlicht: (2025)
Beyond Autoregression: Discrete Diffusion for Complex Reasoning and Planning
von: Ye, Jiacheng, et al.
Veröffentlicht: (2024)
von: Ye, Jiacheng, et al.
Veröffentlicht: (2024)
Scaling Diffusion Language Models via Adaptation from Autoregressive Models
von: Gong, Shansan, et al.
Veröffentlicht: (2024)
von: Gong, Shansan, et al.
Veröffentlicht: (2024)
BBA: Bi-Modal Behavioral Alignment for Reasoning with Large Vision-Language Models
von: Zhao, Xueliang, et al.
Veröffentlicht: (2024)
von: Zhao, Xueliang, et al.
Veröffentlicht: (2024)
Temporal Reasoning Transfer from Text to Video
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
Reasoning Does Not Necessarily Improve Role-Playing Ability
von: Feng, Xiachong, et al.
Veröffentlicht: (2025)
von: Feng, Xiachong, et al.
Veröffentlicht: (2025)
Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
von: Ye, Jiacheng, et al.
Veröffentlicht: (2025)
von: Ye, Jiacheng, et al.
Veröffentlicht: (2025)
ParallelComp: Parallel Long-Context Compressor for Length Extrapolation
von: Xiong, Jing, et al.
Veröffentlicht: (2025)
von: Xiong, Jing, et al.
Veröffentlicht: (2025)
Dream-Coder 7B: An Open Diffusion Language Model for Code
von: Xie, Zhihui, et al.
Veröffentlicht: (2025)
von: Xie, Zhihui, et al.
Veröffentlicht: (2025)
DreamOn: Diffusion Language Models For Code Infilling Beyond Fixed-size Canvas
von: Wu, Zirui, et al.
Veröffentlicht: (2026)
von: Wu, Zirui, et al.
Veröffentlicht: (2026)
Intrinsic Entropy of Context Length Scaling in LLMs
von: Shi, Jingzhe, et al.
Veröffentlicht: (2025)
von: Shi, Jingzhe, et al.
Veröffentlicht: (2025)
Multilingual Machine Translation with Large Language Models: Empirical Results and Analysis
von: Zhu, Wenhao, et al.
Veröffentlicht: (2023)
von: Zhu, Wenhao, et al.
Veröffentlicht: (2023)
GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
von: Li, Qintong, et al.
Veröffentlicht: (2024)
von: Li, Qintong, et al.
Veröffentlicht: (2024)
Long-Short Alignment for Effective Long-Context Modeling in LLMs
von: Du, Tianqi, et al.
Veröffentlicht: (2025)
von: Du, Tianqi, et al.
Veröffentlicht: (2025)
Haste Makes Waste: Evaluating Planning Abilities of LLMs for Efficient and Feasible Multitasking with Time Constraints Between Actions
von: Wu, Zirui, et al.
Veröffentlicht: (2025)
von: Wu, Zirui, et al.
Veröffentlicht: (2025)
Effective Length Extrapolation via Dimension-Wise Positional Embeddings Manipulation
von: Lu, Yi, et al.
Veröffentlicht: (2025)
von: Lu, Yi, et al.
Veröffentlicht: (2025)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
von: Jiang, Botian, et al.
Veröffentlicht: (2024)
von: Jiang, Botian, et al.
Veröffentlicht: (2024)
Diffusion of Thoughts: Chain-of-Thought Reasoning in Diffusion Language Models
von: Ye, Jiacheng, et al.
Veröffentlicht: (2024)
von: Ye, Jiacheng, et al.
Veröffentlicht: (2024)
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
von: Kim, Jeonghye, et al.
Veröffentlicht: (2026)
von: Kim, Jeonghye, et al.
Veröffentlicht: (2026)
A Reparameterized Discrete Diffusion Model for Text Generation
von: Zheng, Lin, et al.
Veröffentlicht: (2023)
von: Zheng, Lin, et al.
Veröffentlicht: (2023)
Why Does New Knowledge Create Messy Ripple Effects in LLMs?
von: Qin, Jiaxin, et al.
Veröffentlicht: (2024)
von: Qin, Jiaxin, et al.
Veröffentlicht: (2024)
Understanding the RoPE Extensions of Long-Context LLMs: An Attention Perspective
von: Zhong, Meizhi, et al.
Veröffentlicht: (2024)
von: Zhong, Meizhi, et al.
Veröffentlicht: (2024)
Length Controlled Generation for Black-box LLMs
von: Gu, Yuxuan, et al.
Veröffentlicht: (2024)
von: Gu, Yuxuan, et al.
Veröffentlicht: (2024)
Teaching Language Models to Critique via Reinforcement Learning
von: Xie, Zhihui, et al.
Veröffentlicht: (2025)
von: Xie, Zhihui, et al.
Veröffentlicht: (2025)
Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks
von: Li, Qintong, et al.
Veröffentlicht: (2023)
von: Li, Qintong, et al.
Veröffentlicht: (2023)
Activation-aware Probe-Query: Effective Key-Value Retrieval for Long-Context LLMs Inference
von: Xiao, Qingfa, et al.
Veröffentlicht: (2025)
von: Xiao, Qingfa, et al.
Veröffentlicht: (2025)
Does RAG Really Perform Bad For Long-Context Processing?
von: Luo, Kun, et al.
Veröffentlicht: (2025)
von: Luo, Kun, et al.
Veröffentlicht: (2025)
Where Does Long-Context Supervision Actually Go? Effective-Context Exposure Balancing
von: Zhu, Jinchang, et al.
Veröffentlicht: (2026)
von: Zhu, Jinchang, et al.
Veröffentlicht: (2026)
Self-Consistency Falls Short! The Adverse Effects of Positional Bias on Long-Context Problems
von: Byerly, Adam, et al.
Veröffentlicht: (2024)
von: Byerly, Adam, et al.
Veröffentlicht: (2024)
PromptCoT 2.0: Scaling Prompt Synthesis for Large Language Model Reasoning
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
Linguistic Frameworks Go Toe-to-Toe at Neuro-Symbolic Language Modeling
von: Prange, Jakob, et al.
Veröffentlicht: (2021)
von: Prange, Jakob, et al.
Veröffentlicht: (2021)
Effective In-Context Example Selection through Data Compression
von: Sun, Zhongxiang, et al.
Veröffentlicht: (2024)
von: Sun, Zhongxiang, et al.
Veröffentlicht: (2024)
Are LLMs Effective Backbones for Fine-tuning? An Experimental Investigation of Supervised LLMs on Chinese Short Text Matching
von: Liu, Shulin, et al.
Veröffentlicht: (2024)
von: Liu, Shulin, et al.
Veröffentlicht: (2024)
Structured Prompt Language: Declarative Context Management for LLMs
von: Gong, Wen G.
Veröffentlicht: (2026)
von: Gong, Wen G.
Veröffentlicht: (2026)
How Well Do LLMs Handle Cantonese? Benchmarking Cantonese Capabilities of Large Language Models
von: Jiang, Jiyue, et al.
Veröffentlicht: (2024)
von: Jiang, Jiyue, et al.
Veröffentlicht: (2024)
Base of RoPE Bounds Context Length
von: Men, Xin, et al.
Veröffentlicht: (2024)
von: Men, Xin, et al.
Veröffentlicht: (2024)
Forewarned is Forearmed: Leveraging LLMs for Data Synthesis through Failure-Inducing Exploration
von: Li, Qintong, et al.
Veröffentlicht: (2024)
von: Li, Qintong, et al.
Veröffentlicht: (2024)
FACTTRACK: Time-Aware World State Tracking in Story Outlines
von: Lyu, Zhiheng, et al.
Veröffentlicht: (2024)
von: Lyu, Zhiheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Training-Free Long-Context Scaling of Large Language Models
von: An, Chenxin, et al.
Veröffentlicht: (2024) -
GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models
von: Li, Mukai, et al.
Veröffentlicht: (2024) -
DiffuCoder: Understanding and Improving Masked Diffusion Models for Code Generation
von: Gong, Shansan, et al.
Veröffentlicht: (2025) -
Beyond Autoregression: Discrete Diffusion for Complex Reasoning and Planning
von: Ye, Jiacheng, et al.
Veröffentlicht: (2024) -
Scaling Diffusion Language Models via Adaptation from Autoregressive Models
von: Gong, Shansan, et al.
Veröffentlicht: (2024)