Curse of High Dimensionality Issue in Transformer for Long-context Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Shuhai, You, Zeng, Chen, Yaofo, Wen, Zhiquan, Wang, Qianyue, Qiu, Zhijie, Li, Yuanqing, Tan, Mingkui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Core Context Aware Transformers for Long Context Language Modeling
by: Chen, Yaofo, et al.
Published: (2024)
by: Chen, Yaofo, et al.
Published: (2024)
Training-free Context-adaptive Attention for Efficient Long Context Modeling
by: You, Zeng, et al.
Published: (2025)
by: You, Zeng, et al.
Published: (2025)
Latent-Condensed Transformer for Efficient Long Context Modeling
by: You, Zeng, et al.
Published: (2026)
by: You, Zeng, et al.
Published: (2026)
Detecting Machine-Generated Texts by Multi-Population Aware Optimization for Maximum Mean Discrepancy
by: Zhang, Shuhai, et al.
Published: (2024)
by: Zhang, Shuhai, et al.
Published: (2024)
Precedent-Informed Reasoning: Mitigating Overthinking in Large Reasoning Models via Test-Time Precedent Learning
by: Wang, Qianyue, et al.
Published: (2026)
by: Wang, Qianyue, et al.
Published: (2026)
Test-Time Learning for Large Language Models
by: Hu, Jinwu, et al.
Published: (2025)
by: Hu, Jinwu, et al.
Published: (2025)
Adapt in the Wild: Test-Time Entropy Minimization with Sharpness and Feature Regularization
by: Niu, Shuaicheng, et al.
Published: (2025)
by: Niu, Shuaicheng, et al.
Published: (2025)
Towards Long Video Understanding via Fine-detailed Video Story Generation
by: You, Zeng, et al.
Published: (2024)
by: You, Zeng, et al.
Published: (2024)
Uncertainty-Calibrated Test-Time Model Adaptation without Forgetting
by: Tan, Mingkui, et al.
Published: (2024)
by: Tan, Mingkui, et al.
Published: (2024)
In-context KV-Cache Eviction for LLMs via Attention-Gate
by: Zeng, Zihao, et al.
Published: (2024)
by: Zeng, Zihao, et al.
Published: (2024)
An Analysis and Mitigation of the Reversal Curse
by: Lv, Ang, et al.
Published: (2023)
by: Lv, Ang, et al.
Published: (2023)
Beyond Model Scaling: Test-Time Intervention for Efficient Deep Reasoning
by: Wang, Qianyue, et al.
Published: (2026)
by: Wang, Qianyue, et al.
Published: (2026)
Inference-Cost-Aware Dynamic Tree Construction for Efficient Inference in Large Language Models
by: Hong, Yinrong, et al.
Published: (2025)
by: Hong, Yinrong, et al.
Published: (2025)
Eliciting In-context Retrieval and Reasoning for Long-context Large Language Models
by: Qiu, Yifu, et al.
Published: (2025)
by: Qiu, Yifu, et al.
Published: (2025)
Deep Electromagnetic Structure Design Under Limited Evaluation Budgets
by: Zheng, Shijian, et al.
Published: (2025)
by: Zheng, Shijian, et al.
Published: (2025)
Graph of Records: Boosting Retrieval Augmented Generation for Long-context Summarization with Graphs
by: Zhang, Haozhen, et al.
Published: (2024)
by: Zhang, Haozhen, et al.
Published: (2024)
RNNs are not Transformers (Yet): The Key Bottleneck on In-context Retrieval
by: Wen, Kaiyue, et al.
Published: (2024)
by: Wen, Kaiyue, et al.
Published: (2024)
What is Wrong with Perplexity for Long-context Language Modeling?
by: Fang, Lizhe, et al.
Published: (2024)
by: Fang, Lizhe, et al.
Published: (2024)
LongReward: Improving Long-context Large Language Models with AI Feedback
by: Zhang, Jiajie, et al.
Published: (2024)
by: Zhang, Jiajie, et al.
Published: (2024)
The Information of Large Language Model Geometry
by: Tan, Zhiquan, et al.
Published: (2024)
by: Tan, Zhiquan, et al.
Published: (2024)
Can I understand what I create? Self-Knowledge Evaluation of Large Language Models
by: Tan, Zhiquan, et al.
Published: (2024)
by: Tan, Zhiquan, et al.
Published: (2024)
Mitigating Reversal Curse in Large Language Models via Semantic-aware Permutation Training
by: Guo, Qingyan, et al.
Published: (2024)
by: Guo, Qingyan, et al.
Published: (2024)
Is the Reversal Curse a Binding Problem? Uncovering Limitations of Transformers from a Basic Generalization Failure
by: Wang, Boshi, et al.
Published: (2025)
by: Wang, Boshi, et al.
Published: (2025)
The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More
by: Kitouni, Ouail, et al.
Published: (2024)
by: Kitouni, Ouail, et al.
Published: (2024)
Transformers are Universal In-context Learners
by: Furuya, Takashi, et al.
Published: (2024)
by: Furuya, Takashi, et al.
Published: (2024)
Towards a Theoretical Understanding of the 'Reversal Curse' via Training Dynamics
by: Zhu, Hanlin, et al.
Published: (2024)
by: Zhu, Hanlin, et al.
Published: (2024)
The Blessing and Curse of Dimensionality in Safety Alignment
by: Teo, Rachel S. Y., et al.
Published: (2025)
by: Teo, Rachel S. Y., et al.
Published: (2025)
Adversarial Purification by Consistency-aware Latent Space Optimization on Data Manifolds
by: Zhang, Shuhai, et al.
Published: (2024)
by: Zhang, Shuhai, et al.
Published: (2024)
The Position Curse: LLMs Struggle to Locate the Last Few Items in a List
by: Zhang, Zhanqi, et al.
Published: (2026)
by: Zhang, Zhanqi, et al.
Published: (2026)
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
by: Wang, Junxuan, et al.
Published: (2025)
by: Wang, Junxuan, et al.
Published: (2025)
Domaino1s: Guiding LLM Reasoning for Explainable Answers in High-Stakes Domains
by: Chu, Xu, et al.
Published: (2025)
by: Chu, Xu, et al.
Published: (2025)
ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
by: Longpre, Shayne, et al.
Published: (2025)
by: Longpre, Shayne, et al.
Published: (2025)
Do Depth-Grown Models Overcome the Curse of Depth? An In-Depth Analysis
by: Kapl, Ferdinand, et al.
Published: (2025)
by: Kapl, Ferdinand, et al.
Published: (2025)
R-Stitch: Dynamic Trajectory Stitching for Efficient Reasoning
by: Chen, Zhuokun, et al.
Published: (2025)
by: Chen, Zhuokun, et al.
Published: (2025)
Long-context Reference-based MT Quality Estimation
by: Haq, Sami Ul, et al.
Published: (2025)
by: Haq, Sami Ul, et al.
Published: (2025)
On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency
by: Wang, Yiming, et al.
Published: (2026)
by: Wang, Yiming, et al.
Published: (2026)
Beyond Fast and Slow: Cognitive-Inspired Elastic Reasoning for Large Language Models
by: Hu, Jinwu, et al.
Published: (2025)
by: Hu, Jinwu, et al.
Published: (2025)
Linking In-context Learning in Transformers to Human Episodic Memory
by: Ji-An, Li, et al.
Published: (2024)
by: Ji-An, Li, et al.
Published: (2024)
Diversity Enhances an LLM's Performance in RAG and Long-context Task
by: Wang, Zhichao, et al.
Published: (2025)
by: Wang, Zhichao, et al.
Published: (2025)
A Decomposition Perspective to Long-context Reasoning for LLMs
by: Xiao, Yanling, et al.
Published: (2026)
by: Xiao, Yanling, et al.
Published: (2026)
Similar Items
-
Core Context Aware Transformers for Long Context Language Modeling
by: Chen, Yaofo, et al.
Published: (2024) -
Training-free Context-adaptive Attention for Efficient Long Context Modeling
by: You, Zeng, et al.
Published: (2025) -
Latent-Condensed Transformer for Efficient Long Context Modeling
by: You, Zeng, et al.
Published: (2026) -
Detecting Machine-Generated Texts by Multi-Population Aware Optimization for Maximum Mean Discrepancy
by: Zhang, Shuhai, et al.
Published: (2024) -
Precedent-Informed Reasoning: Mitigating Overthinking in Large Reasoning Models via Test-Time Precedent Learning
by: Wang, Qianyue, et al.
Published: (2026)