On the Limitations and Capabilities of Position Embeddings for Length Generalization
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yang, Liang, Yitao, Lin, Zhouchen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DIGIC: Domain Generalizable Imitation Learning by Causal Discovery
by: Chen, Yang, et al.
Published: (2024)
by: Chen, Yang, et al.
Published: (2024)
Low-Dimension-to-High-Dimension Generalization And Its Implications for Length Generalization
by: Chen, Yang, et al.
Published: (2024)
by: Chen, Yang, et al.
Published: (2024)
Attention Sinks Induce Gradient Sinks: Massive Activations as Gradient Regulators in Transformers
by: Chen, Yihong, et al.
Published: (2026)
by: Chen, Yihong, et al.
Published: (2026)
Relational Learning in Pre-Trained Models: A Theory from Hypergraph Recovery Perspective
by: Chen, Yang, et al.
Published: (2024)
by: Chen, Yang, et al.
Published: (2024)
RulE: Knowledge Graph Reasoning with Rule Embedding
by: Tang, Xiaojuan, et al.
Published: (2022)
by: Tang, Xiaojuan, et al.
Published: (2022)
On the Adversarial Transferability of Generalized "Skip Connections"
by: Wang, Yisen, et al.
Published: (2024)
by: Wang, Yisen, et al.
Published: (2024)
Proximity Matters: Local Proximity Enhanced Balancing for Treatment Effect Estimation
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Positional versus Symbolic Attention Heads: Learning Dynamics, RoPE Geometry, and Length Generalization
by: Urrutia, Felipe, et al.
Published: (2026)
by: Urrutia, Felipe, et al.
Published: (2026)
Mamba Modulation: On the Length Generalization of Mamba
by: Lu, Peng, et al.
Published: (2025)
by: Lu, Peng, et al.
Published: (2025)
FedAdamW: A Communication-Efficient Optimizer with Convergence and Generalization Guarantees for Federated Large Models
by: Liu, Junkang, et al.
Published: (2025)
by: Liu, Junkang, et al.
Published: (2025)
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
by: Cho, Hanseul, et al.
Published: (2024)
by: Cho, Hanseul, et al.
Published: (2024)
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure
by: Cho, Hanseul, et al.
Published: (2024)
by: Cho, Hanseul, et al.
Published: (2024)
Rethinking RoPE: A Mathematical Blueprint for N-dimensional Positional Embedding
by: Liu, Haiping, et al.
Published: (2025)
by: Liu, Haiping, et al.
Published: (2025)
MEP: Multiple Kernel Learning Enhancing Relative Positional Encoding Length Extrapolation
by: Gao, Weiguo
Published: (2024)
by: Gao, Weiguo
Published: (2024)
Position: Capability Control Should be a Separate Goal From Alignment
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2026)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2026)
On the Theoretical Limitations of Embedding-based Link Prediction
by: Badreddine, Samy, et al.
Published: (2025)
by: Badreddine, Samy, et al.
Published: (2025)
Think Then Embed: Generative Context Improves Multimodal Embedding
by: Cui, Xuanming, et al.
Published: (2025)
by: Cui, Xuanming, et al.
Published: (2025)
On Vanishing Variance in Transformer Length Generalization
by: Li, Ruining, et al.
Published: (2025)
by: Li, Ruining, et al.
Published: (2025)
Temporal Spiking Neural Networks with Synaptic Delay for Graph Reasoning
by: Xiao, Mingqing, et al.
Published: (2024)
by: Xiao, Mingqing, et al.
Published: (2024)
GL-Fusion: Rethinking the Combination of Graph Neural Network and Large Language model
by: Yang, Haotong, et al.
Published: (2024)
by: Yang, Haotong, et al.
Published: (2024)
A Tractable Inference Perspective of Offline RL
by: Liu, Xuejie, et al.
Published: (2023)
by: Liu, Xuejie, et al.
Published: (2023)
Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
by: Wu, Zhoutong, et al.
Published: (2025)
by: Wu, Zhoutong, et al.
Published: (2025)
Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation
by: He, Zhenyu, et al.
Published: (2024)
by: He, Zhenyu, et al.
Published: (2024)
Hierarchical Position Embedding of Graphs with Landmarks and Clustering for Link Prediction
by: Kim, Minsang, et al.
Published: (2024)
by: Kim, Minsang, et al.
Published: (2024)
Generalist++: A Meta-learning Framework for Mitigating Trade-off in Adversarial Training
by: Wang, Yisen, et al.
Published: (2025)
by: Wang, Yisen, et al.
Published: (2025)
An Empirical Study on Context Length for Open-Domain Dialog Generation
by: Shen, Xinyi, et al.
Published: (2024)
by: Shen, Xinyi, et al.
Published: (2024)
Croppable Knowledge Graph Embedding
by: Zhu, Yushan, et al.
Published: (2024)
by: Zhu, Yushan, et al.
Published: (2024)
TFG-Flow: Training-free Guidance in Multimodal Generative Flow
by: Lin, Haowei, et al.
Published: (2025)
by: Lin, Haowei, et al.
Published: (2025)
Efficient Generative Model Training via Embedded Representation Warmup
by: Liu, Deyuan, et al.
Published: (2025)
by: Liu, Deyuan, et al.
Published: (2025)
Continuous-Time Linear Positional Embedding for Irregular Time Series Forecasting
by: Kim, Byunghyun, et al.
Published: (2024)
by: Kim, Byunghyun, et al.
Published: (2024)
Generalization Capability for Imitation Learning
by: Wang, Yixiao
Published: (2025)
by: Wang, Yixiao
Published: (2025)
Embedding Hidden Adversarial Capabilities in Pre-Trained Diffusion Models
by: Beerens, Lucas, et al.
Published: (2025)
by: Beerens, Lucas, et al.
Published: (2025)
Model-Based Reinforcement Learning with Multi-Task Offline Pretraining
by: Pan, Minting, et al.
Published: (2023)
by: Pan, Minting, et al.
Published: (2023)
Position: AI Evaluations Should be Grounded on a Theory of Capability
by: Jo, Nathanael, et al.
Published: (2025)
by: Jo, Nathanael, et al.
Published: (2025)
GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents
by: Cai, Shaofei, et al.
Published: (2024)
by: Cai, Shaofei, et al.
Published: (2024)
Transformers Can Achieve Length Generalization But Not Robustly
by: Zhou, Yongchao, et al.
Published: (2024)
by: Zhou, Yongchao, et al.
Published: (2024)
Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models
by: Phan, Hoang, et al.
Published: (2025)
by: Phan, Hoang, et al.
Published: (2025)
Generalizing Knowledge Graph Embedding with Universal Orthogonal Parameterization
by: Li, Rui, et al.
Published: (2024)
by: Li, Rui, et al.
Published: (2024)
Hebbian Learning based Orthogonal Projection for Continual Learning of Spiking Neural Networks
by: Xiao, Mingqing, et al.
Published: (2024)
by: Xiao, Mingqing, et al.
Published: (2024)
Online Pseudo-Zeroth-Order Training of Neuromorphic Spiking Neural Networks
by: Xiao, Mingqing, et al.
Published: (2024)
by: Xiao, Mingqing, et al.
Published: (2024)
Similar Items
-
DIGIC: Domain Generalizable Imitation Learning by Causal Discovery
by: Chen, Yang, et al.
Published: (2024) -
Low-Dimension-to-High-Dimension Generalization And Its Implications for Length Generalization
by: Chen, Yang, et al.
Published: (2024) -
Attention Sinks Induce Gradient Sinks: Massive Activations as Gradient Regulators in Transformers
by: Chen, Yihong, et al.
Published: (2026) -
Relational Learning in Pre-Trained Models: A Theory from Hypergraph Recovery Perspective
by: Chen, Yang, et al.
Published: (2024) -
RulE: Knowledge Graph Reasoning with Rule Embedding
by: Tang, Xiaojuan, et al.
Published: (2022)