PaTH Attention: Position Encoding via Accumulating Householder Transformations
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Songlin, Shen, Yikang, Wen, Kaiyue, Tan, Shawn, Mishra, Mayank, Ren, Liliang, Panda, Rameswar, Kim, Yoon |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Gated Linear Attention Transformers with Hardware-Efficient Training
by: Yang, Songlin, et al.
Published: (2023)
by: Yang, Songlin, et al.
Published: (2023)
Distilling to Hybrid Attention Models via KL-Guided Layer Selection
by: Li, Yanhong, et al.
Published: (2025)
by: Li, Yanhong, et al.
Published: (2025)
Scaling Stick-Breaking Attention: An Efficient Implementation and In-depth Study
by: Tan, Shawn, et al.
Published: (2024)
by: Tan, Shawn, et al.
Published: (2024)
Scattered Mixture-of-Experts Implementation
by: Tan, Shawn, et al.
Published: (2024)
by: Tan, Shawn, et al.
Published: (2024)
FlashFormer: Whole-Model Kernels for Efficient Low-Batch Inference
by: Nrusimha, Aniruddha, et al.
Published: (2025)
by: Nrusimha, Aniruddha, et al.
Published: (2025)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
by: Brandon, William, et al.
Published: (2024)
by: Brandon, William, et al.
Published: (2024)
Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler
by: Shen, Yikang, et al.
Published: (2024)
by: Shen, Yikang, et al.
Published: (2024)
Mitigating the Impact of Outlier Channels for Language Model Quantization with Activation Regularization
by: Nrusimha, Aniruddha, et al.
Published: (2024)
by: Nrusimha, Aniruddha, et al.
Published: (2024)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
by: Pan, Bowen, et al.
Published: (2024)
by: Pan, Bowen, et al.
Published: (2024)
Finding the Minimal Parameter Budget for Implicit Reasoning: A Data Complexity Driven Scaling Law for Language Models
by: Wang, Xinyi, et al.
Published: (2025)
by: Wang, Xinyi, et al.
Published: (2025)
Diversity Measurement and Subset Selection for Instruction Tuning Datasets
by: Wang, Peiqi, et al.
Published: (2024)
by: Wang, Peiqi, et al.
Published: (2024)
Parallelizing Linear Transformers with the Delta Rule over Sequence Length
by: Yang, Songlin, et al.
Published: (2024)
by: Yang, Songlin, et al.
Published: (2024)
API Pack: A Massive Multi-Programming Language Dataset for API Call Generation
by: Guo, Zhen, et al.
Published: (2024)
by: Guo, Zhen, et al.
Published: (2024)
Structured Code Representations Enable Data-Efficient Adaptation of Code Language Models
by: Agarwal, Mayank, et al.
Published: (2024)
by: Agarwal, Mayank, et al.
Published: (2024)
Positional Encoding via Token-Aware Phase Attention
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
SITAR: Semi-supervised Image Transformer for Action Recognition
by: Iqbal, Owais, et al.
Published: (2024)
by: Iqbal, Owais, et al.
Published: (2024)
M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
by: Mishra, Mayank, et al.
Published: (2026)
by: Mishra, Mayank, et al.
Published: (2026)
TOUCAN: Synthesizing 1.5M Tool-Agentic Data from Real-World MCP Environments
by: Xu, Zhangchen, et al.
Published: (2025)
by: Xu, Zhangchen, et al.
Published: (2025)
RelFlexformer: Efficient Attention 3D-Transformers for Integrable Relative Positional Encodings
by: Kim, Byeongchan, et al.
Published: (2026)
by: Kim, Byeongchan, et al.
Published: (2026)
Data Engineering for Scaling Language Models to 128K Context
by: Fu, Yao, et al.
Published: (2024)
by: Fu, Yao, et al.
Published: (2024)
From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency
by: Wen, Kaiyue, et al.
Published: (2024)
by: Wen, Kaiyue, et al.
Published: (2024)
Energy-Gated Attention and Wavelet Positional Encoding: Complementary Inductive Biases for Transformer Attention
by: Zeris, Athanasios
Published: (2026)
by: Zeris, Athanasios
Published: (2026)
Comparing Graph Transformers via Positional Encodings
by: Black, Mitchell, et al.
Published: (2024)
by: Black, Mitchell, et al.
Published: (2024)
Toward Relative Positional Encoding in Spiking Transformers
by: Lv, Changze, et al.
Published: (2025)
by: Lv, Changze, et al.
Published: (2025)
LangNav: Language as a Perceptual Representation for Navigation
by: Pan, Bowen, et al.
Published: (2023)
by: Pan, Bowen, et al.
Published: (2023)
Dynamical Properties of Tokens in Self-Attention and Effects of Positional Encoding
by: Pham, Duy-Tung, et al.
Published: (2025)
by: Pham, Duy-Tung, et al.
Published: (2025)
PRISM: Demystifying Retention and Interaction in Mid-Training
by: Runwal, Bharat, et al.
Published: (2026)
by: Runwal, Bharat, et al.
Published: (2026)
On the Geometry of Positional Encodings in Transformers
by: Cirrincione, Giansalvo
Published: (2026)
by: Cirrincione, Giansalvo
Published: (2026)
Improving Position Encoding of Transformers for Multivariate Time Series Classification
by: Foumani, Navid Mohammadi, et al.
Published: (2023)
by: Foumani, Navid Mohammadi, et al.
Published: (2023)
Analysis of Attention in Video Diffusion Transformers
by: Wen, Yuxin, et al.
Published: (2025)
by: Wen, Yuxin, et al.
Published: (2025)
PaReGTA: An LLM-based EHR Data Encoding Approach to Capture Temporal Information
by: Yoon, Kihyuk, et al.
Published: (2026)
by: Yoon, Kihyuk, et al.
Published: (2026)
LEDiT: Your Length-Extrapolatable Diffusion Transformer without Positional Encoding
by: Zhang, Shen, et al.
Published: (2025)
by: Zhang, Shen, et al.
Published: (2025)
Log-Linear Attention
by: Guo, Han, et al.
Published: (2025)
by: Guo, Han, et al.
Published: (2025)
Lightweight Spatio-Temporal Attention Network with Graph Embedding and Rotational Position Encoding for Traffic Forecasting
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
RNNs are not Transformers (Yet): The Key Bottleneck on In-context Retrieval
by: Wen, Kaiyue, et al.
Published: (2024)
by: Wen, Kaiyue, et al.
Published: (2024)
Spiking Transformer with Spatial-Temporal Attention
by: Lee, Donghyun, et al.
Published: (2024)
by: Lee, Donghyun, et al.
Published: (2024)
Graph Attention for Heterogeneous Graphs with Positional Encoding
by: Nayak, Nikhil Shivakumar
Published: (2025)
by: Nayak, Nikhil Shivakumar
Published: (2025)
Weierstrass Positional Encoding for Vision Transformers
by: Xin, Zhihang, et al.
Published: (2026)
by: Xin, Zhihang, et al.
Published: (2026)
Graph Transformers without Positional Encodings
by: Garg, Ayush
Published: (2024)
by: Garg, Ayush
Published: (2024)
Integrating Physics Inspired Features with Graph Convolution
by: Sahu, Rameswar
Published: (2024)
by: Sahu, Rameswar
Published: (2024)
Similar Items
-
Gated Linear Attention Transformers with Hardware-Efficient Training
by: Yang, Songlin, et al.
Published: (2023) -
Distilling to Hybrid Attention Models via KL-Guided Layer Selection
by: Li, Yanhong, et al.
Published: (2025) -
Scaling Stick-Breaking Attention: An Efficient Implementation and In-depth Study
by: Tan, Shawn, et al.
Published: (2024) -
Scattered Mixture-of-Experts Implementation
by: Tan, Shawn, et al.
Published: (2024) -
FlashFormer: Whole-Model Kernels for Efficient Low-Batch Inference
by: Nrusimha, Aniruddha, et al.
Published: (2025)