CoCA: Fusing Position Embedding with Collinear Constrained Attention in Transformers for Long Context Window Extending
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Shiyi, Ye, Jing, Jiang, Wei, Xue, Siqiao, Zhang, Qi, Wu, Yifan, Li, Jianguo |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoCA: Cooperative Component Analysis
by: Ding, Daisy Yi, et al.
Published: (2024)
by: Ding, Daisy Yi, et al.
Published: (2024)
CoCA: Regaining Safety-awareness of Multimodal Large Language Models with Constitutional Calibration
by: Gao, Jiahui, et al.
Published: (2024)
by: Gao, Jiahui, et al.
Published: (2024)
LongEmbed: Extending Embedding Models for Long Context Retrieval
by: Zhu, Dawei, et al.
Published: (2024)
by: Zhu, Dawei, et al.
Published: (2024)
Gradual Forgetting: Logarithmic Compression for Extending Transformer Context Windows
by: Dickson, Billy, et al.
Published: (2025)
by: Dickson, Billy, et al.
Published: (2025)
PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise Training
by: Zhu, Dawei, et al.
Published: (2023)
by: Zhu, Dawei, et al.
Published: (2023)
Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings
by: Gelberg, Yoav, et al.
Published: (2025)
by: Gelberg, Yoav, et al.
Published: (2025)
LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning
by: Jin, Hongye, et al.
Published: (2024)
by: Jin, Hongye, et al.
Published: (2024)
An Evaluation of Context Length Extrapolation in Long Code via Positional Embeddings and Efficient Attention
by: Ghosh, Madhusudan, et al.
Published: (2026)
by: Ghosh, Madhusudan, et al.
Published: (2026)
Neighborhood Attention Transformer with Progressive Channel Fusion for Speaker Verification
by: Li, Nian, et al.
Published: (2024)
by: Li, Nian, et al.
Published: (2024)
LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens
by: Ding, Yiran, et al.
Published: (2024)
by: Ding, Yiran, et al.
Published: (2024)
Extending LLMs' Context Window with 100 Samples
by: Zhang, Yikai, et al.
Published: (2024)
by: Zhang, Yikai, et al.
Published: (2024)
Lightweight Vision Transformer with Window and Spatial Attention for Food Image Classification
by: Gao, Xinle, et al.
Published: (2025)
by: Gao, Xinle, et al.
Published: (2025)
Fourier Position Embedding: Enhancing Attention's Periodic Extension for Length Generalization
by: Hua, Ermo, et al.
Published: (2024)
by: Hua, Ermo, et al.
Published: (2024)
SWAA: Sliding Window Attention Adaptation for Efficient and Quality Preserving Long Context Processing
by: Yu, Yijiong, et al.
Published: (2025)
by: Yu, Yijiong, et al.
Published: (2025)
Rotary Positional Embeddings as Phase Modulation: Theoretical Bounds on the RoPE Base for Long-Context Transformers
by: Liu, Feilong
Published: (2026)
by: Liu, Feilong
Published: (2026)
PSC: Extending Context Window of Large Language Models via Phase Shift Calibration
by: Zhu, Wenqiao, et al.
Published: (2025)
by: Zhu, Wenqiao, et al.
Published: (2025)
Fovea Transformer: Efficient Long-Context Modeling with Structured Fine-to-Coarse Attention
by: He, Ziwei, et al.
Published: (2023)
by: He, Ziwei, et al.
Published: (2023)
Open-TQ-Metal: Fused Compressed-Domain Attention for Long-Context LLM Inference on Apple Silicon
by: Vegasena, Sai
Published: (2026)
by: Vegasena, Sai
Published: (2026)
Visual Context Window Extension: A New Perspective for Long Video Understanding
by: Wei, Hongchen, et al.
Published: (2024)
by: Wei, Hongchen, et al.
Published: (2024)
FuseMax: Leveraging Extended Einsums to Optimize Attention Accelerator Design
by: Nayak, Nandeeka, et al.
Published: (2024)
by: Nayak, Nandeeka, et al.
Published: (2024)
Efficient Context Scaling with LongCat ZigZag Attention
by: Zhang, Chen, et al.
Published: (2025)
by: Zhang, Chen, et al.
Published: (2025)
Functional Interpolation for Relative Positions Improves Long Context Transformers
by: Li, Shanda, et al.
Published: (2023)
by: Li, Shanda, et al.
Published: (2023)
Short Data, Long Context: Distilling Positional Knowledge in Transformers
by: Huber, Patrick, et al.
Published: (2026)
by: Huber, Patrick, et al.
Published: (2026)
OWL: Overcoming Window Length-Dependence in Speculative Decoding for Long-Context Inputs
by: Lee, Jaeseong, et al.
Published: (2025)
by: Lee, Jaeseong, et al.
Published: (2025)
HoPE: Hybrid of Position Embedding for Long Context Vision-Language Models
by: Li, Haoran, et al.
Published: (2025)
by: Li, Haoran, et al.
Published: (2025)
Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs
by: Liu, Xiaoran, et al.
Published: (2025)
by: Liu, Xiaoran, et al.
Published: (2025)
HiFlow: Hierarchical Feedback-Driven Optimization for Constrained Long-Form Text Generation
by: Zhu, Yifan, et al.
Published: (2026)
by: Zhu, Yifan, et al.
Published: (2026)
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
by: Lu, Yi, et al.
Published: (2024)
by: Lu, Yi, et al.
Published: (2024)
Positional Biases Shift as Inputs Approach Context Window Limits
by: Veseli, Blerta, et al.
Published: (2025)
by: Veseli, Blerta, et al.
Published: (2025)
CoMeT: Collaborative Memory Transformer for Efficient Long Context Modeling
by: Zhao, Runsong, et al.
Published: (2026)
by: Zhao, Runsong, et al.
Published: (2026)
LMAE4Eth: Generalizable and Robust Ethereum Fraud Detection by Exploring Transaction Semantics and Masked Graph Embedding
by: Jia, Yifan, et al.
Published: (2025)
by: Jia, Yifan, et al.
Published: (2025)
ReCA: Multi-Shot Long Video Extrapolation via Recursive Context Allocation
by: Liu, Akide, et al.
Published: (2026)
by: Liu, Akide, et al.
Published: (2026)
Found in the Middle: Calibrating Positional Attention Bias Improves Long Context Utilization
by: Hsieh, Cheng-Yu, et al.
Published: (2024)
by: Hsieh, Cheng-Yu, et al.
Published: (2024)
Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers
by: Horton, Mark, et al.
Published: (2025)
by: Horton, Mark, et al.
Published: (2025)
Infinite Retrieval: Attention Enhanced LLMs in Long-Context Processing
by: Ye, Xiaoju, et al.
Published: (2025)
by: Ye, Xiaoju, et al.
Published: (2025)
Adamas: Hadamard Sparse Attention for Efficient Long-Context Inference
by: Yan, Siyuan, et al.
Published: (2025)
by: Yan, Siyuan, et al.
Published: (2025)
Cross-level Attention with Overlapped Windows for Camouflaged Object Detection
by: Li, Jiepan, et al.
Published: (2023)
by: Li, Jiepan, et al.
Published: (2023)
Collinear anatomy
by: Belitsky, A. V.
Published: (2024)
by: Belitsky, A. V.
Published: (2024)
Context-aware Rotary Position Embedding
by: Veisi, Ali, et al.
Published: (2025)
by: Veisi, Ali, et al.
Published: (2025)
InfiniteICL: Breaking the Limit of Context Window Size via Long Short-term Memory Transformation
by: Cao, Bowen, et al.
Published: (2025)
by: Cao, Bowen, et al.
Published: (2025)
Similar Items
-
CoCA: Cooperative Component Analysis
by: Ding, Daisy Yi, et al.
Published: (2024) -
CoCA: Regaining Safety-awareness of Multimodal Large Language Models with Constitutional Calibration
by: Gao, Jiahui, et al.
Published: (2024) -
LongEmbed: Extending Embedding Models for Long Context Retrieval
by: Zhu, Dawei, et al.
Published: (2024) -
Gradual Forgetting: Logarithmic Compression for Extending Transformer Context Windows
by: Dickson, Billy, et al.
Published: (2025) -
PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise Training
by: Zhu, Dawei, et al.
Published: (2023)