Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure
Fuente:
arXiv
Saved in:
| Main Authors: | Cho, Hanseul, Cha, Jaeyoung, Awasthi, Pranjal, Bhojanapalli, Srinadh, Gupta, Anupam, Yun, Chulhee |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
by: Cho, Hanseul, et al.
Published: (2024)
by: Cho, Hanseul, et al.
Published: (2024)
DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of Plasticity
by: Shin, Baekrok, et al.
Published: (2024)
by: Shin, Baekrok, et al.
Published: (2024)
The Limits of Preference Data for Post-Training
by: Zhao, Eric, et al.
Published: (2025)
by: Zhao, Eric, et al.
Published: (2025)
On student-teacher deviations in distillation: does it pay to disobey?
by: Nagarajan, Vaishnavh, et al.
Published: (2023)
by: Nagarajan, Vaishnavh, et al.
Published: (2023)
Characterizing Pattern Matching and Its Limits on Compositional Task Structures
by: Chang, Hoyeon, et al.
Published: (2025)
by: Chang, Hoyeon, et al.
Published: (2025)
E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition
by: Gupta, Aryan, et al.
Published: (2025)
by: Gupta, Aryan, et al.
Published: (2025)
Extrapolation by Association: Length Generalization Transfer in Transformers
by: Cai, Ziyang, et al.
Published: (2025)
by: Cai, Ziyang, et al.
Published: (2025)
Task Arithmetic for Language Expansion in Speech Translation
by: Cheng, Yao-Fei, et al.
Published: (2024)
by: Cheng, Yao-Fei, et al.
Published: (2024)
Fourier Position Embedding: Enhancing Attention's Periodic Extension for Length Generalization
by: Hua, Ermo, et al.
Published: (2024)
by: Hua, Ermo, et al.
Published: (2024)
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
by: Gambardella, Andrew, et al.
Published: (2024)
by: Gambardella, Andrew, et al.
Published: (2024)
Structured Thinking Matters: Improving LLMs Generalization in Causal Inference Tasks
by: Sun, Wentao, et al.
Published: (2025)
by: Sun, Wentao, et al.
Published: (2025)
Fundamental Benefit of Alternating Updates in Minimax Optimization
by: Lee, Jaewook, et al.
Published: (2024)
by: Lee, Jaewook, et al.
Published: (2024)
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
by: Jung, Hyunji, et al.
Published: (2025)
by: Jung, Hyunji, et al.
Published: (2025)
The Role of Sparsity for Length Generalization in Transformers
by: Golowich, Noah, et al.
Published: (2025)
by: Golowich, Noah, et al.
Published: (2025)
MATA: Multi-Agent Framework for Reliable and Flexible Table Question Answering
by: Hyeon, Sieun, et al.
Published: (2026)
by: Hyeon, Sieun, et al.
Published: (2026)
Transformers Can Achieve Length Generalization But Not Robustly
by: Zhou, Yongchao, et al.
Published: (2024)
by: Zhou, Yongchao, et al.
Published: (2024)
Bias Vector: Mitigating Biases in Language Models with Task Arithmetic Approach
by: Shirafuji, Daiki, et al.
Published: (2024)
by: Shirafuji, Daiki, et al.
Published: (2024)
PositionID: LLMs can Control Lengths, Copy and Paste with Explicit Positional Awareness
by: Wang, Zekun, et al.
Published: (2024)
by: Wang, Zekun, et al.
Published: (2024)
Can LLMs Generate High-Quality Task-Specific Conversations?
by: Li, Shengqi, et al.
Published: (2025)
by: Li, Shengqi, et al.
Published: (2025)
Resonance RoPE: Improving Context Length Generalization of Large Language Models
by: Wang, Suyuchen, et al.
Published: (2024)
by: Wang, Suyuchen, et al.
Published: (2024)
Long-Span Question-Answering: Automatic Question Generation and QA-System Ranking via Side-by-Side Evaluation
by: Bohnet, Bernd, et al.
Published: (2024)
by: Bohnet, Bernd, et al.
Published: (2024)
MIDGARD: Self-Consistency Using Minimum Description Length for Structured Commonsense Reasoning
by: Nair, Inderjeet, et al.
Published: (2024)
by: Nair, Inderjeet, et al.
Published: (2024)
Don't Let It Fade: Preserving Edits in Diffusion Language Models via Token Timestep Allocation
by: Kim, Woojin, et al.
Published: (2025)
by: Kim, Woojin, et al.
Published: (2025)
Arithmetic Reasoning with LLM: Prolog Generation & Permutation
by: Yang, Xiaocheng, et al.
Published: (2024)
by: Yang, Xiaocheng, et al.
Published: (2024)
Effective Length Extrapolation via Dimension-Wise Positional Embeddings Manipulation
by: Lu, Yi, et al.
Published: (2025)
by: Lu, Yi, et al.
Published: (2025)
Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
by: Leng, Jiaqi, et al.
Published: (2025)
by: Leng, Jiaqi, et al.
Published: (2025)
Provable Benefit of Cutout and CutMix for Feature Learning
by: Oh, Junsoo, et al.
Published: (2024)
by: Oh, Junsoo, et al.
Published: (2024)
Autoregressive Ranking: Bridging the Gap Between Dual and Cross Encoders
by: Rozonoyer, Benjamin, et al.
Published: (2026)
by: Rozonoyer, Benjamin, et al.
Published: (2026)
Multi-Task Pre-Finetuning of Lightweight Transformer Encoders for Text Classification and NER
by: Zhu, Junyi, et al.
Published: (2025)
by: Zhu, Junyi, et al.
Published: (2025)
Lightweight Dataset Pruning without Full Training via Example Difficulty and Prediction Uncertainty
by: Cho, Yeseul, et al.
Published: (2025)
by: Cho, Yeseul, et al.
Published: (2025)
SALAD: Improving Robustness and Generalization through Contrastive Learning with Structure-Aware and LLM-Driven Augmented Data
by: Bae, Suyoung, et al.
Published: (2025)
by: Bae, Suyoung, et al.
Published: (2025)
Sample, Scrutinize and Scale: Effective Inference-Time Search by Scaling Verification
by: Zhao, Eric, et al.
Published: (2025)
by: Zhao, Eric, et al.
Published: (2025)
CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
by: Dong, Qihua, et al.
Published: (2025)
by: Dong, Qihua, et al.
Published: (2025)
Incremental Gradient Descent with Small Epoch Counts is Surprisingly Slow on Ill-Conditioned Problems
by: Kim, Yujun, et al.
Published: (2025)
by: Kim, Yujun, et al.
Published: (2025)
Position as Probability: Self-Supervised Transformers that Think Past Their Training for Length Extrapolation
by: Lee, Philip Heejun
Published: (2025)
by: Lee, Philip Heejun
Published: (2025)
RFEval: Benchmarking Reasoning Faithfulness under Counterfactual Reasoning Intervention in Large Reasoning Models
by: Han, Yunseok, et al.
Published: (2026)
by: Han, Yunseok, et al.
Published: (2026)
HiRE: High Recall Approximate Top-$k$ Estimation for Efficient LLM Inference
by: L, Yashas Samaga B, et al.
Published: (2024)
by: L, Yashas Samaga B, et al.
Published: (2024)
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
by: Aggarwal, Pranjal, et al.
Published: (2025)
by: Aggarwal, Pranjal, et al.
Published: (2025)
Integrating Arithmetic Learning Improves Mathematical Reasoning in Smaller Models
by: Gangwar, Neeraj, et al.
Published: (2025)
by: Gangwar, Neeraj, et al.
Published: (2025)
Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic
by: Bhardwaj, Rishabh, et al.
Published: (2024)
by: Bhardwaj, Rishabh, et al.
Published: (2024)
Similar Items
-
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
by: Cho, Hanseul, et al.
Published: (2024) -
DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of Plasticity
by: Shin, Baekrok, et al.
Published: (2024) -
The Limits of Preference Data for Post-Training
by: Zhao, Eric, et al.
Published: (2025) -
On student-teacher deviations in distillation: does it pay to disobey?
by: Nagarajan, Vaishnavh, et al.
Published: (2023) -
Characterizing Pattern Matching and Its Limits on Compositional Task Structures
by: Chang, Hoyeon, et al.
Published: (2025)