Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Nayoung, Cai, Ziyang, Schwarzschild, Avi, Lee, Kangwook, Papailiopoulos, Dimitris |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Extrapolation by Association: Length Generalization Transfer in Transformers
di: Cai, Ziyang, et al.
Pubblicazione: (2025)
di: Cai, Ziyang, et al.
Pubblicazione: (2025)
From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic Data
di: Xiong, Zheyang, et al.
Pubblicazione: (2024)
di: Xiong, Zheyang, et al.
Pubblicazione: (2024)
Easy2Hard-Bench: Standardized Difficulty Labels for Profiling LLM Performance and Generalization
di: Ding, Mucong, et al.
Pubblicazione: (2024)
di: Ding, Mucong, et al.
Pubblicazione: (2024)
Looped Transformers are Better at Learning Learning Algorithms
di: Yang, Liu, et al.
Pubblicazione: (2023)
di: Yang, Liu, et al.
Pubblicazione: (2023)
How Well Can Transformers Emulate In-context Newton's Method?
di: Giannou, Angeliki, et al.
Pubblicazione: (2024)
di: Giannou, Angeliki, et al.
Pubblicazione: (2024)
Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks
di: Park, Jongho, et al.
Pubblicazione: (2024)
di: Park, Jongho, et al.
Pubblicazione: (2024)
Benchmarking ChatGPT on Algorithmic Reasoning
di: McLeish, Sean, et al.
Pubblicazione: (2024)
di: McLeish, Sean, et al.
Pubblicazione: (2024)
Everything Everywhere All at Once: LLMs can In-Context Learn Multiple Tasks in Superposition
di: Xiong, Zheyang, et al.
Pubblicazione: (2024)
di: Xiong, Zheyang, et al.
Pubblicazione: (2024)
Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
di: Yang, Seongjun, et al.
Pubblicazione: (2023)
di: Yang, Seongjun, et al.
Pubblicazione: (2023)
The Expressive Power of Low-Rank Adaptation
di: Zeng, Yuchen, et al.
Pubblicazione: (2023)
di: Zeng, Yuchen, et al.
Pubblicazione: (2023)
Task Vectors in In-Context Learning: Emergence, Formation, and Benefit
di: Yang, Liu, et al.
Pubblicazione: (2025)
di: Yang, Liu, et al.
Pubblicazione: (2025)
Looped Transformers for Length Generalization
di: Fan, Ying, et al.
Pubblicazione: (2024)
di: Fan, Ying, et al.
Pubblicazione: (2024)
Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
di: Sun, Zhiqing, et al.
Pubblicazione: (2024)
di: Sun, Zhiqing, et al.
Pubblicazione: (2024)
Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning
di: Parashar, Shubham, et al.
Pubblicazione: (2025)
di: Parashar, Shubham, et al.
Pubblicazione: (2025)
Transformers Can Do Arithmetic with the Right Embeddings
di: McLeish, Sean, et al.
Pubblicazione: (2024)
di: McLeish, Sean, et al.
Pubblicazione: (2024)
More Consistent Accuracy PINN via Alternating Easy-Hard Training
di: Gao, Zhaoqian, et al.
Pubblicazione: (2025)
di: Gao, Zhaoqian, et al.
Pubblicazione: (2025)
Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
di: Hans, Abhimanyu, et al.
Pubblicazione: (2024)
di: Hans, Abhimanyu, et al.
Pubblicazione: (2024)
Position as Probability: Self-Supervised Transformers that Think Past Their Training for Length Extrapolation
di: Lee, Philip Heejun
Pubblicazione: (2025)
di: Lee, Philip Heejun
Pubblicazione: (2025)
On Vanishing Variance in Transformer Length Generalization
di: Li, Ruining, et al.
Pubblicazione: (2025)
di: Li, Ruining, et al.
Pubblicazione: (2025)
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
di: Cho, Hanseul, et al.
Pubblicazione: (2024)
di: Cho, Hanseul, et al.
Pubblicazione: (2024)
Enhancing Anomaly Detection via Generating Diversified and Hard-to-distinguish Synthetic Anomalies
di: Kim, Hyuntae, et al.
Pubblicazione: (2024)
di: Kim, Hyuntae, et al.
Pubblicazione: (2024)
LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation
di: Ahn, Jinwoo, et al.
Pubblicazione: (2026)
di: Ahn, Jinwoo, et al.
Pubblicazione: (2026)
The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
di: Hase, Peter, et al.
Pubblicazione: (2024)
di: Hase, Peter, et al.
Pubblicazione: (2024)
Shorter but not Worse: Frugal Reasoning via Easy Samples as Length Regularizers in Math RLVR
di: Bounhar, Abdelaziz, et al.
Pubblicazione: (2025)
di: Bounhar, Abdelaziz, et al.
Pubblicazione: (2025)
Learning from Teaching Regularization: Generalizable Correlations Should be Easy to Imitate
di: Jin, Can, et al.
Pubblicazione: (2024)
di: Jin, Can, et al.
Pubblicazione: (2024)
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
di: Gambardella, Andrew, et al.
Pubblicazione: (2024)
di: Gambardella, Andrew, et al.
Pubblicazione: (2024)
Dynamic Relational Priming Improves Transformer in Multivariate Time Series
di: Lee, Hunjae, et al.
Pubblicazione: (2025)
di: Lee, Hunjae, et al.
Pubblicazione: (2025)
Scalability Matters: Overcoming Challenges in InstructGLM with Similarity-Degree-Based Sampling
di: Lee, Hyun, et al.
Pubblicazione: (2025)
di: Lee, Hyun, et al.
Pubblicazione: (2025)
Repurformer: Transformers for Repurposing-Aware Molecule Generation
di: Lee, Changhun, et al.
Pubblicazione: (2024)
di: Lee, Changhun, et al.
Pubblicazione: (2024)
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure
di: Cho, Hanseul, et al.
Pubblicazione: (2024)
di: Cho, Hanseul, et al.
Pubblicazione: (2024)
Self-supervised Graph Transformer with Contrastive Learning for Brain Connectivity Analysis towards Improving Autism Detection
di: Leng, Yicheng, et al.
Pubblicazione: (2025)
di: Leng, Yicheng, et al.
Pubblicazione: (2025)
The Role of Sparsity for Length Generalization in Transformers
di: Golowich, Noah, et al.
Pubblicazione: (2025)
di: Golowich, Noah, et al.
Pubblicazione: (2025)
MEMENTO: Teaching LLMs to Manage Their Own Context
di: Kontonis, Vasilis, et al.
Pubblicazione: (2026)
di: Kontonis, Vasilis, et al.
Pubblicazione: (2026)
Self-Attribution Bias: When AI Monitors Go Easy on Themselves
di: Khullar, Dipika, et al.
Pubblicazione: (2026)
di: Khullar, Dipika, et al.
Pubblicazione: (2026)
MILES: Making Imitation Learning Easy with Self-Supervision
di: Papagiannis, Georgios, et al.
Pubblicazione: (2024)
di: Papagiannis, Georgios, et al.
Pubblicazione: (2024)
Transformers Can Achieve Length Generalization But Not Robustly
di: Zhou, Yongchao, et al.
Pubblicazione: (2024)
di: Zhou, Yongchao, et al.
Pubblicazione: (2024)
HardCore Generation: Generating Hard UNSAT Problems for Data Augmentation
di: Cotnareanu, Joseph, et al.
Pubblicazione: (2024)
di: Cotnareanu, Joseph, et al.
Pubblicazione: (2024)
SAFE: Finding Sparse and Flat Minima to Improve Pruning
di: Lee, Dongyeop, et al.
Pubblicazione: (2025)
di: Lee, Dongyeop, et al.
Pubblicazione: (2025)
Graph Convolutions Enrich the Self-Attention in Transformers!
di: Choi, Jeongwhan, et al.
Pubblicazione: (2023)
di: Choi, Jeongwhan, et al.
Pubblicazione: (2023)
Existing Large Language Model Unlearning Evaluations Are Inconclusive
di: Feng, Zhili, et al.
Pubblicazione: (2025)
di: Feng, Zhili, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Extrapolation by Association: Length Generalization Transfer in Transformers
di: Cai, Ziyang, et al.
Pubblicazione: (2025) -
From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic Data
di: Xiong, Zheyang, et al.
Pubblicazione: (2024) -
Easy2Hard-Bench: Standardized Difficulty Labels for Profiling LLM Performance and Generalization
di: Ding, Mucong, et al.
Pubblicazione: (2024) -
Looped Transformers are Better at Learning Learning Algorithms
di: Yang, Liu, et al.
Pubblicazione: (2023) -
How Well Can Transformers Emulate In-context Newton's Method?
di: Giannou, Angeliki, et al.
Pubblicazione: (2024)