Extrapolation by Association: Length Generalization Transfer in Transformers
Fuente:
arXiv
Salvato in:
| Autori principali: | Cai, Ziyang, Lee, Nayoung, Schwarzschild, Avi, Oymak, Samet, Papailiopoulos, Dimitris |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
di: Lee, Nayoung, et al.
Pubblicazione: (2025)
di: Lee, Nayoung, et al.
Pubblicazione: (2025)
Everything Everywhere All at Once: LLMs can In-Context Learn Multiple Tasks in Superposition
di: Xiong, Zheyang, et al.
Pubblicazione: (2024)
di: Xiong, Zheyang, et al.
Pubblicazione: (2024)
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
di: Ildiz, M. Emrullah, et al.
Pubblicazione: (2024)
di: Ildiz, M. Emrullah, et al.
Pubblicazione: (2024)
Transformers as Support Vector Machines
di: Tarzanagh, Davoud Ataee, et al.
Pubblicazione: (2023)
di: Tarzanagh, Davoud Ataee, et al.
Pubblicazione: (2023)
From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic Data
di: Xiong, Zheyang, et al.
Pubblicazione: (2024)
di: Xiong, Zheyang, et al.
Pubblicazione: (2024)
Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
di: Ball, Sarah, et al.
Pubblicazione: (2025)
di: Ball, Sarah, et al.
Pubblicazione: (2025)
Benchmarking ChatGPT on Algorithmic Reasoning
di: McLeish, Sean, et al.
Pubblicazione: (2024)
di: McLeish, Sean, et al.
Pubblicazione: (2024)
Fine-grained Analysis of In-context Linear Estimation: Data, Architecture, and Beyond
di: Li, Yingcong, et al.
Pubblicazione: (2024)
di: Li, Yingcong, et al.
Pubblicazione: (2024)
Position as Probability: Self-Supervised Transformers that Think Past Their Training for Length Extrapolation
di: Lee, Philip Heejun
Pubblicazione: (2025)
di: Lee, Philip Heejun
Pubblicazione: (2025)
Provable Benefits of Task-Specific Prompts for In-context Learning
di: Chang, Xiangyu, et al.
Pubblicazione: (2025)
di: Chang, Xiangyu, et al.
Pubblicazione: (2025)
Enabling Intrinsic Reasoning over Dense Geospatial Embeddings with DFR-Gemma
di: Zhang, Xuechen, et al.
Pubblicazione: (2026)
di: Zhang, Xuechen, et al.
Pubblicazione: (2026)
Effective Length Extrapolation via Dimension-Wise Positional Embeddings Manipulation
di: Lu, Yi, et al.
Pubblicazione: (2025)
di: Lu, Yi, et al.
Pubblicazione: (2025)
SmartChunk Retrieval: Query-Aware Chunk Compression with Planning for Efficient Document RAG
di: Zhang, Xuechen, et al.
Pubblicazione: (2025)
di: Zhang, Xuechen, et al.
Pubblicazione: (2025)
Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
di: Hans, Abhimanyu, et al.
Pubblicazione: (2024)
di: Hans, Abhimanyu, et al.
Pubblicazione: (2024)
A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation (GALI)
di: Li, Yan, et al.
Pubblicazione: (2025)
di: Li, Yan, et al.
Pubblicazione: (2025)
Efficient Contextual LLM Cascades through Budget-Constrained Policy Learning
di: Zhang, Xuechen, et al.
Pubblicazione: (2024)
di: Zhang, Xuechen, et al.
Pubblicazione: (2024)
Antidistillation Sampling
di: Savani, Yash, et al.
Pubblicazione: (2025)
di: Savani, Yash, et al.
Pubblicazione: (2025)
Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
di: Li, Yingcong, et al.
Pubblicazione: (2025)
di: Li, Yingcong, et al.
Pubblicazione: (2025)
Mechanics of Next Token Prediction with Self-Attention
di: Li, Yingcong, et al.
Pubblicazione: (2024)
di: Li, Yingcong, et al.
Pubblicazione: (2024)
Extrapolation Merging: Keep Improving With Extrapolation and Merging
di: Lin, Yiguan, et al.
Pubblicazione: (2025)
di: Lin, Yiguan, et al.
Pubblicazione: (2025)
Softplus Attention with Re-weighting Boosts Length Extrapolation in Large Language Models
di: Gao, Bo, et al.
Pubblicazione: (2025)
di: Gao, Bo, et al.
Pubblicazione: (2025)
When and How Unlabeled Data Provably Improve In-Context Learning
di: Li, Yingcong, et al.
Pubblicazione: (2025)
di: Li, Yingcong, et al.
Pubblicazione: (2025)
The Role of Sparsity for Length Generalization in Transformers
di: Golowich, Noah, et al.
Pubblicazione: (2025)
di: Golowich, Noah, et al.
Pubblicazione: (2025)
Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation
di: He, Zhenyu, et al.
Pubblicazione: (2024)
di: He, Zhenyu, et al.
Pubblicazione: (2024)
DYCP: Dynamic Context Pruning for Long-Form Dialogue with LLMs
di: Choi, Nayoung, et al.
Pubblicazione: (2026)
di: Choi, Nayoung, et al.
Pubblicazione: (2026)
Transformers Can Achieve Length Generalization But Not Robustly
di: Zhou, Yongchao, et al.
Pubblicazione: (2024)
di: Zhou, Yongchao, et al.
Pubblicazione: (2024)
Easy2Hard-Bench: Standardized Difficulty Labels for Profiling LLM Performance and Generalization
di: Ding, Mucong, et al.
Pubblicazione: (2024)
di: Ding, Mucong, et al.
Pubblicazione: (2024)
Prompt-Based One-Shot Exact Length-Controlled Generation with LLMs
di: Xie, Juncheng, et al.
Pubblicazione: (2025)
di: Xie, Juncheng, et al.
Pubblicazione: (2025)
Existing Large Language Model Unlearning Evaluations Are Inconclusive
di: Feng, Zhili, et al.
Pubblicazione: (2025)
di: Feng, Zhili, et al.
Pubblicazione: (2025)
Beyond One-Size-Fits-All Summarization: Customizing Summaries for Diverse Users
di: Duran, Mehmet Samet, et al.
Pubblicazione: (2025)
di: Duran, Mehmet Samet, et al.
Pubblicazione: (2025)
Is Grokking Worthwhile? Functional Analysis and Transferability of Generalization Circuits in Transformers
di: He, Kaiyu, et al.
Pubblicazione: (2026)
di: He, Kaiyu, et al.
Pubblicazione: (2026)
Scaling Laws of RoPE-based Extrapolation
di: Liu, Xiaoran, et al.
Pubblicazione: (2023)
di: Liu, Xiaoran, et al.
Pubblicazione: (2023)
RLPR: Extrapolating RLVR to General Domains without Verifiers
di: Yu, Tianyu, et al.
Pubblicazione: (2025)
di: Yu, Tianyu, et al.
Pubblicazione: (2025)
Generating Effective Ensembles for Sentiment Analysis
di: Etelis, Itay, et al.
Pubblicazione: (2024)
di: Etelis, Itay, et al.
Pubblicazione: (2024)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
di: Wu, Wei, et al.
Pubblicazione: (2024)
di: Wu, Wei, et al.
Pubblicazione: (2024)
Large Language Models for Extrapolative Modeling of Manufacturing Processes
di: Khanghah, Kiarash Naghavi, et al.
Pubblicazione: (2025)
di: Khanghah, Kiarash Naghavi, et al.
Pubblicazione: (2025)
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure
di: Cho, Hanseul, et al.
Pubblicazione: (2024)
di: Cho, Hanseul, et al.
Pubblicazione: (2024)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
di: Yang, Wenkai, et al.
Pubblicazione: (2026)
di: Yang, Wenkai, et al.
Pubblicazione: (2026)
Language as Mathematical Structure: Examining Semantic Field Theory Against Language Games
di: Vartziotis, Dimitris
Pubblicazione: (2026)
di: Vartziotis, Dimitris
Pubblicazione: (2026)
Controlled Diversity: Length-optimized Natural Language Generation
di: Schenke, Diana Marie, et al.
Pubblicazione: (2025)
di: Schenke, Diana Marie, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
di: Lee, Nayoung, et al.
Pubblicazione: (2025) -
Everything Everywhere All at Once: LLMs can In-Context Learn Multiple Tasks in Superposition
di: Xiong, Zheyang, et al.
Pubblicazione: (2024) -
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
di: Ildiz, M. Emrullah, et al.
Pubblicazione: (2024) -
Transformers as Support Vector Machines
di: Tarzanagh, Davoud Ataee, et al.
Pubblicazione: (2023) -
From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic Data
di: Xiong, Zheyang, et al.
Pubblicazione: (2024)