A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation (GALI)
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Yan, Zhang, Tianyi, Li, Zechuan, Han, Soyeon Caren |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Multimodal Commonsense Knowledge Distillation for Visual Question Answering
di: Yang, Shuo, et al.
Pubblicazione: (2024)
di: Yang, Shuo, et al.
Pubblicazione: (2024)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
di: Xiao, Chaojun, et al.
Pubblicazione: (2024)
di: Xiao, Chaojun, et al.
Pubblicazione: (2024)
Inclusion-of-Thoughts: Mitigating Preference Instability via Purifying the Decision Space
di: Madani, Mohammad Reza Ghasemi, et al.
Pubblicazione: (2026)
di: Madani, Mohammad Reza Ghasemi, et al.
Pubblicazione: (2026)
Extrapolation by Association: Length Generalization Transfer in Transformers
di: Cai, Ziyang, et al.
Pubblicazione: (2025)
di: Cai, Ziyang, et al.
Pubblicazione: (2025)
Effective Length Extrapolation via Dimension-Wise Positional Embeddings Manipulation
di: Lu, Yi, et al.
Pubblicazione: (2025)
di: Lu, Yi, et al.
Pubblicazione: (2025)
ReAttention: Training-Free Infinite Context with Finite Attention Scope
di: Liu, Xiaoran, et al.
Pubblicazione: (2024)
di: Liu, Xiaoran, et al.
Pubblicazione: (2024)
Softplus Attention with Re-weighting Boosts Length Extrapolation in Large Language Models
di: Gao, Bo, et al.
Pubblicazione: (2025)
di: Gao, Bo, et al.
Pubblicazione: (2025)
The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
di: Song, Yifan, et al.
Pubblicazione: (2024)
di: Song, Yifan, et al.
Pubblicazione: (2024)
Token Prepending: A Training-Free Approach for Eliciting Better Sentence Embeddings from LLMs
di: Fu, Yuchen, et al.
Pubblicazione: (2024)
di: Fu, Yuchen, et al.
Pubblicazione: (2024)
Logit Arithmetic Elicits Long Reasoning Capabilities Without Training
di: Zhang, Yunxiang, et al.
Pubblicazione: (2025)
di: Zhang, Yunxiang, et al.
Pubblicazione: (2025)
Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models
di: Qin, Zhen, et al.
Pubblicazione: (2024)
di: Qin, Zhen, et al.
Pubblicazione: (2024)
Extrapolation Merging: Keep Improving With Extrapolation and Merging
di: Lin, Yiguan, et al.
Pubblicazione: (2025)
di: Lin, Yiguan, et al.
Pubblicazione: (2025)
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
di: Li, Jin, et al.
Pubblicazione: (2025)
di: Li, Jin, et al.
Pubblicazione: (2025)
Position as Probability: Self-Supervised Transformers that Think Past Their Training for Length Extrapolation
di: Lee, Philip Heejun
Pubblicazione: (2025)
di: Lee, Philip Heejun
Pubblicazione: (2025)
What Happened in LLMs Layers when Trained for Fast vs. Slow Thinking: A Gradient Perspective
di: Li, Ming, et al.
Pubblicazione: (2024)
di: Li, Ming, et al.
Pubblicazione: (2024)
From Attribution to Abstention: Training-Free Attention-Based Auditing for Clinical Summarization
di: Yan, Qianqi, et al.
Pubblicazione: (2026)
di: Yan, Qianqi, et al.
Pubblicazione: (2026)
VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding
di: Ding, Yihao, et al.
Pubblicazione: (2025)
di: Ding, Yihao, et al.
Pubblicazione: (2025)
Correlation-Aware Select and Merge Attention for Efficient Fine-Tuning and Context Length Extension
di: Wang, Ning, et al.
Pubblicazione: (2024)
di: Wang, Ning, et al.
Pubblicazione: (2024)
ROME: Memorization Insights from Text, Logits and Representation
di: Li, Bo, et al.
Pubblicazione: (2024)
di: Li, Bo, et al.
Pubblicazione: (2024)
Scaling Laws of RoPE-based Extrapolation
di: Liu, Xiaoran, et al.
Pubblicazione: (2023)
di: Liu, Xiaoran, et al.
Pubblicazione: (2023)
Debiasing LLMs by Masking Unfairness-Driving Attention Heads
di: Han, Tingxu, et al.
Pubblicazione: (2025)
di: Han, Tingxu, et al.
Pubblicazione: (2025)
How Well do LLMs Compress Their Own Chain-of-Thought? A Token Complexity Approach
di: Lee, Ayeong, et al.
Pubblicazione: (2025)
di: Lee, Ayeong, et al.
Pubblicazione: (2025)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
di: Wu, Wei, et al.
Pubblicazione: (2024)
di: Wu, Wei, et al.
Pubblicazione: (2024)
LongSkywork: A Training Recipe for Efficiently Extending Context Length in Large Language Models
di: Zhao, Liang, et al.
Pubblicazione: (2024)
di: Zhao, Liang, et al.
Pubblicazione: (2024)
Attention-Aligned Reasoning for Large Language Models
di: Zhang, Hongxiang, et al.
Pubblicazione: (2025)
di: Zhang, Hongxiang, et al.
Pubblicazione: (2025)
Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation
di: He, Zhenyu, et al.
Pubblicazione: (2024)
di: He, Zhenyu, et al.
Pubblicazione: (2024)
DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs
di: Papi, Sara, et al.
Pubblicazione: (2026)
di: Papi, Sara, et al.
Pubblicazione: (2026)
Summarize Before You Speak with ARACH: A Training-Free Inference-Time Plug-In for Enhancing LLMs via Global Attention Reallocation
di: Wang, Jingtao, et al.
Pubblicazione: (2026)
di: Wang, Jingtao, et al.
Pubblicazione: (2026)
Strategic Deflection: Defending LLMs from Logit Manipulation
di: Rachidy, Yassine, et al.
Pubblicazione: (2025)
di: Rachidy, Yassine, et al.
Pubblicazione: (2025)
TOFA: Training-Free One-Shot Federated Adaptation for Vision-Language Models
di: Zhang, Li, et al.
Pubblicazione: (2025)
di: Zhang, Li, et al.
Pubblicazione: (2025)
Optimizing Length Compression in Large Reasoning Models
di: Cheng, Zhengxiang, et al.
Pubblicazione: (2025)
di: Cheng, Zhengxiang, et al.
Pubblicazione: (2025)
Steering Language Models Before They Speak: Logit-Level Interventions
di: An, Hyeseon, et al.
Pubblicazione: (2026)
di: An, Hyeseon, et al.
Pubblicazione: (2026)
UltraLLaDA: Scaling the Context Length to 128K for Diffusion Large Language Models
di: He, Guangxin, et al.
Pubblicazione: (2025)
di: He, Guangxin, et al.
Pubblicazione: (2025)
EpiCoDe: Boosting Model Performance Beyond Training with Extrapolation and Contrastive Decoding
di: Tao, Mingxu, et al.
Pubblicazione: (2025)
di: Tao, Mingxu, et al.
Pubblicazione: (2025)
BiLD: Bi-directional Logits Difference Loss for Large Language Model Distillation
di: Li, Minchong, et al.
Pubblicazione: (2024)
di: Li, Minchong, et al.
Pubblicazione: (2024)
Fourier Position Embedding: Enhancing Attention's Periodic Extension for Length Generalization
di: Hua, Ermo, et al.
Pubblicazione: (2024)
di: Hua, Ermo, et al.
Pubblicazione: (2024)
Divide, Reweight, and Conquer: A Logit Arithmetic Approach for In-Context Learning
di: Huang, Chengsong, et al.
Pubblicazione: (2024)
di: Huang, Chengsong, et al.
Pubblicazione: (2024)
When More Is Less: A Systematic Analysis of Spatial and Commonsense Information for Visual Spatial Reasoning
di: Akasaka, Muku, et al.
Pubblicazione: (2026)
di: Akasaka, Muku, et al.
Pubblicazione: (2026)
Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps
di: Zhou, Yanke, et al.
Pubblicazione: (2026)
di: Zhou, Yanke, et al.
Pubblicazione: (2026)
GenEOL: Harnessing the Generative Power of LLMs for Training-Free Sentence Embeddings
di: Thirukovalluru, Raghuveer, et al.
Pubblicazione: (2024)
di: Thirukovalluru, Raghuveer, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Multimodal Commonsense Knowledge Distillation for Visual Question Answering
di: Yang, Shuo, et al.
Pubblicazione: (2024) -
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
di: Xiao, Chaojun, et al.
Pubblicazione: (2024) -
Inclusion-of-Thoughts: Mitigating Preference Instability via Purifying the Decision Space
di: Madani, Mohammad Reza Ghasemi, et al.
Pubblicazione: (2026) -
Extrapolation by Association: Length Generalization Transfer in Transformers
di: Cai, Ziyang, et al.
Pubblicazione: (2025) -
Effective Length Extrapolation via Dimension-Wise Positional Embeddings Manipulation
di: Lu, Yi, et al.
Pubblicazione: (2025)