Information Entropy Invariance: Enhancing Length Extrapolation in Attention Mechanisms
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Kewei, Kong, Yanwen, Xu, Yiping, Su, Jianlin, Huang, Lan, Zhang, Ruochi, Zhou, Fengfeng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DAPE V2: Process Attention Score as Feature Map for Length Extrapolation
di: Zheng, Chuanyang, et al.
Pubblicazione: (2024)
di: Zheng, Chuanyang, et al.
Pubblicazione: (2024)
A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation (GALI)
di: Li, Yan, et al.
Pubblicazione: (2025)
di: Li, Yan, et al.
Pubblicazione: (2025)
DAPE: Data-Adaptive Positional Encoding for Length Extrapolation
di: Zheng, Chuanyang, et al.
Pubblicazione: (2024)
di: Zheng, Chuanyang, et al.
Pubblicazione: (2024)
ParallelComp: Parallel Long-Context Compressor for Length Extrapolation
di: Xiong, Jing, et al.
Pubblicazione: (2025)
di: Xiong, Jing, et al.
Pubblicazione: (2025)
A Global-Local Attention Mechanism for Relation Classification
di: Sun, Yiping
Pubblicazione: (2024)
di: Sun, Yiping
Pubblicazione: (2024)
Context-aware Biases for Length Extrapolation
di: Veisi, Ali, et al.
Pubblicazione: (2025)
di: Veisi, Ali, et al.
Pubblicazione: (2025)
Bayesian Attention Mechanism: A Probabilistic Framework for Positional Encoding and Context Length Extrapolation
di: Bianchessi, Arthur S., et al.
Pubblicazione: (2025)
di: Bianchessi, Arthur S., et al.
Pubblicazione: (2025)
CLEX: Continuous Length Extrapolation for Large Language Models
di: Chen, Guanzheng, et al.
Pubblicazione: (2023)
di: Chen, Guanzheng, et al.
Pubblicazione: (2023)
Effective Length Extrapolation via Dimension-Wise Positional Embeddings Manipulation
di: Lu, Yi, et al.
Pubblicazione: (2025)
di: Lu, Yi, et al.
Pubblicazione: (2025)
Enhancing Length Extrapolation in Sequential Models with Pointer-Augmented Neural Memory
di: Le, Hung, et al.
Pubblicazione: (2024)
di: Le, Hung, et al.
Pubblicazione: (2024)
Softplus Attention with Re-weighting Boosts Length Extrapolation in Large Language Models
di: Gao, Bo, et al.
Pubblicazione: (2025)
di: Gao, Bo, et al.
Pubblicazione: (2025)
From Interpolation to Extrapolation: Complete Length Generalization for Arithmetic Transformers
di: Duan, Shaoxiong, et al.
Pubblicazione: (2023)
di: Duan, Shaoxiong, et al.
Pubblicazione: (2023)
Length Extrapolation of Transformers: A Survey from the Perspective of Positional Encoding
di: Zhao, Liang, et al.
Pubblicazione: (2023)
di: Zhao, Liang, et al.
Pubblicazione: (2023)
Efficient Length-Generalizable Attention via Causal Retrieval for Long-Context Language Modeling
di: Hu, Xiang, et al.
Pubblicazione: (2024)
di: Hu, Xiang, et al.
Pubblicazione: (2024)
Extrapolation by Association: Length Generalization Transfer in Transformers
di: Cai, Ziyang, et al.
Pubblicazione: (2025)
di: Cai, Ziyang, et al.
Pubblicazione: (2025)
Measuring Grammatical Diversity from Small Corpora: Derivational Entropy Rates, Mean Length of Utterances, and Annotation Invariance
di: Martin, Fermin Moscoso del Prado
Pubblicazione: (2024)
di: Martin, Fermin Moscoso del Prado
Pubblicazione: (2024)
DCIS: Efficient Length Extrapolation of LLMs via Divide-and-Conquer Scaling Factor Search
di: Yang, Lei, et al.
Pubblicazione: (2024)
di: Yang, Lei, et al.
Pubblicazione: (2024)
A transformer-BiGRU-based framework with data augmentation and confident learning for network intrusion detection
di: Zhang, Jiale, et al.
Pubblicazione: (2025)
di: Zhang, Jiale, et al.
Pubblicazione: (2025)
Extrapolation Merging: Keep Improving With Extrapolation and Merging
di: Lin, Yiguan, et al.
Pubblicazione: (2025)
di: Lin, Yiguan, et al.
Pubblicazione: (2025)
Fourier Position Embedding: Enhancing Attention's Periodic Extension for Length Generalization
di: Hua, Ermo, et al.
Pubblicazione: (2024)
di: Hua, Ermo, et al.
Pubblicazione: (2024)
Entropy Guided Extrapolative Decoding to Improve Factuality in Large Language Models
di: Das, Souvik, et al.
Pubblicazione: (2024)
di: Das, Souvik, et al.
Pubblicazione: (2024)
Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation
di: He, Zhenyu, et al.
Pubblicazione: (2024)
di: He, Zhenyu, et al.
Pubblicazione: (2024)
Dynamic Attention-Guided Context Decoding for Mitigating Context Faithfulness Hallucinations in Large Language Models
di: Huang, Yanwen, et al.
Pubblicazione: (2025)
di: Huang, Yanwen, et al.
Pubblicazione: (2025)
RoT: Enhancing Large Language Models with Reflection on Search Trees
di: Hui, Wenyang, et al.
Pubblicazione: (2024)
di: Hui, Wenyang, et al.
Pubblicazione: (2024)
DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration
di: Zhang, Hanzhi, et al.
Pubblicazione: (2025)
di: Zhang, Hanzhi, et al.
Pubblicazione: (2025)
Can LLMs Track Their Output Length? A Dynamic Feedback Mechanism for Precise Length Regulation
di: Xiao, Meiman, et al.
Pubblicazione: (2026)
di: Xiao, Meiman, et al.
Pubblicazione: (2026)
Enhancing Multimodal Sentiment Analysis for Missing Modality through Self-Distillation and Unified Modality Cross-Attention
di: Weng, Yuzhe, et al.
Pubblicazione: (2024)
di: Weng, Yuzhe, et al.
Pubblicazione: (2024)
Intrinsic Entropy of Context Length Scaling in LLMs
di: Shi, Jingzhe, et al.
Pubblicazione: (2025)
di: Shi, Jingzhe, et al.
Pubblicazione: (2025)
ReviewGuard: Enhancing Deficient Peer Review Detection via LLM-Driven Data Augmentation
di: Zhang, Haoxuan, et al.
Pubblicazione: (2025)
di: Zhang, Haoxuan, et al.
Pubblicazione: (2025)
Why Does the Effective Context Length of LLMs Fall Short?
di: An, Chenxin, et al.
Pubblicazione: (2024)
di: An, Chenxin, et al.
Pubblicazione: (2024)
Unveiling the Merits and Defects of LLMs in Automatic Review Generation for Scientific Papers
di: Li, Ruochi, et al.
Pubblicazione: (2025)
di: Li, Ruochi, et al.
Pubblicazione: (2025)
Position as Probability: Self-Supervised Transformers that Think Past Their Training for Length Extrapolation
di: Lee, Philip Heejun
Pubblicazione: (2025)
di: Lee, Philip Heejun
Pubblicazione: (2025)
Knocking-Heads Attention
di: Zhou, Zhanchao, et al.
Pubblicazione: (2025)
di: Zhou, Zhanchao, et al.
Pubblicazione: (2025)
GiLT: Augmenting Transformer Language Models with Dependency Graphs
di: Huang, Tianyu, et al.
Pubblicazione: (2026)
di: Huang, Tianyu, et al.
Pubblicazione: (2026)
YOCO++: Enhancing YOCO with KV Residual Connections for Efficient LLM Inference
di: Wu, You, et al.
Pubblicazione: (2026)
di: Wu, You, et al.
Pubblicazione: (2026)
LMStyle Benchmark: Evaluating Text Style Transfer for Chatbots
di: Chen, Jianlin
Pubblicazione: (2024)
di: Chen, Jianlin
Pubblicazione: (2024)
Preserving Knowledge Invariance: Rethinking Robustness Evaluation of Open Information Extraction
di: Qi, Ji, et al.
Pubblicazione: (2023)
di: Qi, Ji, et al.
Pubblicazione: (2023)
LaPA$^2$: Length-Aware Prefix and Prompt Attention Augmentation for Long-Form Controllable Text Generation
di: Yang, Jiabing, et al.
Pubblicazione: (2025)
di: Yang, Jiabing, et al.
Pubblicazione: (2025)
Various Lengths, Constant Speed: Efficient Language Modeling with Lightning Attention
di: Qin, Zhen, et al.
Pubblicazione: (2024)
di: Qin, Zhen, et al.
Pubblicazione: (2024)
Attention Consistency for LLMs Explanation
di: Lan, Tian, et al.
Pubblicazione: (2025)
di: Lan, Tian, et al.
Pubblicazione: (2025)
Documenti analoghi
-
DAPE V2: Process Attention Score as Feature Map for Length Extrapolation
di: Zheng, Chuanyang, et al.
Pubblicazione: (2024) -
A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation (GALI)
di: Li, Yan, et al.
Pubblicazione: (2025) -
DAPE: Data-Adaptive Positional Encoding for Length Extrapolation
di: Zheng, Chuanyang, et al.
Pubblicazione: (2024) -
ParallelComp: Parallel Long-Context Compressor for Length Extrapolation
di: Xiong, Jing, et al.
Pubblicazione: (2025) -
A Global-Local Attention Mechanism for Relation Classification
di: Sun, Yiping
Pubblicazione: (2024)