On Provable Length and Compositional Generalization
Fuente:
arXiv
Salvato in:
| Autori principali: | Ahuja, Kartik, Mansouri, Amin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Provable Length Generalization in Sequence Prediction via Spectral Filtering
di: Marsden, Annie, et al.
Pubblicazione: (2024)
di: Marsden, Annie, et al.
Pubblicazione: (2024)
Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability
di: Sundaram, Shobhita, et al.
Pubblicazione: (2026)
di: Sundaram, Shobhita, et al.
Pubblicazione: (2026)
Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification
di: Donhauser, Konstantin, et al.
Pubblicazione: (2025)
di: Donhauser, Konstantin, et al.
Pubblicazione: (2025)
Synthetic Data Generation and Joint Learning for Robust Code-Mixed Translation
di: Kartik, Kartik, et al.
Pubblicazione: (2024)
di: Kartik, Kartik, et al.
Pubblicazione: (2024)
Improving Variable-Length Generation in Diffusion Language Models via Length Regularization
di: Cheng, Zicong, et al.
Pubblicazione: (2026)
di: Cheng, Zicong, et al.
Pubblicazione: (2026)
Provably Learning from Language Feedback
di: Xu, Wanqiao, et al.
Pubblicazione: (2025)
di: Xu, Wanqiao, et al.
Pubblicazione: (2025)
In-Context Learning through the Bayesian Prism
di: Panwar, Madhur, et al.
Pubblicazione: (2023)
di: Panwar, Madhur, et al.
Pubblicazione: (2023)
From Interpolation to Extrapolation: Complete Length Generalization for Arithmetic Transformers
di: Duan, Shaoxiong, et al.
Pubblicazione: (2023)
di: Duan, Shaoxiong, et al.
Pubblicazione: (2023)
Explicitly Encoding Structural Symmetry is Key to Length Generalization in Arithmetic Tasks
di: Sabbaghi, Mahdi, et al.
Pubblicazione: (2024)
di: Sabbaghi, Mahdi, et al.
Pubblicazione: (2024)
Task-Aware Calibration: Provably Optimal Decoding in LLMs
di: Tomov, Tim, et al.
Pubblicazione: (2026)
di: Tomov, Tim, et al.
Pubblicazione: (2026)
Attention with Trained Embeddings Provably Selects Important Tokens
di: Wu, Diyuan, et al.
Pubblicazione: (2025)
di: Wu, Diyuan, et al.
Pubblicazione: (2025)
Provable Knowledge Acquisition and Extraction in One-Layer Transformers
di: Xu, Ruichen, et al.
Pubblicazione: (2025)
di: Xu, Ruichen, et al.
Pubblicazione: (2025)
The Role of Sparsity for Length Generalization in Transformers
di: Golowich, Noah, et al.
Pubblicazione: (2025)
di: Golowich, Noah, et al.
Pubblicazione: (2025)
Learning Syntax Without Planting Trees: Understanding Hierarchical Generalization in Transformers
di: Ahuja, Kabir, et al.
Pubblicazione: (2024)
di: Ahuja, Kabir, et al.
Pubblicazione: (2024)
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
di: Jung, Jaehun, et al.
Pubblicazione: (2024)
di: Jung, Jaehun, et al.
Pubblicazione: (2024)
Provably Robust DPO: Aligning Language Models with Noisy Feedback
di: Chowdhury, Sayak Ray, et al.
Pubblicazione: (2024)
di: Chowdhury, Sayak Ray, et al.
Pubblicazione: (2024)
OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention
di: Zhou, Chenyu, et al.
Pubblicazione: (2026)
di: Zhou, Chenyu, et al.
Pubblicazione: (2026)
Transformers Can Achieve Length Generalization But Not Robustly
di: Zhou, Yongchao, et al.
Pubblicazione: (2024)
di: Zhou, Yongchao, et al.
Pubblicazione: (2024)
RadLite: Multi-Task LoRA Fine-Tuning of Small Language Models for CPU-Deployable Radiology AI
di: Gupta, Pankaj, et al.
Pubblicazione: (2026)
di: Gupta, Pankaj, et al.
Pubblicazione: (2026)
COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamics
di: Sharma, Kartik, et al.
Pubblicazione: (2026)
di: Sharma, Kartik, et al.
Pubblicazione: (2026)
Length Desensitization in Direct Preference Optimization
di: Liu, Wei, et al.
Pubblicazione: (2024)
di: Liu, Wei, et al.
Pubblicazione: (2024)
Efficient Knowledge Probing of Large Language Models by Adapting Pre-trained Embeddings
di: Sharma, Kartik, et al.
Pubblicazione: (2025)
di: Sharma, Kartik, et al.
Pubblicazione: (2025)
Diffusion Language Models Are Natively Length-Aware
di: Rossi, Vittorio, et al.
Pubblicazione: (2026)
di: Rossi, Vittorio, et al.
Pubblicazione: (2026)
Intrinsic Entropy of Context Length Scaling in LLMs
di: Shi, Jingzhe, et al.
Pubblicazione: (2025)
di: Shi, Jingzhe, et al.
Pubblicazione: (2025)
Escaping the Mode Lottery: Multi-Response Training Improves Language Model Generalization
di: Amin, Hasan, et al.
Pubblicazione: (2026)
di: Amin, Hasan, et al.
Pubblicazione: (2026)
An Empirical Study on Context Length for Open-Domain Dialog Generation
di: Shen, Xinyi, et al.
Pubblicazione: (2024)
di: Shen, Xinyi, et al.
Pubblicazione: (2024)
Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
di: Leng, Jiaqi, et al.
Pubblicazione: (2025)
di: Leng, Jiaqi, et al.
Pubblicazione: (2025)
Explaining Length Bias in LLM-Based Preference Evaluations
di: Hu, Zhengyu, et al.
Pubblicazione: (2024)
di: Hu, Zhengyu, et al.
Pubblicazione: (2024)
Disentangling Length from Quality in Direct Preference Optimization
di: Park, Ryan, et al.
Pubblicazione: (2024)
di: Park, Ryan, et al.
Pubblicazione: (2024)
Controlling Summarization Length Through EOS Token Weighting
di: Belligoli, Zeno, et al.
Pubblicazione: (2025)
di: Belligoli, Zeno, et al.
Pubblicazione: (2025)
TRA: Better Length Generalisation with Threshold Relative Attention
di: Opper, Mattia, et al.
Pubblicazione: (2025)
di: Opper, Mattia, et al.
Pubblicazione: (2025)
Provable Interactive Learning with Hindsight Instruction Feedback
di: Misra, Dipendra, et al.
Pubblicazione: (2024)
di: Misra, Dipendra, et al.
Pubblicazione: (2024)
Reinforcement Learning for Compositional Generalization with Outcome-Level Optimization
di: Fu, Xiyan, et al.
Pubblicazione: (2026)
di: Fu, Xiyan, et al.
Pubblicazione: (2026)
Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
di: Ma, Xuezhe, et al.
Pubblicazione: (2024)
di: Ma, Xuezhe, et al.
Pubblicazione: (2024)
Hansel: Output Length Controlling Framework for Large Language Models
di: Song, Seoha, et al.
Pubblicazione: (2024)
di: Song, Seoha, et al.
Pubblicazione: (2024)
Parallelizing Linear Transformers with the Delta Rule over Sequence Length
di: Yang, Songlin, et al.
Pubblicazione: (2024)
di: Yang, Songlin, et al.
Pubblicazione: (2024)
SELF: Self-Extend the Context Length With Logistic Growth Function
di: Dang, Phat Thanh, et al.
Pubblicazione: (2025)
di: Dang, Phat Thanh, et al.
Pubblicazione: (2025)
A Minimum Description Length Approach to Regularization in Neural Networks
di: Abudy, Matan, et al.
Pubblicazione: (2025)
di: Abudy, Matan, et al.
Pubblicazione: (2025)
Diffusion LMs Can Approximate Optimal Infilling Lengths Implicitly
di: Liu, Hengchang, et al.
Pubblicazione: (2026)
di: Liu, Hengchang, et al.
Pubblicazione: (2026)
A Long Way to Go: Investigating Length Correlations in RLHF
di: Singhal, Prasann, et al.
Pubblicazione: (2023)
di: Singhal, Prasann, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Provable Length Generalization in Sequence Prediction via Spectral Filtering
di: Marsden, Annie, et al.
Pubblicazione: (2024) -
Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability
di: Sundaram, Shobhita, et al.
Pubblicazione: (2026) -
Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification
di: Donhauser, Konstantin, et al.
Pubblicazione: (2025) -
Synthetic Data Generation and Joint Learning for Robust Code-Mixed Translation
di: Kartik, Kartik, et al.
Pubblicazione: (2024) -
Improving Variable-Length Generation in Diffusion Language Models via Length Regularization
di: Cheng, Zicong, et al.
Pubblicazione: (2026)