Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for Transformers
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Shaw, Peter, Cohan, James, Eisenstein, Jacob, Toutanova, Kristina |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ALTA: Compiler-Based Analysis of Transformers
par: Shaw, Peter, et autres
Publié: (2024)
par: Shaw, Peter, et autres
Publié: (2024)
BgGPT 1.0: Extending English-centric LLMs to other languages
par: Alexandrov, Anton, et autres
Publié: (2024)
par: Alexandrov, Anton, et autres
Publié: (2024)
The Role of Sparsity for Length Generalization in Transformers
par: Golowich, Noah, et autres
Publié: (2025)
par: Golowich, Noah, et autres
Publié: (2025)
On the Optimal Reasoning Length for RL-Trained Language Models
par: Nohara, Daisuke, et autres
Publié: (2026)
par: Nohara, Daisuke, et autres
Publié: (2026)
Transformers Can Achieve Length Generalization But Not Robustly
par: Zhou, Yongchao, et autres
Publié: (2024)
par: Zhou, Yongchao, et autres
Publié: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
par: Liu, Yixin, et autres
Publié: (2025)
par: Liu, Yixin, et autres
Publié: (2025)
Reward-free Alignment for Conflicting Objectives
par: Chen, Peter, et autres
Publié: (2026)
par: Chen, Peter, et autres
Publié: (2026)
Algorithmic Capabilities of Random Transformers
par: Zhong, Ziqian, et autres
Publié: (2024)
par: Zhong, Ziqian, et autres
Publié: (2024)
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure
par: Cho, Hanseul, et autres
Publié: (2024)
par: Cho, Hanseul, et autres
Publié: (2024)
Observable Propagation: Uncovering Feature Vectors in Transformers
par: Dunefsky, Jacob, et autres
Publié: (2023)
par: Dunefsky, Jacob, et autres
Publié: (2023)
On the Hidden Objective Biases of Group-based Reinforcement Learning
par: Fontana, Aleksandar, et autres
Publié: (2026)
par: Fontana, Aleksandar, et autres
Publié: (2026)
Calibrating Long-form Generations from Large Language Models
par: Huang, Yukun, et autres
Publié: (2024)
par: Huang, Yukun, et autres
Publié: (2024)
Language and Experience: A Computational Model of Social Learning in Complex Tasks
par: Colas, Cédric, et autres
Publié: (2025)
par: Colas, Cédric, et autres
Publié: (2025)
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
par: Liu, Wei, et autres
Publié: (2025)
par: Liu, Wei, et autres
Publié: (2025)
How Does Response Length Affect Long-Form Factuality
par: Zhao, James Xu, et autres
Publié: (2025)
par: Zhao, James Xu, et autres
Publié: (2025)
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models
par: Belenki, Lior, et autres
Publié: (2025)
par: Belenki, Lior, et autres
Publié: (2025)
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference
par: Gao, Mingqi, et autres
Publié: (2024)
par: Gao, Mingqi, et autres
Publié: (2024)
BIRCO: A Benchmark of Information Retrieval Tasks with Complex Objectives
par: Wang, Xiaoyue, et autres
Publié: (2024)
par: Wang, Xiaoyue, et autres
Publié: (2024)
References Improve LLM Alignment in Non-Verifiable Domains
par: Shi, Kejian, et autres
Publié: (2026)
par: Shi, Kejian, et autres
Publié: (2026)
Length-MAX Tokenizer for Language Models
par: Dong, Dong, et autres
Publié: (2025)
par: Dong, Dong, et autres
Publié: (2025)
EMORL: Ensemble Multi-Objective Reinforcement Learning for Efficient and Flexible LLM Fine-Tuning
par: Kong, Lingxiao, et autres
Publié: (2025)
par: Kong, Lingxiao, et autres
Publié: (2025)
More is not always better? Enhancing Many-Shot In-Context Learning with Differentiated and Reweighting Objectives
par: Zhang, Xiaoqing, et autres
Publié: (2025)
par: Zhang, Xiaoqing, et autres
Publié: (2025)
Effective Reasoning Chains Reduce Intrinsic Dimensionality
par: Prasad, Archiki, et autres
Publié: (2026)
par: Prasad, Archiki, et autres
Publié: (2026)
Which Attention Heads Matter for In-Context Learning?
par: Yin, Kayo, et autres
Publié: (2025)
par: Yin, Kayo, et autres
Publié: (2025)
RLP: Reinforcement as a Pretraining Objective
par: Hatamizadeh, Ali, et autres
Publié: (2025)
par: Hatamizadeh, Ali, et autres
Publié: (2025)
SpanNorm: Reconciling Training Stability and Performance in Deep Transformers
par: Wang, Chao, et autres
Publié: (2026)
par: Wang, Chao, et autres
Publié: (2026)
Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning
par: Hwang, Jaedong, et autres
Publié: (2025)
par: Hwang, Jaedong, et autres
Publié: (2025)
The Anxiety of Influence: Bloom Filters in Transformer Attention Heads
par: Balogh, Peter
Publié: (2026)
par: Balogh, Peter
Publié: (2026)
ComplexityNet: Increasing LLM Inference Efficiency by Learning Task Complexity
par: Bae, Henry, et autres
Publié: (2023)
par: Bae, Henry, et autres
Publié: (2023)
Multi-Objective Large Language Model Unlearning
par: Pan, Zibin, et autres
Publié: (2024)
par: Pan, Zibin, et autres
Publié: (2024)
Pareto Multi-Objective Alignment for Language Models
par: He, Qiang, et autres
Publié: (2025)
par: He, Qiang, et autres
Publié: (2025)
DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning
par: Liu, Shih-Yang, et autres
Publié: (2025)
par: Liu, Shih-Yang, et autres
Publié: (2025)
Stepwise Penalization for Length-Efficient Chain-of-Thought Reasoning
par: Li, Xintong, et autres
Publié: (2026)
par: Li, Xintong, et autres
Publié: (2026)
Position as Probability: Self-Supervised Transformers that Think Past Their Training for Length Extrapolation
par: Lee, Philip Heejun
Publié: (2025)
par: Lee, Philip Heejun
Publié: (2025)
Bridging the Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs
par: Kumar, Somnath, et autres
Publié: (2024)
par: Kumar, Somnath, et autres
Publié: (2024)
Transformers Struggle to Learn to Search
par: Saparov, Abulhair, et autres
Publié: (2024)
par: Saparov, Abulhair, et autres
Publié: (2024)
When More is Less: Understanding Chain-of-Thought Length in LLMs
par: Wu, Yuyang, et autres
Publié: (2025)
par: Wu, Yuyang, et autres
Publié: (2025)
Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
par: Leng, Jiaqi, et autres
Publié: (2025)
par: Leng, Jiaqi, et autres
Publié: (2025)
An Empirical Study on Context Length for Open-Domain Dialog Generation
par: Shen, Xinyi, et autres
Publié: (2024)
par: Shen, Xinyi, et autres
Publié: (2024)
Provable Length Generalization in Sequence Prediction via Spectral Filtering
par: Marsden, Annie, et autres
Publié: (2024)
par: Marsden, Annie, et autres
Publié: (2024)
Documents similaires
-
ALTA: Compiler-Based Analysis of Transformers
par: Shaw, Peter, et autres
Publié: (2024) -
BgGPT 1.0: Extending English-centric LLMs to other languages
par: Alexandrov, Anton, et autres
Publié: (2024) -
The Role of Sparsity for Length Generalization in Transformers
par: Golowich, Noah, et autres
Publié: (2025) -
On the Optimal Reasoning Length for RL-Trained Language Models
par: Nohara, Daisuke, et autres
Publié: (2026) -
Transformers Can Achieve Length Generalization But Not Robustly
par: Zhou, Yongchao, et autres
Publié: (2024)