Non-Asymptotic Length Generalization
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Thomas, Ma, Tengyu, Li, Zhiyuan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Chain of Thought Empowers Transformers to Solve Inherently Serial Problems
por: Li, Zhiyuan, et al.
Publicado: (2024)
por: Li, Zhiyuan, et al.
Publicado: (2024)
Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
por: Liu, Hong, et al.
Publicado: (2023)
por: Liu, Hong, et al.
Publicado: (2023)
Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective
por: Wen, Kaiyue, et al.
Publicado: (2024)
por: Wen, Kaiyue, et al.
Publicado: (2024)
Linguistic Calibration of Long-Form Generations
por: Band, Neil, et al.
Publicado: (2024)
por: Band, Neil, et al.
Publicado: (2024)
Configuration-to-Performance Scaling Law with Neural Ansatz
por: Zhang, Huaqing, et al.
Publicado: (2026)
por: Zhang, Huaqing, et al.
Publicado: (2026)
Formal Theorem Proving by Rewarding LLMs to Decompose Proofs Hierarchically
por: Dong, Kefan, et al.
Publicado: (2024)
por: Dong, Kefan, et al.
Publicado: (2024)
Divide-and-Conquer CoT: RL for Reducing Latency via Parallel Reasoning
por: Mahankali, Arvind, et al.
Publicado: (2026)
por: Mahankali, Arvind, et al.
Publicado: (2026)
STP: Self-play LLM Theorem Provers with Iterative Conjecturing and Proving
por: Dong, Kefan, et al.
Publicado: (2025)
por: Dong, Kefan, et al.
Publicado: (2025)
A Theoretical Framework for Self-Play Theorem Proving Algorithms
por: Chen, Thomas, et al.
Publicado: (2026)
por: Chen, Thomas, et al.
Publicado: (2026)
Towards Faster Non-Asymptotic Convergence for Diffusion-Based Generative Models
por: Li, Gen, et al.
Publicado: (2023)
por: Li, Gen, et al.
Publicado: (2023)
Fantastic Pretraining Optimizers and Where to Find Them
por: Wen, Kaiyue, et al.
Publicado: (2025)
por: Wen, Kaiyue, et al.
Publicado: (2025)
Scaling Self-Play with Self-Guidance
por: Bailey, Luke, et al.
Publicado: (2026)
por: Bailey, Luke, et al.
Publicado: (2026)
Low-Dimension-to-High-Dimension Generalization And Its Implications for Length Generalization
por: Chen, Yang, et al.
Publicado: (2024)
por: Chen, Yang, et al.
Publicado: (2024)
Looped Transformers for Length Generalization
por: Fan, Ying, et al.
Publicado: (2024)
por: Fan, Ying, et al.
Publicado: (2024)
Non-Asymptotic Analysis of (Sticky) Track-and-Stop
por: Poiani, Riccardo, et al.
Publicado: (2025)
por: Poiani, Riccardo, et al.
Publicado: (2025)
Non-Asymptotic Analysis of Efficiency in Conformalized Regression
por: Yao, Yunzhen, et al.
Publicado: (2025)
por: Yao, Yunzhen, et al.
Publicado: (2025)
On Vanishing Variance in Transformer Length Generalization
por: Li, Ruining, et al.
Publicado: (2025)
por: Li, Ruining, et al.
Publicado: (2025)
Large Language Models as Tool Makers
por: Cai, Tianle, et al.
Publicado: (2023)
por: Cai, Tianle, et al.
Publicado: (2023)
Non-Asymptotic Convergence of Stochastic Iterative Algorithms: A Lyapunov Framework
por: Chen, Zaiwei, et al.
Publicado: (2026)
por: Chen, Zaiwei, et al.
Publicado: (2026)
Asymptotics of Non-Convex Generalized Linear Models in High-Dimensions: A proof of the replica formula
por: Vilucchio, Matteo, et al.
Publicado: (2025)
por: Vilucchio, Matteo, et al.
Publicado: (2025)
Quantitative Bounds for Length Generalization in Transformers
por: Izzo, Zachary, et al.
Publicado: (2025)
por: Izzo, Zachary, et al.
Publicado: (2025)
Universal Length Generalization with Turing Programs
por: Hou, Kaiying, et al.
Publicado: (2024)
por: Hou, Kaiying, et al.
Publicado: (2024)
Improving Variable-Length Generation in Diffusion Language Models via Length Regularization
por: Cheng, Zicong, et al.
Publicado: (2026)
por: Cheng, Zicong, et al.
Publicado: (2026)
On the Limitations and Capabilities of Position Embeddings for Length Generalization
por: Chen, Yang, et al.
Publicado: (2025)
por: Chen, Yang, et al.
Publicado: (2025)
Mamba Modulation: On the Length Generalization of Mamba
por: Lu, Peng, et al.
Publicado: (2025)
por: Lu, Peng, et al.
Publicado: (2025)
Flight Trajectory Prediction Using an Enhanced CNN-LSTM Network
por: Hao, Qinzhi, et al.
Publicado: (2024)
por: Hao, Qinzhi, et al.
Publicado: (2024)
Fighter flight trajectory prediction based on spatio-temporal graphcial attention network
por: Sun, Yao, et al.
Publicado: (2024)
por: Sun, Yao, et al.
Publicado: (2024)
Non-Asymptotic Optimization and Generalization Bounds for Stochastic Gauss-Newton in Overparameterized Models
por: Cayci, Semih
Publicado: (2025)
por: Cayci, Semih
Publicado: (2025)
Comparative Study on Semi-supervised Learning Applied for Anomaly Detection in Hydraulic Condition Monitoring System
por: Dong, Yongqi, et al.
Publicado: (2023)
por: Dong, Yongqi, et al.
Publicado: (2023)
Understanding and Improving Length Generalization in Recurrent Models
por: Ruiz, Ricardo Buitrago, et al.
Publicado: (2025)
por: Ruiz, Ricardo Buitrago, et al.
Publicado: (2025)
Learning Variable-Length Tokenization for Generative Recommendation
por: Wang, Minhao, et al.
Publicado: (2026)
por: Wang, Minhao, et al.
Publicado: (2026)
Length Generalization with Log-Depth Recurrent Units
por: Pert, Charles, et al.
Publicado: (2026)
por: Pert, Charles, et al.
Publicado: (2026)
No Free Lunch: Non-Asymptotic Analysis of Prediction-Powered Inference
por: Mani, Pranav, et al.
Publicado: (2025)
por: Mani, Pranav, et al.
Publicado: (2025)
Understanding Inverse Reinforcement Learning under Overparameterization: Non-Asymptotic Analysis and Global Optimality
por: Zhang, Ruijia, et al.
Publicado: (2025)
por: Zhang, Ruijia, et al.
Publicado: (2025)
A Tale of Two Geometries: Adaptive Optimizers and Non-Euclidean Descent
por: Xie, Shuo, et al.
Publicado: (2025)
por: Xie, Shuo, et al.
Publicado: (2025)
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
por: Cho, Hanseul, et al.
Publicado: (2024)
por: Cho, Hanseul, et al.
Publicado: (2024)
Pseudo-Formalization for Automatic Proof Verification
por: Barkallah, Slim, et al.
Publicado: (2026)
por: Barkallah, Slim, et al.
Publicado: (2026)
On Provable Length and Compositional Generalization
por: Ahuja, Kartik, et al.
Publicado: (2024)
por: Ahuja, Kartik, et al.
Publicado: (2024)
Provably Minimum-Length Conformal Prediction Sets for Ordinal Classification
por: Zhang, Zijian, et al.
Publicado: (2025)
por: Zhang, Zijian, et al.
Publicado: (2025)
Non-Asymptotic Global Convergence of PPO-Clip
por: Liu, Yin, et al.
Publicado: (2025)
por: Liu, Yin, et al.
Publicado: (2025)
Ejemplares similares
-
Chain of Thought Empowers Transformers to Solve Inherently Serial Problems
por: Li, Zhiyuan, et al.
Publicado: (2024) -
Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
por: Liu, Hong, et al.
Publicado: (2023) -
Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective
por: Wen, Kaiyue, et al.
Publicado: (2024) -
Linguistic Calibration of Long-Form Generations
por: Band, Neil, et al.
Publicado: (2024) -
Configuration-to-Performance Scaling Law with Neural Ansatz
por: Zhang, Huaqing, et al.
Publicado: (2026)