Sharper Generalization Bounds for Transformer
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yawen, Hu, Tao, Lian, Zhouhui, Tian, Wan, Peng, Yijie, Zhang, Huiming, Li, Zhongyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Sharper Picture of Generalization in Transformers
von: Lintilhac, Paul, et al.
Veröffentlicht: (2026)
von: Lintilhac, Paul, et al.
Veröffentlicht: (2026)
Sharper Error Bounds in Late Fusion Multi-view Clustering Using Eigenvalue Proportion
von: Du, Liang, et al.
Veröffentlicht: (2024)
von: Du, Liang, et al.
Veröffentlicht: (2024)
Tail-Aware Information-Theoretic Generalization for RLHF and SGLD
von: Zhang, Huiming, et al.
Veröffentlicht: (2026)
von: Zhang, Huiming, et al.
Veröffentlicht: (2026)
Machine Learning-Assisted High-Dimensional Matrix Estimation
von: Tian, Wan, et al.
Veröffentlicht: (2026)
von: Tian, Wan, et al.
Veröffentlicht: (2026)
From Small to Large: Generalization Bounds for Transformers on Variable-Size Inputs
von: Alokhina, Anastasiia, et al.
Veröffentlicht: (2025)
von: Alokhina, Anastasiia, et al.
Veröffentlicht: (2025)
LLM-Inspired Pretrain-Then-Finetune for Small-Data, Large-Scale Optimization
von: Zhang, Zishi, et al.
Veröffentlicht: (2026)
von: Zhang, Zishi, et al.
Veröffentlicht: (2026)
Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails
von: Jin, Ruinan, et al.
Veröffentlicht: (2026)
von: Jin, Ruinan, et al.
Veröffentlicht: (2026)
RiskPO: Risk-based Policy Optimization via Verifiable Reward for LLM Post-Training
von: Ren, Tao, et al.
Veröffentlicht: (2025)
von: Ren, Tao, et al.
Veröffentlicht: (2025)
FLOPS: Forward Learning with OPtimal Sampling
von: Ren, Tao, et al.
Veröffentlicht: (2024)
von: Ren, Tao, et al.
Veröffentlicht: (2024)
Disentanglement of Variations with Multimodal Generative Modeling
von: Zhang, Yijie, et al.
Veröffentlicht: (2025)
von: Zhang, Yijie, et al.
Veröffentlicht: (2025)
CAST-CKT: Chaos-Aware Spatio-Temporal and Cross-City Knowledge Transfer for Traffic Flow Prediction
von: Fofanah, Abdul Joseph, et al.
Veröffentlicht: (2026)
von: Fofanah, Abdul Joseph, et al.
Veröffentlicht: (2026)
HELM: Harness-Enhanced Long-horizon Memory for Vision-Language-Action Manipulation
von: Zeng, Zijian, et al.
Veröffentlicht: (2026)
von: Zeng, Zijian, et al.
Veröffentlicht: (2026)
Representational Homomorphism Predicts and Improves Compositional Generalization In Transformer Language Model
von: An, Zhiyu, et al.
Veröffentlicht: (2026)
von: An, Zhiyu, et al.
Veröffentlicht: (2026)
Towards Mitigation of Hallucination for LLM-empowered Agents: Progressive Generalization Bound Exploration and Watchdog Monitor
von: Liu, Siyuan, et al.
Veröffentlicht: (2025)
von: Liu, Siyuan, et al.
Veröffentlicht: (2025)
Traj-Transformer: Diffusion Models with Transformer for GPS Trajectory Generation
von: Zhang, Zhiyang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhiyang, et al.
Veröffentlicht: (2025)
Efficient Generative Model Training via Embedded Representation Warmup
von: Liu, Deyuan, et al.
Veröffentlicht: (2025)
von: Liu, Deyuan, et al.
Veröffentlicht: (2025)
Task-Aware Harmony Multi-Task Decision Transformer for Offline Reinforcement Learning
von: Fan, Ziqing, et al.
Veröffentlicht: (2024)
von: Fan, Ziqing, et al.
Veröffentlicht: (2024)
From Uniform to Adaptive: General Skip-Block Mechanisms for Efficient PDE Neural Operators
von: Liu, Lei, et al.
Veröffentlicht: (2025)
von: Liu, Lei, et al.
Veröffentlicht: (2025)
LIFT: Interpretable truck driving risk prediction with literature-informed fine-tuned LLMs
von: Hu, Xiao, et al.
Veröffentlicht: (2025)
von: Hu, Xiao, et al.
Veröffentlicht: (2025)
CHARM: Calibrating Reward Models With Chatbot Arena Scores
von: Zhu, Xiao, et al.
Veröffentlicht: (2025)
von: Zhu, Xiao, et al.
Veröffentlicht: (2025)
Selective Reviews of Bandit Problems in AI via a Statistical View
von: Zhou, Pengjie, et al.
Veröffentlicht: (2024)
von: Zhou, Pengjie, et al.
Veröffentlicht: (2024)
Prompt Tuning with Diffusion for Few-Shot Pre-trained Policy Generalization
von: Hu, Shengchao, et al.
Veröffentlicht: (2024)
von: Hu, Shengchao, et al.
Veröffentlicht: (2024)
Deep Reinforcement Learning for Solving Management Problems: Towards A Large Management Mode
von: Jiang, Jinyang, et al.
Veröffentlicht: (2024)
von: Jiang, Jinyang, et al.
Veröffentlicht: (2024)
On Vanishing Variance in Transformer Length Generalization
von: Li, Ruining, et al.
Veröffentlicht: (2025)
von: Li, Ruining, et al.
Veröffentlicht: (2025)
Beyond the Lower Bound: Bridging Regret Minimization and Best Arm Identification in Lexicographic Bandits
von: Xue, Bo, et al.
Veröffentlicht: (2025)
von: Xue, Bo, et al.
Veröffentlicht: (2025)
Adaptive Robust Estimator for Multi-Agent Reinforcement Learning
von: Li, Zhongyi, et al.
Veröffentlicht: (2026)
von: Li, Zhongyi, et al.
Veröffentlicht: (2026)
JTreeformer: Graph-Transformer via Latent-Diffusion Model for Molecular Generation
von: Shi, Ji, et al.
Veröffentlicht: (2025)
von: Shi, Ji, et al.
Veröffentlicht: (2025)
COGNOS: Universal Enhancement for Time Series Anomaly Detection via Constrained Gaussian-Noise Optimization and Smoothing
von: Shang, Wenlong, et al.
Veröffentlicht: (2025)
von: Shang, Wenlong, et al.
Veröffentlicht: (2025)
Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era
von: Li, Dawei, et al.
Veröffentlicht: (2025)
von: Li, Dawei, et al.
Veröffentlicht: (2025)
Embedding Reliability Verification Constraints into Generation Expansion Planning
von: Liu, Peng, et al.
Veröffentlicht: (2025)
von: Liu, Peng, et al.
Veröffentlicht: (2025)
Towards Sharper Risk Bounds for Minimax Problems
von: Zhu, Bowei, et al.
Veröffentlicht: (2024)
von: Zhu, Bowei, et al.
Veröffentlicht: (2024)
Rethinking Reinforcement fine-tuning of LLMs: A Multi-armed Bandit Learning Perspective
von: Hu, Xiao, et al.
Veröffentlicht: (2026)
von: Hu, Xiao, et al.
Veröffentlicht: (2026)
Humanoid-inspired Causal Representation Learning for Domain Generalization
von: Tao, Ze, et al.
Veröffentlicht: (2025)
von: Tao, Ze, et al.
Veröffentlicht: (2025)
RDesign: Hierarchical Data-efficient Representation Learning for Tertiary Structure-based RNA Design
von: Tan, Cheng, et al.
Veröffentlicht: (2023)
von: Tan, Cheng, et al.
Veröffentlicht: (2023)
Learning Non-Vacuous Generalization Bounds from Optimization
von: Tan, Chengli, et al.
Veröffentlicht: (2022)
von: Tan, Chengli, et al.
Veröffentlicht: (2022)
Why Transformers Need Adam: A Hessian Perspective
von: Zhang, Yushun, et al.
Veröffentlicht: (2024)
von: Zhang, Yushun, et al.
Veröffentlicht: (2024)
Towards Theoretical Understandings of Self-Consuming Generative Models
von: Fu, Shi, et al.
Veröffentlicht: (2024)
von: Fu, Shi, et al.
Veröffentlicht: (2024)
CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models
von: Zhu, Xiao, et al.
Veröffentlicht: (2026)
von: Zhu, Xiao, et al.
Veröffentlicht: (2026)
TrafficGPT: Towards Multi-Scale Traffic Analysis and Generation with Spatial-Temporal Agent Framework
von: Ouyang, Jinhui, et al.
Veröffentlicht: (2024)
von: Ouyang, Jinhui, et al.
Veröffentlicht: (2024)
LAC: Graph Contrastive Learning with Learnable Augmentation in Continuous Space
von: Lin, Zhenyu, et al.
Veröffentlicht: (2024)
von: Lin, Zhenyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Sharper Picture of Generalization in Transformers
von: Lintilhac, Paul, et al.
Veröffentlicht: (2026) -
Sharper Error Bounds in Late Fusion Multi-view Clustering Using Eigenvalue Proportion
von: Du, Liang, et al.
Veröffentlicht: (2024) -
Tail-Aware Information-Theoretic Generalization for RLHF and SGLD
von: Zhang, Huiming, et al.
Veröffentlicht: (2026) -
Machine Learning-Assisted High-Dimensional Matrix Estimation
von: Tian, Wan, et al.
Veröffentlicht: (2026) -
From Small to Large: Generalization Bounds for Transformers on Variable-Size Inputs
von: Alokhina, Anastasiia, et al.
Veröffentlicht: (2025)