Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Bu, Dake, Huang, Wei, Han, Andi, Nitanda, Atsushi, Suzuki, Taiji, Zhang, Qingfu, Wong, Hau-San |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Provable In-Context Vector Arithmetic via Retrieving Task Concepts
di: Bu, Dake, et al.
Pubblicazione: (2025)
di: Bu, Dake, et al.
Pubblicazione: (2025)
Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training
di: Bu, Dake, et al.
Pubblicazione: (2025)
di: Bu, Dake, et al.
Pubblicazione: (2025)
DPRM: A Plug-in Doob h transform-induced Token-Ordering Module for Diffusion Language Models
di: Bu, Dake, et al.
Pubblicazione: (2026)
di: Bu, Dake, et al.
Pubblicazione: (2026)
Post-Training as Reweighting: A Stochastic View of Reasoning Trajectories in Language Models
di: Bu, Dake, et al.
Pubblicazione: (2025)
di: Bu, Dake, et al.
Pubblicazione: (2025)
Provably Neural Active Learning Succeeds via Prioritizing Perplexing Samples
di: Bu, Dake, et al.
Pubblicazione: (2024)
di: Bu, Dake, et al.
Pubblicazione: (2024)
Direct Distributional Optimization for Provable Alignment of Diffusion Models
di: Kawata, Ryotaro, et al.
Pubblicazione: (2025)
di: Kawata, Ryotaro, et al.
Pubblicazione: (2025)
Slowly Annealed Langevin Dynamics: Theory and Applications to Training-Free Guided Generation
di: Nitanda, Atsushi, et al.
Pubblicazione: (2026)
di: Nitanda, Atsushi, et al.
Pubblicazione: (2026)
Transformers Provably Solve Parity Efficiently with Chain of Thought
di: Kim, Juno, et al.
Pubblicazione: (2024)
di: Kim, Juno, et al.
Pubblicazione: (2024)
Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization
di: Chen, Zonghao, et al.
Pubblicazione: (2025)
di: Chen, Zonghao, et al.
Pubblicazione: (2025)
In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
di: Wakayama, Tomoya, et al.
Pubblicazione: (2025)
di: Wakayama, Tomoya, et al.
Pubblicazione: (2025)
Intrinsic Wasserstein Rates for Score-Based Generative Models on Smooth Manifolds
di: Fu, Guoji, et al.
Pubblicazione: (2026)
di: Fu, Guoji, et al.
Pubblicazione: (2026)
Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine
di: Huang, Wei, et al.
Pubblicazione: (2026)
di: Huang, Wei, et al.
Pubblicazione: (2026)
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
di: Nishikawa, Naoki, et al.
Pubblicazione: (2024)
di: Nishikawa, Naoki, et al.
Pubblicazione: (2024)
Koopman-based generalization bound: New aspect for full-rank weights
di: Hashimoto, Yuka, et al.
Pubblicazione: (2023)
di: Hashimoto, Yuka, et al.
Pubblicazione: (2023)
On the Comparison between Multi-modal and Single-modal Contrastive Learning
di: Huang, Wei, et al.
Pubblicazione: (2024)
di: Huang, Wei, et al.
Pubblicazione: (2024)
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
di: Kim, Juno, et al.
Pubblicazione: (2024)
di: Kim, Juno, et al.
Pubblicazione: (2024)
Improved Particle Approximation Error for Mean Field Neural Networks
di: Nitanda, Atsushi
Pubblicazione: (2024)
di: Nitanda, Atsushi
Pubblicazione: (2024)
Propagation of Chaos for Mean-Field Langevin Dynamics and its Application to Model Ensemble
di: Nitanda, Atsushi, et al.
Pubblicazione: (2025)
di: Nitanda, Atsushi, et al.
Pubblicazione: (2025)
Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
di: Oh, Junsoo, et al.
Pubblicazione: (2025)
di: Oh, Junsoo, et al.
Pubblicazione: (2025)
On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent
di: Li, Bingrui, et al.
Pubblicazione: (2024)
di: Li, Bingrui, et al.
Pubblicazione: (2024)
How Does Preconditioning Guide Feature Learning in Deep Neural Networks?
di: Yoshida, Kotaro, et al.
Pubblicazione: (2025)
di: Yoshida, Kotaro, et al.
Pubblicazione: (2025)
Transformers are Minimax Optimal Nonparametric In-Context Learners
di: Kim, Juno, et al.
Pubblicazione: (2024)
di: Kim, Juno, et al.
Pubblicazione: (2024)
Generalization Bound of Gradient Flow through Training Trajectory and Data-dependent Kernel
di: Chen, Yilan, et al.
Pubblicazione: (2025)
di: Chen, Yilan, et al.
Pubblicazione: (2025)
Uniform convergence of the smooth calibration error and its relationship with functional gradient
di: Futami, Futoshi, et al.
Pubblicazione: (2025)
di: Futami, Futoshi, et al.
Pubblicazione: (2025)
Alternating Diffusion for Proximal Sampling with Zeroth Order Queries
di: Takagi, Hirohane, et al.
Pubblicazione: (2026)
di: Takagi, Hirohane, et al.
Pubblicazione: (2026)
On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD
di: Zhang, Tongcheng, et al.
Pubblicazione: (2026)
di: Zhang, Tongcheng, et al.
Pubblicazione: (2026)
Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning
di: Qiu, Shuang, et al.
Pubblicazione: (2024)
di: Qiu, Shuang, et al.
Pubblicazione: (2024)
How do Transformers perform In-Context Autoregressive Learning?
di: Sander, Michael E., et al.
Pubblicazione: (2024)
di: Sander, Michael E., et al.
Pubblicazione: (2024)
Mixture of Experts Provably Detect and Learn the Latent Cluster Structure in Gradient-Based Learning
di: Kawata, Ryotaro, et al.
Pubblicazione: (2025)
di: Kawata, Ryotaro, et al.
Pubblicazione: (2025)
On the Role of Label Noise in the Feature Learning Process
di: Han, Andi, et al.
Pubblicazione: (2025)
di: Han, Andi, et al.
Pubblicazione: (2025)
How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?
di: Huang, Wei, et al.
Pubblicazione: (2025)
di: Huang, Wei, et al.
Pubblicazione: (2025)
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
di: Kawata, Ryotaro, et al.
Pubblicazione: (2026)
di: Kawata, Ryotaro, et al.
Pubblicazione: (2026)
Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input
di: Takakura, Shokichi, et al.
Pubblicazione: (2023)
di: Takakura, Shokichi, et al.
Pubblicazione: (2023)
Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation
di: Kim, Juno, et al.
Pubblicazione: (2025)
di: Kim, Juno, et al.
Pubblicazione: (2025)
Mirror Descent Policy Optimisation for Robust Constrained Markov Decision Processes
di: Bossens, David M., et al.
Pubblicazione: (2025)
di: Bossens, David M., et al.
Pubblicazione: (2025)
Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression
di: Jiang, Jiarui, et al.
Pubblicazione: (2025)
di: Jiang, Jiarui, et al.
Pubblicazione: (2025)
Statistical Analysis of the Sinkhorn Iterations for Two-Sample Schrödinger Bridge Estimation
di: Maeda, Ibuki, et al.
Pubblicazione: (2025)
di: Maeda, Ibuki, et al.
Pubblicazione: (2025)
Provable In-Context Learning of Nonlinear Regression with Transformers
di: Li, Hongbo, et al.
Pubblicazione: (2025)
di: Li, Hongbo, et al.
Pubblicazione: (2025)
Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization
di: Jiang, Jiarui, et al.
Pubblicazione: (2024)
di: Jiang, Jiarui, et al.
Pubblicazione: (2024)
Deep Two-Way Matrix Reordering for Relational Data Analysis
di: Watanabe, Chihiro, et al.
Pubblicazione: (2021)
di: Watanabe, Chihiro, et al.
Pubblicazione: (2021)
Documenti analoghi
-
Provable In-Context Vector Arithmetic via Retrieving Task Concepts
di: Bu, Dake, et al.
Pubblicazione: (2025) -
Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training
di: Bu, Dake, et al.
Pubblicazione: (2025) -
DPRM: A Plug-in Doob h transform-induced Token-Ordering Module for Diffusion Language Models
di: Bu, Dake, et al.
Pubblicazione: (2026) -
Post-Training as Reweighting: A Stochastic View of Reasoning Trajectories in Language Models
di: Bu, Dake, et al.
Pubblicazione: (2025) -
Provably Neural Active Learning Succeeds via Prioritizing Perplexing Samples
di: Bu, Dake, et al.
Pubblicazione: (2024)