Learning to Think from Multiple Thinkers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Joshi, Nirmit, Magen, Roey, Srebro, Nathan, Tsilivis, Nikolaos, Vardi, Gal |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Theory of Learning with Autoregressive Chain of Thought
von: Joshi, Nirmit, et al.
Veröffentlicht: (2025)
von: Joshi, Nirmit, et al.
Veröffentlicht: (2025)
On the Hardness of Learning Regular Expressions
von: Attias, Idan, et al.
Veröffentlicht: (2025)
von: Attias, Idan, et al.
Veröffentlicht: (2025)
Noisy Interpolation Learning with Shallow Univariate ReLU Networks
von: Joshi, Nirmit, et al.
Veröffentlicht: (2023)
von: Joshi, Nirmit, et al.
Veröffentlicht: (2023)
Transformers are almost optimal metalearners for linear classification
von: Magen, Roey, et al.
Veröffentlicht: (2025)
von: Magen, Roey, et al.
Veröffentlicht: (2025)
Learning to Answer from Correct Demonstrations
von: Joshi, Nirmit, et al.
Veröffentlicht: (2025)
von: Joshi, Nirmit, et al.
Veröffentlicht: (2025)
On the Complexity of Learning Sparse Functions with Statistical and Gradient Queries
von: Joshi, Nirmit, et al.
Veröffentlicht: (2024)
von: Joshi, Nirmit, et al.
Veröffentlicht: (2024)
Provable Failure of Language Models in Learning Majority Boolean Logic via Gradient Descent
von: Chen, Bo, et al.
Veröffentlicht: (2025)
von: Chen, Bo, et al.
Veröffentlicht: (2025)
Mathematical Algorithm Design for Deep Learning under Societal and Judicial Constraints: The Algorithmic Transparency Requirement
von: Boche, Holger, et al.
Veröffentlicht: (2024)
von: Boche, Holger, et al.
Veröffentlicht: (2024)
Have Large Language Models Learned to Reason? A Characterization via 3-SAT Phase Transition
von: Hazra, Rishi, et al.
Veröffentlicht: (2025)
von: Hazra, Rishi, et al.
Veröffentlicht: (2025)
Learning Tree Pattern Transformations
von: Neider, Daniel, et al.
Veröffentlicht: (2024)
von: Neider, Daniel, et al.
Veröffentlicht: (2024)
Overfitting Behaviour of Gaussian Kernel Ridgeless Regression: Varying Bandwidth or Dimensionality
von: Medvedev, Marko, et al.
Veröffentlicht: (2024)
von: Medvedev, Marko, et al.
Veröffentlicht: (2024)
Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural Networks
von: Tsilivis, Nikolaos, et al.
Veröffentlicht: (2024)
von: Tsilivis, Nikolaos, et al.
Veröffentlicht: (2024)
A Provable Expressiveness Hierarchy in Hybrid Linear-Full Attention
von: Ye, Xiaowei, et al.
Veröffentlicht: (2026)
von: Ye, Xiaowei, et al.
Veröffentlicht: (2026)
On the Expressive Power and Limitations of Multi-Layer SSMs
von: Zubić, Nikola, et al.
Veröffentlicht: (2026)
von: Zubić, Nikola, et al.
Veröffentlicht: (2026)
How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers
von: Wang, Xiao
Veröffentlicht: (2026)
von: Wang, Xiao
Veröffentlicht: (2026)
A Quantitative Definition of Intelligence
von: Choi, Kang-Sin
Veröffentlicht: (2026)
von: Choi, Kang-Sin
Veröffentlicht: (2026)
Computational Limits of Low-Rank Adaptation (LoRA) Fine-Tuning for Transformer Models
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
Lossless Model Compression via Joint Low-Rank Factorization Optimization
von: Zhang, Boyang, et al.
Veröffentlicht: (2024)
von: Zhang, Boyang, et al.
Veröffentlicht: (2024)
Provably Overwhelming Transformer Models with Designed Inputs
von: Stambler, Lev, et al.
Veröffentlicht: (2025)
von: Stambler, Lev, et al.
Veröffentlicht: (2025)
Limitations on Accurate, Trusted, Human-level Reasoning
von: Panigrahy, Rina, et al.
Veröffentlicht: (2025)
von: Panigrahy, Rina, et al.
Veröffentlicht: (2025)
Looped ReLU MLPs May Be All You Need as Practical Programmable Computers
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
Mathematical Formalism for Memory Compression in Selective State Space Models
von: Bhat, Siddhanth
Veröffentlicht: (2024)
von: Bhat, Siddhanth
Veröffentlicht: (2024)
When Can We Solve the Weighted Low Rank Approximation Problem in Truly Subquadratic Time?
von: Li, Chenyang, et al.
Veröffentlicht: (2025)
von: Li, Chenyang, et al.
Veröffentlicht: (2025)
A Unified Approach for Maximizing Continuous DR-submodular Functions
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2023)
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2023)
On the Computational Capability of Graph Neural Networks: A Circuit Complexity Bound Perspective
von: Li, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Li, Xiaoyu, et al.
Veröffentlicht: (2025)
The Price of Implicit Bias in Adversarially Robust Generalization
von: Tsilivis, Nikolaos, et al.
Veröffentlicht: (2024)
von: Tsilivis, Nikolaos, et al.
Veröffentlicht: (2024)
Learning single-index models via harmonic decomposition
von: Joshi, Nirmit, et al.
Veröffentlicht: (2025)
von: Joshi, Nirmit, et al.
Veröffentlicht: (2025)
Near-Optimal Learning and Planning in Separated Latent MDPs
von: Chen, Fan, et al.
Veröffentlicht: (2024)
von: Chen, Fan, et al.
Veröffentlicht: (2024)
Nearest Neighbor CCP-Based Molecular Sequence Analysis
von: Ali, Sarwan, et al.
Veröffentlicht: (2024)
von: Ali, Sarwan, et al.
Veröffentlicht: (2024)
Reinforced Generation of Combinatorial Structures: Hardness of Approximation
von: Nagda, Ansh, et al.
Veröffentlicht: (2025)
von: Nagda, Ansh, et al.
Veröffentlicht: (2025)
A Measure-Theoretic Analysis of Reasoning: Structural Generalization and Approximation Limits
von: Zhang, Yuyang, et al.
Veröffentlicht: (2026)
von: Zhang, Yuyang, et al.
Veröffentlicht: (2026)
Demystifying the unreasonable effectiveness of online alignment methods
von: Kang, Enoch Hyunwook
Veröffentlicht: (2026)
von: Kang, Enoch Hyunwook
Veröffentlicht: (2026)
RoPE Attention Can Be Trained in Almost Linear Time
von: Cao, Yang, et al.
Veröffentlicht: (2024)
von: Cao, Yang, et al.
Veröffentlicht: (2024)
The Alignment Trap: Complexity Barriers
von: Yao, Jasper
Veröffentlicht: (2025)
von: Yao, Jasper
Veröffentlicht: (2025)
Modern Hopfield Networks Require Chain-of-Thought to Solve $\mathsf{NC}^1$-Hard Problems
von: Cao, Yang, et al.
Veröffentlicht: (2024)
von: Cao, Yang, et al.
Veröffentlicht: (2024)
The Computational Limits of State-Space Models and Mamba via the Lens of Circuit Complexity
von: Chen, Yifang, et al.
Veröffentlicht: (2024)
von: Chen, Yifang, et al.
Veröffentlicht: (2024)
NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes
von: Fan, Lizhou, et al.
Veröffentlicht: (2023)
von: Fan, Lizhou, et al.
Veröffentlicht: (2023)
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
von: Li, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Li, Xiaoyu, et al.
Veröffentlicht: (2024)
A Complexity Map of Probabilistic Reasoning for Neurosymbolic Classification Techniques
von: Ledaguenel, Arthur, et al.
Veröffentlicht: (2024)
von: Ledaguenel, Arthur, et al.
Veröffentlicht: (2024)
Neural Algorithmic Reasoning for Hypergraphs with Looped Transformers
von: Huang, Zekai, et al.
Veröffentlicht: (2025)
von: Huang, Zekai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Theory of Learning with Autoregressive Chain of Thought
von: Joshi, Nirmit, et al.
Veröffentlicht: (2025) -
On the Hardness of Learning Regular Expressions
von: Attias, Idan, et al.
Veröffentlicht: (2025) -
Noisy Interpolation Learning with Shallow Univariate ReLU Networks
von: Joshi, Nirmit, et al.
Veröffentlicht: (2023) -
Transformers are almost optimal metalearners for linear classification
von: Magen, Roey, et al.
Veröffentlicht: (2025) -
Learning to Answer from Correct Demonstrations
von: Joshi, Nirmit, et al.
Veröffentlicht: (2025)