Inference Scaling vs Reasoning: An Empirical Analysis of Compute-Optimal LLM Problem-Solving
Fuente:
arXiv
Guardado en:
| Autores principales: | AbdElhameed, Marwan, Halim, Pavly |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Modern Hopfield Networks Require Chain-of-Thought to Solve $\mathsf{NC}^1$-Hard Problems
por: Cao, Yang, et al.
Publicado: (2024)
por: Cao, Yang, et al.
Publicado: (2024)
Position: Scaling LLM Agents Requires Asymptotic Analysis with LLM Primitives
por: Meyerson, Elliot, et al.
Publicado: (2025)
por: Meyerson, Elliot, et al.
Publicado: (2025)
A Measure-Theoretic Analysis of Reasoning: Structural Generalization and Approximation Limits
por: Zhang, Yuyang, et al.
Publicado: (2026)
por: Zhang, Yuyang, et al.
Publicado: (2026)
Chain of Thought Empowers Transformers to Solve Inherently Serial Problems
por: Li, Zhiyuan, et al.
Publicado: (2024)
por: Li, Zhiyuan, et al.
Publicado: (2024)
Neural Algorithmic Reasoning for Hypergraphs with Looped Transformers
por: Huang, Zekai, et al.
Publicado: (2025)
por: Huang, Zekai, et al.
Publicado: (2025)
When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning
por: Singhi, Nishad, et al.
Publicado: (2025)
por: Singhi, Nishad, et al.
Publicado: (2025)
Unlocking the Theory Behind Scaling 1-Bit Neural Networks
por: Daliri, Majid, et al.
Publicado: (2024)
por: Daliri, Majid, et al.
Publicado: (2024)
NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes
por: Fan, Lizhou, et al.
Publicado: (2023)
por: Fan, Lizhou, et al.
Publicado: (2023)
The Computational Limits of State-Space Models and Mamba via the Lens of Circuit Complexity
por: Chen, Yifang, et al.
Publicado: (2024)
por: Chen, Yifang, et al.
Publicado: (2024)
On Efficiently Representing Regular Languages as RNNs
por: Svete, Anej, et al.
Publicado: (2024)
por: Svete, Anej, et al.
Publicado: (2024)
Fundamental Limitations on Subquadratic Alternatives to Transformers
por: Alman, Josh, et al.
Publicado: (2024)
por: Alman, Josh, et al.
Publicado: (2024)
The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought
por: Brösamle, Moritz, et al.
Publicado: (2026)
por: Brösamle, Moritz, et al.
Publicado: (2026)
Perfect diffusion is $\mathsf{TC}^0$ -- Bad diffusion is Turing-complete
por: Liu, Yuxi
Publicado: (2025)
por: Liu, Yuxi
Publicado: (2025)
Compute Optimal Scaling of Skills: Knowledge vs Reasoning
por: Roberts, Nicholas, et al.
Publicado: (2025)
por: Roberts, Nicholas, et al.
Publicado: (2025)
The Fine-Grained Complexity of Gradient Computation for Training Large Language Models
por: Alman, Josh, et al.
Publicado: (2024)
por: Alman, Josh, et al.
Publicado: (2024)
Local vs. Global Interpretability: A Computational Complexity Perspective
por: Bassan, Shahaf, et al.
Publicado: (2024)
por: Bassan, Shahaf, et al.
Publicado: (2024)
When Can We Solve the Weighted Low Rank Approximation Problem in Truly Subquadratic Time?
por: Li, Chenyang, et al.
Publicado: (2025)
por: Li, Chenyang, et al.
Publicado: (2025)
RoPE Attention Can Be Trained in Almost Linear Time
por: Cao, Yang, et al.
Publicado: (2024)
por: Cao, Yang, et al.
Publicado: (2024)
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
por: Li, Xiaoyu, et al.
Publicado: (2024)
por: Li, Xiaoyu, et al.
Publicado: (2024)
Circuit Complexity Bounds for RoPE-based Transformer Architecture
por: Chen, Bo, et al.
Publicado: (2024)
por: Chen, Bo, et al.
Publicado: (2024)
On Fine-Grained I/O Complexity of Attention Backward Passes
por: Li, Xiaoyu, et al.
Publicado: (2024)
por: Li, Xiaoyu, et al.
Publicado: (2024)
Demystifying the unreasonable effectiveness of online alignment methods
por: Kang, Enoch Hyunwook
Publicado: (2026)
por: Kang, Enoch Hyunwook
Publicado: (2026)
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
por: Chen, Yifang, et al.
Publicado: (2025)
por: Chen, Yifang, et al.
Publicado: (2025)
Circuit Complexity Bounds for Visual Autoregressive Model
por: Ke, Yekun, et al.
Publicado: (2025)
por: Ke, Yekun, et al.
Publicado: (2025)
The Expressive Power of Transformers with Chain of Thought
por: Merrill, William, et al.
Publicado: (2023)
por: Merrill, William, et al.
Publicado: (2023)
Certifiable Boolean Reasoning Is Universal
por: Li, Wenhao, et al.
Publicado: (2026)
por: Li, Wenhao, et al.
Publicado: (2026)
The Illusion of State in State-Space Models
por: Merrill, William, et al.
Publicado: (2024)
por: Merrill, William, et al.
Publicado: (2024)
On the Computational Hardness of Transformers
por: Saha, Barna, et al.
Publicado: (2026)
por: Saha, Barna, et al.
Publicado: (2026)
Polynomial-Time Optimal Group Selection via the Double-Commutator Eigenvalue Problem
por: Thornton, Mitchell A.
Publicado: (2026)
por: Thornton, Mitchell A.
Publicado: (2026)
Context-Free Recognition with Transformers
por: Jerad, Selim, et al.
Publicado: (2026)
por: Jerad, Selim, et al.
Publicado: (2026)
Why Are Linear RNNs More Parallelizable?
por: Merrill, William, et al.
Publicado: (2026)
por: Merrill, William, et al.
Publicado: (2026)
Unique Hard Attention: A Tale of Two Sides
por: Jerad, Selim, et al.
Publicado: (2025)
por: Jerad, Selim, et al.
Publicado: (2025)
Diffusion Language Models are Provably Optimal Parallel Samplers
por: Jiang, Haozhe, et al.
Publicado: (2025)
por: Jiang, Haozhe, et al.
Publicado: (2025)
Lower Bounds for Chain-of-Thought Reasoning in Hard-Attention Transformers
por: Amiri, Alireza, et al.
Publicado: (2025)
por: Amiri, Alireza, et al.
Publicado: (2025)
Optimizing Computational-Statistical Runtime for Wasserstein Distance Estimation
por: Jacobs, Peter Matthew, et al.
Publicado: (2026)
por: Jacobs, Peter Matthew, et al.
Publicado: (2026)
Additive Models Explained: A Computational Complexity Approach
por: Bassan, Shahaf, et al.
Publicado: (2025)
por: Bassan, Shahaf, et al.
Publicado: (2025)
Statistical and Computational Guarantees of Kernel Max-Sliced Wasserstein Distances
por: Wang, Jie, et al.
Publicado: (2024)
por: Wang, Jie, et al.
Publicado: (2024)
Necessary and Sufficient Oracles: Toward a Computational Taxonomy For Reinforcement Learning
por: Rohatgi, Dhruv, et al.
Publicado: (2025)
por: Rohatgi, Dhruv, et al.
Publicado: (2025)
Deep Learning as a Convex Paradigm of Computation: Minimizing Circuit Size with ResNets
por: Jacot, Arthur
Publicado: (2025)
por: Jacot, Arthur
Publicado: (2025)
Smoothed Analysis for Learning Concepts with Low Intrinsic Dimension
por: Chandrasekaran, Gautam, et al.
Publicado: (2024)
por: Chandrasekaran, Gautam, et al.
Publicado: (2024)
Ejemplares similares
-
Modern Hopfield Networks Require Chain-of-Thought to Solve $\mathsf{NC}^1$-Hard Problems
por: Cao, Yang, et al.
Publicado: (2024) -
Position: Scaling LLM Agents Requires Asymptotic Analysis with LLM Primitives
por: Meyerson, Elliot, et al.
Publicado: (2025) -
A Measure-Theoretic Analysis of Reasoning: Structural Generalization and Approximation Limits
por: Zhang, Yuyang, et al.
Publicado: (2026) -
Chain of Thought Empowers Transformers to Solve Inherently Serial Problems
por: Li, Zhiyuan, et al.
Publicado: (2024) -
Neural Algorithmic Reasoning for Hypergraphs with Looped Transformers
por: Huang, Zekai, et al.
Publicado: (2025)