Saved in:
| Main Authors: | Curth, Alicia, Lawrence, Rachel, Karmalkar, Sushrut, Prasad, Niranjani |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.12426 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Parallel Sampling from Masked Diffusion Models via Conditional Independence Testing
by: Azangulov, Iskander, et al.
Published: (2025)
by: Azangulov, Iskander, et al.
Published: (2025)
Better Think Thrice: Learning to Reason Causally with Double Counterfactual Consistency
by: Lin, Victoria, et al.
Published: (2026)
by: Lin, Victoria, et al.
Published: (2026)
Efficient Knowledge Distillation via Curriculum Extraction
by: Gupta, Shivam, et al.
Published: (2025)
by: Gupta, Shivam, et al.
Published: (2025)
Depth-Width tradeoffs in Algorithmic Reasoning of Graph Tasks with Transformers
by: Yehudai, Gilad, et al.
Published: (2025)
by: Yehudai, Gilad, et al.
Published: (2025)
A Fourier Space Perspective on Diffusion Models
by: Falck, Fabian, et al.
Published: (2025)
by: Falck, Fabian, et al.
Published: (2025)
Classical Statistical (In-Sample) Intuitions Don't Generalize Well: A Note on Bias-Variance Tradeoffs, Overfitting and Moving from Fixed to Random Designs
by: Curth, Alicia
Published: (2024)
by: Curth, Alicia
Published: (2024)
Forward-Learned Discrete Diffusion: Learning how to noise to denoise faster
by: Bartosh, Grigory, et al.
Published: (2026)
by: Bartosh, Grigory, et al.
Published: (2026)
Sum-of-squares lower bounds for Non-Gaussian Component Analysis
by: Diakonikolas, Ilias, et al.
Published: (2024)
by: Diakonikolas, Ilias, et al.
Published: (2024)
Transformer See, Transformer Do: Copying as an Intermediate Step in Learning Analogical Reasoning
by: Hellwig, Philipp, et al.
Published: (2026)
by: Hellwig, Philipp, et al.
Published: (2026)
Robust Learning of a Group DRO Neuron
by: Cao, Guyang, et al.
Published: (2026)
by: Cao, Guyang, et al.
Published: (2026)
Mathematical Derivation Graphs: A Relation Extraction Task in STEM Manuscripts
by: Prasad, Vishesh, et al.
Published: (2024)
by: Prasad, Vishesh, et al.
Published: (2024)
Learning a Single Neuron Robustly to Distributional Shifts and Adversarial Label Noise
by: Li, Shuyao, et al.
Published: (2024)
by: Li, Shuyao, et al.
Published: (2024)
Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers
by: Kohli, Harsh, et al.
Published: (2026)
by: Kohli, Harsh, et al.
Published: (2026)
Universal Transformers Need Memory: Depth-State Trade-offs in Adaptive Recursive Reasoning
by: Sapunov, Grigory
Published: (2026)
by: Sapunov, Grigory
Published: (2026)
Do Depth-Grown Models Overcome the Curse of Depth? An In-Depth Analysis
by: Kapl, Ferdinand, et al.
Published: (2025)
by: Kapl, Ferdinand, et al.
Published: (2025)
Adaptive Prompting for Continual Relation Extraction: A Within-Task Variance Perspective
by: Le, Minh, et al.
Published: (2024)
by: Le, Minh, et al.
Published: (2024)
Flexi-LoRA with Input-Adaptive Ranks: Efficient Finetuning for Speech and Reasoning Tasks
by: Li, Zongqian, et al.
Published: (2026)
by: Li, Zongqian, et al.
Published: (2026)
What Evidence Do Language Models Find Convincing?
by: Wan, Alexander, et al.
Published: (2024)
by: Wan, Alexander, et al.
Published: (2024)
Cog-DRIFT: Exploration on Adaptively Reformulated Instances Enables Learning from Hard Reasoning Problems
by: Chen, Justin Chih-Yao, et al.
Published: (2026)
by: Chen, Justin Chih-Yao, et al.
Published: (2026)
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
by: Gambardella, Andrew, et al.
Published: (2024)
by: Gambardella, Andrew, et al.
Published: (2024)
Why do Random Forests Work? Understanding Tree Ensembles as Self-Regularizing Adaptive Smoothers
by: Curth, Alicia, et al.
Published: (2024)
by: Curth, Alicia, et al.
Published: (2024)
Pause Tokens Strictly Increase the Expressivity of Constant-Depth Transformers
by: London, Charles, et al.
Published: (2025)
by: London, Charles, et al.
Published: (2025)
Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks
by: Jang, Lawrence Keunho, et al.
Published: (2026)
by: Jang, Lawrence Keunho, et al.
Published: (2026)
Task Vector Geometry Underlies Dual Modes of Task Inference in Transformers
by: Yan, Hao, et al.
Published: (2026)
by: Yan, Hao, et al.
Published: (2026)
Reasoning in Transformers -- Mitigating Spurious Correlations and Reasoning Shortcuts
by: Enström, Daniel, et al.
Published: (2024)
by: Enström, Daniel, et al.
Published: (2024)
Adaptive Task Vectors for Large Language Models
by: Kang, Joonseong, et al.
Published: (2025)
by: Kang, Joonseong, et al.
Published: (2025)
What Do Language Models Learn in Context? The Structured Task Hypothesis
by: Li, Jiaoda, et al.
Published: (2024)
by: Li, Jiaoda, et al.
Published: (2024)
You Do Not Fully Utilize Transformer's Representation Capacity
by: Gerasimov, Gleb, et al.
Published: (2025)
by: Gerasimov, Gleb, et al.
Published: (2025)
Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation
by: Bae, Sangmin, et al.
Published: (2025)
by: Bae, Sangmin, et al.
Published: (2025)
The NLP Task Effectiveness of Long-Range Transformers
by: Qin, Guanghui, et al.
Published: (2022)
by: Qin, Guanghui, et al.
Published: (2022)
LASeR: Learning to Adaptively Select Reward Models with Multi-Armed Bandits
by: Nguyen, Duy, et al.
Published: (2024)
by: Nguyen, Duy, et al.
Published: (2024)
Benchmarking Large Language Models for Math Reasoning Tasks
by: Seßler, Kathrin, et al.
Published: (2024)
by: Seßler, Kathrin, et al.
Published: (2024)
DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging
by: Pagliardini, Matteo, et al.
Published: (2024)
by: Pagliardini, Matteo, et al.
Published: (2024)
Do Sentence Transformers Learn Quasi-Geospatial Concepts from General Text?
by: Ilyankou, Ilya, et al.
Published: (2024)
by: Ilyankou, Ilya, et al.
Published: (2024)
Transformers meet Neural Algorithmic Reasoners
by: Bounsi, Wilfried, et al.
Published: (2024)
by: Bounsi, Wilfried, et al.
Published: (2024)
Do Multilingual VLMs Reason Equally? A Cross-Lingual Visual Reasoning Audit for Indian Languages
by: R, Swastik
Published: (2026)
by: R, Swastik
Published: (2026)
Transformers on Markov Data: Constant Depth Suffices
by: Rajaraman, Nived, et al.
Published: (2024)
by: Rajaraman, Nived, et al.
Published: (2024)
Language Models Use Trigonometry to Do Addition
by: Kantamneni, Subhash, et al.
Published: (2025)
by: Kantamneni, Subhash, et al.
Published: (2025)
Direct-Scoring NLG Evaluators Can Use Pairwise Comparisons Too
by: Lawrence, Logan, et al.
Published: (2025)
by: Lawrence, Logan, et al.
Published: (2025)
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
by: Geiping, Jonas, et al.
Published: (2025)
by: Geiping, Jonas, et al.
Published: (2025)
Similar Items
-
Parallel Sampling from Masked Diffusion Models via Conditional Independence Testing
by: Azangulov, Iskander, et al.
Published: (2025) -
Better Think Thrice: Learning to Reason Causally with Double Counterfactual Consistency
by: Lin, Victoria, et al.
Published: (2026) -
Efficient Knowledge Distillation via Curriculum Extraction
by: Gupta, Shivam, et al.
Published: (2025) -
Depth-Width tradeoffs in Algorithmic Reasoning of Graph Tasks with Transformers
by: Yehudai, Gilad, et al.
Published: (2025) -
A Fourier Space Perspective on Diffusion Models
by: Falck, Fabian, et al.
Published: (2025)