Saved in:
| Main Authors: | Strobl, Lena, Angluin, Dana, Frank, Robert |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2503.22076 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What Formal Languages Can Transformers Express? A Survey
by: Strobl, Lena, et al.
Published: (2023)
by: Strobl, Lena, et al.
Published: (2023)
Transformers as Transducers
by: Strobl, Lena, et al.
Published: (2024)
by: Strobl, Lena, et al.
Published: (2024)
Simulating Hard Attention Using Soft Attention
by: Yang, Andy, et al.
Published: (2024)
by: Yang, Andy, et al.
Published: (2024)
Constructing Concise Characteristic Samples for Acceptors of Omega Regular Languages
by: Angluin, Dana, et al.
Published: (2022)
by: Angluin, Dana, et al.
Published: (2022)
Masked Hard-Attention Transformers Recognize Exactly the Star-Free Languages
by: Yang, Andy, et al.
Published: (2023)
by: Yang, Andy, et al.
Published: (2023)
Transformers Can Do Bayesian Inference
by: Müller, Samuel, et al.
Published: (2021)
by: Müller, Samuel, et al.
Published: (2021)
Can GNNs Learn Link Heuristics? A Concise Review and Evaluation of Link Prediction Methods
by: Liang, Shuming, et al.
Published: (2024)
by: Liang, Shuming, et al.
Published: (2024)
Reasoning Models Sometimes Output Illegible Chains of Thought
by: Jose, Arun
Published: (2025)
by: Jose, Arun
Published: (2025)
Can Transformers Do Enumerative Geometry?
by: Hashemi, Baran, et al.
Published: (2024)
by: Hashemi, Baran, et al.
Published: (2024)
Why Less is More (Sometimes): A Theory of Data Curation
by: Dohmatob, Elvis, et al.
Published: (2025)
by: Dohmatob, Elvis, et al.
Published: (2025)
Transformers Can Do Arithmetic with the Right Embeddings
by: McLeish, Sean, et al.
Published: (2024)
by: McLeish, Sean, et al.
Published: (2024)
A Single-Layer Model Can Do Language Modeling
by: Wang, Zanmin
Published: (2026)
by: Wang, Zanmin
Published: (2026)
Learning Causally Predictable Outcomes from Psychiatric Longitudinal Data
by: Strobl, Eric V.
Published: (2025)
by: Strobl, Eric V.
Published: (2025)
Global Interpretability via Automated Preprocessing: A Framework Inspired by Psychiatric Questionnaires
by: Strobl, Eric V.
Published: (2026)
by: Strobl, Eric V.
Published: (2026)
(Sometimes) Less is More: Mitigating the Complexity of Rule-based Representation for Interpretable Classification
by: Bergamin, Luca, et al.
Published: (2025)
by: Bergamin, Luca, et al.
Published: (2025)
Provable Knowledge Acquisition and Extraction in One-Layer Transformers
by: Xu, Ruichen, et al.
Published: (2025)
by: Xu, Ruichen, et al.
Published: (2025)
What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov Chains
by: Ekbote, Chanakya, et al.
Published: (2025)
by: Ekbote, Chanakya, et al.
Published: (2025)
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
by: Kim, Jeonghye, et al.
Published: (2026)
by: Kim, Jeonghye, et al.
Published: (2026)
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
by: Li, Zihao, et al.
Published: (2024)
by: Li, Zihao, et al.
Published: (2024)
Probabilistic Graphical Models: A Concise Tutorial
by: Maasch, Jacqueline, et al.
Published: (2025)
by: Maasch, Jacqueline, et al.
Published: (2025)
Sometimes I am a Tree: Data Drives Unstable Hierarchical Generalization
by: Qin, Tian, et al.
Published: (2024)
by: Qin, Tian, et al.
Published: (2024)
Alternative Loss Function in Evaluation of Transformer Models
by: Michańków, Jakub, et al.
Published: (2025)
by: Michańków, Jakub, et al.
Published: (2025)
Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws
by: Kamigaito, Hidetaka, et al.
Published: (2025)
by: Kamigaito, Hidetaka, et al.
Published: (2025)
A Concise Review of Hallucinations in LLMs and their Mitigation
by: Pulkundwar, Parth, et al.
Published: (2025)
by: Pulkundwar, Parth, et al.
Published: (2025)
LayerMatch: Do Pseudo-labels Benefit All Layers?
by: Liang, Chaoqi, et al.
Published: (2024)
by: Liang, Chaoqi, et al.
Published: (2024)
Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?
by: Öncel, Fırat, et al.
Published: (2024)
by: Öncel, Fırat, et al.
Published: (2024)
Multi-Layer Transformers Gradient Can be Approximated in Almost Linear Time
by: Liang, Yingyu, et al.
Published: (2024)
by: Liang, Yingyu, et al.
Published: (2024)
ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models
by: Dumitru, Razvan-Gabriel, et al.
Published: (2025)
by: Dumitru, Razvan-Gabriel, et al.
Published: (2025)
Probabilistic Robustness in Deep Learning: A Concise yet Comprehensive Guide
by: Zhao, Xingyu
Published: (2025)
by: Zhao, Xingyu
Published: (2025)
mlr3summary: Concise and interpretable summaries for machine learning models
by: Dandl, Susanne, et al.
Published: (2024)
by: Dandl, Susanne, et al.
Published: (2024)
Model Checking for Reinforcement Learning in Autonomous Driving: One Can Do More Than You Think!
by: Gu, Rong
Published: (2024)
by: Gu, Rong
Published: (2024)
HomeAdam: Adam and AdamW Algorithms Sometimes Go Home to Obtain Better Provable Generalization
by: Huang, Feihu, et al.
Published: (2026)
by: Huang, Feihu, et al.
Published: (2026)
SIMformer: Single-Layer Vanilla Transformer Can Learn Free-Space Trajectory Similarity
by: Yang, Chuang, et al.
Published: (2024)
by: Yang, Chuang, et al.
Published: (2024)
Extracting Root-Causal Brain Activity Driving Psychopathology from Resting State fMRI
by: Strobl, Eric V.
Published: (2026)
by: Strobl, Eric V.
Published: (2026)
FairPFN: Transformers Can do Counterfactual Fairness
by: Robertson, Jake, et al.
Published: (2024)
by: Robertson, Jake, et al.
Published: (2024)
Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors
by: Didolkar, Aniket, et al.
Published: (2025)
by: Didolkar, Aniket, et al.
Published: (2025)
A Concise Mathematical Description of Active Inference in Discrete Time
by: van Oostrum, Jesse, et al.
Published: (2024)
by: van Oostrum, Jesse, et al.
Published: (2024)
Tabular Foundation Models Can Do Survival Analysis
by: Kim, Da In, et al.
Published: (2026)
by: Kim, Da In, et al.
Published: (2026)
MAUNet-Light: A Concise MAUNet Architecture for Bias Correction and Downscaling of Precipitation Estimates
by: Sharma, Sumanta Chandra Mishra, et al.
Published: (2026)
by: Sharma, Sumanta Chandra Mishra, et al.
Published: (2026)
Just One Layer Norm Guarantees Stable Extrapolation
by: Ziomek, Juliusz, et al.
Published: (2025)
by: Ziomek, Juliusz, et al.
Published: (2025)
Similar Items
-
What Formal Languages Can Transformers Express? A Survey
by: Strobl, Lena, et al.
Published: (2023) -
Transformers as Transducers
by: Strobl, Lena, et al.
Published: (2024) -
Simulating Hard Attention Using Soft Attention
by: Yang, Andy, et al.
Published: (2024) -
Constructing Concise Characteristic Samples for Acceptors of Omega Regular Languages
by: Angluin, Dana, et al.
Published: (2022) -
Masked Hard-Attention Transformers Recognize Exactly the Star-Free Languages
by: Yang, Andy, et al.
Published: (2023)