A mathematical theory of balancing relational generalization and memorization
Fuente:
arXiv
Saved in:
| Main Authors: | Cheng, Luke, Lippl, Samuel |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When does compositional structure yield compositional generalization? A kernel theory
by: Lippl, Samuel, et al.
Published: (2024)
by: Lippl, Samuel, et al.
Published: (2024)
Short window attention enables long-term memorization
by: Cabannes, Loïc, et al.
Published: (2025)
by: Cabannes, Loïc, et al.
Published: (2025)
Math Takes Two: A test for emergent mathematical reasoning in communication
by: Cooper, Michael, et al.
Published: (2026)
by: Cooper, Michael, et al.
Published: (2026)
Algorithmic Primitives and Compositional Geometry of Reasoning in Language Models
by: Lippl, Samuel, et al.
Published: (2025)
by: Lippl, Samuel, et al.
Published: (2025)
Deep sequence models tend to memorize geometrically; it is unclear why
by: Noroozizadeh, Shahriar, et al.
Published: (2025)
by: Noroozizadeh, Shahriar, et al.
Published: (2025)
A new membership inference attack that spots memorization in generative and predictive models: Loss-Based with Reference Model algorithm (LBRM)
by: Taleb, Faiz, et al.
Published: (2025)
by: Taleb, Faiz, et al.
Published: (2025)
Self-rewarding correction for mathematical reasoning
by: Xiong, Wei, et al.
Published: (2025)
by: Xiong, Wei, et al.
Published: (2025)
GoldenTransformer: A Modular Fault Injection Framework for Transformer Robustness Research
by: Howard, Luke
Published: (2025)
by: Howard, Luke
Published: (2025)
Putnam-like dataset summary: LLMs as mathematical competition contestants
by: Bieganowski, Bartosz, et al.
Published: (2025)
by: Bieganowski, Bartosz, et al.
Published: (2025)
Differential learning kinetics govern the transition from memorization to generalization during in-context learning
by: Nguyen, Alex, et al.
Published: (2024)
by: Nguyen, Alex, et al.
Published: (2024)
Infinite Width Models That Work: Why Feature Learning Doesn't Matter as Much as You Think
by: Sernau, Luke
Published: (2024)
by: Sernau, Luke
Published: (2024)
Ratio law: mathematical descriptions for a universal relationship between AI performance and input samples
by: Kang, Boming, et al.
Published: (2024)
by: Kang, Boming, et al.
Published: (2024)
Time Matters: Scaling Laws for Any Budget
by: Inbar, Itay, et al.
Published: (2024)
by: Inbar, Itay, et al.
Published: (2024)
Int2Int: a framework for mathematics with transformers
by: Charton, François
Published: (2025)
by: Charton, François
Published: (2025)
Linear Mode Connectivity in Sparse Neural Networks
by: McDermott, Luke, et al.
Published: (2023)
by: McDermott, Luke, et al.
Published: (2023)
Reinforcement Learning with $ω$-Regular Objectives and Constraints
by: Wagner, Dominik, et al.
Published: (2025)
by: Wagner, Dominik, et al.
Published: (2025)
SB-TRPO: Towards Safe Reinforcement Learning with Hard Constraints
by: Wagner, Dominik, et al.
Published: (2025)
by: Wagner, Dominik, et al.
Published: (2025)
An explainable transformer circuit for compositional generalization
by: Tang, Cheng, et al.
Published: (2025)
by: Tang, Cheng, et al.
Published: (2025)
Meta-Gradient Search Control: A Method for Improving the Efficiency of Dyna-style Planning
by: Burega, Bradley, et al.
Published: (2024)
by: Burega, Bradley, et al.
Published: (2024)
AMPED: Adaptive Multi-objective Projection for balancing Exploration and skill Diversification
by: Cho, Geonwoo, et al.
Published: (2025)
by: Cho, Geonwoo, et al.
Published: (2025)
Inductive biases of multi-task learning and finetuning: multiple regimes of feature reuse
by: Lippl, Samuel, et al.
Published: (2023)
by: Lippl, Samuel, et al.
Published: (2023)
All Random Features Representations are Equivalent
by: Sernau, Luke, et al.
Published: (2024)
by: Sernau, Luke, et al.
Published: (2024)
Tucker Attention: A generalization of approximate attention mechanisms
by: Klein, Timon, et al.
Published: (2026)
by: Klein, Timon, et al.
Published: (2026)
In-Context Black-Box Optimization with Unreliable Feedback
by: Blumer, Nicolas Samuel, et al.
Published: (2026)
by: Blumer, Nicolas Samuel, et al.
Published: (2026)
A Survey of State Representation Learning for Deep Reinforcement Learning
by: Echchahed, Ayoub, et al.
Published: (2025)
by: Echchahed, Ayoub, et al.
Published: (2025)
A multi-algorithm approach for operational human resources workload balancing in a last mile urban delivery system
by: Moreno-Saavedra, Luis M., et al.
Published: (2025)
by: Moreno-Saavedra, Luis M., et al.
Published: (2025)
On the Performance of Imputation Techniques for Missing Values on Healthcare Datasets
by: Joel, Luke Oluwaseye, et al.
Published: (2024)
by: Joel, Luke Oluwaseye, et al.
Published: (2024)
MAD-SmaAt-GNet: A Multimodal Advection-Guided Neural Network for Precipitation Nowcasting
by: van Wonderen, Samuel, et al.
Published: (2026)
by: van Wonderen, Samuel, et al.
Published: (2026)
The Formalism-Implementation Gap in Reinforcement Learning Research
by: Castro, Pablo Samuel
Published: (2025)
by: Castro, Pablo Samuel
Published: (2025)
Diffusion model for relational inference
by: Zheng, Shuhan, et al.
Published: (2024)
by: Zheng, Shuhan, et al.
Published: (2024)
JaxUED: A simple and useable UED library in Jax
by: Coward, Samuel, et al.
Published: (2024)
by: Coward, Samuel, et al.
Published: (2024)
Data-Aware Random Feature Kernel for Transformers
by: Farzam, Amirhossein, et al.
Published: (2026)
by: Farzam, Amirhossein, et al.
Published: (2026)
Feature maps for the Laplacian kernel and its generalizations
by: Ahir, Sudhendu, et al.
Published: (2025)
by: Ahir, Sudhendu, et al.
Published: (2025)
Towards a theory of out-of-distribution learning
by: Dey, Jayanta, et al.
Published: (2021)
by: Dey, Jayanta, et al.
Published: (2021)
Improving MoE Compute Efficiency by Composing Weight and Data Sparsity
by: Kilian, Maciej, et al.
Published: (2026)
by: Kilian, Maciej, et al.
Published: (2026)
Calibration in Deep Learning: A Survey of the State-of-the-Art
by: Wang, Cheng
Published: (2023)
by: Wang, Cheng
Published: (2023)
Outcome-based Reinforcement Learning to Predict the Future
by: Turtel, Benjamin, et al.
Published: (2025)
by: Turtel, Benjamin, et al.
Published: (2025)
Continuous Thought Machines
by: Darlow, Luke, et al.
Published: (2025)
by: Darlow, Luke, et al.
Published: (2025)
Towards Understanding the Role of Sharpness-Aware Minimization Algorithms for Out-of-Distribution Generalization
by: Schapiro, Samuel, et al.
Published: (2024)
by: Schapiro, Samuel, et al.
Published: (2024)
Causal Masking on Spatial Data: An Information-Theoretic Case for Learning Spatial Datasets with Unimodal Language Models
by: Junkin, Jared, et al.
Published: (2025)
by: Junkin, Jared, et al.
Published: (2025)
Similar Items
-
When does compositional structure yield compositional generalization? A kernel theory
by: Lippl, Samuel, et al.
Published: (2024) -
Short window attention enables long-term memorization
by: Cabannes, Loïc, et al.
Published: (2025) -
Math Takes Two: A test for emergent mathematical reasoning in communication
by: Cooper, Michael, et al.
Published: (2026) -
Algorithmic Primitives and Compositional Geometry of Reasoning in Language Models
by: Lippl, Samuel, et al.
Published: (2025) -
Deep sequence models tend to memorize geometrically; it is unclear why
by: Noroozizadeh, Shahriar, et al.
Published: (2025)