Nested Learning: The Illusion of Deep Learning Architectures
Fuente:
arXiv
Saved in:
| Main Authors: | Behrouz, Ali, Razaviyayn, Meisam, Zhong, Peilin, Mirrokni, Vahab |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
by: Behrouz, Ali, et al.
Published: (2025)
by: Behrouz, Ali, et al.
Published: (2025)
Memory Caching: RNNs with Growing Memory
by: Behrouz, Ali, et al.
Published: (2026)
by: Behrouz, Ali, et al.
Published: (2026)
Titans: Learning to Memorize at Test Time
by: Behrouz, Ali, et al.
Published: (2024)
by: Behrouz, Ali, et al.
Published: (2024)
TNT: Improving Chunkwise Training for Test-Time Memorization
by: Li, Zeman, et al.
Published: (2025)
by: Li, Zeman, et al.
Published: (2025)
Sampling and Loss Weights in Multi-Domain Training
by: Salmani, Mahdi, et al.
Published: (2025)
by: Salmani, Mahdi, et al.
Published: (2025)
PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts
by: Li, Zeman, et al.
Published: (2025)
by: Li, Zeman, et al.
Published: (2025)
ATLAS: Learning to Optimally Memorize the Context at Test Time
by: Behrouz, Ali, et al.
Published: (2025)
by: Behrouz, Ali, et al.
Published: (2025)
Addax: Utilizing Zeroth-Order Gradients to Improve Memory Efficiency and Performance of SGD for Fine-Tuning Language Models
by: Li, Zeman, et al.
Published: (2024)
by: Li, Zeman, et al.
Published: (2024)
Lattice: Learning to Efficiently Compress the Memory
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
Less is More: Convergence Benefits of Fewer Data Weight Updates over Longer Horizon
by: Das, Rudrajit, et al.
Published: (2026)
by: Das, Rudrajit, et al.
Published: (2026)
MS-SSM: A Multi-Scale State Space Model for Efficient Sequence Modeling
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
Optimistic Rates for Learning from Label Proportions
by: Li, Gene, et al.
Published: (2024)
by: Li, Gene, et al.
Published: (2024)
Understanding the Role of Training Data in Test-Time Scaling
by: Javanmard, Adel, et al.
Published: (2025)
by: Javanmard, Adel, et al.
Published: (2025)
Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
by: Javanmard, Adel, et al.
Published: (2026)
by: Javanmard, Adel, et al.
Published: (2026)
Trellis: Learning to Compress Key-Value Memory in Attention Models
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels
by: Kacham, Praneeth, et al.
Published: (2023)
by: Kacham, Praneeth, et al.
Published: (2023)
Learning from Aggregate responses: Instance Level versus Bag Level Loss Functions
by: Javanmard, Adel, et al.
Published: (2024)
by: Javanmard, Adel, et al.
Published: (2024)
Synthetic Text Generation for Training Large Language Models via Gradient Matching
by: Nguyen, Dang, et al.
Published: (2025)
by: Nguyen, Dang, et al.
Published: (2025)
DiSK: Differentially Private Optimizer with Simplified Kalman Filter for Noise Reduction
by: Zhang, Xinwei, et al.
Published: (2024)
by: Zhang, Xinwei, et al.
Published: (2024)
Differentially Private Synthetic Data Release for Topics API Outputs
by: Dick, Travis, et al.
Published: (2025)
by: Dick, Travis, et al.
Published: (2025)
ECO: Quantized Training without Full-Precision Master Weights
by: Nikdan, Mahdi, et al.
Published: (2026)
by: Nikdan, Mahdi, et al.
Published: (2026)
PolarQuant: Quantizing KV Caches with Polar Transformation
by: Han, Insu, et al.
Published: (2025)
by: Han, Insu, et al.
Published: (2025)
Early Stopping for Large Reasoning Models via Confidence Dynamics
by: Hosseini, Parsa, et al.
Published: (2026)
by: Hosseini, Parsa, et al.
Published: (2026)
SubGen: Token Generation in Sublinear Time and Memory
by: Zandieh, Amir, et al.
Published: (2024)
by: Zandieh, Amir, et al.
Published: (2024)
TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate
by: Zandieh, Amir, et al.
Published: (2025)
by: Zandieh, Amir, et al.
Published: (2025)
Federated Nested Learning: Collaborative Training of Self-Referential Memories for Test-Time Adaptation
by: Chen, Hong, et al.
Published: (2026)
by: Chen, Hong, et al.
Published: (2026)
Differentially Private Graph Learning via Sensitivity-Bounded Personalized PageRank
by: Epasto, Alessandro, et al.
Published: (2022)
by: Epasto, Alessandro, et al.
Published: (2022)
Private Federated Learning Without a Trusted Server: Optimal Algorithms for Convex Losses
by: Lowy, Andrew, et al.
Published: (2021)
by: Lowy, Andrew, et al.
Published: (2021)
Understanding Transformer Reasoning Capabilities via Graph Algorithms
by: Sanford, Clayton, et al.
Published: (2024)
by: Sanford, Clayton, et al.
Published: (2024)
Chimera: Effectively Modeling Multivariate Time Series with 2-Dimensional State Space Models
by: Behrouz, Ali, et al.
Published: (2024)
by: Behrouz, Ali, et al.
Published: (2024)
Cross-Modal Bayesian Low-Rank Adaptation for Uncertainty-Aware Multimodal Learning
by: Naderi, Habibeh, et al.
Published: (2026)
by: Naderi, Habibeh, et al.
Published: (2026)
Output Perturbation for Differentially Private Convex Optimization: Faster and More General
by: Lowy, Andrew, et al.
Published: (2021)
by: Lowy, Andrew, et al.
Published: (2021)
Best of Both Worlds: Advantages of Hybrid Graph Sequence Models
by: Behrouz, Ali, et al.
Published: (2024)
by: Behrouz, Ali, et al.
Published: (2024)
Bridging Reasoning to Learning: Unmasking Illusions using Complexity Out of Distribution Generalization
by: Paqaleh, Mohammad Mahdi Samiei, et al.
Published: (2025)
by: Paqaleh, Mohammad Mahdi Samiei, et al.
Published: (2025)
Product Interaction: An Algebraic Formalism for Deep Learning Architectures
by: Dong, Haonan, et al.
Published: (2026)
by: Dong, Haonan, et al.
Published: (2026)
The Importance of Architecture Choice in Deep Learning for Climate Applications
by: Dräger, Simon, et al.
Published: (2024)
by: Dräger, Simon, et al.
Published: (2024)
DREaM: Drug-Drug Relation Extraction via Transfer Learning Method
by: Fata, Ali, et al.
Published: (2025)
by: Fata, Ali, et al.
Published: (2025)
Injecting Imbalance Sensitivity for Multi-Task Learning
by: Zhou, Zhipeng, et al.
Published: (2025)
by: Zhou, Zhipeng, et al.
Published: (2025)
Private Stochastic Optimization With Large Worst-Case Lipschitz Parameter
by: Lowy, Andrew, et al.
Published: (2022)
by: Lowy, Andrew, et al.
Published: (2022)
Are Flat Minima an Illusion?
by: Bennett, Michael Timothy
Published: (2026)
by: Bennett, Michael Timothy
Published: (2026)
Similar Items
-
It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
by: Behrouz, Ali, et al.
Published: (2025) -
Memory Caching: RNNs with Growing Memory
by: Behrouz, Ali, et al.
Published: (2026) -
Titans: Learning to Memorize at Test Time
by: Behrouz, Ali, et al.
Published: (2024) -
TNT: Improving Chunkwise Training for Test-Time Memorization
by: Li, Zeman, et al.
Published: (2025) -
Sampling and Loss Weights in Multi-Domain Training
by: Salmani, Mahdi, et al.
Published: (2025)