Optimizers Qualitatively Alter Solutions And We Should Leverage This
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pascanu, Razvan, Lyle, Clare, Modoranu, Ionut-Vlad, Borras, Naima Elosegui, Alistarh, Dan, Velickovic, Petar, Chandar, Sarath, De, Soham, Martens, James |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026)
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026)
LDAdam: Adaptive Optimization from Low-Dimensional Gradient Statistics
von: Robert, Thomas, et al.
Veröffentlicht: (2024)
von: Robert, Thomas, et al.
Veröffentlicht: (2024)
The Iterative Optimal Brain Surgeon: Faster Sparse Recovery by Leveraging Second-Order Information
von: Wu, Diyuan, et al.
Veröffentlicht: (2024)
von: Wu, Diyuan, et al.
Veröffentlicht: (2024)
Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models
von: Gu, Xiangming, et al.
Veröffentlicht: (2026)
von: Gu, Xiangming, et al.
Veröffentlicht: (2026)
Latent Space Representations of Neural Algorithmic Reasoners
von: Mirjanić, Vladimir V., et al.
Veröffentlicht: (2023)
von: Mirjanić, Vladimir V., et al.
Veröffentlicht: (2023)
Error Feedback Can Accurately Compress Preconditioners
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2023)
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2023)
DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026)
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026)
Round and Round We Go! What makes Rotary Positional Encodings useful?
von: Barbero, Federico, et al.
Veröffentlicht: (2024)
von: Barbero, Federico, et al.
Veröffentlicht: (2024)
FFT-based Dynamic Subspace Selection for Low-Rank Adaptive Optimization of Large Language Models
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2025)
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2025)
Softmax is not Enough (for Sharp Size Generalisation)
von: Veličković, Petar, et al.
Veröffentlicht: (2024)
von: Veličković, Petar, et al.
Veröffentlicht: (2024)
The Illusion of Stochasticity in LLMs
von: Gu, Xiangming, et al.
Veröffentlicht: (2026)
von: Gu, Xiangming, et al.
Veröffentlicht: (2026)
MicroAdam: Accurate Adaptive Optimization with Low Space Overhead and Provable Convergence
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2024)
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2024)
Unified Scaling Laws for Compressed Representations
von: Panferov, Andrei, et al.
Veröffentlicht: (2025)
von: Panferov, Andrei, et al.
Veröffentlicht: (2025)
Asynchronous Algorithmic Alignment with Cocycles
von: Dudzik, Andrew, et al.
Veröffentlicht: (2023)
von: Dudzik, Andrew, et al.
Veröffentlicht: (2023)
Perplexity Cannot Always Tell Right from Wrong
von: Veličković, Petar, et al.
Veröffentlicht: (2026)
von: Veličković, Petar, et al.
Veröffentlicht: (2026)
Should We Attend More or Less? Modulating Attention for Fairness
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2023)
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2023)
Mining Generalizable Activation Functions
von: Vitvitskyi, Alex, et al.
Veröffentlicht: (2026)
von: Vitvitskyi, Alex, et al.
Veröffentlicht: (2026)
Filter Equivariant Functions: A symmetric account of length-general extrapolation on lists
von: Lewis, Owen, et al.
Veröffentlicht: (2025)
von: Lewis, Owen, et al.
Veröffentlicht: (2025)
What Can Grokking Teach Us About Learning Under Nonstationarity?
von: Lyle, Clare, et al.
Veröffentlicht: (2025)
von: Lyle, Clare, et al.
Veröffentlicht: (2025)
LoRDO: Distributed Low-Rank Optimization with Infrequent Communication
von: Jovanović, Andrej, et al.
Veröffentlicht: (2026)
von: Jovanović, Andrej, et al.
Veröffentlicht: (2026)
An Analysis of Human Alignment of Latent Diffusion Models
von: Linhardt, Lorenz, et al.
Veröffentlicht: (2024)
von: Linhardt, Lorenz, et al.
Veröffentlicht: (2024)
Leveraging Classical Algorithms for Graph Neural Networks
von: Wu, Jason, et al.
Veröffentlicht: (2025)
von: Wu, Jason, et al.
Veröffentlicht: (2025)
Disentangling the Causes of Plasticity Loss in Neural Networks
von: Lyle, Clare, et al.
Veröffentlicht: (2024)
von: Lyle, Clare, et al.
Veröffentlicht: (2024)
Normalization and effective learning rates in reinforcement learning
von: Lyle, Clare, et al.
Veröffentlicht: (2024)
von: Lyle, Clare, et al.
Veröffentlicht: (2024)
Why do LLMs attend to the first token?
von: Barbero, Federico, et al.
Veröffentlicht: (2025)
von: Barbero, Federico, et al.
Veröffentlicht: (2025)
Torque-Aware Momentum
von: Malviya, Pranshu, et al.
Veröffentlicht: (2024)
von: Malviya, Pranshu, et al.
Veröffentlicht: (2024)
Transformers meet Neural Algorithmic Reasoners
von: Bounsi, Wilfried, et al.
Veröffentlicht: (2024)
von: Bounsi, Wilfried, et al.
Veröffentlicht: (2024)
Dynamic Event-based Optical Identification and Communication
von: von Arnim, Axel, et al.
Veröffentlicht: (2023)
von: von Arnim, Axel, et al.
Veröffentlicht: (2023)
Promoting Exploration in Memory-Augmented Adam using Critical Momenta
von: Malviya, Pranshu, et al.
Veröffentlicht: (2023)
von: Malviya, Pranshu, et al.
Veröffentlicht: (2023)
Transformers need glasses! Information over-squashing in language tasks
von: Barbero, Federico, et al.
Veröffentlicht: (2024)
von: Barbero, Federico, et al.
Veröffentlicht: (2024)
Fine-Tuned In-Context Learners for Efficient Adaptation
von: Bornschein, Jorg, et al.
Veröffentlicht: (2025)
von: Bornschein, Jorg, et al.
Veröffentlicht: (2025)
Recurrent Aggregators in Neural Algorithmic Reasoning
von: Xu, Kaijia, et al.
Veröffentlicht: (2024)
von: Xu, Kaijia, et al.
Veröffentlicht: (2024)
Reducing Memorisation in Generative Models via Riemannian Bayesian Inference
von: Gegenfurtner, Johanna Marie, et al.
Veröffentlicht: (2026)
von: Gegenfurtner, Johanna Marie, et al.
Veröffentlicht: (2026)
Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
von: Orvieto, Antonio, et al.
Veröffentlicht: (2023)
von: Orvieto, Antonio, et al.
Veröffentlicht: (2023)
Hybrid Decentralized Optimization: Leveraging Both First- and Zeroth-Order Optimizers for Faster Convergence
von: Ansaripour, Matin, et al.
Veröffentlicht: (2022)
von: Ansaripour, Matin, et al.
Veröffentlicht: (2022)
How do language models learn facts? Dynamics, curricula and hallucinations
von: Zucchet, Nicolas, et al.
Veröffentlicht: (2025)
von: Zucchet, Nicolas, et al.
Veröffentlicht: (2025)
Lattice: Learning to Efficiently Compress the Memory
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
Deep Grokking: Would Deep Neural Networks Generalize Better?
von: Fan, Simin, et al.
Veröffentlicht: (2024)
von: Fan, Simin, et al.
Veröffentlicht: (2024)
Non-Stationary Learning of Neural Networks with Automatic Soft Parameter Reset
von: Galashov, Alexandre, et al.
Veröffentlicht: (2024)
von: Galashov, Alexandre, et al.
Veröffentlicht: (2024)
Faithfulness Measurable Masked Language Models
von: Madsen, Andreas, et al.
Veröffentlicht: (2023)
von: Madsen, Andreas, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026) -
LDAdam: Adaptive Optimization from Low-Dimensional Gradient Statistics
von: Robert, Thomas, et al.
Veröffentlicht: (2024) -
The Iterative Optimal Brain Surgeon: Faster Sparse Recovery by Leveraging Second-Order Information
von: Wu, Diyuan, et al.
Veröffentlicht: (2024) -
Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models
von: Gu, Xiangming, et al.
Veröffentlicht: (2026) -
Latent Space Representations of Neural Algorithmic Reasoners
von: Mirjanić, Vladimir V., et al.
Veröffentlicht: (2023)