Saved in:
Bibliographic Details
Main Author: Di Sipio, Riccardo
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2506.15830
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • Optimization in large language models (LLMs) unfolds over high-dimensional parameter spaces with non-Euclidean structure. Information geometry frames this landscape using the Fisher information metric, enabling more principled learning via natural gradient descent. Though often impractical, this geometric lens clarifies phenomena such as sharp minima, generalization, and observed scaling laws. We argue that curvature-based approaches deepen our understanding of LLM training. Finally, we speculate on quantum analogies based on the Fubini-Study metric and Quantum Fisher Information, hinting at efficient optimization in quantum-enhanced systems.