Hessian of Perplexity for Large Language Models by PyTorch autograd (Open Source)
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Ilin, Ivan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A resource-efficient model for deep kernel learning
von: D'Amore, Luisa
Veröffentlicht: (2024)
von: D'Amore, Luisa
Veröffentlicht: (2024)
Hybrid Least Squares/Gradient Descent Methods for DeepONets
von: Choi, Jun, et al.
Veröffentlicht: (2025)
von: Choi, Jun, et al.
Veröffentlicht: (2025)
Bare-Metal Tensor Virtualization: Overcoming the Memory Wall in Edge-AI Inference on ARM64
von: Kilictas, Bugra, et al.
Veröffentlicht: (2026)
von: Kilictas, Bugra, et al.
Veröffentlicht: (2026)
Randomized Matrix Sketching for Neural Network Training and Gradient Monitoring
von: Antil, Harbir, et al.
Veröffentlicht: (2025)
von: Antil, Harbir, et al.
Veröffentlicht: (2025)
Layer-Parallel Training for Transformers
von: Jiang, Shuai, et al.
Veröffentlicht: (2026)
von: Jiang, Shuai, et al.
Veröffentlicht: (2026)
A Low-complexity Structured Neural Network to Realize States of Dynamical Systems
von: Aluvihare, Hansaka, et al.
Veröffentlicht: (2025)
von: Aluvihare, Hansaka, et al.
Veröffentlicht: (2025)
Deep Unfolding Network for Nonlinear Multi-Frequency Electrical Impedance Tomography
von: Alberti, Giovanni S., et al.
Veröffentlicht: (2025)
von: Alberti, Giovanni S., et al.
Veröffentlicht: (2025)
Multilevel Training for Kolmogorov Arnold Networks
von: Southworth, Ben S., et al.
Veröffentlicht: (2026)
von: Southworth, Ben S., et al.
Veröffentlicht: (2026)
Approximation of the Proximal Operator of the $\ell_\infty$ Norm Using a Neural Network
von: Linehan, Kathryn, et al.
Veröffentlicht: (2024)
von: Linehan, Kathryn, et al.
Veröffentlicht: (2024)
Recent Advances in Non-convex Smoothness Conditions and Applicability to Deep Linear Neural Networks
von: Patel, Vivak, et al.
Veröffentlicht: (2024)
von: Patel, Vivak, et al.
Veröffentlicht: (2024)
To be or not to be stable, that is the question: understanding neural networks for inverse problems
von: Evangelista, Davide, et al.
Veröffentlicht: (2022)
von: Evangelista, Davide, et al.
Veröffentlicht: (2022)
Space-time parallel scaling of Parareal with a physics-informed Fourier Neural Operator coarse propagator applied to the Black-Scholes equation
von: Ibrahim, Abdul Qadir, et al.
Veröffentlicht: (2024)
von: Ibrahim, Abdul Qadir, et al.
Veröffentlicht: (2024)
Randomized Forward Mode of Automatic Differentiation For Optimization Algorithms
von: Shukla, Khemraj, et al.
Veröffentlicht: (2023)
von: Shukla, Khemraj, et al.
Veröffentlicht: (2023)
Learning Nonlinear Finite Element Solution Operators using Multilayer Perceptrons and Energy Minimization
von: Larson, Mats G., et al.
Veröffentlicht: (2024)
von: Larson, Mats G., et al.
Veröffentlicht: (2024)
A network based approach for unbalanced optimal transport on surfaces
von: Pan, Jiangong, et al.
Veröffentlicht: (2024)
von: Pan, Jiangong, et al.
Veröffentlicht: (2024)
Approximation and Gradient Descent Training with Neural Networks
von: Welper, G.
Veröffentlicht: (2024)
von: Welper, G.
Veröffentlicht: (2024)
Neural-network methods for two-dimensional finite-source reflector design
von: Hacking, Roel, et al.
Veröffentlicht: (2026)
von: Hacking, Roel, et al.
Veröffentlicht: (2026)
RandNet-Parareal: a time-parallel PDE solver using Random Neural Networks
von: Gattiglio, Guglielmo, et al.
Veröffentlicht: (2024)
von: Gattiglio, Guglielmo, et al.
Veröffentlicht: (2024)
Objective Value Change and Shape-Based Accelerated Optimization for the Neural Network Approximation
von: Xie, Pengcheng, et al.
Veröffentlicht: (2025)
von: Xie, Pengcheng, et al.
Veröffentlicht: (2025)
Progressive Power Homotopy for Non-convex Optimization
von: Xu, Chen
Veröffentlicht: (2026)
von: Xu, Chen
Veröffentlicht: (2026)
Kourkoutas-Beta: A Sunspike-Driven Adam Optimizer with Desert Flair
von: Kassinos, Stavros C.
Veröffentlicht: (2025)
von: Kassinos, Stavros C.
Veröffentlicht: (2025)
Neural Preconditioning via Krylov Subspace Geometry
von: Dimola, Nunzio, et al.
Veröffentlicht: (2025)
von: Dimola, Nunzio, et al.
Veröffentlicht: (2025)
The Pontryagin Maximum Principle for Training Convolutional Neural Networks
von: Hofmann, Sebastian, et al.
Veröffentlicht: (2025)
von: Hofmann, Sebastian, et al.
Veröffentlicht: (2025)
Inter-Layer Hessian Analysis of Neural Networks with DAG Architectures
von: Bolshim, Maxim, et al.
Veröffentlicht: (2026)
von: Bolshim, Maxim, et al.
Veröffentlicht: (2026)
Local properties of neural networks through the lens of layer-wise Hessians
von: Bolshim, Maxim, et al.
Veröffentlicht: (2025)
von: Bolshim, Maxim, et al.
Veröffentlicht: (2025)
Stochastic Estimation of the Layer-wise Hessian Trace for Monitoring Neural-network Training
von: Bolshim, Maxim, et al.
Veröffentlicht: (2026)
von: Bolshim, Maxim, et al.
Veröffentlicht: (2026)
Derivative-Informed Fourier Neural Operator: Universal Approximation and Applications to PDE-Constrained Optimization
von: Yao, Boyuan, et al.
Veröffentlicht: (2025)
von: Yao, Boyuan, et al.
Veröffentlicht: (2025)
High-performance matrix-free unfitted finite element operator evaluation
von: Bergbauer, Maximilian, et al.
Veröffentlicht: (2024)
von: Bergbauer, Maximilian, et al.
Veröffentlicht: (2024)
Matrix-Free Evaluation of High-Order Shifted Boundary Finite Element Operators
von: Wichrowski, Michał
Veröffentlicht: (2025)
von: Wichrowski, Michał
Veröffentlicht: (2025)
Polyharmonic Cascade
von: Bakhvalov, Yuriy N.
Veröffentlicht: (2025)
von: Bakhvalov, Yuriy N.
Veröffentlicht: (2025)
SplInterp: Improving our Understanding and Training of Sparse Autoencoders
von: Budd, Jeremy, et al.
Veröffentlicht: (2025)
von: Budd, Jeremy, et al.
Veröffentlicht: (2025)
PETScML: Second-order solvers for training regression problems in Scientific Machine Learning
von: Zampini, Stefano, et al.
Veröffentlicht: (2024)
von: Zampini, Stefano, et al.
Veröffentlicht: (2024)
A Judge Agent Closes the Reliability Gap in AI-Generated Scientific Simulation
von: Yang, Chengshuai
Veröffentlicht: (2026)
von: Yang, Chengshuai
Veröffentlicht: (2026)
Sequential Least-Squares Estimators with Fast Randomized Sketching for Linear Statistical Models
von: Chen, Guan-Yu, et al.
Veröffentlicht: (2025)
von: Chen, Guan-Yu, et al.
Veröffentlicht: (2025)
Improving Matrix Exponential for Generative AI Flows: A Taylor-Based Approach Beyond Paterson--Stockmeyer
von: Sastre, Jorge, et al.
Veröffentlicht: (2025)
von: Sastre, Jorge, et al.
Veröffentlicht: (2025)
A Neumann-Neumann Acceleration with Coarse Space for Domain Decomposition of Extreme Learning Machines
von: Lee, Chang-Ock, et al.
Veröffentlicht: (2025)
von: Lee, Chang-Ock, et al.
Veröffentlicht: (2025)
Parallel-in-time Multilevel Krylov Methods: A Prototype
von: Erlangga, Yogi A.
Veröffentlicht: (2023)
von: Erlangga, Yogi A.
Veröffentlicht: (2023)
Self2Seg: Single-Image Self-Supervised Joint Segmentation and Denoising
von: Gruber, Nadja, et al.
Veröffentlicht: (2023)
von: Gruber, Nadja, et al.
Veröffentlicht: (2023)
Neural Network-Based Parameter Estimation for Non-Autonomous Differential Equations with Discontinuous Signals
von: Jo, Hyeontae, et al.
Veröffentlicht: (2025)
von: Jo, Hyeontae, et al.
Veröffentlicht: (2025)
SDFs from Unoriented Point Clouds using Neural Variational Heat Distances
von: Weidemaier, Samuel, et al.
Veröffentlicht: (2025)
von: Weidemaier, Samuel, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A resource-efficient model for deep kernel learning
von: D'Amore, Luisa
Veröffentlicht: (2024) -
Hybrid Least Squares/Gradient Descent Methods for DeepONets
von: Choi, Jun, et al.
Veröffentlicht: (2025) -
Bare-Metal Tensor Virtualization: Overcoming the Memory Wall in Edge-AI Inference on ARM64
von: Kilictas, Bugra, et al.
Veröffentlicht: (2026) -
Randomized Matrix Sketching for Neural Network Training and Gradient Monitoring
von: Antil, Harbir, et al.
Veröffentlicht: (2025) -
Layer-Parallel Training for Transformers
von: Jiang, Shuai, et al.
Veröffentlicht: (2026)