Generalizing the Geometry of Model Merging Through Frechet Averages
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | da Silva, Marvin F., Adnan, Mohammed, Dangel, Felix, Oore, Sageev |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hide & Seek: Transformer Symmetries Obscure Sharpness & Riemannian Geometry Finds It
von: da Silva, Marvin F., et al.
Veröffentlicht: (2025)
von: da Silva, Marvin F., et al.
Veröffentlicht: (2025)
DiffAug: A Diffuse-and-Denoise Augmentation for Training Robust Classifiers
von: Sastry, Chandramouli, et al.
Veröffentlicht: (2023)
von: Sastry, Chandramouli, et al.
Veröffentlicht: (2023)
Test-Time Training for Depression Detection
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
Convolutions and More as Einsum: A Tensor Network Perspective with Advances for Second-Order Methods
von: Dangel, Felix
Veröffentlicht: (2023)
von: Dangel, Felix
Veröffentlicht: (2023)
Sensitivity of Generative VLMs to Semantically and Lexically Altered Prompts
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
Self-Distillation of Hidden Layers for Self-Supervised Representation Learning
von: Lowe, Scott C., et al.
Veröffentlicht: (2026)
von: Lowe, Scott C., et al.
Veröffentlicht: (2026)
Lowering PyTorch's Memory Consumption for Selective Differentiation
von: Bhatia, Samarth, et al.
Veröffentlicht: (2024)
von: Bhatia, Samarth, et al.
Veröffentlicht: (2024)
Self-Supervised Embeddings for Detecting Individual Symptoms of Depression
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
Predicting Individual Depression Symptoms from Acoustic Features During Speech
von: Rodriguez, Sebastian, et al.
Veröffentlicht: (2024)
von: Rodriguez, Sebastian, et al.
Veröffentlicht: (2024)
Symbolic Music Generation with Non-Differentiable Rule Guided Diffusion
von: Huang, Yujia, et al.
Veröffentlicht: (2024)
von: Huang, Yujia, et al.
Veröffentlicht: (2024)
Efficient Bilevel Optimization with KFAC-Based Hypergradients
von: Liao, Disen, et al.
Veröffentlicht: (2026)
von: Liao, Disen, et al.
Veröffentlicht: (2026)
On the Disconnect Between Theory and Practice of Neural Networks: Limits of the NTK Perspective
von: Wenger, Jonathan, et al.
Veröffentlicht: (2023)
von: Wenger, Jonathan, et al.
Veröffentlicht: (2023)
An Empirical Study into Clustering of Unseen Datasets with Self-Supervised Encoders
von: Lowe, Scott C., et al.
Veröffentlicht: (2024)
von: Lowe, Scott C., et al.
Veröffentlicht: (2024)
SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
Revisiting Weight Averaging for Model Merging
von: Choi, Jiho, et al.
Veröffentlicht: (2024)
von: Choi, Jiho, et al.
Veröffentlicht: (2024)
What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis
von: Ormaniec, Weronika, et al.
Veröffentlicht: (2024)
von: Ormaniec, Weronika, et al.
Veröffentlicht: (2024)
Kronecker-Factored Approximate Curvature for Physics-Informed Neural Networks
von: Dangel, Felix, et al.
Veröffentlicht: (2024)
von: Dangel, Felix, et al.
Veröffentlicht: (2024)
Kronecker-factored Approximate Curvature (KFAC) From Scratch
von: Dangel, Felix, et al.
Veröffentlicht: (2025)
von: Dangel, Felix, et al.
Veröffentlicht: (2025)
Collapsing Taylor Mode Automatic Differentiation
von: Dangel, Felix, et al.
Veröffentlicht: (2025)
von: Dangel, Felix, et al.
Veröffentlicht: (2025)
Improving Energy Natural Gradient Descent through Woodbury, Momentum, and Randomization
von: Guzmán-Cordero, Andrés, et al.
Veröffentlicht: (2025)
von: Guzmán-Cordero, Andrés, et al.
Veröffentlicht: (2025)
Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator
von: Li, YuXin, et al.
Veröffentlicht: (2025)
von: Li, YuXin, et al.
Veröffentlicht: (2025)
Theoretical and Practical Analysis of Fréchet Regression via Comparison Geometry
von: Kimura, Masanari, et al.
Veröffentlicht: (2025)
von: Kimura, Masanari, et al.
Veröffentlicht: (2025)
Sketching Low-Rank Plus Diagonal Matrices
von: Fernandez, Andres, et al.
Veröffentlicht: (2025)
von: Fernandez, Andres, et al.
Veröffentlicht: (2025)
Parameter-Efficient Checkpoint Merging via Metrics-Weighted Averaging
von: Yu, Shi Jie, et al.
Veröffentlicht: (2025)
von: Yu, Shi Jie, et al.
Veröffentlicht: (2025)
Revisiting Scalable Hessian Diagonal Approximations for Applications in Reinforcement Learning
von: Elsayed, Mohamed, et al.
Veröffentlicht: (2024)
von: Elsayed, Mohamed, et al.
Veröffentlicht: (2024)
Model Merging on Loss Landscape: A Geometry Perspective
von: Lu, Juanwu, et al.
Veröffentlicht: (2026)
von: Lu, Juanwu, et al.
Veröffentlicht: (2026)
Reparametrizing Shampoo and SOAP for Subspace Basis Updates and BFloat16 Storage
von: Milligan, Alan, et al.
Veröffentlicht: (2026)
von: Milligan, Alan, et al.
Veröffentlicht: (2026)
Position: Curvature Matrices Should Be Democratized via Linear Operators
von: Dangel, Felix, et al.
Veröffentlicht: (2025)
von: Dangel, Felix, et al.
Veröffentlicht: (2025)
Soft Merging of Experts with Adaptive Routing
von: Muqeeth, Mohammed, et al.
Veröffentlicht: (2023)
von: Muqeeth, Mohammed, et al.
Veröffentlicht: (2023)
Merging in a Bottle: Differentiable Adaptive Merging (DAM) and the Path from Averaging to Automation
von: Gauthier-Caron, Thomas, et al.
Veröffentlicht: (2024)
von: Gauthier-Caron, Thomas, et al.
Veröffentlicht: (2024)
Deep Fréchet Regression
von: Iao, Su I, et al.
Veröffentlicht: (2024)
von: Iao, Su I, et al.
Veröffentlicht: (2024)
Fréchet Geodesic Boosting
von: Zhou, Yidong, et al.
Veröffentlicht: (2025)
von: Zhou, Yidong, et al.
Veröffentlicht: (2025)
Merge and Guide: Unifying Model Merging and Guided Decoding for Controllable Multi-Objective Generation
von: Xie, Guofu, et al.
Veröffentlicht: (2025)
von: Xie, Guofu, et al.
Veröffentlicht: (2025)
Merging Smarter, Generalizing Better: Enhancing Model Merging on OOD Data
von: Zhang, Bingjie, et al.
Veröffentlicht: (2025)
von: Zhang, Bingjie, et al.
Veröffentlicht: (2025)
Spectral-factorized Positive-definite Curvature Learning for NN Training
von: Lin, Wu, et al.
Veröffentlicht: (2025)
von: Lin, Wu, et al.
Veröffentlicht: (2025)
Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization
von: Lin, Wu, et al.
Veröffentlicht: (2025)
von: Lin, Wu, et al.
Veröffentlicht: (2025)
Tracking Universal Features Through Fine-Tuning and Model Merging
von: Horn, Niels, et al.
Veröffentlicht: (2024)
von: Horn, Niels, et al.
Veröffentlicht: (2024)
Bridging Training and Merging Through Momentum-Aware Optimization
von: Moayedikia, Alireza, et al.
Veröffentlicht: (2025)
von: Moayedikia, Alireza, et al.
Veröffentlicht: (2025)
FedMerge: Federated Personalization via Model Merging
von: Chen, Shutong, et al.
Veröffentlicht: (2025)
von: Chen, Shutong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Hide & Seek: Transformer Symmetries Obscure Sharpness & Riemannian Geometry Finds It
von: da Silva, Marvin F., et al.
Veröffentlicht: (2025) -
DiffAug: A Diffuse-and-Denoise Augmentation for Training Robust Classifiers
von: Sastry, Chandramouli, et al.
Veröffentlicht: (2023) -
Test-Time Training for Depression Detection
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024) -
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024) -
Convolutions and More as Einsum: A Tensor Network Perspective with Advances for Second-Order Methods
von: Dangel, Felix
Veröffentlicht: (2023)