Do We Really Need Permutations? Impact of Model Width on Linear Mode Connectivity
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ito, Akira, Yamada, Masanori, Chijiwa, Daiki, Kumagai, Atsutoshi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Analysis of Linear Mode Connectivity via Permutation-Based Weight Matching: With Insights into Other Permutation Search Methods
von: Ito, Akira, et al.
Veröffentlicht: (2024)
von: Ito, Akira, et al.
Veröffentlicht: (2024)
Transfer Learning with Pre-trained Conditional Generative Models
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2022)
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2022)
Toward Data Efficient Model Merging between Different Datasets without Performance Degradation
von: Yamada, Masanori, et al.
Veröffentlicht: (2023)
von: Yamada, Masanori, et al.
Veröffentlicht: (2023)
Do We Really Even Need Data?
von: Hoffman, Kentaro, et al.
Veröffentlicht: (2024)
von: Hoffman, Kentaro, et al.
Veröffentlicht: (2024)
Relative Density Ratio Optimization for Stable and Statistically Consistent Model Alignment
von: Takahashi, Hiroshi, et al.
Veröffentlicht: (2026)
von: Takahashi, Hiroshi, et al.
Veröffentlicht: (2026)
Analyzing the Role of Permutation Invariance in Linear Mode Connectivity
von: Zhan, Keyao, et al.
Veröffentlicht: (2025)
von: Zhan, Keyao, et al.
Veröffentlicht: (2025)
Meta-learning for Positive-unlabeled Classification
von: Kumagai, Atsutoshi, et al.
Veröffentlicht: (2024)
von: Kumagai, Atsutoshi, et al.
Veröffentlicht: (2024)
Covariance-aware Feature Alignment with Pre-computed Source Statistics for Test-time Adaptation to Multiple Image Corruptions
von: Adachi, Kazuki, et al.
Veröffentlicht: (2022)
von: Adachi, Kazuki, et al.
Veröffentlicht: (2022)
MambaOut: Do We Really Need Mamba for Vision?
von: Yu, Weihao, et al.
Veröffentlicht: (2024)
von: Yu, Weihao, et al.
Veröffentlicht: (2024)
Is the Hard-Label Cryptanalytic Model Extraction Really Polynomial?
von: Ito, Akira, et al.
Veröffentlicht: (2025)
von: Ito, Akira, et al.
Veröffentlicht: (2025)
Do We Really Even Need Data? A Modern Look at Drawing Inference with Predicted Data
von: Salerno, Stephen, et al.
Veröffentlicht: (2025)
von: Salerno, Stephen, et al.
Veröffentlicht: (2025)
Deep Positive-Unlabeled Anomaly Detection for Contaminated Unlabeled Data
von: Takahashi, Hiroshi, et al.
Veröffentlicht: (2024)
von: Takahashi, Hiroshi, et al.
Veröffentlicht: (2024)
Positive-Unlabeled Diffusion Models for Preventing Sensitive Data Generation
von: Takahashi, Hiroshi, et al.
Veröffentlicht: (2025)
von: Takahashi, Hiroshi, et al.
Veröffentlicht: (2025)
Do We Really Need to Design New Byzantine-robust Aggregation Rules?
von: Fang, Minghong, et al.
Veröffentlicht: (2025)
von: Fang, Minghong, et al.
Veröffentlicht: (2025)
Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models
von: Yamashita, Tomoya, et al.
Veröffentlicht: (2025)
von: Yamashita, Tomoya, et al.
Veröffentlicht: (2025)
Meta-learning Representations for Learning from Multiple Annotators
von: Kumagai, Atsutoshi, et al.
Veröffentlicht: (2025)
von: Kumagai, Atsutoshi, et al.
Veröffentlicht: (2025)
Test-time Adaptation for Regression by Subspace Alignment
von: Adachi, Kazuki, et al.
Veröffentlicht: (2024)
von: Adachi, Kazuki, et al.
Veröffentlicht: (2024)
Do Large Language Models (Really) Need Statistical Foundations?
von: Su, Weijie
Veröffentlicht: (2025)
von: Su, Weijie
Veröffentlicht: (2025)
Triple-BERT: Do We Really Need MARL for Order Dispatch on Ride-Sharing Platforms?
von: Zhao, Zijian, et al.
Veröffentlicht: (2025)
von: Zhao, Zijian, et al.
Veröffentlicht: (2025)
Portable Reward Tuning: Towards Reusable Fine-Tuning across Different Pretrained Models
von: Chijiwa, Daiki, et al.
Veröffentlicht: (2025)
von: Chijiwa, Daiki, et al.
Veröffentlicht: (2025)
Do We Really Need Quantum Machine Learning?: A Multidimensional Empirical Study
von: Vhaduri, Sudip, et al.
Veröffentlicht: (2026)
von: Vhaduri, Sudip, et al.
Veröffentlicht: (2026)
Rationale-Enhanced Decoding for Multi-modal Chain-of-Thought
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2025)
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2025)
Do We Really Need Graph Convolution During Training? Light Post-Training Graph-ODE for Efficient Recommendation
von: Zhang, Weizhi, et al.
Veröffentlicht: (2024)
von: Zhang, Weizhi, et al.
Veröffentlicht: (2024)
Landscaping Linear Mode Connectivity
von: Singh, Sidak Pal, et al.
Veröffentlicht: (2024)
von: Singh, Sidak Pal, et al.
Veröffentlicht: (2024)
Zero-shot Concept Bottleneck Models
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2025)
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2025)
Parallel In-context Learning for Large Vision Language Models
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2026)
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2026)
Generalized Linear Mode Connectivity for Transformers
von: Theus, Alexander, et al.
Veröffentlicht: (2025)
von: Theus, Alexander, et al.
Veröffentlicht: (2025)
Layer-wise Linear Mode Connectivity
von: Adilova, Linara, et al.
Veröffentlicht: (2023)
von: Adilova, Linara, et al.
Veröffentlicht: (2023)
Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?
von: Wang, Yanbo, et al.
Veröffentlicht: (2025)
von: Wang, Yanbo, et al.
Veröffentlicht: (2025)
On Linear Mode Connectivity of Mixture-of-Experts Architectures
von: Tran, Viet-Hoang, et al.
Veröffentlicht: (2025)
von: Tran, Viet-Hoang, et al.
Veröffentlicht: (2025)
Linear Mode Connectivity in Differentiable Tree Ensembles
von: Kanoh, Ryuichi, et al.
Veröffentlicht: (2024)
von: Kanoh, Ryuichi, et al.
Veröffentlicht: (2024)
Adaptive Random Feature Regularization on Fine-tuning Deep Neural Networks
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2024)
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2024)
Post-pre-training for Modality Alignment in Vision-Language Foundation Models
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2025)
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2025)
Do We Need Transformers to Play FPS Video Games?
von: Batth, Karmanbir, et al.
Veröffentlicht: (2025)
von: Batth, Karmanbir, et al.
Veröffentlicht: (2025)
Do We Need Frontier Models to Verify Mathematical Proofs?
von: Naik, Aaditya, et al.
Veröffentlicht: (2026)
von: Naik, Aaditya, et al.
Veröffentlicht: (2026)
The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms
von: Otsuka, Hikari, et al.
Veröffentlicht: (2025)
von: Otsuka, Hikari, et al.
Veröffentlicht: (2025)
Linear Mode Connectivity in Sparse Neural Networks
von: McDermott, Luke, et al.
Veröffentlicht: (2023)
von: McDermott, Luke, et al.
Veröffentlicht: (2023)
Why Do We Need Weight Decay in Modern Deep Learning?
von: D'Angelo, Francesco, et al.
Veröffentlicht: (2023)
von: D'Angelo, Francesco, et al.
Veröffentlicht: (2023)
Proving Linear Mode Connectivity of Neural Networks via Optimal Transport
von: Ferbach, Damien, et al.
Veröffentlicht: (2023)
von: Ferbach, Damien, et al.
Veröffentlicht: (2023)
Do You Really Need Public Data? Surrogate Public Data for Differential Privacy on Tabular Data
von: Hod, Shlomi, et al.
Veröffentlicht: (2025)
von: Hod, Shlomi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Analysis of Linear Mode Connectivity via Permutation-Based Weight Matching: With Insights into Other Permutation Search Methods
von: Ito, Akira, et al.
Veröffentlicht: (2024) -
Transfer Learning with Pre-trained Conditional Generative Models
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2022) -
Toward Data Efficient Model Merging between Different Datasets without Performance Degradation
von: Yamada, Masanori, et al.
Veröffentlicht: (2023) -
Do We Really Even Need Data?
von: Hoffman, Kentaro, et al.
Veröffentlicht: (2024) -
Relative Density Ratio Optimization for Stable and Statistically Consistent Model Alignment
von: Takahashi, Hiroshi, et al.
Veröffentlicht: (2026)