Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gan, Yulu, Isola, Phillip |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ReasonCACHE: Teaching LLMs To Reason Without Weight Updates
von: Gupta, Sharut, et al.
Veröffentlicht: (2026)
von: Gupta, Sharut, et al.
Veröffentlicht: (2026)
Pretraining on Sleep Data Improves non-Sleep Biosignal Tasks
von: Lehn-Schiøler, William, et al.
Veröffentlicht: (2026)
von: Lehn-Schiøler, William, et al.
Veröffentlicht: (2026)
Training Neural Networks from Scratch with Parallel Low-Rank Adapters
von: Huh, Minyoung, et al.
Veröffentlicht: (2024)
von: Huh, Minyoung, et al.
Veröffentlicht: (2024)
Fairness Aware Reward Optimization
von: Choi, Ching Lam, et al.
Veröffentlicht: (2026)
von: Choi, Ching Lam, et al.
Veröffentlicht: (2026)
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
von: Panda, Ashwinee, et al.
Veröffentlicht: (2025)
von: Panda, Ashwinee, et al.
Veröffentlicht: (2025)
Data Diversity as Implicit Regularization: How Does Diversity Shape the Weight Space of Deep Neural Networks?
von: Ba, Yang, et al.
Veröffentlicht: (2024)
von: Ba, Yang, et al.
Veröffentlicht: (2024)
Mixture of Diverse Size Experts
von: Sun, Manxi, et al.
Veröffentlicht: (2024)
von: Sun, Manxi, et al.
Veröffentlicht: (2024)
On the Emergence of Cross-Task Linearity in the Pretraining-Finetuning Paradigm
von: Zhou, Zhanpeng, et al.
Veröffentlicht: (2024)
von: Zhou, Zhanpeng, et al.
Veröffentlicht: (2024)
Neural Architecture for Fast and Reliable Coagulation Assessment in Clinical Settings: Leveraging Thromboelastography
von: Wang, Yulu, et al.
Veröffentlicht: (2026)
von: Wang, Yulu, et al.
Veröffentlicht: (2026)
Adaptive Length Image Tokenization via Recurrent Allocation
von: Duggal, Shivam, et al.
Veröffentlicht: (2024)
von: Duggal, Shivam, et al.
Veröffentlicht: (2024)
Dense and Diverse Goal Coverage in Multi Goal Reinforcement Learning
von: Singh, Sagalpreet, et al.
Veröffentlicht: (2025)
von: Singh, Sagalpreet, et al.
Veröffentlicht: (2025)
When Losses Align: Gradient-Based Composite Loss Weighting for Efficient Pretraining
von: Karpukhin, Ivan, et al.
Veröffentlicht: (2026)
von: Karpukhin, Ivan, et al.
Veröffentlicht: (2026)
Densely Multiplied Physics Informed Neural Networks
von: Jiang, Feilong, et al.
Veröffentlicht: (2024)
von: Jiang, Feilong, et al.
Veröffentlicht: (2024)
Optimizing Dense Feed-Forward Neural Networks
von: Balderas, Luis, et al.
Veröffentlicht: (2023)
von: Balderas, Luis, et al.
Veröffentlicht: (2023)
Single-pass Adaptive Image Tokenization for Minimum Program Search
von: Duggal, Shivam, et al.
Veröffentlicht: (2025)
von: Duggal, Shivam, et al.
Veröffentlicht: (2025)
Geometric Regularization in Mixture-of-Experts: The Disconnect Between Weights and Activations
von: Kim, Hyunjun
Veröffentlicht: (2026)
von: Kim, Hyunjun
Veröffentlicht: (2026)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
Mixture of Experts (MoE): A Big Data Perspective
von: Gan, Wensheng, et al.
Veröffentlicht: (2025)
von: Gan, Wensheng, et al.
Veröffentlicht: (2025)
MLPMoE: Zero-Shot Architectural Metamorphosis of Dense LLM MLPs into Static Mixture-of-Experts
von: Novikov, Ivan
Veröffentlicht: (2025)
von: Novikov, Ivan
Veröffentlicht: (2025)
Neural Additive Experts: Context-Gated Experts for Controllable Model Additivity
von: Xiong, Guangzhi, et al.
Veröffentlicht: (2026)
von: Xiong, Guangzhi, et al.
Veröffentlicht: (2026)
Beyond Redundancy: Diverse and Specialized Multi-Expert Sparse Autoencoder
von: Xu, Zhen, et al.
Veröffentlicht: (2025)
von: Xu, Zhen, et al.
Veröffentlicht: (2025)
Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
von: O'Brien, Kyle, et al.
Veröffentlicht: (2025)
von: O'Brien, Kyle, et al.
Veröffentlicht: (2025)
Model-Based Reinforcement Learning with Multi-Task Offline Pretraining
von: Pan, Minting, et al.
Veröffentlicht: (2023)
von: Pan, Minting, et al.
Veröffentlicht: (2023)
Each Rank Could be an Expert: Single-Ranked Mixture of Experts LoRA for Multi-Task Learning
von: Zhao, Ziyu, et al.
Veröffentlicht: (2025)
von: Zhao, Ziyu, et al.
Veröffentlicht: (2025)
The Platonic Representation Hypothesis
von: Huh, Minyoung, et al.
Veröffentlicht: (2024)
von: Huh, Minyoung, et al.
Veröffentlicht: (2024)
Learning with Expert Abstractions for Efficient Multi-Task Continuous Control
von: Jewett, Jeff, et al.
Veröffentlicht: (2025)
von: Jewett, Jeff, et al.
Veröffentlicht: (2025)
DrS: Learning Reusable Dense Rewards for Multi-Stage Tasks
von: Mu, Tongzhou, et al.
Veröffentlicht: (2024)
von: Mu, Tongzhou, et al.
Veröffentlicht: (2024)
TAIL: Task-specific Adapters for Imitation Learning with Large Pretrained Models
von: Liu, Zuxin, et al.
Veröffentlicht: (2023)
von: Liu, Zuxin, et al.
Veröffentlicht: (2023)
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
von: Fan, Dongyang, et al.
Veröffentlicht: (2025)
von: Fan, Dongyang, et al.
Veröffentlicht: (2025)
MVMoE: Multi-Task Vehicle Routing Solver with Mixture-of-Experts
von: Zhou, Jianan, et al.
Veröffentlicht: (2024)
von: Zhou, Jianan, et al.
Veröffentlicht: (2024)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
von: Meterez, Alexandru, et al.
Veröffentlicht: (2026)
von: Meterez, Alexandru, et al.
Veröffentlicht: (2026)
Weight Scope Alignment: A Frustratingly Easy Method for Model Merging
von: Xu, Yichu, et al.
Veröffentlicht: (2024)
von: Xu, Yichu, et al.
Veröffentlicht: (2024)
Unlock the Potential of Large Language Models for Predictive Tabular Tasks in Data Science with Table-Specific Pretraining
von: Yang, Yazheng, et al.
Veröffentlicht: (2024)
von: Yang, Yazheng, et al.
Veröffentlicht: (2024)
Neural Inhibition Improves Dynamic Routing and Mixture of Experts
von: Zou, Will Y., et al.
Veröffentlicht: (2025)
von: Zou, Will Y., et al.
Veröffentlicht: (2025)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
von: Hui, Tingfeng, et al.
Veröffentlicht: (2024)
von: Hui, Tingfeng, et al.
Veröffentlicht: (2024)
Stochastic Weight Sharing for Bayesian Neural Networks
von: Lin, Moule, et al.
Veröffentlicht: (2025)
von: Lin, Moule, et al.
Veröffentlicht: (2025)
Diffusion-Based Neural Network Weights Generation
von: Soro, Bedionita, et al.
Veröffentlicht: (2024)
von: Soro, Bedionita, et al.
Veröffentlicht: (2024)
OSDTW: Optimal Shared Depth and Task Weighting for Long-Tailed Recognition
von: Chu, Chang, et al.
Veröffentlicht: (2026)
von: Chu, Chang, et al.
Veröffentlicht: (2026)
Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation Learning
von: Yang, Hanlin, et al.
Veröffentlicht: (2024)
von: Yang, Hanlin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ReasonCACHE: Teaching LLMs To Reason Without Weight Updates
von: Gupta, Sharut, et al.
Veröffentlicht: (2026) -
Pretraining on Sleep Data Improves non-Sleep Biosignal Tasks
von: Lehn-Schiøler, William, et al.
Veröffentlicht: (2026) -
Training Neural Networks from Scratch with Parallel Low-Rank Adapters
von: Huh, Minyoung, et al.
Veröffentlicht: (2024) -
Fairness Aware Reward Optimization
von: Choi, Ching Lam, et al.
Veröffentlicht: (2026) -
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
von: Panda, Ashwinee, et al.
Veröffentlicht: (2025)