Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights
Fuente:
arXiv
Salvato in:
| Autori principali: | Gan, Yulu, Isola, Phillip |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ReasonCACHE: Teaching LLMs To Reason Without Weight Updates
di: Gupta, Sharut, et al.
Pubblicazione: (2026)
di: Gupta, Sharut, et al.
Pubblicazione: (2026)
Pretraining on Sleep Data Improves non-Sleep Biosignal Tasks
di: Lehn-Schiøler, William, et al.
Pubblicazione: (2026)
di: Lehn-Schiøler, William, et al.
Pubblicazione: (2026)
Training Neural Networks from Scratch with Parallel Low-Rank Adapters
di: Huh, Minyoung, et al.
Pubblicazione: (2024)
di: Huh, Minyoung, et al.
Pubblicazione: (2024)
Fairness Aware Reward Optimization
di: Choi, Ching Lam, et al.
Pubblicazione: (2026)
di: Choi, Ching Lam, et al.
Pubblicazione: (2026)
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
di: Panda, Ashwinee, et al.
Pubblicazione: (2025)
di: Panda, Ashwinee, et al.
Pubblicazione: (2025)
Data Diversity as Implicit Regularization: How Does Diversity Shape the Weight Space of Deep Neural Networks?
di: Ba, Yang, et al.
Pubblicazione: (2024)
di: Ba, Yang, et al.
Pubblicazione: (2024)
Mixture of Diverse Size Experts
di: Sun, Manxi, et al.
Pubblicazione: (2024)
di: Sun, Manxi, et al.
Pubblicazione: (2024)
On the Emergence of Cross-Task Linearity in the Pretraining-Finetuning Paradigm
di: Zhou, Zhanpeng, et al.
Pubblicazione: (2024)
di: Zhou, Zhanpeng, et al.
Pubblicazione: (2024)
Neural Architecture for Fast and Reliable Coagulation Assessment in Clinical Settings: Leveraging Thromboelastography
di: Wang, Yulu, et al.
Pubblicazione: (2026)
di: Wang, Yulu, et al.
Pubblicazione: (2026)
Adaptive Length Image Tokenization via Recurrent Allocation
di: Duggal, Shivam, et al.
Pubblicazione: (2024)
di: Duggal, Shivam, et al.
Pubblicazione: (2024)
Dense and Diverse Goal Coverage in Multi Goal Reinforcement Learning
di: Singh, Sagalpreet, et al.
Pubblicazione: (2025)
di: Singh, Sagalpreet, et al.
Pubblicazione: (2025)
When Losses Align: Gradient-Based Composite Loss Weighting for Efficient Pretraining
di: Karpukhin, Ivan, et al.
Pubblicazione: (2026)
di: Karpukhin, Ivan, et al.
Pubblicazione: (2026)
Densely Multiplied Physics Informed Neural Networks
di: Jiang, Feilong, et al.
Pubblicazione: (2024)
di: Jiang, Feilong, et al.
Pubblicazione: (2024)
Optimizing Dense Feed-Forward Neural Networks
di: Balderas, Luis, et al.
Pubblicazione: (2023)
di: Balderas, Luis, et al.
Pubblicazione: (2023)
Single-pass Adaptive Image Tokenization for Minimum Program Search
di: Duggal, Shivam, et al.
Pubblicazione: (2025)
di: Duggal, Shivam, et al.
Pubblicazione: (2025)
Geometric Regularization in Mixture-of-Experts: The Disconnect Between Weights and Activations
di: Kim, Hyunjun
Pubblicazione: (2026)
di: Kim, Hyunjun
Pubblicazione: (2026)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
di: Kim, Junhyuck, et al.
Pubblicazione: (2026)
di: Kim, Junhyuck, et al.
Pubblicazione: (2026)
Mixture of Experts (MoE): A Big Data Perspective
di: Gan, Wensheng, et al.
Pubblicazione: (2025)
di: Gan, Wensheng, et al.
Pubblicazione: (2025)
MLPMoE: Zero-Shot Architectural Metamorphosis of Dense LLM MLPs into Static Mixture-of-Experts
di: Novikov, Ivan
Pubblicazione: (2025)
di: Novikov, Ivan
Pubblicazione: (2025)
Neural Additive Experts: Context-Gated Experts for Controllable Model Additivity
di: Xiong, Guangzhi, et al.
Pubblicazione: (2026)
di: Xiong, Guangzhi, et al.
Pubblicazione: (2026)
Beyond Redundancy: Diverse and Specialized Multi-Expert Sparse Autoencoder
di: Xu, Zhen, et al.
Pubblicazione: (2025)
di: Xu, Zhen, et al.
Pubblicazione: (2025)
Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
di: O'Brien, Kyle, et al.
Pubblicazione: (2025)
di: O'Brien, Kyle, et al.
Pubblicazione: (2025)
Model-Based Reinforcement Learning with Multi-Task Offline Pretraining
di: Pan, Minting, et al.
Pubblicazione: (2023)
di: Pan, Minting, et al.
Pubblicazione: (2023)
Each Rank Could be an Expert: Single-Ranked Mixture of Experts LoRA for Multi-Task Learning
di: Zhao, Ziyu, et al.
Pubblicazione: (2025)
di: Zhao, Ziyu, et al.
Pubblicazione: (2025)
The Platonic Representation Hypothesis
di: Huh, Minyoung, et al.
Pubblicazione: (2024)
di: Huh, Minyoung, et al.
Pubblicazione: (2024)
Learning with Expert Abstractions for Efficient Multi-Task Continuous Control
di: Jewett, Jeff, et al.
Pubblicazione: (2025)
di: Jewett, Jeff, et al.
Pubblicazione: (2025)
DrS: Learning Reusable Dense Rewards for Multi-Stage Tasks
di: Mu, Tongzhou, et al.
Pubblicazione: (2024)
di: Mu, Tongzhou, et al.
Pubblicazione: (2024)
TAIL: Task-specific Adapters for Imitation Learning with Large Pretrained Models
di: Liu, Zuxin, et al.
Pubblicazione: (2023)
di: Liu, Zuxin, et al.
Pubblicazione: (2023)
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
di: Fan, Dongyang, et al.
Pubblicazione: (2025)
di: Fan, Dongyang, et al.
Pubblicazione: (2025)
MVMoE: Multi-Task Vehicle Routing Solver with Mixture-of-Experts
di: Zhou, Jianan, et al.
Pubblicazione: (2024)
di: Zhou, Jianan, et al.
Pubblicazione: (2024)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
di: Meterez, Alexandru, et al.
Pubblicazione: (2026)
di: Meterez, Alexandru, et al.
Pubblicazione: (2026)
Weight Scope Alignment: A Frustratingly Easy Method for Model Merging
di: Xu, Yichu, et al.
Pubblicazione: (2024)
di: Xu, Yichu, et al.
Pubblicazione: (2024)
Unlock the Potential of Large Language Models for Predictive Tabular Tasks in Data Science with Table-Specific Pretraining
di: Yang, Yazheng, et al.
Pubblicazione: (2024)
di: Yang, Yazheng, et al.
Pubblicazione: (2024)
Neural Inhibition Improves Dynamic Routing and Mixture of Experts
di: Zou, Will Y., et al.
Pubblicazione: (2025)
di: Zou, Will Y., et al.
Pubblicazione: (2025)
Stochastic Weight Sharing for Bayesian Neural Networks
di: Lin, Moule, et al.
Pubblicazione: (2025)
di: Lin, Moule, et al.
Pubblicazione: (2025)
Diffusion-Based Neural Network Weights Generation
di: Soro, Bedionita, et al.
Pubblicazione: (2024)
di: Soro, Bedionita, et al.
Pubblicazione: (2024)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
di: Pan, Bowen, et al.
Pubblicazione: (2024)
di: Pan, Bowen, et al.
Pubblicazione: (2024)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
di: Hui, Tingfeng, et al.
Pubblicazione: (2024)
di: Hui, Tingfeng, et al.
Pubblicazione: (2024)
OSDTW: Optimal Shared Depth and Task Weighting for Long-Tailed Recognition
di: Chu, Chang, et al.
Pubblicazione: (2026)
di: Chu, Chang, et al.
Pubblicazione: (2026)
Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation Learning
di: Yang, Hanlin, et al.
Pubblicazione: (2024)
di: Yang, Hanlin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
ReasonCACHE: Teaching LLMs To Reason Without Weight Updates
di: Gupta, Sharut, et al.
Pubblicazione: (2026) -
Pretraining on Sleep Data Improves non-Sleep Biosignal Tasks
di: Lehn-Schiøler, William, et al.
Pubblicazione: (2026) -
Training Neural Networks from Scratch with Parallel Low-Rank Adapters
di: Huh, Minyoung, et al.
Pubblicazione: (2024) -
Fairness Aware Reward Optimization
di: Choi, Ching Lam, et al.
Pubblicazione: (2026) -
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
di: Panda, Ashwinee, et al.
Pubblicazione: (2025)