Sparse Mixture-of-Experts for Compositional Generalization: Empirical Evidence and Theoretical Foundations of Optimal Sparsity
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Jinze, Wang, Peihao, Yang, Junjie, Cai, Ruisi, Liu, Gaowen, Srinivasa, Jayanth, Kompella, Ramana Rao, Liang, Yingbin, Wang, Zhangyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generalization Error Analysis for Sparse Mixture-of-Experts: A Preliminary Study
von: Zhao, Jinze, et al.
Veröffentlicht: (2024)
von: Zhao, Jinze, et al.
Veröffentlicht: (2024)
Meta ControlNet: Enhancing Task Adaptation via Meta Learning
von: Yang, Junjie, et al.
Veröffentlicht: (2023)
von: Yang, Junjie, et al.
Veröffentlicht: (2023)
Context Bootstrapped Reinforcement Learning
von: Agashe, Saaket, et al.
Veröffentlicht: (2026)
von: Agashe, Saaket, et al.
Veröffentlicht: (2026)
Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design
von: Cai, Ruisi, et al.
Veröffentlicht: (2024)
von: Cai, Ruisi, et al.
Veröffentlicht: (2024)
VLA Knows Its Limits
von: Wang, Haoxuan, et al.
Veröffentlicht: (2026)
von: Wang, Haoxuan, et al.
Veröffentlicht: (2026)
Why Neural Network Can Discover Symbolic Structures with Gradient-based Training: An Algebraic and Geometric Foundation for Neurosymbolic Reasoning
von: Wang, Peihao, et al.
Veröffentlicht: (2025)
von: Wang, Peihao, et al.
Veröffentlicht: (2025)
Efficient Multitask Dense Predictor via Binarization
von: Shang, Yuzhang, et al.
Veröffentlicht: (2024)
von: Shang, Yuzhang, et al.
Veröffentlicht: (2024)
Targeted Forgetting of Image Subgroups in CLIP Models
von: Zhang, Zeliang, et al.
Veröffentlicht: (2025)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2025)
Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR
von: Fan, Chongyu, et al.
Veröffentlicht: (2026)
von: Fan, Chongyu, et al.
Veröffentlicht: (2026)
GIFSplat: Generative Prior-Guided Iterative Feed-Forward 3D Gaussian Splatting from Sparse Views
von: Chen, Tianyu, et al.
Veröffentlicht: (2026)
von: Chen, Tianyu, et al.
Veröffentlicht: (2026)
Orientation-anchored Hyper-Gaussian for 4D Reconstruction from Casual Videos
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
Motion Marionette: Rethinking Rigid Motion Transfer via Prior Guidance
von: Wang, Haoxuan, et al.
Veröffentlicht: (2025)
von: Wang, Haoxuan, et al.
Veröffentlicht: (2025)
Real-Time Robot Execution with Masked Action Chunking
von: Wang, Haoxuan, et al.
Veröffentlicht: (2026)
von: Wang, Haoxuan, et al.
Veröffentlicht: (2026)
Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2025)
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2025)
Diverse Score Distillation
von: Xu, Yanbo, et al.
Veröffentlicht: (2024)
von: Xu, Yanbo, et al.
Veröffentlicht: (2024)
Attention Reveals More Than Tokens: Training-Free Long-Context Reasoning with Attention-guided Retrieval
von: Zhang, Yuwei, et al.
Veröffentlicht: (2025)
von: Zhang, Yuwei, et al.
Veröffentlicht: (2025)
SwiftNDC: Fast Neural Depth Correction for High-Fidelity 3D Reconstruction
von: Han, Kang, et al.
Veröffentlicht: (2026)
von: Han, Kang, et al.
Veröffentlicht: (2026)
Urban Scene Diffusion through Semantic Occupancy Map
von: Zhang, Junge, et al.
Veröffentlicht: (2024)
von: Zhang, Junge, et al.
Veröffentlicht: (2024)
Towards Vector Optimization on Low-Dimensional Vector Symbolic Architecture
von: Duan, Shijin, et al.
Veröffentlicht: (2025)
von: Duan, Shijin, et al.
Veröffentlicht: (2025)
Riemannian Multinomial Logistics Regression for SPD Neural Networks
von: Chen, Ziheng, et al.
Veröffentlicht: (2023)
von: Chen, Ziheng, et al.
Veröffentlicht: (2023)
Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding
von: Zhu, Jiajun, et al.
Veröffentlicht: (2025)
von: Zhu, Jiajun, et al.
Veröffentlicht: (2025)
A Neurosymbolic Agent System for Compositional Visual Reasoning
von: Xu, Yichang, et al.
Veröffentlicht: (2025)
von: Xu, Yichang, et al.
Veröffentlicht: (2025)
Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing
von: Wang, Peihao, et al.
Veröffentlicht: (2024)
von: Wang, Peihao, et al.
Veröffentlicht: (2024)
$\nabla$-Reasoner: LLM Reasoning via Test-Time Gradient Descent in Latent Space
von: Wang, Peihao, et al.
Veröffentlicht: (2026)
von: Wang, Peihao, et al.
Veröffentlicht: (2026)
TIER: Trajectory-Invariant Execution Rewards for Multi-Step Tool Composition
von: Kulkarni, Anay, et al.
Veröffentlicht: (2026)
von: Kulkarni, Anay, et al.
Veröffentlicht: (2026)
Rethinking PGD Attack: Is Sign Function Necessary?
von: Yang, Junjie, et al.
Veröffentlicht: (2023)
von: Yang, Junjie, et al.
Veröffentlicht: (2023)
Enabling Elastic Model Serving with MultiWorld
von: Lee, Myungjin, et al.
Veröffentlicht: (2024)
von: Lee, Myungjin, et al.
Veröffentlicht: (2024)
Neural Networks with Sparse Activation Induced by Large Bias: Tighter Analysis with Bias-Generalized NTK
von: Yang, Hongru, et al.
Veröffentlicht: (2023)
von: Yang, Hongru, et al.
Veröffentlicht: (2023)
Vision Verification Enhanced Fusion of VLMs for Efficient Visual Reasoning
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2026)
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2026)
Reversing the Forget-Retain Objectives: An Efficient LLM Unlearning Framework from Logit Difference
von: Ji, Jiabao, et al.
Veröffentlicht: (2024)
von: Ji, Jiabao, et al.
Veröffentlicht: (2024)
ProDiF: Protecting Domain-Invariant Features to Secure Pre-Trained Models Against Extraction
von: Zhou, Tong, et al.
Veröffentlicht: (2025)
von: Zhou, Tong, et al.
Veröffentlicht: (2025)
Navigating the Clutter: Waypoint-Based Bi-Level Planning for Multi-Robot Systems
von: Ji, Jiabao, et al.
Veröffentlicht: (2026)
von: Ji, Jiabao, et al.
Veröffentlicht: (2026)
Collision- and Reachability-Aware Multi-Robot Control with Grounded LLM Planners
von: Ji, Jiabao, et al.
Veröffentlicht: (2025)
von: Ji, Jiabao, et al.
Veröffentlicht: (2025)
ConQuER: Modular Architectures for Control and Bias Mitigation in IQP Quantum Generative Models
von: Zou, Xiaocheng, et al.
Veröffentlicht: (2025)
von: Zou, Xiaocheng, et al.
Veröffentlicht: (2025)
The Quantum-Cryptographic Co-evolution
von: Kundu, Ashish, et al.
Veröffentlicht: (2026)
von: Kundu, Ashish, et al.
Veröffentlicht: (2026)
Training-Free Semantic Segmentation via LLM-Supervision
von: Sun, Wenfang, et al.
Veröffentlicht: (2024)
von: Sun, Wenfang, et al.
Veröffentlicht: (2024)
From Trojan Horses to Castle Walls: Unveiling Bilateral Data Poisoning Effects in Diffusion Models
von: Pan, Zhuoshi, et al.
Veröffentlicht: (2023)
von: Pan, Zhuoshi, et al.
Veröffentlicht: (2023)
Position: Weight Space Should Be a First-Class Generative AI Modality
von: Wang, Zhangyang, et al.
Veröffentlicht: (2026)
von: Wang, Zhangyang, et al.
Veröffentlicht: (2026)
LoCoCo: Dropping In Convolutions for Long Context Compression
von: Cai, Ruisi, et al.
Veröffentlicht: (2024)
von: Cai, Ruisi, et al.
Veröffentlicht: (2024)
Answer is All You Need: Instruction-following Text Embedding via Answering the Question
von: Peng, Letian, et al.
Veröffentlicht: (2024)
von: Peng, Letian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Generalization Error Analysis for Sparse Mixture-of-Experts: A Preliminary Study
von: Zhao, Jinze, et al.
Veröffentlicht: (2024) -
Meta ControlNet: Enhancing Task Adaptation via Meta Learning
von: Yang, Junjie, et al.
Veröffentlicht: (2023) -
Context Bootstrapped Reinforcement Learning
von: Agashe, Saaket, et al.
Veröffentlicht: (2026) -
Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design
von: Cai, Ruisi, et al.
Veröffentlicht: (2024) -
VLA Knows Its Limits
von: Wang, Haoxuan, et al.
Veröffentlicht: (2026)