Autonomy-of-Experts Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lv, Ang, Xie, Ruobing, Qian, Yining, Wu, Songhao, Sun, Xingwu, Kang, Zhanhui, Wang, Di, Yan, Rui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Language Models "Grok" to Copy
von: Lv, Ang, et al.
Veröffentlicht: (2024)
von: Lv, Ang, et al.
Veröffentlicht: (2024)
More Expressive Attention with Negative Weights
von: Lv, Ang, et al.
Veröffentlicht: (2024)
von: Lv, Ang, et al.
Veröffentlicht: (2024)
Proximal Supervised Fine-Tuning
von: Zhu, Wenhong, et al.
Veröffentlicht: (2025)
von: Zhu, Wenhong, et al.
Veröffentlicht: (2025)
Self-Distillation for Multi-Token Prediction
von: Zhao, Guoliang, et al.
Veröffentlicht: (2026)
von: Zhao, Guoliang, et al.
Veröffentlicht: (2026)
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason
von: Lv, Ang, et al.
Veröffentlicht: (2025)
von: Lv, Ang, et al.
Veröffentlicht: (2025)
Towards a Comprehensive Scaling Law of Mixture-of-Experts
von: Zhao, Guoliang, et al.
Veröffentlicht: (2025)
von: Zhao, Guoliang, et al.
Veröffentlicht: (2025)
Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization
von: Fu, Yuhan, et al.
Veröffentlicht: (2024)
von: Fu, Yuhan, et al.
Veröffentlicht: (2024)
Lossless KV Cache Compression to 2%
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
HMoE: Heterogeneous Mixture of Experts for Language Modeling
von: Wang, An, et al.
Veröffentlicht: (2024)
von: Wang, An, et al.
Veröffentlicht: (2024)
Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models
von: Lv, Ang, et al.
Veröffentlicht: (2024)
von: Lv, Ang, et al.
Veröffentlicht: (2024)
PEAR: Position-Embedding-Agnostic Attention Re-weighting Enhances Retrieval-Augmented Generation with Zero Inference Overhead
von: Tan, Tao, et al.
Veröffentlicht: (2024)
von: Tan, Tao, et al.
Veröffentlicht: (2024)
An Analysis and Mitigation of the Reversal Curse
von: Lv, Ang, et al.
Veröffentlicht: (2023)
von: Lv, Ang, et al.
Veröffentlicht: (2023)
Magnifier Prompt: Tackling Multimodal Hallucination via Extremely Simple Instructions
von: Fu, Yuhan, et al.
Veröffentlicht: (2024)
von: Fu, Yuhan, et al.
Veröffentlicht: (2024)
Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use
von: Chen, Yuhan, et al.
Veröffentlicht: (2023)
von: Chen, Yuhan, et al.
Veröffentlicht: (2023)
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
von: Zhang, Kaiyi, et al.
Veröffentlicht: (2025)
von: Zhang, Kaiyi, et al.
Veröffentlicht: (2025)
Scaling Embeddings Outperforms Scaling Experts in Language Models
von: Liu, Hong, et al.
Veröffentlicht: (2026)
von: Liu, Hong, et al.
Veröffentlicht: (2026)
Masked Thought: Simply Masking Partial Reasoning Steps Can Improve Mathematical Reasoning Learning of Language Models
von: Chen, Changyu, et al.
Veröffentlicht: (2024)
von: Chen, Changyu, et al.
Veröffentlicht: (2024)
E-Sparse: Boosting the Large Language Model Inference through Entropy-based N:M Sparsity
von: Li, Yun, et al.
Veröffentlicht: (2023)
von: Li, Yun, et al.
Veröffentlicht: (2023)
dLLM: Simple Diffusion Language Modeling
von: Zhou, Zhanhui, et al.
Veröffentlicht: (2026)
von: Zhou, Zhanhui, et al.
Veröffentlicht: (2026)
Truth Forest: Toward Multi-Scale Truthfulness in Large Language Models through Intervention without Tuning
von: Chen, Zhongzhi, et al.
Veröffentlicht: (2023)
von: Chen, Zhongzhi, et al.
Veröffentlicht: (2023)
Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts
von: He, Shwai, et al.
Veröffentlicht: (2025)
von: He, Shwai, et al.
Veröffentlicht: (2025)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
Mixtures of SubExperts for Large Language Continual Learning
von: Kang, Haeyong
Veröffentlicht: (2025)
von: Kang, Haeyong
Veröffentlicht: (2025)
QAVA: Query-Agnostic Visual Attack to Large Vision-Language Models
von: Zhang, Yudong, et al.
Veröffentlicht: (2025)
von: Zhang, Yudong, et al.
Veröffentlicht: (2025)
Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models
von: Zhou, Zhanhui, et al.
Veröffentlicht: (2024)
von: Zhou, Zhanhui, et al.
Veröffentlicht: (2024)
Exploring Forgetting in Large Language Model Pre-Training
von: Liao, Chonghua, et al.
Veröffentlicht: (2024)
von: Liao, Chonghua, et al.
Veröffentlicht: (2024)
HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and Extrapolation
von: Chen, Yuhan, et al.
Veröffentlicht: (2024)
von: Chen, Yuhan, et al.
Veröffentlicht: (2024)
Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2024)
von: Wang, Zihan, et al.
Veröffentlicht: (2024)
More is not always better? Enhancing Many-Shot In-Context Learning with Differentiated and Reweighting Objectives
von: Zhang, Xiaoqing, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoqing, et al.
Veröffentlicht: (2025)
LaSeR: Reinforcement Learning with Last-Token Self-Rewarding
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
UltraFeedback: Boosting Language Models with Scaled AI Feedback
von: Cui, Ganqu, et al.
Veröffentlicht: (2023)
von: Cui, Ganqu, et al.
Veröffentlicht: (2023)
Multi-Head Mixture-of-Experts
von: Wu, Xun, et al.
Veröffentlicht: (2024)
von: Wu, Xun, et al.
Veröffentlicht: (2024)
MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
von: Xie, Yanyue, et al.
Veröffentlicht: (2024)
von: Xie, Yanyue, et al.
Veröffentlicht: (2024)
PRISM: Parametrically Refactoring Inference for Speculative Sampling Draft Models
von: Wang, Xuliang, et al.
Veröffentlicht: (2026)
von: Wang, Xuliang, et al.
Veröffentlicht: (2026)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
von: Yang, Wenkai, et al.
Veröffentlicht: (2026)
von: Yang, Wenkai, et al.
Veröffentlicht: (2026)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
von: Hui, Tingfeng, et al.
Veröffentlicht: (2024)
von: Hui, Tingfeng, et al.
Veröffentlicht: (2024)
Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!
von: Zhou, Zhanhui, et al.
Veröffentlicht: (2024)
von: Zhou, Zhanhui, et al.
Veröffentlicht: (2024)
Upcycling Large Language Models into Mixture of Experts
von: He, Ethan, et al.
Veröffentlicht: (2024)
von: He, Ethan, et al.
Veröffentlicht: (2024)
HERGC: Heterogeneous Experts Representation and Generative Completion for Multimodal Knowledge Graphs
von: Xiao, Yongkang, et al.
Veröffentlicht: (2025)
von: Xiao, Yongkang, et al.
Veröffentlicht: (2025)
RePO: Replay-Enhanced Policy Optimization
von: Li, Siheng, et al.
Veröffentlicht: (2025)
von: Li, Siheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Language Models "Grok" to Copy
von: Lv, Ang, et al.
Veröffentlicht: (2024) -
More Expressive Attention with Negative Weights
von: Lv, Ang, et al.
Veröffentlicht: (2024) -
Proximal Supervised Fine-Tuning
von: Zhu, Wenhong, et al.
Veröffentlicht: (2025) -
Self-Distillation for Multi-Token Prediction
von: Zhao, Guoliang, et al.
Veröffentlicht: (2026) -
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason
von: Lv, Ang, et al.
Veröffentlicht: (2025)