Guardado en:
| Autores principales: | Sun, Hanchi, Liu, Yixin, Wu, Yonghui, Sun, Lichao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2603.11535 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FinLLM-B: When Large Language Models Meet Financial Breakout Trading
por: Zhang, Kang, et al.
Publicado: (2024)
por: Zhang, Kang, et al.
Publicado: (2024)
MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU
por: Yuan, Zhengqing, et al.
Publicado: (2026)
por: Yuan, Zhengqing, et al.
Publicado: (2026)
Agentic AutoSurvey: Let LLMs Survey LLMs
por: Liu, Yixin, et al.
Publicado: (2025)
por: Liu, Yixin, et al.
Publicado: (2025)
Instruction Mining: Instruction Data Selection for Tuning Large Language Models
por: Cao, Yihan, et al.
Publicado: (2023)
por: Cao, Yihan, et al.
Publicado: (2023)
Self-Cognition in Large Language Models: An Exploratory Study
por: Chen, Dongping, et al.
Publicado: (2024)
por: Chen, Dongping, et al.
Publicado: (2024)
SpecHub: Provable Acceleration to Multi-Draft Speculative Decoding
por: Sun, Ryan, et al.
Publicado: (2024)
por: Sun, Ryan, et al.
Publicado: (2024)
Cosine-Similarity Routing with Semantic Anchors for Interpretable Mixture-of-Experts Language Models
por: Ternovtsii, Ivan, et al.
Publicado: (2025)
por: Ternovtsii, Ivan, et al.
Publicado: (2025)
Beyond Spurious Signals: Debiasing Multimodal Large Language Models via Counterfactual Inference and Adaptive Expert Routing
por: Wu, Zichen, et al.
Publicado: (2025)
por: Wu, Zichen, et al.
Publicado: (2025)
HonestLLM: Toward an Honest and Helpful Large Language Model
por: Gao, Chujie, et al.
Publicado: (2024)
por: Gao, Chujie, et al.
Publicado: (2024)
EfficientLLM: Efficiency in Large Language Models
por: Yuan, Zhengqing, et al.
Publicado: (2025)
por: Yuan, Zhengqing, et al.
Publicado: (2025)
Routing-Aligned Fine-Tuning for Multilingual Downstream Tasks in Mixture-of-Experts Models
por: Deng, Guanzhi, et al.
Publicado: (2026)
por: Deng, Guanzhi, et al.
Publicado: (2026)
Large Language Model Is Not a Good Few-shot Information Extractor, but a Good Reranker for Hard Samples!
por: Ma, Yubo, et al.
Publicado: (2023)
por: Ma, Yubo, et al.
Publicado: (2023)
An Expert is Worth One Token: Synergizing Multiple Expert LLMs as Generalist via Expert Token Routing
por: Chai, Ziwei, et al.
Publicado: (2024)
por: Chai, Ziwei, et al.
Publicado: (2024)
Routing-Free Mixture-of-Experts
por: Liu, Yilun, et al.
Publicado: (2026)
por: Liu, Yilun, et al.
Publicado: (2026)
LD-MoLE: Learnable Dynamic Routing for Mixture of LoRA Experts
por: Zhuang, Yuan, et al.
Publicado: (2025)
por: Zhuang, Yuan, et al.
Publicado: (2025)
ALPINE: Unveiling the Planning Capability of Autoregressive Learning in Language Models
por: Wang, Siwei, et al.
Publicado: (2024)
por: Wang, Siwei, et al.
Publicado: (2024)
RecycleGPT: An Autoregressive Language Model with Recyclable Module
por: Jiang, Yufan, et al.
Publicado: (2023)
por: Jiang, Yufan, et al.
Publicado: (2023)
MoSEs: Uncertainty-Aware AI-Generated Text Detection via Mixture of Stylistics Experts with Conditional Thresholds
por: Wu, Junxi, et al.
Publicado: (2025)
por: Wu, Junxi, et al.
Publicado: (2025)
Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?
por: Fan, Chenrui, et al.
Publicado: (2025)
por: Fan, Chenrui, et al.
Publicado: (2025)
Routing by Analogy: kNN-Augmented Expert Assignment for Mixture-of-Experts
por: Lyu, Boxuan, et al.
Publicado: (2026)
por: Lyu, Boxuan, et al.
Publicado: (2026)
Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination
por: Zheng, Haojie, et al.
Publicado: (2024)
por: Zheng, Haojie, et al.
Publicado: (2024)
SpectR: Dynamically Composing LM Experts with Spectral Routing
por: Fleshman, William, et al.
Publicado: (2025)
por: Fleshman, William, et al.
Publicado: (2025)
Balancing Rigor and Utility: Mitigating Cognitive Biases in Large Language Models for Multiple-Choice Questions
por: Zhong, Hanyang, et al.
Publicado: (2024)
por: Zhong, Hanyang, et al.
Publicado: (2024)
Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model
por: Li, Chong, et al.
Publicado: (2025)
por: Li, Chong, et al.
Publicado: (2025)
Scaling Embeddings Outperforms Scaling Experts in Language Models
por: Liu, Hong, et al.
Publicado: (2026)
por: Liu, Hong, et al.
Publicado: (2026)
Autoregressive + Chain of Thought = Recurrent: Recurrence's Role in Language Models' Computability and a Revisit of Recurrent Transformer
por: Zhang, Xiang, et al.
Publicado: (2024)
por: Zhang, Xiang, et al.
Publicado: (2024)
Multilingual Routing in Mixture-of-Experts
por: Bandarkar, Lucas, et al.
Publicado: (2025)
por: Bandarkar, Lucas, et al.
Publicado: (2025)
Understanding Dynamic Compute Allocation in Recurrent Transformers
por: Moosa, Ibraheem Muhammad, et al.
Publicado: (2026)
por: Moosa, Ibraheem Muhammad, et al.
Publicado: (2026)
SciDFM: A Large Language Model with Mixture-of-Experts for Science
por: Sun, Liangtai, et al.
Publicado: (2024)
por: Sun, Liangtai, et al.
Publicado: (2024)
Probing Semantic Routing in Large Mixture-of-Expert Models
por: Olson, Matthew Lyle, et al.
Publicado: (2025)
por: Olson, Matthew Lyle, et al.
Publicado: (2025)
GPT-SW3: An Autoregressive Language Model for the Nordic Languages
por: Ekgren, Ariel, et al.
Publicado: (2023)
por: Ekgren, Ariel, et al.
Publicado: (2023)
RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty
por: Zhang, Ziqian, et al.
Publicado: (2026)
por: Zhang, Ziqian, et al.
Publicado: (2026)
Autonomy-of-Experts Models
por: Lv, Ang, et al.
Publicado: (2025)
por: Lv, Ang, et al.
Publicado: (2025)
A Study of Large Language Models for Patient Information Extraction: Model Architecture, Fine-Tuning Strategy, and Multi-task Instruction Tuning
por: Peng, Cheng, et al.
Publicado: (2025)
por: Peng, Cheng, et al.
Publicado: (2025)
Countering Catastrophic Forgetting of Large Language Models for Better Instruction Following via Weight-Space Model Merging
por: Lyu, Mengxian, et al.
Publicado: (2026)
por: Lyu, Mengxian, et al.
Publicado: (2026)
BiomedGPT: A Generalist Vision-Language Foundation Model for Diverse Biomedical Tasks
por: Zhang, Kai, et al.
Publicado: (2023)
por: Zhang, Kai, et al.
Publicado: (2023)
Continuous Autoregressive Language Models
por: Shao, Chenze, et al.
Publicado: (2025)
por: Shao, Chenze, et al.
Publicado: (2025)
SynapseRoute: An Auto-Route Switching Framework on Dual-State Large Language Model
por: Zhang, Wencheng, et al.
Publicado: (2025)
por: Zhang, Wencheng, et al.
Publicado: (2025)
MoxE: Mixture of xLSTM Experts with Entropy-Aware Routing for Efficient Language Modeling
por: Thiombiano, Abdoul Majid O., et al.
Publicado: (2025)
por: Thiombiano, Abdoul Majid O., et al.
Publicado: (2025)
Differences in Text Generated by Diffusion and Autoregressive Language Models
por: Zhang, Zeyang, et al.
Publicado: (2026)
por: Zhang, Zeyang, et al.
Publicado: (2026)
Ejemplares similares
-
FinLLM-B: When Large Language Models Meet Financial Breakout Trading
por: Zhang, Kang, et al.
Publicado: (2024) -
MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU
por: Yuan, Zhengqing, et al.
Publicado: (2026) -
Agentic AutoSurvey: Let LLMs Survey LLMs
por: Liu, Yixin, et al.
Publicado: (2025) -
Instruction Mining: Instruction Data Selection for Tuning Large Language Models
por: Cao, Yihan, et al.
Publicado: (2023) -
Self-Cognition in Large Language Models: An Exploratory Study
por: Chen, Dongping, et al.
Publicado: (2024)