Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM
Fuente:
arXiv
Saved in:
| Main Authors: | Codefuse, Team, Ling, :, Cai, Wenting, Cao, Yuchen, Chen, Chaoyu, Chen, Chen, Chen, Siba, Cui, Qing, Di, Peng, Fang, Junpeng, Gong, Zi, Guo, Ting, He, Zhengyu, Huang, Yang, Li, Cong, Li, Jianguo, Li, Zheng, Lian, Shijie, Liu, BingChang, Luo, Songshan, Mao, Shuo, Shen, Min, Wu, Jian, Yang, Jiaolong, Yang, Wenjie, Ye, Tong, Yu, Hang, Zhang, Wei, Zhang, Zhenduo, Zhao, Hailin, Zheng, Xunjin, Zhou, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Double-Estimation-Friendly Inference for High-Dimensional Measurement Error Models with Non-Sparse Adaptability
by: Cui, Shijie, et al.
Published: (2024)
by: Cui, Shijie, et al.
Published: (2024)
Is Graph Convolution Always Beneficial For Every Feature?
by: Zheng, Yilun, et al.
Published: (2024)
by: Zheng, Yilun, et al.
Published: (2024)
FreeEnricher: Enriching Face Landmarks without Additional Cost
by: Huang, Yangyu, et al.
Published: (2022)
by: Huang, Yangyu, et al.
Published: (2022)
Maternal educational anxiety and mother–child closeness in China: Mediating roles of gatekeeping and father involvement
by: Wenting Wu, et al.
Published: (2026)
by: Wenting Wu, et al.
Published: (2026)
MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMs
by: Chen, Xiaodong, et al.
Published: (2025)
by: Chen, Xiaodong, et al.
Published: (2025)
Non-thermal distributions of charm and charmonium in relativistic heavy-ion collisions
by: Pan, Chaoyu, et al.
Published: (2023)
by: Pan, Chaoyu, et al.
Published: (2023)
Organic Frameworks Memristor: An Emerging Candidate for Data Storage, Artificial Synapse, and Neuromorphic Device
by: Zheng Xu, et al.
Published: (2024)
by: Zheng Xu, et al.
Published: (2024)
Distributed Convoluted Rank Regression for Non-Shareable Data under Non-Additive Losses
by: Zhang, Wen, et al.
Published: (2026)
by: Zhang, Wen, et al.
Published: (2026)
Sparse Rank Regression for Restricted-Access Economic Data
by: Zhang, Wen, et al.
Published: (2026)
by: Zhang, Wen, et al.
Published: (2026)
Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs
by: Ling Team, et al.
Published: (2025)
by: Ling Team, et al.
Published: (2025)
Efficient Partition-based Approaches for Diversified Top-k Subgraph Matching
by: Chen, Liuyi, et al.
Published: (2025)
by: Chen, Liuyi, et al.
Published: (2025)
Multi-Task Learning in Natural Language Processing: An Overview
by: Chen, Shijie, et al.
Published: (2021)
by: Chen, Shijie, et al.
Published: (2021)
iJTyper: An Iterative Type Inference Framework for Java by Integrating Constraint- and Statistically-based Methods
by: Chen, Zhixiang, et al.
Published: (2024)
by: Chen, Zhixiang, et al.
Published: (2024)
A New Medium That Promotes Mass Sporulation of Trichophyton rubrum In Vitro
by: Hailin Zheng, et al.
Published: (2025)
by: Hailin Zheng, et al.
Published: (2025)
Weissella paramesenteroides NRIC1542 inhibits dextran sodium sulfate‐induced colitis in mice through regulating gut microbiota and SIRT1 / NF ‐ κB signaling pathway
by: Shuang Li, et al.
Published: (2024)
by: Shuang Li, et al.
Published: (2024)
CoBa: Convergence Balancer for Multitask Finetuning of Large Language Models
by: Gong, Zi, et al.
Published: (2024)
by: Gong, Zi, et al.
Published: (2024)
Toward a New Understanding of Graphene Oxide Photolysis: The Role of Photoreduction in Degradation Pathway
by: Yuchen Yang, et al.
Published: (2025)
by: Yuchen Yang, et al.
Published: (2025)
Regulation of alternative splicing in the ABA signaling pathway of plants
by: Shijie Ma, et al.
Published: (2025)
by: Shijie Ma, et al.
Published: (2025)
Make Every Draft Count: Hidden State based Speculative Decoding
by: Chen, Yuetao, et al.
Published: (2026)
by: Chen, Yuetao, et al.
Published: (2026)
Co-optimizing Physical Reconfiguration Parameters and Controllers for an Origami-inspired Reconfigurable Manipulator
by: Chen, Zhe, et al.
Published: (2025)
by: Chen, Zhe, et al.
Published: (2025)
TAT-R1: Terminology-Aware Translation with Reinforcement Learning and Word Alignment
by: Li, Zheng, et al.
Published: (2025)
by: Li, Zheng, et al.
Published: (2025)
Few-shot Unknown Class Discovery of Hyperspectral Images with Prototype Learning and Clustering
by: Liu, Chun, et al.
Published: (2025)
by: Liu, Chun, et al.
Published: (2025)
Schwarz-Pick type lemma and Landau type theorem for $α$-harmonic mappings
by: Arora, Vibhuti, et al.
Published: (2025)
by: Arora, Vibhuti, et al.
Published: (2025)
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
by: Yang, Lijie, et al.
Published: (2024)
by: Yang, Lijie, et al.
Published: (2024)
Every Step Evolves: Scaling Reinforcement Learning for Trillion-Scale Thinking Model
by: Ling Team, et al.
Published: (2025)
by: Ling Team, et al.
Published: (2025)
Unveiling Hidden Collaboration within Mixture-of-Experts in Large Language Models
by: Tang, Yuanbo, et al.
Published: (2025)
by: Tang, Yuanbo, et al.
Published: (2025)
Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
by: Ling Team, et al.
Published: (2025)
by: Ling Team, et al.
Published: (2025)
MoPD: Mixture-of-Prompts Distillation for Vision-Language Models
by: Chen, Yang, et al.
Published: (2024)
by: Chen, Yang, et al.
Published: (2024)
Towards Accurate Camouflaged Object Detection with Mixture Convolution and Interactive Fusion
by: Chen, Geng, et al.
Published: (2021)
by: Chen, Geng, et al.
Published: (2021)
Complex semiclassical theory for non-Hermitian quantum systems
by: Yang, Guang, et al.
Published: (2023)
by: Yang, Guang, et al.
Published: (2023)
Robust Inference for High-dimensional Linear Models with Heavy-tailed Errors via Partial Gini Covariance
by: Zhang, Yilin, et al.
Published: (2024)
by: Zhang, Yilin, et al.
Published: (2024)
State-of-the-Art Approaches to Enhancing Privacy Preservation of Machine Learning Datasets: A Survey
by: Zhang, Chaoyu, et al.
Published: (2024)
by: Zhang, Chaoyu, et al.
Published: (2024)
Optimal blood pressure control with fewer antihypertensive medications: Achieved mostly in low‐risk hypertensive patients
by: Zhanyang Zhou, et al.
Published: (2024)
by: Zhanyang Zhou, et al.
Published: (2024)
ProgressVLA: Progress-Guided Diffusion Policy for Vision-Language Robotic Manipulation
by: Yan, Hongyu, et al.
Published: (2026)
by: Yan, Hongyu, et al.
Published: (2026)
An efficient proximal algorithm for squared L1 over L2 regularized sparse recovery
by: Zhang, Na, et al.
Published: (2025)
by: Zhang, Na, et al.
Published: (2025)
IL-7 enhances the protective efficacy of inactivated grass carp reovirus vaccine as adjuvant in grass carp (Ctenopharyngodon idella).
by: Wang, Shijie, et al.
Published: (2026)
by: Wang, Shijie, et al.
Published: (2026)
AP-Loss for Accurate One-Stage Object Detection
by: Chen, Kean, et al.
Published: (2020)
by: Chen, Kean, et al.
Published: (2020)
Unifying the Perspectives of NLP and Software Engineering: A Survey on Language Models for Code
by: Zhang, Ziyin, et al.
Published: (2023)
by: Zhang, Ziyin, et al.
Published: (2023)
Collaborative Inference for Sparse High-Dimensional Models with Non-Shared Data
by: Gu, Yifan, et al.
Published: (2025)
by: Gu, Yifan, et al.
Published: (2025)
CTAGE: Curvature-Based Topology-Aware Graph Embedding for Learning Molecular Representations
by: Chen, Yili, et al.
Published: (2023)
by: Chen, Yili, et al.
Published: (2023)
Similar Items
-
Double-Estimation-Friendly Inference for High-Dimensional Measurement Error Models with Non-Sparse Adaptability
by: Cui, Shijie, et al.
Published: (2024) -
Is Graph Convolution Always Beneficial For Every Feature?
by: Zheng, Yilun, et al.
Published: (2024) -
FreeEnricher: Enriching Face Landmarks without Additional Cost
by: Huang, Yangyu, et al.
Published: (2022) -
Maternal educational anxiety and mother–child closeness in China: Mediating roles of gatekeeping and father involvement
by: Wenting Wu, et al.
Published: (2026) -
MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMs
by: Chen, Xiaodong, et al.
Published: (2025)