On-Device Collaborative Language Modeling via a Mixture of Generalists and Specialists
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fan, Dongyang, Messmer, Bettina, Doikov, Nikita, Jaggi, Martin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards an empirical understanding of MoE design choices
von: Fan, Dongyang, et al.
Veröffentlicht: (2024)
von: Fan, Dongyang, et al.
Veröffentlicht: (2024)
Personalized Collaborative Fine-Tuning for On-Device Large Language Models
von: Wagner, Nicolas, et al.
Veröffentlicht: (2024)
von: Wagner, Nicolas, et al.
Veröffentlicht: (2024)
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
von: Messmer, Bettina, et al.
Veröffentlicht: (2025)
von: Messmer, Bettina, et al.
Veröffentlicht: (2025)
Gradient-Normalized Smoothness for Optimization with Approximate Hessians
von: Semenov, Andrei, et al.
Veröffentlicht: (2025)
von: Semenov, Andrei, et al.
Veröffentlicht: (2025)
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
von: Fan, Dongyang, et al.
Veröffentlicht: (2025)
von: Fan, Dongyang, et al.
Veröffentlicht: (2025)
Analyzing & Reducing the Need for Learning Rate Warmup in GPT Training
von: Kosson, Atli, et al.
Veröffentlicht: (2024)
von: Kosson, Atli, et al.
Veröffentlicht: (2024)
Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks
von: Kosson, Atli, et al.
Veröffentlicht: (2023)
von: Kosson, Atli, et al.
Veröffentlicht: (2023)
Improving Stochastic Cubic Newton with Momentum
von: Chayti, El Mahdi, et al.
Veröffentlicht: (2024)
von: Chayti, El Mahdi, et al.
Veröffentlicht: (2024)
Unified Convergence Theory of Stochastic and Variance-Reduced Cubic Newton Methods
von: Chayti, El Mahdi, et al.
Veröffentlicht: (2023)
von: Chayti, El Mahdi, et al.
Veröffentlicht: (2023)
TiMoE: Time-Aware Mixture of Language Experts
von: Faro, Robin, et al.
Veröffentlicht: (2025)
von: Faro, Robin, et al.
Veröffentlicht: (2025)
From Generalist to Specialist: A Survey of Large Language Models for Chemistry
von: Han, Yang, et al.
Veröffentlicht: (2024)
von: Han, Yang, et al.
Veröffentlicht: (2024)
Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
von: Fan, Dongyang, et al.
Veröffentlicht: (2025)
von: Fan, Dongyang, et al.
Veröffentlicht: (2025)
On Convergence of Incremental Gradient for Non-Convex Smooth Functions
von: Koloskova, Anastasia, et al.
Veröffentlicht: (2023)
von: Koloskova, Anastasia, et al.
Veröffentlicht: (2023)
Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection
von: Turki, Yassine, et al.
Veröffentlicht: (2026)
von: Turki, Yassine, et al.
Veröffentlicht: (2026)
DoGE: Domain Reweighting with Generalization Estimation
von: Fan, Simin, et al.
Veröffentlicht: (2023)
von: Fan, Simin, et al.
Veröffentlicht: (2023)
DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging
von: Pagliardini, Matteo, et al.
Veröffentlicht: (2024)
von: Pagliardini, Matteo, et al.
Veröffentlicht: (2024)
CoTFormer: A Chain-of-Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference
von: Mohtashami, Amirkeivan, et al.
Veröffentlicht: (2023)
von: Mohtashami, Amirkeivan, et al.
Veröffentlicht: (2023)
A Survey on Mixture of Experts in Large Language Models
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
URLs Help, Topics Guide: Understanding Metadata Utility in LLM Training
von: Fan, Dongyang, et al.
Veröffentlicht: (2025)
von: Fan, Dongyang, et al.
Veröffentlicht: (2025)
Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents
von: Agashe, Saaket, et al.
Veröffentlicht: (2025)
von: Agashe, Saaket, et al.
Veröffentlicht: (2025)
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
von: Ablin, Pierre, et al.
Veröffentlicht: (2025)
von: Ablin, Pierre, et al.
Veröffentlicht: (2025)
Table-LLM-Specialist: Language Model Specialists for Tables using Iterative Generator-Validator Fine-tuning
von: Xing, Junjie, et al.
Veröffentlicht: (2024)
von: Xing, Junjie, et al.
Veröffentlicht: (2024)
Fast Large Language Model Collaborative Decoding via Speculation
von: Fu, Jiale, et al.
Veröffentlicht: (2025)
von: Fu, Jiale, et al.
Veröffentlicht: (2025)
No Need to Talk: Asynchronous Mixture of Language Models
von: Filippova, Anastasiia, et al.
Veröffentlicht: (2024)
von: Filippova, Anastasiia, et al.
Veröffentlicht: (2024)
Bayesian Mixture of Experts For Large Language Models
von: Dialameh, Maryam, et al.
Veröffentlicht: (2025)
von: Dialameh, Maryam, et al.
Veröffentlicht: (2025)
Mixture-of-Personas Language Models for Population Simulation
von: Bui, Ngoc, et al.
Veröffentlicht: (2025)
von: Bui, Ngoc, et al.
Veröffentlicht: (2025)
Not All Denoising Steps Are Equal: Model Scheduling for Faster Masked Diffusion Language Models
von: Sedykh, Ivan, et al.
Veröffentlicht: (2026)
von: Sedykh, Ivan, et al.
Veröffentlicht: (2026)
BTS: Harmonizing Specialized Experts into a Generalist LLM
von: Zhang, Qizhen, et al.
Veröffentlicht: (2025)
von: Zhang, Qizhen, et al.
Veröffentlicht: (2025)
Upcycling Large Language Models into Mixture of Experts
von: He, Ethan, et al.
Veröffentlicht: (2024)
von: He, Ethan, et al.
Veröffentlicht: (2024)
HMoE: Heterogeneous Mixture of Experts for Language Modeling
von: Wang, An, et al.
Veröffentlicht: (2024)
von: Wang, An, et al.
Veröffentlicht: (2024)
Scaling Laws for Upcycling Mixture-of-Experts Language Models
von: Liew, Seng Pei, et al.
Veröffentlicht: (2025)
von: Liew, Seng Pei, et al.
Veröffentlicht: (2025)
Leveraging the true depth of LLMs
von: González, Ramón Calvo, et al.
Veröffentlicht: (2025)
von: González, Ramón Calvo, et al.
Veröffentlicht: (2025)
Inference-Time Scaling for Generalist Reward Modeling
von: Liu, Zijun, et al.
Veröffentlicht: (2025)
von: Liu, Zijun, et al.
Veröffentlicht: (2025)
DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices
von: Song, Chenyang, et al.
Veröffentlicht: (2026)
von: Song, Chenyang, et al.
Veröffentlicht: (2026)
A Closer Look into Mixture-of-Experts in Large Language Models
von: Lo, Ka Man, et al.
Veröffentlicht: (2024)
von: Lo, Ka Man, et al.
Veröffentlicht: (2024)
Optimal Splitting of Language Models from Mixtures to Specialized Domains
von: Seto, Skyler, et al.
Veröffentlicht: (2026)
von: Seto, Skyler, et al.
Veröffentlicht: (2026)
Analyze Feature Flow to Enhance Interpretation and Steering in Language Models
von: Laptev, Daniil, et al.
Veröffentlicht: (2025)
von: Laptev, Daniil, et al.
Veröffentlicht: (2025)
Empowering Many, Biasing a Few: Generalist Credit Scoring through Large Language Models
von: Feng, Duanyu, et al.
Veröffentlicht: (2023)
von: Feng, Duanyu, et al.
Veröffentlicht: (2023)
CoBo: Collaborative Learning via Bilevel Optimization
von: Hashemi, Diba, et al.
Veröffentlicht: (2024)
von: Hashemi, Diba, et al.
Veröffentlicht: (2024)
Learning to Decode Collaboratively with Multiple Language Models
von: Shen, Shannon Zejiang, et al.
Veröffentlicht: (2024)
von: Shen, Shannon Zejiang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Towards an empirical understanding of MoE design choices
von: Fan, Dongyang, et al.
Veröffentlicht: (2024) -
Personalized Collaborative Fine-Tuning for On-Device Large Language Models
von: Wagner, Nicolas, et al.
Veröffentlicht: (2024) -
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
von: Messmer, Bettina, et al.
Veröffentlicht: (2025) -
Gradient-Normalized Smoothness for Optimization with Approximate Hessians
von: Semenov, Andrei, et al.
Veröffentlicht: (2025) -
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
von: Fan, Dongyang, et al.
Veröffentlicht: (2025)