Guardado en:
| Autores principales: | Wang, Weixuan, Wu, Minghao, Haddow, Barry, Birch, Alexandra |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2505.12313 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Demystifying Multilingual Chain-of-Thought in Process Reward Modeling
por: Wang, Weixuan, et al.
Publicado: (2025)
por: Wang, Weixuan, et al.
Publicado: (2025)
HBO: Hierarchical Balancing Optimization for Fine-Tuning Large Language Models
por: Wang, Weixuan, et al.
Publicado: (2025)
por: Wang, Weixuan, et al.
Publicado: (2025)
Learning to Summarize by Learning to Quiz: Adversarial Agentic Collaboration for Long Document Summarization
por: Wang, Weixuan, et al.
Publicado: (2025)
por: Wang, Weixuan, et al.
Publicado: (2025)
Bridging the Language Gaps in Large Language Models with Inference-Time Cross-Lingual Intervention
por: Wang, Weixuan, et al.
Publicado: (2024)
por: Wang, Weixuan, et al.
Publicado: (2024)
Sharing Matters: Analysing Neurons Across Languages and Tasks in LLMs
por: Wang, Weixuan, et al.
Publicado: (2024)
por: Wang, Weixuan, et al.
Publicado: (2024)
Multilingual Retrieval-Augmented Generation for Knowledge-Intensive Task
por: Ranaldi, Leonardo, et al.
Publicado: (2025)
por: Ranaldi, Leonardo, et al.
Publicado: (2025)
MGen: Millions of Naturally Occurring Generics in Context
por: Cilleruelo, Gustavo, et al.
Publicado: (2025)
por: Cilleruelo, Gustavo, et al.
Publicado: (2025)
When Does Monolingual Data Help Multilingual Translation: The Role of Domain and Model Scale
por: Baziotis, Christos, et al.
Publicado: (2023)
por: Baziotis, Christos, et al.
Publicado: (2023)
The Ups and Downs of Large Language Model Inference with Vocabulary Trimming by Language Heuristics
por: Bogoychev, Nikolay, et al.
Publicado: (2023)
por: Bogoychev, Nikolay, et al.
Publicado: (2023)
Compact Speech Translation Models via Discrete Speech Units Pretraining
por: Lam, Tsz Kin, et al.
Publicado: (2024)
por: Lam, Tsz Kin, et al.
Publicado: (2024)
Improving Multilingual Retrieval-Augmented Language Models through Dialectic Reasoning Argumentations
por: Ranaldi, Leonardo, et al.
Publicado: (2025)
por: Ranaldi, Leonardo, et al.
Publicado: (2025)
The Prosody of Emojis
por: Zhou, Giulio, et al.
Publicado: (2025)
por: Zhou, Giulio, et al.
Publicado: (2025)
Prosody in Cascade and Direct Speech-to-Text Translation: a case study on Korean Wh-Phrases
por: Zhou, Giulio, et al.
Publicado: (2024)
por: Zhou, Giulio, et al.
Publicado: (2024)
Liaozhai through the Looking-Glass: On Paratextual Explicitation of Culture-Bound Terms in Machine Translation
por: Shen, Sherrie, et al.
Publicado: (2025)
por: Shen, Sherrie, et al.
Publicado: (2025)
Generics are puzzling. Can language models find the missing piece?
por: Calderón, Gustavo Cilleruelo, et al.
Publicado: (2024)
por: Calderón, Gustavo Cilleruelo, et al.
Publicado: (2024)
Quality or Quantity? On Data Scale and Diversity in Adapting Large Language Models for Low-Resource Translation
por: Iyer, Vivek, et al.
Publicado: (2024)
por: Iyer, Vivek, et al.
Publicado: (2024)
Understanding Multilingualism in Mixture-of-Experts LLMs: Routing Mechanism, Expert Specialization, and Layerwise Steering
por: Chen, Yuxin, et al.
Publicado: (2026)
por: Chen, Yuxin, et al.
Publicado: (2026)
Understanding and Leveraging the Expert Specialization of Context Faithfulness in Mixture-of-Experts LLMs
por: Bai, Jun, et al.
Publicado: (2025)
por: Bai, Jun, et al.
Publicado: (2025)
Semantics-Adaptive Activation Intervention for LLMs via Dynamic Steering Vectors
por: Wang, Weixuan, et al.
Publicado: (2024)
por: Wang, Weixuan, et al.
Publicado: (2024)
Steering MoE LLMs via Expert (De)Activation
por: Fayyaz, Mohsen, et al.
Publicado: (2025)
por: Fayyaz, Mohsen, et al.
Publicado: (2025)
Is Modularity Transferable? A Case Study through the Lens of Knowledge Distillation
por: Klimaszewski, Mateusz, et al.
Publicado: (2024)
por: Klimaszewski, Mateusz, et al.
Publicado: (2024)
From Beginner to Expert: Modeling Medical Knowledge into General LLMs
por: Li, Qiang, et al.
Publicado: (2023)
por: Li, Qiang, et al.
Publicado: (2023)
An Expert is Worth One Token: Synergizing Multiple Expert LLMs as Generalist via Expert Token Routing
por: Chai, Ziwei, et al.
Publicado: (2024)
por: Chai, Ziwei, et al.
Publicado: (2024)
Context and System Fusion in Post-ASR Emotion Recognition with Large Language Models
por: Stepachev, Pavel, et al.
Publicado: (2024)
por: Stepachev, Pavel, et al.
Publicado: (2024)
Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs
por: Zhou, Yixiao, et al.
Publicado: (2025)
por: Zhou, Yixiao, et al.
Publicado: (2025)
Integrating Expert Knowledge into Logical Programs via LLMs
por: Górski, Franciszek, et al.
Publicado: (2025)
por: Górski, Franciszek, et al.
Publicado: (2025)
LF-Steering: Latent Feature Activation Steering for Enhancing Semantic Consistency in Large Language Models
por: Yang, Jingyuan, et al.
Publicado: (2025)
por: Yang, Jingyuan, et al.
Publicado: (2025)
Is It Good Data for Multilingual Instruction Tuning or Just Bad Multilingual Evaluation for Large Language Models?
por: Chen, Pinzhen, et al.
Publicado: (2024)
por: Chen, Pinzhen, et al.
Publicado: (2024)
Iterative Translation Refinement with Large Language Models
por: Chen, Pinzhen, et al.
Publicado: (2023)
por: Chen, Pinzhen, et al.
Publicado: (2023)
Mixture of Experts for Low-Resource LLMs
por: Joseph, Ori Bar, et al.
Publicado: (2026)
por: Joseph, Ori Bar, et al.
Publicado: (2026)
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
por: Zhang, Zeliang, et al.
Publicado: (2024)
por: Zhang, Zeliang, et al.
Publicado: (2024)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
por: Zhuang, Haomin, et al.
Publicado: (2024)
por: Zhuang, Haomin, et al.
Publicado: (2024)
DiEP: Adaptive Mixture-of-Experts Compression through Differentiable Expert Pruning
por: Bai, Sikai, et al.
Publicado: (2025)
por: Bai, Sikai, et al.
Publicado: (2025)
Continuously Steering LLMs Sensitivity to Contextual Knowledge with Proxy Models
por: Wang, Yilin, et al.
Publicado: (2025)
por: Wang, Yilin, et al.
Publicado: (2025)
Test-Time Steering for Lossless Text Compression via Weighted Product of Experts
por: Zhang, Qihang, et al.
Publicado: (2025)
por: Zhang, Qihang, et al.
Publicado: (2025)
Knowledge Localization in Mixture-of-Experts LLMs Using Cross-Lingual Inconsistency
por: Bandarkar, Lucas, et al.
Publicado: (2026)
por: Bandarkar, Lucas, et al.
Publicado: (2026)
MatheMagic: Generating Dynamic Mathematics Benchmarks Robust to Memorization
por: O'Brien, Dayyán, et al.
Publicado: (2025)
por: O'Brien, Dayyán, et al.
Publicado: (2025)
Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM
por: Sukhbaatar, Sainbayar, et al.
Publicado: (2024)
por: Sukhbaatar, Sainbayar, et al.
Publicado: (2024)
dMoE: dLLMs with Learnable Block Experts
por: Feng, Sicheng, et al.
Publicado: (2026)
por: Feng, Sicheng, et al.
Publicado: (2026)
CartesianMoE: Boosting Knowledge Sharing among Experts via Cartesian Product Routing in Mixture-of-Experts
por: Su, Zhenpeng, et al.
Publicado: (2024)
por: Su, Zhenpeng, et al.
Publicado: (2024)
Ejemplares similares
-
Demystifying Multilingual Chain-of-Thought in Process Reward Modeling
por: Wang, Weixuan, et al.
Publicado: (2025) -
HBO: Hierarchical Balancing Optimization for Fine-Tuning Large Language Models
por: Wang, Weixuan, et al.
Publicado: (2025) -
Learning to Summarize by Learning to Quiz: Adversarial Agentic Collaboration for Long Document Summarization
por: Wang, Weixuan, et al.
Publicado: (2025) -
Bridging the Language Gaps in Large Language Models with Inference-Time Cross-Lingual Intervention
por: Wang, Weixuan, et al.
Publicado: (2024) -
Sharing Matters: Analysing Neurons Across Languages and Tasks in LLMs
por: Wang, Weixuan, et al.
Publicado: (2024)