Monet: Mixture of Monosemantic Experts for Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Park, Jungwoo, Ahn, Young Jin, Kim, Kee-Eung, Kang, Jaewoo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization
von: Ahn, Young Jin, et al.
Veröffentlicht: (2024)
von: Ahn, Young Jin, et al.
Veröffentlicht: (2024)
ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
von: Park, Yein, et al.
Veröffentlicht: (2025)
von: Park, Yein, et al.
Veröffentlicht: (2025)
Does Time Have Its Place? Temporal Heads: Where Language Models Recall Time-specific Information
von: Park, Yein, et al.
Veröffentlicht: (2025)
von: Park, Yein, et al.
Veröffentlicht: (2025)
ChroKnowledge: Unveiling Chronological Knowledge of Language Models in Multiple Domains
von: Park, Yein, et al.
Veröffentlicht: (2024)
von: Park, Yein, et al.
Veröffentlicht: (2024)
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
von: Park, Jungwoo, et al.
Veröffentlicht: (2025)
von: Park, Jungwoo, et al.
Veröffentlicht: (2025)
Polysemantic Experts, Monosemantic Paths: Routing as Control in MoEs
von: Ye, Charles, et al.
Veröffentlicht: (2026)
von: Ye, Charles, et al.
Veröffentlicht: (2026)
GDPO: Learning to Directly Align Language Models with Diversity Using GFlowNets
von: Kwon, Oh Joon, et al.
Veröffentlicht: (2024)
von: Kwon, Oh Joon, et al.
Veröffentlicht: (2024)
Stitching Sub-Trajectories with Conditional Diffusion Model for Goal-Conditioned Offline RL
von: Kim, Sungyoon, et al.
Veröffentlicht: (2024)
von: Kim, Sungyoon, et al.
Veröffentlicht: (2024)
Encourage or Inhibit Monosemanticity? Revisit Monosemanticity from a Feature Decorrelation Perspective
von: Yan, Hanqi, et al.
Veröffentlicht: (2024)
von: Yan, Hanqi, et al.
Veröffentlicht: (2024)
Thinking Sparks!: Emergent Attention Heads in Reasoning Models During Post Training
von: Park, Yein, et al.
Veröffentlicht: (2025)
von: Park, Yein, et al.
Veröffentlicht: (2025)
How Many Experts Are Enough? Towards Optimal Semantic Specialization for Mixture-of-Experts
von: Park, Sumin, et al.
Veröffentlicht: (2025)
von: Park, Sumin, et al.
Veröffentlicht: (2025)
Rethinking LLM Inference Bottlenecks: Insights from Latent Attention and Mixture-of-Experts
von: Yun, Sungmin, et al.
Veröffentlicht: (2025)
von: Yun, Sungmin, et al.
Veröffentlicht: (2025)
Tracing Mathematical Proficiency Through Problem-Solving Processes
von: Park, Jungyang, et al.
Veröffentlicht: (2025)
von: Park, Jungyang, et al.
Veröffentlicht: (2025)
Adapting Text-based Dialogue State Tracker for Spoken Dialogues
von: Yoon, Jaeseok, et al.
Veröffentlicht: (2023)
von: Yoon, Jaeseok, et al.
Veröffentlicht: (2023)
ToxReason: A Benchmark for Mechanistic Chemical Toxicity Reasoning via Adverse Outcome Pathway
von: Park, Jueon, et al.
Veröffentlicht: (2026)
von: Park, Jueon, et al.
Veröffentlicht: (2026)
PMoE: Progressive Mixture of Experts with Asymmetric Transformer for Continual Learning
von: Jung, Min Jae, et al.
Veröffentlicht: (2024)
von: Jung, Min Jae, et al.
Veröffentlicht: (2024)
A Monosemantic Attribution Framework for Stable Interpretability in Clinical Neuroscience Transformer-Based Language Models
von: Mamalakis, Michail, et al.
Veröffentlicht: (2026)
von: Mamalakis, Michail, et al.
Veröffentlicht: (2026)
Beyond Interpretability: The Gains of Feature Monosemanticity on Model Robustness
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
von: Templeton, Adly, et al.
Veröffentlicht: (2026)
von: Templeton, Adly, et al.
Veröffentlicht: (2026)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
von: Lim, Junyoung, et al.
Veröffentlicht: (2025)
von: Lim, Junyoung, et al.
Veröffentlicht: (2025)
SCRIPTMIND: Crime Script Inference and Cognitive Evaluation for LLM-based Social Engineering Scam Detection System
von: Kim, Heedou, et al.
Veröffentlicht: (2026)
von: Kim, Heedou, et al.
Veröffentlicht: (2026)
MoDEx: Mixture of Depth-specific Experts for Multivariate Long-term Time Series Forecasting
von: Yoon, Hyekyung, et al.
Veröffentlicht: (2026)
von: Yoon, Hyekyung, et al.
Veröffentlicht: (2026)
MolDeTox: Evaluating Language Model's Stepwise Fragment Editing for Molecular Detoxification
von: Park, Jueon, et al.
Veröffentlicht: (2026)
von: Park, Jueon, et al.
Veröffentlicht: (2026)
On the Spatial Structure of Mixture-of-Experts in Transformers
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2025)
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2025)
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
von: Lee, Haanvid, et al.
Veröffentlicht: (2024)
von: Lee, Haanvid, et al.
Veröffentlicht: (2024)
Detecting Time Series Anomalies Like an Expert: A Multi-Agent LLM Framework with Specialized Analyzers
von: Kang, Hyeongwon, et al.
Veröffentlicht: (2026)
von: Kang, Hyeongwon, et al.
Veröffentlicht: (2026)
Extracting Interaction-Aware Monosemantic Concepts in Recommender Systems
von: Arviv, Dor, et al.
Veröffentlicht: (2025)
von: Arviv, Dor, et al.
Veröffentlicht: (2025)
Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
Diffusion Model Patching via Mixture-of-Prompts
von: Ham, Seokil, et al.
Veröffentlicht: (2024)
von: Ham, Seokil, et al.
Veröffentlicht: (2024)
Mixtures of SubExperts for Large Language Continual Learning
von: Kang, Haeyong
Veröffentlicht: (2025)
von: Kang, Haeyong
Veröffentlicht: (2025)
Dynamic Mixture of Experts Against Severe Distribution Shifts
von: Kim, Donghu
Veröffentlicht: (2025)
von: Kim, Donghu
Veröffentlicht: (2025)
FlipConcept: Tuning-Free Multi-Concept Personalization for Text-to-Image Generation
von: Woo, Young Beom, et al.
Veröffentlicht: (2025)
von: Woo, Young Beom, et al.
Veröffentlicht: (2025)
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
von: Jin, Peng, et al.
Veröffentlicht: (2024)
von: Jin, Peng, et al.
Veröffentlicht: (2024)
Generalization and Scaling Laws for Mixture-of-Experts Transformers
von: Mayaki, Mansour Zoubeirou a
Veröffentlicht: (2026)
von: Mayaki, Mansour Zoubeirou a
Veröffentlicht: (2026)
Monetizing Currency Pair Sentiments through LLM Explainability
von: Limonad, Lior, et al.
Veröffentlicht: (2024)
von: Limonad, Lior, et al.
Veröffentlicht: (2024)
Superposition in Transformers: A Novel Way of Building Mixture of Experts
von: Chaliah, Ayoub Ben, et al.
Veröffentlicht: (2024)
von: Chaliah, Ayoub Ben, et al.
Veröffentlicht: (2024)
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
von: Bambhaniya, Abhimanyu, et al.
Veröffentlicht: (2026)
von: Bambhaniya, Abhimanyu, et al.
Veröffentlicht: (2026)
Geometric Regularization in Mixture-of-Experts: The Disconnect Between Weights and Activations
von: Kim, Hyunjun
Veröffentlicht: (2026)
von: Kim, Hyunjun
Veröffentlicht: (2026)
MoEUT: Mixture-of-Experts Universal Transformers
von: Csordás, Róbert, et al.
Veröffentlicht: (2024)
von: Csordás, Róbert, et al.
Veröffentlicht: (2024)
ECG-MoE: Mixture-of-Expert Electrocardiogram Foundation Model
von: Xu, Yuhao, et al.
Veröffentlicht: (2026)
von: Xu, Yuhao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization
von: Ahn, Young Jin, et al.
Veröffentlicht: (2024) -
ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
von: Park, Yein, et al.
Veröffentlicht: (2025) -
Does Time Have Its Place? Temporal Heads: Where Language Models Recall Time-specific Information
von: Park, Yein, et al.
Veröffentlicht: (2025) -
ChroKnowledge: Unveiling Chronological Knowledge of Language Models in Multiple Domains
von: Park, Yein, et al.
Veröffentlicht: (2024) -
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
von: Park, Jungwoo, et al.
Veröffentlicht: (2025)