Routing in Sparsely-gated Language Models responds to Context
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Arnold, Stefan, Fietta, Marian, Yesilbas, Dilara |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Memorization in Language Models through the Lens of Intrinsic Dimension
von: Arnold, Stefan
Veröffentlicht: (2025)
von: Arnold, Stefan
Veröffentlicht: (2025)
Steering Prepositional Phrases in Language Models: A Case of with-headed Adjectival and Adverbial Complements in Gemma-2
von: Arnold, Stefan, et al.
Veröffentlicht: (2025)
von: Arnold, Stefan, et al.
Veröffentlicht: (2025)
Women Are Beautiful, Men Are Leaders: Gender Stereotypes in Machine Translation and Language Modeling
von: Pikuliak, Matúš, et al.
Veröffentlicht: (2023)
von: Pikuliak, Matúš, et al.
Veröffentlicht: (2023)
Differentially-Private Text Rewriting reshapes Linguistic Style
von: Arnold, Stefan
Veröffentlicht: (2026)
von: Arnold, Stefan
Veröffentlicht: (2026)
Inspecting the Representation Manifold of Differentially-Private Text
von: Arnold, Stefan
Veröffentlicht: (2025)
von: Arnold, Stefan
Veröffentlicht: (2025)
Lookahead Routing for Large Language Models
von: Huang, Canbin, et al.
Veröffentlicht: (2025)
von: Huang, Canbin, et al.
Veröffentlicht: (2025)
SparseD: Sparse Attention for Diffusion Language Models
von: Wang, Zeqing, et al.
Veröffentlicht: (2025)
von: Wang, Zeqing, et al.
Veröffentlicht: (2025)
Documentation Practices of Artificial Intelligence
von: Arnold, Stefan, et al.
Veröffentlicht: (2024)
von: Arnold, Stefan, et al.
Veröffentlicht: (2024)
LPC-SM: Local Predictive Coding and Sparse Memory for Long-Context Language Modeling
von: Xie, Keqin
Veröffentlicht: (2026)
von: Xie, Keqin
Veröffentlicht: (2026)
SPLA: Block Sparse Plus Linear Attention for Long Context Modeling
von: Wang, Bailin, et al.
Veröffentlicht: (2026)
von: Wang, Bailin, et al.
Veröffentlicht: (2026)
Towards Generalizable Implicit In-Context Learning with Attention Routing
von: Li, Jiaqian, et al.
Veröffentlicht: (2025)
von: Li, Jiaqian, et al.
Veröffentlicht: (2025)
Soft Language Prompts for Language Transfer
von: Vykopal, Ivan, et al.
Veröffentlicht: (2024)
von: Vykopal, Ivan, et al.
Veröffentlicht: (2024)
Text-Routed Sparse Mixture-of-Experts Model with Explanation and Temporal Alignment for Multi-Modal Sentiment Analysis
von: Rao, Dongning, et al.
Veröffentlicht: (2025)
von: Rao, Dongning, et al.
Veröffentlicht: (2025)
Towards Enabling FAIR Dataspaces Using Large Language Models
von: Arnold, Benedikt T., et al.
Veröffentlicht: (2024)
von: Arnold, Benedikt T., et al.
Veröffentlicht: (2024)
Route Before Retrieve: Activating Latent Routing Abilities of LLMs for RAG vs. Long-Context Selection
von: Chen, Yiwen, et al.
Veröffentlicht: (2026)
von: Chen, Yiwen, et al.
Veröffentlicht: (2026)
Long-Context Language Modeling with Parallel Context Encoding
von: Yen, Howard, et al.
Veröffentlicht: (2024)
von: Yen, Howard, et al.
Veröffentlicht: (2024)
Generative Large Language Models in Automated Fact-Checking: A Survey
von: Vykopal, Ivan, et al.
Veröffentlicht: (2024)
von: Vykopal, Ivan, et al.
Veröffentlicht: (2024)
Sparse Reward Subsystem in Large Language Models
von: Xu, Guowei, et al.
Veröffentlicht: (2026)
von: Xu, Guowei, et al.
Veröffentlicht: (2026)
Understanding Refusal in Language Models with Sparse Autoencoders
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025)
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025)
Sparse Subnetwork Enhancement for Underrepresented Languages in Large Language Models
von: Gurgurov, Daniil, et al.
Veröffentlicht: (2025)
von: Gurgurov, Daniil, et al.
Veröffentlicht: (2025)
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
von: Piękos, Piotr, et al.
Veröffentlicht: (2025)
von: Piękos, Piotr, et al.
Veröffentlicht: (2025)
Lag-Relative Sparse Attention In Long Context Training
von: Liang, Manlai, et al.
Veröffentlicht: (2025)
von: Liang, Manlai, et al.
Veröffentlicht: (2025)
Building Efficient and Effective OpenQA Systems for Low-Resource Languages
von: Budur, Emrah, et al.
Veröffentlicht: (2024)
von: Budur, Emrah, et al.
Veröffentlicht: (2024)
UniBERT: Adversarial Training for Language-Universal Representations
von: Avram, Andrei-Marius, et al.
Veröffentlicht: (2025)
von: Avram, Andrei-Marius, et al.
Veröffentlicht: (2025)
MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning
von: Shen, Jingyan, et al.
Veröffentlicht: (2025)
von: Shen, Jingyan, et al.
Veröffentlicht: (2025)
On-Policy Context Distillation for Language Models
von: Ye, Tianzhu, et al.
Veröffentlicht: (2026)
von: Ye, Tianzhu, et al.
Veröffentlicht: (2026)
In-Context Watermarks for Large Language Models
von: Liu, Yepeng, et al.
Veröffentlicht: (2025)
von: Liu, Yepeng, et al.
Veröffentlicht: (2025)
Sparse Matrix in Large Language Model Fine-tuning
von: He, Haoze, et al.
Veröffentlicht: (2024)
von: He, Haoze, et al.
Veröffentlicht: (2024)
Long-Context Generalization with Sparse Attention
von: Vasylenko, Pavlo, et al.
Veröffentlicht: (2025)
von: Vasylenko, Pavlo, et al.
Veröffentlicht: (2025)
In-Context Former: Lightning-fast Compressing Context for Large Language Model
von: Wang, Xiangfeng, et al.
Veröffentlicht: (2024)
von: Wang, Xiangfeng, et al.
Veröffentlicht: (2024)
RAGRouter: Learning to Route Queries to Multiple Retrieval-Augmented Language Models
von: Zhang, Jiarui, et al.
Veröffentlicht: (2025)
von: Zhang, Jiarui, et al.
Veröffentlicht: (2025)
$π$-Attention: Periodic Sparse Transformers for Efficient Long-Context Modeling
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
Adamas: Hadamard Sparse Attention for Efficient Long-Context Inference
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
SparseEval: Efficient Evaluation of Large Language Models by Sparse Optimization
von: Zhang, Taolin, et al.
Veröffentlicht: (2026)
von: Zhang, Taolin, et al.
Veröffentlicht: (2026)
Unveiling Language-Specific Features in Large Language Models via Sparse Autoencoders
von: Deng, Boyi, et al.
Veröffentlicht: (2025)
von: Deng, Boyi, et al.
Veröffentlicht: (2025)
NestedKV: Nested Memory Routing for Long-Context KV Cache Compression
von: Chen, Hong, et al.
Veröffentlicht: (2026)
von: Chen, Hong, et al.
Veröffentlicht: (2026)
Routing Absorption in Sparse Attention: Why Random Gates Are Hard to Beat
von: Aquino-Michaels, Keston
Veröffentlicht: (2026)
von: Aquino-Michaels, Keston
Veröffentlicht: (2026)
RULER: What's the Real Context Size of Your Long-Context Language Models?
von: Hsieh, Cheng-Ping, et al.
Veröffentlicht: (2024)
von: Hsieh, Cheng-Ping, et al.
Veröffentlicht: (2024)
Efficiently Computing Susceptibility to Context in Language Models
von: Liu, Tianyu, et al.
Veröffentlicht: (2024)
von: Liu, Tianyu, et al.
Veröffentlicht: (2024)
Can Large Language Models Understand Context?
von: Zhu, Yilun, et al.
Veröffentlicht: (2024)
von: Zhu, Yilun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Memorization in Language Models through the Lens of Intrinsic Dimension
von: Arnold, Stefan
Veröffentlicht: (2025) -
Steering Prepositional Phrases in Language Models: A Case of with-headed Adjectival and Adverbial Complements in Gemma-2
von: Arnold, Stefan, et al.
Veröffentlicht: (2025) -
Women Are Beautiful, Men Are Leaders: Gender Stereotypes in Machine Translation and Language Modeling
von: Pikuliak, Matúš, et al.
Veröffentlicht: (2023) -
Differentially-Private Text Rewriting reshapes Linguistic Style
von: Arnold, Stefan
Veröffentlicht: (2026) -
Inspecting the Representation Manifold of Differentially-Private Text
von: Arnold, Stefan
Veröffentlicht: (2025)