N-vium: Mixture-of-Exits Transformer for Accelerated Exact Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Lorenc, Aleksander, Berdoz, Frédéric, Mathys, Joël, Wattenhofer, Roger |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WorldSpeech: A Multilingual Speech Corpus from Around the World
by: Asonitis, Antonis, et al.
Published: (2026)
by: Asonitis, Antonis, et al.
Published: (2026)
Can an AI Agent Safely Run a Government? Existence of Probably Approximately Aligned Policies
by: Berdoz, Frédéric, et al.
Published: (2024)
by: Berdoz, Frédéric, et al.
Published: (2024)
Beyond Interpolation: Extrapolative Reasoning with Reinforcement Learning and Graph Neural Networks
by: Grillo, Niccolò, et al.
Published: (2025)
by: Grillo, Niccolò, et al.
Published: (2025)
Steering Pretrained Drafters during Speculative Decoding
by: Berdoz, Frédéric, et al.
Published: (2025)
by: Berdoz, Frédéric, et al.
Published: (2025)
Can AI Agents Agree?
by: Berdoz, Frédéric, et al.
Published: (2026)
by: Berdoz, Frédéric, et al.
Published: (2026)
Benchmarking Positional Encodings for GNNs and Graph Transformers
by: Grötschla, Florian, et al.
Published: (2024)
by: Grötschla, Florian, et al.
Published: (2024)
Subliminal Signals in Preference Labels
by: Magistrali, Isotta, et al.
Published: (2026)
by: Magistrali, Isotta, et al.
Published: (2026)
Recommender Systems for Democracy: Toward Adversarial Robustness in Voting Advice Applications
by: Berdoz, Frédéric, et al.
Published: (2025)
by: Berdoz, Frédéric, et al.
Published: (2025)
Reasoning Boosts Opinion Alignment in LLMs
by: Berdoz, Frédéric, et al.
Published: (2026)
by: Berdoz, Frédéric, et al.
Published: (2026)
On the Expressive Power of GNNs for Boolean Satisfiability
by: Peltonen, Saku, et al.
Published: (2026)
by: Peltonen, Saku, et al.
Published: (2026)
Alignment-Aware Decoding
by: Berdoz, Frédéric, et al.
Published: (2025)
by: Berdoz, Frédéric, et al.
Published: (2025)
GraphFSA: A Finite State Automaton Framework for Algorithmic Learning on Graphs
by: Grötschla, Florian, et al.
Published: (2024)
by: Grötschla, Florian, et al.
Published: (2024)
Flood and Echo Net: Algorithmically Aligned GNNs that Generalize
by: Mathys, Joël, et al.
Published: (2023)
by: Mathys, Joël, et al.
Published: (2023)
Next Level Message-Passing with Hierarchical Support Graphs
by: Vonessen, Carlos, et al.
Published: (2024)
by: Vonessen, Carlos, et al.
Published: (2024)
SpecExit: Accelerating Large Reasoning Model via Speculative Exit
by: Yang, Rubing, et al.
Published: (2025)
by: Yang, Rubing, et al.
Published: (2025)
From Message-Passing to Linearized Graph Sequence Models
by: Mathys, Joël, et al.
Published: (2026)
by: Mathys, Joël, et al.
Published: (2026)
Scalable Evaluation and Neural Models for Compositional Generalization
by: Camposampiero, Giacomo, et al.
Published: (2025)
by: Camposampiero, Giacomo, et al.
Published: (2025)
High-Fidelity Speech Enhancement via Discrete Audio Tokens
by: Lanzendörfer, Luca A., et al.
Published: (2025)
by: Lanzendörfer, Luca A., et al.
Published: (2025)
Text-to-Scene with Large Reasoning Models
by: Berdoz, Frédéric, et al.
Published: (2025)
by: Berdoz, Frédéric, et al.
Published: (2025)
I-RAVEN-X: Benchmarking Generalization and Robustness of Analogical and Mathematical Reasoning in Large Language and Reasoning Models
by: Camposampiero, Giacomo, et al.
Published: (2025)
by: Camposampiero, Giacomo, et al.
Published: (2025)
CoRe-GD: A Hierarchical Framework for Scalable Graph Visualization with GNNs
by: Grötschla, Florian, et al.
Published: (2024)
by: Grötschla, Florian, et al.
Published: (2024)
Parametric Neural Amp Modeling with Active Learning
by: Grötschla, Florian, et al.
Published: (2025)
by: Grötschla, Florian, et al.
Published: (2025)
PUZZLES: A Benchmark for Neural Algorithmic Reasoning
by: Estermann, Benjamin, et al.
Published: (2024)
by: Estermann, Benjamin, et al.
Published: (2024)
Beyond Greedy Exits: Improved Early Exit Decisions for Risk Control and Reliability
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
Position Paper: Rethinking Privacy in RL for Sequential Decision-making in the Age of LLMs
by: Fan, Flint Xiaofeng, et al.
Published: (2025)
by: Fan, Flint Xiaofeng, et al.
Published: (2025)
Can Large Reasoning Models do Analogical Reasoning under Perceptual Uncertainty?
by: Camposampiero, Giacomo, et al.
Published: (2025)
by: Camposampiero, Giacomo, et al.
Published: (2025)
MixGCN: Scalable GCN Training by Mixture of Parallelism and Mixture of Accelerators
by: Wan, Cheng, et al.
Published: (2025)
by: Wan, Cheng, et al.
Published: (2025)
FLIP Reasoning Challenge
by: Plesner, Andreas, et al.
Published: (2025)
by: Plesner, Andreas, et al.
Published: (2025)
AEBNAS: Strengthening Exit Branches in Early-Exit Networks through Hardware-Aware Neural Architecture Search
by: Robben, Oscar, et al.
Published: (2025)
by: Robben, Oscar, et al.
Published: (2025)
Speculating Experts Accelerates Inference for Mixture-of-Experts
by: Madan, Vivan, et al.
Published: (2026)
by: Madan, Vivan, et al.
Published: (2026)
EEG-Bench: A Benchmark for EEG Foundation Models in Clinical Applications
by: Kastrati, Ard, et al.
Published: (2025)
by: Kastrati, Ard, et al.
Published: (2025)
One Jump Is All You Need: Short-Cutting Transformers for Early Exit Prediction with One Jump to Fit All Exit Levels
by: Seshadri, Amrit Diggavi
Published: (2025)
by: Seshadri, Amrit Diggavi
Published: (2025)
LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference
by: Kapadia, Shashank, et al.
Published: (2026)
by: Kapadia, Shashank, et al.
Published: (2026)
Are Transformers More Robust? Towards Exact Robustness Verification for Transformers
by: Liao, Brian Hsuan-Cheng, et al.
Published: (2022)
by: Liao, Brian Hsuan-Cheng, et al.
Published: (2022)
Provably Powerful Graph Neural Networks for Directed Multigraphs
by: Egressy, Béni, et al.
Published: (2023)
by: Egressy, Béni, et al.
Published: (2023)
Generalization and Scaling Laws for Mixture-of-Experts Transformers
by: Mayaki, Mansour Zoubeirou a
Published: (2026)
by: Mayaki, Mansour Zoubeirou a
Published: (2026)
Exact Attention Sensitivity and the Geometry of Transformer Stability
by: Emadi, Seyed Morteza
Published: (2026)
by: Emadi, Seyed Morteza
Published: (2026)
Early-Exit Neural Networks with Nested Prediction Sets
by: Jazbec, Metod, et al.
Published: (2023)
by: Jazbec, Metod, et al.
Published: (2023)
Towards Learning to Reason: Comparing LLMs with Neuro-Symbolic on Arithmetic Relations in Abstract Reasoning
by: Hersche, Michael, et al.
Published: (2024)
by: Hersche, Michael, et al.
Published: (2024)
Graph Dimension Attention Networks for Enterprise Credit Assessment
by: Wei, Shaopeng, et al.
Published: (2024)
by: Wei, Shaopeng, et al.
Published: (2024)
Similar Items
-
WorldSpeech: A Multilingual Speech Corpus from Around the World
by: Asonitis, Antonis, et al.
Published: (2026) -
Can an AI Agent Safely Run a Government? Existence of Probably Approximately Aligned Policies
by: Berdoz, Frédéric, et al.
Published: (2024) -
Beyond Interpolation: Extrapolative Reasoning with Reinforcement Learning and Graph Neural Networks
by: Grillo, Niccolò, et al.
Published: (2025) -
Steering Pretrained Drafters during Speculative Decoding
by: Berdoz, Frédéric, et al.
Published: (2025) -
Can AI Agents Agree?
by: Berdoz, Frédéric, et al.
Published: (2026)