Multi-Mixer Models: Flexible Sequence Modeling with Shared Representations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Kevin Y., Trockman, Asher, Suresh, Ananda Theertha, Sun, Ziteng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CafeQ: Calibration-free Quantization via Learned Transformations and Adaptive Rounding
von: Sun, Ziteng, et al.
Veröffentlicht: (2025)
von: Sun, Ziteng, et al.
Veröffentlicht: (2025)
The importance of feature preprocessing for differentially private linear optimization
von: Sun, Ziteng, et al.
Veröffentlicht: (2023)
von: Sun, Ziteng, et al.
Veröffentlicht: (2023)
Asymptotics of Language Model Alignment
von: Yang, Joy Qiping, et al.
Veröffentlicht: (2024)
von: Yang, Joy Qiping, et al.
Veröffentlicht: (2024)
Subset-Based Instance Optimality in Private Estimation
von: Dick, Travis, et al.
Veröffentlicht: (2023)
von: Dick, Travis, et al.
Veröffentlicht: (2023)
CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization
von: Kwon, Soo Min, et al.
Veröffentlicht: (2026)
von: Kwon, Soo Min, et al.
Veröffentlicht: (2026)
Rate of Model Collapse in Recursive Training
von: Suresh, Ananda Theertha, et al.
Veröffentlicht: (2024)
von: Suresh, Ananda Theertha, et al.
Veröffentlicht: (2024)
On Robust Hypothesis Testing with respect to the Hellinger Distance
von: Modak, Eeshan, et al.
Veröffentlicht: (2025)
von: Modak, Eeshan, et al.
Veröffentlicht: (2025)
Coupling without Communication and Drafter-Invariant Speculative Decoding
von: Daliri, Majid, et al.
Veröffentlicht: (2024)
von: Daliri, Majid, et al.
Veröffentlicht: (2024)
SpecTr: Fast Speculative Decoding via Optimal Transport
von: Sun, Ziteng, et al.
Veröffentlicht: (2023)
von: Sun, Ziteng, et al.
Veröffentlicht: (2023)
Mimetic Initialization of MLPs
von: Trockman, Asher, et al.
Veröffentlicht: (2026)
von: Trockman, Asher, et al.
Veröffentlicht: (2026)
Mean estimation in the add-remove model of differential privacy
von: Kulesza, Alex, et al.
Veröffentlicht: (2023)
von: Kulesza, Alex, et al.
Veröffentlicht: (2023)
In-Context Credit Assignment via the Core
von: Harris, Keegan, et al.
Veröffentlicht: (2026)
von: Harris, Keegan, et al.
Veröffentlicht: (2026)
Block Verification Accelerates Speculative Decoding
von: Sun, Ziteng, et al.
Veröffentlicht: (2024)
von: Sun, Ziteng, et al.
Veröffentlicht: (2024)
Mimetic Initialization Helps State Space Models Learn to Recall
von: Trockman, Asher, et al.
Veröffentlicht: (2024)
von: Trockman, Asher, et al.
Veröffentlicht: (2024)
Efficient Language Model Architectures for Differentially Private Federated Learning
von: Ro, Jae Hun, et al.
Veröffentlicht: (2024)
von: Ro, Jae Hun, et al.
Veröffentlicht: (2024)
Multi-Group Fairness Evaluation via Conditional Value-at-Risk Testing
von: Paes, Lucas Monteiro, et al.
Veröffentlicht: (2023)
von: Paes, Lucas Monteiro, et al.
Veröffentlicht: (2023)
Exploring and Improving Drafts in Blockwise Parallel Decoding
von: Kim, Taehyeon, et al.
Veröffentlicht: (2024)
von: Kim, Taehyeon, et al.
Veröffentlicht: (2024)
Hierarchical Retrieval: The Geometry and a Pretrain-Finetune Recipe
von: You, Chong, et al.
Veröffentlicht: (2025)
von: You, Chong, et al.
Veröffentlicht: (2025)
Private federated discovery of out-of-vocabulary words for Gboard
von: Sun, Ziteng, et al.
Veröffentlicht: (2024)
von: Sun, Ziteng, et al.
Veröffentlicht: (2024)
InfAlign: Inference-aware language model alignment
von: Balashankar, Ananth, et al.
Veröffentlicht: (2024)
von: Balashankar, Ananth, et al.
Veröffentlicht: (2024)
Toeplitz MLP Mixers are Low Complexity, Information-Rich Sequence Models
von: Badger, Benjamin L., et al.
Veröffentlicht: (2026)
von: Badger, Benjamin L., et al.
Veröffentlicht: (2026)
Theoretical guarantees on the best-of-n alignment policy
von: Beirami, Ahmad, et al.
Veröffentlicht: (2024)
von: Beirami, Ahmad, et al.
Veröffentlicht: (2024)
Structured Recurrent Mixers for Massively Parallelized Sequence Generation
von: Badger, Benjamin L.
Veröffentlicht: (2026)
von: Badger, Benjamin L.
Veröffentlicht: (2026)
Capability-Aware Shared Hypernetworks for Flexible Heterogeneous Multi-Robot Coordination
von: Fu, Kevin, et al.
Veröffentlicht: (2025)
von: Fu, Kevin, et al.
Veröffentlicht: (2025)
U-Mixer: An Unet-Mixer Architecture with Stationarity Correction for Time Series Forecasting
von: Ma, Xiang, et al.
Veröffentlicht: (2024)
von: Ma, Xiang, et al.
Veröffentlicht: (2024)
tLoRA: Efficient Multi-LoRA Training with Elastic Shared Super-Models
von: Li, Kevin, et al.
Veröffentlicht: (2026)
von: Li, Kevin, et al.
Veröffentlicht: (2026)
Orchid: Flexible and Data-Dependent Convolution for Sequence Modeling
von: Karami, Mahdi, et al.
Veröffentlicht: (2024)
von: Karami, Mahdi, et al.
Veröffentlicht: (2024)
WindowMixer: Intra-Window and Inter-Window Modeling for Time Series Forecasting
von: Liu, Quangao, et al.
Veröffentlicht: (2024)
von: Liu, Quangao, et al.
Veröffentlicht: (2024)
BlockGen: Flexible Blockwise Sequence Modeling with Hybrid Samplers
von: Deschenaux, Justin, et al.
Veröffentlicht: (2026)
von: Deschenaux, Justin, et al.
Veröffentlicht: (2026)
Contextualizing MLP-Mixers Spatiotemporally for Urban Data Forecast at Scale
von: Nie, Tong, et al.
Veröffentlicht: (2023)
von: Nie, Tong, et al.
Veröffentlicht: (2023)
To Pool or Not To Pool: Analyzing the Regularizing Effects of Group-Fair Training on Shared Models
von: Cousins, Cyrus, et al.
Veröffentlicht: (2024)
von: Cousins, Cyrus, et al.
Veröffentlicht: (2024)
JustDense: Just using Dense instead of Sequence Mixer for Time Series analysis
von: Park, TaekHyun, et al.
Veröffentlicht: (2025)
von: Park, TaekHyun, et al.
Veröffentlicht: (2025)
Antidistillation Fingerprinting
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2026)
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2026)
MalMixer: Few-Shot Malware Classification with Retrieval-Augmented Semi-Supervised Learning
von: Li, Jiliang, et al.
Veröffentlicht: (2024)
von: Li, Jiliang, et al.
Veröffentlicht: (2024)
Mamba-3: Improved Sequence Modeling using State Space Principles
von: Lahoti, Aakash, et al.
Veröffentlicht: (2026)
von: Lahoti, Aakash, et al.
Veröffentlicht: (2026)
Fast Jet Tagging with MLP-Mixers on FPGAs
von: Sun, Chang, et al.
Veröffentlicht: (2025)
von: Sun, Chang, et al.
Veröffentlicht: (2025)
A Multi-Scale Decomposition MLP-Mixer for Time Series Analysis
von: Zhong, Shuhan, et al.
Veröffentlicht: (2023)
von: Zhong, Shuhan, et al.
Veröffentlicht: (2023)
Learning Shared Representations for Multi-Task Linear Bandits
von: Lin, Jiabin, et al.
Veröffentlicht: (2026)
von: Lin, Jiabin, et al.
Veröffentlicht: (2026)
MixerFlow: MLP-Mixer meets Normalising Flows
von: English, Eshant, et al.
Veröffentlicht: (2023)
von: English, Eshant, et al.
Veröffentlicht: (2023)
Hydra: Bidirectional State Space Models Through Generalized Matrix Mixers
von: Hwang, Sukjun, et al.
Veröffentlicht: (2024)
von: Hwang, Sukjun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CafeQ: Calibration-free Quantization via Learned Transformations and Adaptive Rounding
von: Sun, Ziteng, et al.
Veröffentlicht: (2025) -
The importance of feature preprocessing for differentially private linear optimization
von: Sun, Ziteng, et al.
Veröffentlicht: (2023) -
Asymptotics of Language Model Alignment
von: Yang, Joy Qiping, et al.
Veröffentlicht: (2024) -
Subset-Based Instance Optimality in Private Estimation
von: Dick, Travis, et al.
Veröffentlicht: (2023) -
CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization
von: Kwon, Soo Min, et al.
Veröffentlicht: (2026)