Structured Recurrent Mixers for Massively Parallelized Sequence Generation
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Badger, Benjamin L. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Masked Mixers for Language Generation and Retrieval
von: Badger, Benjamin L.
Veröffentlicht: (2024)
von: Badger, Benjamin L.
Veröffentlicht: (2024)
Toeplitz MLP Mixers are Low Complexity, Information-Rich Sequence Models
von: Badger, Benjamin L., et al.
Veröffentlicht: (2026)
von: Badger, Benjamin L., et al.
Veröffentlicht: (2026)
Language Model Memory and Memory Models for Language
von: Badger, Benjamin L.
Veröffentlicht: (2026)
von: Badger, Benjamin L.
Veröffentlicht: (2026)
Know Your Limits: Entropy Estimation Modeling for Compression and Generalization
von: Badger, Benjamin L., et al.
Veröffentlicht: (2025)
von: Badger, Benjamin L., et al.
Veröffentlicht: (2025)
Linear Attention Sequence Parallelism
von: Sun, Weigao, et al.
Veröffentlicht: (2024)
von: Sun, Weigao, et al.
Veröffentlicht: (2024)
Cubit: Token Mixer with Kernel Ridge Regression
von: Zheng, Chuanyang, et al.
Veröffentlicht: (2026)
von: Zheng, Chuanyang, et al.
Veröffentlicht: (2026)
Free Energy Mixer
von: Lu, Jiecheng, et al.
Veröffentlicht: (2026)
von: Lu, Jiecheng, et al.
Veröffentlicht: (2026)
Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts
von: Sivtsov, Danil, et al.
Veröffentlicht: (2025)
von: Sivtsov, Danil, et al.
Veröffentlicht: (2025)
FourierNAT: A Fourier-Mixing-Based Non-Autoregressive Transformer for Parallel Sequence Generation
von: Kiruluta, Andrew, et al.
Veröffentlicht: (2025)
von: Kiruluta, Andrew, et al.
Veröffentlicht: (2025)
Efficient Parallel Samplers for Recurrent-Depth Models and Their Connection to Diffusion Language Models
von: Geiping, Jonas, et al.
Veröffentlicht: (2025)
von: Geiping, Jonas, et al.
Veröffentlicht: (2025)
Parallelizing Linear Transformers with the Delta Rule over Sequence Length
von: Yang, Songlin, et al.
Veröffentlicht: (2024)
von: Yang, Songlin, et al.
Veröffentlicht: (2024)
LLM-Mixer: Multiscale Mixing in LLMs for Time Series Forecasting
von: Kowsher, Md, et al.
Veröffentlicht: (2024)
von: Kowsher, Md, et al.
Veröffentlicht: (2024)
State Stream Transformer (SST) V2: Parallel Training of Nonlinear Recurrence for Latent Space Reasoning
von: Aviss, Thea
Veröffentlicht: (2026)
von: Aviss, Thea
Veröffentlicht: (2026)
Gated Associative Memory: A Parallel O(N) Architecture for Efficient Sequence Modeling
von: Acharya, Rishiraj
Veröffentlicht: (2025)
von: Acharya, Rishiraj
Veröffentlicht: (2025)
360-LLaMA-Factory: Plug & Play Sequence Parallelism for Long Post-Training
von: Zou, Haosheng, et al.
Veröffentlicht: (2025)
von: Zou, Haosheng, et al.
Veröffentlicht: (2025)
Gumbel Distillation for Parallel Text Generation
von: Zhang, Chi, et al.
Veröffentlicht: (2026)
von: Zhang, Chi, et al.
Veröffentlicht: (2026)
Associative-State Universal Transformers: Sparse Retrieval Meets Structured Recurrence
von: Xiao, Liu
Veröffentlicht: (2026)
von: Xiao, Liu
Veröffentlicht: (2026)
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
Task Structure Reverses Layerwise State Encoding in Sequence Models
von: Jiang, Yuhang
Veröffentlicht: (2026)
von: Jiang, Yuhang
Veröffentlicht: (2026)
SpikingSSMs: Learning Long Sequences with Sparse and Parallel Spiking State Space Models
von: Shen, Shuaijie, et al.
Veröffentlicht: (2024)
von: Shen, Shuaijie, et al.
Veröffentlicht: (2024)
Multistage Collaborative Knowledge Distillation from a Large Language Model for Semi-Supervised Sequence Generation
von: Zhao, Jiachen, et al.
Veröffentlicht: (2023)
von: Zhao, Jiachen, et al.
Veröffentlicht: (2023)
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
von: Rodionov, Gleb, et al.
Veröffentlicht: (2025)
von: Rodionov, Gleb, et al.
Veröffentlicht: (2025)
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
Dissecting Linear Recurrent Models: How Different Gating Strategies Drive Selectivity and Generalization
von: Bouhadjar, Younes, et al.
Veröffentlicht: (2026)
von: Bouhadjar, Younes, et al.
Veröffentlicht: (2026)
Low-Perplexity LLM-Generated Sequences and Where To Find Them
von: Wuhrmann, Arthur, et al.
Veröffentlicht: (2025)
von: Wuhrmann, Arthur, et al.
Veröffentlicht: (2025)
Insertion Language Models: Sequence Generation with Arbitrary-Position Insertions
von: Patel, Dhruvesh, et al.
Veröffentlicht: (2025)
von: Patel, Dhruvesh, et al.
Veröffentlicht: (2025)
ProPD: Dynamic Token Tree Pruning and Generation for LLM Parallel Decoding
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2024)
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2024)
DLM-One: Diffusion Language Models for One-Step Sequence Generation
von: Chen, Tianqi, et al.
Veröffentlicht: (2025)
von: Chen, Tianqi, et al.
Veröffentlicht: (2025)
HiGen: Hierarchy-Aware Sequence Generation for Hierarchical Text Classification
von: Jain, Vidit, et al.
Veröffentlicht: (2024)
von: Jain, Vidit, et al.
Veröffentlicht: (2024)
Liger: Linearizing Large Language Models to Gated Recurrent Structures
von: Lan, Disen, et al.
Veröffentlicht: (2025)
von: Lan, Disen, et al.
Veröffentlicht: (2025)
Transforming Chatbot Text: A Sequence-to-Sequence Approach
von: Reddy, Natesh, et al.
Veröffentlicht: (2025)
von: Reddy, Natesh, et al.
Veröffentlicht: (2025)
Massive Activations in Large Language Models
von: Sun, Mingjie, et al.
Veröffentlicht: (2024)
von: Sun, Mingjie, et al.
Veröffentlicht: (2024)
TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
Generative Language Models on Nucleotide Sequences of Human Genes
von: Ihtiyar, Musa Nuri, et al.
Veröffentlicht: (2023)
von: Ihtiyar, Musa Nuri, et al.
Veröffentlicht: (2023)
Distil-xLSTM: Learning Attention Mechanisms through Recurrent Structures
von: Thiombiano, Abdoul Majid O., et al.
Veröffentlicht: (2025)
von: Thiombiano, Abdoul Majid O., et al.
Veröffentlicht: (2025)
Parallel Structures in Pre-training Data Yield In-Context Learning
von: Chen, Yanda, et al.
Veröffentlicht: (2024)
von: Chen, Yanda, et al.
Veröffentlicht: (2024)
From Word Sequences to Behavioral Sequences: Adapting Modeling and Evaluation Paradigms for Longitudinal NLP
von: Ganesan, Adithya V, et al.
Veröffentlicht: (2026)
von: Ganesan, Adithya V, et al.
Veröffentlicht: (2026)
On the Representational Capacity of Recurrent Neural Language Models
von: Nowak, Franz, et al.
Veröffentlicht: (2023)
von: Nowak, Franz, et al.
Veröffentlicht: (2023)
Parallel Scaling Law for Language Models
von: Chen, Mouxiang, et al.
Veröffentlicht: (2025)
von: Chen, Mouxiang, et al.
Veröffentlicht: (2025)
Parallel Token Prediction for Language Models
von: Draxler, Felix, et al.
Veröffentlicht: (2025)
von: Draxler, Felix, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Masked Mixers for Language Generation and Retrieval
von: Badger, Benjamin L.
Veröffentlicht: (2024) -
Toeplitz MLP Mixers are Low Complexity, Information-Rich Sequence Models
von: Badger, Benjamin L., et al.
Veröffentlicht: (2026) -
Language Model Memory and Memory Models for Language
von: Badger, Benjamin L.
Veröffentlicht: (2026) -
Know Your Limits: Entropy Estimation Modeling for Compression and Generalization
von: Badger, Benjamin L., et al.
Veröffentlicht: (2025) -
Linear Attention Sequence Parallelism
von: Sun, Weigao, et al.
Veröffentlicht: (2024)