Transformers are Multi-State RNNs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Oren, Matanel, Hassid, Michael, Yarden, Nir, Adi, Yossi, Schwartz, Roy |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning
von: Hassid, Michael, et al.
Veröffentlicht: (2025)
von: Hassid, Michael, et al.
Veröffentlicht: (2025)
The Larger the Better? Improved LLM Code-Generation via Budget Reallocation
von: Hassid, Michael, et al.
Veröffentlicht: (2024)
von: Hassid, Michael, et al.
Veröffentlicht: (2024)
On Pruning State-Space LLMs
von: Ghattas, Tamer, et al.
Veröffentlicht: (2025)
von: Ghattas, Tamer, et al.
Veröffentlicht: (2025)
From Tokens to Words: On the Inner Lexicon of LLMs
von: Kaplan, Guy, et al.
Veröffentlicht: (2024)
von: Kaplan, Guy, et al.
Veröffentlicht: (2024)
Scaling Analysis of Interleaved Speech-Text Language Models
von: Maimon, Gallil, et al.
Veröffentlicht: (2025)
von: Maimon, Gallil, et al.
Veröffentlicht: (2025)
Self-Execution Simulation Improves Coding Models
von: Maimon, Gallil, et al.
Veröffentlicht: (2026)
von: Maimon, Gallil, et al.
Veröffentlicht: (2026)
LaMI: Augmenting Large Language Models via Late Multi-Image Fusion
von: Yariv, Guy, et al.
Veröffentlicht: (2024)
von: Yariv, Guy, et al.
Veröffentlicht: (2024)
More Documents, Same Length: Isolating the Challenge of Multiple Documents in RAG
von: Levy, Shahar, et al.
Veröffentlicht: (2025)
von: Levy, Shahar, et al.
Veröffentlicht: (2025)
Textually Pretrained Speech Language Models
von: Hassid, Michael, et al.
Veröffentlicht: (2023)
von: Hassid, Michael, et al.
Veröffentlicht: (2023)
LAST: Language Model Aware Speech Tokenization
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
GmSLM : Generative Marmoset Spoken Language Modeling
von: Sternberg, Talia, et al.
Veröffentlicht: (2025)
von: Sternberg, Talia, et al.
Veröffentlicht: (2025)
HGRN2: Gated Linear RNNs with State Expansion
von: Qin, Zhen, et al.
Veröffentlicht: (2024)
von: Qin, Zhen, et al.
Veröffentlicht: (2024)
Does Transformer Interpretability Transfer to RNNs?
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024)
LLMs versus the Halting Problem: Characterizing Program Termination Reasoning
von: Sultan, Oren, et al.
Veröffentlicht: (2026)
von: Sultan, Oren, et al.
Veröffentlicht: (2026)
RNNs are not Transformers (Yet): The Key Bottleneck on In-context Retrieval
von: Wen, Kaiyue, et al.
Veröffentlicht: (2024)
von: Wen, Kaiyue, et al.
Veröffentlicht: (2024)
Compositional Reasoning with Transformers, RNNs, and Chain of Thought
von: Yehudai, Gilad, et al.
Veröffentlicht: (2025)
von: Yehudai, Gilad, et al.
Veröffentlicht: (2025)
Salmon: A Suite for Acoustic Language Model Evaluation
von: Maimon, Gallil, et al.
Veröffentlicht: (2024)
von: Maimon, Gallil, et al.
Veröffentlicht: (2024)
A Language Modeling Approach to Diacritic-Free Hebrew TTS
von: Roth, Amit, et al.
Veröffentlicht: (2024)
von: Roth, Amit, et al.
Veröffentlicht: (2024)
WHISTRESS: Enriching Transcriptions with Sentence Stress Detection
von: Yosha, Iddo, et al.
Veröffentlicht: (2025)
von: Yosha, Iddo, et al.
Veröffentlicht: (2025)
StressTest: Can YOUR Speech LM Handle the Stress?
von: Yosha, Iddo, et al.
Veröffentlicht: (2025)
von: Yosha, Iddo, et al.
Veröffentlicht: (2025)
Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues
von: Grazzi, Riccardo, et al.
Veröffentlicht: (2024)
von: Grazzi, Riccardo, et al.
Veröffentlicht: (2024)
SpeLLM: Character-Level Multi-Head Decoding
von: Ben-Artzy, Amit, et al.
Veröffentlicht: (2025)
von: Ben-Artzy, Amit, et al.
Veröffentlicht: (2025)
Learning State-Tracking from Code Using Linear RNNs
von: Siems, Julien, et al.
Veröffentlicht: (2026)
von: Siems, Julien, et al.
Veröffentlicht: (2026)
PAST: Phonetic-Acoustic Speech Tokenizer
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
Why Are Linear RNNs More Parallelizable?
von: Merrill, William, et al.
Veröffentlicht: (2026)
von: Merrill, William, et al.
Veröffentlicht: (2026)
DeltaProduct: Improving State-Tracking in Linear RNNs via Householder Products
von: Siems, Julien, et al.
Veröffentlicht: (2025)
von: Siems, Julien, et al.
Veröffentlicht: (2025)
Unsupervised Speech Segmentation: A General Approach Using Speech Language Models
von: Elmakies, Avishai, et al.
Veröffentlicht: (2025)
von: Elmakies, Avishai, et al.
Veröffentlicht: (2025)
Slamming: Training a Speech Language Model on One GPU in a Day
von: Maimon, Gallil, et al.
Veröffentlicht: (2025)
von: Maimon, Gallil, et al.
Veröffentlicht: (2025)
Learning to (Learn at Test Time): RNNs with Expressive Hidden States
von: Sun, Yu, et al.
Veröffentlicht: (2024)
von: Sun, Yu, et al.
Veröffentlicht: (2024)
Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark
von: Turetzky, Arnon, et al.
Veröffentlicht: (2026)
von: Turetzky, Arnon, et al.
Veröffentlicht: (2026)
On Efficiently Representing Regular Languages as RNNs
von: Svete, Anej, et al.
Veröffentlicht: (2024)
von: Svete, Anej, et al.
Veröffentlicht: (2024)
Enhancing Transformer RNNs with Multiple Temporal Perspectives
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2024)
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2024)
Dynamic Speculation Lookahead Accelerates Speculative Decoding of Large Language Models
von: Mamou, Jonathan, et al.
Veröffentlicht: (2024)
von: Mamou, Jonathan, et al.
Veröffentlicht: (2024)
Beyond Performance: Quantifying and Mitigating Label Bias in LLMs
von: Reif, Yuval, et al.
Veröffentlicht: (2024)
von: Reif, Yuval, et al.
Veröffentlicht: (2024)
Arabic Dialect Classification using RNNs, Transformers, and Large Language Models: A Comparative Analysis
von: Essameldin, Omar A., et al.
Veröffentlicht: (2025)
von: Essameldin, Omar A., et al.
Veröffentlicht: (2025)
Attend First, Consolidate Later: On the Importance of Attention in Different LLM Layers
von: Ben-Artzy, Amit, et al.
Veröffentlicht: (2024)
von: Ben-Artzy, Amit, et al.
Veröffentlicht: (2024)
BabyHGRN: Exploring RNNs for Sample-Efficient Training of Language Models
von: Haller, Patrick, et al.
Veröffentlicht: (2024)
von: Haller, Patrick, et al.
Veröffentlicht: (2024)
Comba: Improving Bilinear RNNs with Closed-loop Control
von: Hu, Jiaxi, et al.
Veröffentlicht: (2025)
von: Hu, Jiaxi, et al.
Veröffentlicht: (2025)
Selective Attention Improves Transformer
von: Leviathan, Yaniv, et al.
Veröffentlicht: (2024)
von: Leviathan, Yaniv, et al.
Veröffentlicht: (2024)
Vocab Diet: Reshaping the Vocabulary of LLMs via Vector Arithmetic
von: Reif, Yuval, et al.
Veröffentlicht: (2025)
von: Reif, Yuval, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning
von: Hassid, Michael, et al.
Veröffentlicht: (2025) -
The Larger the Better? Improved LLM Code-Generation via Budget Reallocation
von: Hassid, Michael, et al.
Veröffentlicht: (2024) -
On Pruning State-Space LLMs
von: Ghattas, Tamer, et al.
Veröffentlicht: (2025) -
From Tokens to Words: On the Inner Lexicon of LLMs
von: Kaplan, Guy, et al.
Veröffentlicht: (2024) -
Scaling Analysis of Interleaved Speech-Text Language Models
von: Maimon, Gallil, et al.
Veröffentlicht: (2025)