Gespeichert in:
| Hauptverfasser: | Wei, Xiuying, Yadav, Anunay, Pascanu, Razvan, Gulcehre, Caglar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2507.04416 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Investigating Low-Rank Training in Transformer Language Models: Efficiency and Scaling Analysis
von: Wei, Xiuying, et al.
Veröffentlicht: (2024)
von: Wei, Xiuying, et al.
Veröffentlicht: (2024)
Building on Efficient Foundations: Effectively Training LLMs with Structured Feedforward Layers
von: Wei, Xiuying, et al.
Veröffentlicht: (2024)
von: Wei, Xiuying, et al.
Veröffentlicht: (2024)
RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference
von: Wei, Xiuying, et al.
Veröffentlicht: (2026)
von: Wei, Xiuying, et al.
Veröffentlicht: (2026)
Augmenting Attention with Exponentially Decaying Memory Improves Query-Aware KV Sparsity
von: Wei, Xiuying, et al.
Veröffentlicht: (2026)
von: Wei, Xiuying, et al.
Veröffentlicht: (2026)
Beyond Autoregression: Fast LLMs via Self-Distillation Through Time
von: Deschenaux, Justin, et al.
Veröffentlicht: (2024)
von: Deschenaux, Justin, et al.
Veröffentlicht: (2024)
Promises, Outlooks and Challenges of Diffusion Language Modeling
von: Deschenaux, Justin, et al.
Veröffentlicht: (2024)
von: Deschenaux, Justin, et al.
Veröffentlicht: (2024)
Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
von: Orvieto, Antonio, et al.
Veröffentlicht: (2023)
von: Orvieto, Antonio, et al.
Veröffentlicht: (2023)
The Emergence of Chunking Structures with Hierarchical RNN
von: Wu, Zijun, et al.
Veröffentlicht: (2023)
von: Wu, Zijun, et al.
Veröffentlicht: (2023)
No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPO
von: Moalla, Skander, et al.
Veröffentlicht: (2024)
von: Moalla, Skander, et al.
Veröffentlicht: (2024)
From Markov to Laplace: How Mamba In-Context Learns Markov Chains
von: Bondaschi, Marco, et al.
Veröffentlicht: (2025)
von: Bondaschi, Marco, et al.
Veröffentlicht: (2025)
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
von: De, Soham, et al.
Veröffentlicht: (2024)
von: De, Soham, et al.
Veröffentlicht: (2024)
Aligning Large Language Models with Diverse Political Viewpoints
von: Stammbach, Dominik, et al.
Veröffentlicht: (2024)
von: Stammbach, Dominik, et al.
Veröffentlicht: (2024)
Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models
von: Gu, Xiangming, et al.
Veröffentlicht: (2026)
von: Gu, Xiangming, et al.
Veröffentlicht: (2026)
Context-Aware Toxicity Detection in Multiplayer Games: Integrating Domain-Adaptive Pretraining and Match Metadata
von: Schurger-Foy, Adrien, et al.
Veröffentlicht: (2025)
von: Schurger-Foy, Adrien, et al.
Veröffentlicht: (2025)
BlockGen: Flexible Blockwise Sequence Modeling with Hybrid Samplers
von: Deschenaux, Justin, et al.
Veröffentlicht: (2026)
von: Deschenaux, Justin, et al.
Veröffentlicht: (2026)
Fleet of Agents: Coordinated Problem Solving with Large Language Models
von: Klein, Lars, et al.
Veröffentlicht: (2024)
von: Klein, Lars, et al.
Veröffentlicht: (2024)
Self-Recognition in Language Models
von: Davidson, Tim R., et al.
Veröffentlicht: (2024)
von: Davidson, Tim R., et al.
Veröffentlicht: (2024)
Round and Round We Go! What makes Rotary Positional Encodings useful?
von: Barbero, Federico, et al.
Veröffentlicht: (2024)
von: Barbero, Federico, et al.
Veröffentlicht: (2024)
Perplexity Cannot Always Tell Right from Wrong
von: Veličković, Petar, et al.
Veröffentlicht: (2026)
von: Veličković, Petar, et al.
Veröffentlicht: (2026)
How do language models learn facts? Dynamics, curricula and hallucinations
von: Zucchet, Nicolas, et al.
Veröffentlicht: (2025)
von: Zucchet, Nicolas, et al.
Veröffentlicht: (2025)
The Illusion of Stochasticity in LLMs
von: Gu, Xiangming, et al.
Veröffentlicht: (2026)
von: Gu, Xiangming, et al.
Veröffentlicht: (2026)
CItruS: Chunked Instruction-aware State Eviction for Long Sequence Modeling
von: Bai, Yu, et al.
Veröffentlicht: (2024)
von: Bai, Yu, et al.
Veröffentlicht: (2024)
Why do LLMs attend to the first token?
von: Barbero, Federico, et al.
Veröffentlicht: (2025)
von: Barbero, Federico, et al.
Veröffentlicht: (2025)
Rodimus*: Breaking the Accuracy-Efficiency Trade-Off with Efficient Attentions
von: He, Zhihao, et al.
Veröffentlicht: (2024)
von: He, Zhihao, et al.
Veröffentlicht: (2024)
LLM-as-RNN: A Recurrent Language Model for Memory Updates and Sequence Prediction
von: Lu, Yuxing, et al.
Veröffentlicht: (2026)
von: Lu, Yuxing, et al.
Veröffentlicht: (2026)
A Study of the Plausibility of Attention between RNN Encoders in Natural Language Inference
von: Nguyen, Duc Hau, et al.
Veröffentlicht: (2025)
von: Nguyen, Duc Hau, et al.
Veröffentlicht: (2025)
ARWKV: Pretrain is not what we need, an RNN-Attention-Based Language Model Born from Transformer
von: Yueyu, Lin, et al.
Veröffentlicht: (2025)
von: Yueyu, Lin, et al.
Veröffentlicht: (2025)
Revisiting Dynamic Evaluation: Online Adaptation for Large Language Models
von: Rannen-Triki, Amal, et al.
Veröffentlicht: (2024)
von: Rannen-Triki, Amal, et al.
Veröffentlicht: (2024)
CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection
von: Song, Jiwon, et al.
Veröffentlicht: (2026)
von: Song, Jiwon, et al.
Veröffentlicht: (2026)
Chunking, Retrieval, and Re-ranking: An Empirical Evaluation of RAG Architectures for Policy Document Question Answering
von: Maharjan, Anuj, et al.
Veröffentlicht: (2026)
von: Maharjan, Anuj, et al.
Veröffentlicht: (2026)
Dynamic Chunking for Diffusion Language Models
von: Zhu, Yichen, et al.
Veröffentlicht: (2026)
von: Zhu, Yichen, et al.
Veröffentlicht: (2026)
MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
von: von Oswald, Johannes, et al.
Veröffentlicht: (2025)
von: von Oswald, Johannes, et al.
Veröffentlicht: (2025)
Transformers meet Neural Algorithmic Reasoners
von: Bounsi, Wilfried, et al.
Veröffentlicht: (2024)
von: Bounsi, Wilfried, et al.
Veröffentlicht: (2024)
Chunk-Distilled Language Modeling
von: Li, Yanhong, et al.
Veröffentlicht: (2024)
von: Li, Yanhong, et al.
Veröffentlicht: (2024)
ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition
von: Ye, Lu, et al.
Veröffentlicht: (2024)
von: Ye, Lu, et al.
Veröffentlicht: (2024)
ChunkRAG: Novel LLM-Chunk Filtering Method for RAG Systems
von: Singh, Ishneet Sukhvinder, et al.
Veröffentlicht: (2024)
von: Singh, Ishneet Sukhvinder, et al.
Veröffentlicht: (2024)
Learning Transductions and Alignments with RNN Seq2seq Models
von: Wang, Zhengxiang
Veröffentlicht: (2023)
von: Wang, Zhengxiang
Veröffentlicht: (2023)
Transformers need glasses! Information over-squashing in language tasks
von: Barbero, Federico, et al.
Veröffentlicht: (2024)
von: Barbero, Federico, et al.
Veröffentlicht: (2024)
Beyond Chunk-Local Extraction: Cross-Chunk Graph Augmentation for GraphRAG
von: Zhang, Jiaming, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2026)
GhostRNN: Reducing State Redundancy in RNN with Cheap Operations
von: Zhou, Hang, et al.
Veröffentlicht: (2024)
von: Zhou, Hang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Investigating Low-Rank Training in Transformer Language Models: Efficiency and Scaling Analysis
von: Wei, Xiuying, et al.
Veröffentlicht: (2024) -
Building on Efficient Foundations: Effectively Training LLMs with Structured Feedforward Layers
von: Wei, Xiuying, et al.
Veröffentlicht: (2024) -
RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference
von: Wei, Xiuying, et al.
Veröffentlicht: (2026) -
Augmenting Attention with Exponentially Decaying Memory Improves Query-Aware KV Sparsity
von: Wei, Xiuying, et al.
Veröffentlicht: (2026) -
Beyond Autoregression: Fast LLMs via Self-Distillation Through Time
von: Deschenaux, Justin, et al.
Veröffentlicht: (2024)