From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
Fuente:
arXiv
Guardado en:
| Autores principales: | Ildiz, M. Emrullah, Huang, Yixiao, Li, Yingcong, Rawat, Ankit Singh, Oymak, Samet |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Mechanics of Next Token Prediction with Self-Attention
por: Li, Yingcong, et al.
Publicado: (2024)
por: Li, Yingcong, et al.
Publicado: (2024)
Fine-grained Analysis of In-context Linear Estimation: Data, Architecture, and Beyond
por: Li, Yingcong, et al.
Publicado: (2024)
por: Li, Yingcong, et al.
Publicado: (2024)
Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
por: Li, Yingcong, et al.
Publicado: (2025)
por: Li, Yingcong, et al.
Publicado: (2025)
TimePFN: Effective Multivariate Time Series Forecasting with Synthetic Data
por: Taga, Ege Onur, et al.
Publicado: (2025)
por: Taga, Ege Onur, et al.
Publicado: (2025)
Retrieval Augmented Time Series Forecasting
por: Tire, Kutay, et al.
Publicado: (2024)
por: Tire, Kutay, et al.
Publicado: (2024)
Learning to Correct: Calibrated Reinforcement Learning for Multi-Attempt Chain-of-Thought
por: Ildiz, Muhammed Emrullah, et al.
Publicado: (2026)
por: Ildiz, Muhammed Emrullah, et al.
Publicado: (2026)
Transformers as Support Vector Machines
por: Tarzanagh, Davoud Ataee, et al.
Publicado: (2023)
por: Tarzanagh, Davoud Ataee, et al.
Publicado: (2023)
Continuous Chain of Thought Enables Parallel Exploration and Reasoning
por: Gozeten, Halil Alperen, et al.
Publicado: (2025)
por: Gozeten, Halil Alperen, et al.
Publicado: (2025)
When and How Unlabeled Data Provably Improve In-Context Learning
por: Li, Yingcong, et al.
Publicado: (2025)
por: Li, Yingcong, et al.
Publicado: (2025)
On the Power of Convolution Augmented Transformer
por: Li, Mingchen, et al.
Publicado: (2024)
por: Li, Mingchen, et al.
Publicado: (2024)
A Statistical Framework for Data-dependent Retrieval-Augmented Models
por: Basu, Soumya, et al.
Publicado: (2024)
por: Basu, Soumya, et al.
Publicado: (2024)
High-dimensional Analysis of Knowledge Distillation: Weak-to-Strong Generalization and Scaling Laws
por: Ildiz, M. Emrullah, et al.
Publicado: (2024)
por: Ildiz, M. Emrullah, et al.
Publicado: (2024)
Test-Time Training Provably Improves Transformers as In-context Learners
por: Gozeten, Halil Alperen, et al.
Publicado: (2025)
por: Gozeten, Halil Alperen, et al.
Publicado: (2025)
Efficient Contextual LLM Cascades through Budget-Constrained Policy Learning
por: Zhang, Xuechen, et al.
Publicado: (2024)
por: Zhang, Xuechen, et al.
Publicado: (2024)
SmartChunk Retrieval: Query-Aware Chunk Compression with Planning for Efficient Document RAG
por: Zhang, Xuechen, et al.
Publicado: (2025)
por: Zhang, Xuechen, et al.
Publicado: (2025)
Provable Benefits of Task-Specific Prompts for In-context Learning
por: Chang, Xiangyu, et al.
Publicado: (2025)
por: Chang, Xiangyu, et al.
Publicado: (2025)
Can Transformers Learn Optimal Filtering for Unknown Systems?
por: Balim, Haldun, et al.
Publicado: (2023)
por: Balim, Haldun, et al.
Publicado: (2023)
Language Model Cascades: Token-level uncertainty and beyond
por: Gupta, Neha, et al.
Publicado: (2024)
por: Gupta, Neha, et al.
Publicado: (2024)
Think before you speak: Training Language Models With Pause Tokens
por: Goyal, Sachin, et al.
Publicado: (2023)
por: Goyal, Sachin, et al.
Publicado: (2023)
Evolutionary Multi-Task Optimization for LLM-Guided Program Discovery
por: Gozeten, Halil Alperen, et al.
Publicado: (2026)
por: Gozeten, Halil Alperen, et al.
Publicado: (2026)
In-Context Learning Under Regime Change
por: Dudley, Carson, et al.
Publicado: (2026)
por: Dudley, Carson, et al.
Publicado: (2026)
Extrapolation by Association: Length Generalization Transfer in Transformers
por: Cai, Ziyang, et al.
Publicado: (2025)
por: Cai, Ziyang, et al.
Publicado: (2025)
Selective Attention: Enhancing Transformer through Principled Context Control
por: Zhang, Xuechen, et al.
Publicado: (2024)
por: Zhang, Xuechen, et al.
Publicado: (2024)
Attention with Trained Embeddings Provably Selects Important Tokens
por: Wu, Diyuan, et al.
Publicado: (2025)
por: Wu, Diyuan, et al.
Publicado: (2025)
Faster Cascades via Speculative Decoding
por: Narasimhan, Harikrishna, et al.
Publicado: (2024)
por: Narasimhan, Harikrishna, et al.
Publicado: (2024)
Mixture of Chapters: Scaling Learnt Memory in Transformers
por: Tibrewal, Tasmay Pankaj, et al.
Publicado: (2026)
por: Tibrewal, Tasmay Pankaj, et al.
Publicado: (2026)
RecurFormer: Not All Transformer Heads Need Self-Attention
por: Yan, Ruiqing, et al.
Publicado: (2024)
por: Yan, Ruiqing, et al.
Publicado: (2024)
Recent Advances in Generative AI and Large Language Models: Current Status, Challenges, and Perspectives
por: Hagos, Desta Haileselassie, et al.
Publicado: (2024)
por: Hagos, Desta Haileselassie, et al.
Publicado: (2024)
DistillSpec: Improving Speculative Decoding via Knowledge Distillation
por: Zhou, Yongchao, et al.
Publicado: (2023)
por: Zhou, Yongchao, et al.
Publicado: (2023)
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
por: Su, Jingtong, et al.
Publicado: (2025)
por: Su, Jingtong, et al.
Publicado: (2025)
Memory as a Markov Matrix: Sample Efficient Knowledge Expansion via Token-to-Dictionary Mapping
por: Pethkar, Kaustubh, et al.
Publicado: (2026)
por: Pethkar, Kaustubh, et al.
Publicado: (2026)
Integrating Locality-Aware Attention with Transformers for General Geometry PDEs
por: Koh, Minsu, et al.
Publicado: (2025)
por: Koh, Minsu, et al.
Publicado: (2025)
Dynamics of Spontaneous Topic Changes in Next Token Prediction with Self-Attention
por: Jia, Mumin, et al.
Publicado: (2025)
por: Jia, Mumin, et al.
Publicado: (2025)
IDEA Prune: An Integrated Enlarge-and-Prune Pipeline in Generative Language Model Pretraining
por: Li, Yixiao, et al.
Publicado: (2025)
por: Li, Yixiao, et al.
Publicado: (2025)
JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
por: Tian, Yuandong, et al.
Publicado: (2023)
por: Tian, Yuandong, et al.
Publicado: (2023)
Plug-and-Play Transformer Modules for Test-Time Adaptation
por: Chang, Xiangyu, et al.
Publicado: (2024)
por: Chang, Xiangyu, et al.
Publicado: (2024)
Empirical Capacity Model for Self-Attention Neural Networks
por: Härmä, Aki, et al.
Publicado: (2024)
por: Härmä, Aki, et al.
Publicado: (2024)
What Matters in Transformers? Not All Attention is Needed
por: He, Shwai, et al.
Publicado: (2024)
por: He, Shwai, et al.
Publicado: (2024)
Selective Attention Improves Transformer
por: Leviathan, Yaniv, et al.
Publicado: (2024)
por: Leviathan, Yaniv, et al.
Publicado: (2024)
Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention
por: Guo, Zhenyu, et al.
Publicado: (2025)
por: Guo, Zhenyu, et al.
Publicado: (2025)
Ejemplares similares
-
Mechanics of Next Token Prediction with Self-Attention
por: Li, Yingcong, et al.
Publicado: (2024) -
Fine-grained Analysis of In-context Linear Estimation: Data, Architecture, and Beyond
por: Li, Yingcong, et al.
Publicado: (2024) -
Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
por: Li, Yingcong, et al.
Publicado: (2025) -
TimePFN: Effective Multivariate Time Series Forecasting with Synthetic Data
por: Taga, Ege Onur, et al.
Publicado: (2025) -
Retrieval Augmented Time Series Forecasting
por: Tire, Kutay, et al.
Publicado: (2024)