Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Makkuva, Ashok Vardhan, Bondaschi, Marco, Girish, Adway, Nagle, Alliot, Jaggi, Martin, Kim, Hyeji, Gastpar, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Fundamental Limits of Prompt Compression: A Rate-Distortion Framework for Black-Box Language Models
von: Nagle, Alliot, et al.
Veröffentlicht: (2024)
von: Nagle, Alliot, et al.
Veröffentlicht: (2024)
Local to Global: Learning Dynamics and Effect of Initialization for Transformers
von: Makkuva, Ashok Vardhan, et al.
Veröffentlicht: (2024)
von: Makkuva, Ashok Vardhan, et al.
Veröffentlicht: (2024)
Transformers on Markov Data: Constant Depth Suffices
von: Rajaraman, Nived, et al.
Veröffentlicht: (2024)
von: Rajaraman, Nived, et al.
Veröffentlicht: (2024)
TERMINATOR: Learning Optimal Exit Points for Early Stopping in Chain-of-Thought Reasoning
von: Nagle, Alliot, et al.
Veröffentlicht: (2026)
von: Nagle, Alliot, et al.
Veröffentlicht: (2026)
LASER: Linear Compression in Wireless Distributed Optimization
von: Makkuva, Ashok Vardhan, et al.
Veröffentlicht: (2023)
von: Makkuva, Ashok Vardhan, et al.
Veröffentlicht: (2023)
From Markov to Laplace: How Mamba In-Context Learns Markov Chains
von: Bondaschi, Marco, et al.
Veröffentlicht: (2025)
von: Bondaschi, Marco, et al.
Veröffentlicht: (2025)
What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov Chains
von: Ekbote, Chanakya, et al.
Veröffentlicht: (2025)
von: Ekbote, Chanakya, et al.
Veröffentlicht: (2025)
Alpha-NML Universal Predictors
von: Bondaschi, Marco, et al.
Veröffentlicht: (2022)
von: Bondaschi, Marco, et al.
Veröffentlicht: (2022)
Batch Universal Prediction
von: Bondaschi, Marco, et al.
Veröffentlicht: (2024)
von: Bondaschi, Marco, et al.
Veröffentlicht: (2024)
The Conditional Regret-Capacity Theorem for Batch Universal Prediction
von: Bondaschi, Marco, et al.
Veröffentlicht: (2025)
von: Bondaschi, Marco, et al.
Veröffentlicht: (2025)
Neural Distributed Source Coding
von: Whang, Jay, et al.
Veröffentlicht: (2021)
von: Whang, Jay, et al.
Veröffentlicht: (2021)
On entropy-constrained Gaussian channel capacity via the moment problem
von: Girish, Adway, et al.
Veröffentlicht: (2025)
von: Girish, Adway, et al.
Veröffentlicht: (2025)
Open-Source Reproduction and Explainability Analysis of Corrective Retrieval Augmented Generation
von: Yalavarthi, Surya Vardhan
Veröffentlicht: (2026)
von: Yalavarthi, Surya Vardhan
Veröffentlicht: (2026)
High signal-to-noise ratio asymptotics of entropy-constrained Gaussian channel capacity
von: Girish, Adway, et al.
Veröffentlicht: (2026)
von: Girish, Adway, et al.
Veröffentlicht: (2026)
Principled Gradient-based Markov Chain Monte Carlo for Text Generation
von: Du, Li, et al.
Veröffentlicht: (2023)
von: Du, Li, et al.
Veröffentlicht: (2023)
Virtual Quantum Markov Chains
von: Chen, Yu-Ao, et al.
Veröffentlicht: (2023)
von: Chen, Yu-Ao, et al.
Veröffentlicht: (2023)
Batch Normalization Decomposed
von: Nachum, Ido, et al.
Veröffentlicht: (2024)
von: Nachum, Ido, et al.
Veröffentlicht: (2024)
Learning Extrapolative Sequence Transformations from Markov Chains
von: Hager, Sophia, et al.
Veröffentlicht: (2025)
von: Hager, Sophia, et al.
Veröffentlicht: (2025)
Geometry and Duality of Alternating Markov Chains
von: Mithal, Deven, et al.
Veröffentlicht: (2024)
von: Mithal, Deven, et al.
Veröffentlicht: (2024)
Identifying the Source of Information Spread in Networks via Markov Chains
von: Sabato, Yael, et al.
Veröffentlicht: (2024)
von: Sabato, Yael, et al.
Veröffentlicht: (2024)
The Markov-Chain Polytope with Applications
von: Golin, Mordecai J., et al.
Veröffentlicht: (2024)
von: Golin, Mordecai J., et al.
Veröffentlicht: (2024)
On the suboptimality of linear codes for binary distributed hypothesis testing
von: Girish, Adway, et al.
Veröffentlicht: (2026)
von: Girish, Adway, et al.
Veröffentlicht: (2026)
Towards Detecting LLMs Hallucination via Markov Chain-based Multi-agent Debate Framework
von: Sun, Xiaoxi, et al.
Veröffentlicht: (2024)
von: Sun, Xiaoxi, et al.
Veröffentlicht: (2024)
Shared Information for a Markov Chain on a Tree
von: Bhattacharya, Sagnik, et al.
Veröffentlicht: (2023)
von: Bhattacharya, Sagnik, et al.
Veröffentlicht: (2023)
GRACE: Generative Recommendation via Journey-Aware Sparse Attention on Chain-of-Thought Tokenization
von: Ma, Luyi, et al.
Veröffentlicht: (2025)
von: Ma, Luyi, et al.
Veröffentlicht: (2025)
Converse for Multi-Server Single-Message PIR with Side Information
von: Li, Su, et al.
Veröffentlicht: (2018)
von: Li, Su, et al.
Veröffentlicht: (2018)
Shannon Bounds for Quadratic Rate-Distortion Problems
von: Gastpar, Michael, et al.
Veröffentlicht: (2024)
von: Gastpar, Michael, et al.
Veröffentlicht: (2024)
An Information-Theoretic Approach to Understanding Transformers' In-Context Learning of Variable-Order Markov Chains
von: Zhou, Ruida, et al.
Veröffentlicht: (2024)
von: Zhou, Ruida, et al.
Veröffentlicht: (2024)
Near-Optimal Clustering in Mixture of Markov Chains
von: Lee, Junghyun, et al.
Veröffentlicht: (2025)
von: Lee, Junghyun, et al.
Veröffentlicht: (2025)
Contraction of Markovian Operators in Orlicz Spaces and Error Bounds for Markov Chain Monte Carlo
von: Esposito, Amedeo Roberto, et al.
Veröffentlicht: (2024)
von: Esposito, Amedeo Roberto, et al.
Veröffentlicht: (2024)
Single-Period Portfolio Selection via Information Projection
von: Yang, Bo-Yu, et al.
Veröffentlicht: (2026)
von: Yang, Bo-Yu, et al.
Veröffentlicht: (2026)
Item Association Factorization Mixed Markov Chains for Sequential Recommendation
von: Du, DongYu, et al.
Veröffentlicht: (2024)
von: Du, DongYu, et al.
Veröffentlicht: (2024)
Algorithmic Randomness in Continuous-Time Markov Chains
von: Huang, Xiang, et al.
Veröffentlicht: (2019)
von: Huang, Xiang, et al.
Veröffentlicht: (2019)
Large Language Models as Markov Chains
von: Zekri, Oussama, et al.
Veröffentlicht: (2024)
von: Zekri, Oussama, et al.
Veröffentlicht: (2024)
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
von: Ildiz, M. Emrullah, et al.
Veröffentlicht: (2024)
von: Ildiz, M. Emrullah, et al.
Veröffentlicht: (2024)
Model non-collapse: Minimax bounds for recursive discrete distribution estimation
von: Kanabar, Millen, et al.
Veröffentlicht: (2025)
von: Kanabar, Millen, et al.
Veröffentlicht: (2025)
Absorbing Markov Chain-Based Analysis of Age of Information in Discrete-Time Dual-Queue Systems
von: Feng, Yifan, et al.
Veröffentlicht: (2025)
von: Feng, Yifan, et al.
Veröffentlicht: (2025)
Markov Chain of Thought for Efficient Mathematical Reasoning
von: Yang, Wen, et al.
Veröffentlicht: (2024)
von: Yang, Wen, et al.
Veröffentlicht: (2024)
Beyond Decisiveness of Infinite Markov Chains
von: Barbot, Benoît, et al.
Veröffentlicht: (2024)
von: Barbot, Benoît, et al.
Veröffentlicht: (2024)
Evolutionary Cooperation with Game Transitions via Markov Decision Chain in Networked Population
von: Luo, Chaoyang, et al.
Veröffentlicht: (2025)
von: Luo, Chaoyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Fundamental Limits of Prompt Compression: A Rate-Distortion Framework for Black-Box Language Models
von: Nagle, Alliot, et al.
Veröffentlicht: (2024) -
Local to Global: Learning Dynamics and Effect of Initialization for Transformers
von: Makkuva, Ashok Vardhan, et al.
Veröffentlicht: (2024) -
Transformers on Markov Data: Constant Depth Suffices
von: Rajaraman, Nived, et al.
Veröffentlicht: (2024) -
TERMINATOR: Learning Optimal Exit Points for Early Stopping in Chain-of-Thought Reasoning
von: Nagle, Alliot, et al.
Veröffentlicht: (2026) -
LASER: Linear Compression in Wireless Distributed Optimization
von: Makkuva, Ashok Vardhan, et al.
Veröffentlicht: (2023)