From Markov to Laplace: How Mamba In-Context Learns Markov Chains
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bondaschi, Marco, Rajaraman, Nived, Wei, Xiuying, Ramchandran, Kannan, Pascanu, Razvan, Gulcehre, Caglar, Gastpar, Michael, Makkuva, Ashok Vardhan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Transformers on Markov Data: Constant Depth Suffices
von: Rajaraman, Nived, et al.
Veröffentlicht: (2024)
von: Rajaraman, Nived, et al.
Veröffentlicht: (2024)
Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains
von: Makkuva, Ashok Vardhan, et al.
Veröffentlicht: (2024)
von: Makkuva, Ashok Vardhan, et al.
Veröffentlicht: (2024)
What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov Chains
von: Ekbote, Chanakya, et al.
Veröffentlicht: (2025)
von: Ekbote, Chanakya, et al.
Veröffentlicht: (2025)
Fundamental Limits of Prompt Compression: A Rate-Distortion Framework for Black-Box Language Models
von: Nagle, Alliot, et al.
Veröffentlicht: (2024)
von: Nagle, Alliot, et al.
Veröffentlicht: (2024)
LASER: Linear Compression in Wireless Distributed Optimization
von: Makkuva, Ashok Vardhan, et al.
Veröffentlicht: (2023)
von: Makkuva, Ashok Vardhan, et al.
Veröffentlicht: (2023)
Local to Global: Learning Dynamics and Effect of Initialization for Transformers
von: Makkuva, Ashok Vardhan, et al.
Veröffentlicht: (2024)
von: Makkuva, Ashok Vardhan, et al.
Veröffentlicht: (2024)
Statistical Complexity and Optimal Algorithms for Non-linear Ridge Bandits
von: Rajaraman, Nived, et al.
Veröffentlicht: (2023)
von: Rajaraman, Nived, et al.
Veröffentlicht: (2023)
Alpha-NML Universal Predictors
von: Bondaschi, Marco, et al.
Veröffentlicht: (2022)
von: Bondaschi, Marco, et al.
Veröffentlicht: (2022)
Building on Efficient Foundations: Effectively Training LLMs with Structured Feedforward Layers
von: Wei, Xiuying, et al.
Veröffentlicht: (2024)
von: Wei, Xiuying, et al.
Veröffentlicht: (2024)
Investigating Low-Rank Training in Transformer Language Models: Efficiency and Scaling Analysis
von: Wei, Xiuying, et al.
Veröffentlicht: (2024)
von: Wei, Xiuying, et al.
Veröffentlicht: (2024)
RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence Modeling
von: Wei, Xiuying, et al.
Veröffentlicht: (2025)
von: Wei, Xiuying, et al.
Veröffentlicht: (2025)
Batch Universal Prediction
von: Bondaschi, Marco, et al.
Veröffentlicht: (2024)
von: Bondaschi, Marco, et al.
Veröffentlicht: (2024)
The Conditional Regret-Capacity Theorem for Batch Universal Prediction
von: Bondaschi, Marco, et al.
Veröffentlicht: (2025)
von: Bondaschi, Marco, et al.
Veröffentlicht: (2025)
Toward a Theory of Tokenization in LLMs
von: Rajaraman, Nived, et al.
Veröffentlicht: (2024)
von: Rajaraman, Nived, et al.
Veröffentlicht: (2024)
Augmenting Attention with Exponentially Decaying Memory Improves Query-Aware KV Sparsity
von: Wei, Xiuying, et al.
Veröffentlicht: (2026)
von: Wei, Xiuying, et al.
Veröffentlicht: (2026)
RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference
von: Wei, Xiuying, et al.
Veröffentlicht: (2026)
von: Wei, Xiuying, et al.
Veröffentlicht: (2026)
TERMINATOR: Learning Optimal Exit Points for Early Stopping in Chain-of-Thought Reasoning
von: Nagle, Alliot, et al.
Veröffentlicht: (2026)
von: Nagle, Alliot, et al.
Veröffentlicht: (2026)
No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPO
von: Moalla, Skander, et al.
Veröffentlicht: (2024)
von: Moalla, Skander, et al.
Veröffentlicht: (2024)
Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
von: Orvieto, Antonio, et al.
Veröffentlicht: (2023)
von: Orvieto, Antonio, et al.
Veröffentlicht: (2023)
Markov Chains Approximate Message Passing
von: Rajaraman, Amit, et al.
Veröffentlicht: (2025)
von: Rajaraman, Amit, et al.
Veröffentlicht: (2025)
Virtual Quantum Markov Chains
von: Chen, Yu-Ao, et al.
Veröffentlicht: (2023)
von: Chen, Yu-Ao, et al.
Veröffentlicht: (2023)
Batch Normalization Decomposed
von: Nachum, Ido, et al.
Veröffentlicht: (2024)
von: Nachum, Ido, et al.
Veröffentlicht: (2024)
Quantifying Positional Biases in Text Embedding Models
von: Lee, Reagan J., et al.
Veröffentlicht: (2024)
von: Lee, Reagan J., et al.
Veröffentlicht: (2024)
Geometry and Duality of Alternating Markov Chains
von: Mithal, Deven, et al.
Veröffentlicht: (2024)
von: Mithal, Deven, et al.
Veröffentlicht: (2024)
The Markov-Chain Polytope with Applications
von: Golin, Mordecai J., et al.
Veröffentlicht: (2024)
von: Golin, Mordecai J., et al.
Veröffentlicht: (2024)
An Information-Theoretic Approach to Understanding Transformers' In-Context Learning of Variable-Order Markov Chains
von: Zhou, Ruida, et al.
Veröffentlicht: (2024)
von: Zhou, Ruida, et al.
Veröffentlicht: (2024)
Shared Information for a Markov Chain on a Tree
von: Bhattacharya, Sagnik, et al.
Veröffentlicht: (2023)
von: Bhattacharya, Sagnik, et al.
Veröffentlicht: (2023)
Converse for Multi-Server Single-Message PIR with Side Information
von: Li, Su, et al.
Veröffentlicht: (2018)
von: Li, Su, et al.
Veröffentlicht: (2018)
Shannon Bounds for Quadratic Rate-Distortion Problems
von: Gastpar, Michael, et al.
Veröffentlicht: (2024)
von: Gastpar, Michael, et al.
Veröffentlicht: (2024)
Softmax is not Enough (for Sharp Size Generalisation)
von: Veličković, Petar, et al.
Veröffentlicht: (2024)
von: Veličković, Petar, et al.
Veröffentlicht: (2024)
Near-Optimal Clustering in Mixture of Markov Chains
von: Lee, Junghyun, et al.
Veröffentlicht: (2025)
von: Lee, Junghyun, et al.
Veröffentlicht: (2025)
Contraction of Markovian Operators in Orlicz Spaces and Error Bounds for Markov Chain Monte Carlo
von: Esposito, Amedeo Roberto, et al.
Veröffentlicht: (2024)
von: Esposito, Amedeo Roberto, et al.
Veröffentlicht: (2024)
Item Association Factorization Mixed Markov Chains for Sequential Recommendation
von: Du, DongYu, et al.
Veröffentlicht: (2024)
von: Du, DongYu, et al.
Veröffentlicht: (2024)
Algorithmic Randomness in Continuous-Time Markov Chains
von: Huang, Xiang, et al.
Veröffentlicht: (2019)
von: Huang, Xiang, et al.
Veröffentlicht: (2019)
Model non-collapse: Minimax bounds for recursive discrete distribution estimation
von: Kanabar, Millen, et al.
Veröffentlicht: (2025)
von: Kanabar, Millen, et al.
Veröffentlicht: (2025)
Identifying the Source of Information Spread in Networks via Markov Chains
von: Sabato, Yael, et al.
Veröffentlicht: (2024)
von: Sabato, Yael, et al.
Veröffentlicht: (2024)
Interactive Learning of Single-Index Models via Stochastic Gradient Descent
von: Rajaraman, Nived, et al.
Veröffentlicht: (2026)
von: Rajaraman, Nived, et al.
Veröffentlicht: (2026)
Gradient-Based Markov Chain Monte Carlo for MIMO Detection
von: Zhou, Xingyu, et al.
Veröffentlicht: (2023)
von: Zhou, Xingyu, et al.
Veröffentlicht: (2023)
$\texttt{SPECS}$: Faster Test-Time Scaling through Speculative Drafts
von: Cemri, Mert, et al.
Veröffentlicht: (2025)
von: Cemri, Mert, et al.
Veröffentlicht: (2025)
Fastest Mixing Reversible Markov Chain: Clique Lifted Graphs and Subgraphs
von: Jafarizadeh, Saber
Veröffentlicht: (2025)
von: Jafarizadeh, Saber
Veröffentlicht: (2025)
Ähnliche Einträge
-
Transformers on Markov Data: Constant Depth Suffices
von: Rajaraman, Nived, et al.
Veröffentlicht: (2024) -
Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains
von: Makkuva, Ashok Vardhan, et al.
Veröffentlicht: (2024) -
What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov Chains
von: Ekbote, Chanakya, et al.
Veröffentlicht: (2025) -
Fundamental Limits of Prompt Compression: A Rate-Distortion Framework for Black-Box Language Models
von: Nagle, Alliot, et al.
Veröffentlicht: (2024) -
LASER: Linear Compression in Wireless Distributed Optimization
von: Makkuva, Ashok Vardhan, et al.
Veröffentlicht: (2023)