Learning Extrapolative Sequence Transformations from Markov Chains
Fuente:
arXiv
Saved in:
| Main Authors: | Hager, Sophia, Khan, Aleem, Wang, Andrew, Andrews, Nicholas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning to Generate Text in Arbitrary Writing Styles
by: Khan, Aleem, et al.
Published: (2023)
by: Khan, Aleem, et al.
Published: (2023)
Uncertainty Distillation: Teaching Language Models to Express Semantic Confidence
by: Hager, Sophia, et al.
Published: (2025)
by: Hager, Sophia, et al.
Published: (2025)
Few-Shot Detection of Machine-Generated Text using Style Representations
by: Soto, Rafael Rivera, et al.
Published: (2024)
by: Soto, Rafael Rivera, et al.
Published: (2024)
Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains
by: Makkuva, Ashok Vardhan, et al.
Published: (2024)
by: Makkuva, Ashok Vardhan, et al.
Published: (2024)
Inducing Artificial Uncertainty in Language Models
by: Hager, Sophia, et al.
Published: (2026)
by: Hager, Sophia, et al.
Published: (2026)
AdapterSwap: Continuous Training of LLMs with Data Removal and Access-Control Guarantees
by: Fleshman, William, et al.
Published: (2024)
by: Fleshman, William, et al.
Published: (2024)
Most Likely Sequence Generation for $n$-Grams, Transformers, HMMs, and Markov Chains, by Using Rollout Algorithms
by: Li, Yuchao, et al.
Published: (2024)
by: Li, Yuchao, et al.
Published: (2024)
From Interpolation to Extrapolation: Complete Length Generalization for Arithmetic Transformers
by: Duan, Shaoxiong, et al.
Published: (2023)
by: Duan, Shaoxiong, et al.
Published: (2023)
Hell or High Water: Evaluating Agentic Recovery from External Failures
by: Wang, Andrew, et al.
Published: (2025)
by: Wang, Andrew, et al.
Published: (2025)
Automatic Differential Diagnosis using Transformer-Based Multi-Label Sequence Classification
by: Sadi, Abu Adnan, et al.
Published: (2024)
by: Sadi, Abu Adnan, et al.
Published: (2024)
Highly Fast Text Segmentation With Pairwise Markov Chains
by: Azeraf, Elie, et al.
Published: (2021)
by: Azeraf, Elie, et al.
Published: (2021)
Encoding Agent Trajectories as Representations with Sequence Transformers
by: Tsiligkaridis, Athanasios, et al.
Published: (2024)
by: Tsiligkaridis, Athanasios, et al.
Published: (2024)
Transforming Chatbot Text: A Sequence-to-Sequence Approach
by: Reddy, Natesh, et al.
Published: (2025)
by: Reddy, Natesh, et al.
Published: (2025)
FourierNAT: A Fourier-Mixing-Based Non-Autoregressive Transformer for Parallel Sequence Generation
by: Kiruluta, Andrew, et al.
Published: (2025)
by: Kiruluta, Andrew, et al.
Published: (2025)
Large Language Models as Markov Chains
by: Zekri, Oussama, et al.
Published: (2024)
by: Zekri, Oussama, et al.
Published: (2024)
Principled Gradient-based Markov Chain Monte Carlo for Text Generation
by: Du, Li, et al.
Published: (2023)
by: Du, Li, et al.
Published: (2023)
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
by: Setlur, Amrith, et al.
Published: (2025)
by: Setlur, Amrith, et al.
Published: (2025)
The Extrapolation Cliff in On-Policy Distillation of Near-Deterministic Structured Outputs
by: Li, Xin, et al.
Published: (2026)
by: Li, Xin, et al.
Published: (2026)
Model Extrapolation Expedites Alignment
by: Zheng, Chujie, et al.
Published: (2024)
by: Zheng, Chujie, et al.
Published: (2024)
Parallelizing Linear Transformers with the Delta Rule over Sequence Length
by: Yang, Songlin, et al.
Published: (2024)
by: Yang, Songlin, et al.
Published: (2024)
Can Authorship Attribution Models Distinguish Speakers in Speech Transcripts?
by: Aggazzotti, Cristina, et al.
Published: (2023)
by: Aggazzotti, Cristina, et al.
Published: (2023)
Mitigating Paraphrase Attacks on Machine-Text Detectors via Paraphrase Inversion
by: Soto, Rafael Rivera, et al.
Published: (2024)
by: Soto, Rafael Rivera, et al.
Published: (2024)
Large Language Models as Interpolated and Extrapolated Event Predictors
by: Zhang, Libo, et al.
Published: (2024)
by: Zhang, Libo, et al.
Published: (2024)
Dissociating Decodability and Causal Use in Bracket-Sequence Transformers
by: Sharma, Aryan, et al.
Published: (2026)
by: Sharma, Aryan, et al.
Published: (2026)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
by: Yang, Wenkai, et al.
Published: (2026)
by: Yang, Wenkai, et al.
Published: (2026)
How do Large Language Models Learn In-Context? Query and Key Matrices of In-Context Heads are Two Towers for Metric Learning
by: Yu, Zeping, et al.
Published: (2024)
by: Yu, Zeping, et al.
Published: (2024)
Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
Enhancing Length Extrapolation in Sequential Models with Pointer-Augmented Neural Memory
by: Le, Hung, et al.
Published: (2024)
by: Le, Hung, et al.
Published: (2024)
The Impact of Automatic Speech Transcription on Speaker Attribution
by: Aggazzotti, Cristina, et al.
Published: (2025)
by: Aggazzotti, Cristina, et al.
Published: (2025)
Transformers on Markov Data: Constant Depth Suffices
by: Rajaraman, Nived, et al.
Published: (2024)
by: Rajaraman, Nived, et al.
Published: (2024)
Position as Probability: Self-Supervised Transformers that Think Past Their Training for Length Extrapolation
by: Lee, Philip Heejun
Published: (2025)
by: Lee, Philip Heejun
Published: (2025)
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
by: Dong, Yihe, et al.
Published: (2025)
by: Dong, Yihe, et al.
Published: (2025)
RLPR: Extrapolating RLVR to General Domains without Verifiers
by: Yu, Tianyu, et al.
Published: (2025)
by: Yu, Tianyu, et al.
Published: (2025)
Rule Extrapolation in Language Models: A Study of Compositional Generalization on OOD Prompts
by: Mészáros, Anna, et al.
Published: (2024)
by: Mészáros, Anna, et al.
Published: (2024)
Sentiment Classification of Gaza War Headlines: A Comparative Analysis of Large Language Models and Arabic Fine-Tuned BERT Models
by: Eleraqi, Amr, et al.
Published: (2026)
by: Eleraqi, Amr, et al.
Published: (2026)
Evolving Demonstration Optimization for Chain-of-Thought Feature Transformation
by: Wang, Xinyuan, et al.
Published: (2026)
by: Wang, Xinyuan, et al.
Published: (2026)
CIEGAD: Cluster-Conditioned Interpolative and Extrapolative Framework for Geometry-Aware and Domain-Aligned Data Augmentation
by: Inoshita, Keito, et al.
Published: (2025)
by: Inoshita, Keito, et al.
Published: (2025)
You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories
by: Wei, Zhepei, et al.
Published: (2026)
by: Wei, Zhepei, et al.
Published: (2026)
Multistage Collaborative Knowledge Distillation from a Large Language Model for Semi-Supervised Sequence Generation
by: Zhao, Jiachen, et al.
Published: (2023)
by: Zhao, Jiachen, et al.
Published: (2023)
Changes by Butterflies: Farsighted Forecasting with Group Reservoir Transformer
by: Kowsher, Md, et al.
Published: (2024)
by: Kowsher, Md, et al.
Published: (2024)
Similar Items
-
Learning to Generate Text in Arbitrary Writing Styles
by: Khan, Aleem, et al.
Published: (2023) -
Uncertainty Distillation: Teaching Language Models to Express Semantic Confidence
by: Hager, Sophia, et al.
Published: (2025) -
Few-Shot Detection of Machine-Generated Text using Style Representations
by: Soto, Rafael Rivera, et al.
Published: (2024) -
Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains
by: Makkuva, Ashok Vardhan, et al.
Published: (2024) -
Inducing Artificial Uncertainty in Language Models
by: Hager, Sophia, et al.
Published: (2026)