Why Any-Order Autoregressive Models Need Two-Stream Attention: A Structural-Semantic Tradeoff
Fuente:
arXiv
Saved in:
| Main Authors: | Pynadath, Patrick, Zhang, Ruqi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generative Frontiers: Why Evaluation Matters for Diffusion Language Models
by: Pynadath, Patrick, et al.
Published: (2026)
by: Pynadath, Patrick, et al.
Published: (2026)
Controlled LLM Decoding via Discrete Auto-regressive Biasing
by: Pynadath, Patrick, et al.
Published: (2025)
by: Pynadath, Patrick, et al.
Published: (2025)
VERA: Variational Inference Framework for Jailbreaking Large Language Models
by: Lochab, Anamika, et al.
Published: (2025)
by: Lochab, Anamika, et al.
Published: (2025)
CANDI: Hybrid Discrete-Continuous Diffusion Models
by: Pynadath, Patrick, et al.
Published: (2025)
by: Pynadath, Patrick, et al.
Published: (2025)
AR-Omni: A Unified Autoregressive Model for Any-to-Any Generation
by: Cheng, Dongjie, et al.
Published: (2026)
by: Cheng, Dongjie, et al.
Published: (2026)
Single-Step Consistent Diffusion Samplers
by: Jutras-Dubé, Pascal, et al.
Published: (2025)
by: Jutras-Dubé, Pascal, et al.
Published: (2025)
Energy-Based Reward Models for Robust Language Model Alignment
by: Lochab, Anamika, et al.
Published: (2025)
by: Lochab, Anamika, et al.
Published: (2025)
Reviving Any-Subset Autoregressive Models with Principled Parallel Sampling and Speculative Decoding
by: Guo, Gabe, et al.
Published: (2025)
by: Guo, Gabe, et al.
Published: (2025)
Gradient-based Discrete Sampling with Automatic Cyclical Scheduling
by: Pynadath, Patrick, et al.
Published: (2024)
by: Pynadath, Patrick, et al.
Published: (2024)
Tokenization Tradeoffs in Structured EHR Foundation Models
by: Guo, Lin Lawrence, et al.
Published: (2026)
by: Guo, Lin Lawrence, et al.
Published: (2026)
CoT-UQ: Improving Response-wise Uncertainty Quantification in LLMs with Chain-of-Thought
by: Zhang, Boxuan, et al.
Published: (2025)
by: Zhang, Boxuan, et al.
Published: (2025)
Sherlock: Self-Correcting Reasoning in Vision-Language Models
by: Ding, Yi, et al.
Published: (2025)
by: Ding, Yi, et al.
Published: (2025)
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity
by: Lochab, Anamika, et al.
Published: (2026)
by: Lochab, Anamika, et al.
Published: (2026)
Why Softmax Attention Outperforms Linear Attention
by: Deng, Yichuan, et al.
Published: (2023)
by: Deng, Yichuan, et al.
Published: (2023)
Learning Semantic Structure through First-Order-Logic Translation
by: Chaturvedi, Akshay, et al.
Published: (2024)
by: Chaturvedi, Akshay, et al.
Published: (2024)
Mask-Enhanced Autoregressive Prediction: Pay Less Attention to Learn More
by: Zhuang, Xialie, et al.
Published: (2025)
by: Zhuang, Xialie, et al.
Published: (2025)
Why Are Positional Encodings Nonessential for Deep Autoregressive Transformers? Revisiting a Petroglyph
by: Irie, Kazuki
Published: (2024)
by: Irie, Kazuki
Published: (2024)
A Pseudo-Semantic Loss for Autoregressive Models with Logical Constraints
by: Ahmed, Kareem, et al.
Published: (2023)
by: Ahmed, Kareem, et al.
Published: (2023)
ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time
by: Ding, Yi, et al.
Published: (2024)
by: Ding, Yi, et al.
Published: (2024)
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
by: Tyukin, Georgy, et al.
Published: (2024)
by: Tyukin, Georgy, et al.
Published: (2024)
You Need Better Attention Priors
by: Litman, Elon, et al.
Published: (2026)
by: Litman, Elon, et al.
Published: (2026)
Attention Is Not All You Need: The Importance of Feedforward Networks in Transformer Models
by: Gerber, Isaac
Published: (2025)
by: Gerber, Isaac
Published: (2025)
Learning Self-Correction in Vision-Language Models via Rollout Augmentation
by: Ding, Yi, et al.
Published: (2026)
by: Ding, Yi, et al.
Published: (2026)
SALSA: Single-pass Autoregressive LLM Structured Classification
by: Berdichevsky, Ruslan, et al.
Published: (2025)
by: Berdichevsky, Ruslan, et al.
Published: (2025)
VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models
by: Liao, Qilin, et al.
Published: (2025)
by: Liao, Qilin, et al.
Published: (2025)
Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization
by: Zhang, Zheyuan, et al.
Published: (2026)
by: Zhang, Zheyuan, et al.
Published: (2026)
Cascade Reward Sampling for Efficient Decoding-Time Alignment
by: Li, Bolian, et al.
Published: (2024)
by: Li, Bolian, et al.
Published: (2024)
Attention Needs to Focus: A Unified Perspective on Attention Allocation
by: Fu, Zichuan, et al.
Published: (2026)
by: Fu, Zichuan, et al.
Published: (2026)
Routing Absorption in Sparse Attention: Why Random Gates Are Hard to Beat
by: Aquino-Michaels, Keston
Published: (2026)
by: Aquino-Michaels, Keston
Published: (2026)
Tensor Product Attention Is All You Need
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive Flows
by: Zhang, Ruixiang, et al.
Published: (2025)
by: Zhang, Ruixiang, et al.
Published: (2025)
Semantic Structure of Feature Space in Large Language Models
by: Kozlowski, Austin C., et al.
Published: (2026)
by: Kozlowski, Austin C., et al.
Published: (2026)
Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control
by: Li, Bolian, et al.
Published: (2026)
by: Li, Bolian, et al.
Published: (2026)
One Pass Streaming Algorithm for Super Long Token Attention Approximation in Sublinear Space
by: Addanki, Raghav, et al.
Published: (2023)
by: Addanki, Raghav, et al.
Published: (2023)
The Flexibility Trap: Why Arbitrary Order Limits Reasoning Potential in Diffusion Language Models
by: Ni, Zanlin, et al.
Published: (2026)
by: Ni, Zanlin, et al.
Published: (2026)
NExT-GPT: Any-to-Any Multimodal LLM
by: Wu, Shengqiong, et al.
Published: (2023)
by: Wu, Shengqiong, et al.
Published: (2023)
Feeding Two Birds or Favoring One? Adequacy-Fluency Tradeoffs in Evaluation and Meta-Evaluation of Machine Translation
by: Shayegh, Behzad, et al.
Published: (2025)
by: Shayegh, Behzad, et al.
Published: (2025)
Attention Smoothing Is All You Need For Unlearning
by: Zade, Saleh Zare, et al.
Published: (2026)
by: Zade, Saleh Zare, et al.
Published: (2026)
What Matters in Transformers? Not All Attention is Needed
by: He, Shwai, et al.
Published: (2024)
by: He, Shwai, et al.
Published: (2024)
Universe Routing: Why Self-Evolving Agents Need Epistemic Control
by: Wang, Zhaohui Geoffrey
Published: (2026)
by: Wang, Zhaohui Geoffrey
Published: (2026)
Similar Items
-
Generative Frontiers: Why Evaluation Matters for Diffusion Language Models
by: Pynadath, Patrick, et al.
Published: (2026) -
Controlled LLM Decoding via Discrete Auto-regressive Biasing
by: Pynadath, Patrick, et al.
Published: (2025) -
VERA: Variational Inference Framework for Jailbreaking Large Language Models
by: Lochab, Anamika, et al.
Published: (2025) -
CANDI: Hybrid Discrete-Continuous Diffusion Models
by: Pynadath, Patrick, et al.
Published: (2025) -
AR-Omni: A Unified Autoregressive Model for Any-to-Any Generation
by: Cheng, Dongjie, et al.
Published: (2026)