An Analysis of Attention via the Lens of Exchangeability and Latent Variable Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yufeng, Liu, Boyi, Cai, Qi, Wang, Lingxiao, Wang, Zhaoran |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Embed to Control Partially Observed Systems: Representation Learning with Provable Sample Efficiency
by: Wang, Lingxiao, et al.
Published: (2022)
by: Wang, Lingxiao, et al.
Published: (2022)
Can Temporal-Difference and Q-Learning Learn Representation? A Mean-Field Theory
by: Zhang, Yufeng, et al.
Published: (2020)
by: Zhang, Yufeng, et al.
Published: (2020)
Double Duality: Variational Primal-Dual Policy Optimization for Constrained Reinforcement Learning
by: Li, Zihao, et al.
Published: (2024)
by: Li, Zihao, et al.
Published: (2024)
Exchangeable Sequence Models Quantify Uncertainty Over Latent Concepts
by: Ye, Naimeng, et al.
Published: (2024)
by: Ye, Naimeng, et al.
Published: (2024)
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs
by: Zhang, Shenao, et al.
Published: (2024)
by: Zhang, Shenao, et al.
Published: (2024)
Reinforcement Learning from Partial Observation: Linear Function Approximation with Provable Sample Efficiency
by: Cai, Qi, et al.
Published: (2022)
by: Cai, Qi, et al.
Published: (2022)
Contrastive UCB: Provably Efficient Contrastive Self-Supervised Learning in Online Reinforcement Learning
by: Qiu, Shuang, et al.
Published: (2022)
by: Qiu, Shuang, et al.
Published: (2022)
Rethinking Bivariate Causal Discovery Through the Lens of Exchangeability
by: Brogueira, Tiago, et al.
Published: (2025)
by: Brogueira, Tiago, et al.
Published: (2025)
A Mean-Field Analysis of Neural Stochastic Gradient Descent-Ascent for Functional Minimax Optimization
by: Zhu, Yuchen, et al.
Published: (2024)
by: Zhu, Yuchen, et al.
Published: (2024)
Wasserstein Flow Meets Replicator Dynamics: A Mean-Field Analysis of Representation Learning in Actor-Critic
by: Zhang, Yufeng, et al.
Published: (2021)
by: Zhang, Yufeng, et al.
Published: (2021)
Variational Transport: A Convergent Particle-BasedAlgorithm for Distributional Optimization
by: Yang, Zhuoran, et al.
Published: (2020)
by: Yang, Zhuoran, et al.
Published: (2020)
Federated Offline Reinforcement Learning
by: Zhou, Doudou, et al.
Published: (2022)
by: Zhou, Doudou, et al.
Published: (2022)
Reason for Future, Act for Now: A Principled Framework for Autonomous LLM Agents with Provable Sample Efficiency
by: Liu, Zhihan, et al.
Published: (2023)
by: Liu, Zhihan, et al.
Published: (2023)
Provably Efficient Exploration in Policy Optimization
by: Cai, Qi, et al.
Published: (2019)
by: Cai, Qi, et al.
Published: (2019)
Identifiable Latent Polynomial Causal Models Through the Lens of Change
by: Liu, Yuhang, et al.
Published: (2023)
by: Liu, Yuhang, et al.
Published: (2023)
Training-Free Adaptation of Diffusion Models via Doob's $h$-Transform
by: Zhu, Qijie, et al.
Published: (2026)
by: Zhu, Qijie, et al.
Published: (2026)
Instrumental Variable Value Iteration for Causal Offline Reinforcement Learning
by: Liao, Luofeng, et al.
Published: (2021)
by: Liao, Luofeng, et al.
Published: (2021)
Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer
by: Liu, Zhihan, et al.
Published: (2024)
by: Liu, Zhihan, et al.
Published: (2024)
Diffusion Models are Secretly Exchangeable: Parallelizing DDPMs via Autospeculation
by: Hu, Hengyuan, et al.
Published: (2025)
by: Hu, Hengyuan, et al.
Published: (2025)
Power-based Partial Attention: Bridging Linear-Complexity and Full Attention
by: Huang, Yufeng
Published: (2026)
by: Huang, Yufeng
Published: (2026)
Architectural and Inferential Inductive Biases For Exchangeable Sequence Modeling
by: Mittal, Daksh, et al.
Published: (2025)
by: Mittal, Daksh, et al.
Published: (2025)
Stabilizing Self-Consuming Diffusion Models with Latent Space Filtering
by: Cai, Zhongteng, et al.
Published: (2025)
by: Cai, Zhongteng, et al.
Published: (2025)
Generalization Error Analysis for Selective State-Space Models Through the Lens of Attention
by: Honarpisheh, Arya, et al.
Published: (2025)
by: Honarpisheh, Arya, et al.
Published: (2025)
Network Causal Effect Estimation In Graphical Models Of Contagion And Latent Confounding
by: Wu, Yufeng, et al.
Published: (2024)
by: Wu, Yufeng, et al.
Published: (2024)
Causal Discovery for Linear DAGs with Dependent Latent Variables via Higher-order Cumulants
by: Cai, Ming, et al.
Published: (2025)
by: Cai, Ming, et al.
Published: (2025)
Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
by: Zuo, Yifei, et al.
Published: (2025)
by: Zuo, Yifei, et al.
Published: (2025)
Delving into Non-Exchangeability for Conformal Prediction in Graph-Structured Multivariate Time Series
by: Guo, Ruichao, et al.
Published: (2026)
by: Guo, Ruichao, et al.
Published: (2026)
Parallax: Parameterized Local Linear Attention for Language Modeling
by: Zuo, Yifei, et al.
Published: (2026)
by: Zuo, Yifei, et al.
Published: (2026)
Superlinear Multi-Step Attention
by: Huang, Yufeng
Published: (2026)
by: Huang, Yufeng
Published: (2026)
Identifiable Deep Latent Variable Models for MNAR Data
by: Xie, Huiming, et al.
Published: (2026)
by: Xie, Huiming, et al.
Published: (2026)
Non-Exchangeable Conformal Risk Control
by: Farinhas, António, et al.
Published: (2023)
by: Farinhas, António, et al.
Published: (2023)
Do Finetti: On Causal Effects for Exchangeable Data
by: Guo, Siyuan, et al.
Published: (2024)
by: Guo, Siyuan, et al.
Published: (2024)
OmniLens++: Blind Lens Aberration Correction via Large LensLib Pre-Training and Latent PSF Representation
by: Jiang, Qi, et al.
Published: (2025)
by: Jiang, Qi, et al.
Published: (2025)
Let Models Speak Ciphers: Multiagent Debate through Embeddings
by: Pham, Chau, et al.
Published: (2023)
by: Pham, Chau, et al.
Published: (2023)
Scalable Random Feature Latent Variable Models
by: Li, Ying, et al.
Published: (2024)
by: Li, Ying, et al.
Published: (2024)
Preventing Model Collapse in Gaussian Process Latent Variable Models
by: Li, Ying, et al.
Published: (2024)
by: Li, Ying, et al.
Published: (2024)
Are Transformers Able to Reason by Connecting Separated Knowledge in Training Data?
by: Yin, Yutong, et al.
Published: (2025)
by: Yin, Yutong, et al.
Published: (2025)
Efficient Video Diffusion Models via Content-Frame Motion-Latent Decomposition
by: Yu, Sihyun, et al.
Published: (2024)
by: Yu, Sihyun, et al.
Published: (2024)
Particle Dynamics for Latent-Variable Energy-Based Models
by: Tang, Shiqin, et al.
Published: (2025)
by: Tang, Shiqin, et al.
Published: (2025)
LaX: Boosting Low-Rank Training of Foundation Models via Latent Crossing
by: Zhang, Ruijie, et al.
Published: (2025)
by: Zhang, Ruijie, et al.
Published: (2025)
Similar Items
-
Embed to Control Partially Observed Systems: Representation Learning with Provable Sample Efficiency
by: Wang, Lingxiao, et al.
Published: (2022) -
Can Temporal-Difference and Q-Learning Learn Representation? A Mean-Field Theory
by: Zhang, Yufeng, et al.
Published: (2020) -
Double Duality: Variational Primal-Dual Policy Optimization for Constrained Reinforcement Learning
by: Li, Zihao, et al.
Published: (2024) -
Exchangeable Sequence Models Quantify Uncertainty Over Latent Concepts
by: Ye, Naimeng, et al.
Published: (2024) -
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs
by: Zhang, Shenao, et al.
Published: (2024)