Transformers Simulate MLE for Sequence Generation in Bayesian Networks
Fuente:
arXiv
Salvato in:
| Autori principali: | Cao, Yuan, He, Yihan, Wu, Dennis, Chen, Hong-Yu, Fan, Jianqing, Liu, Han |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Learning Spectral Methods by Transformers
di: He, Yihan, et al.
Pubblicazione: (2025)
di: He, Yihan, et al.
Pubblicazione: (2025)
Transformers versus the EM Algorithm in Multi-class Clustering
di: He, Yihan, et al.
Pubblicazione: (2025)
di: He, Yihan, et al.
Pubblicazione: (2025)
Transformers and Their Roles as Time Series Foundation Models
di: Wu, Dennis, et al.
Pubblicazione: (2025)
di: Wu, Dennis, et al.
Pubblicazione: (2025)
Uncertainty Quantification of MLE for Entity Ranking with Covariates
di: Fan, Jianqing, et al.
Pubblicazione: (2022)
di: Fan, Jianqing, et al.
Pubblicazione: (2022)
Global Convergence in Training Large-Scale Transformers
di: Gao, Cheng, et al.
Pubblicazione: (2024)
di: Gao, Cheng, et al.
Pubblicazione: (2024)
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
di: Li, Zihao, et al.
Pubblicazione: (2024)
di: Li, Zihao, et al.
Pubblicazione: (2024)
MLE-Smith: Scaling MLE Tasks with Automated Multi-Agent Pipeline
di: Qiang, Rushi, et al.
Pubblicazione: (2025)
di: Qiang, Rushi, et al.
Pubblicazione: (2025)
Benign Overfitting in Out-of-Distribution Generalization of Linear Models
di: Tang, Shange, et al.
Pubblicazione: (2024)
di: Tang, Shange, et al.
Pubblicazione: (2024)
Graph External Attention Enhanced Transformer
di: Liang, Jianqing, et al.
Pubblicazione: (2024)
di: Liang, Jianqing, et al.
Pubblicazione: (2024)
FigBO: A Generalized Acquisition Function Framework with Look-Ahead Capability for Bayesian Optimization
di: Chen, Hui, et al.
Pubblicazione: (2025)
di: Chen, Hui, et al.
Pubblicazione: (2025)
Coordinate ascent neural Kalman-MLE for state estimation
di: Hanlon, Bettina, et al.
Pubblicazione: (2025)
di: Hanlon, Bettina, et al.
Pubblicazione: (2025)
When Models Don't Collapse: On the Consistency of Iterative MLE
di: Barzilai, Daniel, et al.
Pubblicazione: (2025)
di: Barzilai, Daniel, et al.
Pubblicazione: (2025)
Nonparametric MLE for Gaussian Location Mixtures: Certified Computation and Generic Behavior
di: Polyanskiy, Yury, et al.
Pubblicazione: (2025)
di: Polyanskiy, Yury, et al.
Pubblicazione: (2025)
Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning
di: Chai, Jinhang, et al.
Pubblicazione: (2025)
di: Chai, Jinhang, et al.
Pubblicazione: (2025)
AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
di: Toledo, Edan, et al.
Pubblicazione: (2025)
di: Toledo, Edan, et al.
Pubblicazione: (2025)
Covariates-Adjusted Mixed-Membership Estimation: A Novel Network Model with Optimal Guarantees
di: Fan, Jianqing, et al.
Pubblicazione: (2025)
di: Fan, Jianqing, et al.
Pubblicazione: (2025)
An Overview of Diffusion Models: Applications, Guided Generation, Statistical Rates and Optimization
di: Chen, Minshuo, et al.
Pubblicazione: (2024)
di: Chen, Minshuo, et al.
Pubblicazione: (2024)
MLE-STAR: Machine Learning Engineering Agent via Search and Targeted Refinement
di: Nam, Jaehyun, et al.
Pubblicazione: (2025)
di: Nam, Jaehyun, et al.
Pubblicazione: (2025)
Asymptotic Theory of Eigenvectors for Latent Embeddings with Generalized Laplacian Matrices
di: Fan, Jianqing, et al.
Pubblicazione: (2025)
di: Fan, Jianqing, et al.
Pubblicazione: (2025)
Reasoning as Gradient: Scaling MLE Agents Beyond Tree Search
di: Zhang, Yifei, et al.
Pubblicazione: (2026)
di: Zhang, Yifei, et al.
Pubblicazione: (2026)
Marked Temporal Bayesian Flow Point Processes
di: Chen, Hui, et al.
Pubblicazione: (2024)
di: Chen, Hui, et al.
Pubblicazione: (2024)
Structured Matrix Learning under Arbitrary Entrywise Dependence and Estimation of Markov Transition Kernel
di: Chai, Jinhang, et al.
Pubblicazione: (2024)
di: Chai, Jinhang, et al.
Pubblicazione: (2024)
ParamReL: Learning Parameter Space Representation via Progressively Encoding Bayesian Flow Networks
di: Wu, Zhangkai, et al.
Pubblicazione: (2024)
di: Wu, Zhangkai, et al.
Pubblicazione: (2024)
k-MLE, k-Bregman, k-VARs: Theory, Convergence, Computation
di: Yue, Zuogong, et al.
Pubblicazione: (2024)
di: Yue, Zuogong, et al.
Pubblicazione: (2024)
Flatten Graphs as Sequences: Transformers are Scalable Graph Generators
di: Chen, Dexiong, et al.
Pubblicazione: (2025)
di: Chen, Dexiong, et al.
Pubblicazione: (2025)
Covariate Assisted Entity Ranking with Sparse Intrinsic Scores
di: Fan, Jianqing, et al.
Pubblicazione: (2024)
di: Fan, Jianqing, et al.
Pubblicazione: (2024)
Variational Bayesian Flow Network for Graph Generation
di: Xiong, Yida, et al.
Pubblicazione: (2026)
di: Xiong, Yida, et al.
Pubblicazione: (2026)
Spectral Ranking Inferences based on General Multiway Comparisons
di: Fan, Jianqing, et al.
Pubblicazione: (2023)
di: Fan, Jianqing, et al.
Pubblicazione: (2023)
MLE-Dojo: Interactive Environments for Empowering LLM Agents in Machine Learning Engineering
di: Qiang, Rushi, et al.
Pubblicazione: (2025)
di: Qiang, Rushi, et al.
Pubblicazione: (2025)
GNN-Transformer Cooperative Architecture for Trustworthy Graph Contrastive Learning
di: Liang, Jianqing, et al.
Pubblicazione: (2024)
di: Liang, Jianqing, et al.
Pubblicazione: (2024)
Exploring Topological Bias in Heterogeneous Graph Neural Networks
di: Zhang, Yihan
Pubblicazione: (2025)
di: Zhang, Yihan
Pubblicazione: (2025)
Understanding the Benefits of SimCLR Pre-Training in Two-Layer Convolutional Neural Networks
di: Zhang, Han, et al.
Pubblicazione: (2024)
di: Zhang, Han, et al.
Pubblicazione: (2024)
Feature Augmentations for High-Dimensional Learning
di: Zhu, Xiaonan, et al.
Pubblicazione: (2025)
di: Zhu, Xiaonan, et al.
Pubblicazione: (2025)
Beyond MLE: Investigating SEARNN for Low-Resourced Neural Machine Translation
di: Emezue, Chris
Pubblicazione: (2024)
di: Emezue, Chris
Pubblicazione: (2024)
Mini-Sequence Transformer: Optimizing Intermediate Memory for Long Sequences Training
di: Luo, Cheng, et al.
Pubblicazione: (2024)
di: Luo, Cheng, et al.
Pubblicazione: (2024)
TAH-QUANT: Effective Activation Quantization in Pipeline Parallelism over Slow Network
di: He, Guangxin, et al.
Pubblicazione: (2025)
di: He, Guangxin, et al.
Pubblicazione: (2025)
Inference for Heteroskedastic PCA with Missing Data
di: Yan, Yuling, et al.
Pubblicazione: (2021)
di: Yan, Yuling, et al.
Pubblicazione: (2021)
Looped Transformers with Layer Normalization Provably Learn the Power Method
di: Wu, Lyumin, et al.
Pubblicazione: (2026)
di: Wu, Lyumin, et al.
Pubblicazione: (2026)
Dynamic Graph Unlearning: A General and Efficient Post-Processing Method via Gradient Transformation
di: Zhang, He, et al.
Pubblicazione: (2024)
di: Zhang, He, et al.
Pubblicazione: (2024)
Random pairing MLE for estimation of item parameters in Rasch model
di: Yang, Yuepeng, et al.
Pubblicazione: (2024)
di: Yang, Yuepeng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Learning Spectral Methods by Transformers
di: He, Yihan, et al.
Pubblicazione: (2025) -
Transformers versus the EM Algorithm in Multi-class Clustering
di: He, Yihan, et al.
Pubblicazione: (2025) -
Transformers and Their Roles as Time Series Foundation Models
di: Wu, Dennis, et al.
Pubblicazione: (2025) -
Uncertainty Quantification of MLE for Entity Ranking with Covariates
di: Fan, Jianqing, et al.
Pubblicazione: (2022) -
Global Convergence in Training Large-Scale Transformers
di: Gao, Cheng, et al.
Pubblicazione: (2024)