Salvato in:
| Autori principali: | Cannella, Nick, Teh, Anzo, Han, Yanjun, Polyanskiy, Yury |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2602.15136 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Solving Empirical Bayes via Transformers
di: Teh, Anzo, et al.
Pubblicazione: (2025)
di: Teh, Anzo, et al.
Pubblicazione: (2025)
Function estimation in the empirical Bayes setting
di: Kang, Benjamin, et al.
Pubblicazione: (2026)
di: Kang, Benjamin, et al.
Pubblicazione: (2026)
The Sample Complexity of Approximate Rejection Sampling with Applications to Smoothed Online Learning
di: Block, Adam, et al.
Pubblicazione: (2023)
di: Block, Adam, et al.
Pubblicazione: (2023)
High-Rate Quantized Matrix Multiplication II
di: Ordentlich, Or, et al.
Pubblicazione: (2026)
di: Ordentlich, Or, et al.
Pubblicazione: (2026)
Optimal Quantization for Matrix Multiplication
di: Ordentlich, Or, et al.
Pubblicazione: (2024)
di: Ordentlich, Or, et al.
Pubblicazione: (2024)
Nonparametric MLE for Gaussian Location Mixtures: Certified Computation and Generic Behavior
di: Polyanskiy, Yury, et al.
Pubblicazione: (2025)
di: Polyanskiy, Yury, et al.
Pubblicazione: (2025)
On the Minimax Regret of Sequential Probability Assignment via Square-Root Entropy
di: Jia, Zeyu, et al.
Pubblicazione: (2025)
di: Jia, Zeyu, et al.
Pubblicazione: (2025)
Price of universality in vector quantization is at most 0.11 bit
di: Harbuzova, Alina, et al.
Pubblicazione: (2026)
di: Harbuzova, Alina, et al.
Pubblicazione: (2026)
Representation Alignment Rests on Linear Structure
di: Bangachev, Kiril, et al.
Pubblicazione: (2026)
di: Bangachev, Kiril, et al.
Pubblicazione: (2026)
A Gapped Scale-Sensitive Dimension and Lower Bounds for Offset Rademacher Complexity
di: Jia, Zeyu, et al.
Pubblicazione: (2025)
di: Jia, Zeyu, et al.
Pubblicazione: (2025)
Optimal empirical Bayes estimation for the Poisson model via minimum-distance methods
di: Jana, Soham, et al.
Pubblicazione: (2022)
di: Jana, Soham, et al.
Pubblicazione: (2022)
YuriiFormer: A Suite of Nesterov-Accelerated Transformers
di: Zimin, Aleksandr, et al.
Pubblicazione: (2026)
di: Zimin, Aleksandr, et al.
Pubblicazione: (2026)
WaterSIC: information-theoretically (near) optimal linear layer quantization
di: Lifar, Egor, et al.
Pubblicazione: (2026)
di: Lifar, Egor, et al.
Pubblicazione: (2026)
Synchronization of mean-field models on the circle
di: Polyanskiy, Yury, et al.
Pubblicazione: (2025)
di: Polyanskiy, Yury, et al.
Pubblicazione: (2025)
NestQuant: Nested Lattice Quantization for Matrix Products and LLMs
di: Savkin, Semyon, et al.
Pubblicazione: (2025)
di: Savkin, Semyon, et al.
Pubblicazione: (2025)
The emergence of clusters in self-attention dynamics
di: Geshkovski, Borjan, et al.
Pubblicazione: (2023)
di: Geshkovski, Borjan, et al.
Pubblicazione: (2023)
Global Minimizers of Sigmoid Contrastive Loss
di: Bangachev, Kiril, et al.
Pubblicazione: (2025)
di: Bangachev, Kiril, et al.
Pubblicazione: (2025)
Is Dimensionality a Barrier for Retrieval Models?
di: Bangachev, Kiril, et al.
Pubblicazione: (2026)
di: Bangachev, Kiril, et al.
Pubblicazione: (2026)
Measure-to-measure Regression with Transformers
di: Vandergrift, Matthew, et al.
Pubblicazione: (2026)
di: Vandergrift, Matthew, et al.
Pubblicazione: (2026)
A mathematical perspective on Transformers
di: Geshkovski, Borjan, et al.
Pubblicazione: (2023)
di: Geshkovski, Borjan, et al.
Pubblicazione: (2023)
Dynamic metastability in the self-attention model
di: Geshkovski, Borjan, et al.
Pubblicazione: (2024)
di: Geshkovski, Borjan, et al.
Pubblicazione: (2024)
Quantitative Clustering in Mean-Field Transformer Models
di: Chen, Shi, et al.
Pubblicazione: (2025)
di: Chen, Shi, et al.
Pubblicazione: (2025)
Data-driven informative priors for Bayesian inference with quasi-periodic data
di: Lopez-Santiago, Javier, et al.
Pubblicazione: (2025)
di: Lopez-Santiago, Javier, et al.
Pubblicazione: (2025)
Gradient descent inference in empirical risk minimization
di: Han, Qiyang, et al.
Pubblicazione: (2024)
di: Han, Qiyang, et al.
Pubblicazione: (2024)
Critical attention scaling in long-context transformers
di: Chen, Shi, et al.
Pubblicazione: (2025)
di: Chen, Shi, et al.
Pubblicazione: (2025)
Weak neural variational inference for solving Bayesian inverse problems without forward models: applications in elastography
di: Scholz, Vincent C., et al.
Pubblicazione: (2024)
di: Scholz, Vincent C., et al.
Pubblicazione: (2024)
Residual connections provably mitigate oversmoothing in graph neural networks
di: Chen, Ziang, et al.
Pubblicazione: (2025)
di: Chen, Ziang, et al.
Pubblicazione: (2025)
Clustering in Causal Attention Masking
di: Karagodin, Nikita, et al.
Pubblicazione: (2024)
di: Karagodin, Nikita, et al.
Pubblicazione: (2024)
Optimal score estimation via empirical Bayes smoothing
di: Wibisono, Andre, et al.
Pubblicazione: (2024)
di: Wibisono, Andre, et al.
Pubblicazione: (2024)
The Radio-Frequency Transformer for Signal Separation
di: Lifar, Egor, et al.
Pubblicazione: (2026)
di: Lifar, Egor, et al.
Pubblicazione: (2026)
Continuous First, Discrete Later: VQ-VAEs Without Dimensional Collapse
di: Zhao, Xinyu, et al.
Pubblicazione: (2026)
di: Zhao, Xinyu, et al.
Pubblicazione: (2026)
Scaling Limits of Long-Context Transformers
di: Bruno, Giuseppe, et al.
Pubblicazione: (2026)
di: Bruno, Giuseppe, et al.
Pubblicazione: (2026)
The effectiveness of MAE pre-pretraining for billion-scale pretraining
di: Singh, Mannat, et al.
Pubblicazione: (2023)
di: Singh, Mannat, et al.
Pubblicazione: (2023)
Minimax optimal testing by classification
di: Gerber, Patrik Róbert, et al.
Pubblicazione: (2023)
di: Gerber, Patrik Róbert, et al.
Pubblicazione: (2023)
Evolution Strategies for Deep RL pretraining
di: Martínez, Adrian, et al.
Pubblicazione: (2026)
di: Martínez, Adrian, et al.
Pubblicazione: (2026)
Synthetic continued pretraining
di: Yang, Zitong, et al.
Pubblicazione: (2024)
di: Yang, Zitong, et al.
Pubblicazione: (2024)
Learning to solve Bayesian inverse problems: An amortized variational inference approach using Gaussian and Flow guides
di: Karumuri, Sharmila, et al.
Pubblicazione: (2023)
di: Karumuri, Sharmila, et al.
Pubblicazione: (2023)
Interactive Learning of Single-Index Models via Stochastic Gradient Descent
di: Rajaraman, Nived, et al.
Pubblicazione: (2026)
di: Rajaraman, Nived, et al.
Pubblicazione: (2026)
On Uniform, Bayesian, and PAC-Bayesian Deep Ensembles
di: Hauptvogel, Nick, et al.
Pubblicazione: (2024)
di: Hauptvogel, Nick, et al.
Pubblicazione: (2024)
Normalization in Attention Dynamics
di: Karagodin, Nikita, et al.
Pubblicazione: (2025)
di: Karagodin, Nikita, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Solving Empirical Bayes via Transformers
di: Teh, Anzo, et al.
Pubblicazione: (2025) -
Function estimation in the empirical Bayes setting
di: Kang, Benjamin, et al.
Pubblicazione: (2026) -
The Sample Complexity of Approximate Rejection Sampling with Applications to Smoothed Online Learning
di: Block, Adam, et al.
Pubblicazione: (2023) -
High-Rate Quantized Matrix Multiplication II
di: Ordentlich, Or, et al.
Pubblicazione: (2026) -
Optimal Quantization for Matrix Multiplication
di: Ordentlich, Or, et al.
Pubblicazione: (2024)