Measure-to-measure Regression with Transformers
Fuente:
arXiv
Salvato in:
| Autori principali: | Vandergrift, Matthew, White, Martha, Polyanskiy, Yury, Rigollet, Philippe, Atanackovic, Lazar |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
YuriiFormer: A Suite of Nesterov-Accelerated Transformers
di: Zimin, Aleksandr, et al.
Pubblicazione: (2026)
di: Zimin, Aleksandr, et al.
Pubblicazione: (2026)
A mathematical perspective on Transformers
di: Geshkovski, Borjan, et al.
Pubblicazione: (2023)
di: Geshkovski, Borjan, et al.
Pubblicazione: (2023)
Quantitative Clustering in Mean-Field Transformer Models
di: Chen, Shi, et al.
Pubblicazione: (2025)
di: Chen, Shi, et al.
Pubblicazione: (2025)
Synchronization of mean-field models on the circle
di: Polyanskiy, Yury, et al.
Pubblicazione: (2025)
di: Polyanskiy, Yury, et al.
Pubblicazione: (2025)
The emergence of clusters in self-attention dynamics
di: Geshkovski, Borjan, et al.
Pubblicazione: (2023)
di: Geshkovski, Borjan, et al.
Pubblicazione: (2023)
Scaling Limits of Long-Context Transformers
di: Bruno, Giuseppe, et al.
Pubblicazione: (2026)
di: Bruno, Giuseppe, et al.
Pubblicazione: (2026)
Dynamic metastability in the self-attention model
di: Geshkovski, Borjan, et al.
Pubblicazione: (2024)
di: Geshkovski, Borjan, et al.
Pubblicazione: (2024)
Clustering in Causal Attention Masking
di: Karagodin, Nikita, et al.
Pubblicazione: (2024)
di: Karagodin, Nikita, et al.
Pubblicazione: (2024)
Critical attention scaling in long-context transformers
di: Chen, Shi, et al.
Pubblicazione: (2025)
di: Chen, Shi, et al.
Pubblicazione: (2025)
Residual connections provably mitigate oversmoothing in graph neural networks
di: Chen, Ziang, et al.
Pubblicazione: (2025)
di: Chen, Ziang, et al.
Pubblicazione: (2025)
Investigating Generalization Behaviours of Generative Flow Networks
di: Atanackovic, Lazar, et al.
Pubblicazione: (2024)
di: Atanackovic, Lazar, et al.
Pubblicazione: (2024)
Normalization in Attention Dynamics
di: Karagodin, Nikita, et al.
Pubblicazione: (2025)
di: Karagodin, Nikita, et al.
Pubblicazione: (2025)
Splat Regression Models
di: Daniels, Mara, et al.
Pubblicazione: (2025)
di: Daniels, Mara, et al.
Pubblicazione: (2025)
Measure-to-measure interpolation using Transformers
di: Geshkovski, Borjan, et al.
Pubblicazione: (2024)
di: Geshkovski, Borjan, et al.
Pubblicazione: (2024)
The Mean-Field Dynamics of Transformers
di: Rigollet, Philippe
Pubblicazione: (2025)
di: Rigollet, Philippe
Pubblicazione: (2025)
Solving Empirical Bayes via Transformers
di: Teh, Anzo, et al.
Pubblicazione: (2025)
di: Teh, Anzo, et al.
Pubblicazione: (2025)
Homogenized Transformers
di: Koubbi, Hugo, et al.
Pubblicazione: (2026)
di: Koubbi, Hugo, et al.
Pubblicazione: (2026)
The Sample Complexity of Approximate Rejection Sampling with Applications to Smoothed Online Learning
di: Block, Adam, et al.
Pubblicazione: (2023)
di: Block, Adam, et al.
Pubblicazione: (2023)
High-Rate Quantized Matrix Multiplication II
di: Ordentlich, Or, et al.
Pubblicazione: (2026)
di: Ordentlich, Or, et al.
Pubblicazione: (2026)
A Call to Lagrangian Action: Learning Population Mechanics from Temporal Snapshots
di: Guan, Vincent, et al.
Pubblicazione: (2026)
di: Guan, Vincent, et al.
Pubblicazione: (2026)
Nonparametric MLE for Gaussian Location Mixtures: Certified Computation and Generic Behavior
di: Polyanskiy, Yury, et al.
Pubblicazione: (2025)
di: Polyanskiy, Yury, et al.
Pubblicazione: (2025)
Optimal Quantization for Matrix Multiplication
di: Ordentlich, Or, et al.
Pubblicazione: (2024)
di: Ordentlich, Or, et al.
Pubblicazione: (2024)
The Superposition of Diffusion Models Using the Itô Density Estimator
di: Skreta, Marta, et al.
Pubblicazione: (2024)
di: Skreta, Marta, et al.
Pubblicazione: (2024)
Price of universality in vector quantization is at most 0.11 bit
di: Harbuzova, Alina, et al.
Pubblicazione: (2026)
di: Harbuzova, Alina, et al.
Pubblicazione: (2026)
The power of fine-grained experts: Granularity boosts expressivity in Mixture of Experts
di: Boix-Adsera, Enric, et al.
Pubblicazione: (2025)
di: Boix-Adsera, Enric, et al.
Pubblicazione: (2025)
Representation Alignment Rests on Linear Structure
di: Bangachev, Kiril, et al.
Pubblicazione: (2026)
di: Bangachev, Kiril, et al.
Pubblicazione: (2026)
On the Minimax Regret of Sequential Probability Assignment via Square-Root Entropy
di: Jia, Zeyu, et al.
Pubblicazione: (2025)
di: Jia, Zeyu, et al.
Pubblicazione: (2025)
A Gapped Scale-Sensitive Dimension and Lower Bounds for Offset Rademacher Complexity
di: Jia, Zeyu, et al.
Pubblicazione: (2025)
di: Jia, Zeyu, et al.
Pubblicazione: (2025)
Gaussian mixture layers for neural networks
di: Chewi, Sinho, et al.
Pubblicazione: (2025)
di: Chewi, Sinho, et al.
Pubblicazione: (2025)
On the number of modes of Gaussian kernel density estimators
di: Geshkovski, Borjan, et al.
Pubblicazione: (2024)
di: Geshkovski, Borjan, et al.
Pubblicazione: (2024)
The Radio-Frequency Transformer for Signal Separation
di: Lifar, Egor, et al.
Pubblicazione: (2026)
di: Lifar, Egor, et al.
Pubblicazione: (2026)
A Computational Framework for Solving Wasserstein Lagrangian Flows
di: Neklyudov, Kirill, et al.
Pubblicazione: (2023)
di: Neklyudov, Kirill, et al.
Pubblicazione: (2023)
Universal priors: solving empirical Bayes via Bayesian inference and pretraining
di: Cannella, Nick, et al.
Pubblicazione: (2026)
di: Cannella, Nick, et al.
Pubblicazione: (2026)
WaterSIC: information-theoretically (near) optimal linear layer quantization
di: Lifar, Egor, et al.
Pubblicazione: (2026)
di: Lifar, Egor, et al.
Pubblicazione: (2026)
Statistical optimal transport
di: Chewi, Sinho, et al.
Pubblicazione: (2024)
di: Chewi, Sinho, et al.
Pubblicazione: (2024)
Global Minimizers of Sigmoid Contrastive Loss
di: Bangachev, Kiril, et al.
Pubblicazione: (2025)
di: Bangachev, Kiril, et al.
Pubblicazione: (2025)
NestQuant: Nested Lattice Quantization for Matrix Products and LLMs
di: Savkin, Semyon, et al.
Pubblicazione: (2025)
di: Savkin, Semyon, et al.
Pubblicazione: (2025)
Simulation-free Schrödinger bridges via score and flow matching
di: Tong, Alexander, et al.
Pubblicazione: (2023)
di: Tong, Alexander, et al.
Pubblicazione: (2023)
Is Dimensionality a Barrier for Retrieval Models?
di: Bangachev, Kiril, et al.
Pubblicazione: (2026)
di: Bangachev, Kiril, et al.
Pubblicazione: (2026)
Meta Flow Matching: Integrating Vector Fields on the Wasserstein Manifold
di: Atanackovic, Lazar, et al.
Pubblicazione: (2024)
di: Atanackovic, Lazar, et al.
Pubblicazione: (2024)
Documenti analoghi
-
YuriiFormer: A Suite of Nesterov-Accelerated Transformers
di: Zimin, Aleksandr, et al.
Pubblicazione: (2026) -
A mathematical perspective on Transformers
di: Geshkovski, Borjan, et al.
Pubblicazione: (2023) -
Quantitative Clustering in Mean-Field Transformer Models
di: Chen, Shi, et al.
Pubblicazione: (2025) -
Synchronization of mean-field models on the circle
di: Polyanskiy, Yury, et al.
Pubblicazione: (2025) -
The emergence of clusters in self-attention dynamics
di: Geshkovski, Borjan, et al.
Pubblicazione: (2023)