Gaussian Match-and-Copy: A Minimalist Benchmark for Studying Transformer Induction
Fuente:
arXiv
Salvato in:
| Autori principali: | Gonon, Antoine, Cordonnier, Alexandre, Boumal, Nicolas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Insights on Muon from Simple Quadratics
di: Gonon, Antoine, et al.
Pubblicazione: (2026)
di: Gonon, Antoine, et al.
Pubblicazione: (2026)
Induction Signatures Are Not Enough: A Matched-Compute Study of Load-Bearing Structure in In-Context Learning
di: Sabry, Mohammed, et al.
Pubblicazione: (2025)
di: Sabry, Mohammed, et al.
Pubblicazione: (2025)
A Minimalist Method for Fine-tuning Text-to-Image Diffusion Models
di: Miao, Yanting, et al.
Pubblicazione: (2025)
di: Miao, Yanting, et al.
Pubblicazione: (2025)
MINTS: Minimalist Thompson Sampling
di: Wang, Kaizheng
Pubblicazione: (2026)
di: Wang, Kaizheng
Pubblicazione: (2026)
ReplaySCM: A Benchmark for Executable Causal Mechanism Induction from Interventions
di: Batzoglou, Serafim
Pubblicazione: (2026)
di: Batzoglou, Serafim
Pubblicazione: (2026)
Minimalist Softmax Attention Provably Learns Constrained Boolean Functions
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2025)
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2025)
Sparsity in neural networks can improve their privacy
di: Gonon, Antoine, et al.
Pubblicazione: (2023)
di: Gonon, Antoine, et al.
Pubblicazione: (2023)
A Minimalist Bayesian Framework for Stochastic Optimization
di: Wang, Kaizheng
Pubblicazione: (2025)
di: Wang, Kaizheng
Pubblicazione: (2025)
Repeat After Me: Transformers are Better than State Space Models at Copying
di: Jelassi, Samy, et al.
Pubblicazione: (2024)
di: Jelassi, Samy, et al.
Pubblicazione: (2024)
Context is Key: A Benchmark for Forecasting with Essential Textual Information
di: Williams, Andrew Robert, et al.
Pubblicazione: (2024)
di: Williams, Andrew Robert, et al.
Pubblicazione: (2024)
A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
di: Xiong, Wei, et al.
Pubblicazione: (2025)
di: Xiong, Wei, et al.
Pubblicazione: (2025)
Improving Generalization by Permutation Routing Across Model Copies
di: Kashiwamura, Shuhei, et al.
Pubblicazione: (2026)
di: Kashiwamura, Shuhei, et al.
Pubblicazione: (2026)
SKADA-Bench: Benchmarking Unsupervised Domain Adaptation Methods with Realistic Validation On Diverse Modalities
di: Lalou, Yanis, et al.
Pubblicazione: (2024)
di: Lalou, Yanis, et al.
Pubblicazione: (2024)
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
di: Glentis, Athanasios, et al.
Pubblicazione: (2025)
di: Glentis, Athanasios, et al.
Pubblicazione: (2025)
Flow Matching with Gaussian Process Priors for Probabilistic Time Series Forecasting
di: Kollovieh, Marcel, et al.
Pubblicazione: (2024)
di: Kollovieh, Marcel, et al.
Pubblicazione: (2024)
Discrete Diffusion Schrödinger Bridge Matching for Graph Transformation
di: Kim, Jun Hyeong, et al.
Pubblicazione: (2024)
di: Kim, Jun Hyeong, et al.
Pubblicazione: (2024)
Language Models "Grok" to Copy
di: Lv, Ang, et al.
Pubblicazione: (2024)
di: Lv, Ang, et al.
Pubblicazione: (2024)
Optimal Transport for Domain Adaptation through Gaussian Mixture Models
di: Montesuma, Eduardo Fernandes, et al.
Pubblicazione: (2024)
di: Montesuma, Eduardo Fernandes, et al.
Pubblicazione: (2024)
Equalized Generative Treatment: Matching f-divergences for Fairness in Generative Models
di: Verine, Alexandre, et al.
Pubblicazione: (2026)
di: Verine, Alexandre, et al.
Pubblicazione: (2026)
A Benchmark Study on Calibration
di: Tao, Linwei, et al.
Pubblicazione: (2023)
di: Tao, Linwei, et al.
Pubblicazione: (2023)
Foundation Models and Transformers for Anomaly Detection: A Survey
di: Ammar, Mouïn Ben, et al.
Pubblicazione: (2025)
di: Ammar, Mouïn Ben, et al.
Pubblicazione: (2025)
What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov Chains
di: Ekbote, Chanakya, et al.
Pubblicazione: (2025)
di: Ekbote, Chanakya, et al.
Pubblicazione: (2025)
Keep Everyone Happy: Online Fair Division of Numerous Items with Few Copies
di: Verma, Arun, et al.
Pubblicazione: (2024)
di: Verma, Arun, et al.
Pubblicazione: (2024)
Program-Based Strategy Induction for Reinforcement Learning
di: Correa, Carlos G., et al.
Pubblicazione: (2024)
di: Correa, Carlos G., et al.
Pubblicazione: (2024)
Modeling Matches as Language: A Generative Transformer Approach for Counterfactual Player Valuation in Football
di: Hong, Miru, et al.
Pubblicazione: (2026)
di: Hong, Miru, et al.
Pubblicazione: (2026)
Benchmarking Positional Encodings for GNNs and Graph Transformers
di: Grötschla, Florian, et al.
Pubblicazione: (2024)
di: Grötschla, Florian, et al.
Pubblicazione: (2024)
Using Artificial Intuition in Distinct, Minimalist Classification of Scientific Abstracts for Management of Technology Portfolios
di: Ranka, Prateek, et al.
Pubblicazione: (2025)
di: Ranka, Prateek, et al.
Pubblicazione: (2025)
Differentiable Rule Induction from Raw Sequence Inputs
di: Gao, Kun, et al.
Pubblicazione: (2026)
di: Gao, Kun, et al.
Pubblicazione: (2026)
Weight Copy and Low-Rank Adaptation for Few-Shot Distillation of Vision Transformers
di: Grigore, Diana-Nicoleta, et al.
Pubblicazione: (2024)
di: Grigore, Diana-Nicoleta, et al.
Pubblicazione: (2024)
BIRD: Behavior Induction via Representation-structure Distillation
di: Pogoncheff, Galen, et al.
Pubblicazione: (2025)
di: Pogoncheff, Galen, et al.
Pubblicazione: (2025)
Amortized Active Causal Induction with Deep Reinforcement Learning
di: Annadani, Yashas, et al.
Pubblicazione: (2024)
di: Annadani, Yashas, et al.
Pubblicazione: (2024)
The Value of Covariance Matching in Gaussian DDPMs and the Lanczos Sampler
di: Akhtar, Md Sahil, et al.
Pubblicazione: (2026)
di: Akhtar, Md Sahil, et al.
Pubblicazione: (2026)
Nano World Models: A Minimalist Implementation of Future Video Prediction
di: Huang, Siqiao, et al.
Pubblicazione: (2026)
di: Huang, Siqiao, et al.
Pubblicazione: (2026)
Task Agnostic Architecture for Algorithm Induction via Implicit Composition
di: Sindhi, Sahil J., et al.
Pubblicazione: (2024)
di: Sindhi, Sahil J., et al.
Pubblicazione: (2024)
Function Induction and Task Generalization: An Interpretability Study with Off-by-One Addition
di: Ye, Qinyuan, et al.
Pubblicazione: (2025)
di: Ye, Qinyuan, et al.
Pubblicazione: (2025)
Bayesian Inverse Problems Meet Flow Matching: Efficient and Flexible Inference via Transformers
di: Sherki, Daniil, et al.
Pubblicazione: (2025)
di: Sherki, Daniil, et al.
Pubblicazione: (2025)
Knot So Simple: A Minimalistic Environment for Spatial Reasoning
di: Chen, Zizhao, et al.
Pubblicazione: (2025)
di: Chen, Zizhao, et al.
Pubblicazione: (2025)
Deep Reinforcement Learning for Personalized Diagnostic Decision Pathways Using Electronic Health Records: A Comparative Study on Anemia and Systemic Lupus Erythematosus
di: Muyama, Lillian, et al.
Pubblicazione: (2024)
di: Muyama, Lillian, et al.
Pubblicazione: (2024)
Fault Analysis And Predictive Maintenance Of Induction Motor Using Machine Learning
di: Venkatesh, Kavana, et al.
Pubblicazione: (2024)
di: Venkatesh, Kavana, et al.
Pubblicazione: (2024)
Reliable Evaluation and Benchmarks for Statement Autoformalization
di: Poiroux, Auguste, et al.
Pubblicazione: (2024)
di: Poiroux, Auguste, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Insights on Muon from Simple Quadratics
di: Gonon, Antoine, et al.
Pubblicazione: (2026) -
Induction Signatures Are Not Enough: A Matched-Compute Study of Load-Bearing Structure in In-Context Learning
di: Sabry, Mohammed, et al.
Pubblicazione: (2025) -
A Minimalist Method for Fine-tuning Text-to-Image Diffusion Models
di: Miao, Yanting, et al.
Pubblicazione: (2025) -
MINTS: Minimalist Thompson Sampling
di: Wang, Kaizheng
Pubblicazione: (2026) -
ReplaySCM: A Benchmark for Executable Causal Mechanism Induction from Interventions
di: Batzoglou, Serafim
Pubblicazione: (2026)