Transformers know more than they can tell -- Learning the Collatz sequence
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Charton, François, Narayanan, Ashvni |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning Euler Factors of Elliptic Curves
von: Babei, Angelica, et al.
Veröffentlicht: (2025)
von: Babei, Angelica, et al.
Veröffentlicht: (2025)
Learning the greatest common divisor: explaining transformer predictions
von: Charton, François
Veröffentlicht: (2023)
von: Charton, François
Veröffentlicht: (2023)
Int2Int: a framework for mathematics with transformers
von: Charton, François
Veröffentlicht: (2025)
von: Charton, François
Veröffentlicht: (2025)
Emergent properties with repeated examples
von: Charton, François, et al.
Veröffentlicht: (2024)
von: Charton, François, et al.
Veröffentlicht: (2024)
Revealing economic facts: LLMs know more than they say
von: Buckmann, Marcus, et al.
Veröffentlicht: (2025)
von: Buckmann, Marcus, et al.
Veröffentlicht: (2025)
Global Lyapunov functions: a long-standing open problem in mathematics, with symbolic transformers
von: Alfarano, Alberto, et al.
Veröffentlicht: (2024)
von: Alfarano, Alberto, et al.
Veröffentlicht: (2024)
Transforming Calabi-Yau Constructions: Generating New Calabi-Yau Manifolds with Transformers
von: Yip, Jacky H. T., et al.
Veröffentlicht: (2025)
von: Yip, Jacky H. T., et al.
Veröffentlicht: (2025)
Increasing biases can be more efficient than increasing weights
von: Metta, Carlo, et al.
Veröffentlicht: (2023)
von: Metta, Carlo, et al.
Veröffentlicht: (2023)
TAPAS: Datasets for Learning the Learning with Errors Problem
von: Saxena, Eshika, et al.
Veröffentlicht: (2025)
von: Saxena, Eshika, et al.
Veröffentlicht: (2025)
Learning Regularizers: Learning Optimizers that can Regularize
von: Sahoo, Suraj Kumar, et al.
Veröffentlicht: (2025)
von: Sahoo, Suraj Kumar, et al.
Veröffentlicht: (2025)
Instruction Diversity Drives Generalization To Unseen Tasks
von: Zhang, Dylan, et al.
Veröffentlicht: (2024)
von: Zhang, Dylan, et al.
Veröffentlicht: (2024)
$\textbf{Only-IF}$:Revealing the Decisive Effect of Instruction Diversity on Generalization
von: Zhang, Dylan, et al.
Veröffentlicht: (2024)
von: Zhang, Dylan, et al.
Veröffentlicht: (2024)
Extrapolating Jet Radiation with Autoregressive Transformers
von: Butter, Anja, et al.
Veröffentlicht: (2024)
von: Butter, Anja, et al.
Veröffentlicht: (2024)
Context information can be more important than reasoning for time series forecasting with a large language model
von: Yang, Janghoon
Veröffentlicht: (2025)
von: Yang, Janghoon
Veröffentlicht: (2025)
There is more to graphs than meets the eye: Learning universal features with self-supervision
von: Das, Laya, et al.
Veröffentlicht: (2023)
von: Das, Laya, et al.
Veröffentlicht: (2023)
From Symbolic Tasks to Code Generation: Diversification Yields Better Task Performers
von: Zhang, Dylan, et al.
Veröffentlicht: (2024)
von: Zhang, Dylan, et al.
Veröffentlicht: (2024)
Mixture of Parrots: Experts improve memorization more than reasoning
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
Improving ML Attacks on LWE with Data Repetition and Stepwise Regression
von: Alfarano, Alberto, et al.
Veröffentlicht: (2026)
von: Alfarano, Alberto, et al.
Veröffentlicht: (2026)
Beyond Model Collapse: Scaling Up with Synthesized Data Requires Verification
von: Feng, Yunzhen, et al.
Veröffentlicht: (2024)
von: Feng, Yunzhen, et al.
Veröffentlicht: (2024)
Weight Decay may matter more than muP for Learning Rate Transfer in Practice
von: Kosson, Atli, et al.
Veröffentlicht: (2025)
von: Kosson, Atli, et al.
Veröffentlicht: (2025)
Painful intelligence: What AI can tell us about human suffering
von: Hyvärinen, Aapo
Veröffentlicht: (2022)
von: Hyvärinen, Aapo
Veröffentlicht: (2022)
Ask more, know better: Reinforce-Learned Prompt Questions for Decision Making with Large Language Models
von: Yan, Xue, et al.
Veröffentlicht: (2023)
von: Yan, Xue, et al.
Veröffentlicht: (2023)
Are uGLAD? Time will tell!
von: Imani, Shima, et al.
Veröffentlicht: (2023)
von: Imani, Shima, et al.
Veröffentlicht: (2023)
Salsa Fresca: Angular Embeddings and Pre-Training for ML Attacks on Learning With Errors
von: Stevens, Samuel, et al.
Veröffentlicht: (2024)
von: Stevens, Samuel, et al.
Veröffentlicht: (2024)
Smoothed Online Classification can be Harder than Batch Classification
von: Raman, Vinod, et al.
Veröffentlicht: (2024)
von: Raman, Vinod, et al.
Veröffentlicht: (2024)
Making Hard Problems Easier with Custom Data Distributions and Loss Regularization: A Case Study in Modular Arithmetic
von: Saxena, Eshika, et al.
Veröffentlicht: (2024)
von: Saxena, Eshika, et al.
Veröffentlicht: (2024)
Regularization can make diffusion models more efficient
von: Taheri, Mahsa, et al.
Veröffentlicht: (2025)
von: Taheri, Mahsa, et al.
Veröffentlicht: (2025)
In-Context Learning can distort the relationship between sequence likelihoods and biological fitness
von: Kantroo, Pranav, et al.
Veröffentlicht: (2025)
von: Kantroo, Pranav, et al.
Veröffentlicht: (2025)
A Tale of Tails: Model Collapse as a Change of Scaling Laws
von: Dohmatob, Elvis, et al.
Veröffentlicht: (2024)
von: Dohmatob, Elvis, et al.
Veröffentlicht: (2024)
The boosted HP filter is more general than you might think
von: Mei, Ziwei, et al.
Veröffentlicht: (2022)
von: Mei, Ziwei, et al.
Veröffentlicht: (2022)
Less can be more for predicting properties with large language models
von: Alampara, Nawaf, et al.
Veröffentlicht: (2024)
von: Alampara, Nawaf, et al.
Veröffentlicht: (2024)
Mixer is more than just a model
von: Ji, Qingfeng, et al.
Veröffentlicht: (2024)
von: Ji, Qingfeng, et al.
Veröffentlicht: (2024)
The Free Transformer
von: Fleuret, François
Veröffentlicht: (2025)
von: Fleuret, François
Veröffentlicht: (2025)
On Limitations of the Transformer Architecture
von: Peng, Binghui, et al.
Veröffentlicht: (2024)
von: Peng, Binghui, et al.
Veröffentlicht: (2024)
Iteration Head: A Mechanistic Study of Chain-of-Thought
von: Cabannes, Vivien, et al.
Veröffentlicht: (2024)
von: Cabannes, Vivien, et al.
Veröffentlicht: (2024)
Transforming the Bootstrap: Using Transformers to Compute Scattering Amplitudes in Planar N = 4 Super Yang-Mills Theory
von: Cai, Tianji, et al.
Veröffentlicht: (2024)
von: Cai, Tianji, et al.
Veröffentlicht: (2024)
Quality In / Quality Out: Data quality more relevant than model choice in anomaly detection with the UGR'16
von: Camacho, José, et al.
Veröffentlicht: (2023)
von: Camacho, José, et al.
Veröffentlicht: (2023)
Phase Transition for Stochastic Block Model with more than $\sqrt{n}$ Communities
von: Carpentier, Alexandra, et al.
Veröffentlicht: (2025)
von: Carpentier, Alexandra, et al.
Veröffentlicht: (2025)
Phase Transition for Stochastic Block Model with more than $\sqrt{n}$ Communities (II)
von: Carpentier, Alexandra, et al.
Veröffentlicht: (2025)
von: Carpentier, Alexandra, et al.
Veröffentlicht: (2025)
Estimating the expected output of wide random MLPs more efficiently than sampling
von: Wu, Wilson, et al.
Veröffentlicht: (2026)
von: Wu, Wilson, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Learning Euler Factors of Elliptic Curves
von: Babei, Angelica, et al.
Veröffentlicht: (2025) -
Learning the greatest common divisor: explaining transformer predictions
von: Charton, François
Veröffentlicht: (2023) -
Int2Int: a framework for mathematics with transformers
von: Charton, François
Veröffentlicht: (2025) -
Emergent properties with repeated examples
von: Charton, François, et al.
Veröffentlicht: (2024) -
Revealing economic facts: LLMs know more than they say
von: Buckmann, Marcus, et al.
Veröffentlicht: (2025)