(How) Can Transformers Predict Pseudo-Random Numbers?
Fuente:
arXiv
Saved in:
| Main Authors: | Tao, Tao, Doshi, Darshil, Kalra, Dayal Singh, He, Tianyu, Barkeshli, Maissam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Pseudorandom Numbers with Transformers: Permuted Congruential Generators, Curricula, and Interpretability
by: Tao, Tao, et al.
Published: (2025)
by: Tao, Tao, et al.
Published: (2025)
Why Warmup the Learning Rate? Underlying Mechanisms and Improvements
by: Kalra, Dayal Singh, et al.
Published: (2024)
by: Kalra, Dayal Singh, et al.
Published: (2024)
Universal Sharpness Dynamics in Neural Network Training: Fixed Point Analysis, Edge of Stability, and Route to Chaos
by: Kalra, Dayal Singh, et al.
Published: (2023)
by: Kalra, Dayal Singh, et al.
Published: (2023)
Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate
by: Kalra, Dayal Singh, et al.
Published: (2026)
by: Kalra, Dayal Singh, et al.
Published: (2026)
When Can You Get Away with Low Memory Adam?
by: Kalra, Dayal Singh, et al.
Published: (2025)
by: Kalra, Dayal Singh, et al.
Published: (2025)
To grok or not to grok: Disentangling generalization and memorization on corrupted algorithmic datasets
by: Doshi, Darshil, et al.
Published: (2023)
by: Doshi, Darshil, et al.
Published: (2023)
On the origin of neural scaling laws: from random graphs to natural language
by: Barkeshli, Maissam, et al.
Published: (2026)
by: Barkeshli, Maissam, et al.
Published: (2026)
On the existence of consistent adversarial attacks in high-dimensional linear classification
by: Vilucchio, Matteo, et al.
Published: (2025)
by: Vilucchio, Matteo, et al.
Published: (2025)
Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasks
by: He, Tianyu, et al.
Published: (2024)
by: He, Tianyu, et al.
Published: (2024)
Adaptive Variation-Resilient Random Number Generator for Embedded Encryption
by: Zahoor, Furqan, et al.
Published: (2025)
by: Zahoor, Furqan, et al.
Published: (2025)
Grokking Modular Polynomials
by: Doshi, Darshil, et al.
Published: (2024)
by: Doshi, Darshil, et al.
Published: (2024)
Spectral Feature Extraction for Robust Network Intrusion Detection Using MFCCs
by: Lee, HyeYoung, et al.
Published: (2025)
by: Lee, HyeYoung, et al.
Published: (2025)
Are Neural Networks Collision Resistant?
by: Benedetti, Marco, et al.
Published: (2025)
by: Benedetti, Marco, et al.
Published: (2025)
Quantum-activated neural reservoirs on-chip open up large hardware security models for resilient authentication
by: He, Zhao, et al.
Published: (2024)
by: He, Zhao, et al.
Published: (2024)
Random-Energy Secret Sharing via Extreme Synergy
by: Ngampruetikorn, Vudtiwat, et al.
Published: (2023)
by: Ngampruetikorn, Vudtiwat, et al.
Published: (2023)
A Scalable Measure of Loss Landscape Curvature for Analyzing the Training Dynamics of LLMs
by: Kalra, Dayal Singh, et al.
Published: (2026)
by: Kalra, Dayal Singh, et al.
Published: (2026)
How Feature Learning Can Improve Neural Scaling Laws
by: Bordelon, Blake, et al.
Published: (2024)
by: Bordelon, Blake, et al.
Published: (2024)
Small Singular Values Matter: A Random Matrix Analysis of Transformer Models
by: Staats, Max, et al.
Published: (2024)
by: Staats, Max, et al.
Published: (2024)
How Deep Networks Learn Sparse and Hierarchical Data: the Sparse Random Hierarchy Model
by: Tomasini, Umberto, et al.
Published: (2024)
by: Tomasini, Umberto, et al.
Published: (2024)
Geometric Dynamics of Signal Propagation Predict Trainability of Transformers
by: Cowsik, Aditya, et al.
Published: (2024)
by: Cowsik, Aditya, et al.
Published: (2024)
Random features and polynomial rules
by: Aguirre-López, Fabián, et al.
Published: (2024)
by: Aguirre-López, Fabián, et al.
Published: (2024)
Analog Physical Systems Can Exhibit Double Descent
by: Dillavou, Sam, et al.
Published: (2025)
by: Dillavou, Sam, et al.
Published: (2025)
Exact Learning Dynamics of In-Context Learning in Linear Transformers and Its Application to Non-Linear Transformers
by: Mainali, Nischal, et al.
Published: (2025)
by: Mainali, Nischal, et al.
Published: (2025)
A Random-Matrix Criterion for Initializing Gated Recurrent Neural Networks
by: Fioratti, Tommaso, et al.
Published: (2026)
by: Fioratti, Tommaso, et al.
Published: (2026)
Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning
by: Bordelon, Blake, et al.
Published: (2026)
by: Bordelon, Blake, et al.
Published: (2026)
EB-RANSAC: Random Sample Consensus based on Energy-Based Model
by: Yasuda, Muneki, et al.
Published: (2026)
by: Yasuda, Muneki, et al.
Published: (2026)
Towards Distributed Neural Architectures
by: Cowsik, Aditya, et al.
Published: (2025)
by: Cowsik, Aditya, et al.
Published: (2025)
Infinite Limits of Multi-head Transformer Dynamics
by: Bordelon, Blake, et al.
Published: (2024)
by: Bordelon, Blake, et al.
Published: (2024)
No Free Lunch From Random Feature Ensembles: Scaling Laws and Near-Optimality Conditions
by: Ruben, Benjamin S., et al.
Published: (2024)
by: Ruben, Benjamin S., et al.
Published: (2024)
Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model
by: Bordelon, Blake, et al.
Published: (2026)
by: Bordelon, Blake, et al.
Published: (2026)
How does training shape the Riemannian geometry of neural network representations?
by: Zavatone-Veth, Jacob A., et al.
Published: (2023)
by: Zavatone-Veth, Jacob A., et al.
Published: (2023)
Initial Guessing Bias: How Untrained Networks Favor Some Classes
by: Francazi, Emanuele, et al.
Published: (2023)
by: Francazi, Emanuele, et al.
Published: (2023)
Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
by: Bordelon, Blake, et al.
Published: (2025)
by: Bordelon, Blake, et al.
Published: (2025)
Towards Understanding Inductive Bias in Transformers: A View From Infinity
by: Lavie, Itay, et al.
Published: (2024)
by: Lavie, Itay, et al.
Published: (2024)
Algorithmic Task Capture, Computational Complexity, and Inductive Bias of Infinite Transformers
by: Davidovich, Orit, et al.
Published: (2026)
by: Davidovich, Orit, et al.
Published: (2026)
Quantum cryptographic protocols with dual messaging system via 2D alternate quantum walk of a genuine single-photon entangled state
by: Panda, Dinesh Kumar, et al.
Published: (2024)
by: Panda, Dinesh Kumar, et al.
Published: (2024)
Building Conformal Prediction Intervals with Approximate Message Passing
by: Clarté, Lucas, et al.
Published: (2024)
by: Clarté, Lucas, et al.
Published: (2024)
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
by: Cagnetta, Francesco, et al.
Published: (2025)
by: Cagnetta, Francesco, et al.
Published: (2025)
Graph Neural Network Approach to Predicting Magnetization in Quasi-One-Dimensional Ising Systems
by: Slavin, V., et al.
Published: (2025)
by: Slavin, V., et al.
Published: (2025)
Siamese Neural Network for Label-Efficient Critical Phenomena Prediction in 3D Percolation Models
by: Wang, Shanshan, et al.
Published: (2025)
by: Wang, Shanshan, et al.
Published: (2025)
Similar Items
-
Learning Pseudorandom Numbers with Transformers: Permuted Congruential Generators, Curricula, and Interpretability
by: Tao, Tao, et al.
Published: (2025) -
Why Warmup the Learning Rate? Underlying Mechanisms and Improvements
by: Kalra, Dayal Singh, et al.
Published: (2024) -
Universal Sharpness Dynamics in Neural Network Training: Fixed Point Analysis, Edge of Stability, and Route to Chaos
by: Kalra, Dayal Singh, et al.
Published: (2023) -
Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate
by: Kalra, Dayal Singh, et al.
Published: (2026) -
When Can You Get Away with Low Memory Adam?
by: Kalra, Dayal Singh, et al.
Published: (2025)