Pre-trained Language Models Learn Remarkably Accurate Representations of Numbers
Fuente:
arXiv
Saved in:
| Main Authors: | Kadlčík, Marek, Štefánik, Michal, Mickus, Timothee, Spiegel, Michal, Kuchař, Josef |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Language Models Learn Universal Representations of Numbers and Here's Why You Should Care
by: Štefánik, Michal, et al.
Published: (2025)
by: Štefánik, Michal, et al.
Published: (2025)
Can Out-of-Distribution Evaluations Uncover Reliance on Shortcuts? A Case Study in Question Answering
by: Štefánik, Michal, et al.
Published: (2025)
by: Štefánik, Michal, et al.
Published: (2025)
VectorEdits: A Dataset and Benchmark for Instruction-Based Editing of Vector Graphics
by: Kuchař, Josef, et al.
Published: (2025)
by: Kuchař, Josef, et al.
Published: (2025)
Attend or Perish: Benchmarking Attention in Algorithmic Reasoning
by: Spiegel, Michal, et al.
Published: (2025)
by: Spiegel, Michal, et al.
Published: (2025)
SpikeGPT: Generative Pre-trained Language Model with Spiking Neural Networks
by: Zhu, Rui-Jie, et al.
Published: (2023)
by: Zhu, Rui-Jie, et al.
Published: (2023)
Self-training Language Models for Arithmetic Reasoning
by: Kadlčík, Marek, et al.
Published: (2024)
by: Kadlčík, Marek, et al.
Published: (2024)
Large Language Models for Tuning Evolution Strategies
by: Kramer, Oliver
Published: (2024)
by: Kramer, Oliver
Published: (2024)
Sorbet: A Neuromorphic Hardware-Compatible Transformer-Based Spiking Language Model
by: Tang, Kaiwen, et al.
Published: (2024)
by: Tang, Kaiwen, et al.
Published: (2024)
Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models
by: Dong, Peijie, et al.
Published: (2024)
by: Dong, Peijie, et al.
Published: (2024)
Genetic Instruct: Scaling up Synthetic Generation of Coding Instructions for Large Language Models
by: Majumdar, Somshubra, et al.
Published: (2024)
by: Majumdar, Somshubra, et al.
Published: (2024)
Availability of Perfect Decomposition in Statistical Linkage Learning for Unitation-based Function Concatenations
by: Prusik, Michal, et al.
Published: (2025)
by: Prusik, Michal, et al.
Published: (2025)
Large Language Models Suffer From Their Own Output: An Analysis of the Self-Consuming Training Loop
by: Briesch, Martin, et al.
Published: (2023)
by: Briesch, Martin, et al.
Published: (2023)
Parametric-Task MAP-Elites
by: Anne, Timothée, et al.
Published: (2024)
by: Anne, Timothée, et al.
Published: (2024)
SpikeLM: Towards General Spike-Driven Language Modeling via Elastic Bi-Spiking Mechanisms
by: Xing, Xingrun, et al.
Published: (2024)
by: Xing, Xingrun, et al.
Published: (2024)
SpikeLLM: Scaling up Spiking Neural Network to Large Language Models via Saliency-based Spiking
by: Xing, Xingrun, et al.
Published: (2024)
by: Xing, Xingrun, et al.
Published: (2024)
SpikingSSMs: Learning Long Sequences with Sparse and Parallel Spiking State Space Models
by: Shen, Shuaijie, et al.
Published: (2024)
by: Shen, Shuaijie, et al.
Published: (2024)
Improving Language Plasticity via Pretraining with Active Forgetting
by: Chen, Yihong, et al.
Published: (2023)
by: Chen, Yihong, et al.
Published: (2023)
HAT: Hardware-Aware Transformers for Efficient Natural Language Processing
by: Wang, Hanrui, et al.
Published: (2020)
by: Wang, Hanrui, et al.
Published: (2020)
The Space Between: On Folding, Symmetries and Sampling
by: Lewandowski, Michal, et al.
Published: (2025)
by: Lewandowski, Michal, et al.
Published: (2025)
An enhanced Teaching-Learning-Based Optimization (TLBO) with Grey Wolf Optimizer (GWO) for text feature selection and clustering
by: Azarshab, Mahsa, et al.
Published: (2024)
by: Azarshab, Mahsa, et al.
Published: (2024)
Negation: A Pink Elephant in the Large Language Models' Room?
by: Vrabcová, Tereza, et al.
Published: (2025)
by: Vrabcová, Tereza, et al.
Published: (2025)
Improving Sequence-to-Sequence Models for Abstractive Text Summarization Using Meta Heuristic Approaches
by: Saxena, Aditya, et al.
Published: (2024)
by: Saxena, Aditya, et al.
Published: (2024)
B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory
by: Zancato, Luca, et al.
Published: (2024)
by: Zancato, Luca, et al.
Published: (2024)
AP-BMM: Approximating Capability-Cost Pareto Sets of LLMs via Asynchronous Prior-Guided Bayesian Model Merging
by: Chen, Kesheng, et al.
Published: (2025)
by: Chen, Kesheng, et al.
Published: (2025)
On Space Folds of ReLU Neural Networks
by: Lewandowski, Michal, et al.
Published: (2025)
by: Lewandowski, Michal, et al.
Published: (2025)
Recent Advances in Federated Learning Driven Large Language Models: A Survey on Architecture, Performance, and Security
by: Qu, Youyang, et al.
Published: (2024)
by: Qu, Youyang, et al.
Published: (2024)
Learning Self-Growth Maps for Fast and Accurate Imbalanced Streaming Data Clustering
by: Zhang, Yiqun, et al.
Published: (2024)
by: Zhang, Yiqun, et al.
Published: (2024)
Diffusion Language Models for Speech Recognition
by: Naveriani, Davyd, et al.
Published: (2026)
by: Naveriani, Davyd, et al.
Published: (2026)
Elastic Architecture Search for Efficient Language Models
by: Wang, Shang
Published: (2025)
by: Wang, Shang
Published: (2025)
Hyperbolic Fine-Tuning for Large Language Models
by: Yang, Menglin, et al.
Published: (2024)
by: Yang, Menglin, et al.
Published: (2024)
EvoMerge: Neuroevolution for Large Language Models
by: Jiang, Yushu
Published: (2024)
by: Jiang, Yushu
Published: (2024)
Solve the Loop: Attractor Models for Language and Reasoning
by: Fein-Ashley, Jacob, et al.
Published: (2026)
by: Fein-Ashley, Jacob, et al.
Published: (2026)
H-Node Attack and Defense in Large Language Models
by: Yocam, Eric, et al.
Published: (2026)
by: Yocam, Eric, et al.
Published: (2026)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
by: Yu, Bohan, et al.
Published: (2025)
by: Yu, Bohan, et al.
Published: (2025)
Intelligent Neural Networks: From Layered Architectures to Graph-Organized Intelligence
by: Salomon, Antoine
Published: (2025)
by: Salomon, Antoine
Published: (2025)
Decomposing Evolutionary Mixture-of-LoRA Architectures: The Routing Lever, the Lifecycle Penalty, and a Substrate-Conditional Boundary
by: Kumaresan, Ramchand
Published: (2026)
by: Kumaresan, Ramchand
Published: (2026)
An In-depth Walkthrough on Evolution of Neural Machine Translation
by: Jagtap, Rohan, et al.
Published: (2020)
by: Jagtap, Rohan, et al.
Published: (2020)
Hysteresis Activation Function for Efficient Inference
by: Kimhi, Moshe, et al.
Published: (2024)
by: Kimhi, Moshe, et al.
Published: (2024)
On the Power of Convolution Augmented Transformer
by: Li, Mingchen, et al.
Published: (2024)
by: Li, Mingchen, et al.
Published: (2024)
A Hormone-inspired Emotion Layer for Transformer language models (HELT)
by: Reda, Eslam, et al.
Published: (2026)
by: Reda, Eslam, et al.
Published: (2026)
Similar Items
-
Language Models Learn Universal Representations of Numbers and Here's Why You Should Care
by: Štefánik, Michal, et al.
Published: (2025) -
Can Out-of-Distribution Evaluations Uncover Reliance on Shortcuts? A Case Study in Question Answering
by: Štefánik, Michal, et al.
Published: (2025) -
VectorEdits: A Dataset and Benchmark for Instruction-Based Editing of Vector Graphics
by: Kuchař, Josef, et al.
Published: (2025) -
Attend or Perish: Benchmarking Attention in Algorithmic Reasoning
by: Spiegel, Michal, et al.
Published: (2025) -
SpikeGPT: Generative Pre-trained Language Model with Spiking Neural Networks
by: Zhu, Rui-Jie, et al.
Published: (2023)