Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Gambardella, Andrew, Iwasawa, Yusuke, Matsuo, Yutaka |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
by: Gambardella, Andrew, et al.
Published: (2025)
by: Gambardella, Andrew, et al.
Published: (2025)
Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models
by: Cao, Qi, et al.
Published: (2026)
by: Cao, Qi, et al.
Published: (2026)
The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
by: Hase, Peter, et al.
Published: (2024)
by: Hase, Peter, et al.
Published: (2024)
Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning
by: Parashar, Shubham, et al.
Published: (2025)
by: Parashar, Shubham, et al.
Published: (2025)
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
by: Minegishi, Gouki, et al.
Published: (2025)
by: Minegishi, Gouki, et al.
Published: (2025)
Exposing Limitations of Language Model Agents in Sequential-Task Compositions on the Web
by: Furuta, Hiroki, et al.
Published: (2023)
by: Furuta, Hiroki, et al.
Published: (2023)
Which Programming Language and What Features at Pre-training Stage Affect Downstream Logical Inference Performance?
by: Uchiyama, Fumiya, et al.
Published: (2024)
by: Uchiyama, Fumiya, et al.
Published: (2024)
Self-Harmony: Learning to Harmonize Self-Supervision and Self-Play in Test-Time Reinforcement Learning
by: Wang, Ru, et al.
Published: (2025)
by: Wang, Ru, et al.
Published: (2025)
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure
by: Cho, Hanseul, et al.
Published: (2024)
by: Cho, Hanseul, et al.
Published: (2024)
Safe Transformer: An Explicit Safety Bit For Interpretable And Controllable Alignment
by: Feng, Jingyuan, et al.
Published: (2026)
by: Feng, Jingyuan, et al.
Published: (2026)
Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
by: Sun, Zhiqing, et al.
Published: (2024)
by: Sun, Zhiqing, et al.
Published: (2024)
Large Language Models as Theory of Mind Aware Generative Agents with Counterfactual Reflection
by: Yang, Bo, et al.
Published: (2025)
by: Yang, Bo, et al.
Published: (2025)
Easy2Hard-Bench: Standardized Difficulty Labels for Profiling LLM Performance and Generalization
by: Ding, Mucong, et al.
Published: (2024)
by: Ding, Mucong, et al.
Published: (2024)
Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models
by: Gu, Xiaojie, et al.
Published: (2026)
by: Gu, Xiaojie, et al.
Published: (2026)
RevOrder: A Novel Method for Enhanced Arithmetic in Language Models
by: Shen, Si, et al.
Published: (2024)
by: Shen, Si, et al.
Published: (2024)
A Careful Examination of Large Language Model Performance on Grade School Arithmetic
by: Zhang, Hugh, et al.
Published: (2024)
by: Zhang, Hugh, et al.
Published: (2024)
OccamLLM: Fast and Exact Language Model Arithmetic in a Single Step
by: Dugan, Owen, et al.
Published: (2024)
by: Dugan, Owen, et al.
Published: (2024)
Task Arithmetic for Language Expansion in Speech Translation
by: Cheng, Yao-Fei, et al.
Published: (2024)
by: Cheng, Yao-Fei, et al.
Published: (2024)
Task Arithmetic can Mitigate Synthetic-to-Real Gap in Automatic Speech Recognition
by: Su, Hsuan, et al.
Published: (2024)
by: Su, Hsuan, et al.
Published: (2024)
Integrating Arithmetic Learning Improves Mathematical Reasoning in Smaller Models
by: Gangwar, Neeraj, et al.
Published: (2025)
by: Gangwar, Neeraj, et al.
Published: (2025)
Mathematical Reasoning in Large Language Models: Assessing Logical and Arithmetic Errors across Wide Numerical Ranges
by: Shrestha, Safal, et al.
Published: (2025)
by: Shrestha, Safal, et al.
Published: (2025)
Towards Empirical Interpretation of Internal Circuits and Properties in Grokked Transformers on Modular Polynomials
by: Furuta, Hiroki, et al.
Published: (2024)
by: Furuta, Hiroki, et al.
Published: (2024)
Bias Vector: Mitigating Biases in Language Models with Task Arithmetic Approach
by: Shirafuji, Daiki, et al.
Published: (2024)
by: Shirafuji, Daiki, et al.
Published: (2024)
How Numerical Precision Affects Arithmetical Reasoning Capabilities of LLMs
by: Feng, Guhao, et al.
Published: (2024)
by: Feng, Guhao, et al.
Published: (2024)
Soft Tokens, Hard Truths
by: Butt, Natasha, et al.
Published: (2025)
by: Butt, Natasha, et al.
Published: (2025)
Transformers Can Do Arithmetic with the Right Embeddings
by: McLeish, Sean, et al.
Published: (2024)
by: McLeish, Sean, et al.
Published: (2024)
EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning
by: Seth, Ashish, et al.
Published: (2024)
by: Seth, Ashish, et al.
Published: (2024)
Large Language Models as Generalizable Policies for Embodied Tasks
by: Szot, Andrew, et al.
Published: (2023)
by: Szot, Andrew, et al.
Published: (2023)
Divide, Reweight, and Conquer: A Logit Arithmetic Approach for In-Context Learning
by: Huang, Chengsong, et al.
Published: (2024)
by: Huang, Chengsong, et al.
Published: (2024)
(How) Do Language Models Track State?
by: Li, Belinda Z., et al.
Published: (2025)
by: Li, Belinda Z., et al.
Published: (2025)
Language Models Use Trigonometry to Do Addition
by: Kantamneni, Subhash, et al.
Published: (2025)
by: Kantamneni, Subhash, et al.
Published: (2025)
ConsistentEE: A Consistent and Hardness-Guided Early Exiting Method for Accelerating Language Models Inference
by: Zeng, Ziqian, et al.
Published: (2023)
by: Zeng, Ziqian, et al.
Published: (2023)
QED-Nano: Teaching a Tiny Model to Prove Hard Theorems
by: LM-Provers, et al.
Published: (2026)
by: LM-Provers, et al.
Published: (2026)
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
by: Minegishi, Gouki, et al.
Published: (2025)
by: Minegishi, Gouki, et al.
Published: (2025)
AttributionBench: How Hard is Automatic Attribution Evaluation?
by: Li, Yifei, et al.
Published: (2024)
by: Li, Yifei, et al.
Published: (2024)
Verifier-Backed Hard Problem Generation for Mathematical Reasoning
by: Lai, Yuhang, et al.
Published: (2026)
by: Lai, Yuhang, et al.
Published: (2026)
Reuse your FLOPs: Scaling RL on Hard Problems by Conditioning on Very Off-Policy Prefixes
by: Setlur, Amrith, et al.
Published: (2026)
by: Setlur, Amrith, et al.
Published: (2026)
Understanding Emergent Misalignment via Feature Superposition Geometry
by: Minegishi, Gouki, et al.
Published: (2026)
by: Minegishi, Gouki, et al.
Published: (2026)
Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training
by: Matsutani, Kohsei, et al.
Published: (2026)
by: Matsutani, Kohsei, et al.
Published: (2026)
Do Large Language Models Show Biases in Causal Learning?
by: Carro, Maria Victoria, et al.
Published: (2024)
by: Carro, Maria Victoria, et al.
Published: (2024)
Similar Items
-
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
by: Gambardella, Andrew, et al.
Published: (2025) -
Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models
by: Cao, Qi, et al.
Published: (2026) -
The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
by: Hase, Peter, et al.
Published: (2024) -
Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning
by: Parashar, Shubham, et al.
Published: (2025) -
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
by: Minegishi, Gouki, et al.
Published: (2025)