A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Hugh, Da, Jeff, Lee, Dean, Robinson, Vaughn, Wu, Catherine, Song, Will, Zhao, Tiffany, Raja, Pranav, Zhuang, Charlotte, Slack, Dylan, Lyu, Qin, Hendryx, Sean, Kaplan, Russell, Lunati, Michele, Yue, Summer |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning Goal-Conditioned Representations for Language Reward Models
von: Nath, Vaskar, et al.
Veröffentlicht: (2024)
von: Nath, Vaskar, et al.
Veröffentlicht: (2024)
ToolComp: A Multi-Tool Reasoning & Process Supervision Benchmark
von: Nath, Vaskar, et al.
Veröffentlicht: (2025)
von: Nath, Vaskar, et al.
Veröffentlicht: (2025)
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
von: Da, Jeff, et al.
Veröffentlicht: (2025)
von: Da, Jeff, et al.
Veröffentlicht: (2025)
Planning In Natural Language Improves LLM Search For Code Generation
von: Wang, Evan, et al.
Veröffentlicht: (2024)
von: Wang, Evan, et al.
Veröffentlicht: (2024)
Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents
von: Kumar, Priyanshu, et al.
Veröffentlicht: (2024)
von: Kumar, Priyanshu, et al.
Veröffentlicht: (2024)
Revisiting the Superficial Alignment Hypothesis
von: Raghavendra, Mohit, et al.
Veröffentlicht: (2024)
von: Raghavendra, Mohit, et al.
Veröffentlicht: (2024)
Pre-Training Multimodal Hallucination Detectors with Corrupted Grounding Data
von: Whitehead, Spencer, et al.
Veröffentlicht: (2024)
von: Whitehead, Spencer, et al.
Veröffentlicht: (2024)
EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges
von: Wang, Clinton J., et al.
Veröffentlicht: (2025)
von: Wang, Clinton J., et al.
Veröffentlicht: (2025)
Progress over Points: Reframing LM Benchmarks Around Scientific Objectives
von: Jin, Alwin, et al.
Veröffentlicht: (2025)
von: Jin, Alwin, et al.
Veröffentlicht: (2025)
Early Detection and Reduction of Memorisation for Domain Adaptation and Instruction Tuning
von: Slack, Dean L., et al.
Veröffentlicht: (2025)
von: Slack, Dean L., et al.
Veröffentlicht: (2025)
Going 3D with Technology: An Overarching Approach for Language Teachers
von: Jason D. Hendryx
Veröffentlicht: (2016)
von: Jason D. Hendryx
Veröffentlicht: (2016)
A Baseline Analysis of Reward Models' Ability To Accurately Analyze Foundation Models Under Distribution Shift
von: LeVine, Will, et al.
Veröffentlicht: (2023)
von: LeVine, Will, et al.
Veröffentlicht: (2023)
The discrete empirical interpolation method in class identification and data summarization
von: Emily P. Hendryx Lyons
Veröffentlicht: (2024)
von: Emily P. Hendryx Lyons
Veröffentlicht: (2024)
Surrogate Trajectories Along Probability Flows: Pseudo Markovian Alternative to Mori Zwanzig
von: Stauffer, Noé, et al.
Veröffentlicht: (2025)
von: Stauffer, Noé, et al.
Veröffentlicht: (2025)
A Respiratory Simulator for the Study of Pathogen Transmission in Indoor Environments
von: Claudio Mucignat, et al.
Veröffentlicht: (2024)
von: Claudio Mucignat, et al.
Veröffentlicht: (2024)
Arithmeticity and commensurability of links in thickened surfaces
von: Futer, David, et al.
Veröffentlicht: (2024)
von: Futer, David, et al.
Veröffentlicht: (2024)
Viajes violentos : la transformación de la migración clandestina hacia Sonora y Arizona / Jeremy Slack, Scott Whiteford
von: Slack, Jeremy
Veröffentlicht: (2010)
von: Slack, Jeremy
Veröffentlicht: (2010)
Learning to Let Go: Disenfranchisement from Academia
von: Hannah Slack
Veröffentlicht: (2024)
von: Hannah Slack
Veröffentlicht: (2024)
Learning to Let Go: Disenfranchisement from Academia
von: Hannah Slack
Veröffentlicht: (2024)
von: Hannah Slack
Veröffentlicht: (2024)
Viajes violentos: la transformación de la migración clandestina hacia Sonora y Arizona
von: Jeremy Slack
Veröffentlicht: (2010)
von: Jeremy Slack
Veröffentlicht: (2010)
Long‐Term Effects of Low‐Drop Grade Control Structures on Channel Evolution in the Yazoo River Basin
von: Nicky M. Faucheux, et al.
Veröffentlicht: (2025)
von: Nicky M. Faucheux, et al.
Veröffentlicht: (2025)
Matryoshka Quantization
von: Nair, Pranav, et al.
Veröffentlicht: (2025)
von: Nair, Pranav, et al.
Veröffentlicht: (2025)
The Time Machine. Grade 4-6.
von: Petersen, Jeff
Veröffentlicht: (1986)
von: Petersen, Jeff
Veröffentlicht: (1986)
Shxwelí li te shxwelítemelh xíts'etáwtxw: The museum's confinement of Indigenous kin
von: Dylan Robinson
Veröffentlicht: (2024)
von: Dylan Robinson
Veröffentlicht: (2024)
Periodic autoregressive moving average models for the analysis of streamflow [abstract]
von: Slack, J.R.
Veröffentlicht: (1988)
von: Slack, J.R.
Veröffentlicht: (1988)
Vocab Diet: Reshaping the Vocabulary of LLMs via Vector Arithmetic
von: Reif, Yuval, et al.
Veröffentlicht: (2025)
von: Reif, Yuval, et al.
Veröffentlicht: (2025)
How Consistent are Course Grades? An Examination of Differential Grading
von: Samuel Rauschenberg
Veröffentlicht: (2014)
von: Samuel Rauschenberg
Veröffentlicht: (2014)
Precision-Graded Cohomology and Arithmetic Persistence for Network Sheaves
von: Ghrist, Robert, et al.
Veröffentlicht: (2025)
von: Ghrist, Robert, et al.
Veröffentlicht: (2025)
Graded Lie Algebras, Compactified Jacobians and Arithmetic Statistics
von: Laga, Jef
Veröffentlicht: (2022)
von: Laga, Jef
Veröffentlicht: (2022)
An In‐Depth Examination of Learning Behaviors in Autistic Middle‐Schoolers Without Intellectual Disability
von: Amie Duncan, et al.
Veröffentlicht: (2025)
von: Amie Duncan, et al.
Veröffentlicht: (2025)
Embedding Elites: Examining the Use of Tweets Embedded in Online News Articles across Reliable and Fringe Outlets
von: Horne, Benjamin D., et al.
Veröffentlicht: (2024)
von: Horne, Benjamin D., et al.
Veröffentlicht: (2024)
Investigating Task Arithmetic for Zero-Shot Information Retrieval
von: Braga, Marco, et al.
Veröffentlicht: (2025)
von: Braga, Marco, et al.
Veröffentlicht: (2025)
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
von: Nath, Vaskar, et al.
Veröffentlicht: (2025)
von: Nath, Vaskar, et al.
Veröffentlicht: (2025)
Arithmetic Geometric Model for the Renormalisation of Bi-critical Irrationally Indifferent Attractors
von: Russell, Jocelyn Finbar
Veröffentlicht: (2025)
von: Russell, Jocelyn Finbar
Veröffentlicht: (2025)
Arithmetic regularity as an alternative to transference
von: Chow, Sam, et al.
Veröffentlicht: (2026)
von: Chow, Sam, et al.
Veröffentlicht: (2026)
Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers
von: Slack, Dean L, et al.
Veröffentlicht: (2025)
von: Slack, Dean L, et al.
Veröffentlicht: (2025)
Orientalizing New Spain: Perspectives on Asian Influence in Colonial Mexico
von: Edward R. Slack, Jr.
Veröffentlicht: (2012)
von: Edward R. Slack, Jr.
Veröffentlicht: (2012)
Out-of-Distribution Detection & Applications With Ablated Learned Temperature Energy
von: LeVine, Will, et al.
Veröffentlicht: (2024)
von: LeVine, Will, et al.
Veröffentlicht: (2024)
The time for decision / Welles Summer
von: Welles, Summer
Veröffentlicht: (1944)
von: Welles, Summer
Veröffentlicht: (1944)
Reconfiguring and Remediating Social Media as Alternative Media: Exploring Youth Activists’ Digital Media Ecology in El Salvador
von: Summer Harlow
Veröffentlicht: (2016)
von: Summer Harlow
Veröffentlicht: (2016)
Ähnliche Einträge
-
Learning Goal-Conditioned Representations for Language Reward Models
von: Nath, Vaskar, et al.
Veröffentlicht: (2024) -
ToolComp: A Multi-Tool Reasoning & Process Supervision Benchmark
von: Nath, Vaskar, et al.
Veröffentlicht: (2025) -
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
von: Da, Jeff, et al.
Veröffentlicht: (2025) -
Planning In Natural Language Improves LLM Search For Code Generation
von: Wang, Evan, et al.
Veröffentlicht: (2024) -
Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents
von: Kumar, Priyanshu, et al.
Veröffentlicht: (2024)