A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Hugh, Da, Jeff, Lee, Dean, Robinson, Vaughn, Wu, Catherine, Song, Will, Zhao, Tiffany, Raja, Pranav, Zhuang, Charlotte, Slack, Dylan, Lyu, Qin, Hendryx, Sean, Kaplan, Russell, Lunati, Michele, Yue, Summer |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Learning Goal-Conditioned Representations for Language Reward Models
di: Nath, Vaskar, et al.
Pubblicazione: (2024)
di: Nath, Vaskar, et al.
Pubblicazione: (2024)
ToolComp: A Multi-Tool Reasoning & Process Supervision Benchmark
di: Nath, Vaskar, et al.
Pubblicazione: (2025)
di: Nath, Vaskar, et al.
Pubblicazione: (2025)
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
di: Da, Jeff, et al.
Pubblicazione: (2025)
di: Da, Jeff, et al.
Pubblicazione: (2025)
Planning In Natural Language Improves LLM Search For Code Generation
di: Wang, Evan, et al.
Pubblicazione: (2024)
di: Wang, Evan, et al.
Pubblicazione: (2024)
Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents
di: Kumar, Priyanshu, et al.
Pubblicazione: (2024)
di: Kumar, Priyanshu, et al.
Pubblicazione: (2024)
Revisiting the Superficial Alignment Hypothesis
di: Raghavendra, Mohit, et al.
Pubblicazione: (2024)
di: Raghavendra, Mohit, et al.
Pubblicazione: (2024)
Pre-Training Multimodal Hallucination Detectors with Corrupted Grounding Data
di: Whitehead, Spencer, et al.
Pubblicazione: (2024)
di: Whitehead, Spencer, et al.
Pubblicazione: (2024)
EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges
di: Wang, Clinton J., et al.
Pubblicazione: (2025)
di: Wang, Clinton J., et al.
Pubblicazione: (2025)
Progress over Points: Reframing LM Benchmarks Around Scientific Objectives
di: Jin, Alwin, et al.
Pubblicazione: (2025)
di: Jin, Alwin, et al.
Pubblicazione: (2025)
Early Detection and Reduction of Memorisation for Domain Adaptation and Instruction Tuning
di: Slack, Dean L., et al.
Pubblicazione: (2025)
di: Slack, Dean L., et al.
Pubblicazione: (2025)
Going 3D with Technology: An Overarching Approach for Language Teachers
di: Jason D. Hendryx
Pubblicazione: (2016)
di: Jason D. Hendryx
Pubblicazione: (2016)
A Baseline Analysis of Reward Models' Ability To Accurately Analyze Foundation Models Under Distribution Shift
di: LeVine, Will, et al.
Pubblicazione: (2023)
di: LeVine, Will, et al.
Pubblicazione: (2023)
The discrete empirical interpolation method in class identification and data summarization
di: Emily P. Hendryx Lyons
Pubblicazione: (2024)
di: Emily P. Hendryx Lyons
Pubblicazione: (2024)
Surrogate Trajectories Along Probability Flows: Pseudo Markovian Alternative to Mori Zwanzig
di: Stauffer, Noé, et al.
Pubblicazione: (2025)
di: Stauffer, Noé, et al.
Pubblicazione: (2025)
A Respiratory Simulator for the Study of Pathogen Transmission in Indoor Environments
di: Claudio Mucignat, et al.
Pubblicazione: (2024)
di: Claudio Mucignat, et al.
Pubblicazione: (2024)
Arithmeticity and commensurability of links in thickened surfaces
di: Futer, David, et al.
Pubblicazione: (2024)
di: Futer, David, et al.
Pubblicazione: (2024)
Viajes violentos : la transformación de la migración clandestina hacia Sonora y Arizona / Jeremy Slack, Scott Whiteford
di: Slack, Jeremy
Pubblicazione: (2010)
di: Slack, Jeremy
Pubblicazione: (2010)
Learning to Let Go: Disenfranchisement from Academia
di: Hannah Slack
Pubblicazione: (2024)
di: Hannah Slack
Pubblicazione: (2024)
Learning to Let Go: Disenfranchisement from Academia
di: Hannah Slack
Pubblicazione: (2024)
di: Hannah Slack
Pubblicazione: (2024)
Viajes violentos: la transformación de la migración clandestina hacia Sonora y Arizona
di: Jeremy Slack
Pubblicazione: (2010)
di: Jeremy Slack
Pubblicazione: (2010)
Long‐Term Effects of Low‐Drop Grade Control Structures on Channel Evolution in the Yazoo River Basin
di: Nicky M. Faucheux, et al.
Pubblicazione: (2025)
di: Nicky M. Faucheux, et al.
Pubblicazione: (2025)
Matryoshka Quantization
di: Nair, Pranav, et al.
Pubblicazione: (2025)
di: Nair, Pranav, et al.
Pubblicazione: (2025)
The Time Machine. Grade 4-6.
di: Petersen, Jeff
Pubblicazione: (1986)
di: Petersen, Jeff
Pubblicazione: (1986)
Shxwelí li te shxwelítemelh xíts'etáwtxw: The museum's confinement of Indigenous kin
di: Dylan Robinson
Pubblicazione: (2024)
di: Dylan Robinson
Pubblicazione: (2024)
Periodic autoregressive moving average models for the analysis of streamflow [abstract]
di: Slack, J.R.
Pubblicazione: (1988)
di: Slack, J.R.
Pubblicazione: (1988)
Vocab Diet: Reshaping the Vocabulary of LLMs via Vector Arithmetic
di: Reif, Yuval, et al.
Pubblicazione: (2025)
di: Reif, Yuval, et al.
Pubblicazione: (2025)
How Consistent are Course Grades? An Examination of Differential Grading
di: Samuel Rauschenberg
Pubblicazione: (2014)
di: Samuel Rauschenberg
Pubblicazione: (2014)
Precision-Graded Cohomology and Arithmetic Persistence for Network Sheaves
di: Ghrist, Robert, et al.
Pubblicazione: (2025)
di: Ghrist, Robert, et al.
Pubblicazione: (2025)
Graded Lie Algebras, Compactified Jacobians and Arithmetic Statistics
di: Laga, Jef
Pubblicazione: (2022)
di: Laga, Jef
Pubblicazione: (2022)
An In‐Depth Examination of Learning Behaviors in Autistic Middle‐Schoolers Without Intellectual Disability
di: Amie Duncan, et al.
Pubblicazione: (2025)
di: Amie Duncan, et al.
Pubblicazione: (2025)
Embedding Elites: Examining the Use of Tweets Embedded in Online News Articles across Reliable and Fringe Outlets
di: Horne, Benjamin D., et al.
Pubblicazione: (2024)
di: Horne, Benjamin D., et al.
Pubblicazione: (2024)
Investigating Task Arithmetic for Zero-Shot Information Retrieval
di: Braga, Marco, et al.
Pubblicazione: (2025)
di: Braga, Marco, et al.
Pubblicazione: (2025)
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
di: Nath, Vaskar, et al.
Pubblicazione: (2025)
di: Nath, Vaskar, et al.
Pubblicazione: (2025)
Arithmetic Geometric Model for the Renormalisation of Bi-critical Irrationally Indifferent Attractors
di: Russell, Jocelyn Finbar
Pubblicazione: (2025)
di: Russell, Jocelyn Finbar
Pubblicazione: (2025)
Arithmetic regularity as an alternative to transference
di: Chow, Sam, et al.
Pubblicazione: (2026)
di: Chow, Sam, et al.
Pubblicazione: (2026)
Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers
di: Slack, Dean L, et al.
Pubblicazione: (2025)
di: Slack, Dean L, et al.
Pubblicazione: (2025)
Orientalizing New Spain: Perspectives on Asian Influence in Colonial Mexico
di: Edward R. Slack, Jr.
Pubblicazione: (2012)
di: Edward R. Slack, Jr.
Pubblicazione: (2012)
Out-of-Distribution Detection & Applications With Ablated Learned Temperature Energy
di: LeVine, Will, et al.
Pubblicazione: (2024)
di: LeVine, Will, et al.
Pubblicazione: (2024)
The time for decision / Welles Summer
di: Welles, Summer
Pubblicazione: (1944)
di: Welles, Summer
Pubblicazione: (1944)
Reconfiguring and Remediating Social Media as Alternative Media: Exploring Youth Activists’ Digital Media Ecology in El Salvador
di: Summer Harlow
Pubblicazione: (2016)
di: Summer Harlow
Pubblicazione: (2016)
Documenti analoghi
-
Learning Goal-Conditioned Representations for Language Reward Models
di: Nath, Vaskar, et al.
Pubblicazione: (2024) -
ToolComp: A Multi-Tool Reasoning & Process Supervision Benchmark
di: Nath, Vaskar, et al.
Pubblicazione: (2025) -
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
di: Da, Jeff, et al.
Pubblicazione: (2025) -
Planning In Natural Language Improves LLM Search For Code Generation
di: Wang, Evan, et al.
Pubblicazione: (2024) -
Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents
di: Kumar, Priyanshu, et al.
Pubblicazione: (2024)