Code-enabled language models can outperform reasoning models on diverse tasks
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Cedegao E., Colas, Cédric, Poesia, Gabriel, Tenenbaum, Joshua B., Andreas, Jacob |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Dissociating language and thought in large language models
par: Mahowald, Kyle, et autres
Publié: (2023)
par: Mahowald, Kyle, et autres
Publié: (2023)
Superhuman performance of a large language model on the reasoning tasks of a physician
par: Brodeur, Peter G., et autres
Publié: (2024)
par: Brodeur, Peter G., et autres
Publié: (2024)
Shoot First, Ask Questions Later? Building Rational Agents that Explore and Act Like People
par: Grand, Gabriel, et autres
Publié: (2025)
par: Grand, Gabriel, et autres
Publié: (2025)
Loose LIPS Sink Ships: Asking Questions in Battleship with Language-Informed Program Sampling
par: Grand, Gabriel, et autres
Publié: (2024)
par: Grand, Gabriel, et autres
Publié: (2024)
Language and Experience: A Computational Model of Social Learning in Complex Tasks
par: Colas, Cédric, et autres
Publié: (2025)
par: Colas, Cédric, et autres
Publié: (2025)
Evidence from counterfactual tasks supports emergent analogical reasoning in large language models
par: Webb, Taylor, et autres
Publié: (2024)
par: Webb, Taylor, et autres
Publié: (2024)
Scaling up the think-aloud method
par: Wurgaft, Daniel, et autres
Publié: (2025)
par: Wurgaft, Daniel, et autres
Publié: (2025)
LINC: A Neurosymbolic Approach for Logical Reasoning by Combining Language Models with First-Order Logic Provers
par: Olausson, Theo X., et autres
Publié: (2023)
par: Olausson, Theo X., et autres
Publié: (2023)
LILO: Learning Interpretable Libraries by Compressing and Documenting Code
par: Grand, Gabriel, et autres
Publié: (2023)
par: Grand, Gabriel, et autres
Publié: (2023)
Retrieval-augmented reasoning with lean language models
par: Chan, Ryan Sze-Yin, et autres
Publié: (2025)
par: Chan, Ryan Sze-Yin, et autres
Publié: (2025)
Slm-mux: Orchestrating small language models for reasoning
par: Wang, Chenyu, et autres
Publié: (2025)
par: Wang, Chenyu, et autres
Publié: (2025)
Response: Emergent analogical reasoning in large language models
par: Hodel, Damian, et autres
Publié: (2023)
par: Hodel, Damian, et autres
Publié: (2023)
Self-Steering Language Models
par: Grand, Gabriel, et autres
Publié: (2025)
par: Grand, Gabriel, et autres
Publié: (2025)
Policy Learning with a Language Bottleneck
par: Srivastava, Megha, et autres
Publié: (2024)
par: Srivastava, Megha, et autres
Publié: (2024)
Reinforcement learning fine-tuning of language model for instruction following and math reasoning
par: Han, Yifu, et autres
Publié: (2025)
par: Han, Yifu, et autres
Publié: (2025)
MathDivide: Improved mathematical reasoning by large language models
par: Srivastava, Saksham Sahai, et autres
Publié: (2024)
par: Srivastava, Saksham Sahai, et autres
Publié: (2024)
Large language models show fragile cognitive reasoning about human emotions
par: Bhattacharyya, Sree, et autres
Publié: (2025)
par: Bhattacharyya, Sree, et autres
Publié: (2025)
People use fast, goal-directed simulation to reason about novel games
par: Zhang, Cedegao E., et autres
Publié: (2024)
par: Zhang, Cedegao E., et autres
Publié: (2024)
Enhancing reasoning accuracy in large language models during inference time
par: Sharma, Vinay, et autres
Publié: (2026)
par: Sharma, Vinay, et autres
Publié: (2026)
ZNO-Eval: Benchmarking reasoning capabilities of large language models in Ukrainian
par: Syromiatnikov, Mykyta, et autres
Publié: (2025)
par: Syromiatnikov, Mykyta, et autres
Publié: (2025)
Modeling Open-World Cognition as On-Demand Synthesis of Probabilistic Models
par: Wong, Lionel, et autres
Publié: (2025)
par: Wong, Lionel, et autres
Publié: (2025)
Just-in-time and distributed task representations in language models
par: Li, Yuxuan, et autres
Publié: (2025)
par: Li, Yuxuan, et autres
Publié: (2025)
From Next-Token to Mathematics: The Learning Dynamics of Mathematical Reasoning in Language Models
par: Mishra, Shubhra, et autres
Publié: (2024)
par: Mishra, Shubhra, et autres
Publié: (2024)
ThoughtSource: A central hub for large language model reasoning data
par: Ott, Simon, et autres
Publié: (2023)
par: Ott, Simon, et autres
Publié: (2023)
Language-Informed Synthesis of Rational Agent Models for Grounded Theory-of-Mind Reasoning On-The-Fly
par: Ying, Lance, et autres
Publié: (2025)
par: Ying, Lance, et autres
Publié: (2025)
Auxiliary task demands mask the capabilities of smaller language models
par: Hu, Jennifer, et autres
Publié: (2024)
par: Hu, Jennifer, et autres
Publié: (2024)
Language models show human-like content effects on reasoning tasks
par: Dasgupta, Ishita, et autres
Publié: (2022)
par: Dasgupta, Ishita, et autres
Publié: (2022)
What can large language models do for sustainable food?
par: Thomas, Anna T., et autres
Publié: (2025)
par: Thomas, Anna T., et autres
Publié: (2025)
Cognitive models can reveal interpretable value trade-offs in language models
par: Murthy, Sonia K., et autres
Publié: (2025)
par: Murthy, Sonia K., et autres
Publié: (2025)
Investigating the interaction of linguistic and mathematical reasoning in language models using multilingual number puzzles
par: Bhattacharya, Antara Raaghavi, et autres
Publié: (2025)
par: Bhattacharya, Antara Raaghavi, et autres
Publié: (2025)
Social preferences with unstable interactive reasoning: Large language models in economic trust games
par: Jiamin, Ou, et autres
Publié: (2025)
par: Jiamin, Ou, et autres
Publié: (2025)
When can transformers reason with abstract symbols?
par: Boix-Adsera, Enric, et autres
Publié: (2023)
par: Boix-Adsera, Enric, et autres
Publié: (2023)
Predict the Next Word: Humans exhibit uncertainty in this task and language models _____
par: Ilia, Evgenia, et autres
Publié: (2024)
par: Ilia, Evgenia, et autres
Publié: (2024)
ClinicalGPT-R1: Pushing reasoning capability of generalist disease diagnosis with large language model
par: Lan, Wuyang, et autres
Publié: (2025)
par: Lan, Wuyang, et autres
Publié: (2025)
Multi-step retrieval and reasoning improves radiology question answering with large language models
par: Wind, Sebastian, et autres
Publié: (2025)
par: Wind, Sebastian, et autres
Publié: (2025)
Automated stereotactic radiosurgery planning using a human-in-the-loop reasoning large language model agent
par: Nusrat, Humza, et autres
Publié: (2025)
par: Nusrat, Humza, et autres
Publié: (2025)
CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
par: Yu, Ping, et autres
Publié: (2025)
par: Yu, Ping, et autres
Publié: (2025)
Vocabulary embeddings organize linguistic structure early in language model training
par: Papadimitriou, Isabel, et autres
Publié: (2025)
par: Papadimitriou, Isabel, et autres
Publié: (2025)
Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines
par: Jørgensen, Mikkel Godsk, et autres
Publié: (2026)
par: Jørgensen, Mikkel Godsk, et autres
Publié: (2026)
Functional Subspace, where language models can use vector algebra to solve problems
par: Lee, Jung H., et autres
Publié: (2026)
par: Lee, Jung H., et autres
Publié: (2026)
Documents similaires
-
Dissociating language and thought in large language models
par: Mahowald, Kyle, et autres
Publié: (2023) -
Superhuman performance of a large language model on the reasoning tasks of a physician
par: Brodeur, Peter G., et autres
Publié: (2024) -
Shoot First, Ask Questions Later? Building Rational Agents that Explore and Act Like People
par: Grand, Gabriel, et autres
Publié: (2025) -
Loose LIPS Sink Ships: Asking Questions in Battleship with Language-Informed Program Sampling
par: Grand, Gabriel, et autres
Publié: (2024) -
Language and Experience: A Computational Model of Social Learning in Complex Tasks
par: Colas, Cédric, et autres
Publié: (2025)