Why LLMs Aren't Scientists Yet: Lessons from Four Autonomous Research Attempts
Fuente:
arXiv
Salvato in:
| Autori principali: | Trehan, Dhruv, Chopra, Paras |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
METIS: Mentoring Engine for Thoughtful Inquiry & Solutions
di: Kumar, Abhinav Rajeev, et al.
Pubblicazione: (2026)
di: Kumar, Abhinav Rajeev, et al.
Pubblicazione: (2026)
The Sequential Edge: Inverse-Entropy Voting Beats Parallel Self-Consistency at Matched Compute
di: Sharma, Aman, et al.
Pubblicazione: (2025)
di: Sharma, Aman, et al.
Pubblicazione: (2025)
Think Just Enough: Sequence-Level Entropy as a Confidence Signal for LLM Reasoning
di: Sharma, Aman, et al.
Pubblicazione: (2025)
di: Sharma, Aman, et al.
Pubblicazione: (2025)
EsoLang-Bench: Evaluating Genuine Reasoning in Large Language Models via Esoteric Programming Languages
di: Sharma, Aman, et al.
Pubblicazione: (2026)
di: Sharma, Aman, et al.
Pubblicazione: (2026)
Hybrid Neural World Models
di: Lakshmanan, Pranav, et al.
Pubblicazione: (2026)
di: Lakshmanan, Pranav, et al.
Pubblicazione: (2026)
Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability
di: Winata, Genta Indra, et al.
Pubblicazione: (2025)
di: Winata, Genta Indra, et al.
Pubblicazione: (2025)
Discovering Reinforcement Learning Interfaces with Large Language Models
di: Jaswal, Akshat Singh, et al.
Pubblicazione: (2026)
di: Jaswal, Akshat Singh, et al.
Pubblicazione: (2026)
Building Interpretable Models for Moral Decision-Making
di: Goel, Mayank, et al.
Pubblicazione: (2026)
di: Goel, Mayank, et al.
Pubblicazione: (2026)
Who Said Neural Networks Aren't Linear?
di: Berman, Nimrod, et al.
Pubblicazione: (2025)
di: Berman, Nimrod, et al.
Pubblicazione: (2025)
Detecting Jailbreak Attempts in Clinical Training LLMs Through Automated Linguistic Feature Extraction
di: Nguyen, Tri, et al.
Pubblicazione: (2026)
di: Nguyen, Tri, et al.
Pubblicazione: (2026)
ALAS: Autonomous Learning Agent for Self-Updating Language Models
di: Atreja, Dhruv
Pubblicazione: (2025)
di: Atreja, Dhruv
Pubblicazione: (2025)
Scaling Is All You Need: Autonomous Driving with JAX-Accelerated Reinforcement Learning
di: Harmel, Moritz, et al.
Pubblicazione: (2023)
di: Harmel, Moritz, et al.
Pubblicazione: (2023)
Future Is Unevenly Distributed: Forecasting Ability of LLMs Depends on What We're Asking
di: Karkar, Chinmay, et al.
Pubblicazione: (2025)
di: Karkar, Chinmay, et al.
Pubblicazione: (2025)
Zero-Shot Confidence Estimation for Small LLMs: When Supervised Baselines Aren't Worth Training
di: Nguyen, Luong N.
Pubblicazione: (2026)
di: Nguyen, Luong N.
Pubblicazione: (2026)
CauScientist: Teaching LLMs to Respect Data for Causal Discovery
di: Peng, Bo, et al.
Pubblicazione: (2026)
di: Peng, Bo, et al.
Pubblicazione: (2026)
Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
di: Xu, Wanghan, et al.
Pubblicazione: (2025)
di: Xu, Wanghan, et al.
Pubblicazione: (2025)
Stateless Yet Not Forgetful: Implicit Memory as a Hidden Channel in LLMs
di: Salem, Ahmed, et al.
Pubblicazione: (2026)
di: Salem, Ahmed, et al.
Pubblicazione: (2026)
Towards a Medical AI Scientist
di: Wu, Hongtao, et al.
Pubblicazione: (2026)
di: Wu, Hongtao, et al.
Pubblicazione: (2026)
AlphaLab: Autonomous Multi-Agent Research Across Optimization Domains with Frontier LLMs
di: Hogan, Brendan R., et al.
Pubblicazione: (2026)
di: Hogan, Brendan R., et al.
Pubblicazione: (2026)
Do Two AI Scientists Agree?
di: Fu, Xinghong, et al.
Pubblicazione: (2025)
di: Fu, Xinghong, et al.
Pubblicazione: (2025)
Learning to Correct: Calibrated Reinforcement Learning for Multi-Attempt Chain-of-Thought
di: Ildiz, Muhammed Emrullah, et al.
Pubblicazione: (2026)
di: Ildiz, Muhammed Emrullah, et al.
Pubblicazione: (2026)
Robust Yet Efficient Conformal Prediction Sets
di: Zargarbashi, Soroush H., et al.
Pubblicazione: (2024)
di: Zargarbashi, Soroush H., et al.
Pubblicazione: (2024)
Learning from Failures in Multi-Attempt Reinforcement Learning
di: Chung, Stephen, et al.
Pubblicazione: (2025)
di: Chung, Stephen, et al.
Pubblicazione: (2025)
Why Aren't There More Whistleblowers?
di: Robert A. Aronowitz
Pubblicazione: (2024)
di: Robert A. Aronowitz
Pubblicazione: (2024)
MaD-Scientist: AI-based Scientist solving Convection-Diffusion-Reaction Equations Using Massive PINN-Based Prior Data
di: Kang, Mingu, et al.
Pubblicazione: (2024)
di: Kang, Mingu, et al.
Pubblicazione: (2024)
The AI Data Scientist
di: Akimov, Farkhad, et al.
Pubblicazione: (2025)
di: Akimov, Farkhad, et al.
Pubblicazione: (2025)
Wiring the 'Why': A Unified Taxonomy and Survey of Abductive Reasoning in LLMs
di: Salimi, Moein, et al.
Pubblicazione: (2026)
di: Salimi, Moein, et al.
Pubblicazione: (2026)
Discovering Spoofing Attempts on Language Model Watermarks
di: Gloaguen, Thibaud, et al.
Pubblicazione: (2024)
di: Gloaguen, Thibaud, et al.
Pubblicazione: (2024)
Causal Feature Selection for Responsible Machine Learning
di: Moraffah, Raha, et al.
Pubblicazione: (2024)
di: Moraffah, Raha, et al.
Pubblicazione: (2024)
Geometry of Human Perceptual Domains Emerges Transiently in LLM Representations
di: Singh, Simardeep, et al.
Pubblicazione: (2026)
di: Singh, Simardeep, et al.
Pubblicazione: (2026)
See, Symbolize, Act: Grounding VLMs with Spatial Representations for Better Gameplay
di: Baghel, Ashish, et al.
Pubblicazione: (2026)
di: Baghel, Ashish, et al.
Pubblicazione: (2026)
MILP-SAT-GNN: Yet Another Neural SAT Solver
di: Cardillo, Franco Alberto, et al.
Pubblicazione: (2025)
di: Cardillo, Franco Alberto, et al.
Pubblicazione: (2025)
Can LLMs Effectively Leverage Graph Structural Information through Prompts, and Why?
di: Huang, Jin, et al.
Pubblicazione: (2023)
di: Huang, Jin, et al.
Pubblicazione: (2023)
Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks
di: Li, Ang, et al.
Pubblicazione: (2025)
di: Li, Ang, et al.
Pubblicazione: (2025)
Not Yet AlphaFold for the Mind: Evaluating Centaur as a Synthetic Participant
di: Namazova, Sabrina, et al.
Pubblicazione: (2025)
di: Namazova, Sabrina, et al.
Pubblicazione: (2025)
ARIES: Autonomous Reasoning with LLMs on Interactive Thought Graph Environments
di: Gimenes, Pedro, et al.
Pubblicazione: (2025)
di: Gimenes, Pedro, et al.
Pubblicazione: (2025)
When Visuals Aren't the Problem: Evaluating Vision-Language Models on Misleading Data Visualizations
di: Lalai, Harsh Nishant, et al.
Pubblicazione: (2026)
di: Lalai, Harsh Nishant, et al.
Pubblicazione: (2026)
Retrieval Models Aren't Tool-Savvy: Benchmarking Tool Retrieval for Large Language Models
di: Shi, Zhengliang, et al.
Pubblicazione: (2025)
di: Shi, Zhengliang, et al.
Pubblicazione: (2025)
ThinkSwitch: Context Distillation with LoRA and Weight Interpolation for Specific-Purpose Reasoning Tasks
di: Saini, Dhruv, et al.
Pubblicazione: (2026)
di: Saini, Dhruv, et al.
Pubblicazione: (2026)
Collaborative Yet Personalized Policy Training: Single-Timescale Federated Actor-Critic
di: Wang, Leo Muxing, et al.
Pubblicazione: (2026)
di: Wang, Leo Muxing, et al.
Pubblicazione: (2026)
Documenti analoghi
-
METIS: Mentoring Engine for Thoughtful Inquiry & Solutions
di: Kumar, Abhinav Rajeev, et al.
Pubblicazione: (2026) -
The Sequential Edge: Inverse-Entropy Voting Beats Parallel Self-Consistency at Matched Compute
di: Sharma, Aman, et al.
Pubblicazione: (2025) -
Think Just Enough: Sequence-Level Entropy as a Confidence Signal for LLM Reasoning
di: Sharma, Aman, et al.
Pubblicazione: (2025) -
EsoLang-Bench: Evaluating Genuine Reasoning in Large Language Models via Esoteric Programming Languages
di: Sharma, Aman, et al.
Pubblicazione: (2026) -
Hybrid Neural World Models
di: Lakshmanan, Pranav, et al.
Pubblicazione: (2026)