How Far Are AI Scientists from Changing the World?
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Qiujie, Weng, Yixuan, Zhu, Minjun, Shen, Fuchen, Huang, Shulin, Lin, Zhen, Zhou, Jiahui, Mao, Zilan, Yang, Zijie, Yang, Linyi, Wu, Jian, Zhang, Yue |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI Scientists Fail Without Strong Implementation Capability
by: Zhu, Minjun, et al.
Published: (2025)
by: Zhu, Minjun, et al.
Published: (2025)
DeepScientist: Advancing Frontier-Pushing Scientific Findings Progressively
by: Weng, Yixuan, et al.
Published: (2025)
by: Weng, Yixuan, et al.
Published: (2025)
Personality Alignment of Large Language Models
by: Zhu, Minjun, et al.
Published: (2024)
by: Zhu, Minjun, et al.
Published: (2024)
DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process
by: Zhu, Minjun, et al.
Published: (2025)
by: Zhu, Minjun, et al.
Published: (2025)
Linear Reasoning vs. Proof by Cases: Obstacles for Large Language Models in FOL Problem Solving
by: Ji, Yuliang, et al.
Published: (2026)
by: Ji, Yuliang, et al.
Published: (2026)
ResearStudio: A Human-Intervenable Framework for Building Controllable Deep-Research Agents
by: Yang, Linyi, et al.
Published: (2025)
by: Yang, Linyi, et al.
Published: (2025)
AiraXiv: An AI-Driven Open-Access Platform for Human and AI Scientists
by: Pan, Junshu, et al.
Published: (2026)
by: Pan, Junshu, et al.
Published: (2026)
AutoFigure: Generating and Refining Publication-Ready Scientific Illustrations
by: Zhu, Minjun, et al.
Published: (2026)
by: Zhu, Minjun, et al.
Published: (2026)
CycleResearcher: Improving Automated Research via Automated Review
by: Weng, Yixuan, et al.
Published: (2024)
by: Weng, Yixuan, et al.
Published: (2024)
DeepReviewer 2.0: A Traceable Agentic System for Auditable Scientific Peer Review
by: Weng, Yixuan, et al.
Published: (2026)
by: Weng, Yixuan, et al.
Published: (2026)
MCEval: A Dynamic Framework for Fair Multilingual Cultural Evaluation of LLMs
by: Huang, Shulin, et al.
Published: (2025)
by: Huang, Shulin, et al.
Published: (2025)
AI-Generated Text is Non-Stationary: Detection via Temporal Tomography
by: West, Alva, et al.
Published: (2025)
by: West, Alva, et al.
Published: (2025)
SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts
by: Xin, Yuan, et al.
Published: (2026)
by: Xin, Yuan, et al.
Published: (2026)
Abduct, Act, Predict: Scaffolding Causal Inference for Automated Failure Attribution in Multi-Agent Systems
by: West, Alva, et al.
Published: (2025)
by: West, Alva, et al.
Published: (2025)
Locking Down the Finetuned LLMs Safety
by: Zhu, Minjun, et al.
Published: (2024)
by: Zhu, Minjun, et al.
Published: (2024)
An Empirical Analysis of Uncertainty in Large Language Model Evaluations
by: Xie, Qiujie, et al.
Published: (2025)
by: Xie, Qiujie, et al.
Published: (2025)
AutoFigure-Edit: Generating Editable Scientific Illustration
by: Lin, Zhen, et al.
Published: (2026)
by: Lin, Zhen, et al.
Published: (2026)
A Built-in Crypto Expert for Artificial Intelligence: How Far is the Horizon?
by: Weng, Jiasi, et al.
Published: (2026)
by: Weng, Jiasi, et al.
Published: (2026)
Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?
by: Xin, Yuan, et al.
Published: (2025)
by: Xin, Yuan, et al.
Published: (2025)
How Far is Video Generation from World Model: A Physical Law Perspective
by: Kang, Bingyi, et al.
Published: (2024)
by: Kang, Bingyi, et al.
Published: (2024)
T-Detect: Tail-Aware Statistical Normalization for Robust Detection of Adversarial Machine-Generated Text
by: West, Alva, et al.
Published: (2025)
by: West, Alva, et al.
Published: (2025)
Cofca: A Step-Wise Counterfactual Multi-hop QA benchmark
by: Wu, Jian, et al.
Published: (2024)
by: Wu, Jian, et al.
Published: (2024)
Triple Path Enhanced Neural Architecture Search for Multimodal Fake News Detection
by: Xu, Bo, et al.
Published: (2025)
by: Xu, Bo, et al.
Published: (2025)
Early‐Life Famine Experience and Households’ Diversity of Financial Asset Portfolio: Evidence From China
by: Shulin Xu, et al.
Published: (2025)
by: Shulin Xu, et al.
Published: (2025)
LingGym: How Far Are LLMs from Thinking Like Field Linguists?
by: Yang, Changbing, et al.
Published: (2025)
by: Yang, Changbing, et al.
Published: (2025)
Towards a Medical AI Scientist
by: Wu, Hongtao, et al.
Published: (2026)
by: Wu, Hongtao, et al.
Published: (2026)
How Likely Do LLMs with CoT Mimic Human Reasoning?
by: Bao, Guangsheng, et al.
Published: (2024)
by: Bao, Guangsheng, et al.
Published: (2024)
OmniScientist: Toward a Co-evolving Ecosystem of Human and AI Scientists
by: Shao, Chenyang, et al.
Published: (2025)
by: Shao, Chenyang, et al.
Published: (2025)
How Far are AI-generated Videos from Simulating the 3D Visual World: A Learned 3D Evaluation Approach
by: Chang, Chirui, et al.
Published: (2024)
by: Chang, Chirui, et al.
Published: (2024)
How Far Are VLMs from Privacy Awareness in the Physical World? An Empirical Study
by: Wang, Junran, et al.
Published: (2026)
by: Wang, Junran, et al.
Published: (2026)
The AI Hippocampus: How Far are We From Human Memory?
by: Jia, Zixia, et al.
Published: (2026)
by: Jia, Zixia, et al.
Published: (2026)
How Far Has AI Come in Liver Fibrosis Staging? A Large-Scale Real-World Dataset and Benchmark
by: Liu, Yuanye, et al.
Published: (2026)
by: Liu, Yuanye, et al.
Published: (2026)
Managerial Nationalism and Firm Innovation: Social Identity Perspective
by: Zhen Yang, et al.
Published: (2025)
by: Zhen Yang, et al.
Published: (2025)
Delving into the Reversal Curse: How Far Can Large Language Models Generalize?
by: Lin, Zhengkai, et al.
Published: (2024)
by: Lin, Zhengkai, et al.
Published: (2024)
Mastering Symbolic Operations: Augmenting Language Models with Compiled Neural Networks
by: Weng, Yixuan, et al.
Published: (2023)
by: Weng, Yixuan, et al.
Published: (2023)
Orthosis Management in Knee Osteoarthritis: Evaluating Existing Recommendations and Achieving Consensus on Implementation Through the Delphi Method
by: Zilan Bazancir‐Apaydin
Published: (2024)
by: Zilan Bazancir‐Apaydin
Published: (2024)
Automatically Detecting Checked-In Secrets in Android Apps: How Far Are We?
by: Li, Kevin, et al.
Published: (2024)
by: Li, Kevin, et al.
Published: (2024)
Seismic vulnerability analysis of a high‐rise steel frame structure equipped with novel displacement‐amplified viscoelastic dampers
by: Mao Ye, et al.
Published: (2024)
by: Mao Ye, et al.
Published: (2024)
LLMs for Relational Reasoning: How Far are We?
by: Li, Zhiming, et al.
Published: (2024)
by: Li, Zhiming, et al.
Published: (2024)
Water Infrastructure and Grain Yield: Evidence From the South‐to‐North Water Diversion Project
by: Jindong Pang, et al.
Published: (2025)
by: Jindong Pang, et al.
Published: (2025)
Similar Items
-
AI Scientists Fail Without Strong Implementation Capability
by: Zhu, Minjun, et al.
Published: (2025) -
DeepScientist: Advancing Frontier-Pushing Scientific Findings Progressively
by: Weng, Yixuan, et al.
Published: (2025) -
Personality Alignment of Large Language Models
by: Zhu, Minjun, et al.
Published: (2024) -
DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process
by: Zhu, Minjun, et al.
Published: (2025) -
Linear Reasoning vs. Proof by Cases: Obstacles for Large Language Models in FOL Problem Solving
by: Ji, Yuliang, et al.
Published: (2026)