Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tang, Xiangru, Jin, Qiao, Zhu, Kunlun, Yuan, Tongxin, Zhang, Yichi, Zhou, Wangchunshu, Qu, Meng, Zhao, Yilun, Tang, Jian, Zhang, Zhuosheng, Cohan, Arman, Lu, Zhiyong, Gerstein, Mark |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data?
von: Tang, Xiangru, et al.
Veröffentlicht: (2023)
von: Tang, Xiangru, et al.
Veröffentlicht: (2023)
Investigating Data Contamination in Modern Benchmarks for Large Language Models
von: Deng, Chunyuan, et al.
Veröffentlicht: (2023)
von: Deng, Chunyuan, et al.
Veröffentlicht: (2023)
MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning
von: Tang, Xiangru, et al.
Veröffentlicht: (2023)
von: Tang, Xiangru, et al.
Veröffentlicht: (2023)
MIMIR: A Streamlined Platform for Personalized Agent Tuning in Domain Expertise
von: Deng, Chunyuan, et al.
Veröffentlicht: (2024)
von: Deng, Chunyuan, et al.
Veröffentlicht: (2024)
ChemAgent: Self-updating Library in Large Language Models Improves Chemical Reasoning
von: Tang, Xiangru, et al.
Veröffentlicht: (2025)
von: Tang, Xiangru, et al.
Veröffentlicht: (2025)
Step-Back Profiling: Distilling User History for Personalized Scientific Writing
von: Tang, Xiangru, et al.
Veröffentlicht: (2024)
von: Tang, Xiangru, et al.
Veröffentlicht: (2024)
Unveiling the Spectrum of Data Contamination in Language Models: A Survey from Detection to Remediation
von: Deng, Chunyuan, et al.
Veröffentlicht: (2024)
von: Deng, Chunyuan, et al.
Veröffentlicht: (2024)
MedAgentsBench: Benchmarking Thinking Models and Agent Frameworks for Complex Medical Reasoning
von: Tang, Xiangru, et al.
Veröffentlicht: (2025)
von: Tang, Xiangru, et al.
Veröffentlicht: (2025)
ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain
von: Zhao, Haochen, et al.
Veröffentlicht: (2024)
von: Zhao, Haochen, et al.
Veröffentlicht: (2024)
Generalizable Chain-of-Thought Prompting in Mixed-task Scenarios with Large Language Models
von: Zou, Anni, et al.
Veröffentlicht: (2023)
von: Zou, Anni, et al.
Veröffentlicht: (2023)
DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Long and Specialized Documents
von: Zhao, Yilun, et al.
Veröffentlicht: (2023)
von: Zhao, Yilun, et al.
Veröffentlicht: (2023)
Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
von: Yu, Zhaojian, et al.
Veröffentlicht: (2024)
von: Yu, Zhaojian, et al.
Veröffentlicht: (2024)
FinDVer: Explainable Claim Verification over Long and Hybrid-Content Financial Documents
von: Zhao, Yilun, et al.
Veröffentlicht: (2024)
von: Zhao, Yilun, et al.
Veröffentlicht: (2024)
ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code
von: Tang, Xiangru, et al.
Veröffentlicht: (2023)
von: Tang, Xiangru, et al.
Veröffentlicht: (2023)
MCTS-RAG: Enhancing Retrieval-Augmented Generation with Monte Carlo Tree Search
von: Hu, Yunhai, et al.
Veröffentlicht: (2025)
von: Hu, Yunhai, et al.
Veröffentlicht: (2025)
Table-R1: Inference-Time Scaling for Table Reasoning
von: Yang, Zheyuan, et al.
Veröffentlicht: (2025)
von: Yang, Zheyuan, et al.
Veröffentlicht: (2025)
D-Flow: Multi-modality Flow Matching for D-peptide Design
von: Wu, Fang, et al.
Veröffentlicht: (2024)
von: Wu, Fang, et al.
Veröffentlicht: (2024)
SUCEA: Reasoning-Intensive Retrieval for Adversarial Fact-checking through Claim Decomposition and Editing
von: Liu, Hongjun, et al.
Veröffentlicht: (2025)
von: Liu, Hongjun, et al.
Veröffentlicht: (2025)
LimRank: Less is More for Reasoning-Intensive Information Reranking
von: Song, Tingyu, et al.
Veröffentlicht: (2025)
von: Song, Tingyu, et al.
Veröffentlicht: (2025)
SAGE: Benchmarking and Improving Retrieval for Deep Research Agents
von: Hu, Tiansheng, et al.
Veröffentlicht: (2026)
von: Hu, Tiansheng, et al.
Veröffentlicht: (2026)
Z1: Efficient Test-time Scaling with Code
von: Yu, Zhaojian, et al.
Veröffentlicht: (2025)
von: Yu, Zhaojian, et al.
Veröffentlicht: (2025)
BioCoder: A Benchmark for Bioinformatics Code Generation with Large Language Models
von: Tang, Xiangru, et al.
Veröffentlicht: (2023)
von: Tang, Xiangru, et al.
Veröffentlicht: (2023)
FinanceMath: Knowledge-Intensive Math Reasoning in Finance Domains
von: Zhao, Yilun, et al.
Veröffentlicht: (2023)
von: Zhao, Yilun, et al.
Veröffentlicht: (2023)
Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers
von: Xu, Zhijian, et al.
Veröffentlicht: (2025)
von: Xu, Zhijian, et al.
Veröffentlicht: (2025)
Patient-Similarity Cohort Reasoning in Clinical Text-to-SQL
von: Shen, Yifei, et al.
Veröffentlicht: (2026)
von: Shen, Yifei, et al.
Veröffentlicht: (2026)
SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification
von: Wang, Chengye, et al.
Veröffentlicht: (2025)
von: Wang, Chengye, et al.
Veröffentlicht: (2025)
FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
von: Long, Yitao, et al.
Veröffentlicht: (2025)
von: Long, Yitao, et al.
Veröffentlicht: (2025)
Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems
von: Zhao, Yilun, et al.
Veröffentlicht: (2026)
von: Zhao, Yilun, et al.
Veröffentlicht: (2026)
AlphaResearch: Accelerating New Algorithm Discovery with Language Models
von: Yu, Zhaojian, et al.
Veröffentlicht: (2025)
von: Yu, Zhaojian, et al.
Veröffentlicht: (2025)
Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective
von: Zhang, Siyue, et al.
Veröffentlicht: (2025)
von: Zhang, Siyue, et al.
Veröffentlicht: (2025)
SciRAG: Adaptive, Citation-Aware, and Outline-Guided Retrieval and Synthesis for Scientific Literature
von: Ding, Hang, et al.
Veröffentlicht: (2025)
von: Ding, Hang, et al.
Veröffentlicht: (2025)
Observable Propagation: Uncovering Feature Vectors in Transformers
von: Dunefsky, Jacob, et al.
Veröffentlicht: (2023)
von: Dunefsky, Jacob, et al.
Veröffentlicht: (2023)
SciMDR: Advancing Scientific Multimodal Document Reasoning
von: Chen, Ziyu, et al.
Veröffentlicht: (2026)
von: Chen, Ziyu, et al.
Veröffentlicht: (2026)
ChatCell: Facilitating Single-Cell Analysis with Natural Language
von: Fang, Yin, et al.
Veröffentlicht: (2024)
von: Fang, Yin, et al.
Veröffentlicht: (2024)
MSRS: Evaluating Multi-Source Retrieval-Augmented Generation
von: Phanse, Rohan, et al.
Veröffentlicht: (2025)
von: Phanse, Rohan, et al.
Veröffentlicht: (2025)
Agent-ScanKit: Unraveling Memory and Reasoning of Multimodal Agents via Sensitivity Perturbations
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2025)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2025)
Perturbation Analysis of Randomized SVD and its Applications to Statistics
von: Zhang, Yichi, et al.
Veröffentlicht: (2022)
von: Zhang, Yichi, et al.
Veröffentlicht: (2022)
CellForge: Agentic Design of Virtual Cell Models
von: Tang, Xiangru, et al.
Veröffentlicht: (2025)
von: Tang, Xiangru, et al.
Veröffentlicht: (2025)
LocAgent: Graph-Guided LLM Agents for Code Localization
von: Chen, Zhaoling, et al.
Veröffentlicht: (2025)
von: Chen, Zhaoling, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data?
von: Tang, Xiangru, et al.
Veröffentlicht: (2023) -
Investigating Data Contamination in Modern Benchmarks for Large Language Models
von: Deng, Chunyuan, et al.
Veröffentlicht: (2023) -
MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning
von: Tang, Xiangru, et al.
Veröffentlicht: (2023) -
MIMIR: A Streamlined Platform for Personalized Agent Tuning in Domain Expertise
von: Deng, Chunyuan, et al.
Veröffentlicht: (2024) -
ChemAgent: Self-updating Library in Large Language Models Improves Chemical Reasoning
von: Tang, Xiangru, et al.
Veröffentlicht: (2025)