The AI Data Scientist
Fuente:
arXiv
Saved in:
| Main Authors: | Akimov, Farkhad, Nwadike, Munachiso Samuel, Iklassov, Zangir, Takáč, Martin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Measuring AI Reasoning: A Guide for Researchers
by: Nwadike, Munachiso Samuel, et al.
Published: (2026)
by: Nwadike, Munachiso Samuel, et al.
Published: (2026)
Self-Guiding Exploration for Combinatorial Problems
by: Iklassov, Zangir, et al.
Published: (2024)
by: Iklassov, Zangir, et al.
Published: (2024)
RECALL: Library-Like Behavior In Language Models is Enhanced by Self-Referencing Causal Cycles
by: Nwadike, Munachiso, et al.
Published: (2025)
by: Nwadike, Munachiso, et al.
Published: (2025)
SVRPBench: A Realistic Benchmark for Stochastic Vehicle Routing Problem
by: Heakl, Ahmed, et al.
Published: (2025)
by: Heakl, Ahmed, et al.
Published: (2025)
Improving Personalisation in Valence and Arousal Prediction using Data Augmentation
by: Nwadike, Munachiso, et al.
Published: (2024)
by: Nwadike, Munachiso, et al.
Published: (2024)
LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs
by: Choukrani, Omar, et al.
Published: (2025)
by: Choukrani, Omar, et al.
Published: (2025)
Reinforcement Learning for Solving Stochastic Vehicle Routing Problem with Time Windows
by: Iklassov, Zangir, et al.
Published: (2024)
by: Iklassov, Zangir, et al.
Published: (2024)
AI Scientists Fail Without Strong Implementation Capability
by: Zhu, Minjun, et al.
Published: (2025)
by: Zhu, Minjun, et al.
Published: (2025)
AiraXiv: An AI-Driven Open-Access Platform for Human and AI Scientists
by: Pan, Junshu, et al.
Published: (2026)
by: Pan, Junshu, et al.
Published: (2026)
Latent Reasoning in TRMs is Secretly a Policy Improvement Operator
by: Asadulaev, Arip, et al.
Published: (2025)
by: Asadulaev, Arip, et al.
Published: (2025)
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
by: Lu, Chris, et al.
Published: (2024)
by: Lu, Chris, et al.
Published: (2024)
CauScientist: Teaching LLMs to Respect Data for Causal Discovery
by: Peng, Bo, et al.
Published: (2026)
by: Peng, Bo, et al.
Published: (2026)
Sycophancy Hides Linearly in the Attention Heads
by: Genadi, Rifo, et al.
Published: (2026)
by: Genadi, Rifo, et al.
Published: (2026)
Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy
by: Tang, Xiangru, et al.
Published: (2024)
by: Tang, Xiangru, et al.
Published: (2024)
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems
by: Geng, Jiayi, et al.
Published: (2025)
by: Geng, Jiayi, et al.
Published: (2025)
The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
by: Yamada, Yutaro, et al.
Published: (2025)
by: Yamada, Yutaro, et al.
Published: (2025)
Mechanistic Insights into Grokking from the Embedding Layer
by: AlquBoj, H. V., et al.
Published: (2025)
by: AlquBoj, H. V., et al.
Published: (2025)
Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
by: Xu, Wanghan, et al.
Published: (2025)
by: Xu, Wanghan, et al.
Published: (2025)
BEExAI: Benchmark to Evaluate Explainable AI
by: Sithakoul, Samuel, et al.
Published: (2024)
by: Sithakoul, Samuel, et al.
Published: (2024)
Automated test generation to evaluate tool-augmented LLMs as conversational AI agents
by: Arcadinho, Samuel, et al.
Published: (2024)
by: Arcadinho, Samuel, et al.
Published: (2024)
Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
by: Miyai, Atsuyuki, et al.
Published: (2025)
by: Miyai, Atsuyuki, et al.
Published: (2025)
Consent in Crisis: The Rapid Decline of the AI Data Commons
by: Longpre, Shayne, et al.
Published: (2024)
by: Longpre, Shayne, et al.
Published: (2024)
Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
by: Tice, Cameron, et al.
Published: (2026)
by: Tice, Cameron, et al.
Published: (2026)
Speaking the Same Language: Leveraging LLMs in Standardizing Clinical Data for AI
by: Sett, Arindam, et al.
Published: (2024)
by: Sett, Arindam, et al.
Published: (2024)
Mitigating Watermark Forgery in Generative Models via Randomized Key Selection
by: Aremu, Toluwani, et al.
Published: (2025)
by: Aremu, Toluwani, et al.
Published: (2025)
Data Agent: A Holistic Architecture for Orchestrating Data+AI Ecosystems
by: Sun, Zhaoyan, et al.
Published: (2025)
by: Sun, Zhaoyan, et al.
Published: (2025)
FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts
by: Shin, Hagyeong, et al.
Published: (2025)
by: Shin, Hagyeong, et al.
Published: (2025)
Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy
by: Liu, Chris Yuhao, et al.
Published: (2025)
by: Liu, Chris Yuhao, et al.
Published: (2025)
Quantization Dominates Rank Reduction for KV-Cache Compression
by: Salfati, Samuel
Published: (2026)
by: Salfati, Samuel
Published: (2026)
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
by: Toshniwal, Shubham, et al.
Published: (2024)
by: Toshniwal, Shubham, et al.
Published: (2024)
When Bad Data Leads to Good Models
by: Li, Kenneth, et al.
Published: (2025)
by: Li, Kenneth, et al.
Published: (2025)
Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
by: Treutlein, Johannes, et al.
Published: (2024)
by: Treutlein, Johannes, et al.
Published: (2024)
LABBench2: An Improved Benchmark for AI Systems Performing Biology Research
by: Laurent, Jon M, et al.
Published: (2026)
by: Laurent, Jon M, et al.
Published: (2026)
Can Interpretation Predict Behavior on Unseen Data?
by: Li, Victoria R., et al.
Published: (2025)
by: Li, Victoria R., et al.
Published: (2025)
Impacts of Racial Bias in Historical Training Data for News AI
by: Bhargava, Rahul, et al.
Published: (2025)
by: Bhargava, Rahul, et al.
Published: (2025)
AI and Generative AI for Research Discovery and Summarization
by: Glickman, Mark, et al.
Published: (2024)
by: Glickman, Mark, et al.
Published: (2024)
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
by: van der Weij, Teun, et al.
Published: (2024)
by: van der Weij, Teun, et al.
Published: (2024)
AgentRxiv: Towards Collaborative Autonomous Research
by: Schmidgall, Samuel, et al.
Published: (2025)
by: Schmidgall, Samuel, et al.
Published: (2025)
Robustly Improving LLM Fairness in Realistic Settings via Interpretability
by: Karvonen, Adam, et al.
Published: (2025)
by: Karvonen, Adam, et al.
Published: (2025)
AI-rithmetic
by: Bie, Alex, et al.
Published: (2026)
by: Bie, Alex, et al.
Published: (2026)
Similar Items
-
Measuring AI Reasoning: A Guide for Researchers
by: Nwadike, Munachiso Samuel, et al.
Published: (2026) -
Self-Guiding Exploration for Combinatorial Problems
by: Iklassov, Zangir, et al.
Published: (2024) -
RECALL: Library-Like Behavior In Language Models is Enhanced by Self-Referencing Causal Cycles
by: Nwadike, Munachiso, et al.
Published: (2025) -
SVRPBench: A Realistic Benchmark for Stochastic Vehicle Routing Problem
by: Heakl, Ahmed, et al.
Published: (2025) -
Improving Personalisation in Valence and Arousal Prediction using Data Augmentation
by: Nwadike, Munachiso, et al.
Published: (2024)