HeurekaBench: A Benchmarking Framework for AI Co-scientist
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Panigrahi, Siba Smarak, Videnović, Jovana, Brbić, Maria |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unsupervised Process Reward Models
von: Gadetsky, Artyom, et al.
Veröffentlicht: (2026)
von: Gadetsky, Artyom, et al.
Veröffentlicht: (2026)
Improved Canonicalization for Model Agnostic Equivariance
von: Panigrahi, Siba Smarak, et al.
Veröffentlicht: (2024)
von: Panigrahi, Siba Smarak, et al.
Veröffentlicht: (2024)
Efficient Dynamics Modeling in Interactive Environments with Koopman Theory
von: Mondal, Arnab Kumar, et al.
Veröffentlicht: (2023)
von: Mondal, Arnab Kumar, et al.
Veröffentlicht: (2023)
SymmCD: Symmetry-Preserving Crystal Generation with Diffusion Models
von: Levy, Daniel, et al.
Veröffentlicht: (2025)
von: Levy, Daniel, et al.
Veröffentlicht: (2025)
Cross-domain Open-world Discovery
von: Wen, Shuo, et al.
Veröffentlicht: (2024)
von: Wen, Shuo, et al.
Veröffentlicht: (2024)
Let Go of Your Labels with Unsupervised Transfer
von: Gadetsky, Artyom, et al.
Veröffentlicht: (2024)
von: Gadetsky, Artyom, et al.
Veröffentlicht: (2024)
NeuralBench: A Unifying Framework to Benchmark NeuroAI Models
von: Banville, Hubert, et al.
Veröffentlicht: (2026)
von: Banville, Hubert, et al.
Veröffentlicht: (2026)
Fine-grained Classes and How to Find Them
von: Grcić, Matej, et al.
Veröffentlicht: (2024)
von: Grcić, Matej, et al.
Veröffentlicht: (2024)
Weak-to-Strong Generalization under Distribution Shifts
von: Jeon, Myeongho, et al.
Veröffentlicht: (2025)
von: Jeon, Myeongho, et al.
Veröffentlicht: (2025)
Democratizing AI scientists using ToolUniverse
von: Gao, Shanghua, et al.
Veröffentlicht: (2025)
von: Gao, Shanghua, et al.
Veröffentlicht: (2025)
TopoBench: A Framework for Benchmarking Topological Deep Learning
von: Telyatnikov, Lev, et al.
Veröffentlicht: (2024)
von: Telyatnikov, Lev, et al.
Veröffentlicht: (2024)
Fast ML-driven Analog Circuit Layout using Reinforcement Learning and Steiner Trees
von: Basso, Davide, et al.
Veröffentlicht: (2024)
von: Basso, Davide, et al.
Veröffentlicht: (2024)
Revisiting the Platonic Representation Hypothesis: An Aristotelian View
von: Gröger, Fabian, et al.
Veröffentlicht: (2026)
von: Gröger, Fabian, et al.
Veröffentlicht: (2026)
NeuCo-Bench: A Novel Benchmark Framework for Neural Embeddings in Earth Observation
von: Vinge, Rikard, et al.
Veröffentlicht: (2025)
von: Vinge, Rikard, et al.
Veröffentlicht: (2025)
Meta-RL Induces Exploration in Language Agents
von: Jiang, Yulun, et al.
Veröffentlicht: (2025)
von: Jiang, Yulun, et al.
Veröffentlicht: (2025)
AI scientists produce results without reasoning scientifically
von: Ríos-García, Martiño, et al.
Veröffentlicht: (2026)
von: Ríos-García, Martiño, et al.
Veröffentlicht: (2026)
AICD Bench: A Challenging Benchmark for AI-Generated Code Detection
von: Orel, Daniil, et al.
Veröffentlicht: (2026)
von: Orel, Daniil, et al.
Veröffentlicht: (2026)
Introducing CausalBench: A Flexible Benchmark Framework for Causal Analysis and Machine Learning
von: Kapkiç, Ahmet, et al.
Veröffentlicht: (2024)
von: Kapkiç, Ahmet, et al.
Veröffentlicht: (2024)
ExplainBench: A Benchmark Framework for Local Model Explanations in Fairness-Critical Applications
von: Afful, James
Veröffentlicht: (2025)
von: Afful, James
Veröffentlicht: (2025)
With Limited Data for Multimodal Alignment, Let the STRUCTURE Guide You
von: Gröger, Fabian, et al.
Veröffentlicht: (2025)
von: Gröger, Fabian, et al.
Veröffentlicht: (2025)
LLM-Inference-Bench: Inference Benchmarking of Large Language Models on AI Accelerators
von: Chitty-Venkata, Krishna Teja, et al.
Veröffentlicht: (2024)
von: Chitty-Venkata, Krishna Teja, et al.
Veröffentlicht: (2024)
Advancing Routing-Awareness in Analog ICs Floorplanning
von: Basso, Davide, et al.
Veröffentlicht: (2025)
von: Basso, Davide, et al.
Veröffentlicht: (2025)
ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities
von: Karger, Ezra, et al.
Veröffentlicht: (2024)
von: Karger, Ezra, et al.
Veröffentlicht: (2024)
A Distractor-Aware Memory for Visual Object Tracking with SAM2
von: Videnovic, Jovana, et al.
Veröffentlicht: (2024)
von: Videnovic, Jovana, et al.
Veröffentlicht: (2024)
Towards an AI co-scientist
von: Gottweis, Juraj, et al.
Veröffentlicht: (2025)
von: Gottweis, Juraj, et al.
Veröffentlicht: (2025)
A Benchmarking Framework for AI models in Automotive Aerodynamics
von: Tangsali, Kaustubh, et al.
Veröffentlicht: (2025)
von: Tangsali, Kaustubh, et al.
Veröffentlicht: (2025)
YRC-Bench: A Benchmark for Learning to Coordinate with Experts
von: Danesh, Mohamad H., et al.
Veröffentlicht: (2025)
von: Danesh, Mohamad H., et al.
Veröffentlicht: (2025)
Constructing artificial life and materials scientists with accelerated AI using Deep AndersoNN
von: Dajani, Saleem Abdul Fattah Ahmed Al, et al.
Veröffentlicht: (2024)
von: Dajani, Saleem Abdul Fattah Ahmed Al, et al.
Veröffentlicht: (2024)
Effective Analog ICs Floorplanning with Relational Graph Neural Networks and Reinforcement Learning
von: Basso, Davide, et al.
Veröffentlicht: (2024)
von: Basso, Davide, et al.
Veröffentlicht: (2024)
BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices
von: Reuel, Anka, et al.
Veröffentlicht: (2024)
von: Reuel, Anka, et al.
Veröffentlicht: (2024)
Large (Vision) Language Models are Unsupervised In-Context Learners
von: Gadetsky, Artyom, et al.
Veröffentlicht: (2025)
von: Gadetsky, Artyom, et al.
Veröffentlicht: (2025)
Distractor-Aware Memory-Based Visual Object Tracking
von: Videnovic, Jovana, et al.
Veröffentlicht: (2025)
von: Videnovic, Jovana, et al.
Veröffentlicht: (2025)
MergeBench: A Benchmark for Merging Domain-Specialized LLMs
von: He, Yifei, et al.
Veröffentlicht: (2025)
von: He, Yifei, et al.
Veröffentlicht: (2025)
DHG-Bench: A Comprehensive Benchmark for Deep Hypergraph Learning
von: Li, Fan, et al.
Veröffentlicht: (2025)
von: Li, Fan, et al.
Veröffentlicht: (2025)
HinTel-AlignBench: A Framework and Benchmark for Hindi-Telugu with English-Aligned Samples
von: Chigrupaatii, Rishikant, et al.
Veröffentlicht: (2025)
von: Chigrupaatii, Rishikant, et al.
Veröffentlicht: (2025)
UserSumBench: A Benchmark Framework for Evaluating User Summarization Approaches
von: Wang, Chao, et al.
Veröffentlicht: (2024)
von: Wang, Chao, et al.
Veröffentlicht: (2024)
CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2024)
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2024)
PeopleSearchBench: A Multi-Dimensional Benchmark for Evaluating AI-Powered People Search Platforms
von: Wang, Wei, et al.
Veröffentlicht: (2026)
von: Wang, Wei, et al.
Veröffentlicht: (2026)
LJ-Bench: Ontology-Based Benchmark for U.S. Crime
von: Tseng, Hung Yun, et al.
Veröffentlicht: (2026)
von: Tseng, Hung Yun, et al.
Veröffentlicht: (2026)
ODP-Bench: Benchmarking Out-of-Distribution Performance Prediction
von: Yu, Han, et al.
Veröffentlicht: (2025)
von: Yu, Han, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Unsupervised Process Reward Models
von: Gadetsky, Artyom, et al.
Veröffentlicht: (2026) -
Improved Canonicalization for Model Agnostic Equivariance
von: Panigrahi, Siba Smarak, et al.
Veröffentlicht: (2024) -
Efficient Dynamics Modeling in Interactive Environments with Koopman Theory
von: Mondal, Arnab Kumar, et al.
Veröffentlicht: (2023) -
SymmCD: Symmetry-Preserving Crystal Generation with Diffusion Models
von: Levy, Daniel, et al.
Veröffentlicht: (2025) -
Cross-domain Open-world Discovery
von: Wen, Shuo, et al.
Veröffentlicht: (2024)