LABBench2: An Improved Benchmark for AI Systems Performing Biology Research
Fuente:
arXiv
Salvato in:
| Autori principali: | Laurent, Jon M, Bou, Albert, Pieler, Michael, Igoe, Conor, Andonian, Alex, Narayanan, Siddharth, Braza, James, Vassopoulos, Alexandros Sanchez, Steenwyk, Jacob L, Lash, Blake, White, Andrew D, Rodriques, Samuel G |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology
di: Mitchener, Ludovico, et al.
Pubblicazione: (2025)
di: Mitchener, Ludovico, et al.
Pubblicazione: (2025)
Aviary: training language agents on challenging scientific tasks
di: Narayanan, Siddharth, et al.
Pubblicazione: (2024)
di: Narayanan, Siddharth, et al.
Pubblicazione: (2024)
Training a Scientific Reasoning Model for Chemistry
di: Narayanan, Siddharth M., et al.
Pubblicazione: (2025)
di: Narayanan, Siddharth M., et al.
Pubblicazione: (2025)
LAB-Bench: Measuring Capabilities of Language Models for Biology Research
di: Laurent, Jon M., et al.
Pubblicazione: (2024)
di: Laurent, Jon M., et al.
Pubblicazione: (2024)
Language agents achieve superhuman synthesis of scientific knowledge
di: Skarlinski, Michael D., et al.
Pubblicazione: (2024)
di: Skarlinski, Michael D., et al.
Pubblicazione: (2024)
The Nature of the Spectacle
di: Igoe, Jim
Pubblicazione: (2021)
di: Igoe, Jim
Pubblicazione: (2021)
New in the Old Northwest
di: Igoe, James
Pubblicazione: (1969)
di: Igoe, James
Pubblicazione: (1969)
Lebenssoziologie [sociología de la vida/vitalista]:Georg Simmel en la era de la información
di: Scott Lash
Pubblicazione: (2003)
di: Scott Lash
Pubblicazione: (2003)
Deepfake Caricatures: Amplifying attention to artifacts increases deepfake detection by humans and machines
di: Fosco, Camilo, et al.
Pubblicazione: (2022)
di: Fosco, Camilo, et al.
Pubblicazione: (2022)
Drones and Support for the Use of Force
di: Walsh, James Igoe, et al.
Pubblicazione: (2021)
di: Walsh, James Igoe, et al.
Pubblicazione: (2021)
Drones and Support for the Use of Force
di: Walsh, James Igoe, et al.
Pubblicazione: (2019)
di: Walsh, James Igoe, et al.
Pubblicazione: (2019)
Robin: A multi-agent system for automating scientific discovery
di: Ghareeb, Ali Essam, et al.
Pubblicazione: (2025)
di: Ghareeb, Ali Essam, et al.
Pubblicazione: (2025)
Tail-Sensitive KL and Rényi Convergence of Unadjusted Hamiltonian Monte Carlo via One-Shot Couplings
di: Bou-Rabee, Nawaf, et al.
Pubblicazione: (2026)
di: Bou-Rabee, Nawaf, et al.
Pubblicazione: (2026)
Single-Shot Ionization-Based Transverse Profile Monitor for Pulsed Electron Beams
di: Denham, Paul, et al.
Pubblicazione: (2024)
di: Denham, Paul, et al.
Pubblicazione: (2024)
The Importance of Thinking Outside the (Medical) Box: The Impact of Lifestyle on the Outcomes of Rheumatic and Musculoskeletal Conditions and the Promise of Lifestyle Medicine
di: Patricia Katz, et al.
Pubblicazione: (2025)
di: Patricia Katz, et al.
Pubblicazione: (2025)
Stress‐Normalized Sensitivity as a Comparative Benchmark for Intrinsically Piezoresistive Nanocomposite Materials in Wearable Electronics
di: Conor S. Boland
Pubblicazione: (2026)
di: Conor S. Boland
Pubblicazione: (2026)
Anomalous diffusion and factor ordering in (1+1)-dimensional Lorentzian quantum gravity
di: Sanderson, Elijah, et al.
Pubblicazione: (2024)
di: Sanderson, Elijah, et al.
Pubblicazione: (2024)
LAVIB: A Large-scale Video Interpolation Benchmark
di: Stergiou, Alexandros
Pubblicazione: (2024)
di: Stergiou, Alexandros
Pubblicazione: (2024)
Diversity-seeking Jump Games in Networks
di: Narayanan, Lata, et al.
Pubblicazione: (2023)
di: Narayanan, Lata, et al.
Pubblicazione: (2023)
The Cross-environment Hyperparameter Setting Benchmark for Reinforcement Learning
di: Patterson, Andrew, et al.
Pubblicazione: (2024)
di: Patterson, Andrew, et al.
Pubblicazione: (2024)
Alliance models in road infrastructure of the state of Vorarlberg – Feldkirch city tunnel
di: Bernhard Braza, et al.
Pubblicazione: (2024)
di: Bernhard Braza, et al.
Pubblicazione: (2024)
Improved algorithms for learning quantum Hamiltonians, via flat polynomials
di: Narayanan, Shyam
Pubblicazione: (2024)
di: Narayanan, Shyam
Pubblicazione: (2024)
Test-Time Training Scaling Laws for Chemical Exploration in Drug Design
di: Thomas, Morgan, et al.
Pubblicazione: (2025)
di: Thomas, Morgan, et al.
Pubblicazione: (2025)
OptimusKG: Unifying biomedical knowledge in a modern multimodal graph
di: Vittor, Lucas, et al.
Pubblicazione: (2026)
di: Vittor, Lucas, et al.
Pubblicazione: (2026)
Aggregation Artifacts in Subjective Tasks Collapse Large Language Models' Posteriors
di: Chochlakis, Georgios, et al.
Pubblicazione: (2024)
di: Chochlakis, Georgios, et al.
Pubblicazione: (2024)
The Strong Pull of Prior Knowledge in Large Language Models and Its Impact on Emotion Recognition
di: Chochlakis, Georgios, et al.
Pubblicazione: (2024)
di: Chochlakis, Georgios, et al.
Pubblicazione: (2024)
Hallucinations in AlphaFold3 for Intrinsically Disordered Proteins with disorder in Biological Process Residues
di: Gopalan, Shreya, et al.
Pubblicazione: (2025)
di: Gopalan, Shreya, et al.
Pubblicazione: (2025)
Biological Computation as Charge Guided Geometric Resolution
di: White, Kris
Pubblicazione: (2025)
di: White, Kris
Pubblicazione: (2025)
Classification of the anyon sectors of Kitaev's quantum double model
di: Bols, Alex, et al.
Pubblicazione: (2023)
di: Bols, Alex, et al.
Pubblicazione: (2023)
An Underrecognized Problem: Missed and Delayed Carbidopa‐Levodopa Administration in Emergency Department Patients With Parkinson's Disease
di: Natalie M. Elder, et al.
Pubblicazione: (2026)
di: Natalie M. Elder, et al.
Pubblicazione: (2026)
Variety-Seeking Jump Games on Graphs
di: Narayanan, Lata, et al.
Pubblicazione: (2025)
di: Narayanan, Lata, et al.
Pubblicazione: (2025)
ProVox: Personalization and Proactive Planning for Situated Human-Robot Collaboration
di: Grannen, Jennifer, et al.
Pubblicazione: (2025)
di: Grannen, Jennifer, et al.
Pubblicazione: (2025)
Temps i memòria: Camí de sirga i Les veus del Pamano
di: Enric Bou
Pubblicazione: (2013)
di: Enric Bou
Pubblicazione: (2013)
Cyclic Sieving Phenomenon for Independent sets of graphs
di: White, Jacob A
Pubblicazione: (2026)
di: White, Jacob A
Pubblicazione: (2026)
Why Active Empathy Is Essential to Early Childhood Learning in Ireland
di: Hogan, Conor, et al.
Pubblicazione: (2025)
di: Hogan, Conor, et al.
Pubblicazione: (2025)
Banyan: Improved Representation Learning with Explicit Structure
di: Opper, Mattia, et al.
Pubblicazione: (2024)
di: Opper, Mattia, et al.
Pubblicazione: (2024)
How do licensing boards provide oversight? An Idaho case study
di: Conor Norris, et al.
Pubblicazione: (2024)
di: Conor Norris, et al.
Pubblicazione: (2024)
On Machine Learning Approaches for Protein-Ligand Binding Affinity Prediction
di: Schapin, Nikolai, et al.
Pubblicazione: (2024)
di: Schapin, Nikolai, et al.
Pubblicazione: (2024)
An Investigation into the Larvicidal Activity of Biologically Synthesized Silver and Copper Oxide Nanoparticles Against Mosquito Larvae
di: Lakshmanan Narayanan, et al.
Pubblicazione: (2024)
di: Lakshmanan Narayanan, et al.
Pubblicazione: (2024)
Explicit Block Encoding of Difference-of-Gaussian Operators on a Periodic Grid
di: Mahmud, Jishnu, et al.
Pubblicazione: (2026)
di: Mahmud, Jishnu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology
di: Mitchener, Ludovico, et al.
Pubblicazione: (2025) -
Aviary: training language agents on challenging scientific tasks
di: Narayanan, Siddharth, et al.
Pubblicazione: (2024) -
Training a Scientific Reasoning Model for Chemistry
di: Narayanan, Siddharth M., et al.
Pubblicazione: (2025) -
LAB-Bench: Measuring Capabilities of Language Models for Biology Research
di: Laurent, Jon M., et al.
Pubblicazione: (2024) -
Language agents achieve superhuman synthesis of scientific knowledge
di: Skarlinski, Michael D., et al.
Pubblicazione: (2024)