Overconfident Oracles: Limitations of In Silico Sequence Design Benchmarking
Fuente:
arXiv
Saved in:
| Main Authors: | Surana, Shikha, Grinsztajn, Nathan, Atkinson, Timothy, Duckworth, Paul, Barrett, Thomas D. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Combinatorial Optimization with Policy Adaptation using Latent Space Search
by: Chalumeau, Felix, et al.
Published: (2023)
by: Chalumeau, Felix, et al.
Published: (2023)
Metalic: Meta-Learning In-Context with Protein Language Models
by: Beck, Jacob, et al.
Published: (2024)
by: Beck, Jacob, et al.
Published: (2024)
Force-Aware Neural Tangent Kernels for Scalable and Robust Active Learning of MLIPs
by: Varga-Umbrich, Eszter, et al.
Published: (2026)
by: Varga-Umbrich, Eszter, et al.
Published: (2026)
Pretrained Model Representations as Acquisition Signals for Active Learning of MLIPs
by: Varga-Umbrich, Eszter, et al.
Published: (2026)
by: Varga-Umbrich, Eszter, et al.
Published: (2026)
Memory-Enhanced Neural Solvers for Routing Problems
by: Chalumeau, Felix, et al.
Published: (2024)
by: Chalumeau, Felix, et al.
Published: (2024)
Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs
by: Smit, Andries, et al.
Published: (2023)
by: Smit, Andries, et al.
Published: (2023)
Learning the Language of Protein Structure
by: Gaujac, Benoit, et al.
Published: (2024)
by: Gaujac, Benoit, et al.
Published: (2024)
Sample-Efficient Optimisation over the Outputs of Generative Models
by: Willis, Samuel, et al.
Published: (2025)
by: Willis, Samuel, et al.
Published: (2025)
Fair and Calibrated Toxicity Detection with Robust Training and Abstention
by: Surana, Mokshit
Published: (2026)
by: Surana, Mokshit
Published: (2026)
CS-Sum: A Benchmark for Code-Switching Dialogue Summarization and the Limits of Large Language Models
by: Suresh, Sathya Krishnan, et al.
Published: (2025)
by: Suresh, Sathya Krishnan, et al.
Published: (2025)
Better by Default: Strong Pre-Tuned MLPs and Boosted Trees on Tabular Data
by: Holzmüller, David, et al.
Published: (2024)
by: Holzmüller, David, et al.
Published: (2024)
Multi-Objective Quality-Diversity for Crystal Structure Prediction
by: Janmohamed, Hannah, et al.
Published: (2024)
by: Janmohamed, Hannah, et al.
Published: (2024)
CARTE: Pretraining and Transfer for Tabular Learning
by: Kim, Myung Jun, et al.
Published: (2024)
by: Kim, Myung Jun, et al.
Published: (2024)
Agentic Uncertainty Reveals Agentic Overconfidence
by: Kaddour, Jean, et al.
Published: (2026)
by: Kaddour, Jean, et al.
Published: (2026)
Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX
by: Bonnet, Clément, et al.
Published: (2023)
by: Bonnet, Clément, et al.
Published: (2023)
Beyond Overconfidence: Foundation Models Redefine Calibration in Deep Neural Networks
by: Hekler, Achim, et al.
Published: (2025)
by: Hekler, Achim, et al.
Published: (2025)
The Cost of Reasoning: Chain-of-Thought Induces Overconfidence in Vision-Language Models
by: Welch, Robert, et al.
Published: (2026)
by: Welch, Robert, et al.
Published: (2026)
Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes
by: Bereket, Michael, et al.
Published: (2025)
by: Bereket, Michael, et al.
Published: (2025)
AMIGO: Agentic Multi-Image Grounding Oracle Benchmark
by: Wang, Min, et al.
Published: (2026)
by: Wang, Min, et al.
Published: (2026)
SmartOracle -- An Agentic Approach to Mitigate Noise in Differential Oracles
by: Srinivasan, Srinath, et al.
Published: (2026)
by: Srinivasan, Srinath, et al.
Published: (2026)
GROOT: Effective Design of Biological Sequences with Limited Experimental Data
by: Tran, Thanh V. T., et al.
Published: (2024)
by: Tran, Thanh V. T., et al.
Published: (2024)
Mitigating Overconfidence in Out-of-Distribution Detection by Capturing Extreme Activations
by: Azizmalayeri, Mohammad, et al.
Published: (2024)
by: Azizmalayeri, Mohammad, et al.
Published: (2024)
Is Oracle Pruning the True Oracle?
by: Feng, Sicheng, et al.
Published: (2024)
by: Feng, Sicheng, et al.
Published: (2024)
Slimmable NAM: Neural Amp Models with adjustable runtime computational cost
by: Atkinson, Steven
Published: (2025)
by: Atkinson, Steven
Published: (2025)
Optimal Transport-Induced Samples against Out-of-Distribution Overconfidence
by: Tang, Keke, et al.
Published: (2026)
by: Tang, Keke, et al.
Published: (2026)
Overconfident Errors Need Stronger Correction: Asymmetric Confidence Penalties for Reinforcement Learning
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
Measuring and Mitigating Toxicity in Large Language Models: A Comprehensive Replication Study
by: Surana, Mokshit, et al.
Published: (2026)
by: Surana, Mokshit, et al.
Published: (2026)
Data Curation Through the Lens of Spectral Dynamics: Static Limits, Dynamic Acceleration, and Practical Oracles
by: Zhang, Yizhou, et al.
Published: (2025)
by: Zhang, Yizhou, et al.
Published: (2025)
Overconfidence and Calibration in Medical VQA: Empirical Findings and Hallucination-Aware Mitigation
by: Byun, Ji Young, et al.
Published: (2026)
by: Byun, Ji Young, et al.
Published: (2026)
Position: Benchmarking is Limited in Reinforcement Learning Research
by: Jordan, Scott M., et al.
Published: (2024)
by: Jordan, Scott M., et al.
Published: (2024)
Humble your Overconfident Networks: Unlearning Overfitting via Sequential Monte Carlo Tempered Deep Ensembles
by: Millard, Andrew, et al.
Published: (2025)
by: Millard, Andrew, et al.
Published: (2025)
Oracle-Efficient Combinatorial Semi-Bandits
by: Kim, Jung-hun, et al.
Published: (2025)
by: Kim, Jung-hun, et al.
Published: (2025)
SPO: Sequential Monte Carlo Policy Optimisation
by: Macfarlane, Matthew V, et al.
Published: (2024)
by: Macfarlane, Matthew V, et al.
Published: (2024)
DITTO: Offline Imitation Learning with World Models
by: DeMoss, Branton, et al.
Published: (2023)
by: DeMoss, Branton, et al.
Published: (2023)
LLMs are Overconfident: Evaluating Confidence Interval Calibration with FermiEval
by: Epstein, Elliot L., et al.
Published: (2025)
by: Epstein, Elliot L., et al.
Published: (2025)
Known Meets Unknown: Mitigating Overconfidence in Open Set Recognition
by: Zhao, Dongdong, et al.
Published: (2025)
by: Zhao, Dongdong, et al.
Published: (2025)
ConfHit: Conformal Generative Design with Oracle Free Guarantees
by: Laghuvarapu, Siddhartha, et al.
Published: (2026)
by: Laghuvarapu, Siddhartha, et al.
Published: (2026)
Biologically-Grounded Multi-Encoder Architectures as Developability Oracles for Antibody Design
by: Crouzet, Simon J.
Published: (2026)
by: Crouzet, Simon J.
Published: (2026)
ShiQ: Bringing back Bellman to LLMs
by: Clavier, Pierre, et al.
Published: (2025)
by: Clavier, Pierre, et al.
Published: (2025)
Limitations of Sequence-Based Protein Representations for Parkinson's Disease Classification: A Leakage-Free Benchmark
by: Núñez-Prado, César Jesús, et al.
Published: (2026)
by: Núñez-Prado, César Jesús, et al.
Published: (2026)
Similar Items
-
Combinatorial Optimization with Policy Adaptation using Latent Space Search
by: Chalumeau, Felix, et al.
Published: (2023) -
Metalic: Meta-Learning In-Context with Protein Language Models
by: Beck, Jacob, et al.
Published: (2024) -
Force-Aware Neural Tangent Kernels for Scalable and Robust Active Learning of MLIPs
by: Varga-Umbrich, Eszter, et al.
Published: (2026) -
Pretrained Model Representations as Acquisition Signals for Active Learning of MLIPs
by: Varga-Umbrich, Eszter, et al.
Published: (2026) -
Memory-Enhanced Neural Solvers for Routing Problems
by: Chalumeau, Felix, et al.
Published: (2024)