How Reliable are Confidence Estimators for Large Reasoning Models? A Systematic Benchmark on High-Stakes Domains
Fuente:
arXiv
Guardado en:
| Autores principales: | Khanmohammadi, Reza, Miahi, Erfan, Kaur, Simerjot, Brugere, Ivan, Smiley, Charese H., Thind, Kundan, Ghassemi, Mohammad M. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Grounded or Guessing? LVLM Confidence Estimation via Blind-Image Contrastive Ranking
por: Khanmohammadi, Reza, et al.
Publicado: (2026)
por: Khanmohammadi, Reza, et al.
Publicado: (2026)
Calibrating LLM Confidence by Probing Perturbed Representation Stability
por: Khanmohammadi, Reza, et al.
Publicado: (2025)
por: Khanmohammadi, Reza, et al.
Publicado: (2025)
The Influence of Biomedical Research on Future Business Funding: Analyzing Scientific Impact and Content in Industrial Investments
por: Khanmohammadi, Reza, et al.
Publicado: (2024)
por: Khanmohammadi, Reza, et al.
Publicado: (2024)
FinQAPT: Empowering Financial Decisions with End-to-End LLM-driven Question Answering Pipeline
por: Singh, Kuldeep, et al.
Publicado: (2024)
por: Singh, Kuldeep, et al.
Publicado: (2024)
Grounding LLM Reasoning with Knowledge Graphs
por: Amayuelas, Alfonso, et al.
Publicado: (2025)
por: Amayuelas, Alfonso, et al.
Publicado: (2025)
Conservative Bias in Large Language Models: Measuring Relation Predictions
por: Aguda, Toyin, et al.
Publicado: (2025)
por: Aguda, Toyin, et al.
Publicado: (2025)
FinNLI: Novel Dataset for Multi-Genre Financial Natural Language Inference Benchmarking
por: Magomere, Jabez, et al.
Publicado: (2025)
por: Magomere, Jabez, et al.
Publicado: (2025)
Large Language Models as Financial Data Annotators: A Study on Effectiveness and Efficiency
por: Aguda, Toyin, et al.
Publicado: (2024)
por: Aguda, Toyin, et al.
Publicado: (2024)
A Variational Approach for Mitigating Entity Bias in Relation Extraction
por: Mensah, Samuel, et al.
Publicado: (2025)
por: Mensah, Samuel, et al.
Publicado: (2025)
Iterative Prompt Refinement for Radiation Oncology Symptom Extraction Using Teacher-Student Large Language Models
por: Khanmohammadi, Reza, et al.
Publicado: (2024)
por: Khanmohammadi, Reza, et al.
Publicado: (2024)
Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
por: Miahi, Erfan, et al.
Publicado: (2026)
por: Miahi, Erfan, et al.
Publicado: (2026)
Hybrid Student-Teacher Large Language Model Refinement for Cancer Toxicity Symptom Extraction
por: Khanmohammadi, Reza, et al.
Publicado: (2024)
por: Khanmohammadi, Reza, et al.
Publicado: (2024)
Distill and Align Decomposition for Enhanced Claim Verification
por: Magomere, Jabez, et al.
Publicado: (2026)
por: Magomere, Jabez, et al.
Publicado: (2026)
Deep FinResearch Bench: Evaluating AI's Ability to Conduct Professional Financial Investment Research
por: Haque, Mirazul, et al.
Publicado: (2026)
por: Haque, Mirazul, et al.
Publicado: (2026)
AI Analyst: Framework and Comprehensive Evaluation of Large Language Models for Financial Time Series Report Generation
por: Fons, Elizabeth, et al.
Publicado: (2025)
por: Fons, Elizabeth, et al.
Publicado: (2025)
Reduction in preparatory brain activity preceding gait initiation in individuals with chronic ankle instability: A movement‐related cortical potential study
por: Zivar Beyraghi, et al.
Publicado: (2024)
por: Zivar Beyraghi, et al.
Publicado: (2024)
GenPlanX. Generation of Plans and Execution
por: Borrajo, Daniel, et al.
Publicado: (2025)
por: Borrajo, Daniel, et al.
Publicado: (2025)
New Statistical Framework for Extreme Error Probability in High-Stakes Domains for Reliable Machine Learning
por: Michelucci, Umberto, et al.
Publicado: (2025)
por: Michelucci, Umberto, et al.
Publicado: (2025)
Domaino1s: Guiding LLM Reasoning for Explainable Answers in High-Stakes Domains
por: Chu, Xu, et al.
Publicado: (2025)
por: Chu, Xu, et al.
Publicado: (2025)
Luminescent Multifunctional Nanomaterials: Capacitive Removal and Enhanced Detection Efficiency of Heavy Metals Ions for Advanced Water and Wastewater Treatment Application
por: Karim Khanmohammadi Chenab, et al.
Publicado: (2024)
por: Karim Khanmohammadi Chenab, et al.
Publicado: (2024)
PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning
por: Akyürek, Afra Feyza, et al.
Publicado: (2025)
por: Akyürek, Afra Feyza, et al.
Publicado: (2025)
Design of Reliable and Resilient Electric Power Systems for Wide-Body All-Electric Aircraft
por: Ghassemi, Mona
Publicado: (2025)
por: Ghassemi, Mona
Publicado: (2025)
Mainstreaming gender in the BOBLME Project
por: Brugere, Cecile
Publicado: (2012)
por: Brugere, Cecile
Publicado: (2012)
La desaparición de la obra
por: Fabienne Brugère
Publicado: (2007)
por: Fabienne Brugère
Publicado: (2007)
Towards Reliable Medical LLMs: Benchmarking and Enhancing Confidence Estimation of Large Language Models in Medical Consultation
por: Ren, Zhiyao, et al.
Publicado: (2026)
por: Ren, Zhiyao, et al.
Publicado: (2026)
A Comprehensive Forecasting-Based Framework for Time Series Anomaly Detection: Benchmarking on the Numenta Anomaly Benchmark (NAB)
por: Karami, Mohammad, et al.
Publicado: (2025)
por: Karami, Mohammad, et al.
Publicado: (2025)
(How) Do Large Language Models Understand High-Level Message Sequence Charts?
por: Mousavi, Mohammad Reza
Publicado: (2026)
por: Mousavi, Mohammad Reza
Publicado: (2026)
Impact of Social Comparison on Fake News Release and News Credibility
por: Esfidani, Mohammad Rahim, et al.
Publicado: (2023)
por: Esfidani, Mohammad Rahim, et al.
Publicado: (2023)
SINDyG: Sparse Identification of Nonlinear Dynamical Systems from Graph-Structured Data, with Applications to Stuart-Landau Oscillator Networks
por: Basiri, Mohammad Amin, et al.
Publicado: (2024)
por: Basiri, Mohammad Amin, et al.
Publicado: (2024)
The Shift From Neo‐Ottomanism to the Century of Türkiye: Domestic, Regional, and Global Implications
por: Mohammad Hadi Khanmohammadi, et al.
Publicado: (2026)
por: Mohammad Hadi Khanmohammadi, et al.
Publicado: (2026)
Robustness of Transformer-Based Fluence Map Prediction Under Clinically Realistic Perturbations
por: Mgboh, Ujunwa, et al.
Publicado: (2026)
por: Mgboh, Ujunwa, et al.
Publicado: (2026)
FluenceFormer: Transformer-Driven Multi-Beam Fluence Map Regression for Radiotherapy Planning
por: Mgboh, Ujunwa, et al.
Publicado: (2025)
por: Mgboh, Ujunwa, et al.
Publicado: (2025)
FOS: A Large-Scale Temporal Graph Benchmark for Scientific Interdisciplinary Link Prediction
por: Rezaee, Kiyan, et al.
Publicado: (2025)
por: Rezaee, Kiyan, et al.
Publicado: (2025)
Intertextual Parallel Detection in Biblical Hebrew: A Transformer-Based Benchmark
por: Smiley, David M.
Publicado: (2025)
por: Smiley, David M.
Publicado: (2025)
MicroRNA as a Potential Diagnostic and Prognostic Biomarker in Diffuse Large B‐Cell Lymphoma: A Systematic Review and Meta‐Analysis
por: Shaghayegh Khanmohammadi, et al.
Publicado: (2025)
por: Shaghayegh Khanmohammadi, et al.
Publicado: (2025)
Revisiting Reliability in the Reasoning-based Pose Estimation Benchmark
por: Kim, Junsu, et al.
Publicado: (2025)
por: Kim, Junsu, et al.
Publicado: (2025)
Utilizing patient data: A tutorial on predicting second cancer with machine learning models
por: Hossein Sadeghi, et al.
Publicado: (2024)
por: Hossein Sadeghi, et al.
Publicado: (2024)
Energy-Efficient Approximate Full Adders Applying Memristive Serial IMPLY Logic For Image Processing
por: Fatemieh, Seyed Erfan, et al.
Publicado: (2024)
por: Fatemieh, Seyed Erfan, et al.
Publicado: (2024)
Energy‐Efficient Approximate Full Adders Applying Memristive Serial IMPLY Logic for Image Processing
por: Seyed Erfan Fatemieh, et al.
Publicado: (2026)
por: Seyed Erfan Fatemieh, et al.
Publicado: (2026)
Association of Body Roundness Index, Lipid Accumulation Product, and Triglyceride‐Glucose Index With Psoriasis: A Systematic Review of Observational Studies
por: Maryam Yousefiasl, et al.
Publicado: (2026)
por: Maryam Yousefiasl, et al.
Publicado: (2026)
Ejemplares similares
-
Grounded or Guessing? LVLM Confidence Estimation via Blind-Image Contrastive Ranking
por: Khanmohammadi, Reza, et al.
Publicado: (2026) -
Calibrating LLM Confidence by Probing Perturbed Representation Stability
por: Khanmohammadi, Reza, et al.
Publicado: (2025) -
The Influence of Biomedical Research on Future Business Funding: Analyzing Scientific Impact and Content in Industrial Investments
por: Khanmohammadi, Reza, et al.
Publicado: (2024) -
FinQAPT: Empowering Financial Decisions with End-to-End LLM-driven Question Answering Pipeline
por: Singh, Kuldeep, et al.
Publicado: (2024) -
Grounding LLM Reasoning with Knowledge Graphs
por: Amayuelas, Alfonso, et al.
Publicado: (2025)