Strong Reasoning Isn't Enough: Evaluating Evidence Elicitation in Interactive Diagnosis
Fuente:
arXiv
Saved in:
| Main Authors: | Long, Zhuohan, Bao, Zhijie, Wei, Zhongyu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
by: Barkett, Emilio, et al.
Published: (2025)
by: Barkett, Emilio, et al.
Published: (2025)
Recall Isn't Enough: Bounding Commitments in Personalized Language Systems
by: Tang, Rui, et al.
Published: (2026)
by: Tang, Rui, et al.
Published: (2026)
SAIE Framework: Support Alone Isn't Enough -- Advancing LLM Training with Adversarial Remarks
by: Loem, Mengsay, et al.
Published: (2023)
by: Loem, Mengsay, et al.
Published: (2023)
When Fairness Isn't Statistical: The Limits of Machine Learning in Evaluating Legal Reasoning
by: Barale, Claire, et al.
Published: (2025)
by: Barale, Claire, et al.
Published: (2025)
Inverse Scaling: When Bigger Isn't Better
by: McKenzie, Ian R., et al.
Published: (2023)
by: McKenzie, Ian R., et al.
Published: (2023)
Knowing the Answer Isn't Enough: Fixing Reasoning Path Failures in LVLMs
by: Wang, Chaoyang, et al.
Published: (2025)
by: Wang, Chaoyang, et al.
Published: (2025)
Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation
by: Wang, Siyuan, et al.
Published: (2024)
by: Wang, Siyuan, et al.
Published: (2024)
Word Boundary Information Isn't Useful for Encoder Language Models
by: Gow-Smith, Edward, et al.
Published: (2024)
by: Gow-Smith, Edward, et al.
Published: (2024)
Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2026)
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2026)
When Trust Isn't Enough.
by: Behrman, Sara
Published: (1998)
by: Behrman, Sara
Published: (1998)
From LLMs to MLLMs: Exploring the Landscape of Multimodal Jailbreaking
by: Wang, Siyuan, et al.
Published: (2024)
by: Wang, Siyuan, et al.
Published: (2024)
When Bigger Isn't Better: A Comprehensive Fairness Evaluation of Political Bias in Multi-News Summarisation
by: Huang, Nannan, et al.
Published: (2026)
by: Huang, Nannan, et al.
Published: (2026)
When Meaning Isn't Literal: Exploring Idiomatic Meaning Across Languages and Modalities
by: Das, Sarmistha, et al.
Published: (2026)
by: Das, Sarmistha, et al.
Published: (2026)
Being Kind Isn't Always Being Safe: Diagnosing Affective Hallucination in LLMs
by: Kim, Sewon, et al.
Published: (2025)
by: Kim, Sewon, et al.
Published: (2025)
When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions
by: Cao, Zhuo, et al.
Published: (2025)
by: Cao, Zhuo, et al.
Published: (2025)
Mathematics Isn't Culture-Free: Probing Cultural Gaps via Entity and Scenario Perturbations
by: Tomar, Aditya, et al.
Published: (2025)
by: Tomar, Aditya, et al.
Published: (2025)
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
by: Xu, Xiaoyu, et al.
Published: (2025)
by: Xu, Xiaoyu, et al.
Published: (2025)
Talk Isn't Always Cheap: Understanding Failure Modes in Multi-Agent Debate
by: Wynn, Andrea, et al.
Published: (2025)
by: Wynn, Andrea, et al.
Published: (2025)
Surprise! Uniform Information Density Isn't the Whole Story: Predicting Surprisal Contours in Long-form Discourse
by: Tsipidi, Eleftheria, et al.
Published: (2024)
by: Tsipidi, Eleftheria, et al.
Published: (2024)
MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging
by: Bao, Zhijie, et al.
Published: (2026)
by: Bao, Zhijie, et al.
Published: (2026)
Seeing Isn't Believing: Mitigating Belief Inertia via Active Intervention in Embodied Agents
by: Wang, Hanlin, et al.
Published: (2026)
by: Wang, Hanlin, et al.
Published: (2026)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
by: Zhang, Yue, et al.
Published: (2026)
by: Zhang, Yue, et al.
Published: (2026)
Why Synthetic Isn't Real Yet: A Diagnostic Framework for Contact Center Dialogue Generation
by: Devanathan, Rishikesh, et al.
Published: (2025)
by: Devanathan, Rishikesh, et al.
Published: (2025)
When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models
by: Galeone, Cosimo, et al.
Published: (2026)
by: Galeone, Cosimo, et al.
Published: (2026)
Explainable AI Isn't Enough! Rethinking Algorithmic Contestability
by: Freiesleben, Timo, et al.
Published: (2026)
by: Freiesleben, Timo, et al.
Published: (2026)
Access Isn't Enough: Merely Connecting People and Computers Won't Close the Digital Divide.
by: Blau, Andrew
Published: (2002)
by: Blau, Andrew
Published: (2002)
What You Read Isn't What You Hear: Linguistic Sensitivity in Deepfake Speech Detection
by: Nguyen, Binh, et al.
Published: (2025)
by: Nguyen, Binh, et al.
Published: (2025)
From Metacognition to Computable Wisdom: Why "Thinking About Thinking" Isn't Enough for Agentic AI
by: Figurelli, Rogério
Published: (2026)
by: Figurelli, Rogério
Published: (2026)
Relevance Isn't All You Need: Scaling RAG Systems With Inference-Time Compute Via Multi-Criteria Reranking
by: LeVine, Will, et al.
Published: (2025)
by: LeVine, Will, et al.
Published: (2025)
ML Interpretability: Simple Isn't Easy
by: Räz, Tim
Published: (2022)
by: Räz, Tim
Published: (2022)
Weak-to-Strong Elicitation via Mismatched Wrong Drafts
by: Deng, Wei
Published: (2026)
by: Deng, Wei
Published: (2026)
When Alignment Isn't Enough: Response-Path Attacks on LLM Agents
by: Luo, Mingyu, et al.
Published: (2026)
by: Luo, Mingyu, et al.
Published: (2026)
Stop When Enough: Adaptive Early-Stopping for Chain-of-Thought Reasoning
by: Sun, Renliang, et al.
Published: (2025)
by: Sun, Renliang, et al.
Published: (2025)
Stepwise Informativeness Search for Efficient and Effective LLM Reasoning
by: Wang, Siyuan, et al.
Published: (2025)
by: Wang, Siyuan, et al.
Published: (2025)
Weak-eval-Strong: Evaluating and Eliciting Lateral Thinking of LLMs with Situation Puzzles
by: Chen, Qi, et al.
Published: (2024)
by: Chen, Qi, et al.
Published: (2024)
LifeSim: Long-Horizon User Life Simulator for Personalized Assistant Evaluation
by: Duan, Feiyu, et al.
Published: (2026)
by: Duan, Feiyu, et al.
Published: (2026)
When Standard Newborn Screening Isn't Enough: Diagnostic Challenges in the Age of Globalization
by: Marina Ortúzar Menéndez, et al.
Published: (2026)
by: Marina Ortúzar Menéndez, et al.
Published: (2026)
Awake ECMO for Mid‐Tracheal Obstruction: When a Tracheostomy Isn't Enough
by: Jacob Beiriger, et al.
Published: (2026)
by: Jacob Beiriger, et al.
Published: (2026)
Can LLMs Reason with Rules? Logic Scaffolding for Stress-Testing and Improving LLMs
by: Wang, Siyuan, et al.
Published: (2024)
by: Wang, Siyuan, et al.
Published: (2024)
If It Isn't Broken...Break It!
by: Voges, Mickie A.
Published: (2000)
by: Voges, Mickie A.
Published: (2000)
Similar Items
-
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
by: Barkett, Emilio, et al.
Published: (2025) -
Recall Isn't Enough: Bounding Commitments in Personalized Language Systems
by: Tang, Rui, et al.
Published: (2026) -
SAIE Framework: Support Alone Isn't Enough -- Advancing LLM Training with Adversarial Remarks
by: Loem, Mengsay, et al.
Published: (2023) -
When Fairness Isn't Statistical: The Limits of Machine Learning in Evaluating Legal Reasoning
by: Barale, Claire, et al.
Published: (2025) -
Inverse Scaling: When Bigger Isn't Better
by: McKenzie, Ian R., et al.
Published: (2023)