Evaluating LLMs at Detecting Errors in LLM Responses
Fuente:
arXiv
Saved in:
| Main Authors: | Kamoi, Ryo, Das, Sarkar Snigdha Sarathi, Lou, Renze, Ahn, Jihyun Janice, Zhao, Yilun, Lu, Xiaoxin, Zhang, Nan, Zhang, Yusen, Zhang, Ranran Haoran, Vummanthala, Sujeeth Reddy, Dave, Salika, Qin, Shaobo, Cohan, Arman, Yin, Wenpeng, Zhang, Rui |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information
by: Kamoi, Ryo, et al.
Published: (2024)
by: Kamoi, Ryo, et al.
Published: (2024)
Efficient PRM Training Data Synthesis via Formal Verification
by: Kamoi, Ryo, et al.
Published: (2025)
by: Kamoi, Ryo, et al.
Published: (2025)
Direct-Inverse Prompting: Analyzing LLMs' Discriminative Capacity in Self-Improving Generation
by: Ahn, Jihyun Janice, et al.
Published: (2024)
by: Ahn, Jihyun Janice, et al.
Published: (2024)
GReaTer: Gradients over Reasoning Makes Smaller Language Models Strong Prompt Optimizers
by: Das, Sarkar Snigdha Sarathi, et al.
Published: (2024)
by: Das, Sarkar Snigdha Sarathi, et al.
Published: (2024)
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding?
by: Zhang, Yusen, et al.
Published: (2025)
by: Zhang, Yusen, et al.
Published: (2025)
Verbosity $\neq$ Veracity: Demystify Verbosity Compensation Behavior of Large Language Models
by: Zhang, Yusen, et al.
Published: (2024)
by: Zhang, Yusen, et al.
Published: (2024)
Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation
by: Lu, Xiaoxin, et al.
Published: (2025)
by: Lu, Xiaoxin, et al.
Published: (2025)
GREATERPROMPT: A Unified, Customizable, and High-Performing Open-Source Toolkit for Prompt Optimization
by: Zheng, Wenliang, et al.
Published: (2025)
by: Zheng, Wenliang, et al.
Published: (2025)
AAAR-1.0: Assessing AI's Potential to Assist Research
by: Lou, Renze, et al.
Published: (2024)
by: Lou, Renze, et al.
Published: (2024)
Prompt-Reverse Inconsistency: LLM Self-Inconsistency Beyond Generative Randomness and Prompt Paraphrasing
by: Ahn, Jihyun Janice, et al.
Published: (2025)
by: Ahn, Jihyun Janice, et al.
Published: (2025)
Large Language Models for Mathematical Reasoning: Progresses and Challenges
by: Ahn, Janice, et al.
Published: (2024)
by: Ahn, Janice, et al.
Published: (2024)
Large Language Model Instruction Following: A Survey of Progresses and Challenges
by: Lou, Renze, et al.
Published: (2023)
by: Lou, Renze, et al.
Published: (2023)
Toward Zero-Shot Instruction Following
by: Lou, Renze, et al.
Published: (2023)
by: Lou, Renze, et al.
Published: (2023)
When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs
by: Kamoi, Ryo, et al.
Published: (2024)
by: Kamoi, Ryo, et al.
Published: (2024)
MUFFIN: Curating Multi-Faceted Instructions for Improving Instruction-Following
by: Lou, Renze, et al.
Published: (2023)
by: Lou, Renze, et al.
Published: (2023)
DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Long and Specialized Documents
by: Zhao, Yilun, et al.
Published: (2023)
by: Zhao, Yilun, et al.
Published: (2023)
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
by: Yu, Zhaojian, et al.
Published: (2024)
by: Yu, Zhaojian, et al.
Published: (2024)
LimRank: Less is More for Reasoning-Intensive Information Reranking
by: Song, Tingyu, et al.
Published: (2025)
by: Song, Tingyu, et al.
Published: (2025)
SAGE: Benchmarking and Improving Retrieval for Deep Research Agents
by: Hu, Tiansheng, et al.
Published: (2026)
by: Hu, Tiansheng, et al.
Published: (2026)
Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Models
by: Li, Xi, et al.
Published: (2024)
by: Li, Xi, et al.
Published: (2024)
Coverage-based Fairness in Multi-document Summarization
by: Li, Haoyuan, et al.
Published: (2024)
by: Li, Haoyuan, et al.
Published: (2024)
Can Prompt Modifiers Control Bias? A Comparative Analysis of Text-to-Image Generative Models
by: Shin, Philip Wootaek, et al.
Published: (2024)
by: Shin, Philip Wootaek, et al.
Published: (2024)
Fair Abstractive Summarization of Diverse Perspectives
by: Zhang, Yusen, et al.
Published: (2023)
by: Zhang, Yusen, et al.
Published: (2023)
Z1: Efficient Test-time Scaling with Code
by: Yu, Zhaojian, et al.
Published: (2025)
by: Yu, Zhaojian, et al.
Published: (2025)
SciDQA: A Deep Reading Comprehension Dataset over Scientific Papers
by: Singh, Shruti, et al.
Published: (2024)
by: Singh, Shruti, et al.
Published: (2024)
ANCHOR: Branch-Point Data Generation for GUI Agents
by: Wei, Jinbiao, et al.
Published: (2026)
by: Wei, Jinbiao, et al.
Published: (2026)
SUCEA: Reasoning-Intensive Retrieval for Adversarial Fact-checking through Claim Decomposition and Editing
by: Liu, Hongjun, et al.
Published: (2025)
by: Liu, Hongjun, et al.
Published: (2025)
MCTS-RAG: Enhancing Retrieval-Augmented Generation with Monte Carlo Tree Search
by: Hu, Yunhai, et al.
Published: (2025)
by: Hu, Yunhai, et al.
Published: (2025)
Table-R1: Inference-Time Scaling for Table Reasoning
by: Yang, Zheyuan, et al.
Published: (2025)
by: Yang, Zheyuan, et al.
Published: (2025)
Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers
by: Zhao, Yilun, et al.
Published: (2025)
by: Zhao, Yilun, et al.
Published: (2025)
FinanceMath: Knowledge-Intensive Math Reasoning in Finance Domains
by: Zhao, Yilun, et al.
Published: (2023)
by: Zhao, Yilun, et al.
Published: (2023)
Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems
by: Zhao, Yilun, et al.
Published: (2026)
by: Zhao, Yilun, et al.
Published: (2026)
Cycle V — The Reconciliation of Sky and Stone
by: Compa, Salika Bista
Published: (2025)
by: Compa, Salika Bista
Published: (2025)
Bridging the Know-Act Gap via Task-Level Autoregressive Reasoning
by: Ahn, Jihyun Janice, et al.
Published: (2026)
by: Ahn, Jihyun Janice, et al.
Published: (2026)
AlphaResearch: Accelerating New Algorithm Discovery with Language Models
by: Yu, Zhaojian, et al.
Published: (2025)
by: Yu, Zhaojian, et al.
Published: (2025)
Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective
by: Zhang, Siyue, et al.
Published: (2025)
by: Zhang, Siyue, et al.
Published: (2025)
UMIE: Unified Multimodal Information Extraction with Instruction Tuning
by: Sun, Lin, et al.
Published: (2024)
by: Sun, Lin, et al.
Published: (2024)
MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning
by: Tang, Xiangru, et al.
Published: (2023)
by: Tang, Xiangru, et al.
Published: (2023)
MSRS: Evaluating Multi-Source Retrieval-Augmented Generation
by: Phanse, Rohan, et al.
Published: (2025)
by: Phanse, Rohan, et al.
Published: (2025)
Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet
by: Atil, Berk, et al.
Published: (2025)
by: Atil, Berk, et al.
Published: (2025)
Similar Items
-
VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information
by: Kamoi, Ryo, et al.
Published: (2024) -
Efficient PRM Training Data Synthesis via Formal Verification
by: Kamoi, Ryo, et al.
Published: (2025) -
Direct-Inverse Prompting: Analyzing LLMs' Discriminative Capacity in Self-Improving Generation
by: Ahn, Jihyun Janice, et al.
Published: (2024) -
GReaTer: Gradients over Reasoning Makes Smaller Language Models Strong Prompt Optimizers
by: Das, Sarkar Snigdha Sarathi, et al.
Published: (2024) -
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding?
by: Zhang, Yusen, et al.
Published: (2025)