RefuteBench 2.0 -- Agentic Benchmark for Dynamic Evaluation of LLM Responses to Refutation Instruction
Fuente:
arXiv
Salvato in:
| Autori principali: | Yan, Jianhao, Luo, Yun, Zhang, Yue |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RefuteBench: Evaluating Refuting Instruction-Following for Large Language Models
di: Yan, Jianhao, et al.
Pubblicazione: (2024)
di: Yan, Jianhao, et al.
Pubblicazione: (2024)
T-REX: Table -- Refute or Entail eXplainer
di: Horstmann, Tim Luka, et al.
Pubblicazione: (2025)
di: Horstmann, Tim Luka, et al.
Pubblicazione: (2025)
SuperDP: Differential Privacy Refutation via Supermartingales
di: Chatterjee, Krishnendu, et al.
Pubblicazione: (2026)
di: Chatterjee, Krishnendu, et al.
Pubblicazione: (2026)
Refuting Equivalence in Probabilistic Programs with Conditioning
di: Chatterjee, Krishnendu, et al.
Pubblicazione: (2025)
di: Chatterjee, Krishnendu, et al.
Pubblicazione: (2025)
Equivalence and Similarity Refutation for Probabilistic Programs
di: Chatterjee, Krishnendu, et al.
Pubblicazione: (2024)
di: Chatterjee, Krishnendu, et al.
Pubblicazione: (2024)
How We Refute Claims: Automatic Fact-Checking through Flaw Identification and Explanation
di: Kao, Wei-Yu, et al.
Pubblicazione: (2024)
di: Kao, Wei-Yu, et al.
Pubblicazione: (2024)
Support or Refute: Analyzing the Stance of Evidence to Detect Out-of-Context Mis- and Disinformation
di: Yuan, Xin, et al.
Pubblicazione: (2023)
di: Yuan, Xin, et al.
Pubblicazione: (2023)
Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track
di: Schaeffer, Rylan, et al.
Pubblicazione: (2025)
di: Schaeffer, Rylan, et al.
Pubblicazione: (2025)
Towards Generalizable and Faithful Logic Reasoning over Natural Language via Resolution Refutation
di: Sun, Zhouhao, et al.
Pubblicazione: (2024)
di: Sun, Zhouhao, et al.
Pubblicazione: (2024)
Refutability as Recursive as Provability
di: Cattabriga, Paola
Pubblicazione: (2024)
di: Cattabriga, Paola
Pubblicazione: (2024)
Beyond Correlation: Refutation-Validated Aspect-Based Sentiment Analysis for Explainable Energy Market Returns
di: van der Heever, Wihan, et al.
Pubblicazione: (2026)
di: van der Heever, Wihan, et al.
Pubblicazione: (2026)
Dynamics of Instruction Fine-Tuning for Chinese Large Language Models
di: Song, Chiyu, et al.
Pubblicazione: (2023)
di: Song, Chiyu, et al.
Pubblicazione: (2023)
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
di: Perlitz, Yotam, et al.
Pubblicazione: (2024)
di: Perlitz, Yotam, et al.
Pubblicazione: (2024)
IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation
di: Wen, Bosi, et al.
Pubblicazione: (2026)
di: Wen, Bosi, et al.
Pubblicazione: (2026)
Keys to Robust Edits: from Theoretical Insights to Practical Advances
di: Yan, Jianhao, et al.
Pubblicazione: (2024)
di: Yan, Jianhao, et al.
Pubblicazione: (2024)
ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning
di: Huang, Shulin, et al.
Pubblicazione: (2025)
di: Huang, Shulin, et al.
Pubblicazione: (2025)
ELICIT: LLM Augmentation via External In-Context Capability
di: Wang, Futing, et al.
Pubblicazione: (2024)
di: Wang, Futing, et al.
Pubblicazione: (2024)
Strongly Refuting Random CSP without Literals
di: Chan, Siu On, et al.
Pubblicazione: (2026)
di: Chan, Siu On, et al.
Pubblicazione: (2026)
OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding
di: Ding, Deming, et al.
Pubblicazione: (2026)
di: Ding, Deming, et al.
Pubblicazione: (2026)
Refuting Perfect Matchings in Spectral Expanders is Hard
di: Biswas, Ari, et al.
Pubblicazione: (2025)
di: Biswas, Ari, et al.
Pubblicazione: (2025)
A Formal Refutation of the Blockchain Trilemma
di: Wright, Craig
Pubblicazione: (2025)
di: Wright, Craig
Pubblicazione: (2025)
CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency
di: Guo, Jiacheng, et al.
Pubblicazione: (2025)
di: Guo, Jiacheng, et al.
Pubblicazione: (2025)
ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents
di: Jeon, YoungHoon, et al.
Pubblicazione: (2026)
di: Jeon, YoungHoon, et al.
Pubblicazione: (2026)
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels
di: Yan, Jianhao, et al.
Pubblicazione: (2024)
di: Yan, Jianhao, et al.
Pubblicazione: (2024)
Refuting approaches to the log-rank conjecture for XOR functions
di: Hatami, Hamed, et al.
Pubblicazione: (2023)
di: Hatami, Hamed, et al.
Pubblicazione: (2023)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
di: Deng, Shihan, et al.
Pubblicazione: (2024)
di: Deng, Shihan, et al.
Pubblicazione: (2024)
PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
di: Wang, Yidong, et al.
Pubblicazione: (2023)
di: Wang, Yidong, et al.
Pubblicazione: (2023)
FireBench: Evaluating Instruction Following in Enterprise and API-Driven LLM Applications
di: Zhang, Yunfan, et al.
Pubblicazione: (2026)
di: Zhang, Yunfan, et al.
Pubblicazione: (2026)
StructFlowBench: A Structured Flow Benchmark for Multi-turn Instruction Following
di: Li, Jinnan, et al.
Pubblicazione: (2025)
di: Li, Jinnan, et al.
Pubblicazione: (2025)
BenchBench: Benchmarking Automated Benchmark Generation
di: Zheng, Yandan, et al.
Pubblicazione: (2026)
di: Zheng, Yandan, et al.
Pubblicazione: (2026)
NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
di: Moore, Robert J., et al.
Pubblicazione: (2026)
di: Moore, Robert J., et al.
Pubblicazione: (2026)
RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems
di: Lin, Jingru, et al.
Pubblicazione: (2025)
di: Lin, Jingru, et al.
Pubblicazione: (2025)
FactBench: A Dynamic Benchmark for In-the-Wild Language Model Factuality Evaluation
di: Bayat, Farima Fatahi, et al.
Pubblicazione: (2024)
di: Bayat, Farima Fatahi, et al.
Pubblicazione: (2024)
X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic System
di: Wang, Peng, et al.
Pubblicazione: (2025)
di: Wang, Peng, et al.
Pubblicazione: (2025)
CIF-Bench: A Chinese Instruction-Following Benchmark for Evaluating the Generalizability of Large Language Models
di: LI, Yizhi, et al.
Pubblicazione: (2024)
di: LI, Yizhi, et al.
Pubblicazione: (2024)
The Refutability Gap: Challenges in Validating Reasoning by Large Language Models
di: Mossel, Elchanan
Pubblicazione: (2025)
di: Mossel, Elchanan
Pubblicazione: (2025)
Refutation calculi for lattice-based logics: from display to tableaux
di: De Domenico, Andrea, et al.
Pubblicazione: (2026)
di: De Domenico, Andrea, et al.
Pubblicazione: (2026)
Refuting the Metaphysics of Wolfram and Tegmark
di: Natal, Joseph
Pubblicazione: (2024)
di: Natal, Joseph
Pubblicazione: (2024)
Herbrand's Theorem in Refutation Schemata
di: Leitsch, Alexander, et al.
Pubblicazione: (2024)
di: Leitsch, Alexander, et al.
Pubblicazione: (2024)
VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency
di: Liu, Hongcheng, et al.
Pubblicazione: (2025)
di: Liu, Hongcheng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
RefuteBench: Evaluating Refuting Instruction-Following for Large Language Models
di: Yan, Jianhao, et al.
Pubblicazione: (2024) -
T-REX: Table -- Refute or Entail eXplainer
di: Horstmann, Tim Luka, et al.
Pubblicazione: (2025) -
SuperDP: Differential Privacy Refutation via Supermartingales
di: Chatterjee, Krishnendu, et al.
Pubblicazione: (2026) -
Refuting Equivalence in Probabilistic Programs with Conditioning
di: Chatterjee, Krishnendu, et al.
Pubblicazione: (2025) -
Equivalence and Similarity Refutation for Probabilistic Programs
di: Chatterjee, Krishnendu, et al.
Pubblicazione: (2024)