ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yindong, Preiß, Martin, Bugueño, Margarita, Hoffbauer, Jan Vincent, Ghajar, Abdullatif, Buz, Tolga, de Melo, Gerard |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
by: Arad, Dana, et al.
Published: (2023)
by: Arad, Dana, et al.
Published: (2023)
Connecting the Dots: What Graph-Based Text Representations Work Best for Text Classification Using Graph Neural Networks?
by: Bugueño, Margarita, et al.
Published: (2023)
by: Bugueño, Margarita, et al.
Published: (2023)
Rethinking Graph-Based Document Classification: Learning Data-Driven Structures Beyond Heuristic Approaches
by: Bugueño, Margarita, et al.
Published: (2025)
by: Bugueño, Margarita, et al.
Published: (2025)
Investigating Wit, Creativity, and Detectability of Large Language Models in Domain-Specific Writing Style Adaptation of Reddit's Showerthoughts
by: Buz, Tolga, et al.
Published: (2024)
by: Buz, Tolga, et al.
Published: (2024)
GraphLSS: Integrating Lexical, Structural, and Semantic Features for Long Document Extractive Summarization
by: Bugueño, Margarita, et al.
Published: (2024)
by: Bugueño, Margarita, et al.
Published: (2024)
Confabulations from ACL Publications (CAP): A Dataset for Scientific Hallucination Detection
by: Gamba, Federica, et al.
Published: (2025)
by: Gamba, Federica, et al.
Published: (2025)
Hallucination Detection with the Internal Layers of LLMs
by: Preiß, Martin
Published: (2025)
by: Preiß, Martin
Published: (2025)
ReCoQA: A Benchmark for Tool-Augmented and Multi-Step Reasoning in Real Estate Question and Answering
by: Zhang, Yindong, et al.
Published: (2026)
by: Zhang, Yindong, et al.
Published: (2026)
Critical Confabulation: Can LLMs Hallucinate for Social Good?
by: Sui, Peiqi, et al.
Published: (2025)
by: Sui, Peiqi, et al.
Published: (2025)
Can LLMs Detect Their Confabulations? Estimating Reliability in Uncertainty-Aware Language Models
by: Zhou, Tianyi, et al.
Published: (2025)
by: Zhou, Tianyi, et al.
Published: (2025)
Confabulation: The Surprising Value of Large Language Model Hallucinations
by: Sui, Peiqi, et al.
Published: (2024)
by: Sui, Peiqi, et al.
Published: (2024)
Synthetic Fluency: Hallucinations, Confabulations, and the Creation of Irish Words in LLM-Generated Translations
by: Castilho, Sheila, et al.
Published: (2025)
by: Castilho, Sheila, et al.
Published: (2025)
Anchored Confabulation: Partial Evidence Non-Monotonically Amplifies Confident Hallucination in LLMs
by: Lathkar, Ashish Balkishan
Published: (2026)
by: Lathkar, Ashish Balkishan
Published: (2026)
Hidden Measurement Error in LLM Pipelines Distorts Annotation, Evaluation, and Benchmarking
by: Messing, Solomon
Published: (2026)
by: Messing, Solomon
Published: (2026)
Donkii: Can Annotation Error Detection Methods Find Errors in Instruction-Tuning Datasets?
by: Weber-Genzel, Leon, et al.
Published: (2023)
by: Weber-Genzel, Leon, et al.
Published: (2023)
An Annotated Dataset of Errors in Premodern Greek and Baselines for Detecting Them
by: Brooks, Creston, et al.
Published: (2024)
by: Brooks, Creston, et al.
Published: (2024)
MedErrBench: A Fine-Grained Multilingual Benchmark for Medical Error Detection and Correction with Clinical Expert Annotations
by: Ma, Congbo, et al.
Published: (2026)
by: Ma, Congbo, et al.
Published: (2026)
Detecting Reference Errors in Scientific Literature with Large Language Models
by: Zhang, Tianmai M., et al.
Published: (2024)
by: Zhang, Tianmai M., et al.
Published: (2024)
Persona Inconstancy in Multi-Agent LLM Collaboration: Conformity, Confabulation, and Impersonation
by: Baltaji, Razan, et al.
Published: (2024)
by: Baltaji, Razan, et al.
Published: (2024)
Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency
by: Smith, Matthew L., et al.
Published: (2026)
by: Smith, Matthew L., et al.
Published: (2026)
CommitBench: A Benchmark for Commit Message Generation
by: Schall, Maximilian, et al.
Published: (2024)
by: Schall, Maximilian, et al.
Published: (2024)
HUKUKBERT: Domain-Specific Language Model for Turkish Law
by: Öztürk, Mehmet Utku, et al.
Published: (2026)
by: Öztürk, Mehmet Utku, et al.
Published: (2026)
FLAWS: A Benchmark for Error Identification and Localization in Scientific Papers
by: Xi, Sarina, et al.
Published: (2025)
by: Xi, Sarina, et al.
Published: (2025)
Are Online Sports Fan Communities Becoming More Offensive? A Quantitative Review of Topics, Trends, and Toxicity of r/PremierLeague
by: Mazhar, Muhammad Zeeshan, et al.
Published: (2025)
by: Mazhar, Muhammad Zeeshan, et al.
Published: (2025)
Stile extremistischer Tatschreiben
by: Preiss, Ulrike
Published: (2024)
by: Preiss, Ulrike
Published: (2024)
MASSW: A New Dataset and Benchmark Tasks for AI-Assisted Scientific Workflows
by: Zhang, Xingjian, et al.
Published: (2024)
by: Zhang, Xingjian, et al.
Published: (2024)
FACT: Examining the Effectiveness of Iterative Context Rewriting for Multi-fact Retrieval
by: Wang, Jinlin, et al.
Published: (2024)
by: Wang, Jinlin, et al.
Published: (2024)
Consistent Document-Level Relation Extraction via Counterfactuals
by: Modarressi, Ali, et al.
Published: (2024)
by: Modarressi, Ali, et al.
Published: (2024)
Hybrid Human-LLM Corpus Construction and LLM Evaluation for Rare Linguistic Phenomena
by: Weissweiler, Leonie, et al.
Published: (2024)
by: Weissweiler, Leonie, et al.
Published: (2024)
$\textit{BenchIE}^{FL}$ : A Manually Re-Annotated Fact-Based Open Information Extraction Benchmark
by: Lamarche, Fabrice, et al.
Published: (2024)
by: Lamarche, Fabrice, et al.
Published: (2024)
ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection
by: Yan, Yibo, et al.
Published: (2024)
by: Yan, Yibo, et al.
Published: (2024)
ErAConD : Error Annotated Conversational Dialog Dataset for Grammatical Error Correction
by: Yuan, Xun, et al.
Published: (2021)
by: Yuan, Xun, et al.
Published: (2021)
FACT-AUDIT: An Adaptive Multi-Agent Framework for Dynamic Fact-Checking Evaluation of Large Language Models
by: Lin, Hongzhan, et al.
Published: (2025)
by: Lin, Hongzhan, et al.
Published: (2025)
FindTheFlaws: Annotated Errors for Detecting Flawed Reasoning and Scalable Oversight Research
by: Recchia, Gabriel, et al.
Published: (2025)
by: Recchia, Gabriel, et al.
Published: (2025)
Re-examining Sexism and Misogyny Classification with Annotator Attitudes
by: Jiang, Aiqi, et al.
Published: (2024)
by: Jiang, Aiqi, et al.
Published: (2024)
SI-FACT: Mitigating Knowledge Conflict via Self-Improving Faithfulness-Aware Contrastive Tuning
by: Fu, Shengqiang
Published: (2025)
by: Fu, Shengqiang
Published: (2025)
Claim Check-Worthiness Detection: How Well do LLMs Grasp Annotation Guidelines?
by: Majer, Laura, et al.
Published: (2024)
by: Majer, Laura, et al.
Published: (2024)
AR-BENCH: Benchmarking Legal Reasoning with Judgment Error Detection, Classification and Correction
by: Li, Yifei, et al.
Published: (2026)
by: Li, Yifei, et al.
Published: (2026)
Marking: Visual Grading with Highlighting Errors and Annotating Missing Bits
by: Sonkar, Shashank, et al.
Published: (2024)
by: Sonkar, Shashank, et al.
Published: (2024)
FOCUS: Effective Embedding Initialization for Monolingual Specialization of Multilingual Models
by: Dobler, Konstantin, et al.
Published: (2023)
by: Dobler, Konstantin, et al.
Published: (2023)
Similar Items
-
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
by: Arad, Dana, et al.
Published: (2023) -
Connecting the Dots: What Graph-Based Text Representations Work Best for Text Classification Using Graph Neural Networks?
by: Bugueño, Margarita, et al.
Published: (2023) -
Rethinking Graph-Based Document Classification: Learning Data-Driven Structures Beyond Heuristic Approaches
by: Bugueño, Margarita, et al.
Published: (2025) -
Investigating Wit, Creativity, and Detectability of Large Language Models in Domain-Specific Writing Style Adaptation of Reddit's Showerthoughts
by: Buz, Tolga, et al.
Published: (2024) -
GraphLSS: Integrating Lexical, Structural, and Semantic Features for Long Document Extractive Summarization
by: Bugueño, Margarita, et al.
Published: (2024)