SciFaultyQA: Benchmarking LLMs on Faulty Science Question Detection with a GAN-Inspired Approach to Synthetic Dataset Generation
Fuente:
arXiv
Saved in:
| Main Author: | Kundu, Debarshi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tools Fail: Detecting Silent Errors in Faulty Tools
by: Sun, Jimin, et al.
Published: (2024)
by: Sun, Jimin, et al.
Published: (2024)
Classification of User Reports for Detection of Faulty Computer Components using NLP Models: A Case Study
by: Silva, Maria de Lourdes M., et al.
Published: (2025)
by: Silva, Maria de Lourdes M., et al.
Published: (2025)
FoQA: A Faroese Question-Answering Dataset
by: Simonsen, Annika, et al.
Published: (2025)
by: Simonsen, Annika, et al.
Published: (2025)
Minder: Faulty Machine Detection for Large-scale Distributed Model Training
by: Deng, Yangtao, et al.
Published: (2024)
by: Deng, Yangtao, et al.
Published: (2024)
PeruMedQA: Benchmarking Large Language Models (LLMs) on Peruvian Medical Exams -- Dataset Construction and Evaluation
by: Carrillo-Larco, Rodrigo M., et al.
Published: (2025)
by: Carrillo-Larco, Rodrigo M., et al.
Published: (2025)
The Influence of Faulty Labels in Data Sets on Human Pose Estimation
by: Schwarz, Arnold, et al.
Published: (2024)
by: Schwarz, Arnold, et al.
Published: (2024)
From Blind Solvers to Logical Thinkers: Benchmarking LLMs' Logical Integrity on Faulty Mathematical Problems
by: Rahman, A M Muntasir, et al.
Published: (2024)
by: Rahman, A M Muntasir, et al.
Published: (2024)
Synthetic Context Generation for Question Generation
by: Liu, Naiming, et al.
Published: (2024)
by: Liu, Naiming, et al.
Published: (2024)
SyllabusQA: A Course Logistics Question Answering Dataset
by: Fernandez, Nigel, et al.
Published: (2024)
by: Fernandez, Nigel, et al.
Published: (2024)
LaMP-QA: A Benchmark for Personalized Long-form Question Answering
by: Salemi, Alireza, et al.
Published: (2025)
by: Salemi, Alireza, et al.
Published: (2025)
MedConceptsQA: Open Source Medical Concepts QA Benchmark
by: Shoham, Ofir Ben, et al.
Published: (2024)
by: Shoham, Ofir Ben, et al.
Published: (2024)
L3Cube-IndicQuest: A Benchmark Question Answering Dataset for Evaluating Knowledge of LLMs in Indic Context
by: Rohera, Pritika, et al.
Published: (2024)
by: Rohera, Pritika, et al.
Published: (2024)
Synthetic Multimodal Question Generation
by: Wu, Ian, et al.
Published: (2024)
by: Wu, Ian, et al.
Published: (2024)
DataSciBench: An LLM Agent Benchmark for Data Science
by: Zhang, Dan, et al.
Published: (2025)
by: Zhang, Dan, et al.
Published: (2025)
WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
VOYAGER: A Training Free Approach for Generating Diverse Datasets using LLMs
by: Amballa, Avinash, et al.
Published: (2025)
by: Amballa, Avinash, et al.
Published: (2025)
FictionalQA: A Dataset for Studying Memorization and Knowledge Acquisition
by: Kirchenbauer, John, et al.
Published: (2025)
by: Kirchenbauer, John, et al.
Published: (2025)
RLVR Training of LLMs Does Not Improve Thinking Ability for General QA: Evaluation Method and a Simple Solution
by: Li, Kaiyuan, et al.
Published: (2026)
by: Li, Kaiyuan, et al.
Published: (2026)
GINopic: Topic Modeling with Graph Isomorphism Network
by: Adhya, Suman, et al.
Published: (2024)
by: Adhya, Suman, et al.
Published: (2024)
Benchmarking Hindi LLMs: A New Suite of Datasets and a Comparative Analysis
by: Kamath, Anusha, et al.
Published: (2025)
by: Kamath, Anusha, et al.
Published: (2025)
A Benchmark Dataset with Larger Context for Non-Factoid Question Answering over Islamic Text
by: Qamar, Faiza, et al.
Published: (2024)
by: Qamar, Faiza, et al.
Published: (2024)
Balancing Cost and Effectiveness of Synthetic Data Generation Strategies for LLMs
by: Chan, Yung-Chieh, et al.
Published: (2024)
by: Chan, Yung-Chieh, et al.
Published: (2024)
Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework
by: Dycke, Nils, et al.
Published: (2025)
by: Dycke, Nils, et al.
Published: (2025)
ElectroVizQA: How well do Multi-modal LLMs perform in Electronics Visual Question Answering?
by: Meshram, Pragati Shuddhodhan, et al.
Published: (2024)
by: Meshram, Pragati Shuddhodhan, et al.
Published: (2024)
Answering Questions in Stages: Prompt Chaining for Contract QA
by: Roegiest, Adam, et al.
Published: (2024)
by: Roegiest, Adam, et al.
Published: (2024)
MeDiSumQA: Patient-Oriented Question-Answer Generation from Discharge Letters
by: Dada, Amin, et al.
Published: (2025)
by: Dada, Amin, et al.
Published: (2025)
Efficacy of Synthetic Data as a Benchmark
by: Maheshwari, Gaurav, et al.
Published: (2024)
by: Maheshwari, Gaurav, et al.
Published: (2024)
ViQA-COVID: COVID-19 Machine Reading Comprehension Dataset for Vietnamese
by: Nguyen-Phung, Hai-Chung, et al.
Published: (2025)
by: Nguyen-Phung, Hai-Chung, et al.
Published: (2025)
LatentQA: Teaching LLMs to Decode Activations Into Natural Language
by: Pan, Alexander, et al.
Published: (2024)
by: Pan, Alexander, et al.
Published: (2024)
Neural at ArchEHR-QA 2025: Agentic Prompt Optimization for Evidence-Grounded Clinical Question Answering
by: Bogireddy, Sai Prasanna Teja Reddy, et al.
Published: (2025)
by: Bogireddy, Sai Prasanna Teja Reddy, et al.
Published: (2025)
CTG-KrEW: Generating Synthetic Structured Contextually Correlated Content by Conditional Tabular GAN with K-Means Clustering and Efficient Word Embedding
by: Samanta, Riya, et al.
Published: (2024)
by: Samanta, Riya, et al.
Published: (2024)
Integrating Domain Knowledge for Financial QA: A Multi-Retriever RAG Approach with LLMs
by: Zhang, Yukun, et al.
Published: (2025)
by: Zhang, Yukun, et al.
Published: (2025)
QuIM-RAG: Advancing Retrieval-Augmented Generation with Inverted Question Matching for Enhanced QA Performance
by: Saha, Binita, et al.
Published: (2025)
by: Saha, Binita, et al.
Published: (2025)
SciLitLLM: How to Adapt LLMs for Scientific Literature Understanding
by: Li, Sihang, et al.
Published: (2024)
by: Li, Sihang, et al.
Published: (2024)
On the Calibration of Multilingual Question Answering LLMs
by: Yang, Yahan, et al.
Published: (2023)
by: Yang, Yahan, et al.
Published: (2023)
HealthNLP_Retrievers at ArchEHR-QA 2026: Cascaded LLM Pipeline for Grounded Clinical Question Answering
by: Hosen, Md Biplob, et al.
Published: (2026)
by: Hosen, Md Biplob, et al.
Published: (2026)
Exploring the Limits of Model Compression in LLMs: A Knowledge Distillation Study on QA Tasks
by: Datta, Joyeeta, et al.
Published: (2025)
by: Datta, Joyeeta, et al.
Published: (2025)
SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA
by: Xu, Haozhou, et al.
Published: (2025)
by: Xu, Haozhou, et al.
Published: (2025)
FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models
by: Mateega, Spencer, et al.
Published: (2025)
by: Mateega, Spencer, et al.
Published: (2025)
Zero-Shot Conversational Stance Detection: Dataset and Approaches
by: Ding, Yuzhe, et al.
Published: (2025)
by: Ding, Yuzhe, et al.
Published: (2025)
Similar Items
-
Tools Fail: Detecting Silent Errors in Faulty Tools
by: Sun, Jimin, et al.
Published: (2024) -
Classification of User Reports for Detection of Faulty Computer Components using NLP Models: A Case Study
by: Silva, Maria de Lourdes M., et al.
Published: (2025) -
FoQA: A Faroese Question-Answering Dataset
by: Simonsen, Annika, et al.
Published: (2025) -
Minder: Faulty Machine Detection for Large-scale Distributed Model Training
by: Deng, Yangtao, et al.
Published: (2024) -
PeruMedQA: Benchmarking Large Language Models (LLMs) on Peruvian Medical Exams -- Dataset Construction and Evaluation
by: Carrillo-Larco, Rodrigo M., et al.
Published: (2025)