FLAWS: A Benchmark for Error Identification and Localization in Scientific Papers
Fuente:
arXiv
Saved in:
| Main Authors: | Xi, Sarina, Rao, Vishisht, Payan, Justin, Shah, Nihar B. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ML Researchers Support Openness in Peer Review But Are Concerned About Resubmission Bias
by: Rao, Vishisht, et al.
Published: (2025)
by: Rao, Vishisht, et al.
Published: (2025)
Detecting LLM-Generated Peer Reviews
by: Rao, Vishisht, et al.
Published: (2025)
by: Rao, Vishisht, et al.
Published: (2025)
Mapping the Increasing Use of LLMs in Scientific Papers
by: Liang, Weixin, et al.
Published: (2024)
by: Liang, Weixin, et al.
Published: (2024)
LitSearch: A Retrieval Benchmark for Scientific Literature Search
by: Ajith, Anirudh, et al.
Published: (2024)
by: Ajith, Anirudh, et al.
Published: (2024)
Interesting Scientific Idea Generation using Knowledge Graphs and LLMs: Evaluations with 100 Research Group Leaders
by: Gu, Xuemei, et al.
Published: (2024)
by: Gu, Xuemei, et al.
Published: (2024)
Beyond Via: Analysis and Estimation of the Impact of Large Language Models in Academic Papers
by: Geng, Mingmeng, et al.
Published: (2026)
by: Geng, Mingmeng, et al.
Published: (2026)
A Gold Standard Dataset for the Reviewer Assignment Problem
by: Stelmakh, Ivan, et al.
Published: (2023)
by: Stelmakh, Ivan, et al.
Published: (2023)
Accelerating Scientific Discovery with Multi-Document Summarization of Impact-Ranked Papers
by: Koloveas, Paris, et al.
Published: (2025)
by: Koloveas, Paris, et al.
Published: (2025)
OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs
by: Asai, Akari, et al.
Published: (2024)
by: Asai, Akari, et al.
Published: (2024)
FMMD: A multimodal open peer review dataset based on F1000Research
by: Zhuang, Zhenzhen, et al.
Published: (2026)
by: Zhuang, Zhenzhen, et al.
Published: (2026)
Publication Trend Analysis and Synthesis via Large Language Model: A Case Study of Engineering in PNAS
by: Smetana, Mason, et al.
Published: (2025)
by: Smetana, Mason, et al.
Published: (2025)
Is ChatGPT Transforming Academics' Writing Style?
by: Geng, Mingmeng, et al.
Published: (2024)
by: Geng, Mingmeng, et al.
Published: (2024)
'Quis custodiet ipsos custodes?' Who will watch the watchmen? On Detecting AI-generated peer-reviews
by: Kumar, Sandeep, et al.
Published: (2024)
by: Kumar, Sandeep, et al.
Published: (2024)
Efficient Systematic Reviews: Literature Filtering with Transformers & Transfer Learning
by: Hawkins, John, et al.
Published: (2024)
by: Hawkins, John, et al.
Published: (2024)
LitLLMs, LLMs for Literature Review: Are we there yet?
by: Agarwal, Shubham, et al.
Published: (2024)
by: Agarwal, Shubham, et al.
Published: (2024)
TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models
by: Nadas, Mihai, et al.
Published: (2025)
by: Nadas, Mihai, et al.
Published: (2025)
SemEval-2025 Task 5: LLMs4Subjects -- LLM-based Automated Subject Tagging for a National Technical Library's Open-Access Catalog
by: D'Souza, Jennifer, et al.
Published: (2025)
by: D'Souza, Jennifer, et al.
Published: (2025)
Can social media provide early warning of retraction? Evidence from critical tweets identified by human annotation and large language models
by: Zheng, Er-Te, et al.
Published: (2024)
by: Zheng, Er-Te, et al.
Published: (2024)
Mining for Species, Locations, Habitats, and Ecosystems from Scientific Papers in Invasion Biology: A Large-Scale Exploratory Study with Large Language Models
by: D'Souza, Jennifer, et al.
Published: (2025)
by: D'Souza, Jennifer, et al.
Published: (2025)
NLLG Quarterly arXiv Report 09/24: What are the most influential current AI Papers?
by: Leiter, Christoph, et al.
Published: (2024)
by: Leiter, Christoph, et al.
Published: (2024)
How to Find Fantastic AI Papers: Self-Rankings as a Powerful Predictor of Scientific Impact Beyond Peer Review
by: Su, Buxin, et al.
Published: (2025)
by: Su, Buxin, et al.
Published: (2025)
Scientific Statement Classification over arXiv.org
by: Ginev, Deyan, et al.
Published: (2019)
by: Ginev, Deyan, et al.
Published: (2019)
Meursault as a Data Point
by: Pratap, Abhinav
Published: (2025)
by: Pratap, Abhinav
Published: (2025)
Human-LLM Coevolution: Evidence from Academic Writing
by: Geng, Mingmeng, et al.
Published: (2025)
by: Geng, Mingmeng, et al.
Published: (2025)
The Impact of Large Language Models in Academia: from Writing to Speaking
by: Geng, Mingmeng, et al.
Published: (2024)
by: Geng, Mingmeng, et al.
Published: (2024)
The More You Automate, the Less You See: Hidden Pitfalls of AI Scientist Systems
by: Luo, Ziming, et al.
Published: (2025)
by: Luo, Ziming, et al.
Published: (2025)
Usefulness of LLMs as an Author Checklist Assistant for Scientific Papers: NeurIPS'24 Experiment
by: Goldberg, Alexander, et al.
Published: (2024)
by: Goldberg, Alexander, et al.
Published: (2024)
Triples and Knowledge-Infused Embeddings for Clustering and Classification of Scientific Documents
by: Arcan, Mihael
Published: (2025)
by: Arcan, Mihael
Published: (2025)
Annotating Scientific Uncertainty: A comprehensive model using linguistic patterns and comparison with existing approaches
by: Ningrum, Panggih Kusuma, et al.
Published: (2025)
by: Ningrum, Panggih Kusuma, et al.
Published: (2025)
LLMs4Synthesis: Leveraging Large Language Models for Scientific Synthesis
by: Giglou, Hamed Babaei, et al.
Published: (2024)
by: Giglou, Hamed Babaei, et al.
Published: (2024)
PaperWeaver: Enriching Topical Paper Alerts by Contextualizing Recommended Papers with User-collected Papers
by: Lee, Yoonjoo, et al.
Published: (2024)
by: Lee, Yoonjoo, et al.
Published: (2024)
SciHorizon: Benchmarking AI-for-Science Readiness from Scientific Data to Large Language Models
by: Qin, Chuan, et al.
Published: (2025)
by: Qin, Chuan, et al.
Published: (2025)
De-identification of clinical free text using natural language processing: A systematic review of current approaches
by: Kovačević, Aleksandar, et al.
Published: (2023)
by: Kovačević, Aleksandar, et al.
Published: (2023)
LLMs4SchemaDiscovery: A Human-in-the-Loop Workflow for Scientific Schema Mining with Large Language Models
by: Sadruddin, Sameer, et al.
Published: (2025)
by: Sadruddin, Sameer, et al.
Published: (2025)
Context Selection for Hypothesis and Statistical Evidence Extraction from Full-Text Scientific Articles
by: Koneru, Sai, et al.
Published: (2026)
by: Koneru, Sai, et al.
Published: (2026)
Foundation Models for Scientific Discovery: From Paradigm Enhancement to Paradigm Transition
by: Liu, Fan, et al.
Published: (2025)
by: Liu, Fan, et al.
Published: (2025)
On the Readiness of Scientific Data for a Fair and Transparent Use in Machine Learning
by: Giner-Miguelez, Joan, et al.
Published: (2024)
by: Giner-Miguelez, Joan, et al.
Published: (2024)
SC4ANM: Identifying Optimal Section Combinations for Automated Novelty Prediction in Academic Papers
by: Wu, Wenqing, et al.
Published: (2025)
by: Wu, Wenqing, et al.
Published: (2025)
HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL Conferences
by: Sakai, Yusuke, et al.
Published: (2026)
by: Sakai, Yusuke, et al.
Published: (2026)
On the Effectiveness of Large Language Models in Automating Categorization of Scientific Texts
by: Shahi, Gautam Kishore, et al.
Published: (2025)
by: Shahi, Gautam Kishore, et al.
Published: (2025)
Similar Items
-
ML Researchers Support Openness in Peer Review But Are Concerned About Resubmission Bias
by: Rao, Vishisht, et al.
Published: (2025) -
Detecting LLM-Generated Peer Reviews
by: Rao, Vishisht, et al.
Published: (2025) -
Mapping the Increasing Use of LLMs in Scientific Papers
by: Liang, Weixin, et al.
Published: (2024) -
LitSearch: A Retrieval Benchmark for Scientific Literature Search
by: Ajith, Anirudh, et al.
Published: (2024) -
Interesting Scientific Idea Generation using Knowledge Graphs and LLMs: Evaluations with 100 Research Group Leaders
by: Gu, Xuemei, et al.
Published: (2024)