WithdrarXiv: A Large-Scale Dataset for Retraction Study
Fuente:
arXiv
Saved in:
| Main Authors: | Rao, Delip, Young, Jonathan, Dietterich, Thomas, Callison-Burch, Chris |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
What Do Claim Verification Datasets Actually Test? A Reasoning Trace Analysis
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
Autorubric: Unifying Rubric-based LLM Evaluation
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
Reading the unreadable: Creating a dataset of 19th century English newspapers using image-to-text language models
by: Bourne, Jonathan
Published: (2025)
by: Bourne, Jonathan
Published: (2025)
VTechAGP: An Academic-to-General-Audience Text Paraphrase Dataset and Benchmark Models
by: Cheng, Ming, et al.
Published: (2024)
by: Cheng, Ming, et al.
Published: (2024)
On the Effectiveness of Large Language Models in Automating Categorization of Scientific Texts
by: Shahi, Gautam Kishore, et al.
Published: (2025)
by: Shahi, Gautam Kishore, et al.
Published: (2025)
Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
Probabilistic Soundness Guarantees in LLM Reasoning Chains
by: You, Weiqiu, et al.
Published: (2025)
by: You, Weiqiu, et al.
Published: (2025)
FLAWS: A Benchmark for Error Identification and Localization in Scientific Papers
by: Xi, Sarina, et al.
Published: (2025)
by: Xi, Sarina, et al.
Published: (2025)
ThinknCheck: Grounded Claim Verification with Compact, Reasoning-Driven, and Interpretable Models
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
Publication Trend Analysis and Synthesis via Large Language Model: A Case Study of Engineering in PNAS
by: Smetana, Mason, et al.
Published: (2025)
by: Smetana, Mason, et al.
Published: (2025)
NLLG Quarterly arXiv Report 09/24: What are the most influential current AI Papers?
by: Leiter, Christoph, et al.
Published: (2024)
by: Leiter, Christoph, et al.
Published: (2024)
A History of Philosophy in Colombia through Topic Modelling
by: Loaiza, Juan R., et al.
Published: (2024)
by: Loaiza, Juan R., et al.
Published: (2024)
OAG-Bench: A Human-Curated Benchmark for Academic Graph Mining
by: Zhang, Fanjin, et al.
Published: (2024)
by: Zhang, Fanjin, et al.
Published: (2024)
FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
by: Patel, Ajay, et al.
Published: (2026)
by: Patel, Ajay, et al.
Published: (2026)
Falcon 7b for Software Mention Detection in Scholarly Documents
by: Khan, AmeerAli, et al.
Published: (2024)
by: Khan, AmeerAli, et al.
Published: (2024)
Fine-tuning and Prompt Engineering with Cognitive Knowledge Graphs for Scholarly Knowledge Organization
by: Rabby, Gollam, et al.
Published: (2024)
by: Rabby, Gollam, et al.
Published: (2024)
Learning representations of learning representations
by: González-Márquez, Rita, et al.
Published: (2024)
by: González-Márquez, Rita, et al.
Published: (2024)
Post-OCR Text Correction for Bulgarian Historical Documents
by: Beshirov, Angel, et al.
Published: (2024)
by: Beshirov, Angel, et al.
Published: (2024)
Hierarchical Tree-structured Knowledge Graph For Academic Insight Survey
by: Li, Jinghong, et al.
Published: (2024)
by: Li, Jinghong, et al.
Published: (2024)
Automating Violence Detection and Categorization from Ancient Texts
by: Abdelhalim, Alhassan, et al.
Published: (2025)
by: Abdelhalim, Alhassan, et al.
Published: (2025)
Machine Learning Research Has Outpaced Its Communication Norms and NeurIPS Should Act
by: Rangarajan, Ajay Mandyam, et al.
Published: (2026)
by: Rangarajan, Ajay Mandyam, et al.
Published: (2026)
Topics, Authors, and Institutions in Large Language Model Research: Trends from 17K arXiv Papers
by: Movva, Rajiv, et al.
Published: (2023)
by: Movva, Rajiv, et al.
Published: (2023)
C$^2$-Cite: Contextual-Aware Citation Generation for Attributed Large Language Models
by: Yu, Yue, et al.
Published: (2025)
by: Yu, Yue, et al.
Published: (2025)
SciHorizon: Benchmarking AI-for-Science Readiness from Scientific Data to Large Language Models
by: Qin, Chuan, et al.
Published: (2025)
by: Qin, Chuan, et al.
Published: (2025)
Scientific Statement Classification over arXiv.org
by: Ginev, Deyan, et al.
Published: (2019)
by: Ginev, Deyan, et al.
Published: (2019)
NSF-SciFy: Mining the NSF Awards Database for Scientific Claims
by: Rao, Delip, et al.
Published: (2025)
by: Rao, Delip, et al.
Published: (2025)
LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
by: Elazar, Yanai, et al.
Published: (2026)
by: Elazar, Yanai, et al.
Published: (2026)
SyROCCo: Enhancing Systematic Reviews using Machine Learning
by: Fang, Zheng, et al.
Published: (2024)
by: Fang, Zheng, et al.
Published: (2024)
The Impact of Large Language Models in Academia: from Writing to Speaking
by: Geng, Mingmeng, et al.
Published: (2024)
by: Geng, Mingmeng, et al.
Published: (2024)
Unlocking the Archives: Using Large Language Models to Transcribe Handwritten Historical Documents
by: Humphries, Mark, et al.
Published: (2024)
by: Humphries, Mark, et al.
Published: (2024)
Beyond Via: Analysis and Estimation of the Impact of Large Language Models in Academic Papers
by: Geng, Mingmeng, et al.
Published: (2026)
by: Geng, Mingmeng, et al.
Published: (2026)
FMMD: A multimodal open peer review dataset based on F1000Research
by: Zhuang, Zhenzhen, et al.
Published: (2026)
by: Zhuang, Zhenzhen, et al.
Published: (2026)
When Verification Fails: How Compositionally Infeasible Claims Escape Rejection
by: Liu, Muxin, et al.
Published: (2026)
by: Liu, Muxin, et al.
Published: (2026)
DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows
by: Patel, Ajay, et al.
Published: (2024)
by: Patel, Ajay, et al.
Published: (2024)
Combining topic modelling and citation network analysis to study case law from the European Court on Human Rights on the right to respect for private and family life
by: Mohammadi, M., et al.
Published: (2024)
by: Mohammadi, M., et al.
Published: (2024)
Is ChatGPT Transforming Academics' Writing Style?
by: Geng, Mingmeng, et al.
Published: (2024)
by: Geng, Mingmeng, et al.
Published: (2024)
'Quis custodiet ipsos custodes?' Who will watch the watchmen? On Detecting AI-generated peer-reviews
by: Kumar, Sandeep, et al.
Published: (2024)
by: Kumar, Sandeep, et al.
Published: (2024)
Efficient Systematic Reviews: Literature Filtering with Transformers & Transfer Learning
by: Hawkins, John, et al.
Published: (2024)
by: Hawkins, John, et al.
Published: (2024)
Similar Items
-
BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation
by: Rao, Delip, et al.
Published: (2026) -
What Do Claim Verification Datasets Actually Test? A Reasoning Trace Analysis
by: Rao, Delip, et al.
Published: (2026) -
Autorubric: Unifying Rubric-based LLM Evaluation
by: Rao, Delip, et al.
Published: (2026) -
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
by: Rao, Delip, et al.
Published: (2026) -
Reading the unreadable: Creating a dataset of 19th century English newspapers using image-to-text language models
by: Bourne, Jonathan
Published: (2025)