Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Miyai, Atsuyuki, Toyooka, Mashiro, Zhao, Zaiying, Watanabe, Kenta, Yamasaki, Toshihiko, Aizawa, Kiyoharu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
A Highly Clean Recipe Dataset with Ingredient States Annotation for State Probing Task
von: Toyooka, Mashiro, et al.
Veröffentlicht: (2025)
von: Toyooka, Mashiro, et al.
Veröffentlicht: (2025)
PULSE: Practical Evaluation Scenarios for Large Multimodal Model Unlearning
von: Kawakami, Tatsuki, et al.
Veröffentlicht: (2025)
von: Kawakami, Tatsuki, et al.
Veröffentlicht: (2025)
JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
Manga109-v2026: Revisiting Manga109 Annotations for Modern Manga Understanding
von: Baek, Jeonghun, et al.
Veröffentlicht: (2026)
von: Baek, Jeonghun, et al.
Veröffentlicht: (2026)
JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation
von: Onohara, Shota, et al.
Veröffentlicht: (2024)
von: Onohara, Shota, et al.
Veröffentlicht: (2024)
Generalized Out-of-Distribution Detection and Beyond in Vision Language Model Era: A Survey
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
Paper2Agent: Reimagining Research Papers As Interactive and Reliable AI Agents
von: Miao, Jiacheng, et al.
Veröffentlicht: (2025)
von: Miao, Jiacheng, et al.
Veröffentlicht: (2025)
Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection
von: Kulkarni, Atharva, et al.
Veröffentlicht: (2025)
von: Kulkarni, Atharva, et al.
Veröffentlicht: (2025)
Evaluating the Relevance of Uncertainty Estimators for LLM Hallucination
von: Agnimo, Yedidia, et al.
Veröffentlicht: (2026)
von: Agnimo, Yedidia, et al.
Veröffentlicht: (2026)
Harnessing PDF Data for Improving Japanese Large Multimodal Models
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation
von: Mündler, Niels, et al.
Veröffentlicht: (2023)
von: Mündler, Niels, et al.
Veröffentlicht: (2023)
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
von: Sansford, Hannah, et al.
Veröffentlicht: (2024)
von: Sansford, Hannah, et al.
Veröffentlicht: (2024)
PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR
von: Burgess, James, et al.
Veröffentlicht: (2026)
von: Burgess, James, et al.
Veröffentlicht: (2026)
Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models
von: Rahman, Subhey Sadi, et al.
Veröffentlicht: (2025)
von: Rahman, Subhey Sadi, et al.
Veröffentlicht: (2025)
TreeCut: A Synthetic Unanswerable Math Word Problem Dataset for LLM Hallucination Evaluation
von: Ouyang, Jialin
Veröffentlicht: (2025)
von: Ouyang, Jialin
Veröffentlicht: (2025)
BEExAI: Benchmark to Evaluate Explainable AI
von: Sithakoul, Samuel, et al.
Veröffentlicht: (2024)
von: Sithakoul, Samuel, et al.
Veröffentlicht: (2024)
cPAPERS: A Dataset of Situated and Multimodal Interactive Conversations in Scientific Papers
von: Sundar, Anirudh, et al.
Veröffentlicht: (2024)
von: Sundar, Anirudh, et al.
Veröffentlicht: (2024)
PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants
von: Zhao, Zheng, et al.
Veröffentlicht: (2025)
von: Zhao, Zheng, et al.
Veröffentlicht: (2025)
PaperBench: Evaluating AI's Ability to Replicate AI Research
von: Starace, Giulio, et al.
Veröffentlicht: (2025)
von: Starace, Giulio, et al.
Veröffentlicht: (2025)
ReEval: Automatic Hallucination Evaluation for Retrieval-Augmented Large Language Models via Transferable Adversarial Attacks
von: Yu, Xiaodong, et al.
Veröffentlicht: (2023)
von: Yu, Xiaodong, et al.
Veröffentlicht: (2023)
CodeRefine: A Pipeline for Enhancing LLM-Generated Code Implementations of Research Papers
von: Trofimova, Ekaterina, et al.
Veröffentlicht: (2024)
von: Trofimova, Ekaterina, et al.
Veröffentlicht: (2024)
LLM-Powered Ensemble Learning for Paper Source Tracing: A GPU-Free Approach
von: Chen, Kunlong, et al.
Veröffentlicht: (2024)
von: Chen, Kunlong, et al.
Veröffentlicht: (2024)
HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL Conferences
von: Sakai, Yusuke, et al.
Veröffentlicht: (2026)
von: Sakai, Yusuke, et al.
Veröffentlicht: (2026)
DefAn: Definitive Answer Dataset for LLMs Hallucination Evaluation
von: Rahman, A B M Ashikur, et al.
Veröffentlicht: (2024)
von: Rahman, A B M Ashikur, et al.
Veröffentlicht: (2024)
Fine-Tuning and Evaluating Conversational AI for Agricultural Advisory
von: Singh, Sanyam, et al.
Veröffentlicht: (2026)
von: Singh, Sanyam, et al.
Veröffentlicht: (2026)
Toward the Evaluation of Large Language Models Considering Score Variance across Instruction Templates
von: Sakai, Yusuke, et al.
Veröffentlicht: (2024)
von: Sakai, Yusuke, et al.
Veröffentlicht: (2024)
Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks
von: Wang, Jianghui, et al.
Veröffentlicht: (2025)
von: Wang, Jianghui, et al.
Veröffentlicht: (2025)
The Phenomenology of Hallucinations
von: Ruscio, Valeria, et al.
Veröffentlicht: (2026)
von: Ruscio, Valeria, et al.
Veröffentlicht: (2026)
Causal Evaluation of Language Models
von: Chen, Sirui, et al.
Veröffentlicht: (2024)
von: Chen, Sirui, et al.
Veröffentlicht: (2024)
Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking
von: Fang, Zhicheng, et al.
Veröffentlicht: (2026)
von: Fang, Zhicheng, et al.
Veröffentlicht: (2026)
FLAWS: A Benchmark for Error Identification and Localization in Scientific Papers
von: Xi, Sarina, et al.
Veröffentlicht: (2025)
von: Xi, Sarina, et al.
Veröffentlicht: (2025)
SciPIP: An LLM-based Scientific Paper Idea Proposer
von: Wang, Wenxiao, et al.
Veröffentlicht: (2024)
von: Wang, Wenxiao, et al.
Veröffentlicht: (2024)
A Critical Evaluation of AI Feedback for Aligning Large Language Models
von: Sharma, Archit, et al.
Veröffentlicht: (2024)
von: Sharma, Archit, et al.
Veröffentlicht: (2024)
A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI
von: Karagoz, Atahan
Veröffentlicht: (2026)
von: Karagoz, Atahan
Veröffentlicht: (2026)
Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
von: Zhang, Yue, et al.
Veröffentlicht: (2023)
von: Zhang, Yue, et al.
Veröffentlicht: (2023)
APRES: An Agentic Paper Revision and Evaluation System
von: Zhao, Bingchen, et al.
Veröffentlicht: (2026)
von: Zhao, Bingchen, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025) -
WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025) -
A Highly Clean Recipe Dataset with Ingredient States Annotation for State Probing Task
von: Toyooka, Mashiro, et al.
Veröffentlicht: (2025) -
PULSE: Practical Evaluation Scenarios for Large Multimodal Model Unlearning
von: Kawakami, Tatsuki, et al.
Veröffentlicht: (2025) -
JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)