PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tu, Songjun, Ma, Yiwen, Lin, Jiahao, Zhang, Qichao, Lan, Xiangyuan, Li, Junfeng., Xu, Nan, Li, Linjing, Zhao, Dongbin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL
von: Tu, Songjun, et al.
Veröffentlicht: (2025)
von: Tu, Songjun, et al.
Veröffentlicht: (2025)
Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
von: Tu, Songjun, et al.
Veröffentlicht: (2025)
von: Tu, Songjun, et al.
Veröffentlicht: (2025)
Examining Linguistic Shifts in Academic Writing Before and After the Launch of ChatGPT: A Study on Preprint Papers
von: Bao, Tong, et al.
Veröffentlicht: (2025)
von: Bao, Tong, et al.
Veröffentlicht: (2025)
Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
von: Pan, Leyi, et al.
Veröffentlicht: (2025)
von: Pan, Leyi, et al.
Veröffentlicht: (2025)
Emotional Sequential Influence Modeling on False Information
von: Naskar, Debashis, et al.
Veröffentlicht: (2024)
von: Naskar, Debashis, et al.
Veröffentlicht: (2024)
MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare
von: Wang, Yihao, et al.
Veröffentlicht: (2026)
von: Wang, Yihao, et al.
Veröffentlicht: (2026)
IFMTBench: A Comprehensive Benchmark for Multilingual Translation Instruction Following
von: Sun, Mingrui, et al.
Veröffentlicht: (2026)
von: Sun, Mingrui, et al.
Veröffentlicht: (2026)
Co-NAML-LSTUR: A Combined Model with Attentive Multi-View Learning and Long- and Short-term User Representations for News Recommendation
von: Nguyen, Minh Hoang, et al.
Veröffentlicht: (2025)
von: Nguyen, Minh Hoang, et al.
Veröffentlicht: (2025)
Are Non-English Papers Reviewed Fairly? Language-of-Study Bias in NLP Peer Reviews
von: Barkhordar, Ehsan, et al.
Veröffentlicht: (2026)
von: Barkhordar, Ehsan, et al.
Veröffentlicht: (2026)
AI-assisted German Employment Contract Review: A Benchmark Dataset
von: Wardas, Oliver, et al.
Veröffentlicht: (2025)
von: Wardas, Oliver, et al.
Veröffentlicht: (2025)
Compression Method Matters: Benchmark-Dependent Output Dynamics in LLM Prompt Compression
von: Johnson, Warren
Veröffentlicht: (2026)
von: Johnson, Warren
Veröffentlicht: (2026)
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
The Paradox of Poetic Intent in Back-Translation: Evaluating the Quality of Large Language Models in Chinese Translation
von: Weigang, Li, et al.
Veröffentlicht: (2025)
von: Weigang, Li, et al.
Veröffentlicht: (2025)
Copyright Detection in Large Language Models: An Ethical Approach to Generative AI Development
von: Szczecina, David, et al.
Veröffentlicht: (2025)
von: Szczecina, David, et al.
Veröffentlicht: (2025)
Retrieval-Based Multi-Label Legal Annotation: Extensible, Data-Efficient and Hallucination-Free
von: Zhang, Li, et al.
Veröffentlicht: (2026)
von: Zhang, Li, et al.
Veröffentlicht: (2026)
PairCFR: Enhancing Model Training on Paired Counterfactually Augmented Data through Contrastive Learning
von: Qiu, Xiaoqi, et al.
Veröffentlicht: (2024)
von: Qiu, Xiaoqi, et al.
Veröffentlicht: (2024)
Unveiling Attractor Cycles in Large Language Models: A Dynamical Systems View of Successive Paraphrasing
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
Thinking Longer, Not Always Smarter: Evaluating LLM Capabilities in Hierarchical Legal Reasoning
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
New Skills or Sharper Primitives? A Probabilistic Perspective on the Emergence of Reasoning in RLVR
von: Wang, Zhilin, et al.
Veröffentlicht: (2026)
von: Wang, Zhilin, et al.
Veröffentlicht: (2026)
On the Robustness of Document-Level Relation Extraction Models to Entity Name Variations
von: Meng, Shiao, et al.
Veröffentlicht: (2024)
von: Meng, Shiao, et al.
Veröffentlicht: (2024)
NurValues: Real-World Nursing Values Evaluation for Large Language Models in Clinical Context
von: Yao, Ben, et al.
Veröffentlicht: (2025)
von: Yao, Ben, et al.
Veröffentlicht: (2025)
An Unforgeable Publicly Verifiable Watermark for Large Language Models
von: Liu, Aiwei, et al.
Veröffentlicht: (2023)
von: Liu, Aiwei, et al.
Veröffentlicht: (2023)
Empowering Tabular Data Preparation with Language Models: Why and How?
von: Chen, Mengshi, et al.
Veröffentlicht: (2025)
von: Chen, Mengshi, et al.
Veröffentlicht: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
A Survey of Text Watermarking in the Era of Large Language Models
von: Liu, Aiwei, et al.
Veröffentlicht: (2023)
von: Liu, Aiwei, et al.
Veröffentlicht: (2023)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026)
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026)
PathBench: Speech Intelligibility Benchmark for Automatic Pathological Speech Assessment
von: Halpern, Bence Mark, et al.
Veröffentlicht: (2026)
von: Halpern, Bence Mark, et al.
Veröffentlicht: (2026)
MicroRemed: Benchmarking LLMs in Microservices Remediation
von: Zhang, Lingzhe, et al.
Veröffentlicht: (2025)
von: Zhang, Lingzhe, et al.
Veröffentlicht: (2025)
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
von: Ovcharov, Volodymyr
Veröffentlicht: (2026)
von: Ovcharov, Volodymyr
Veröffentlicht: (2026)
The Knesset Corpus: An Annotated Corpus of Hebrew Parliamentary Proceedings
von: Goldin, Gili, et al.
Veröffentlicht: (2024)
von: Goldin, Gili, et al.
Veröffentlicht: (2024)
Math Natural Language Inference: this should be easy!
von: de Paiva, Valeria, et al.
Veröffentlicht: (2025)
von: de Paiva, Valeria, et al.
Veröffentlicht: (2025)
Fast Quiet-STaR: Thinking Without Thought Tokens
von: Huang, Wei, et al.
Veröffentlicht: (2025)
von: Huang, Wei, et al.
Veröffentlicht: (2025)
Pitfalls in Evaluating Interpretability Agents
von: Haklay, Tal, et al.
Veröffentlicht: (2026)
von: Haklay, Tal, et al.
Veröffentlicht: (2026)
d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models
von: Pan, Leyi, et al.
Veröffentlicht: (2025)
von: Pan, Leyi, et al.
Veröffentlicht: (2025)
Towards Effective and Efficient Continual Pre-training of Large Language Models
von: Chen, Jie, et al.
Veröffentlicht: (2024)
von: Chen, Jie, et al.
Veröffentlicht: (2024)
The Unlikely Duel: Evaluating Creative Writing in LLMs through a Unique Scenario
von: Gómez-Rodríguez, Carlos, et al.
Veröffentlicht: (2024)
von: Gómez-Rodríguez, Carlos, et al.
Veröffentlicht: (2024)
Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation
von: Liu, Aiwei, et al.
Veröffentlicht: (2024)
von: Liu, Aiwei, et al.
Veröffentlicht: (2024)
SentiCSE: A Sentiment-aware Contrastive Sentence Embedding Framework with Sentiment-guided Textual Similarity
von: Kim, Jaemin, et al.
Veröffentlicht: (2024)
von: Kim, Jaemin, et al.
Veröffentlicht: (2024)
Distractor Injection Attacks on Large Reasoning Models: Characterization and Defense
von: Zhang, Zhehao, et al.
Veröffentlicht: (2025)
von: Zhang, Zhehao, et al.
Veröffentlicht: (2025)
Exploiting Pre-trained Encoder-Decoder Transformers for Sequence-to-Sequence Constituent Parsing
von: Fernández-González, Daniel, et al.
Veröffentlicht: (2026)
von: Fernández-González, Daniel, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL
von: Tu, Songjun, et al.
Veröffentlicht: (2025) -
Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
von: Tu, Songjun, et al.
Veröffentlicht: (2025) -
Examining Linguistic Shifts in Academic Writing Before and After the Launch of ChatGPT: A Study on Preprint Papers
von: Bao, Tong, et al.
Veröffentlicht: (2025) -
Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
von: Pan, Leyi, et al.
Veröffentlicht: (2025) -
Emotional Sequential Influence Modeling on False Information
von: Naskar, Debashis, et al.
Veröffentlicht: (2024)