The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors
Fuente:
arXiv
Guardado en:
| Autores principales: | Sadallah, Abdelrahman, Baumgärtner, Tim, Gurevych, Iryna, Briscoe, Ted |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PeerQA: A Scientific Question Answering Dataset from Peer Reviews
por: Baumgärtner, Tim, et al.
Publicado: (2025)
por: Baumgärtner, Tim, et al.
Publicado: (2025)
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment
por: Baumgärtner, Tim, et al.
Publicado: (2026)
por: Baumgärtner, Tim, et al.
Publicado: (2026)
Commitment Checklist: Auditing Author Commitments in Peer Review
por: Chen, Chung-Chi, et al.
Publicado: (2026)
por: Chen, Chung-Chi, et al.
Publicado: (2026)
RIRAG: Regulatory Information Retrieval and Answer Generation
por: Gokhan, Tuba, et al.
Publicado: (2024)
por: Gokhan, Tuba, et al.
Publicado: (2024)
Are LLMs Good Cryptic Crossword Solvers?
por: Sadallah, Abdelrahman, et al.
Publicado: (2024)
por: Sadallah, Abdelrahman, et al.
Publicado: (2024)
Reviewing the Reviewer: Elevating Peer Review Quality through LLM-Guided Feedback
por: Purkayastha, Sukannya, et al.
Publicado: (2026)
por: Purkayastha, Sukannya, et al.
Publicado: (2026)
Bad Actor, Good Advisor: Exploring the Role of Large Language Models in Fake News Detection
por: Hu, Beizhe, et al.
Publicado: (2023)
por: Hu, Beizhe, et al.
Publicado: (2023)
What Makes Cryptic Crosswords Challenging for LLMs?
por: Sadallah, Abdelrahman, et al.
Publicado: (2024)
por: Sadallah, Abdelrahman, et al.
Publicado: (2024)
DeepReviewer 2.0: A Traceable Agentic System for Auditable Scientific Peer Review
por: Weng, Yixuan, et al.
Publicado: (2026)
por: Weng, Yixuan, et al.
Publicado: (2026)
The Good, the Bad, and the Ugly: The Role of AI Quality Disclosure in Lie Detection
por: Bhattacharya, Haimanti, et al.
Publicado: (2024)
por: Bhattacharya, Haimanti, et al.
Publicado: (2024)
SPARE: Single-Pass Annotation with Reference-Guided Evaluation for Automatic Process Supervision and Reward Modelling
por: Rizvi, Md Imbesat Hassan, et al.
Publicado: (2025)
por: Rizvi, Md Imbesat Hassan, et al.
Publicado: (2025)
AgentPeerTalk: Empowering Students through Agentic-AI-Driven Discernment of Bullying and Joking in Peer Interactions in Schools
por: Paul, Aditya, et al.
Publicado: (2024)
por: Paul, Aditya, et al.
Publicado: (2024)
Towards Privacy-aware Mental Health AI Models: Advances, Challenges, and Opportunities
por: Mandal, Aishik, et al.
Publicado: (2025)
por: Mandal, Aishik, et al.
Publicado: (2025)
MAGneT: Coordinated Multi-Agent Generation of Synthetic Multi-Turn Mental Health Counseling Sessions
por: Mandal, Aishik, et al.
Publicado: (2025)
por: Mandal, Aishik, et al.
Publicado: (2025)
Preemptive Detection and Correction of Misaligned Actions in LLM Agents
por: Fang, Haishuo, et al.
Publicado: (2024)
por: Fang, Haishuo, et al.
Publicado: (2024)
Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review
por: Lin, Zhicheng
Publicado: (2025)
por: Lin, Zhicheng
Publicado: (2025)
IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code Generators
por: Paul, Indraneil, et al.
Publicado: (2024)
por: Paul, Indraneil, et al.
Publicado: (2024)
Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
por: Kim, Jaeho, et al.
Publicado: (2025)
por: Kim, Jaeho, et al.
Publicado: (2025)
Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable
por: Saha, Rounak, et al.
Publicado: (2026)
por: Saha, Rounak, et al.
Publicado: (2026)
Opacity as Authority: Arbitrariness and the Preclusion of Contestation
por: Kayembe, Naomi Omeonga wa
Publicado: (2025)
por: Kayembe, Naomi Omeonga wa
Publicado: (2025)
Are Large Language Models Good Essay Graders?
por: Kundu, Anindita, et al.
Publicado: (2024)
por: Kundu, Anindita, et al.
Publicado: (2024)
Measuring What Matters -- or What's Convenient?: Robustness of LLM-Based Scoring Systems to Construct-Irrelevant Factors
por: Walsh, Cole, et al.
Publicado: (2026)
por: Walsh, Cole, et al.
Publicado: (2026)
Identifying Aspects in Peer Reviews
por: Lu, Sheng, et al.
Publicado: (2025)
por: Lu, Sheng, et al.
Publicado: (2025)
A Comprehensive Review of Datasets for Clinical Mental Health AI Systems
por: Mandal, Aishik, et al.
Publicado: (2025)
por: Mandal, Aishik, et al.
Publicado: (2025)
Three Models of RLHF Annotation: Extension, Evidence, and Authority
por: Coyne, Steve
Publicado: (2026)
por: Coyne, Steve
Publicado: (2026)
Can AI Be a Good Peer Reviewer? A Survey of Peer Review Process, Evaluation, and the Future
por: Wu, Sihong, et al.
Publicado: (2026)
por: Wu, Sihong, et al.
Publicado: (2026)
Personal Care Utility (PCU): Building the Health Infrastructure for Everyday Insight and Guidance
por: Abbasian, Mahyar, et al.
Publicado: (2025)
por: Abbasian, Mahyar, et al.
Publicado: (2025)
Authorship Without Writing: Large Language Models and the Senior Author Analogy
por: Hurshman, Clint, et al.
Publicado: (2025)
por: Hurshman, Clint, et al.
Publicado: (2025)
Multimodal Large Language Models to Support Real-World Fact-Checking
por: Geng, Jiahui, et al.
Publicado: (2024)
por: Geng, Jiahui, et al.
Publicado: (2024)
Sensitivity, Performance, Robustness: Deconstructing the Effect of Sociodemographic Prompting
por: Beck, Tilman, et al.
Publicado: (2023)
por: Beck, Tilman, et al.
Publicado: (2023)
$\texttt{Droid}$: A Resource Suite for AI-Generated Code Detection
por: Orel, Daniil, et al.
Publicado: (2025)
por: Orel, Daniil, et al.
Publicado: (2025)
Automatic Authorities: Power and AI
por: Lazar, Seth
Publicado: (2024)
por: Lazar, Seth
Publicado: (2024)
Artificial Intelligence Bias on English Language Learners in Automatic Scoring
por: Guo, Shuchen, et al.
Publicado: (2025)
por: Guo, Shuchen, et al.
Publicado: (2025)
Factual Self-Awareness in Language Models: Representation, Robustness, and Scaling
por: Tamoyan, Hovhannes, et al.
Publicado: (2025)
por: Tamoyan, Hovhannes, et al.
Publicado: (2025)
Using GPT-4 to Augment Unbalanced Data for Automatic Scoring
por: Fang, Luyang, et al.
Publicado: (2023)
por: Fang, Luyang, et al.
Publicado: (2023)
Statutory Construction and Interpretation for Artificial Intelligence
por: He, Luxi, et al.
Publicado: (2025)
por: He, Luxi, et al.
Publicado: (2025)
GoodPoint: Learning Constructive Scientific Paper Feedback from Author Responses
por: Mun, Jimin, et al.
Publicado: (2026)
por: Mun, Jimin, et al.
Publicado: (2026)
"I understand why I got this grade": Automatic Short Answer Grading with Feedback
por: Aggarwal, Dishank, et al.
Publicado: (2024)
por: Aggarwal, Dishank, et al.
Publicado: (2024)
Why Don't Prompt-Based Fairness Metrics Correlate?
por: Zayed, Abdelrahman, et al.
Publicado: (2024)
por: Zayed, Abdelrahman, et al.
Publicado: (2024)
Should We Attend More or Less? Modulating Attention for Fairness
por: Zayed, Abdelrahman, et al.
Publicado: (2023)
por: Zayed, Abdelrahman, et al.
Publicado: (2023)
Ejemplares similares
-
PeerQA: A Scientific Question Answering Dataset from Peer Reviews
por: Baumgärtner, Tim, et al.
Publicado: (2025) -
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment
por: Baumgärtner, Tim, et al.
Publicado: (2026) -
Commitment Checklist: Auditing Author Commitments in Peer Review
por: Chen, Chung-Chi, et al.
Publicado: (2026) -
RIRAG: Regulatory Information Retrieval and Answer Generation
por: Gokhan, Tuba, et al.
Publicado: (2024) -
Are LLMs Good Cryptic Crossword Solvers?
por: Sadallah, Abdelrahman, et al.
Publicado: (2024)