Stop Automating Peer Review Without Rigorous Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Baumann, Joachim, Pei, Jiaxin, Koyejo, Sanmi, Hovy, Dirk |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SWE-chat: Coding Agent Interactions From Real Users in the Wild
di: Baumann, Joachim, et al.
Pubblicazione: (2026)
di: Baumann, Joachim, et al.
Pubblicazione: (2026)
The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
di: Wang, Angelina, et al.
Pubblicazione: (2025)
di: Wang, Angelina, et al.
Pubblicazione: (2025)
SCENEBench: An Audio Understanding Benchmark Grounded in Assistive and Industrial Use Cases
di: Iyer, Laya, et al.
Pubblicazione: (2026)
di: Iyer, Laya, et al.
Pubblicazione: (2026)
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
di: Röttger, Paul, et al.
Pubblicazione: (2024)
di: Röttger, Paul, et al.
Pubblicazione: (2024)
A Framework for Objective-Driven Dynamical Stochastic Fields
di: Zhang, Yibo Jacky, et al.
Pubblicazione: (2025)
di: Zhang, Yibo Jacky, et al.
Pubblicazione: (2025)
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
di: Hu, Tiancheng, et al.
Pubblicazione: (2025)
di: Hu, Tiancheng, et al.
Pubblicazione: (2025)
HiFA: High-fidelity Text-to-3D Generation with Advanced Diffusion Guidance
di: Zhu, Junzhe, et al.
Pubblicazione: (2023)
di: Zhu, Junzhe, et al.
Pubblicazione: (2023)
Are Domain Generalization Benchmarks with Accuracy on the Line Misspecified?
di: Salaudeen, Olawale, et al.
Pubblicazione: (2025)
di: Salaudeen, Olawale, et al.
Pubblicazione: (2025)
Reliable and Efficient Amortized Model-based Evaluation
di: Truong, Sang, et al.
Pubblicazione: (2025)
di: Truong, Sang, et al.
Pubblicazione: (2025)
Paper Copilot: Tracking the Evolution of Peer Review in AI Conferences
di: Yang, Jing, et al.
Pubblicazione: (2025)
di: Yang, Jing, et al.
Pubblicazione: (2025)
SycEval: Evaluating LLM Sycophancy
di: Fanous, Aaron, et al.
Pubblicazione: (2025)
di: Fanous, Aaron, et al.
Pubblicazione: (2025)
Do Prompts Reshape Representations? An Empirical Study of Prompting Effects on Embeddings
di: Gonzalez-Gutierrez, Cesar, et al.
Pubblicazione: (2025)
di: Gonzalez-Gutierrez, Cesar, et al.
Pubblicazione: (2025)
Why Do Safety Guardrails Degrade Across Languages?
di: Zhang, Max, et al.
Pubblicazione: (2026)
di: Zhang, Max, et al.
Pubblicazione: (2026)
Logits are All We Need to Adapt Closed Models
di: Hiranandani, Gaurush, et al.
Pubblicazione: (2025)
di: Hiranandani, Gaurush, et al.
Pubblicazione: (2025)
Label Noise Robustness for Domain-Agnostic Fair Corrections via Nearest Neighbors Label Spreading
di: Stromberg, Nathan, et al.
Pubblicazione: (2024)
di: Stromberg, Nathan, et al.
Pubblicazione: (2024)
Interactive Multi-Objective Probabilistic Preference Learning with Soft and Hard Bounds
di: Chen, Edward, et al.
Pubblicazione: (2025)
di: Chen, Edward, et al.
Pubblicazione: (2025)
Extracting books from production language models
di: Ahmed, Ahmed, et al.
Pubblicazione: (2026)
di: Ahmed, Ahmed, et al.
Pubblicazione: (2026)
Optimization and Generalization Guarantees for Weight Normalization
di: Cisneros-Velarde, Pedro, et al.
Pubblicazione: (2024)
di: Cisneros-Velarde, Pedro, et al.
Pubblicazione: (2024)
Position: Model Collapse Does Not Mean What You Think
di: Schaeffer, Rylan, et al.
Pubblicazione: (2025)
di: Schaeffer, Rylan, et al.
Pubblicazione: (2025)
Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed
di: Gupta, Isha, et al.
Pubblicazione: (2025)
di: Gupta, Isha, et al.
Pubblicazione: (2025)
The Optimization Paradox in Clinical AI Multi-Agent Systems
di: Bedi, Suhana, et al.
Pubblicazione: (2025)
di: Bedi, Suhana, et al.
Pubblicazione: (2025)
Defend: Automated Rebuttals for Peer Review with Minimal Author Guidance
di: Khatri, Jyotsana, et al.
Pubblicazione: (2026)
di: Khatri, Jyotsana, et al.
Pubblicazione: (2026)
Quantifying Variance in Evaluation Benchmarks
di: Madaan, Lovish, et al.
Pubblicazione: (2024)
di: Madaan, Lovish, et al.
Pubblicazione: (2024)
Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation
di: Baumann, Joachim, et al.
Pubblicazione: (2025)
di: Baumann, Joachim, et al.
Pubblicazione: (2025)
The Sound of Syntax: Finetuning and Comprehensive Evaluation of Language Models for Speech Pathology
di: Patel, Fagun, et al.
Pubblicazione: (2025)
di: Patel, Fagun, et al.
Pubblicazione: (2025)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
di: Tang, Zeyu, et al.
Pubblicazione: (2026)
di: Tang, Zeyu, et al.
Pubblicazione: (2026)
Lottery Ticket Adaptation: Mitigating Destructive Interference in LLMs
di: Panda, Ashwinee, et al.
Pubblicazione: (2024)
di: Panda, Ashwinee, et al.
Pubblicazione: (2024)
Decision from Suboptimal Classifiers: Excess Risk Pre- and Post-Calibration
di: Perez-Lebel, Alexandre, et al.
Pubblicazione: (2025)
di: Perez-Lebel, Alexandre, et al.
Pubblicazione: (2025)
On the Societal Impact of Machine Learning
di: Baumann, Joachim
Pubblicazione: (2025)
di: Baumann, Joachim
Pubblicazione: (2025)
On Fairness of Low-Rank Adaptation of Large Models
di: Ding, Zhoujie, et al.
Pubblicazione: (2024)
di: Ding, Zhoujie, et al.
Pubblicazione: (2024)
Stop Comparing LLM Agents Without Disclosing the Harness
di: Zhang, Yunbei, et al.
Pubblicazione: (2026)
di: Zhang, Yunbei, et al.
Pubblicazione: (2026)
The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models
di: Russo, Giuseppe, et al.
Pubblicazione: (2025)
di: Russo, Giuseppe, et al.
Pubblicazione: (2025)
Impoverished Language Technology: The Lack of (Social) Class in NLP
di: Curry, Amanda Cercas, et al.
Pubblicazione: (2024)
di: Curry, Amanda Cercas, et al.
Pubblicazione: (2024)
HEART: A Unified Benchmark for Assessing Humans and LLMs in Emotional Support Dialogue
di: Iyer, Laya, et al.
Pubblicazione: (2026)
di: Iyer, Laya, et al.
Pubblicazione: (2026)
Latent Adversarial Regularization for Offline Preference Optimization
di: Jiang, Enyi, et al.
Pubblicazione: (2026)
di: Jiang, Enyi, et al.
Pubblicazione: (2026)
PeerArg: Argumentative Peer Review with LLMs
di: Sukpanichnant, Purin, et al.
Pubblicazione: (2024)
di: Sukpanichnant, Purin, et al.
Pubblicazione: (2024)
ScholarPeer: A Context-Aware Multi-Agent Framework for Automated Peer Review
di: Goyal, Palash, et al.
Pubblicazione: (2026)
di: Goyal, Palash, et al.
Pubblicazione: (2026)
A Sentiment Consolidation Framework for Meta-Review Generation
di: Li, Miao, et al.
Pubblicazione: (2024)
di: Li, Miao, et al.
Pubblicazione: (2024)
Can AI Be a Good Peer Reviewer? A Survey of Peer Review Process, Evaluation, and the Future
di: Wu, Sihong, et al.
Pubblicazione: (2026)
di: Wu, Sihong, et al.
Pubblicazione: (2026)
REMOR: Automated Peer Review Generation with LLM Reasoning and Multi-Objective Reinforcement Learning
di: Taechoyotin, Pawin, et al.
Pubblicazione: (2025)
di: Taechoyotin, Pawin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SWE-chat: Coding Agent Interactions From Real Users in the Wild
di: Baumann, Joachim, et al.
Pubblicazione: (2026) -
The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
di: Wang, Angelina, et al.
Pubblicazione: (2025) -
SCENEBench: An Audio Understanding Benchmark Grounded in Assistive and Industrial Use Cases
di: Iyer, Laya, et al.
Pubblicazione: (2026) -
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
di: Röttger, Paul, et al.
Pubblicazione: (2024) -
A Framework for Objective-Driven Dynamical Stochastic Fields
di: Zhang, Yibo Jacky, et al.
Pubblicazione: (2025)