From Replication to Redesign: Exploring Pairwise Comparisons for LLM-Based Peer Review
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yaohui, Zhang, Haijing, Ji, Wenlong, Hua, Tianyu, Haber, Nick, Cao, Hancheng, Liang, Weixin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Widespread Adoption of Large Language Model-Assisted Writing Across Society
by: Liang, Weixin, et al.
Published: (2025)
by: Liang, Weixin, et al.
Published: (2025)
ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code
by: Hua, Tianyu, et al.
Published: (2025)
by: Hua, Tianyu, et al.
Published: (2025)
Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
by: Liang, Weixin, et al.
Published: (2024)
by: Liang, Weixin, et al.
Published: (2024)
AgentReview: Exploring Peer Review Dynamics with LLM Agents
by: Jin, Yiqiao, et al.
Published: (2024)
by: Jin, Yiqiao, et al.
Published: (2024)
MMReview: A Multidisciplinary and Multimodal Benchmark for LLM-Based Peer Review Automation
by: Gao, Xian, et al.
Published: (2025)
by: Gao, Xian, et al.
Published: (2025)
The Comparative Trap: Pairwise Comparisons Amplifies Biased Preferences of LLM Evaluators
by: Jeong, Hawon, et al.
Published: (2024)
by: Jeong, Hawon, et al.
Published: (2024)
ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
by: Liang, Yesheng, et al.
Published: (2025)
by: Liang, Yesheng, et al.
Published: (2025)
Adaptive Self-improvement LLM Agentic System for ML Library Development
by: Zhang, Genghan, et al.
Published: (2025)
by: Zhang, Genghan, et al.
Published: (2025)
Efficient LLM Comparative Assessment: a Product of Experts Framework for Pairwise Comparisons
by: Liusie, Adian, et al.
Published: (2024)
by: Liusie, Adian, et al.
Published: (2024)
Mapping the Increasing Use of LLMs in Scientific Papers
by: Liang, Weixin, et al.
Published: (2024)
by: Liang, Weixin, et al.
Published: (2024)
Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild
by: Mysore, Sheshera, et al.
Published: (2025)
by: Mysore, Sheshera, et al.
Published: (2025)
ReviewGuard: Enhancing Deficient Peer Review Detection via LLM-Driven Data Augmentation
by: Zhang, Haoxuan, et al.
Published: (2025)
by: Zhang, Haoxuan, et al.
Published: (2025)
Reviewing the Reviewer: Elevating Peer Review Quality through LLM-Guided Feedback
by: Purkayastha, Sukannya, et al.
Published: (2026)
by: Purkayastha, Sukannya, et al.
Published: (2026)
LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models
by: Liusie, Adian, et al.
Published: (2023)
by: Liusie, Adian, et al.
Published: (2023)
PairEval: Open-domain Dialogue Evaluation with Pairwise Comparison
by: Park, ChaeHun, et al.
Published: (2024)
by: Park, ChaeHun, et al.
Published: (2024)
TreeReview: A Dynamic Tree of Questions Framework for Deep and Efficient LLM-based Scientific Peer Review
by: Chang, Yuan, et al.
Published: (2025)
by: Chang, Yuan, et al.
Published: (2025)
Computational Approaches to Understanding Large Language Model Impact on Writing and Information Ecosystems
by: Liang, Weixin
Published: (2025)
by: Liang, Weixin
Published: (2025)
PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers
by: Loc, Ngoc Phan Phuoc, et al.
Published: (2026)
by: Loc, Ngoc Phan Phuoc, et al.
Published: (2026)
PRAIB: Peer Review AI Benchmark of Behaviour of LLM-Assisted Reviewing
by: Żurawicki, Krzysztof, et al.
Published: (2026)
by: Żurawicki, Krzysztof, et al.
Published: (2026)
Advanced Multimodal Deep Learning Architecture for Image-Text Matching
by: Wang, Jinyin, et al.
Published: (2024)
by: Wang, Jinyin, et al.
Published: (2024)
LLM Optimization Unlocks Real-Time Pairwise Reranking
by: Wu, Jingyu, et al.
Published: (2025)
by: Wu, Jingyu, et al.
Published: (2025)
Online Rubrics Elicitation from Pairwise Comparisons
by: Rezaei, MohammadHossein, et al.
Published: (2025)
by: Rezaei, MohammadHossein, et al.
Published: (2025)
Are Optimal Algorithms Still Optimal? Rethinking Sorting in LLM-Based Pairwise Ranking with Batching and Caching
by: Wisznia, Juan, et al.
Published: (2025)
by: Wisznia, Juan, et al.
Published: (2025)
Steel-LLM:From Scratch to Open Source -- A Personal Journey in Building a Chinese-Centric LLM
by: Gu, Qingshui, et al.
Published: (2025)
by: Gu, Qingshui, et al.
Published: (2025)
Can AI Be a Good Peer Reviewer? A Survey of Peer Review Process, Evaluation, and the Future
by: Wu, Sihong, et al.
Published: (2026)
by: Wu, Sihong, et al.
Published: (2026)
Large Language Models Penetration in Scholarly Writing and Peer Review
by: Zhou, Li, et al.
Published: (2025)
by: Zhou, Li, et al.
Published: (2025)
Identifying Aspects in Peer Reviews
by: Lu, Sheng, et al.
Published: (2025)
by: Lu, Sheng, et al.
Published: (2025)
Am I More Pointwise or Pairwise? Revealing Position Bias in Rubric-Based LLM-as-a-Judge
by: Xu, Yuzheng, et al.
Published: (2026)
by: Xu, Yuzheng, et al.
Published: (2026)
PeerPrism: Peer Evaluation Expertise vs Review-writing AI
by: Sadeghian, Soroush, et al.
Published: (2026)
by: Sadeghian, Soroush, et al.
Published: (2026)
Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer Review
by: Yu, Sungduk, et al.
Published: (2025)
by: Yu, Sungduk, et al.
Published: (2025)
Is Your Paper Being Reviewed by an LLM? Investigating AI Text Detectability in Peer Review
by: Yu, Sungduk, et al.
Published: (2024)
by: Yu, Sungduk, et al.
Published: (2024)
Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons
by: Sandan, Isik Baran, et al.
Published: (2025)
by: Sandan, Isik Baran, et al.
Published: (2025)
Estimating the Error of Large Language Models at Pairwise Text Comparison
by: Li, Tianyi
Published: (2025)
by: Li, Tianyi
Published: (2025)
ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences
by: Nguyen, Bang, et al.
Published: (2026)
by: Nguyen, Bang, et al.
Published: (2026)
Exploring Key Point Analysis with Pairwise Generation and Graph Partitioning
by: Li, Xiao, et al.
Published: (2024)
by: Li, Xiao, et al.
Published: (2024)
Disparities in Peer Review Tone and the Role of Reviewer Anonymity
by: Sahakyan, Maria, et al.
Published: (2025)
by: Sahakyan, Maria, et al.
Published: (2025)
LLM Review: Enhancing Creative Writing via Blind Peer Review Feedback
by: Li, Weiyue, et al.
Published: (2026)
by: Li, Weiyue, et al.
Published: (2026)
Concept-Guided Chain-of-Thought Prompting for Pairwise Comparison Scoring of Texts with Large Language Models
by: Wu, Patrick Y., et al.
Published: (2023)
by: Wu, Patrick Y., et al.
Published: (2023)
O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning
by: Huang, Zhongzhen, et al.
Published: (2025)
by: Huang, Zhongzhen, et al.
Published: (2025)
Generating Chain-of-Thoughts with a Pairwise-Comparison Approach to Searching for the Most Promising Intermediate Thought
by: Zhang, Zhen-Yu, et al.
Published: (2024)
by: Zhang, Zhen-Yu, et al.
Published: (2024)
Similar Items
-
The Widespread Adoption of Large Language Model-Assisted Writing Across Society
by: Liang, Weixin, et al.
Published: (2025) -
ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code
by: Hua, Tianyu, et al.
Published: (2025) -
Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
by: Liang, Weixin, et al.
Published: (2024) -
AgentReview: Exploring Peer Review Dynamics with LLM Agents
by: Jin, Yiqiao, et al.
Published: (2024) -
MMReview: A Multidisciplinary and Multimodal Benchmark for LLM-Based Peer Review Automation
by: Gao, Xian, et al.
Published: (2025)