When AI reviews science: Can we trust the referee?
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Jialiang, Liu, Yuchen, Xu, Hang, Hu, Kaichun, Di, Shimin, Ni, Wangze, Yue, Linan, Zhang, Min-Ling, Ren, Kui, Chen, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents
by: Chen, Weiyi, et al.
Published: (2026)
by: Chen, Weiyi, et al.
Published: (2026)
Guided by Trajectories: Repairing and Rewarding Tool-Use Trajectories for Tool-Integrated Reasoning
by: Gong, Siyu, et al.
Published: (2026)
by: Gong, Siyu, et al.
Published: (2026)
RxnNano:Training Compact LLMs for Chemical Reaction and Retrosynthesis Prediction via Hierarchical Curriculum Learning
by: Li, Ran, et al.
Published: (2026)
by: Li, Ran, et al.
Published: (2026)
Code2MCP: Transforming Code Repositories into MCP Services
by: Ouyang, Chaoqian, et al.
Published: (2025)
by: Ouyang, Chaoqian, et al.
Published: (2025)
Can we trust the evaluation on ChatGPT?
by: Aiyappa, Rachith, et al.
Published: (2023)
by: Aiyappa, Rachith, et al.
Published: (2023)
ContextCache: Context-Aware Semantic Cache for Multi-Turn Queries in Large Language Models
by: Yan, Jianxin, et al.
Published: (2025)
by: Yan, Jianxin, et al.
Published: (2025)
ERASER: Machine Unlearning in MLaaS via an Inference Serving-Aware Approach
by: Hu, Yuke, et al.
Published: (2023)
by: Hu, Yuke, et al.
Published: (2023)
Editorial: Why should we care about preparing referee reports?
by: Matheus Albergaria
Published: (2021)
by: Matheus Albergaria
Published: (2021)
Can we disrupt the momentum of the AI colonization of science education?
by: Lucy Avraamidou
Published: (2024)
by: Lucy Avraamidou
Published: (2024)
Can we trust LLM Self-Explanations for Entity Resolution?
by: Teofili, Tommaso, et al.
Published: (2026)
by: Teofili, Tommaso, et al.
Published: (2026)
RAC: Relation-Aware Cache Replacement for Large Language Models
by: Wu, Yuchong, et al.
Published: (2026)
by: Wu, Yuchong, et al.
Published: (2026)
Learning to Compose for Cross-domain Agentic Workflow Generation
by: Wang, Jialiang, et al.
Published: (2026)
by: Wang, Jialiang, et al.
Published: (2026)
FactReview: Evidence-Grounded Peer Review with Execution-Based Claim Verification
by: Yue, Ling, et al.
Published: (2026)
by: Yue, Ling, et al.
Published: (2026)
SJP referees
Published: (2025)
Published: (2025)
When I say … trust in AI
by: Levent Çetinkaya
Published: (2025)
by: Levent Çetinkaya
Published: (2025)
Can we trust AI for acne advice: A double‐blind performance comparison with NICE guidelines
by: Xufei Luo, et al.
Published: (2025)
by: Xufei Luo, et al.
Published: (2025)
Class-aware and Augmentation-free Contrastive Learning from Label Proportion
by: Wang, Jialiang, et al.
Published: (2024)
by: Wang, Jialiang, et al.
Published: (2024)
Bridging Efficiency and Transparency: Explainable CoT Compression in Multimodal Large Reasoning Models
by: Wang, Yizhi, et al.
Published: (2026)
by: Wang, Yizhi, et al.
Published: (2026)
Training Multimodal Large Reasoning Models Needs Better Thoughts: A Three-Stage Framework for Long Chain-of-Thought Synthesis and Selection
by: Wang, Yizhi, et al.
Published: (2025)
by: Wang, Yizhi, et al.
Published: (2025)
SRBench: A Comprehensive Benchmark for Sequential Recommendation with Large Language Models
by: Li, Jianhong, et al.
Published: (2026)
by: Li, Jianhong, et al.
Published: (2026)
ExplainitAI: When do we trust artificial intelligence? The influence of content and explainability in a cross-cultural comparison
by: Kang, Sora, et al.
Published: (2025)
by: Kang, Sora, et al.
Published: (2025)
QCFuse: Query-Centric Cache Fusion for Efficient RAG Inference
by: Yan, Jianxin, et al.
Published: (2026)
by: Yan, Jianxin, et al.
Published: (2026)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
by: Yang, Langqi, et al.
Published: (2025)
by: Yang, Langqi, et al.
Published: (2025)
Can we automatize scientific discovery in the cognitive sciences?
by: Jagadish, Akshay K., et al.
Published: (2026)
by: Jagadish, Akshay K., et al.
Published: (2026)
Misfortunes of a mathematicians' trio using Computer Algebra Systems: Can we trust?
by: Durán, Antonio J., et al.
Published: (2013)
by: Durán, Antonio J., et al.
Published: (2013)
DualBreach: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization
by: Huang, Xinzhe, et al.
Published: (2025)
by: Huang, Xinzhe, et al.
Published: (2025)
Guest editors and referees 2023
Published: (2024)
Published: (2024)
TCT referee recognition 2023
Published: (2024)
Published: (2024)
When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation
by: Min, Anna, et al.
Published: (2025)
by: Min, Anna, et al.
Published: (2025)
Why do referees end their careers and which factors determine the duration of a referee’s career?
by: Christian Rullang
Published: (2017)
by: Christian Rullang
Published: (2017)
Can we coevolve with AI?
by: Joshua E Lerner, et al.
Published: (2024)
by: Joshua E Lerner, et al.
Published: (2024)
Don't Overthink It: A Survey of Efficient R1-style Large Reasoning Models
by: Yue, Linan, et al.
Published: (2025)
by: Yue, Linan, et al.
Published: (2025)
When can we trust untrusted monitoring? A safety case sketch across collusion strategies
by: Gardner-Challis, Nelson, et al.
Published: (2026)
by: Gardner-Challis, Nelson, et al.
Published: (2026)
When should we trust the annotation? Selective prediction for molecular structure retrieval from mass spectra
by: Jürgens, Mira, et al.
Published: (2026)
by: Jürgens, Mira, et al.
Published: (2026)
Scientific writing and referee professional training
by: George Argota Pérez
Published: (2023)
by: George Argota Pérez
Published: (2023)
Can we trust deep learning models diagnosis? The impact of domain shift in chest radiograph classification
by: Pooch, Eduardo H. P., et al.
Published: (2019)
by: Pooch, Eduardo H. P., et al.
Published: (2019)
Oral immunotherapy for Helicobacter pylori: Can it be trusted? A systematic review
by: Mohammad Hossein Peypar, et al.
Published: (2024)
by: Mohammad Hossein Peypar, et al.
Published: (2024)
SpatialJB: How Text Distribution Art Becomes the "Jailbreak Key" for LLM Guardrails
by: Mou, Zhiyi, et al.
Published: (2026)
by: Mou, Zhiyi, et al.
Published: (2026)
Proficient Graph Neural Network Design by Accumulating Knowledge on Large Language Models
by: Wang, Jialiang, et al.
Published: (2024)
by: Wang, Jialiang, et al.
Published: (2024)
Beyond Model Base Retrieval: Weaving Knowledge to Master Fine-grained Neural Network Design
by: Wang, Jialiang, et al.
Published: (2025)
by: Wang, Jialiang, et al.
Published: (2025)
Similar Items
-
TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents
by: Chen, Weiyi, et al.
Published: (2026) -
Guided by Trajectories: Repairing and Rewarding Tool-Use Trajectories for Tool-Integrated Reasoning
by: Gong, Siyu, et al.
Published: (2026) -
RxnNano:Training Compact LLMs for Chemical Reaction and Retrosynthesis Prediction via Hierarchical Curriculum Learning
by: Li, Ran, et al.
Published: (2026) -
Code2MCP: Transforming Code Repositories into MCP Services
by: Ouyang, Chaoqian, et al.
Published: (2025) -
Can we trust the evaluation on ChatGPT?
by: Aiyappa, Rachith, et al.
Published: (2023)