Benchmark on Peer Review Toxic Detection: A Challenging Task with a New Dataset
Fuente:
arXiv
Salvato in:
| Autori principali: | Luo, Man, Peterson, Bradley, Gan, Rafael, Ramalingame, Hari, Gangrade, Navya, Dimarogona, Ariadne, Banerjee, Imon, Howard, Phillip |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer Review
di: Yu, Sungduk, et al.
Pubblicazione: (2025)
di: Yu, Sungduk, et al.
Pubblicazione: (2025)
Is Your Paper Being Reviewed by an LLM? Investigating AI Text Detectability in Peer Review
di: Yu, Sungduk, et al.
Pubblicazione: (2024)
di: Yu, Sungduk, et al.
Pubblicazione: (2024)
Assessing Empathy in Large Language Models with Real-World Physician-Patient Interactions
di: Luo, Man, et al.
Pubblicazione: (2024)
di: Luo, Man, et al.
Pubblicazione: (2024)
CLT and Edgeworth Expansion for m-out-of-n Bootstrap Estimators of The Studentized Median
di: Banerjee, Imon, et al.
Pubblicazione: (2025)
di: Banerjee, Imon, et al.
Pubblicazione: (2025)
A Review on Large Language Models for Visual Analytics
di: Agarwal, Navya Sonal, et al.
Pubblicazione: (2025)
di: Agarwal, Navya Sonal, et al.
Pubblicazione: (2025)
Multi-Perspective LLM Annotations for Valid Analyses in Subjective Tasks
di: Mehrotra, Navya, et al.
Pubblicazione: (2026)
di: Mehrotra, Navya, et al.
Pubblicazione: (2026)
Transformer-Based Temporal Information Extraction and Application: A Review
di: Su, Xin, et al.
Pubblicazione: (2025)
di: Su, Xin, et al.
Pubblicazione: (2025)
Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings
di: Yang, Shujian, et al.
Pubblicazione: (2025)
di: Yang, Shujian, et al.
Pubblicazione: (2025)
Automated Diagnosis of Lower Limb Alignment Deformities Using Deep Learning and Joint Orientation Analysis
di: Aarti Goswami, et al.
Pubblicazione: (2026)
di: Aarti Goswami, et al.
Pubblicazione: (2026)
Training-Free Mitigation of Language Reasoning Degradation After Multimodal Instruction Tuning
di: Ratzlaff, Neale, et al.
Pubblicazione: (2024)
di: Ratzlaff, Neale, et al.
Pubblicazione: (2024)
LazyReview A Dataset for Uncovering Lazy Thinking in NLP Peer Reviews
di: Purkayastha, Sukannya, et al.
Pubblicazione: (2025)
di: Purkayastha, Sukannya, et al.
Pubblicazione: (2025)
SK-VQA: Synthetic Knowledge Generation at Scale for Training Context-Augmented Multimodal LLMs
di: Su, Xin, et al.
Pubblicazione: (2024)
di: Su, Xin, et al.
Pubblicazione: (2024)
Multilingual Open QA on the MIA Shared Task
di: Yarrabelly, Navya, et al.
Pubblicazione: (2025)
di: Yarrabelly, Navya, et al.
Pubblicazione: (2025)
PeerQA: A Scientific Question Answering Dataset from Peer Reviews
di: Baumgärtner, Tim, et al.
Pubblicazione: (2025)
di: Baumgärtner, Tim, et al.
Pubblicazione: (2025)
Position: Restructuring of Categories and Implementation of Guidelines Essential for VLM Adoption in Healthcare
di: Tariq, Amara, et al.
Pubblicazione: (2025)
di: Tariq, Amara, et al.
Pubblicazione: (2025)
PRAIB: Peer Review AI Benchmark of Behaviour of LLM-Assisted Reviewing
di: Żurawicki, Krzysztof, et al.
Pubblicazione: (2026)
di: Żurawicki, Krzysztof, et al.
Pubblicazione: (2026)
Unsupervised Hybrid framework for ANomaly Detection (HAND) -- applied to Screening Mammogram
di: Zhang, Zhemin, et al.
Pubblicazione: (2024)
di: Zhang, Zhemin, et al.
Pubblicazione: (2024)
PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers
di: Loc, Ngoc Phan Phuoc, et al.
Pubblicazione: (2026)
di: Loc, Ngoc Phan Phuoc, et al.
Pubblicazione: (2026)
MASSW: A New Dataset and Benchmark Tasks for AI-Assisted Scientific Workflows
di: Zhang, Xingjian, et al.
Pubblicazione: (2024)
di: Zhang, Xingjian, et al.
Pubblicazione: (2024)
Beyond Detection: Governing GenAI in Academic Peer Review as a Sociotechnical Challenge
di: Chakravorti, Tatiana, et al.
Pubblicazione: (2026)
di: Chakravorti, Tatiana, et al.
Pubblicazione: (2026)
AI and the Future of Academic Peer Review
di: Mann, Sebastian Porsdam, et al.
Pubblicazione: (2025)
di: Mann, Sebastian Porsdam, et al.
Pubblicazione: (2025)
Dataset Representativeness and Downstream Task Fairness
di: Borza, Victor, et al.
Pubblicazione: (2024)
di: Borza, Victor, et al.
Pubblicazione: (2024)
MMReview: A Multidisciplinary and Multimodal Benchmark for LLM-Based Peer Review Automation
di: Gao, Xian, et al.
Pubblicazione: (2025)
di: Gao, Xian, et al.
Pubblicazione: (2025)
Towards Automatic Power Battery Detection: New Challenge, Benchmark Dataset and Baseline
di: Zhao, Xiaoqi, et al.
Pubblicazione: (2023)
di: Zhao, Xiaoqi, et al.
Pubblicazione: (2023)
Looking Beyond the Known: Towards a Data Discovery Guided Open-World Object Detection
di: Majee, Anay, et al.
Pubblicazione: (2025)
di: Majee, Anay, et al.
Pubblicazione: (2025)
Cultural Awareness in Vision-Language Models: A Cross-Country Exploration
di: Madasu, Avinash, et al.
Pubblicazione: (2025)
di: Madasu, Avinash, et al.
Pubblicazione: (2025)
PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies
di: Selch, Lukas, et al.
Pubblicazione: (2025)
di: Selch, Lukas, et al.
Pubblicazione: (2025)
ToxicTone: A Mandarin Audio Dataset Annotated for Toxicity and Toxic Utterance Tonality
di: Luo, Yu-Xiang, et al.
Pubblicazione: (2025)
di: Luo, Yu-Xiang, et al.
Pubblicazione: (2025)
Learn2Reg 2024: New Benchmark Datasets Driving Progress on New Challenges
di: Hansen, Lasse, et al.
Pubblicazione: (2025)
di: Hansen, Lasse, et al.
Pubblicazione: (2025)
Multimodal Multihop Source Retrieval for Web Question Answering
di: Yarrabelly, Navya, et al.
Pubblicazione: (2025)
di: Yarrabelly, Navya, et al.
Pubblicazione: (2025)
Reviewing the Reviewer: Elevating Peer Review Quality through LLM-Guided Feedback
di: Purkayastha, Sukannya, et al.
Pubblicazione: (2026)
di: Purkayastha, Sukannya, et al.
Pubblicazione: (2026)
Bangla Sign Language Translation: Dataset Creation Challenges, Benchmarking and Prospects
di: Rubaiyeat, Husne Ara, et al.
Pubblicazione: (2025)
di: Rubaiyeat, Husne Ara, et al.
Pubblicazione: (2025)
Multi-Dialect Vietnamese: Task, Dataset, Baseline Models and Challenges
di: Van Dinh, Nguyen, et al.
Pubblicazione: (2024)
di: Van Dinh, Nguyen, et al.
Pubblicazione: (2024)
Identifying Aspects in Peer Reviews
di: Lu, Sheng, et al.
Pubblicazione: (2025)
di: Lu, Sheng, et al.
Pubblicazione: (2025)
Multi-Analyte, Swab-based Automated Wound Monitor with AI
di: Sikha, Madhu Babu, et al.
Pubblicazione: (2025)
di: Sikha, Madhu Babu, et al.
Pubblicazione: (2025)
Beyond Toxic: Toxicity Detection Datasets are Not Enough for Brand Safety
di: Korotkova, Elizaveta, et al.
Pubblicazione: (2023)
di: Korotkova, Elizaveta, et al.
Pubblicazione: (2023)
Towards Comprehensive Argument Analysis in Education: Dataset, Tasks, and Method
di: Ren, Yupei, et al.
Pubblicazione: (2025)
di: Ren, Yupei, et al.
Pubblicazione: (2025)
Artificial Intelligence in Mental Health and Well-Being: Evolution, Current Applications, Future Challenges, and Emerging Evidence
di: Pandey, Hari Mohan
Pubblicazione: (2024)
di: Pandey, Hari Mohan
Pubblicazione: (2024)
PeerPrism: Peer Evaluation Expertise vs Review-writing AI
di: Sadeghian, Soroush, et al.
Pubblicazione: (2026)
di: Sadeghian, Soroush, et al.
Pubblicazione: (2026)
Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study
di: Mayer, Leon, et al.
Pubblicazione: (2025)
di: Mayer, Leon, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer Review
di: Yu, Sungduk, et al.
Pubblicazione: (2025) -
Is Your Paper Being Reviewed by an LLM? Investigating AI Text Detectability in Peer Review
di: Yu, Sungduk, et al.
Pubblicazione: (2024) -
Assessing Empathy in Large Language Models with Real-World Physician-Patient Interactions
di: Luo, Man, et al.
Pubblicazione: (2024) -
CLT and Edgeworth Expansion for m-out-of-n Bootstrap Estimators of The Studentized Median
di: Banerjee, Imon, et al.
Pubblicazione: (2025) -
A Review on Large Language Models for Visual Analytics
di: Agarwal, Navya Sonal, et al.
Pubblicazione: (2025)