ReviewEval: An Evaluation Framework for AI-Generated Reviews
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Garg, Madhav Krishan, Prasad, Tejash, Singhal, Tanmay, Kirtani, Chhavi, Mandal, Murari, Kumar, Dhruv |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CricBench: A Multilingual Benchmark for Evaluating LLMs in Cricket Analytics
von: Agarwal, Parth, et al.
Veröffentlicht: (2025)
von: Agarwal, Parth, et al.
Veröffentlicht: (2025)
LLM-as-a-Judge for Time Series Explanations
von: Sivalingam, Preetham, et al.
Veröffentlicht: (2026)
von: Sivalingam, Preetham, et al.
Veröffentlicht: (2026)
When Reject Turns into Accept: Quantifying the Vulnerability of LLM-Based Scientific Reviewers to Indirect Prompt Injection
von: Sahoo, Devanshu, et al.
Veröffentlicht: (2025)
von: Sahoo, Devanshu, et al.
Veröffentlicht: (2025)
Multilingual Mathematical Reasoning: Advancing Open-Source LLMs in Hindi and English
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
Confidence is Not Competence
von: Sanyal, Debdeep, et al.
Veröffentlicht: (2025)
von: Sanyal, Debdeep, et al.
Veröffentlicht: (2025)
Measuring Representation Robustness in Large Language Models for Geometry
von: Jawandhia, Vedant, et al.
Veröffentlicht: (2026)
von: Jawandhia, Vedant, et al.
Veröffentlicht: (2026)
Agents Are All You Need for LLM Unlearning
von: Sanyal, Debdeep, et al.
Veröffentlicht: (2025)
von: Sanyal, Debdeep, et al.
Veröffentlicht: (2025)
The Compliance Paradox: Semantic-Instruction Decoupling in Automated Academic Code Evaluation
von: Sahoo, Devanshu, et al.
Veröffentlicht: (2026)
von: Sahoo, Devanshu, et al.
Veröffentlicht: (2026)
A Critical Review of Monte Carlo Algorithms Balancing Performance and Probabilistic Accuracy with AI Augmented Framework
von: Prasad, Ravi
Veröffentlicht: (2025)
von: Prasad, Ravi
Veröffentlicht: (2025)
Evaluating Novelty in AI-Generated Research Plans Using Multi-Workflow LLM Pipelines
von: Saraogi, Devesh, et al.
Veröffentlicht: (2025)
von: Saraogi, Devesh, et al.
Veröffentlicht: (2025)
A Comprehensive Review of Datasets for Clinical Mental Health AI Systems
von: Mandal, Aishik, et al.
Veröffentlicht: (2025)
von: Mandal, Aishik, et al.
Veröffentlicht: (2025)
QualEval: Qualitative Evaluation for Model Improvement
von: Murahari, Vishvak, et al.
Veröffentlicht: (2023)
von: Murahari, Vishvak, et al.
Veröffentlicht: (2023)
A Hybrid Supervised-LLM Pipeline for Actionable Suggestion Mining in Unstructured Customer Reviews
von: Trivedi, Aakash, et al.
Veröffentlicht: (2026)
von: Trivedi, Aakash, et al.
Veröffentlicht: (2026)
PassiveQA: A Three-Action Framework for Epistemically Calibrated Question Answering via Supervised Finetuning
von: Baidya, Madhav S
Veröffentlicht: (2026)
von: Baidya, Madhav S
Veröffentlicht: (2026)
UnStar: Unlearning with Self-Taught Anti-Sample Reasoning for LLMs
von: Sinha, Yash, et al.
Veröffentlicht: (2024)
von: Sinha, Yash, et al.
Veröffentlicht: (2024)
Nine Ways to Break Copyright Law and Why Our LLM Won't: A Fair Use Aligned Generation Framework
von: Sharma, Aakash Sen, et al.
Veröffentlicht: (2025)
von: Sharma, Aakash Sen, et al.
Veröffentlicht: (2025)
Detecting the Machine: A Comprehensive Benchmark of AI-Generated Text Detectors Across Architectures, Domains, and Adversarial Conditions
von: Baidya, Madhav S., et al.
Veröffentlicht: (2026)
von: Baidya, Madhav S., et al.
Veröffentlicht: (2026)
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language
von: Song, Seyoung, et al.
Veröffentlicht: (2025)
von: Song, Seyoung, et al.
Veröffentlicht: (2025)
Beyond Sentiment: A Multi-Agent Pipeline for Actionable Business Advice from Reviews
von: Bhandari, Kartikey Singh, et al.
Veröffentlicht: (2026)
von: Bhandari, Kartikey Singh, et al.
Veröffentlicht: (2026)
EvalCards: A Framework for Standardized Evaluation Reporting
von: Dhar, Ruchira, et al.
Veröffentlicht: (2025)
von: Dhar, Ruchira, et al.
Veröffentlicht: (2025)
ScholarEval: Research Idea Evaluation Grounded in Literature
von: Moussa, Hanane Nour, et al.
Veröffentlicht: (2025)
von: Moussa, Hanane Nour, et al.
Veröffentlicht: (2025)
Automated Concept Discovery for LLM-as-a-Judge Preference Analysis
von: Wedgwood, James, et al.
Veröffentlicht: (2026)
von: Wedgwood, James, et al.
Veröffentlicht: (2026)
Can AI Be a Good Peer Reviewer? A Survey of Peer Review Process, Evaluation, and the Future
von: Wu, Sihong, et al.
Veröffentlicht: (2026)
von: Wu, Sihong, et al.
Veröffentlicht: (2026)
ArxEval: Evaluating Retrieval and Generation in Language Models for Scientific Literature
von: Sinha, Aarush, et al.
Veröffentlicht: (2025)
von: Sinha, Aarush, et al.
Veröffentlicht: (2025)
NorEval: A Norwegian Language Understanding and Generation Evaluation Benchmark
von: Mikhailov, Vladislav, et al.
Veröffentlicht: (2025)
von: Mikhailov, Vladislav, et al.
Veröffentlicht: (2025)
SurveyEval: Towards Comprehensive Evaluation of LLM-Generated Academic Surveys
von: Zhao, Jiahao, et al.
Veröffentlicht: (2025)
von: Zhao, Jiahao, et al.
Veröffentlicht: (2025)
Infinite Problem Generator: Verifiably Scaling Physics Reasoning Data with Agentic Workflows
von: Sharan, Aditya, et al.
Veröffentlicht: (2026)
von: Sharan, Aditya, et al.
Veröffentlicht: (2026)
Evaluating LLMs for Zeolite Synthesis Event Extraction (ZSEE): A Systematic Analysis of Prompting Strategies
von: Rathore, Charan Prakash, et al.
Veröffentlicht: (2025)
von: Rathore, Charan Prakash, et al.
Veröffentlicht: (2025)
A Sentiment Consolidation Framework for Meta-Review Generation
von: Li, Miao, et al.
Veröffentlicht: (2024)
von: Li, Miao, et al.
Veröffentlicht: (2024)
ZeroSumEval: An Extensible Framework For Scaling LLM Evaluation with Inter-Model Competition
von: Alyahya, Hisham A., et al.
Veröffentlicht: (2025)
von: Alyahya, Hisham A., et al.
Veröffentlicht: (2025)
AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data
von: Wu, JiaRu, et al.
Veröffentlicht: (2025)
von: Wu, JiaRu, et al.
Veröffentlicht: (2025)
VHDL-Eval: A Framework for Evaluating Large Language Models in VHDL Code Generation
von: Vijayaraghavan, Prashanth, et al.
Veröffentlicht: (2024)
von: Vijayaraghavan, Prashanth, et al.
Veröffentlicht: (2024)
FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models
von: Yu, Zhuohao, et al.
Veröffentlicht: (2024)
von: Yu, Zhuohao, et al.
Veröffentlicht: (2024)
IndicEval: A Bilingual Indian Educational Evaluation Framework for Large Language Models
von: Bharti, Saurabh, et al.
Veröffentlicht: (2026)
von: Bharti, Saurabh, et al.
Veröffentlicht: (2026)
SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence
von: Wang, Yiheng, et al.
Veröffentlicht: (2025)
von: Wang, Yiheng, et al.
Veröffentlicht: (2025)
Conversation AI Dialog for Medicare powered by Finetuning and Retrieval Augmented Generation
von: Agrawal, Atharva Mangeshkumar, et al.
Veröffentlicht: (2025)
von: Agrawal, Atharva Mangeshkumar, et al.
Veröffentlicht: (2025)
Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review Composition
von: Tang, Xuemei, et al.
Veröffentlicht: (2024)
von: Tang, Xuemei, et al.
Veröffentlicht: (2024)
Fusion-Eval: Integrating Assistant Evaluators with LLMs
von: Shu, Lei, et al.
Veröffentlicht: (2023)
von: Shu, Lei, et al.
Veröffentlicht: (2023)
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
von: Wang, Ganghua, et al.
Veröffentlicht: (2025)
von: Wang, Ganghua, et al.
Veröffentlicht: (2025)
Graphing the Truth: Structured Visualizations for Automated Hallucination Detection in LLMs
von: Agrawal, Tanmay
Veröffentlicht: (2025)
von: Agrawal, Tanmay
Veröffentlicht: (2025)
Ähnliche Einträge
-
CricBench: A Multilingual Benchmark for Evaluating LLMs in Cricket Analytics
von: Agarwal, Parth, et al.
Veröffentlicht: (2025) -
LLM-as-a-Judge for Time Series Explanations
von: Sivalingam, Preetham, et al.
Veröffentlicht: (2026) -
When Reject Turns into Accept: Quantifying the Vulnerability of LLM-Based Scientific Reviewers to Indirect Prompt Injection
von: Sahoo, Devanshu, et al.
Veröffentlicht: (2025) -
Multilingual Mathematical Reasoning: Advancing Open-Source LLMs in Hindi and English
von: Anand, Avinash, et al.
Veröffentlicht: (2024) -
Confidence is Not Competence
von: Sanyal, Debdeep, et al.
Veröffentlicht: (2025)