Rethinking CyberSecEval: An LLM-Aided Approach to Evaluation Critique
Fuente:
arXiv
Saved in:
| Main Authors: | Hariharan, Suhas, Majid, Zainab Ali, Veuthey, Jaime Raldua, Haimes, Jacob |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
by: Veuthey, Jaime Raldua, et al.
Published: (2025)
by: Veuthey, Jaime Raldua, et al.
Published: (2025)
CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models
by: Bhatt, Manish, et al.
Published: (2024)
by: Bhatt, Manish, et al.
Published: (2024)
View From Above: A Framework for Evaluating Distribution Shifts in Model Behavior
by: Chopra, Tanush, et al.
Published: (2024)
by: Chopra, Tanush, et al.
Published: (2024)
SycEval: Evaluating LLM Sycophancy
by: Fanous, Aaron, et al.
Published: (2025)
by: Fanous, Aaron, et al.
Published: (2025)
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
by: Reese, May Lynn, et al.
Published: (2026)
by: Reese, May Lynn, et al.
Published: (2026)
Toward Virtuous Reinforcement Learning: A Critique and Roadmap
by: Ghasemi, Majid, et al.
Published: (2025)
by: Ghasemi, Majid, et al.
Published: (2025)
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
by: Costarelli, Anthony, et al.
Published: (2024)
by: Costarelli, Anthony, et al.
Published: (2024)
Dataset Featurization: Uncovering Natural Language Features through Unsupervised Data Reconstruction
by: Bravansky, Michal, et al.
Published: (2025)
by: Bravansky, Michal, et al.
Published: (2025)
Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps
by: Chona, Alankrit, et al.
Published: (2026)
by: Chona, Alankrit, et al.
Published: (2026)
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
by: Sturgeon, Benjamin, et al.
Published: (2025)
by: Sturgeon, Benjamin, et al.
Published: (2025)
HonestCyberEval: An AI Cyber Risk Benchmark for Automated Software Exploitation
by: Ristea, Dan, et al.
Published: (2024)
by: Ristea, Dan, et al.
Published: (2024)
CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation
by: Ke, Pei, et al.
Published: (2023)
by: Ke, Pei, et al.
Published: (2023)
Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts
by: Haimes, Jacob, et al.
Published: (2024)
by: Haimes, Jacob, et al.
Published: (2024)
Artificial Intelligence for Optimal Learning: A Comparative Approach towards AI-Enhanced Learning Environments
by: Hariharan, Ananth
Published: (2025)
by: Hariharan, Ananth
Published: (2025)
The potential of LLM-generated reports in DevSecOps
by: Lykousas, Nikolaos, et al.
Published: (2024)
by: Lykousas, Nikolaos, et al.
Published: (2024)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
by: Khatun, Aisha, et al.
Published: (2024)
by: Khatun, Aisha, et al.
Published: (2024)
The Critique of Critique
by: Sun, Shichao, et al.
Published: (2024)
by: Sun, Shichao, et al.
Published: (2024)
ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition
by: Khan, Haidar, et al.
Published: (2025)
by: Khan, Haidar, et al.
Published: (2025)
SurveyEval: Towards Comprehensive Evaluation of LLM-Generated Academic Surveys
by: Zhao, Jiahao, et al.
Published: (2025)
by: Zhao, Jiahao, et al.
Published: (2025)
RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following
by: Pan, Tianjun, et al.
Published: (2026)
by: Pan, Tianjun, et al.
Published: (2026)
TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents
by: Chen, Weiyi, et al.
Published: (2026)
by: Chen, Weiyi, et al.
Published: (2026)
Rethinking Human Preference Evaluation of LLM Rationales
by: Li, Ziang, et al.
Published: (2025)
by: Li, Ziang, et al.
Published: (2025)
Check-Eval: A Checklist-based Approach for Evaluating Text Quality
by: Pereira, Jayr, et al.
Published: (2024)
by: Pereira, Jayr, et al.
Published: (2024)
CySecBench: Generative AI-based CyberSecurity-focused Prompt Dataset for Benchmarking Large Language Models
by: Wahréus, Johan, et al.
Published: (2025)
by: Wahréus, Johan, et al.
Published: (2025)
RxEval: A Prescription-Level Benchmark for Evaluating LLM Medication Recommendation
by: Chen, Shuhao, et al.
Published: (2026)
by: Chen, Shuhao, et al.
Published: (2026)
Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks
by: Chen, Sizhe, et al.
Published: (2025)
by: Chen, Sizhe, et al.
Published: (2025)
Plan Verification for LLM-Based Embodied Task Completion Agents
by: Hariharan, Ananth, et al.
Published: (2025)
by: Hariharan, Ananth, et al.
Published: (2025)
On the Importance of Neural Membrane Potential Leakage for LIDAR-based Robot Obstacle Avoidance using Spiking Neural Networks
by: Ali, Zainab, et al.
Published: (2025)
by: Ali, Zainab, et al.
Published: (2025)
ViLLM-Eval: A Comprehensive Evaluation Suite for Vietnamese Large Language Models
by: Nguyen, Trong-Hieu, et al.
Published: (2024)
by: Nguyen, Trong-Hieu, et al.
Published: (2024)
ZeroSumEval: An Extensible Framework For Scaling LLM Evaluation with Inter-Model Competition
by: Alyahya, Hisham A., et al.
Published: (2025)
by: Alyahya, Hisham A., et al.
Published: (2025)
AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data
by: Wu, JiaRu, et al.
Published: (2025)
by: Wu, JiaRu, et al.
Published: (2025)
Stepwise Think-Critique: A Unified Framework for Robust and Interpretable LLM Reasoning
by: Xu, Jiaqi, et al.
Published: (2025)
by: Xu, Jiaqi, et al.
Published: (2025)
Semantic Mastery: Enhancing LLMs with Advanced Natural Language Understanding
by: Hariharan, Mohanakrishnan
Published: (2025)
by: Hariharan, Mohanakrishnan
Published: (2025)
Reinforcement Learning Integrated Agentic RAG for Software Test Cases Authoring
by: Hariharan, Mohanakrishnan
Published: (2025)
by: Hariharan, Mohanakrishnan
Published: (2025)
LLM4SecHW: Leveraging Domain Specific Large Language Model for Hardware Debugging
by: Fu, Weimin, et al.
Published: (2024)
by: Fu, Weimin, et al.
Published: (2024)
Interactive Critique-Revision Training for Reliable Structured LLM Generation
by: Yu, Fei Xu, et al.
Published: (2026)
by: Yu, Fei Xu, et al.
Published: (2026)
Enhancing LLM Planning Capabilities through Intrinsic Self-Critique
by: Bohnet, Bernd, et al.
Published: (2025)
by: Bohnet, Bernd, et al.
Published: (2025)
Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
by: Ye, Bowen, et al.
Published: (2026)
by: Ye, Bowen, et al.
Published: (2026)
Digital Socrates: Evaluating LLMs through Explanation Critiques
by: Gu, Yuling, et al.
Published: (2023)
by: Gu, Yuling, et al.
Published: (2023)
RAS-Eval: A Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments
by: Fu, Yuchuan, et al.
Published: (2025)
by: Fu, Yuchuan, et al.
Published: (2025)
Similar Items
-
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
by: Veuthey, Jaime Raldua, et al.
Published: (2025) -
CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models
by: Bhatt, Manish, et al.
Published: (2024) -
View From Above: A Framework for Evaluating Distribution Shifts in Model Behavior
by: Chopra, Tanush, et al.
Published: (2024) -
SycEval: Evaluating LLM Sycophancy
by: Fanous, Aaron, et al.
Published: (2025) -
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
by: Reese, May Lynn, et al.
Published: (2026)