Bench4KE: Benchmarking Automated Competency Question Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Lippolis, Anna Sofia, Ragagni, Minh Davide, Ciancarini, Paolo, Nuzzolese, Andrea Giovanni, Presutti, Valentina |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing multimodal analogical reasoning with Logic Augmented Generation
by: Lippolis, Anna Sofia, et al.
Published: (2025)
by: Lippolis, Anna Sofia, et al.
Published: (2025)
Logic Augmented Generation
by: Gangemi, Aldo, et al.
Published: (2024)
by: Gangemi, Aldo, et al.
Published: (2024)
The Medical Metaphors Corpus (MCC)
by: Lippolis, Anna Sofia, et al.
Published: (2025)
by: Lippolis, Anna Sofia, et al.
Published: (2025)
Assessing the Capability of Large Language Models for Domain-Specific Ontology Generation
by: Lippolis, Anna Sofia, et al.
Published: (2025)
by: Lippolis, Anna Sofia, et al.
Published: (2025)
NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
by: Moore, Robert J., et al.
Published: (2026)
by: Moore, Robert J., et al.
Published: (2026)
Large Language Models Assisting Ontology Evaluation
by: Lippolis, Anna Sofia, et al.
Published: (2025)
by: Lippolis, Anna Sofia, et al.
Published: (2025)
Seeing the Intangible: Survey of Image Classification into High-Level and Abstract Categories
by: Pandiani, Delfina Sol Martinez, et al.
Published: (2023)
by: Pandiani, Delfina Sol Martinez, et al.
Published: (2023)
DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis
by: Patel, Liana, et al.
Published: (2025)
by: Patel, Liana, et al.
Published: (2025)
TaskBench: Benchmarking Large Language Models for Task Automation
by: Shen, Yongliang, et al.
Published: (2023)
by: Shen, Yongliang, et al.
Published: (2023)
Ontology Generation using Large Language Models
by: Lippolis, Anna Sofia, et al.
Published: (2025)
by: Lippolis, Anna Sofia, et al.
Published: (2025)
SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia
by: Liu, Chaoqun, et al.
Published: (2025)
by: Liu, Chaoqun, et al.
Published: (2025)
FHIR-AgentBench: Benchmarking LLM Agents for Realistic Interoperable EHR Question Answering
by: Lee, Gyubok, et al.
Published: (2025)
by: Lee, Gyubok, et al.
Published: (2025)
Streamlining Knowledge Graph Creation with PyRML
by: Nuzzolese, Andrea Giovanni
Published: (2025)
by: Nuzzolese, Andrea Giovanni
Published: (2025)
BacktestBench: Benchmarking Large Language Models for Automated Quantitative Strategy Backtesting
by: Wang, Zhensheng, et al.
Published: (2026)
by: Wang, Zhensheng, et al.
Published: (2026)
VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models
by: Dong, Nguyen Tien, et al.
Published: (2025)
by: Dong, Nguyen Tien, et al.
Published: (2025)
Evaluating the Fitness of Ontologies for the Task of Question Generation
by: Alkhuzaey, Samah, et al.
Published: (2025)
by: Alkhuzaey, Samah, et al.
Published: (2025)
$μ$KE: Matryoshka Unstructured Knowledge Editing of Large Language Models
by: Su, Zian, et al.
Published: (2025)
by: Su, Zian, et al.
Published: (2025)
WilKE: Wise-Layer Knowledge Editor for Lifelong Knowledge Editing
by: Hu, Chenhui, et al.
Published: (2024)
by: Hu, Chenhui, et al.
Published: (2024)
DarkBench: Benchmarking Dark Patterns in Large Language Models
by: Kran, Esben, et al.
Published: (2025)
by: Kran, Esben, et al.
Published: (2025)
LongGenBench: Long-context Generation Benchmark
by: Liu, Xiang, et al.
Published: (2024)
by: Liu, Xiang, et al.
Published: (2024)
WritingBench: A Comprehensive Benchmark for Generative Writing
by: Wu, Yuning, et al.
Published: (2025)
by: Wu, Yuning, et al.
Published: (2025)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
by: Tu, Xinming, et al.
Published: (2026)
by: Tu, Xinming, et al.
Published: (2026)
Dynamic Few-Shot Learning for Knowledge Graph Question Answering
by: D'Abramo, Jacopo, et al.
Published: (2024)
by: D'Abramo, Jacopo, et al.
Published: (2024)
QGEval: Benchmarking Multi-dimensional Evaluation for Question Generation
by: Fu, Weiping, et al.
Published: (2024)
by: Fu, Weiping, et al.
Published: (2024)
Automated Question Generation on Tabular Data for Conversational Data Exploration
by: Chaudhuri, Ritwik, et al.
Published: (2024)
by: Chaudhuri, Ritwik, et al.
Published: (2024)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
by: Lin, Bill Yuchen, et al.
Published: (2024)
by: Lin, Bill Yuchen, et al.
Published: (2024)
SlovKE: A Large-Scale Dataset and LLM Evaluation for Slovak Keyphrase Extraction
by: Števaňák, David, et al.
Published: (2026)
by: Števaňák, David, et al.
Published: (2026)
RideKE: Leveraging Low-Resource, User-Generated Twitter Content for Sentiment and Emotion Detection in Kenyan Code-Switched Dataset
by: Etori, Naome A., et al.
Published: (2025)
by: Etori, Naome A., et al.
Published: (2025)
AI Idea Bench 2025: AI Research Idea Generation Benchmark
by: Qiu, Yansheng, et al.
Published: (2025)
by: Qiu, Yansheng, et al.
Published: (2025)
Evaluating Clinical Competencies of Large Language Models with a General Practice Benchmark
by: Li, Zheqing, et al.
Published: (2025)
by: Li, Zheqing, et al.
Published: (2025)
Automated Generation and Tagging of Knowledge Components from Multiple-Choice Questions
by: Moore, Steven, et al.
Published: (2024)
by: Moore, Steven, et al.
Published: (2024)
FinAgentBench: A Benchmark Dataset for Agentic Retrieval in Financial Question Answering
by: Choi, Chanyeol, et al.
Published: (2025)
by: Choi, Chanyeol, et al.
Published: (2025)
TopBench: A Benchmark for Implicit Prediction and Reasoning over Tabular Question Answering
by: Ji, An-Yang, et al.
Published: (2026)
by: Ji, An-Yang, et al.
Published: (2026)
seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMs
by: Ramezanali, Mohammad, et al.
Published: (2025)
by: Ramezanali, Mohammad, et al.
Published: (2025)
ATG: Benchmarking Automated Theorem Generation for Generative Language Models
by: Lin, Xiaohan, et al.
Published: (2024)
by: Lin, Xiaohan, et al.
Published: (2024)
MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems
by: Thakur, Nandan, et al.
Published: (2024)
by: Thakur, Nandan, et al.
Published: (2024)
FDARxBench: Benchmarking Regulatory and Clinical Reasoning on FDA Generic Drug Assessment
by: Xiong, Betty, et al.
Published: (2026)
by: Xiong, Betty, et al.
Published: (2026)
MIRROR: A Novel Approach for the Automated Evaluation of Open-Ended Question Generation
by: Deroy, Aniket, et al.
Published: (2024)
by: Deroy, Aniket, et al.
Published: (2024)
The Future of Learning in the Age of Generative AI: Automated Question Generation and Assessment with Large Language Models
by: Maity, Subhankar, et al.
Published: (2024)
by: Maity, Subhankar, et al.
Published: (2024)
MetaKE: Meta-Learning for Knowledge Editing Toward a Better Accuracy-Editability Trade-off
by: Liu, Shuxin, et al.
Published: (2026)
by: Liu, Shuxin, et al.
Published: (2026)
Similar Items
-
Enhancing multimodal analogical reasoning with Logic Augmented Generation
by: Lippolis, Anna Sofia, et al.
Published: (2025) -
Logic Augmented Generation
by: Gangemi, Aldo, et al.
Published: (2024) -
The Medical Metaphors Corpus (MCC)
by: Lippolis, Anna Sofia, et al.
Published: (2025) -
Assessing the Capability of Large Language Models for Domain-Specific Ontology Generation
by: Lippolis, Anna Sofia, et al.
Published: (2025) -
NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
by: Moore, Robert J., et al.
Published: (2026)