Bench4KE: Benchmarking Automated Competency Question Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lippolis, Anna Sofia, Ragagni, Minh Davide, Ciancarini, Paolo, Nuzzolese, Andrea Giovanni, Presutti, Valentina |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhancing multimodal analogical reasoning with Logic Augmented Generation
von: Lippolis, Anna Sofia, et al.
Veröffentlicht: (2025)
von: Lippolis, Anna Sofia, et al.
Veröffentlicht: (2025)
Logic Augmented Generation
von: Gangemi, Aldo, et al.
Veröffentlicht: (2024)
von: Gangemi, Aldo, et al.
Veröffentlicht: (2024)
The Medical Metaphors Corpus (MCC)
von: Lippolis, Anna Sofia, et al.
Veröffentlicht: (2025)
von: Lippolis, Anna Sofia, et al.
Veröffentlicht: (2025)
Assessing the Capability of Large Language Models for Domain-Specific Ontology Generation
von: Lippolis, Anna Sofia, et al.
Veröffentlicht: (2025)
von: Lippolis, Anna Sofia, et al.
Veröffentlicht: (2025)
NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
von: Moore, Robert J., et al.
Veröffentlicht: (2026)
von: Moore, Robert J., et al.
Veröffentlicht: (2026)
Large Language Models Assisting Ontology Evaluation
von: Lippolis, Anna Sofia, et al.
Veröffentlicht: (2025)
von: Lippolis, Anna Sofia, et al.
Veröffentlicht: (2025)
Seeing the Intangible: Survey of Image Classification into High-Level and Abstract Categories
von: Pandiani, Delfina Sol Martinez, et al.
Veröffentlicht: (2023)
von: Pandiani, Delfina Sol Martinez, et al.
Veröffentlicht: (2023)
DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis
von: Patel, Liana, et al.
Veröffentlicht: (2025)
von: Patel, Liana, et al.
Veröffentlicht: (2025)
TaskBench: Benchmarking Large Language Models for Task Automation
von: Shen, Yongliang, et al.
Veröffentlicht: (2023)
von: Shen, Yongliang, et al.
Veröffentlicht: (2023)
Ontology Generation using Large Language Models
von: Lippolis, Anna Sofia, et al.
Veröffentlicht: (2025)
von: Lippolis, Anna Sofia, et al.
Veröffentlicht: (2025)
SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia
von: Liu, Chaoqun, et al.
Veröffentlicht: (2025)
von: Liu, Chaoqun, et al.
Veröffentlicht: (2025)
FHIR-AgentBench: Benchmarking LLM Agents for Realistic Interoperable EHR Question Answering
von: Lee, Gyubok, et al.
Veröffentlicht: (2025)
von: Lee, Gyubok, et al.
Veröffentlicht: (2025)
Streamlining Knowledge Graph Creation with PyRML
von: Nuzzolese, Andrea Giovanni
Veröffentlicht: (2025)
von: Nuzzolese, Andrea Giovanni
Veröffentlicht: (2025)
BacktestBench: Benchmarking Large Language Models for Automated Quantitative Strategy Backtesting
von: Wang, Zhensheng, et al.
Veröffentlicht: (2026)
von: Wang, Zhensheng, et al.
Veröffentlicht: (2026)
VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models
von: Dong, Nguyen Tien, et al.
Veröffentlicht: (2025)
von: Dong, Nguyen Tien, et al.
Veröffentlicht: (2025)
Evaluating the Fitness of Ontologies for the Task of Question Generation
von: Alkhuzaey, Samah, et al.
Veröffentlicht: (2025)
von: Alkhuzaey, Samah, et al.
Veröffentlicht: (2025)
$μ$KE: Matryoshka Unstructured Knowledge Editing of Large Language Models
von: Su, Zian, et al.
Veröffentlicht: (2025)
von: Su, Zian, et al.
Veröffentlicht: (2025)
WilKE: Wise-Layer Knowledge Editor for Lifelong Knowledge Editing
von: Hu, Chenhui, et al.
Veröffentlicht: (2024)
von: Hu, Chenhui, et al.
Veröffentlicht: (2024)
DarkBench: Benchmarking Dark Patterns in Large Language Models
von: Kran, Esben, et al.
Veröffentlicht: (2025)
von: Kran, Esben, et al.
Veröffentlicht: (2025)
LongGenBench: Long-context Generation Benchmark
von: Liu, Xiang, et al.
Veröffentlicht: (2024)
von: Liu, Xiang, et al.
Veröffentlicht: (2024)
WritingBench: A Comprehensive Benchmark for Generative Writing
von: Wu, Yuning, et al.
Veröffentlicht: (2025)
von: Wu, Yuning, et al.
Veröffentlicht: (2025)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
Dynamic Few-Shot Learning for Knowledge Graph Question Answering
von: D'Abramo, Jacopo, et al.
Veröffentlicht: (2024)
von: D'Abramo, Jacopo, et al.
Veröffentlicht: (2024)
QGEval: Benchmarking Multi-dimensional Evaluation for Question Generation
von: Fu, Weiping, et al.
Veröffentlicht: (2024)
von: Fu, Weiping, et al.
Veröffentlicht: (2024)
Automated Question Generation on Tabular Data for Conversational Data Exploration
von: Chaudhuri, Ritwik, et al.
Veröffentlicht: (2024)
von: Chaudhuri, Ritwik, et al.
Veröffentlicht: (2024)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
SlovKE: A Large-Scale Dataset and LLM Evaluation for Slovak Keyphrase Extraction
von: Števaňák, David, et al.
Veröffentlicht: (2026)
von: Števaňák, David, et al.
Veröffentlicht: (2026)
RideKE: Leveraging Low-Resource, User-Generated Twitter Content for Sentiment and Emotion Detection in Kenyan Code-Switched Dataset
von: Etori, Naome A., et al.
Veröffentlicht: (2025)
von: Etori, Naome A., et al.
Veröffentlicht: (2025)
AI Idea Bench 2025: AI Research Idea Generation Benchmark
von: Qiu, Yansheng, et al.
Veröffentlicht: (2025)
von: Qiu, Yansheng, et al.
Veröffentlicht: (2025)
Evaluating Clinical Competencies of Large Language Models with a General Practice Benchmark
von: Li, Zheqing, et al.
Veröffentlicht: (2025)
von: Li, Zheqing, et al.
Veröffentlicht: (2025)
Automated Generation and Tagging of Knowledge Components from Multiple-Choice Questions
von: Moore, Steven, et al.
Veröffentlicht: (2024)
von: Moore, Steven, et al.
Veröffentlicht: (2024)
FinAgentBench: A Benchmark Dataset for Agentic Retrieval in Financial Question Answering
von: Choi, Chanyeol, et al.
Veröffentlicht: (2025)
von: Choi, Chanyeol, et al.
Veröffentlicht: (2025)
TopBench: A Benchmark for Implicit Prediction and Reasoning over Tabular Question Answering
von: Ji, An-Yang, et al.
Veröffentlicht: (2026)
von: Ji, An-Yang, et al.
Veröffentlicht: (2026)
seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMs
von: Ramezanali, Mohammad, et al.
Veröffentlicht: (2025)
von: Ramezanali, Mohammad, et al.
Veröffentlicht: (2025)
ATG: Benchmarking Automated Theorem Generation for Generative Language Models
von: Lin, Xiaohan, et al.
Veröffentlicht: (2024)
von: Lin, Xiaohan, et al.
Veröffentlicht: (2024)
MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems
von: Thakur, Nandan, et al.
Veröffentlicht: (2024)
von: Thakur, Nandan, et al.
Veröffentlicht: (2024)
FDARxBench: Benchmarking Regulatory and Clinical Reasoning on FDA Generic Drug Assessment
von: Xiong, Betty, et al.
Veröffentlicht: (2026)
von: Xiong, Betty, et al.
Veröffentlicht: (2026)
MIRROR: A Novel Approach for the Automated Evaluation of Open-Ended Question Generation
von: Deroy, Aniket, et al.
Veröffentlicht: (2024)
von: Deroy, Aniket, et al.
Veröffentlicht: (2024)
The Future of Learning in the Age of Generative AI: Automated Question Generation and Assessment with Large Language Models
von: Maity, Subhankar, et al.
Veröffentlicht: (2024)
von: Maity, Subhankar, et al.
Veröffentlicht: (2024)
MetaKE: Meta-Learning for Knowledge Editing Toward a Better Accuracy-Editability Trade-off
von: Liu, Shuxin, et al.
Veröffentlicht: (2026)
von: Liu, Shuxin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Enhancing multimodal analogical reasoning with Logic Augmented Generation
von: Lippolis, Anna Sofia, et al.
Veröffentlicht: (2025) -
Logic Augmented Generation
von: Gangemi, Aldo, et al.
Veröffentlicht: (2024) -
The Medical Metaphors Corpus (MCC)
von: Lippolis, Anna Sofia, et al.
Veröffentlicht: (2025) -
Assessing the Capability of Large Language Models for Domain-Specific Ontology Generation
von: Lippolis, Anna Sofia, et al.
Veröffentlicht: (2025) -
NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
von: Moore, Robert J., et al.
Veröffentlicht: (2026)