Automated Benchmark Generation from Domain Guidelines Informed by Bloom's Taxonomy
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Si, Khiem, Le Huy, Szymanski, Annalisa, Metoyer, Ronald, Hua, Ting, Chawla, Nitesh V. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
di: Li, Peiyu, et al.
Pubblicazione: (2025)
di: Li, Peiyu, et al.
Pubblicazione: (2025)
Building Scaffolding Dialogue Data with LLM-Simulated Novices
di: Chen, Si, et al.
Pubblicazione: (2025)
di: Chen, Si, et al.
Pubblicazione: (2025)
AgentDrug: Utilizing Large Language Models in An Agentic Workflow for Zero-Shot Molecular Editing
di: Le, Khiem, et al.
Pubblicazione: (2024)
di: Le, Khiem, et al.
Pubblicazione: (2024)
FLAME: Towards Federated Fine-Tuning Large Language Models Through Adaptive SMoE
di: Le, Khiem, et al.
Pubblicazione: (2025)
di: Le, Khiem, et al.
Pubblicazione: (2025)
TeachingCoach: A Fine-Tuned Scaffolding Chatbot for Instructional Guidance to Instructors
di: Molnar, Isabel, et al.
Pubblicazione: (2026)
di: Molnar, Isabel, et al.
Pubblicazione: (2026)
Key Considerations for Domain Expert Involvement in LLM Design and Evaluation: An Ethnographic Study
di: Szymanski, Annalisa, et al.
Pubblicazione: (2026)
di: Szymanski, Annalisa, et al.
Pubblicazione: (2026)
Automated Analysis of Learning Outcomes and Exam Questions Based on Bloom's Taxonomy
di: Kumar, Ramya, et al.
Pubblicazione: (2025)
di: Kumar, Ramya, et al.
Pubblicazione: (2025)
Exploring Conversational Design Choices in LLMs for Pedagogical Purposes: Socratic and Narrative Approaches for Improving Instructor's Teaching Practice
di: Chen, Si, et al.
Pubblicazione: (2025)
di: Chen, Si, et al.
Pubblicazione: (2025)
CrochetBench: Can Vision-Language Models Move from Describing to Doing in Crochet Domain?
di: Li, Peiyu, et al.
Pubblicazione: (2025)
di: Li, Peiyu, et al.
Pubblicazione: (2025)
AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration - Learning from Cheap, Optimizing Expensive
di: Guo, Taicheng, et al.
Pubblicazione: (2026)
di: Guo, Taicheng, et al.
Pubblicazione: (2026)
Detecting the Machine: A Comprehensive Benchmark of AI-Generated Text Detectors Across Architectures, Domains, and Adversarial Conditions
di: Baidya, Madhav S., et al.
Pubblicazione: (2026)
di: Baidya, Madhav S., et al.
Pubblicazione: (2026)
Breaking Language Barriers: Equitable Performance in Multilingual Language Models
di: Nagar, Tanay, et al.
Pubblicazione: (2025)
di: Nagar, Tanay, et al.
Pubblicazione: (2025)
Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy
di: Raimondi, Bianca, et al.
Pubblicazione: (2026)
di: Raimondi, Bianca, et al.
Pubblicazione: (2026)
How Effective is GPT-4 Turbo in Generating School-Level Questions from Textbooks Based on Bloom's Revised Taxonomy?
di: Maity, Subhankar, et al.
Pubblicazione: (2024)
di: Maity, Subhankar, et al.
Pubblicazione: (2024)
How Teachers Can Use Large Language Models and Bloom's Taxonomy to Create Educational Quizzes
di: Elkins, Sabina, et al.
Pubblicazione: (2024)
di: Elkins, Sabina, et al.
Pubblicazione: (2024)
Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation
di: Tian, Yijun, et al.
Pubblicazione: (2024)
di: Tian, Yijun, et al.
Pubblicazione: (2024)
RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG
di: Gao, Joshua, et al.
Pubblicazione: (2025)
di: Gao, Joshua, et al.
Pubblicazione: (2025)
Automated Educational Question Generation at Different Bloom's Skill Levels using Large Language Models: Strategies and Evaluation
di: Scaria, Nicy, et al.
Pubblicazione: (2024)
di: Scaria, Nicy, et al.
Pubblicazione: (2024)
Bridging the AI Adoption Gap: Designing an Interactive Pedagogical Agent for Higher Education Instructors
di: Chen, Si, et al.
Pubblicazione: (2025)
di: Chen, Si, et al.
Pubblicazione: (2025)
"Not in My Backyard": LLMs Uncover Online and Offline Social Biases Against Homelessness
di: Karr Jr., Jonathan A., et al.
Pubblicazione: (2025)
di: Karr Jr., Jonathan A., et al.
Pubblicazione: (2025)
Context Attribution with Multi-Armed Bandit Optimization
di: Pan, Deng, et al.
Pubblicazione: (2025)
di: Pan, Deng, et al.
Pubblicazione: (2025)
NGQA: A Nutritional Graph Question Answering Benchmark for Personalized Health-aware Nutritional Reasoning
di: Zhang, Zheyuan, et al.
Pubblicazione: (2024)
di: Zhang, Zheyuan, et al.
Pubblicazione: (2024)
ATG: Benchmarking Automated Theorem Generation for Generative Language Models
di: Lin, Xiaohan, et al.
Pubblicazione: (2024)
di: Lin, Xiaohan, et al.
Pubblicazione: (2024)
A Survey of Event Causality Identification: Taxonomy, Challenges, Assessment, and Prospects
di: Cheng, Qing, et al.
Pubblicazione: (2024)
di: Cheng, Qing, et al.
Pubblicazione: (2024)
Beyond SELECT: A Comprehensive Taxonomy-Guided Benchmark for Real-World Text-to-SQL Translation
di: Wang, Hao, et al.
Pubblicazione: (2025)
di: Wang, Hao, et al.
Pubblicazione: (2025)
Large Language Model based Multi-Agents: A Survey of Progress and Challenges
di: Guo, Taicheng, et al.
Pubblicazione: (2024)
di: Guo, Taicheng, et al.
Pubblicazione: (2024)
Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems
di: Cui, Tianyu, et al.
Pubblicazione: (2024)
di: Cui, Tianyu, et al.
Pubblicazione: (2024)
PPM: Automated Generation of Diverse Programming Problems for Benchmarking Code Generation Models
di: Chen, Simin, et al.
Pubblicazione: (2024)
di: Chen, Simin, et al.
Pubblicazione: (2024)
Bench4KE: Benchmarking Automated Competency Question Generation
di: Lippolis, Anna Sofia, et al.
Pubblicazione: (2025)
di: Lippolis, Anna Sofia, et al.
Pubblicazione: (2025)
MolX: Enhancing Large Language Models for Molecular Understanding With A Multi-Modal Extension
di: Le, Khiem, et al.
Pubblicazione: (2024)
di: Le, Khiem, et al.
Pubblicazione: (2024)
GTA: A Benchmark for General Tool Agents
di: Wang, Jize, et al.
Pubblicazione: (2024)
di: Wang, Jize, et al.
Pubblicazione: (2024)
Graph Neural Prompting with Large Language Models
di: Tian, Yijun, et al.
Pubblicazione: (2023)
di: Tian, Yijun, et al.
Pubblicazione: (2023)
A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future
di: Zhong, Jialun, et al.
Pubblicazione: (2025)
di: Zhong, Jialun, et al.
Pubblicazione: (2025)
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
di: Ye, Jiayi, et al.
Pubblicazione: (2024)
di: Ye, Jiayi, et al.
Pubblicazione: (2024)
Probing and Steering Evaluation Awareness of Language Models
di: Nguyen, Jord, et al.
Pubblicazione: (2025)
di: Nguyen, Jord, et al.
Pubblicazione: (2025)
MedMeta: A Benchmark for LLMs in Synthesizing Meta-Analysis Conclusion from Medical Studies
di: Ha, Huy Hoang, et al.
Pubblicazione: (2026)
di: Ha, Huy Hoang, et al.
Pubblicazione: (2026)
Charting the European LLM Benchmarking Landscape: A New Taxonomy and a Set of Best Practices
di: Vintar, Špela, et al.
Pubblicazione: (2025)
di: Vintar, Špela, et al.
Pubblicazione: (2025)
A Decade-Scale Benchmark Evaluating LLMs' Clinical Practice Guidelines Detection and Adherence in Multi-turn Conversations
di: Tan, Andong, et al.
Pubblicazione: (2026)
di: Tan, Andong, et al.
Pubblicazione: (2026)
FoodTaxo: Generating Food Taxonomies with Large Language Models
di: Wullschleger, Pascal, et al.
Pubblicazione: (2025)
di: Wullschleger, Pascal, et al.
Pubblicazione: (2025)
Towards Automated Smart Contract Generation: Evaluation, Benchmarking, and Retrieval-Augmented Repair
di: Chen, Zaoyu, et al.
Pubblicazione: (2025)
di: Chen, Zaoyu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
di: Li, Peiyu, et al.
Pubblicazione: (2025) -
Building Scaffolding Dialogue Data with LLM-Simulated Novices
di: Chen, Si, et al.
Pubblicazione: (2025) -
AgentDrug: Utilizing Large Language Models in An Agentic Workflow for Zero-Shot Molecular Editing
di: Le, Khiem, et al.
Pubblicazione: (2024) -
FLAME: Towards Federated Fine-Tuning Large Language Models Through Adaptive SMoE
di: Le, Khiem, et al.
Pubblicazione: (2025) -
TeachingCoach: A Fine-Tuned Scaffolding Chatbot for Instructional Guidance to Instructors
di: Molnar, Isabel, et al.
Pubblicazione: (2026)