Guardado en:
| Autores principales: | Moore, Steven, Costello, Eamon, Nguyen, Huy A., Stamper, John |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2405.20529 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Automated Generation and Tagging of Knowledge Components from Multiple-Choice Questions
por: Moore, Steven, et al.
Publicado: (2024)
por: Moore, Steven, et al.
Publicado: (2024)
Cognitive Agent Compilation for Explicit Problem Solver Modeling
por: Moon, Hyeongdon, et al.
Publicado: (2026)
por: Moon, Hyeongdon, et al.
Publicado: (2026)
Automatic Evaluation of Healthcare LLMs Beyond Question-Answering
por: Arias-Duart, Anna, et al.
Publicado: (2025)
por: Arias-Duart, Anna, et al.
Publicado: (2025)
Small but Significant: On the Promise of Small Language Models for Accessible AIED
por: Wei, Yumou, et al.
Publicado: (2025)
por: Wei, Yumou, et al.
Publicado: (2025)
Long-Span Question-Answering: Automatic Question Generation and QA-System Ranking via Side-by-Side Evaluation
por: Bohnet, Bernd, et al.
Publicado: (2024)
por: Bohnet, Bernd, et al.
Publicado: (2024)
AGENT-CQ: Automatic Generation and Evaluation of Clarifying Questions for Conversational Search with LLMs
por: Siro, Clemencia, et al.
Publicado: (2024)
por: Siro, Clemencia, et al.
Publicado: (2024)
Automatic Generation of Inference Making Questions for Reading Comprehension Assessments
por: Ma, Wanjing Anya, et al.
Publicado: (2025)
por: Ma, Wanjing Anya, et al.
Publicado: (2025)
Automatic Dataset Generation for Knowledge Intensive Question Answering Tasks
por: Yuen, Sizhe, et al.
Publicado: (2025)
por: Yuen, Sizhe, et al.
Publicado: (2025)
Hallucination-Free Automatic Question & Answer Generation for Intuitive Learning
por: Wang, Nicholas X., et al.
Publicado: (2026)
por: Wang, Nicholas X., et al.
Publicado: (2026)
Automatic Question & Answer Generation Using Generative Large Language Model (LLM)
por: Ehsan, Md. Alvee, et al.
Publicado: (2025)
por: Ehsan, Md. Alvee, et al.
Publicado: (2025)
WalledEval: A Comprehensive Safety Evaluation Toolkit for Large Language Models
por: Gupta, Prannaya, et al.
Publicado: (2024)
por: Gupta, Prannaya, et al.
Publicado: (2024)
SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence
por: Wang, Yiheng, et al.
Publicado: (2025)
por: Wang, Yiheng, et al.
Publicado: (2025)
Are Large Language Models Consistent over Value-laden Questions?
por: Moore, Jared, et al.
Publicado: (2024)
por: Moore, Jared, et al.
Publicado: (2024)
Confabulation: The Surprising Value of Large Language Model Hallucinations
por: Sui, Peiqi, et al.
Publicado: (2024)
por: Sui, Peiqi, et al.
Publicado: (2024)
Comprehensive and Practical Evaluation of Retrieval-Augmented Generation Systems for Medical Question Answering
por: Ngo, Nghia Trung, et al.
Publicado: (2024)
por: Ngo, Nghia Trung, et al.
Publicado: (2024)
Bias in the Tails: How Name-conditioned Evaluative Framing in Resume Summaries Destabilizes LLM-based Hiring
por: Nghiem, Huy, et al.
Publicado: (2026)
por: Nghiem, Huy, et al.
Publicado: (2026)
CompeteSMoE -- Statistically Guaranteed Mixture of Experts Training via Competition
por: Nguyen, Nam V., et al.
Publicado: (2025)
por: Nguyen, Nam V., et al.
Publicado: (2025)
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
por: Nghiem, Huy, et al.
Publicado: (2025)
por: Nghiem, Huy, et al.
Publicado: (2025)
From Answers to Questions: EQGBench for Evaluating LLMs' Educational Question Generation
por: Zhou, Chengliang, et al.
Publicado: (2025)
por: Zhou, Chengliang, et al.
Publicado: (2025)
R-Eval: A Unified Toolkit for Evaluating Domain Knowledge of Retrieval Augmented Large Language Models
por: Tu, Shangqing, et al.
Publicado: (2024)
por: Tu, Shangqing, et al.
Publicado: (2024)
Automatic Legal Writing Evaluation of LLMs
por: Pires, Ramon, et al.
Publicado: (2025)
por: Pires, Ramon, et al.
Publicado: (2025)
Sketch: A Toolkit for Streamlining LLM Operations
por: Jiang, Xin, et al.
Publicado: (2024)
por: Jiang, Xin, et al.
Publicado: (2024)
The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory
por: Schmucker, Robin, et al.
Publicado: (2025)
por: Schmucker, Robin, et al.
Publicado: (2025)
Automatic Generation of Question Hints for Mathematics Problems using Large Language Models in Educational Technology
por: Tonga, Junior Cedric, et al.
Publicado: (2024)
por: Tonga, Junior Cedric, et al.
Publicado: (2024)
Training Computer Use Agents to Assess the Usability of Graphical User Interfaces
por: Gao, Alice, et al.
Publicado: (2026)
por: Gao, Alice, et al.
Publicado: (2026)
Adaptive Stopping for Multi-Turn LLM Reasoning
por: Zhou, Xiaofan, et al.
Publicado: (2026)
por: Zhou, Xiaofan, et al.
Publicado: (2026)
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
por: Xu, Austin, et al.
Publicado: (2025)
por: Xu, Austin, et al.
Publicado: (2025)
VLSP 2025 MLQA-TSR Challenge: Vietnamese Multimodal Legal Question Answering on Traffic Sign Regulation
por: Luu, Son T., et al.
Publicado: (2025)
por: Luu, Son T., et al.
Publicado: (2025)
Feedback Forensics: A Toolkit to Measure AI Personality
por: Findeis, Arduin, et al.
Publicado: (2025)
por: Findeis, Arduin, et al.
Publicado: (2025)
VLQA: The First Comprehensive, Large, and High-Quality Vietnamese Dataset for Legal Question Answering
por: Nguyen, Tan-Minh, et al.
Publicado: (2025)
por: Nguyen, Tan-Minh, et al.
Publicado: (2025)
Evaluating the Fitness of Ontologies for the Task of Question Generation
por: Alkhuzaey, Samah, et al.
Publicado: (2025)
por: Alkhuzaey, Samah, et al.
Publicado: (2025)
Qworld: Question-Specific Evaluation Criteria for LLMs
por: Gao, Shanghua, et al.
Publicado: (2026)
por: Gao, Shanghua, et al.
Publicado: (2026)
Can we train ASR systems on Code-switch without real code-switch data? Case study for Singapore's languages
por: Nguyen, Tuan, et al.
Publicado: (2025)
por: Nguyen, Tuan, et al.
Publicado: (2025)
Cross-Attention Watermarking of Large Language Models
por: Baldassini, Folco Bertini, et al.
Publicado: (2024)
por: Baldassini, Folco Bertini, et al.
Publicado: (2024)
Jury: A Comprehensive Evaluation Toolkit
por: Cavusoglu, Devrim, et al.
Publicado: (2023)
por: Cavusoglu, Devrim, et al.
Publicado: (2023)
Towards Automatic Evaluation of Task-Oriented Dialogue Flows
por: Mirtaheri, Mehrnoosh, et al.
Publicado: (2024)
por: Mirtaheri, Mehrnoosh, et al.
Publicado: (2024)
Submodular Evaluation Subset Selection in Automatic Prompt Optimization
por: Nian, Jinming, et al.
Publicado: (2026)
por: Nian, Jinming, et al.
Publicado: (2026)
LMFlow: An Extensible Toolkit for Finetuning and Inference of Large Foundation Models
por: Diao, Shizhe, et al.
Publicado: (2023)
por: Diao, Shizhe, et al.
Publicado: (2023)
JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models
por: Ran, Delong, et al.
Publicado: (2024)
por: Ran, Delong, et al.
Publicado: (2024)
GR-NLP-TOOLKIT: An Open-Source NLP Toolkit for Modern Greek
por: Loukas, Lefteris, et al.
Publicado: (2024)
por: Loukas, Lefteris, et al.
Publicado: (2024)
Ejemplares similares
-
Automated Generation and Tagging of Knowledge Components from Multiple-Choice Questions
por: Moore, Steven, et al.
Publicado: (2024) -
Cognitive Agent Compilation for Explicit Problem Solver Modeling
por: Moon, Hyeongdon, et al.
Publicado: (2026) -
Automatic Evaluation of Healthcare LLMs Beyond Question-Answering
por: Arias-Duart, Anna, et al.
Publicado: (2025) -
Small but Significant: On the Promise of Small Language Models for Accessible AIED
por: Wei, Yumou, et al.
Publicado: (2025) -
Long-Span Question-Answering: Automatic Question Generation and QA-System Ranking via Side-by-Side Evaluation
por: Bohnet, Bernd, et al.
Publicado: (2024)