An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Martin-Boyle, Anna, Humphreys, William, Brown, Martha, Leckey, Cara, Kaur, Harmanpreet |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A
di: Martin-Boyle, Anna, et al.
Pubblicazione: (2026)
di: Martin-Boyle, Anna, et al.
Pubblicazione: (2026)
Designing and Evaluating Chain-of-Hints for Scientific Question Answering
di: Jangra, Anubhav, et al.
Pubblicazione: (2025)
di: Jangra, Anubhav, et al.
Pubblicazione: (2025)
A Survey of Large Language Model Agents for Question Answering
di: Yue, Murong
Pubblicazione: (2025)
di: Yue, Murong
Pubblicazione: (2025)
Granuscore: A Reference-Free Measure of Granularity for Text Analysis and Question Answering
di: Ellinger, Lukas, et al.
Pubblicazione: (2026)
di: Ellinger, Lukas, et al.
Pubblicazione: (2026)
RealitySummary: Exploring On-Demand Mixed Reality Text Summarization and Question Answering using Large Language Models
di: Gunturu, Aditya, et al.
Pubblicazione: (2024)
di: Gunturu, Aditya, et al.
Pubblicazione: (2024)
Leveraging the Domain Adaptation of Retrieval Augmented Generation Models for Question Answering and Reducing Hallucination
di: Rakin, Salman, et al.
Pubblicazione: (2024)
di: Rakin, Salman, et al.
Pubblicazione: (2024)
Question Answering for Decisionmaking in Green Building Design: A Multimodal Data Reasoning Method Driven by Large Language Models
di: Li, Yihui, et al.
Pubblicazione: (2024)
di: Li, Yihui, et al.
Pubblicazione: (2024)
Exploring the Potential of Large Language Models for Estimating the Reading Comprehension Question Difficulty
di: Jain, Yoshee, et al.
Pubblicazione: (2025)
di: Jain, Yoshee, et al.
Pubblicazione: (2025)
Wisdom of the Crowd, Without the Crowd: A Socratic LLM for Asynchronous Deliberation on Perspectivist Data
di: Khadar, Malik, et al.
Pubblicazione: (2025)
di: Khadar, Malik, et al.
Pubblicazione: (2025)
Unraveling Entangled Feeds: Rethinking Social Media Design to Enhance User Well-being
di: Milton, Ashlee, et al.
Pubblicazione: (2026)
di: Milton, Ashlee, et al.
Pubblicazione: (2026)
How the Advent of Ubiquitous Large Language Models both Stymie and Turbocharge Dynamic Adversarial Question Generation
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024)
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024)
Evaluating Large Language Models for Health-related Queries with Presuppositions
di: Kaur, Navreet, et al.
Pubblicazione: (2023)
di: Kaur, Navreet, et al.
Pubblicazione: (2023)
Evaluation Of P300 Speller Performance Using Large Language Models Along With Cross-Subject Training
di: Parthasarathy, Nithin, et al.
Pubblicazione: (2024)
di: Parthasarathy, Nithin, et al.
Pubblicazione: (2024)
LOGOS: LLM-driven End-to-End Grounded Theory Development and Schema Induction for Qualitative Research
di: Pi, Xinyu, et al.
Pubblicazione: (2025)
di: Pi, Xinyu, et al.
Pubblicazione: (2025)
Evaluating the Usage of African-American Vernacular English in Large Language Models
di: Dunlap, Deja, et al.
Pubblicazione: (2026)
di: Dunlap, Deja, et al.
Pubblicazione: (2026)
PALLM: Evaluating and Enhancing PALLiative Care Conversations with Large Language Models
di: Wang, Zhiyuan, et al.
Pubblicazione: (2024)
di: Wang, Zhiyuan, et al.
Pubblicazione: (2024)
Social Skill Training with Large Language Models
di: Yang, Diyi, et al.
Pubblicazione: (2024)
di: Yang, Diyi, et al.
Pubblicazione: (2024)
Who's Asking? Simulating Role-Based Questions for Conversational AI Evaluation
di: Kaur, Navreet, et al.
Pubblicazione: (2025)
di: Kaur, Navreet, et al.
Pubblicazione: (2025)
QACP: An Annotated Question Answering Dataset for Assisting Chinese Python Programming Learners
di: Xiao, Rui, et al.
Pubblicazione: (2024)
di: Xiao, Rui, et al.
Pubblicazione: (2024)
Beyond Semantic Similarity: A Component-Wise Evaluation Framework for Medical Question Answering Systems with Health Equity Implications
di: Sakib, Abu Noman Md, et al.
Pubblicazione: (2026)
di: Sakib, Abu Noman Md, et al.
Pubblicazione: (2026)
Prompt4Vis: Prompting Large Language Models with Example Mining and Schema Filtering for Tabular Data Visualization
di: Li, Shuaimin, et al.
Pubblicazione: (2024)
di: Li, Shuaimin, et al.
Pubblicazione: (2024)
Toward Automated Qualitative Analysis: Leveraging Large Language Models for Tutoring Dialogue Evaluation
di: Gu, Megan, et al.
Pubblicazione: (2025)
di: Gu, Megan, et al.
Pubblicazione: (2025)
DiscussLLM: Teaching Large Language Models When to Speak
di: Patel, Deep Anil, et al.
Pubblicazione: (2025)
di: Patel, Deep Anil, et al.
Pubblicazione: (2025)
Evaluating Telugu Proficiency in Large Language Models_ A Comparative Analysis of ChatGPT and Gemini
di: Kishore, Katikela Sreeharsha, et al.
Pubblicazione: (2024)
di: Kishore, Katikela Sreeharsha, et al.
Pubblicazione: (2024)
Generate-Then-Validate: A Novel Question Generation Approach Using Small Language Models
di: Wei, Yumou, et al.
Pubblicazione: (2025)
di: Wei, Yumou, et al.
Pubblicazione: (2025)
AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?
di: Gor, Maharshi, et al.
Pubblicazione: (2026)
di: Gor, Maharshi, et al.
Pubblicazione: (2026)
TaleFrame: An Interactive Story Generation System with Fine-Grained Control and Large Language Models
di: Wang, Yunchao, et al.
Pubblicazione: (2025)
di: Wang, Yunchao, et al.
Pubblicazione: (2025)
Exploring the Impact of Personality Traits on Conversational Recommender Systems: A Simulation with Large Language Models
di: Zhao, Xiaoyan, et al.
Pubblicazione: (2025)
di: Zhao, Xiaoyan, et al.
Pubblicazione: (2025)
Roleplay-doh: Enabling Domain-Experts to Create LLM-simulated Patients via Eliciting and Adhering to Principles
di: Louie, Ryan, et al.
Pubblicazione: (2024)
di: Louie, Ryan, et al.
Pubblicazione: (2024)
Evaluating Large Language Models in Theory of Mind Tasks
di: Kosinski, Michal
Pubblicazione: (2023)
di: Kosinski, Michal
Pubblicazione: (2023)
Towards a Design Guideline for RPA Evaluation: A Survey of Large Language Model-Based Role-Playing Agents
di: Chen, Chaoran, et al.
Pubblicazione: (2025)
di: Chen, Chaoran, et al.
Pubblicazione: (2025)
Educational Question Generation of Children Storybooks via Question Type Distribution Learning and Event-Centric Summarization
di: Zhao, Zhenjie, et al.
Pubblicazione: (2022)
di: Zhao, Zhenjie, et al.
Pubblicazione: (2022)
Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models
di: Ibrahim, Lujain, et al.
Pubblicazione: (2025)
di: Ibrahim, Lujain, et al.
Pubblicazione: (2025)
Mitigating Harmful Erraticism in LLMs Through Dialectical Behavior Therapy Based De-Escalation Strategies
di: Rangarajan, Pooja, et al.
Pubblicazione: (2025)
di: Rangarajan, Pooja, et al.
Pubblicazione: (2025)
Mapping Geopolitical Bias in 11 Large Language Models: A Bilingual, Dual-Framing Analysis of U.S.-China Tensions
di: Guey, William, et al.
Pubblicazione: (2025)
di: Guey, William, et al.
Pubblicazione: (2025)
Are Humans as Brittle as Large Language Models?
di: Li, Jiahui, et al.
Pubblicazione: (2025)
di: Li, Jiahui, et al.
Pubblicazione: (2025)
Do LLMs Make Mistakes Like Students? Exploring Natural Alignment between Language Models and Human Error Patterns
di: Liu, Naiming, et al.
Pubblicazione: (2025)
di: Liu, Naiming, et al.
Pubblicazione: (2025)
Are We There Yet? Revealing the Risks of Utilizing Large Language Models in Scholarly Peer Review
di: Ye, Rui, et al.
Pubblicazione: (2024)
di: Ye, Rui, et al.
Pubblicazione: (2024)
An Evaluation of Estimative Uncertainty in Large Language Models
di: Tang, Zhisheng, et al.
Pubblicazione: (2024)
di: Tang, Zhisheng, et al.
Pubblicazione: (2024)
Evaluating the Prompt Steerability of Large Language Models
di: Miehling, Erik, et al.
Pubblicazione: (2024)
di: Miehling, Erik, et al.
Pubblicazione: (2024)
Documenti analoghi
-
PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A
di: Martin-Boyle, Anna, et al.
Pubblicazione: (2026) -
Designing and Evaluating Chain-of-Hints for Scientific Question Answering
di: Jangra, Anubhav, et al.
Pubblicazione: (2025) -
A Survey of Large Language Model Agents for Question Answering
di: Yue, Murong
Pubblicazione: (2025) -
Granuscore: A Reference-Free Measure of Granularity for Text Analysis and Question Answering
di: Ellinger, Lukas, et al.
Pubblicazione: (2026) -
RealitySummary: Exploring On-Demand Mixed Reality Text Summarization and Question Answering using Large Language Models
di: Gunturu, Aditya, et al.
Pubblicazione: (2024)