How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives
Fuente:
arXiv
Saved in:
| Main Authors: | Ichmoukhamedov, Timour, Hinns, James, Martens, David |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring the generalization of LLM truth directions on conversational formats
by: Ichmoukhamedov, Timour, et al.
Published: (2025)
by: Ichmoukhamedov, Timour, et al.
Published: (2025)
Cash or Comfort? How LLMs Value Your Inconvenience
by: Cedro, Mateusz, et al.
Published: (2025)
by: Cedro, Mateusz, et al.
Published: (2025)
Exposing Image Classifier Shortcuts with Counterfactual Frequency (CoF) Tables
by: Hinns, James, et al.
Published: (2024)
by: Hinns, James, et al.
Published: (2024)
Aggregating Local Saliency Maps for Semi-Global Explainable Image Classification
by: Hinns, James, et al.
Published: (2025)
by: Hinns, James, et al.
Published: (2025)
Tell Me a Story! Narrative-Driven XAI with Large Language Models
by: Martens, David, et al.
Published: (2023)
by: Martens, David, et al.
Published: (2023)
Faithfulness metric fusion: Improving the evaluation of LLM trustworthiness across domains
by: Malin, Ben, et al.
Published: (2025)
by: Malin, Ben, et al.
Published: (2025)
Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator
by: Kirstein, Frederic, et al.
Published: (2024)
by: Kirstein, Frederic, et al.
Published: (2024)
LiTransProQA: an LLM-based Literary Translation evaluation metric with Professional Question Answering
by: Zhang, Ran, et al.
Published: (2025)
by: Zhang, Ran, et al.
Published: (2025)
Multi-Agent LLM Judge: automatic personalized LLM judge design for evaluating natural language generation applications
by: Cao, Hongliu, et al.
Published: (2025)
by: Cao, Hongliu, et al.
Published: (2025)
How good is GPT at writing political speeches for the White House?
by: Savoy, Jacques
Published: (2024)
by: Savoy, Jacques
Published: (2024)
On the Importance and Evaluation of Narrativity in Natural Language AI Explanations
by: Cedro, Mateusz, et al.
Published: (2026)
by: Cedro, Mateusz, et al.
Published: (2026)
An evaluation of LLM code generation capabilities through graded exercises
by: Jiménez, Álvaro Barbero
Published: (2024)
by: Jiménez, Álvaro Barbero
Published: (2024)
Kantian-Utilitarian XAI: Meta-Explained
by: Atf, Zahra, et al.
Published: (2025)
by: Atf, Zahra, et al.
Published: (2025)
Towards Autoformalization of LLM-generated Outputs for Requirement Verification
by: Gupte, Mihir, et al.
Published: (2025)
by: Gupte, Mihir, et al.
Published: (2025)
The illusion of a perfect metric: Why evaluating AI's words is harder than it looks
by: Oliva, Maria Paz, et al.
Published: (2025)
by: Oliva, Maria Paz, et al.
Published: (2025)
DepressLLM: Interpretable domain-adapted language model for depression detection from real-world narratives
by: Moon, Sehwan, et al.
Published: (2025)
by: Moon, Sehwan, et al.
Published: (2025)
Explainable AI: XAI-Guided Context-Aware Data Augmentation
by: Mersha, Melkamu Abay, et al.
Published: (2025)
by: Mersha, Melkamu Abay, et al.
Published: (2025)
Reinforcement learning for path integrals in quantum statistical physics
by: Ichmoukhamedov, Timour, et al.
Published: (2026)
by: Ichmoukhamedov, Timour, et al.
Published: (2026)
Are LLMs good pragmatic speakers?
by: Jian, Mingyue, et al.
Published: (2024)
by: Jian, Mingyue, et al.
Published: (2024)
WHODUNIT: Evaluation benchmark for culprit detection in mystery stories
by: Gupta, Kshitij
Published: (2025)
by: Gupta, Kshitij
Published: (2025)
Towards Geo-Culturally Grounded LLM Generations
by: Lertvittayakumjorn, Piyawat, et al.
Published: (2025)
by: Lertvittayakumjorn, Piyawat, et al.
Published: (2025)
Plancraft: an evaluation dataset for planning with LLM agents
by: Dagan, Gautier, et al.
Published: (2024)
by: Dagan, Gautier, et al.
Published: (2024)
From Black Boxes to Conversations: Incorporating XAI in a Conversational Agent
by: Nguyen, Van Bach, et al.
Published: (2022)
by: Nguyen, Van Bach, et al.
Published: (2022)
How well do LLMs cite relevant medical references? An evaluation framework and analyses
by: Wu, Kevin, et al.
Published: (2024)
by: Wu, Kevin, et al.
Published: (2024)
Gender Bias in LLM-generated Interview Responses
by: Kong, Haein, et al.
Published: (2024)
by: Kong, Haein, et al.
Published: (2024)
Learning to Summarize from LLM-generated Feedback
by: Song, Hwanjun, et al.
Published: (2024)
by: Song, Hwanjun, et al.
Published: (2024)
Benchmark of stylistic variation in LLM-generated texts
by: Milička, Jiří, et al.
Published: (2025)
by: Milička, Jiří, et al.
Published: (2025)
Would a Large Language Model Pay Extra for a View? Inferring Willingness to Pay from Subjective Choices
by: Reusens, Manon, et al.
Published: (2026)
by: Reusens, Manon, et al.
Published: (2026)
XAI4LLM. Let Machine Learning Models and LLMs Collaborate for Enhanced In-Context Learning in Healthcare
by: Nazary, Fatemeh, et al.
Published: (2024)
by: Nazary, Fatemeh, et al.
Published: (2024)
Reliable and diverse evaluation of LLM medical knowledge mastery
by: Zhou, Yuxuan, et al.
Published: (2024)
by: Zhou, Yuxuan, et al.
Published: (2024)
LLMs as annotators of credibility assessment in Danish asylum decisions: evaluating classification performance and errors beyond aggregated metrics
by: Humblot-Renaux, Galadrielle, et al.
Published: (2026)
by: Humblot-Renaux, Galadrielle, et al.
Published: (2026)
Explaining Humour Style Classifications: An XAI Approach to Understanding Computational Humour Analysis
by: Kenneth, Mary Ogbuka, et al.
Published: (2025)
by: Kenneth, Mary Ogbuka, et al.
Published: (2025)
Semantic uncertainty in advanced decoding methods for LLM generation
by: Foodeei, Darius, et al.
Published: (2025)
by: Foodeei, Darius, et al.
Published: (2025)
Automated test generation to evaluate tool-augmented LLMs as conversational AI agents
by: Arcadinho, Samuel, et al.
Published: (2024)
by: Arcadinho, Samuel, et al.
Published: (2024)
LongStory: Coherent, Complete and Length Controlled Long story Generation
by: Park, Kyeongman, et al.
Published: (2023)
by: Park, Kyeongman, et al.
Published: (2023)
Impact of enriched meaning representations for language generation in dialogue tasks: A comprehensive exploration of the relevance of tasks, corpora and metrics
by: Vázquez, Alain, et al.
Published: (2026)
by: Vázquez, Alain, et al.
Published: (2026)
Serialized EHR make for good text representations
by: Chou, Zhirong, et al.
Published: (2025)
by: Chou, Zhirong, et al.
Published: (2025)
AI-generated stories favour stability over change: homogeneity and cultural stereotyping in narratives generated by gpt-4o-mini
by: Rettberg, Jill Walker, et al.
Published: (2025)
by: Rettberg, Jill Walker, et al.
Published: (2025)
Kalahi: A handcrafted, grassroots cultural LLM evaluation suite for Filipino
by: Montalan, Jann Railey, et al.
Published: (2024)
by: Montalan, Jann Railey, et al.
Published: (2024)
LAG-XAI: A Lie-Inspired Affine Geometric Framework for Interpretable Paraphrasing in Transformer Latent Spaces
by: Mazurets, Olexander, et al.
Published: (2026)
by: Mazurets, Olexander, et al.
Published: (2026)
Similar Items
-
Exploring the generalization of LLM truth directions on conversational formats
by: Ichmoukhamedov, Timour, et al.
Published: (2025) -
Cash or Comfort? How LLMs Value Your Inconvenience
by: Cedro, Mateusz, et al.
Published: (2025) -
Exposing Image Classifier Shortcuts with Counterfactual Frequency (CoF) Tables
by: Hinns, James, et al.
Published: (2024) -
Aggregating Local Saliency Maps for Semi-Global Explainable Image Classification
by: Hinns, James, et al.
Published: (2025) -
Tell Me a Story! Narrative-Driven XAI with Large Language Models
by: Martens, David, et al.
Published: (2023)