Reference-Free Evaluation of Taxonomies
Fuente:
arXiv
Guardado en:
| Autores principales: | Wullschleger, Pascal, Zarharan, Majid, Daly, Donnacha, Pouly, Marc, Foster, Jennifer |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FoodTaxo: Generating Food Taxonomies with Large Language Models
por: Wullschleger, Pascal, et al.
Publicado: (2025)
por: Wullschleger, Pascal, et al.
Publicado: (2025)
Tell Me Why: Explainable Public Health Fact-Checking with Large Language Models
por: Zarharan, Majid, et al.
Publicado: (2024)
por: Zarharan, Majid, et al.
Publicado: (2024)
FarExStance: Explainable Stance Detection for Farsi
por: Zarharan, Majid, et al.
Publicado: (2024)
por: Zarharan, Majid, et al.
Publicado: (2024)
Estimating Text Similarity based on Semantic Concept Embeddings
por: der Brück, Tim vor, et al.
Publicado: (2024)
por: der Brück, Tim vor, et al.
Publicado: (2024)
Mitigating the Impact of Reference Quality on Evaluation of Summarization Systems with Reference-Free Metrics
por: Gigant, Théo, et al.
Publicado: (2024)
por: Gigant, Théo, et al.
Publicado: (2024)
Measuring the Robustness of Reference-Free Dialogue Evaluation Systems
por: Vasselli, Justin, et al.
Publicado: (2025)
por: Vasselli, Justin, et al.
Publicado: (2025)
TrustScore: Reference-Free Evaluation of LLM Response Trustworthiness
por: Zheng, Danna, et al.
Publicado: (2024)
por: Zheng, Danna, et al.
Publicado: (2024)
A Taxonomy for Design and Evaluation of Prompt-Based Natural Language Explanations
por: Nejadgholi, Isar, et al.
Publicado: (2025)
por: Nejadgholi, Isar, et al.
Publicado: (2025)
Evaluating the Utility of Grounding Documents with Reference-Free LLM-based Metrics
por: Hua, Yilun, et al.
Publicado: (2026)
por: Hua, Yilun, et al.
Publicado: (2026)
BanglaSummEval: Reference-Free Factual Consistency Evaluation for Bangla Summarization
por: Rafid, Ahmed, et al.
Publicado: (2026)
por: Rafid, Ahmed, et al.
Publicado: (2026)
SCORE: Specificity, Context Utilization, Robustness, and Relevance for Reference-Free LLM Evaluation
por: Shomee, Homaira Huda, et al.
Publicado: (2026)
por: Shomee, Homaira Huda, et al.
Publicado: (2026)
A Dual-Axis Taxonomy of Knowledge Editing for LLMs: From Mechanisms to Functions
por: Salehoof, Amir Mohammad, et al.
Publicado: (2025)
por: Salehoof, Amir Mohammad, et al.
Publicado: (2025)
Evaluation Revisited: A Taxonomy of Evaluation Concerns in Natural Language Processing
por: Dhar, Ruchira, et al.
Publicado: (2026)
por: Dhar, Ruchira, et al.
Publicado: (2026)
CREAM: Comparison-Based Reference-Free ELO-Ranked Automatic Evaluation for Meeting Summarization
por: Gong, Ziwei, et al.
Publicado: (2024)
por: Gong, Ziwei, et al.
Publicado: (2024)
Monotonic Reference-Free Refinement for Autoformalization
por: Zhang, Lan, et al.
Publicado: (2026)
por: Zhang, Lan, et al.
Publicado: (2026)
Taxonomy-based CheckList for Large Language Model Evaluation
por: Zhang, Damin
Publicado: (2023)
por: Zhang, Damin
Publicado: (2023)
LITE: LLM-Impelled efficient Taxonomy Evaluation
por: Zhang, Lin, et al.
Publicado: (2025)
por: Zhang, Lin, et al.
Publicado: (2025)
Evaluating Nuanced Bias in Large Language Model Free Response Answers
por: Healey, Jennifer, et al.
Publicado: (2024)
por: Healey, Jennifer, et al.
Publicado: (2024)
PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs
por: Fu, Xiao, et al.
Publicado: (2025)
por: Fu, Xiao, et al.
Publicado: (2025)
An Online Reference-Free Evaluation Framework for Flowchart Image-to-Code Generation
por: Nguyen, Giang Son, et al.
Publicado: (2026)
por: Nguyen, Giang Son, et al.
Publicado: (2026)
The PICCO Framework for Large Language Model Prompting: A Taxonomy and Reference Architecture for Prompt Structure
por: Cook, David A.
Publicado: (2026)
por: Cook, David A.
Publicado: (2026)
Tutor Move Taxonomy: A Theory-Aligned Framework for Analyzing Instructional Moves in Tutoring
por: Zhou, Zhuqian, et al.
Publicado: (2026)
por: Zhou, Zhuqian, et al.
Publicado: (2026)
Evaluating Optimal Reference Translations
por: Zouhar, Vilém, et al.
Publicado: (2023)
por: Zouhar, Vilém, et al.
Publicado: (2023)
MILE-RefHumEval: A Reference-Free, Multi-Independent LLM Framework for Human-Aligned Evaluation
por: Srun, Nalin, et al.
Publicado: (2026)
por: Srun, Nalin, et al.
Publicado: (2026)
ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization
por: Bae, Suyoung, et al.
Publicado: (2026)
por: Bae, Suyoung, et al.
Publicado: (2026)
LLMs as Function Approximators: Terminology, Taxonomy, and Questions for Evaluation
por: Schlangen, David
Publicado: (2024)
por: Schlangen, David
Publicado: (2024)
SocREval: Large Language Models with the Socratic Method for Reference-Free Reasoning Evaluation
por: He, Hangfeng, et al.
Publicado: (2023)
por: He, Hangfeng, et al.
Publicado: (2023)
Generating Leakage-Free Benchmarks for Robust RAG Evaluation
por: Liu, Jiayi, et al.
Publicado: (2026)
por: Liu, Jiayi, et al.
Publicado: (2026)
Cobra Effect in Reference-Free Image Captioning Metrics
por: Ma, Zheng, et al.
Publicado: (2024)
por: Ma, Zheng, et al.
Publicado: (2024)
Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors
por: Maurya, Kaushal Kumar, et al.
Publicado: (2024)
por: Maurya, Kaushal Kumar, et al.
Publicado: (2024)
An Examination of the Robustness of Reference-Free Image Captioning Evaluation Metrics
por: Ahmadi, Saba, et al.
Publicado: (2023)
por: Ahmadi, Saba, et al.
Publicado: (2023)
Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework
por: Jain, Shomik, et al.
Publicado: (2025)
por: Jain, Shomik, et al.
Publicado: (2025)
References Matter: Investigating the Impact of Reference Set Variation on Summarization Evaluation
por: Casola, Silvia, et al.
Publicado: (2025)
por: Casola, Silvia, et al.
Publicado: (2025)
From Performance to Purpose: A Sociotechnical Taxonomy for Evaluating Large Language Model Utility
por: Levinson, Gavin, et al.
Publicado: (2026)
por: Levinson, Gavin, et al.
Publicado: (2026)
Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory
por: Nejadgholi, Isar, et al.
Publicado: (2026)
por: Nejadgholi, Isar, et al.
Publicado: (2026)
Can Deep Research Agents Retrieve and Organize? Evaluating the Synthesis Gap with Expert Taxonomies
por: Zhang, Ming, et al.
Publicado: (2026)
por: Zhang, Ming, et al.
Publicado: (2026)
References Indeed Matter? Reference-Free Preference Optimization for Conversational Query Reformulation
por: Kim, Doyoung, et al.
Publicado: (2025)
por: Kim, Doyoung, et al.
Publicado: (2025)
Evaluating the performance of state-of-the-art esg domain-specific pre-trained large language models in text classification against existing models and traditional machine learning techniques
por: Chung, Tin Yuet, et al.
Publicado: (2024)
por: Chung, Tin Yuet, et al.
Publicado: (2024)
Evaluating Large Language Models on Time Series Feature Understanding: A Comprehensive Taxonomy and Benchmark
por: Fons, Elizabeth, et al.
Publicado: (2024)
por: Fons, Elizabeth, et al.
Publicado: (2024)
A Unified Taxonomy-Guided Instruction Tuning Framework for Entity Set Expansion and Taxonomy Expansion
por: Shen, Yanzhen, et al.
Publicado: (2024)
por: Shen, Yanzhen, et al.
Publicado: (2024)
Ejemplares similares
-
FoodTaxo: Generating Food Taxonomies with Large Language Models
por: Wullschleger, Pascal, et al.
Publicado: (2025) -
Tell Me Why: Explainable Public Health Fact-Checking with Large Language Models
por: Zarharan, Majid, et al.
Publicado: (2024) -
FarExStance: Explainable Stance Detection for Farsi
por: Zarharan, Majid, et al.
Publicado: (2024) -
Estimating Text Similarity based on Semantic Concept Embeddings
por: der Brück, Tim vor, et al.
Publicado: (2024) -
Mitigating the Impact of Reference Quality on Evaluation of Summarization Systems with Reference-Free Metrics
por: Gigant, Théo, et al.
Publicado: (2024)