The QCET Taxonomy of Standard Quality Criterion Names and Definitions for the Evaluation of NLP Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Belz, Anya, Mille, Simon, Thomson, Craig |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Knowledge Distillation for Large Language Models
by: La Torre, Alejandro Paredes, et al.
Published: (2026)
by: La Torre, Alejandro Paredes, et al.
Published: (2026)
PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models
by: Friedland, Gerald, et al.
Published: (2024)
by: Friedland, Gerald, et al.
Published: (2024)
Semantic Delta: An Interpretable Signal Differentiating Human and LLMs Dialogue
by: Scantamburlo, Riccardo, et al.
Published: (2026)
by: Scantamburlo, Riccardo, et al.
Published: (2026)
ValueDCG: Measuring Comprehensive Human Value Understanding Ability of Language Models
by: Zhang, Zhaowei, et al.
Published: (2023)
by: Zhang, Zhaowei, et al.
Published: (2023)
Enhancing Retrieval-Augmented Generation for Electric Power Industry Customer Support
by: Chan, Hei Yu, et al.
Published: (2025)
by: Chan, Hei Yu, et al.
Published: (2025)
Leveraging Synthetic Data for Question Answering with Multilingual LLMs in the Agricultural Domain
by: Kaur, Rishemjit, et al.
Published: (2025)
by: Kaur, Rishemjit, et al.
Published: (2025)
Exploring Model Invariance with Discrete Search for Ultra-Low-Bit Quantization
by: Wen, Yuqiao, et al.
Published: (2025)
by: Wen, Yuqiao, et al.
Published: (2025)
EBBS: An Ensemble with Bi-Level Beam Search for Zero-Shot Machine Translation
by: Wen, Yuqiao, et al.
Published: (2024)
by: Wen, Yuqiao, et al.
Published: (2024)
The Battle of LLMs: A Comparative Study in Conversational QA Tasks
by: Rangapur, Aryan, et al.
Published: (2024)
by: Rangapur, Aryan, et al.
Published: (2024)
Opinion Mining on Offshore Wind Energy for Environmental Engineering
by: Bittencourt, Isabele, et al.
Published: (2024)
by: Bittencourt, Isabele, et al.
Published: (2024)
QRA++: Quantified Reproducibility Assessment for Common Types of Results in Natural Language Processing
by: Belz, Anya
Published: (2025)
by: Belz, Anya
Published: (2025)
HEDS 3.0: The Human Evaluation Data Sheet Version 3.0
by: Belz, Anya, et al.
Published: (2024)
by: Belz, Anya, et al.
Published: (2024)
APP: Accelerated Path Patching with Task-Specific Pruning
by: Andersen, Frauke, et al.
Published: (2025)
by: Andersen, Frauke, et al.
Published: (2025)
Challenges and Opportunities of NLP for HR Applications: A Discussion Paper
by: Leidner, Jochen L., et al.
Published: (2024)
by: Leidner, Jochen L., et al.
Published: (2024)
LLMs as Signal Detectors: Sensitivity, Bias, and the Temperature-Criterion Analogy
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Resource for Error Analysis in Text Simplification: New Taxonomy and Test Collection
by: Vendeville, Benjamin, et al.
Published: (2025)
by: Vendeville, Benjamin, et al.
Published: (2025)
Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles
by: Budagam, Devichand, et al.
Published: (2024)
by: Budagam, Devichand, et al.
Published: (2024)
Defining and Quantifying Creative Behavior in Popular Image Generators
by: Ramaswamy, Aditi, et al.
Published: (2025)
by: Ramaswamy, Aditi, et al.
Published: (2025)
Robust Explanations for User Trust in Enterprise NLP Systems
by: Zhang, Guilin, et al.
Published: (2026)
by: Zhang, Guilin, et al.
Published: (2026)
Evaluating Vision Transformer Models for Visual Quality Control in Industrial Manufacturing
by: Alber, Miriam, et al.
Published: (2024)
by: Alber, Miriam, et al.
Published: (2024)
When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't
by: Nemitz, Jonathan, et al.
Published: (2026)
by: Nemitz, Jonathan, et al.
Published: (2026)
CodeNER: Code Prompting for Named Entity Recognition
by: Han, Sungwoo, et al.
Published: (2025)
by: Han, Sungwoo, et al.
Published: (2025)
Tuning-less Object Naming with a Foundation Model
by: Lucny, Andrej, et al.
Published: (2023)
by: Lucny, Andrej, et al.
Published: (2023)
Benchmarking the Performance of Pre-trained LLMs across Urdu NLP Tasks
by: Tahir, Munief Hassan, et al.
Published: (2024)
by: Tahir, Munief Hassan, et al.
Published: (2024)
High Throughput Phenotyping of Physician Notes with Large Language and Hybrid NLP Models
by: Munzir, Syed I., et al.
Published: (2024)
by: Munzir, Syed I., et al.
Published: (2024)
Deepfake Technology Unveiled: The Commoditization of AI and Its Impact on Digital Trust
by: Popa, Claudiu, et al.
Published: (2025)
by: Popa, Claudiu, et al.
Published: (2025)
SCOR: A Framework for Responsible AI Innovation in Digital Ecosystems
by: Torkestani, Mohammad Saleh, et al.
Published: (2025)
by: Torkestani, Mohammad Saleh, et al.
Published: (2025)
A global AI community requires language-diverse publishing
by: Lepp, Haley, et al.
Published: (2024)
by: Lepp, Haley, et al.
Published: (2024)
Presumed Cultural Identity: How Names Shape LLM Responses
by: Pawar, Siddhesh, et al.
Published: (2025)
by: Pawar, Siddhesh, et al.
Published: (2025)
Beyond Black-Box Labels: Interpretable Criteria for Diagnosing Subjective NLP Tasks
by: Rair, Nisrine, et al.
Published: (2026)
by: Rair, Nisrine, et al.
Published: (2026)
Few-Shot Learning for Mental Disorder Detection: A Continuous Multi-Prompt Engineering Approach with Medical Knowledge Injection
by: Liu, Haoxin, et al.
Published: (2024)
by: Liu, Haoxin, et al.
Published: (2024)
The Dark Side of AI Transformers: Sentiment Polarization & the Loss of Business Neutrality by NLP Transformers
by: Kumar, Prasanna
Published: (2026)
by: Kumar, Prasanna
Published: (2026)
When Does Data Augmentation Help? Evaluating LLM and Back-Translation Methods for Hausa and Fongbe NLP
by: Adjovi, Mahounan Pericles, et al.
Published: (2026)
by: Adjovi, Mahounan Pericles, et al.
Published: (2026)
Key Algorithms for Keyphrase Generation: Instruction-Based LLMs for Russian Scientific Keyphrases
by: Glazkova, Anna, et al.
Published: (2024)
by: Glazkova, Anna, et al.
Published: (2024)
What Makes a Good AI Review? Concern-Level Diagnostics for AI Peer Review
by: Jin, Ming
Published: (2026)
by: Jin, Ming
Published: (2026)
SIGN: Schema-Induced Games for Naming
by: Zhang, Ryan, et al.
Published: (2025)
by: Zhang, Ryan, et al.
Published: (2025)
Counterfactual Strategies for Markov Decision Processes
by: Kobialka, Paul, et al.
Published: (2025)
by: Kobialka, Paul, et al.
Published: (2025)
Attribution-based Explanations for Markov Decision Processes
by: Kobialka, Paul, et al.
Published: (2026)
by: Kobialka, Paul, et al.
Published: (2026)
Seeing Through Circuits: Faithful Mechanistic Interpretability for Vision Transformers
by: Żukowska, Nina, et al.
Published: (2026)
by: Żukowska, Nina, et al.
Published: (2026)
Measurement Risk in Supervised Financial NLP: Rubric and Metric Sensitivity on JF-ICR
by: Chang, Sidi, et al.
Published: (2026)
by: Chang, Sidi, et al.
Published: (2026)
Similar Items
-
Knowledge Distillation for Large Language Models
by: La Torre, Alejandro Paredes, et al.
Published: (2026) -
PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models
by: Friedland, Gerald, et al.
Published: (2024) -
Semantic Delta: An Interpretable Signal Differentiating Human and LLMs Dialogue
by: Scantamburlo, Riccardo, et al.
Published: (2026) -
ValueDCG: Measuring Comprehensive Human Value Understanding Ability of Language Models
by: Zhang, Zhaowei, et al.
Published: (2023) -
Enhancing Retrieval-Augmented Generation for Electric Power Industry Customer Support
by: Chan, Hei Yu, et al.
Published: (2025)