We Need to Talk About Classification Evaluation Metrics in NLP
Fuente:
arXiv
Saved in:
| Main Authors: | Vickers, Peter, Barrault, Loïc, Monti, Emilio, Aletras, Nikolaos |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Incorporating Attribution Importance for Improving Faithfulness Metrics
by: Zhao, Zhixue, et al.
Published: (2023)
by: Zhao, Zhixue, et al.
Published: (2023)
Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
by: Xue, Huiyin, et al.
Published: (2025)
by: Xue, Huiyin, et al.
Published: (2025)
Improving Multimodal Classification of Social Media Posts by Leveraging Image-Text Auxiliary Tasks
by: Villegas, Danae Sánchez, et al.
Published: (2023)
by: Villegas, Danae Sánchez, et al.
Published: (2023)
Where does output diversity collapse in post-training?
by: Karouzos, Constantinos, et al.
Published: (2026)
by: Karouzos, Constantinos, et al.
Published: (2026)
An Empirical Study on Preference Tuning Generalization and Diversity Under Domain Shift
by: Karouzos, Constantinos, et al.
Published: (2026)
by: Karouzos, Constantinos, et al.
Published: (2026)
Boundary-targeted Membership Inference Attacks on Safety Classifiers
by: Hughes, Anthony, et al.
Published: (2026)
by: Hughes, Anthony, et al.
Published: (2026)
No Need to Talk: Asynchronous Mixture of Language Models
by: Filippova, Anastasiia, et al.
Published: (2024)
by: Filippova, Anastasiia, et al.
Published: (2024)
Privacy Evaluation Benchmarks for NLP Models
by: Huang, Wei, et al.
Published: (2024)
by: Huang, Wei, et al.
Published: (2024)
Still "Talking About Large Language Models": Some Clarifications
by: Shanahan, Murray
Published: (2024)
by: Shanahan, Murray
Published: (2024)
A Multi-Task Text Classification Pipeline with Natural Language Explanations: A User-Centric Evaluation in Sentiment Analysis and Offensive Language Identification in Greek Tweets
by: Mylonas, Nikolaos, et al.
Published: (2024)
by: Mylonas, Nikolaos, et al.
Published: (2024)
Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models
by: Villegas, Danae Sánchez, et al.
Published: (2026)
by: Villegas, Danae Sánchez, et al.
Published: (2026)
Classification of kinetic-related injury in hospital triage data using NLP
by: Shyam, Midhun, et al.
Published: (2025)
by: Shyam, Midhun, et al.
Published: (2025)
A Closer Look at Classification Evaluation Metrics and a Critical Reflection of Common Evaluation Practice
by: Opitz, Juri
Published: (2024)
by: Opitz, Juri
Published: (2024)
What Large Language Models Do Not Talk About: An Empirical Study of Moderation and Censorship Practices
by: Noels, Sander, et al.
Published: (2025)
by: Noels, Sander, et al.
Published: (2025)
NLP-ADBench: NLP Anomaly Detection Benchmark
by: Li, Yuangang, et al.
Published: (2024)
by: Li, Yuangang, et al.
Published: (2024)
Do You Know What You Are Talking About? Characterizing Query-Knowledge Relevance For Reliable Retrieval Augmented Generation
by: Li, Zhuohang, et al.
Published: (2024)
by: Li, Zhuohang, et al.
Published: (2024)
One Model is All You Need: ByT5-Sanskrit, a Unified Model for Sanskrit NLP Tasks
by: Nehrdich, Sebastian, et al.
Published: (2024)
by: Nehrdich, Sebastian, et al.
Published: (2024)
A Small Claims Court for the NLP: Judging Legal Text Classification Strategies With Small Datasets
by: Noguti, Mariana Yukari, et al.
Published: (2024)
by: Noguti, Mariana Yukari, et al.
Published: (2024)
Revisiting Hierarchical Text Classification: Inference and Metrics
by: Plaud, Roman, et al.
Published: (2024)
by: Plaud, Roman, et al.
Published: (2024)
How Far Are We From AGI: Are LLMs All We Need?
by: Feng, Tao, et al.
Published: (2024)
by: Feng, Tao, et al.
Published: (2024)
How Can We Effectively Expand the Vocabulary of LLMs with 0.01GB of Target Language Text?
by: Yamaguchi, Atsuki, et al.
Published: (2024)
by: Yamaguchi, Atsuki, et al.
Published: (2024)
Logits are All We Need to Adapt Closed Models
by: Hiranandani, Gaurush, et al.
Published: (2025)
by: Hiranandani, Gaurush, et al.
Published: (2025)
Do We Need Frontier Models to Verify Mathematical Proofs?
by: Naik, Aaditya, et al.
Published: (2026)
by: Naik, Aaditya, et al.
Published: (2026)
Vocabulary-level Memory Efficiency for Language Model Fine-tuning
by: Williams, Miles, et al.
Published: (2023)
by: Williams, Miles, et al.
Published: (2023)
On the Impact of Calibration Data in Post-training Quantization and Pruning
by: Williams, Miles, et al.
Published: (2023)
by: Williams, Miles, et al.
Published: (2023)
Automated ICD Classification of Psychiatric Diagnoses: From Classical NLP to Large Language Models
by: Ortega, Fernando, et al.
Published: (2026)
by: Ortega, Fernando, et al.
Published: (2026)
Enhancing Traffic Accident Classifications: Application of NLP Methods for City Safety
by: Özeren, Enes, et al.
Published: (2025)
by: Özeren, Enes, et al.
Published: (2025)
Rethinking LLM Evaluation: Can We Evaluate LLMs with 200x Less Data?
by: Wang, Shaobo, et al.
Published: (2025)
by: Wang, Shaobo, et al.
Published: (2025)
The NLP Task Effectiveness of Long-Range Transformers
by: Qin, Guanghui, et al.
Published: (2022)
by: Qin, Guanghui, et al.
Published: (2022)
Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language Models
by: Kumar, Abhishek, et al.
Published: (2024)
by: Kumar, Abhishek, et al.
Published: (2024)
What We Talk About When We Talk About LMs: Implicit Paradigm Shifts and the Ship of Language Models
by: Zhu, Shengqi, et al.
Published: (2024)
by: Zhu, Shengqi, et al.
Published: (2024)
ConvNLP: Image-based AI Text Detection
by: Jambunathan, Suriya Prakash, et al.
Published: (2024)
by: Jambunathan, Suriya Prakash, et al.
Published: (2024)
On Importance of Pruning and Distillation for Efficient Low Resource NLP
by: Mirashi, Aishwarya, et al.
Published: (2024)
by: Mirashi, Aishwarya, et al.
Published: (2024)
Bridging the Gap: From Ad-hoc to Proactive Search in Conversations
by: Meng, Chuan, et al.
Published: (2025)
by: Meng, Chuan, et al.
Published: (2025)
Towards Explainable Evaluation Metrics for Machine Translation
by: Leiter, Christoph, et al.
Published: (2023)
by: Leiter, Christoph, et al.
Published: (2023)
Can We Trust the Performance Evaluation of Uncertainty Estimation Methods in Text Summarization?
by: He, Jianfeng, et al.
Published: (2024)
by: He, Jianfeng, et al.
Published: (2024)
Classification of User Reports for Detection of Faulty Computer Components using NLP Models: A Case Study
by: Silva, Maria de Lourdes M., et al.
Published: (2025)
by: Silva, Maria de Lourdes M., et al.
Published: (2025)
Towards Efficient Active Learning in NLP via Pretrained Representations
by: Vysogorets, Artem, et al.
Published: (2024)
by: Vysogorets, Artem, et al.
Published: (2024)
TiME: Tiny Monolingual Encoders for Efficient NLP Pipelines
by: Schulmeister, David, et al.
Published: (2025)
by: Schulmeister, David, et al.
Published: (2025)
Boosting classification reliability of NLP transformer models in the long run
by: Kmetty, Zoltán, et al.
Published: (2023)
by: Kmetty, Zoltán, et al.
Published: (2023)
Similar Items
-
Incorporating Attribution Importance for Improving Faithfulness Metrics
by: Zhao, Zhixue, et al.
Published: (2023) -
Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
by: Xue, Huiyin, et al.
Published: (2025) -
Improving Multimodal Classification of Social Media Posts by Leveraging Image-Text Auxiliary Tasks
by: Villegas, Danae Sánchez, et al.
Published: (2023) -
Where does output diversity collapse in post-training?
by: Karouzos, Constantinos, et al.
Published: (2026) -
An Empirical Study on Preference Tuning Generalization and Diversity Under Domain Shift
by: Karouzos, Constantinos, et al.
Published: (2026)