Semantic F1 Scores: Fair Evaluation Under Fuzzy Class Boundaries
Fuente:
arXiv
Saved in:
| Main Authors: | Chochlakis, Georgios, Trager, Jackson, Jhaveri, Vedant, Ravichandran, Nikhil, Potamianos, Alexandros, Narayanan, Shrikanth |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Aggregation Artifacts in Subjective Tasks Collapse Large Language Models' Posteriors
by: Chochlakis, Georgios, et al.
Published: (2024)
by: Chochlakis, Georgios, et al.
Published: (2024)
The Strong Pull of Prior Knowledge in Large Language Models and Its Impact on Emotion Recognition
by: Chochlakis, Georgios, et al.
Published: (2024)
by: Chochlakis, Georgios, et al.
Published: (2024)
Larger Language Models Don't Care How You Think: Why Chain-of-Thought Prompting Fails in Subjective Tasks
by: Chochlakis, Georgios, et al.
Published: (2024)
by: Chochlakis, Georgios, et al.
Published: (2024)
Authors Should Label Their Own Documents
by: Ma, Marcus, et al.
Published: (2025)
by: Ma, Marcus, et al.
Published: (2025)
Intelligence Requires Grounding But Not Embodiment
by: Ma, Marcus, et al.
Published: (2026)
by: Ma, Marcus, et al.
Published: (2026)
EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing
by: Sioros, Vassilis, et al.
Published: (2025)
by: Sioros, Vassilis, et al.
Published: (2025)
GF-Score: Certified Class-Conditional Robustness Evaluation with Fairness Guarantees
by: Shah, Arya, et al.
Published: (2026)
by: Shah, Arya, et al.
Published: (2026)
DeepMLF: Multimodal language model with learnable tokens for deep fusion in sentiment analysis
by: Georgiou, Efthymios, et al.
Published: (2025)
by: Georgiou, Efthymios, et al.
Published: (2025)
How to Retrieve Examples in In-context Learning to Improve Conversational Emotion Recognition using Large Language Models?
by: Wang, Mengqi, et al.
Published: (2025)
by: Wang, Mengqi, et al.
Published: (2025)
CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning
by: Kanatas, Angelos-Nikolaos, et al.
Published: (2025)
by: Kanatas, Angelos-Nikolaos, et al.
Published: (2025)
Community-Informed AI Models for Police Accountability
by: Graham, Benjamin A. T., et al.
Published: (2024)
by: Graham, Benjamin A. T., et al.
Published: (2024)
Theory Trace Card: Theory-Driven Socio-Cognitive Evaluation of LLMs
by: Karimi-Malekabadi, Farzan, et al.
Published: (2026)
by: Karimi-Malekabadi, Farzan, et al.
Published: (2026)
Enhancing Fast Feed Forward Networks with Load Balancing and a Master Leaf Node
by: Charalampopoulos, Andreas, et al.
Published: (2024)
by: Charalampopoulos, Andreas, et al.
Published: (2024)
Large Language Models Do Multi-Label Classification Differently
by: Ma, Marcus, et al.
Published: (2025)
by: Ma, Marcus, et al.
Published: (2025)
Online Fair Division for Personalized $2$-Value Instances
by: Amanatidis, Georgios, et al.
Published: (2025)
by: Amanatidis, Georgios, et al.
Published: (2025)
ModalityMirror: Improving Audio Classification in Modality Heterogeneity Federated Learning with Multimodal Distillation
by: Feng, Tiantian, et al.
Published: (2024)
by: Feng, Tiantian, et al.
Published: (2024)
Knowledge-guided EEG Representation Learning
by: Kommineni, Aditya, et al.
Published: (2024)
by: Kommineni, Aditya, et al.
Published: (2024)
CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models
by: Padwal, Vedant
Published: (2026)
by: Padwal, Vedant
Published: (2026)
Is $F_1$ Score Suboptimal for Cybersecurity Models? Introducing $C_{score}$, a Cost-Aware Alternative for Model Assessment
by: Marwah, Manish, et al.
Published: (2024)
by: Marwah, Manish, et al.
Published: (2024)
Evaluating AI Group Fairness: a Fuzzy Logic Perspective
by: Krasanakis, Emmanouil, et al.
Published: (2024)
by: Krasanakis, Emmanouil, et al.
Published: (2024)
Humans Hallucinate Too: Language Models Identify and Correct Subjective Annotation Errors With Label-in-a-Haystack Prompts
by: Chochlakis, Georgios, et al.
Published: (2025)
by: Chochlakis, Georgios, et al.
Published: (2025)
Egocentric Speaker Classification in Child-Adult Dyadic Interactions: From Sensing to Computational Modeling
by: Feng, Tiantian, et al.
Published: (2024)
by: Feng, Tiantian, et al.
Published: (2024)
Learning-free L2-Accented Speech Generation using Phonological Rules
by: Lertpetchpun, Thanathai, et al.
Published: (2026)
by: Lertpetchpun, Thanathai, et al.
Published: (2026)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
by: Tang, Zeyu, et al.
Published: (2026)
by: Tang, Zeyu, et al.
Published: (2026)
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
by: Zhang, Hezhao, et al.
Published: (2026)
by: Zhang, Hezhao, et al.
Published: (2026)
Fuzzy Propositional Formulas under the Stable Model Semantics
by: Lee, Joohyung, et al.
Published: (2025)
by: Lee, Joohyung, et al.
Published: (2025)
ACCeLLiuM: Supervised Fine-Tuning for Automated OpenACC Pragma Generation
by: Jhaveri, Samyak, et al.
Published: (2025)
by: Jhaveri, Samyak, et al.
Published: (2025)
Bridging the Gap: Empowering Small Models in Reliable OpenACC-based Parallelization via GEPA-Optimized Prompting
by: Jhaveri, Samyak, et al.
Published: (2026)
by: Jhaveri, Samyak, et al.
Published: (2026)
Fairness in the Multi-Secretary Problem
by: Papasotiropoulos, Georgios, et al.
Published: (2025)
by: Papasotiropoulos, Georgios, et al.
Published: (2025)
Semantic Fusion with Fuzzy-Membership Features for Controllable Language Modelling
by: Huang, Yongchao, et al.
Published: (2025)
by: Huang, Yongchao, et al.
Published: (2025)
ConPro: Learning Severity Representation for Medical Images using Contrastive Learning and Preference Optimization
by: Nguyen, Hong, et al.
Published: (2024)
by: Nguyen, Hong, et al.
Published: (2024)
DARD: A Multi-Agent Approach for Task-Oriented Dialog Systems
by: Gupta, Aman, et al.
Published: (2024)
by: Gupta, Aman, et al.
Published: (2024)
On The Fairness Impacts of Hardware Selection in Machine Learning
by: Nelaturu, Sree Harsha, et al.
Published: (2023)
by: Nelaturu, Sree Harsha, et al.
Published: (2023)
Fuzzy Categorical Planning: Autonomous Goal Satisfaction with Graded Semantic Constraints
by: Qu, Shuhui
Published: (2026)
by: Qu, Shuhui
Published: (2026)
Can Layer-wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?
by: Sinha, Abhijit, et al.
Published: (2025)
by: Sinha, Abhijit, et al.
Published: (2025)
Can a Machine Distinguish High and Low Amount of Social Creak in Speech?
by: Laukkanen, Anne-Maria, et al.
Published: (2024)
by: Laukkanen, Anne-Maria, et al.
Published: (2024)
LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation
by: Raval, Ananya, et al.
Published: (2025)
by: Raval, Ananya, et al.
Published: (2025)
Position: Safety and Fairness in Agentic AI Depend on Interaction Topology, Not on Model Scale or Alignment
by: Bajaj, Tanav Singh, et al.
Published: (2026)
by: Bajaj, Tanav Singh, et al.
Published: (2026)
Standardized Interpretable Fairness Measures for Continuous Risk Scores
by: Becker, Ann-Kristin, et al.
Published: (2023)
by: Becker, Ann-Kristin, et al.
Published: (2023)
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
by: Pan, Zhiyuan, et al.
Published: (2025)
by: Pan, Zhiyuan, et al.
Published: (2025)
Similar Items
-
Aggregation Artifacts in Subjective Tasks Collapse Large Language Models' Posteriors
by: Chochlakis, Georgios, et al.
Published: (2024) -
The Strong Pull of Prior Knowledge in Large Language Models and Its Impact on Emotion Recognition
by: Chochlakis, Georgios, et al.
Published: (2024) -
Larger Language Models Don't Care How You Think: Why Chain-of-Thought Prompting Fails in Subjective Tasks
by: Chochlakis, Georgios, et al.
Published: (2024) -
Authors Should Label Their Own Documents
by: Ma, Marcus, et al.
Published: (2025) -
Intelligence Requires Grounding But Not Embodiment
by: Ma, Marcus, et al.
Published: (2026)