ECBD: Evidence-Centered Benchmark Design for NLP
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Liu, Yu Lu, Blodgett, Su Lin, Cheung, Jackie Chi Kit, Liao, Q. Vera, Olteanu, Alexandra, Xiao, Ziang |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Challenges to Evaluating the Generalization of Coreference Resolution Models: A Measurement Modeling Perspective
par: Porada, Ian, et autres
Publié: (2023)
par: Porada, Ian, et autres
Publié: (2023)
"One-Size-Fits-All"? Examining Expectations around What Constitute "Fair" or "Good" NLG System Behaviors
par: Lucy, Li, et autres
Publié: (2023)
par: Lucy, Li, et autres
Publié: (2023)
Solving the Challenge Set without Solving the Task: On Winograd Schemas as a Test of Pronominal Coreference Resolution
par: Porada, Ian, et autres
Publié: (2024)
par: Porada, Ian, et autres
Publié: (2024)
"It was 80% me, 20% AI": Seeking Authenticity in Co-Writing with Large Language Models
par: Hwang, Angel Hsing-Chi, et autres
Publié: (2024)
par: Hwang, Angel Hsing-Chi, et autres
Publié: (2024)
From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing Assistants
par: Rismani, Shalaleh, et autres
Publié: (2026)
par: Rismani, Shalaleh, et autres
Publié: (2026)
"I Am the One and Only, Your Cyber BFF": Understanding the Impact of GenAI Requires Understanding the Impact of Anthropomorphic AI
par: Cheng, Myra, et autres
Publié: (2024)
par: Cheng, Myra, et autres
Publié: (2024)
A Taxonomy of Linguistic Expressions That Contribute To Anthropomorphism of Language Technologies
par: DeVrio, Alicia, et autres
Publié: (2025)
par: DeVrio, Alicia, et autres
Publié: (2025)
Dehumanizing Machines: Mitigating Anthropomorphic Behaviors in Text Generation Systems
par: Cheng, Myra, et autres
Publié: (2025)
par: Cheng, Myra, et autres
Publié: (2025)
Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems
par: Harvey, Emma, et autres
Publié: (2025)
par: Harvey, Emma, et autres
Publié: (2025)
A Controlled Reevaluation of Coreference Resolution Models
par: Porada, Ian, et autres
Publié: (2024)
par: Porada, Ian, et autres
Publié: (2024)
PreSumm: Predicting Summarization Performance Without Summarizing
par: Koniaev, Steven, et autres
Publié: (2025)
par: Koniaev, Steven, et autres
Publié: (2025)
Mechanistic Understanding and Mitigation of Language Model Non-Factual Hallucinations
par: Yu, Lei, et autres
Publié: (2024)
par: Yu, Lei, et autres
Publié: (2024)
Improving the Calibration of Confidence Scores in Text Generation Using the Output Distribution's Characteristics
par: Flores, Lorenzo Jaime Yu, et autres
Publié: (2025)
par: Flores, Lorenzo Jaime Yu, et autres
Publié: (2025)
Generative Echo Chamber? Effects of LLM-Powered Search Systems on Diverse Information Seeking
par: Sharma, Nikhil, et autres
Publié: (2024)
par: Sharma, Nikhil, et autres
Publié: (2024)
Can Vision Language Models Be Adaptive in Mathematics Education? A Learner Model-based Rubric Study
par: Gao, Jie, et autres
Publié: (2026)
par: Gao, Jie, et autres
Publié: (2026)
Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning
par: Flores, Lorenzo Jaime Yu, et autres
Publié: (2026)
par: Flores, Lorenzo Jaime Yu, et autres
Publié: (2026)
Stochastic Chameleons: Irrelevant Context Hallucinations Reveal Class-Based (Mis)Generalization in LLMs
par: Cheng, Ziling, et autres
Publié: (2025)
par: Cheng, Ziling, et autres
Publié: (2025)
$\texttt{COSMIC}$: Mutual Information for Task-Agnostic Summarization Evaluation
par: Darrin, Maxime, et autres
Publié: (2024)
par: Darrin, Maxime, et autres
Publié: (2024)
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
par: Chehbouni, Khaoula, et autres
Publié: (2025)
par: Chehbouni, Khaoula, et autres
Publié: (2025)
Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation
par: Cheng, Ziling, et autres
Publié: (2025)
par: Cheng, Ziling, et autres
Publié: (2025)
$(RSA)^2$: A Rhetorical-Strategy-Aware Rational Speech Act Framework for Figurative Language Understanding
par: Piano, Cesare Spinoso-Di, et autres
Publié: (2025)
par: Piano, Cesare Spinoso-Di, et autres
Publié: (2025)
Testing the Assumptions of Active Learning for Translation Tasks with Few Samples
par: Flores, Lorenzo Jaime Yu, et autres
Publié: (2026)
par: Flores, Lorenzo Jaime Yu, et autres
Publié: (2026)
CItruS: Chunked Instruction-aware State Eviction for Long Sequence Modeling
par: Bai, Yu, et autres
Publié: (2024)
par: Bai, Yu, et autres
Publié: (2024)
NLP-ADBench: NLP Anomaly Detection Benchmark
par: Li, Yuangang, et autres
Publié: (2024)
par: Li, Yuangang, et autres
Publié: (2024)
Identifying and Analyzing Performance-Critical Tokens in Large Language Models
par: Bai, Yu, et autres
Publié: (2024)
par: Bai, Yu, et autres
Publié: (2024)
Error Diversity Matters: An Error-Resistant Ensemble Method for Unsupervised Dependency Parsing
par: Shayegh, Behzad, et autres
Publié: (2024)
par: Shayegh, Behzad, et autres
Publié: (2024)
The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels
par: Fleisig, Eve, et autres
Publié: (2024)
par: Fleisig, Eve, et autres
Publié: (2024)
Rethinking Model Evaluation as Narrowing the Socio-Technical Gap
par: Liao, Q. Vera, et autres
Publié: (2023)
par: Liao, Q. Vera, et autres
Publié: (2023)
AI Automatons: AI Systems Intended to Imitate Humans
par: Olteanu, Alexandra, et autres
Publié: (2025)
par: Olteanu, Alexandra, et autres
Publié: (2025)
Synthetic Users, Real Differences: an Evaluation Framework for User Simulation in Multi-Turn Conversations
par: Liu, Yu Lu, et autres
Publié: (2026)
par: Liu, Yu Lu, et autres
Publié: (2026)
Gaps Between Research and Practice When Measuring Representational Harms Caused by LLM-Based Systems
par: Harvey, Emma, et autres
Publié: (2024)
par: Harvey, Emma, et autres
Publié: (2024)
GLIMPSE: Pragmatically Informative Multi-Document Summarization for Scholarly Reviews
par: Darrin, Maxime, et autres
Publié: (2024)
par: Darrin, Maxime, et autres
Publié: (2024)
Modeling Orthographic Variation Improves NLP Performance for Nigerian Pidgin
par: Lin, Pin-Jie, et autres
Publié: (2024)
par: Lin, Pin-Jie, et autres
Publié: (2024)
Safety Alignment in NLP Tasks: Weakly Aligned Summarization as an In-Context Attack
par: Fu, Yu, et autres
Publié: (2023)
par: Fu, Yu, et autres
Publié: (2023)
Privacy Evaluation Benchmarks for NLP Models
par: Huang, Wei, et autres
Publié: (2024)
par: Huang, Wei, et autres
Publié: (2024)
Exploring NLP Benchmarks in an Extremely Low-Resource Setting
par: Nuha, Ulin, et autres
Publié: (2025)
par: Nuha, Ulin, et autres
Publié: (2025)
Benchmarking Retrieval-Augmented Large Language Models in Biomedical NLP: Application, Robustness, and Self-Awareness
par: Li, Mingchen, et autres
Publié: (2024)
par: Li, Mingchen, et autres
Publié: (2024)
Designing NLP Systems That Adapt to Diverse Worldviews
par: Creanga, Claudiu, et autres
Publié: (2024)
par: Creanga, Claudiu, et autres
Publié: (2024)
Collaborative Rational Speech Act: Pragmatic Reasoning for Multi-Turn Dialog
par: Estienne, Lautaro, et autres
Publié: (2025)
par: Estienne, Lautaro, et autres
Publié: (2025)
Benchmarking Large Language Models on Multiple Tasks in Bioinformatics NLP with Prompting
par: Jiang, Jiyue, et autres
Publié: (2025)
par: Jiang, Jiyue, et autres
Publié: (2025)
Documents similaires
-
Challenges to Evaluating the Generalization of Coreference Resolution Models: A Measurement Modeling Perspective
par: Porada, Ian, et autres
Publié: (2023) -
"One-Size-Fits-All"? Examining Expectations around What Constitute "Fair" or "Good" NLG System Behaviors
par: Lucy, Li, et autres
Publié: (2023) -
Solving the Challenge Set without Solving the Task: On Winograd Schemas as a Test of Pronominal Coreference Resolution
par: Porada, Ian, et autres
Publié: (2024) -
"It was 80% me, 20% AI": Seeking Authenticity in Co-Writing with Large Language Models
par: Hwang, Angel Hsing-Chi, et autres
Publié: (2024) -
From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing Assistants
par: Rismani, Shalaleh, et autres
Publié: (2026)