SEER: The Span-based Emotion Evidence Retrieval Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Sampath, Aneesha, Aran, Oya, Provost, Emily Mower |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Finetuning for Dimensional Speech Emotion Recognition in the Age of Transformers
by: Sampath, Aneesha, et al.
Published: (2025)
by: Sampath, Aneesha, et al.
Published: (2025)
What You Feel Is Not What They See: On Predicting Self-Reported Emotion from Third-Party Observer Labels
by: El-Tawil, Yara, et al.
Published: (2026)
by: El-Tawil, Yara, et al.
Published: (2026)
Beyond Binary: Multiclass Paraphasia Detection with Generative Pretrained Transformers and End-to-End Models
by: Perez, Matthew, et al.
Published: (2024)
by: Perez, Matthew, et al.
Published: (2024)
Contrastive Distillation of Emotion Knowledge from LLMs for Zero-Shot Emotion Recognition
by: Niu, Minxue, et al.
Published: (2025)
by: Niu, Minxue, et al.
Published: (2025)
From Text to Emotion: Unveiling the Emotion Annotation Capabilities of LLMs
by: Niu, Minxue, et al.
Published: (2024)
by: Niu, Minxue, et al.
Published: (2024)
Rethinking Emotion Annotations in the Era of Large Language Models
by: Niu, Minxue, et al.
Published: (2024)
by: Niu, Minxue, et al.
Published: (2024)
Span-level Emotion-Cause-Category Triplet Extraction with Instruction Tuning LLMs and Data Augmentation
by: Li, Xiangju, et al.
Published: (2025)
by: Li, Xiangju, et al.
Published: (2025)
From Documents to Spans: Scalable Supervision for Evidence-Based ICD Coding with LLMs
by: Zhang, Xu, et al.
Published: (2026)
by: Zhang, Xu, et al.
Published: (2026)
BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity
by: Diddee, Harshita, et al.
Published: (2026)
by: Diddee, Harshita, et al.
Published: (2026)
Responsible AI in NLP: GUS-Net Span-Level Bias Detection Dataset and Benchmark for Generalizations, Unfairness, and Stereotypes
by: Powers, Maximus, et al.
Published: (2024)
by: Powers, Maximus, et al.
Published: (2024)
Retrieval-Augmented Generation with Conflicting Evidence
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
Cost-efficient Crowdsourcing for Span-based Sequence Labeling: Worker Selection and Data Augmentation
by: Wang, Yujie, et al.
Published: (2023)
by: Wang, Yujie, et al.
Published: (2023)
Benchmarking Retrieval-Augmented Generation for Medicine
by: Xiong, Guangzhi, et al.
Published: (2024)
by: Xiong, Guangzhi, et al.
Published: (2024)
Scaling Bidirectional Spans and Span Violations in Attention Mechanism
by: Kim, Jongwook, et al.
Published: (2025)
by: Kim, Jongwook, et al.
Published: (2025)
Emotion Transcription in Conversation: A Benchmark for Capturing Subtle and Complex Emotional States through Natural Language
by: Tanaka, Yoshiki, et al.
Published: (2026)
by: Tanaka, Yoshiki, et al.
Published: (2026)
Span-Level Machine Translation Meta-Evaluation
by: Perrella, Stefano, et al.
Published: (2026)
by: Perrella, Stefano, et al.
Published: (2026)
MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments
by: Wang, Han, et al.
Published: (2026)
by: Wang, Han, et al.
Published: (2026)
Emotion and Intent Joint Understanding in Multimodal Conversation: A Benchmarking Dataset
by: Liu, Rui, et al.
Published: (2024)
by: Liu, Rui, et al.
Published: (2024)
SER Evals: In-domain and Out-of-domain Benchmarking for Speech Emotion Recognition
by: Osman, Mohamed, et al.
Published: (2024)
by: Osman, Mohamed, et al.
Published: (2024)
EmotionQueen: A Benchmark for Evaluating Empathy of Large Language Models
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
ImpliRet: Benchmarking the Implicit Fact Retrieval Challenge
by: Taghavi, Zeinab Sadat, et al.
Published: (2025)
by: Taghavi, Zeinab Sadat, et al.
Published: (2025)
MRAG: Benchmarking Retrieval-Augmented Generation for Bio-medicine
by: Li, Liz, et al.
Published: (2026)
by: Li, Liz, et al.
Published: (2026)
RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems
by: Friel, Robert, et al.
Published: (2024)
by: Friel, Robert, et al.
Published: (2024)
keepitsimple at SemEval-2025 Task 3: LLM-Uncertainty based Approach for Multilingual Hallucination Span Detection
by: Vemula, Saketh Reddy, et al.
Published: (2025)
by: Vemula, Saketh Reddy, et al.
Published: (2025)
Stateful Evidence-Driven Retrieval-Augmented Generation with Iterative Reasoning
by: Dong, Qi, et al.
Published: (2026)
by: Dong, Qi, et al.
Published: (2026)
Evidence-Driven Retrieval Augmented Response Generation for Online Misinformation
by: Yue, Zhenrui, et al.
Published: (2024)
by: Yue, Zhenrui, et al.
Published: (2024)
Prompt-Counterfactual Explanations for Generative AI System Behavior
by: Goethals, Sofie, et al.
Published: (2026)
by: Goethals, Sofie, et al.
Published: (2026)
EmoBench-M: Benchmarking Emotional Intelligence for Multimodal Large Language Models
by: Hu, He, et al.
Published: (2025)
by: Hu, He, et al.
Published: (2025)
From Joy to Fear: A Benchmark of Emotion Estimation in Pop Song Lyrics
by: Dahary, Shay, et al.
Published: (2025)
by: Dahary, Shay, et al.
Published: (2025)
HEART: A Unified Benchmark for Assessing Humans and LLMs in Emotional Support Dialogue
by: Iyer, Laya, et al.
Published: (2026)
by: Iyer, Laya, et al.
Published: (2026)
The Role of Syntactic Span Preferences in Post-Hoc Explanation Disagreement
by: Kamp, Jonathan, et al.
Published: (2024)
by: Kamp, Jonathan, et al.
Published: (2024)
Spotting AI's Touch: Identifying LLM-Paraphrased Spans in Text
by: Li, Yafu, et al.
Published: (2024)
by: Li, Yafu, et al.
Published: (2024)
CURE: A Multimodal Benchmark for Clinical Understanding and Retrieval Evaluation
by: Gu, Yannian, et al.
Published: (2026)
by: Gu, Yannian, et al.
Published: (2026)
Empaths at SemEval-2025 Task 11: Retrieval-Augmented Approach to Perceived Emotions Prediction
by: Morozov, Lev, et al.
Published: (2025)
by: Morozov, Lev, et al.
Published: (2025)
Evidence from fMRI Supports a Two-Phase Abstraction Process in Language Models
by: Cheng, Emily, et al.
Published: (2024)
by: Cheng, Emily, et al.
Published: (2024)
ES-MemEval: Benchmarking Conversational Agents on Personalized Long-Term Emotional Support
by: Chen, Tiantian, et al.
Published: (2026)
by: Chen, Tiantian, et al.
Published: (2026)
Something Just Like TRuST : Toxicity Recognition of Span and Target
by: Atil, Berk, et al.
Published: (2025)
by: Atil, Berk, et al.
Published: (2025)
Exploring the Performance of Large Language Models on Subjective Span Identification Tasks
by: Dmonte, Alphaeus, et al.
Published: (2026)
by: Dmonte, Alphaeus, et al.
Published: (2026)
A Dataset of Open-Domain Question Answering with Multiple-Span Answers
by: Luo, Zhiyi, et al.
Published: (2024)
by: Luo, Zhiyi, et al.
Published: (2024)
Self-Selected Attention Span for Accelerating Large Language Model Inference
by: Jin, Tian, et al.
Published: (2024)
by: Jin, Tian, et al.
Published: (2024)
Similar Items
-
Efficient Finetuning for Dimensional Speech Emotion Recognition in the Age of Transformers
by: Sampath, Aneesha, et al.
Published: (2025) -
What You Feel Is Not What They See: On Predicting Self-Reported Emotion from Third-Party Observer Labels
by: El-Tawil, Yara, et al.
Published: (2026) -
Beyond Binary: Multiclass Paraphasia Detection with Generative Pretrained Transformers and End-to-End Models
by: Perez, Matthew, et al.
Published: (2024) -
Contrastive Distillation of Emotion Knowledge from LLMs for Zero-Shot Emotion Recognition
by: Niu, Minxue, et al.
Published: (2025) -
From Text to Emotion: Unveiling the Emotion Annotation Capabilities of LLMs
by: Niu, Minxue, et al.
Published: (2024)