Evaluating Webcam-based Gaze Data as an Alternative for Human Rationale Annotations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Brandl, Stephanie, Eberle, Oliver, Ribeiro, Tiago, Søgaard, Anders, Hollenstein, Nora |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WebQAmGaze: A Multilingual Webcam Eye-Tracking-While-Reading Dataset
von: Ribeiro, Tiago, et al.
Veröffentlicht: (2023)
von: Ribeiro, Tiago, et al.
Veröffentlicht: (2023)
Trick or Neat: Adversarial Ambiguity and Language Model Evaluation
von: Karamolegkou, Antonia, et al.
Veröffentlicht: (2025)
von: Karamolegkou, Antonia, et al.
Veröffentlicht: (2025)
Evaluation Revisited: A Taxonomy of Evaluation Concerns in Natural Language Processing
von: Dhar, Ruchira, et al.
Veröffentlicht: (2026)
von: Dhar, Ruchira, et al.
Veröffentlicht: (2026)
Evaluating the Limits of Large Language Models in Multilingual Legal Reasoning
von: Ioannou, Antreas, et al.
Veröffentlicht: (2025)
von: Ioannou, Antreas, et al.
Veröffentlicht: (2025)
Concept Space Alignment in Multilingual LLMs
von: Peng, Qiwei, et al.
Veröffentlicht: (2024)
von: Peng, Qiwei, et al.
Veröffentlicht: (2024)
Revisiting the Othello World Model Hypothesis
von: Yuan, Yifei, et al.
Veröffentlicht: (2025)
von: Yuan, Yifei, et al.
Veröffentlicht: (2025)
Beyond Technocratic XAI: The Who, What & How in Explanation Design
von: Dhar, Ruchira, et al.
Veröffentlicht: (2025)
von: Dhar, Ruchira, et al.
Veröffentlicht: (2025)
Evaluating Adjective-Noun Compositionality in LLMs: Functional vs Representational Perspectives
von: Dhar, Ruchira, et al.
Veröffentlicht: (2026)
von: Dhar, Ruchira, et al.
Veröffentlicht: (2026)
Llama meets EU: Investigating the European Political Spectrum through the Lens of LLMs
von: Chalkidis, Ilias, et al.
Veröffentlicht: (2024)
von: Chalkidis, Ilias, et al.
Veröffentlicht: (2024)
Factual Consistency of Multilingual Pretrained Language Models
von: Fierro, Constanza, et al.
Veröffentlicht: (2022)
von: Fierro, Constanza, et al.
Veröffentlicht: (2022)
Comprehensive Reassessment of Large-Scale Evaluation Outcomes in LLMs: A Multifaceted Statistical Approach
von: Sun, Kun, et al.
Veröffentlicht: (2024)
von: Sun, Kun, et al.
Veröffentlicht: (2024)
From Words to Worlds: Compositionality for Cognitive Architectures
von: Dhar, Ruchira, et al.
Veröffentlicht: (2024)
von: Dhar, Ruchira, et al.
Veröffentlicht: (2024)
Rethinking Human Preference Evaluation of LLM Rationales
von: Li, Ziang, et al.
Veröffentlicht: (2025)
von: Li, Ziang, et al.
Veröffentlicht: (2025)
Explaining Text Similarity in Transformer Models
von: Vasileiou, Alexandros, et al.
Veröffentlicht: (2024)
von: Vasileiou, Alexandros, et al.
Veröffentlicht: (2024)
Does Instruction Tuning Make LLMs More Consistent?
von: Fierro, Constanza, et al.
Veröffentlicht: (2024)
von: Fierro, Constanza, et al.
Veröffentlicht: (2024)
Evaluating Human Alignment and Model Faithfulness of LLM Rationale
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2024)
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2024)
Cat, Rat, Meow: On the Alignment of Language Model and Human Term-Similarity Judgments
von: Linhardt, Lorenz, et al.
Veröffentlicht: (2025)
von: Linhardt, Lorenz, et al.
Veröffentlicht: (2025)
Identifying Fine-grained Forms of Populism in Political Discourse: A Case Study on Donald Trump's Presidential Campaigns
von: Chalkidis, Ilias, et al.
Veröffentlicht: (2025)
von: Chalkidis, Ilias, et al.
Veröffentlicht: (2025)
EvalCards: A Framework for Standardized Evaluation Reporting
von: Dhar, Ruchira, et al.
Veröffentlicht: (2025)
von: Dhar, Ruchira, et al.
Veröffentlicht: (2025)
Fine-Grained Perspectives: Modeling Explanations with Annotator-Specific Rationales
von: Sarumi, Olufunke O., et al.
Veröffentlicht: (2026)
von: Sarumi, Olufunke O., et al.
Veröffentlicht: (2026)
How Do Multilingual Language Models Remember Facts?
von: Fierro, Constanza, et al.
Veröffentlicht: (2024)
von: Fierro, Constanza, et al.
Veröffentlicht: (2024)
Unlocking Markets: A Multilingual Benchmark to Cross-Market Question Answering
von: Yuan, Yifei, et al.
Veröffentlicht: (2024)
von: Yuan, Yifei, et al.
Veröffentlicht: (2024)
Rationale-based Opinion Summarization
von: Li, Haoyuan, et al.
Veröffentlicht: (2024)
von: Li, Haoyuan, et al.
Veröffentlicht: (2024)
Understanding Subword Compositionality of Large Language Models
von: Peng, Qiwei, et al.
Veröffentlicht: (2025)
von: Peng, Qiwei, et al.
Veröffentlicht: (2025)
RORA: Robust Free-Text Rationale Evaluation
von: Jiang, Zhengping, et al.
Veröffentlicht: (2024)
von: Jiang, Zhengping, et al.
Veröffentlicht: (2024)
BiasLab: Toward Explainable Political Bias Detection with Dual-Axis Annotations and Rationale Indicators
von: Solaiman, Kma
Veröffentlicht: (2025)
von: Solaiman, Kma
Veröffentlicht: (2025)
Enhancing Text Annotation through Rationale-Driven Collaborative Few-Shot Prompting
von: Wu, Jianfei, et al.
Veröffentlicht: (2024)
von: Wu, Jianfei, et al.
Veröffentlicht: (2024)
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
von: Calderon, Nitay, et al.
Veröffentlicht: (2025)
von: Calderon, Nitay, et al.
Veröffentlicht: (2025)
Word Order and World Knowledge
von: Zhao, Qinghua, et al.
Veröffentlicht: (2024)
von: Zhao, Qinghua, et al.
Veröffentlicht: (2024)
A Discriminative Latent-Variable Model for Bilingual Lexicon Induction
von: Ruder, Sebastian, et al.
Veröffentlicht: (2018)
von: Ruder, Sebastian, et al.
Veröffentlicht: (2018)
Rationale-Aware Answer Verification by Pairwise Self-Evaluation
von: Kawabata, Akira, et al.
Veröffentlicht: (2024)
von: Kawabata, Akira, et al.
Veröffentlicht: (2024)
Defining Knowledge: Bridging Epistemology and Large Language Models
von: Fierro, Constanza, et al.
Veröffentlicht: (2024)
von: Fierro, Constanza, et al.
Veröffentlicht: (2024)
MuLan: A Study of Fact Mutability in Language Models
von: Fierro, Constanza, et al.
Veröffentlicht: (2024)
von: Fierro, Constanza, et al.
Veröffentlicht: (2024)
Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability
von: Zhu, Chiwei, et al.
Veröffentlicht: (2025)
von: Zhu, Chiwei, et al.
Veröffentlicht: (2025)
Debiasing Multilingual LLMs in Cross-lingual Latent Space
von: Peng, Qiwei, et al.
Veröffentlicht: (2025)
von: Peng, Qiwei, et al.
Veröffentlicht: (2025)
Disagreeing Rationales: Rethinking Classification and Explainability Evaluation in Hate Speech Detection
von: Muscato, Benedetta, et al.
Veröffentlicht: (2026)
von: Muscato, Benedetta, et al.
Veröffentlicht: (2026)
LLMs as Data Annotators: How Close Are We to Human Performance
von: Haq, Muhammad Uzair Ul, et al.
Veröffentlicht: (2025)
von: Haq, Muhammad Uzair Ul, et al.
Veröffentlicht: (2025)
Reasoning Pattern Matters: Learning to Reason without Human Rationales
von: Pang, Chaoxu, et al.
Veröffentlicht: (2025)
von: Pang, Chaoxu, et al.
Veröffentlicht: (2025)
ReasonScaffold: A Scaffolded Reasoning-based Annotation Protocol for Human-AI Co-Annotation
von: Sudheendra, Smitha Muthya, et al.
Veröffentlicht: (2026)
von: Sudheendra, Smitha Muthya, et al.
Veröffentlicht: (2026)
CoAnnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data Annotation
von: Li, Minzhi, et al.
Veröffentlicht: (2023)
von: Li, Minzhi, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
WebQAmGaze: A Multilingual Webcam Eye-Tracking-While-Reading Dataset
von: Ribeiro, Tiago, et al.
Veröffentlicht: (2023) -
Trick or Neat: Adversarial Ambiguity and Language Model Evaluation
von: Karamolegkou, Antonia, et al.
Veröffentlicht: (2025) -
Evaluation Revisited: A Taxonomy of Evaluation Concerns in Natural Language Processing
von: Dhar, Ruchira, et al.
Veröffentlicht: (2026) -
Evaluating the Limits of Large Language Models in Multilingual Legal Reasoning
von: Ioannou, Antreas, et al.
Veröffentlicht: (2025) -
Concept Space Alignment in Multilingual LLMs
von: Peng, Qiwei, et al.
Veröffentlicht: (2024)