WMT24 Test Suite: Gender Resolution in Speaker-Listener Dialogue Roles
Fuente:
arXiv
Saved in:
| Main Authors: | Dawkins, Hillary, Nejadgholi, Isar, Lo, Chi-kiu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Gender-Neutral Machine Translation Strategies in Practice
by: Dawkins, Hillary, et al.
Published: (2025)
by: Dawkins, Hillary, et al.
Published: (2025)
Projective Methods for Mitigating Gender Bias in Pre-trained Language Models
by: Dawkins, Hillary, et al.
Published: (2024)
by: Dawkins, Hillary, et al.
Published: (2024)
Adaptable Moral Stances of Large Language Models on Sexist Content: Implications for Society and Gender Discourse
by: Guo, Rongchen, et al.
Published: (2024)
by: Guo, Rongchen, et al.
Published: (2024)
Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency
by: Fraser, Kathleen C., et al.
Published: (2025)
by: Fraser, Kathleen C., et al.
Published: (2025)
Challenging Negative Gender Stereotypes: A Study on the Effectiveness of Automated Counter-Stereotypes
by: Nejadgholi, Isar, et al.
Published: (2024)
by: Nejadgholi, Isar, et al.
Published: (2024)
WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects
by: Deutsch, Daniel, et al.
Published: (2025)
by: Deutsch, Daniel, et al.
Published: (2025)
Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory
by: Nejadgholi, Isar, et al.
Published: (2026)
by: Nejadgholi, Isar, et al.
Published: (2026)
LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories
by: Vishnubhotla, Krishnapriya, et al.
Published: (2026)
by: Vishnubhotla, Krishnapriya, et al.
Published: (2026)
Semantic Differentiation in Speech Emotion Recognition: Insights from Descriptive and Expressive Speech Roles
by: Guo, Rongchen, et al.
Published: (2025)
by: Guo, Rongchen, et al.
Published: (2025)
Socially Aware Synthetic Data Generation for Suicidal Ideation Detection Using Large Language Models
by: Ghanadian, Hamideh, et al.
Published: (2024)
by: Ghanadian, Hamideh, et al.
Published: (2024)
The crime of being poor
by: Curto, Georgina, et al.
Published: (2023)
by: Curto, Georgina, et al.
Published: (2023)
A Taxonomy for Design and Evaluation of Prompt-Based Natural Language Explanations
by: Nejadgholi, Isar, et al.
Published: (2025)
by: Nejadgholi, Isar, et al.
Published: (2025)
When Detection Fails: The Power of Fine-Tuned Models to Generate Human-Like Social Media Text
by: Dawkins, Hillary, et al.
Published: (2025)
by: Dawkins, Hillary, et al.
Published: (2025)
Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods
by: Fraser, Kathleen C., et al.
Published: (2024)
by: Fraser, Kathleen C., et al.
Published: (2024)
Preliminary WMT24 Ranking of General MT Systems and LLMs
by: Kocmi, Tom, et al.
Published: (2024)
by: Kocmi, Tom, et al.
Published: (2024)
Tackling Social Bias against the Poor: A Dataset and Taxonomy on Aporophobia
by: Curto, Georgina, et al.
Published: (2025)
by: Curto, Georgina, et al.
Published: (2025)
MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task
by: Juraska, Juraj, et al.
Published: (2024)
by: Juraska, Juraj, et al.
Published: (2024)
IKUN for WMT24 General MT Task: LLMs Are here for Multilingual Machine Translation
by: Liao, Baohao, et al.
Published: (2024)
by: Liao, Baohao, et al.
Published: (2024)
Expanding the WMT24++ Benchmark with Rumantsch Grischun, Sursilvan, Sutsilvan, Surmiran, Puter, and Vallader
by: Vamvas, Jannis, et al.
Published: (2025)
by: Vamvas, Jannis, et al.
Published: (2025)
GenderBench: Evaluation Suite for Gender Biases in LLMs
by: Pikuliak, Matúš
Published: (2025)
by: Pikuliak, Matúš
Published: (2025)
Communicating with Speakers and Listeners of Different Pragmatic Levels
by: Naszadi, Kata, et al.
Published: (2024)
by: Naszadi, Kata, et al.
Published: (2024)
Findings of the WMT 2024 Shared Task on Chat Translation
by: Mohammed, Wafaa, et al.
Published: (2024)
by: Mohammed, Wafaa, et al.
Published: (2024)
Listen Then See: Video Alignment with Speaker Attention
by: Agrawal, Aviral, et al.
Published: (2024)
by: Agrawal, Aviral, et al.
Published: (2024)
In2x at WMT25 Translation Task
by: Pang, Lei, et al.
Published: (2025)
by: Pang, Lei, et al.
Published: (2025)
NLIP_Lab-IITH Low-Resource MT System for WMT24 Indic MT Shared Task
by: Sahoo, Pramit, et al.
Published: (2024)
by: Sahoo, Pramit, et al.
Published: (2024)
Lingua Custodi's participation at the WMT 2025 Terminology shared task
by: Liu, Jingshu, et al.
Published: (2025)
by: Liu, Jingshu, et al.
Published: (2025)
Preliminary Ranking of WMT25 General Machine Translation Systems
by: Kocmi, Tom, et al.
Published: (2025)
by: Kocmi, Tom, et al.
Published: (2025)
Social and Ethical Risks Posed by General-Purpose LLMs for Settling Newcomers in Canada
by: Nejadgholi, Isar, et al.
Published: (2024)
by: Nejadgholi, Isar, et al.
Published: (2024)
Cogs in a Machine, Doing What They're Meant to Do -- The AMI Submission to the WMT24 General Translation Task
by: Jasonarson, Atli, et al.
Published: (2024)
by: Jasonarson, Atli, et al.
Published: (2024)
Findings of the WMT 2024 Shared Task on Discourse-Level Literary Translation
by: Wang, Longyue, et al.
Published: (2024)
by: Wang, Longyue, et al.
Published: (2024)
SPECTRUM: Speaker-Enhanced Pre-Training for Long Dialogue Summarization
by: Cho, Sangwoo, et al.
Published: (2024)
by: Cho, Sangwoo, et al.
Published: (2024)
Typing to Listen at the Cocktail Party: Text-Guided Target Speaker Extraction
by: Hao, Xiang, et al.
Published: (2023)
by: Hao, Xiang, et al.
Published: (2023)
SpeakerSleuth: Can Large Audio-Language Models Judge Speaker Consistency across Multi-turn Dialogues?
by: Lee, Jonggeun, et al.
Published: (2026)
by: Lee, Jonggeun, et al.
Published: (2026)
Listening to Patients: A Framework of Detecting and Mitigating Patient Misreport for Medical Dialogue Generation
by: Qin, Lang, et al.
Published: (2024)
by: Qin, Lang, et al.
Published: (2024)
Exploring Parameter-Efficient Fine-Tuning and Backtranslation for the WMT 25 General Translation Task
by: Fujita, Felipe, et al.
Published: (2025)
by: Fujita, Felipe, et al.
Published: (2025)
Detecting Mental Manipulation in Speech via Synthetic Multi-Speaker Dialogue
by: Chen, Run, et al.
Published: (2026)
by: Chen, Run, et al.
Published: (2026)
Contrastive Speaker-Aware Learning for Multi-party Dialogue Generation with LLMs
by: Sun, Tianyu, et al.
Published: (2025)
by: Sun, Tianyu, et al.
Published: (2025)
Advancing Multi-Party Dialogue Framework with Speaker-ware Contrastive Learning
by: Hu, Zhongtian, et al.
Published: (2025)
by: Hu, Zhongtian, et al.
Published: (2025)
Human-Centered AI Applications for Canada's Immigration Settlement Sector
by: Nejadgholi, Isar, et al.
Published: (2024)
by: Nejadgholi, Isar, et al.
Published: (2024)
How Hypocritical Is Your LLM judge? Listener-Speaker Asymmetries in the Pragmatic Competence of Large Language Models
by: Sieker, Judith, et al.
Published: (2026)
by: Sieker, Judith, et al.
Published: (2026)
Similar Items
-
Gender-Neutral Machine Translation Strategies in Practice
by: Dawkins, Hillary, et al.
Published: (2025) -
Projective Methods for Mitigating Gender Bias in Pre-trained Language Models
by: Dawkins, Hillary, et al.
Published: (2024) -
Adaptable Moral Stances of Large Language Models on Sexist Content: Implications for Society and Gender Discourse
by: Guo, Rongchen, et al.
Published: (2024) -
Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency
by: Fraser, Kathleen C., et al.
Published: (2025) -
Challenging Negative Gender Stereotypes: A Study on the Effectiveness of Automated Counter-Stereotypes
by: Nejadgholi, Isar, et al.
Published: (2024)