GAMBIT+: A Challenge Set for Evaluating Gender Bias in Machine Translation Quality Estimation Metrics
Fuente:
arXiv
Saved in:
| Main Authors: | Filandrianos, Giorgos, Mastromichalakis, Orfeas Menis, Mohammed, Wafaa, Attanasio, Giuseppe, Zerva, Chrysoula |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Assumed Identities: Quantifying Gender Bias in Machine Translation of Gender-Ambiguous Occupational Terms
by: Mastromichalakis, Orfeas Menis, et al.
Published: (2025)
by: Mastromichalakis, Orfeas Menis, et al.
Published: (2025)
GOSt-MT: A Knowledge Graph for Occupation-related Gender Biases in Machine Translation
by: Mastromichalakis, Orfeas Menis, et al.
Published: (2024)
by: Mastromichalakis, Orfeas Menis, et al.
Published: (2024)
Unlocking Latent Discourse Translation in LLMs Through Quality-Aware Decoding
by: Mohammed, Wafaa, et al.
Published: (2025)
by: Mohammed, Wafaa, et al.
Published: (2025)
AILS-NTUA at SemEval-2025 Task 4: Parameter-Efficient Unlearning for Large Language Models using Data Chunking
by: Premptis, Iraklis, et al.
Published: (2025)
by: Premptis, Iraklis, et al.
Published: (2025)
"I Never Said That": A dataset, taxonomy and baselines on response clarity classification
by: Thomas, Konstantinos, et al.
Published: (2024)
by: Thomas, Konstantinos, et al.
Published: (2024)
SemEval-2026 Task 6: CLARITY -- Unmasking Political Question Evasions
by: Thomas, Konstantinos, et al.
Published: (2026)
by: Thomas, Konstantinos, et al.
Published: (2026)
Don't Erase, Inform! Detecting and Contextualizing Harmful Language in Cultural Heritage Collections
by: Mastromichalakis, Orfeas Menis, et al.
Published: (2025)
by: Mastromichalakis, Orfeas Menis, et al.
Published: (2025)
Deep Ensemble Art Style Recognition
by: Menis-Mastromichalakis, Orfeas, et al.
Published: (2024)
by: Menis-Mastromichalakis, Orfeas, et al.
Published: (2024)
The Grounding Gap: How LLMs Anchor the Meaning of Abstract Concepts Differently from Humans
by: Chlapanis, Odysseas S., et al.
Published: (2026)
by: Chlapanis, Odysseas S., et al.
Published: (2026)
Semantic Prototypes: Enhancing Transparency Without Black Boxes
by: Menis-Mastromichalakis, Orfeas, et al.
Published: (2024)
by: Menis-Mastromichalakis, Orfeas, et al.
Published: (2024)
Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality Estimation
by: Zaranis, Emmanouil, et al.
Published: (2024)
by: Zaranis, Emmanouil, et al.
Published: (2024)
Beyond One-Size-Fits-All: Adapting Counterfactual Explanations to User Objectives
by: Mastromichalakis, Orfeas Menis, et al.
Published: (2024)
by: Mastromichalakis, Orfeas Menis, et al.
Published: (2024)
Explain the Flag: Contextualizing Hate Speech Beyond Censorship
by: Liartis, Jason, et al.
Published: (2026)
by: Liartis, Jason, et al.
Published: (2026)
Building Bridges: A Dataset for Evaluating Gender-Fair Machine Translation into German
by: Lardelli, Manuel, et al.
Published: (2024)
by: Lardelli, Manuel, et al.
Published: (2024)
A Conformal Risk Control Framework for Granular Word Assessment and Uncertainty Calibration of CLIPScore Quality Estimates
by: Gomes, Gonçalo, et al.
Published: (2025)
by: Gomes, Gonçalo, et al.
Published: (2025)
MusicLIME: Explainable Multimodal Music Understanding
by: Sotirou, Theodoros, et al.
Published: (2024)
by: Sotirou, Theodoros, et al.
Published: (2024)
Evaluation of Multilingual Image Captioning: How far can we get with CLIP models?
by: Gomes, Gonçalo, et al.
Published: (2025)
by: Gomes, Gonçalo, et al.
Published: (2025)
Bias Beware: The Impact of Cognitive Biases on LLM-Driven Product Recommendations
by: Filandrianos, Giorgos, et al.
Published: (2025)
by: Filandrianos, Giorgos, et al.
Published: (2025)
Translating With Feeling: Centering Translator Perspectives within Translation Technologies
by: Chechelnitsky, Daniel, et al.
Published: (2026)
by: Chechelnitsky, Daniel, et al.
Published: (2026)
Reasoning or Pattern Matching? Probing Large Vision-Language Models with Visual Puzzles
by: Lymperaiou, Maria, et al.
Published: (2026)
by: Lymperaiou, Maria, et al.
Published: (2026)
Different Speech Translation Models Encode and Translate Speaker Gender Differently
by: Fucci, Dennis, et al.
Published: (2025)
by: Fucci, Dennis, et al.
Published: (2025)
Evaluating Counterfactual Strategic Reasoning in Large Language Models
by: Georgousis, Dimitrios, et al.
Published: (2026)
by: Georgousis, Dimitrios, et al.
Published: (2026)
Don't Rank, Combine! Combining Machine Translation Hypotheses Using Quality Estimation
by: Vernikos, Giorgos, et al.
Published: (2024)
by: Vernikos, Giorgos, et al.
Published: (2024)
PAKTON: A Multi-Agent Framework for Question Answering in Long Legal Agreements
by: Raptopoulos, Petros, et al.
Published: (2025)
by: Raptopoulos, Petros, et al.
Published: (2025)
AILS-NTUA at SemEval-2026 Task 8: Evaluating Multi-Turn RAG Conversations
by: Athanasiou, Dimosthenis, et al.
Published: (2026)
by: Athanasiou, Dimosthenis, et al.
Published: (2026)
Non-Exchangeable Conformal Language Generation with Nearest Neighbors
by: Ulmer, Dennis, et al.
Published: (2024)
by: Ulmer, Dennis, et al.
Published: (2024)
Gender Inflected or Bias Inflicted: On Using Grammatical Gender Cues for Bias Evaluation in Machine Translation
by: Singh, Pushpdeep
Published: (2023)
by: Singh, Pushpdeep
Published: (2023)
AILS-NTUA at SemEval-2025 Task 3: Leveraging Large Language Models and Translation Strategies for Multilingual Hallucination Detection
by: Karkani, Dimitra, et al.
Published: (2025)
by: Karkani, Dimitra, et al.
Published: (2025)
Context-Aware or Context-Insensitive? Assessing LLMs' Performance in Document-Level Translation
by: Mohammed, Wafaa, et al.
Published: (2024)
by: Mohammed, Wafaa, et al.
Published: (2024)
Enhancing adversarial robustness in Natural Language Inference using explanations
by: Koulakos, Alexandros, et al.
Published: (2024)
by: Koulakos, Alexandros, et al.
Published: (2024)
AILS-NTUA at SemEval-2024 Task 6: Efficient model tuning for hallucination detection and analysis
by: Grigoriadou, Natalia, et al.
Published: (2024)
by: Grigoriadou, Natalia, et al.
Published: (2024)
Optimal and efficient text counterfactuals using Graph Neural Networks
by: Lymperopoulos, Dimitris, et al.
Published: (2024)
by: Lymperopoulos, Dimitris, et al.
Published: (2024)
RISCORE: Enhancing In-Context Riddle Solving in Language Models through Context-Reconstructed Example Augmentation
by: Panagiotopoulos, Ioannis, et al.
Published: (2024)
by: Panagiotopoulos, Ioannis, et al.
Published: (2024)
Puzzle Solving using Reasoning of Large Language Models: A Survey
by: Giadikiaroglou, Panagiotis, et al.
Published: (2024)
by: Giadikiaroglou, Panagiotis, et al.
Published: (2024)
AILS-NTUA at SemEval-2024 Task 9: Cracking Brain Teasers: Transformer Models for Lateral Thinking Puzzles
by: Panagiotopoulos, Ioannis, et al.
Published: (2024)
by: Panagiotopoulos, Ioannis, et al.
Published: (2024)
Gender Bias in English-to-Greek Machine Translation
by: Gkovedarou, Eleni, et al.
Published: (2025)
by: Gkovedarou, Eleni, et al.
Published: (2025)
Machine Translation Meta Evaluation through Translation Accuracy Challenge Sets
by: Moghe, Nikita, et al.
Published: (2024)
by: Moghe, Nikita, et al.
Published: (2024)
FairTranslate: An English-French Dataset for Gender Bias Evaluation in Machine Translation by Overcoming Gender Binarity
by: Jourdan, Fanny, et al.
Published: (2025)
by: Jourdan, Fanny, et al.
Published: (2025)
Mind the Inclusivity Gap: Multilingual Gender-Neutral Translation Evaluation with mGeNTE
by: Savoldi, Beatrice, et al.
Published: (2025)
by: Savoldi, Beatrice, et al.
Published: (2025)
Are All Spanish Doctors Male? Evaluating Gender Bias in German Machine Translation
by: Kappl, Michelle
Published: (2025)
by: Kappl, Michelle
Published: (2025)
Similar Items
-
Assumed Identities: Quantifying Gender Bias in Machine Translation of Gender-Ambiguous Occupational Terms
by: Mastromichalakis, Orfeas Menis, et al.
Published: (2025) -
GOSt-MT: A Knowledge Graph for Occupation-related Gender Biases in Machine Translation
by: Mastromichalakis, Orfeas Menis, et al.
Published: (2024) -
Unlocking Latent Discourse Translation in LLMs Through Quality-Aware Decoding
by: Mohammed, Wafaa, et al.
Published: (2025) -
AILS-NTUA at SemEval-2025 Task 4: Parameter-Efficient Unlearning for Large Language Models using Data Chunking
by: Premptis, Iraklis, et al.
Published: (2025) -
"I Never Said That": A dataset, taxonomy and baselines on response clarity classification
by: Thomas, Konstantinos, et al.
Published: (2024)