GRUFF: LLM Pronoun Fidelity, Reasoning, and Biases in German
Fuente:
arXiv
Saved in:
| Main Authors: | Mewes, Fabian, Lauscher, Anne, Gautam, Vagrant |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robust Pronoun Fidelity with English LLMs: Are they Reasoning, Repeating, or Just Biased?
by: Gautam, Vagrant, et al.
Published: (2024)
by: Gautam, Vagrant, et al.
Published: (2024)
Stop! In the Name of Flaws: Disentangling Personal Names and Sociodemographic Attributes in NLP
by: Gautam, Vagrant, et al.
Published: (2024)
by: Gautam, Vagrant, et al.
Published: (2024)
Aligned Probing: Relating Toxic Behavior and Model Internals
by: Waldis, Andreas, et al.
Published: (2025)
by: Waldis, Andreas, et al.
Published: (2025)
Decoding Multilingual Moral Preferences: Unveiling LLM's Biases Through the Moral Machine Experiment
by: Vida, Karina, et al.
Published: (2024)
by: Vida, Karina, et al.
Published: (2024)
Teaching and Critiquing Conceptualization and Operationalization in NLP
by: Gautam, Vagrant
Published: (2025)
by: Gautam, Vagrant
Published: (2025)
WinoPron: Revisiting English Winogender Schemas for Consistency, Coverage, and Grammatical Case
by: Gautam, Vagrant, et al.
Published: (2024)
by: Gautam, Vagrant, et al.
Published: (2024)
Whose Facts Win? LLM Source Preferences under Knowledge Conflicts
by: Schuster, Jakob, et al.
Published: (2026)
by: Schuster, Jakob, et al.
Published: (2026)
Building Bridges: A Dataset for Evaluating Gender-Fair Machine Translation into German
by: Lardelli, Manuel, et al.
Published: (2024)
by: Lardelli, Manuel, et al.
Published: (2024)
Colombian Waitresses y Jueces canadienses: Gender and Country Biases in Occupation Recommendations from LLMs
by: Rodríguez, Elisa Forcada, et al.
Published: (2025)
by: Rodríguez, Elisa Forcada, et al.
Published: (2025)
The Lou Dataset -- Exploring the Impact of Gender-Fair Language in German Text Classification
by: Waldis, Andreas, et al.
Published: (2024)
by: Waldis, Andreas, et al.
Published: (2024)
Detecting Hallucinations in Authentic LLM-Human Interactions
by: Ren, Yujie, et al.
Published: (2025)
by: Ren, Yujie, et al.
Published: (2025)
Your Multimodal Speech Model Says I Have a Face for Radio
by: Nachesa, Maya K., et al.
Published: (2026)
by: Nachesa, Maya K., et al.
Published: (2026)
Around the World in 24 Hours: Probing LLM Knowledge of Time and Place
by: Holtermann, Carolin, et al.
Published: (2025)
by: Holtermann, Carolin, et al.
Published: (2025)
Understanding "Democratization" in NLP and ML Research
by: Subramonian, Arjun, et al.
Published: (2024)
by: Subramonian, Arjun, et al.
Published: (2024)
Agree to Disagree? A Meta-Evaluation of LLM Misgendering
by: Subramonian, Arjun, et al.
Published: (2025)
by: Subramonian, Arjun, et al.
Published: (2025)
Pronoun Logic
by: Bohrer, Rose, et al.
Published: (2024)
by: Bohrer, Rose, et al.
Published: (2024)
Large Language Models Discriminate Against Speakers of German Dialects
by: Bui, Minh Duc, et al.
Published: (2025)
by: Bui, Minh Duc, et al.
Published: (2025)
How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM Hallucination
by: Islam, Saad Obaid ul, et al.
Published: (2025)
by: Islam, Saad Obaid ul, et al.
Published: (2025)
Mention Attention for Pronoun Translation
by: Tang, Gongbo, et al.
Published: (2024)
by: Tang, Gongbo, et al.
Published: (2024)
The Echoes of Multilinguality: Tracing Cultural Value Shifts during LM Fine-tuning
by: Choenni, Rochelle, et al.
Published: (2024)
by: Choenni, Rochelle, et al.
Published: (2024)
From Insights to Actions: The Impact of Interpretability and Analysis Research on NLP
by: Mosbach, Marius, et al.
Published: (2024)
by: Mosbach, Marius, et al.
Published: (2024)
Reviewing the Reviewer: Elevating Peer Review Quality through LLM-Guided Feedback
by: Purkayastha, Sukannya, et al.
Published: (2026)
by: Purkayastha, Sukannya, et al.
Published: (2026)
Decoding Partial Differential Equations: Cross-Modal Adaptation of Decoder-only Models to PDEs
by: García-de-Herreros, Paloma, et al.
Published: (2025)
by: García-de-Herreros, Paloma, et al.
Published: (2025)
Blind Spots and Biases: Exploring the Role of Annotator Cognitive Biases in NLP
by: Gautam, Sanjana, et al.
Published: (2024)
by: Gautam, Sanjana, et al.
Published: (2024)
AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZ
by: Belouadi, Jonas, et al.
Published: (2023)
by: Belouadi, Jonas, et al.
Published: (2023)
SoS: Analysis of Surface over Semantics in Multilingual Text-To-Image Generation
by: Holtermann, Carolin, et al.
Published: (2026)
by: Holtermann, Carolin, et al.
Published: (2026)
TempViz: On the Evaluation of Temporal Knowledge in Text-to-Image Models
by: Holtermann, Carolin, et al.
Published: (2026)
by: Holtermann, Carolin, et al.
Published: (2026)
GIMMICK -- Globally Inclusive Multimodal Multitask Cultural Knowledge Benchmarking
by: Schneider, Florian, et al.
Published: (2025)
by: Schneider, Florian, et al.
Published: (2025)
Decision-Making with Deliberation: Meta-reviewing as a Document-grounded Dialogue
by: Purkayastha, Sukannya, et al.
Published: (2025)
by: Purkayastha, Sukannya, et al.
Published: (2025)
Local Contrastive Editing of Gender Stereotypes
by: Lutz, Marlene, et al.
Published: (2024)
by: Lutz, Marlene, et al.
Published: (2024)
Multi3Hate: Multimodal, Multilingual, and Multicultural Hate Speech Detection with Vision-Language Models
by: Bui, Minh Duc, et al.
Published: (2024)
by: Bui, Minh Duc, et al.
Published: (2024)
Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases
by: Huang, Hui, et al.
Published: (2026)
by: Huang, Hui, et al.
Published: (2026)
The Curious Case of Factual (Mis)Alignment between LLMs' Short- and Long-Form Answers
by: Islam, Saad Obaid ul, et al.
Published: (2025)
by: Islam, Saad Obaid ul, et al.
Published: (2025)
What the Weight?! A Unified Framework for Zero-Shot Knowledge Composition
by: Holtermann, Carolin, et al.
Published: (2024)
by: Holtermann, Carolin, et al.
Published: (2024)
Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ
by: Holtermann, Carolin, et al.
Published: (2024)
by: Holtermann, Carolin, et al.
Published: (2024)
Sensitivity, Performance, Robustness: Deconstructing the Effect of Sociodemographic Prompting
by: Beck, Tilman, et al.
Published: (2023)
by: Beck, Tilman, et al.
Published: (2023)
Large Language Models for Human-Machine Collaborative Particle Accelerator Tuning through Natural Language
by: Kaiser, Jan, et al.
Published: (2024)
by: Kaiser, Jan, et al.
Published: (2024)
LazyReview A Dataset for Uncovering Lazy Thinking in NLP Peer Reviews
by: Purkayastha, Sukannya, et al.
Published: (2025)
by: Purkayastha, Sukannya, et al.
Published: (2025)
How Much Do LLMs Know About Chinese Zero Pronouns?
by: Li, Yifei, et al.
Published: (2026)
by: Li, Yifei, et al.
Published: (2026)
Headlines You Won't Forget: Can Pronoun Insertion Increase Memorability?
by: Meyer, Selina, et al.
Published: (2026)
by: Meyer, Selina, et al.
Published: (2026)
Similar Items
-
Robust Pronoun Fidelity with English LLMs: Are they Reasoning, Repeating, or Just Biased?
by: Gautam, Vagrant, et al.
Published: (2024) -
Stop! In the Name of Flaws: Disentangling Personal Names and Sociodemographic Attributes in NLP
by: Gautam, Vagrant, et al.
Published: (2024) -
Aligned Probing: Relating Toxic Behavior and Model Internals
by: Waldis, Andreas, et al.
Published: (2025) -
Decoding Multilingual Moral Preferences: Unveiling LLM's Biases Through the Moral Machine Experiment
by: Vida, Karina, et al.
Published: (2024) -
Teaching and Critiquing Conceptualization and Operationalization in NLP
by: Gautam, Vagrant
Published: (2025)