Will the Prince Get True Love's Kiss? On the Model Sensitivity to Gender Perturbation over Fairytale Texts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chance, Christina, Yin, Da, Wang, Dakuo, Chang, Kai-Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging Language Models to Detect Greenwashing
von: Vinella, Avalon, et al.
Veröffentlicht: (2023)
von: Vinella, Avalon, et al.
Veröffentlicht: (2023)
KPEval: Towards Fine-Grained Semantic-Based Keyphrase Evaluation
von: Wu, Di, et al.
Veröffentlicht: (2023)
von: Wu, Di, et al.
Veröffentlicht: (2023)
IYKYK (But AI Doesn't): Automated Content Moderation Does Not Capture Communities' Heterogeneous Attitudes Towards Reclaimed Language
von: Chance, Christina, et al.
Veröffentlicht: (2026)
von: Chance, Christina, et al.
Veröffentlicht: (2026)
FairytaleQA Translated: Enabling Educational Question and Answer Generation in Less-Resourced Languages
von: Leite, Bernardo, et al.
Veröffentlicht: (2024)
von: Leite, Bernardo, et al.
Veröffentlicht: (2024)
Reinforcing Stereotypes of Anger: Emotion AI on African American Vernacular English
von: Dorn, Rebecca, et al.
Veröffentlicht: (2025)
von: Dorn, Rebecca, et al.
Veröffentlicht: (2025)
MOSS-Speech: Towards True Speech-to-Speech Models Without Text Guidance
von: Zhao, Xingjian, et al.
Veröffentlicht: (2025)
von: Zhao, Xingjian, et al.
Veröffentlicht: (2025)
Characterizing Truthfulness in Large Language Model Generations with Local Intrinsic Dimension
von: Yin, Fan, et al.
Veröffentlicht: (2024)
von: Yin, Fan, et al.
Veröffentlicht: (2024)
Leveraging Large Language Models and Topic Modeling for Toxicity Classification
von: Oskouie, Haniyeh Ehsani, et al.
Veröffentlicht: (2024)
von: Oskouie, Haniyeh Ehsani, et al.
Veröffentlicht: (2024)
Who Gets the Callback? Generative AI and Gender Bias
von: Chaturvedi, Sugat, et al.
Veröffentlicht: (2025)
von: Chaturvedi, Sugat, et al.
Veröffentlicht: (2025)
What Gets Unmasked First? Trajectory Analysis of Diffusion Models for Graph-to-Text Generation
von: Wang, Qing, et al.
Veröffentlicht: (2026)
von: Wang, Qing, et al.
Veröffentlicht: (2026)
The Root Shapes the Fruit: On the Persistence of Gender-Exclusive Harms in Aligned Language Models
von: Ovalle, Anaelia, et al.
Veröffentlicht: (2024)
von: Ovalle, Anaelia, et al.
Veröffentlicht: (2024)
Kiss up, Kick down: Exploring Behavioral Changes in Multi-modal Large Language Models with Assigned Visual Personas
von: Sun, Seungjong, et al.
Veröffentlicht: (2024)
von: Sun, Seungjong, et al.
Veröffentlicht: (2024)
ModelCitizens: Representing Community Voices in Online Safety
von: Suvarna, Ashima, et al.
Veröffentlicht: (2025)
von: Suvarna, Ashima, et al.
Veröffentlicht: (2025)
SafeWorld: Geo-Diverse Safety Alignment
von: Yin, Da, et al.
Veröffentlicht: (2024)
von: Yin, Da, et al.
Veröffentlicht: (2024)
VisRet: Visualization Improves Knowledge-Intensive Text-to-Image Retrieval
von: Wu, Di, et al.
Veröffentlicht: (2025)
von: Wu, Di, et al.
Veröffentlicht: (2025)
Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025)
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025)
Guided Perturbation Sensitivity (GPS): Detecting Adversarial Text via Embedding Stability and Word Importance
von: Tuck, Bryan E., et al.
Veröffentlicht: (2025)
von: Tuck, Bryan E., et al.
Veröffentlicht: (2025)
Customer-R1: Personalized Simulation of Human Behaviors via RL-based LLM Agent in Online Shopping
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
CompAlign: Improving Compositional Text-to-Image Generation with a Complex Benchmark and Fine-Grained Feedback
von: Wan, Yixin, et al.
Veröffentlicht: (2025)
von: Wan, Yixin, et al.
Veröffentlicht: (2025)
Understanding and Mitigating Spurious Correlations in Text Classification with Neighborhood Analysis
von: Chew, Oscar, et al.
Veröffentlicht: (2023)
von: Chew, Oscar, et al.
Veröffentlicht: (2023)
PEPPER: Perception-Guided Perturbation for Robust Backdoor Defense in Text-to-Image Diffusion Models
von: Chew, Oscar, et al.
Veröffentlicht: (2025)
von: Chew, Oscar, et al.
Veröffentlicht: (2025)
Who Gets the Mic? Investigating Gender Bias in the Speaker Assignment of a Speech-LLM
von: Puhach, Dariia, et al.
Veröffentlicht: (2025)
von: Puhach, Dariia, et al.
Veröffentlicht: (2025)
ROIC-DM: Robust Text Inference and Classification via Diffusion Model
von: Yuan, Shilong, et al.
Veröffentlicht: (2024)
von: Yuan, Shilong, et al.
Veröffentlicht: (2024)
Are AI-Generated Text Detectors Robust to Adversarial Perturbations?
von: Huang, Guanhua, et al.
Veröffentlicht: (2024)
von: Huang, Guanhua, et al.
Veröffentlicht: (2024)
Fragile Reasoning: A Mechanistic Analysis of LLM Sensitivity to Meaning-Preserving Perturbations
von: Han, Shou-Tzu, et al.
Veröffentlicht: (2026)
von: Han, Shou-Tzu, et al.
Veröffentlicht: (2026)
Effective Unsupervised Constrained Text Generation based on Perturbed Masking
von: Fu, Yingwen, et al.
Veröffentlicht: (2024)
von: Fu, Yingwen, et al.
Veröffentlicht: (2024)
The Factuality Tax of Diversity-Intervened Text-to-Image Generation: Benchmark and Fact-Augmented Intervention
von: Wan, Yixin, et al.
Veröffentlicht: (2024)
von: Wan, Yixin, et al.
Veröffentlicht: (2024)
How Linguistics Learned to Stop Worrying and Love the Language Models
von: Futrell, Richard, et al.
Veröffentlicht: (2025)
von: Futrell, Richard, et al.
Veröffentlicht: (2025)
Beyond Content: How Grammatical Gender Shapes Visual Representation in Text-to-Image Models
von: Saeed, Muhammed, et al.
Veröffentlicht: (2025)
von: Saeed, Muhammed, et al.
Veröffentlicht: (2025)
Adversarial Text Generation with Dynamic Contextual Perturbation
von: Waghela, Hetvi, et al.
Veröffentlicht: (2025)
von: Waghela, Hetvi, et al.
Veröffentlicht: (2025)
Are Language Models Sensitive to Morally Irrelevant Distractors?
von: Shaw, Andrew, et al.
Veröffentlicht: (2026)
von: Shaw, Andrew, et al.
Veröffentlicht: (2026)
Agent-ScanKit: Unraveling Memory and Reasoning of Multimodal Agents via Sensitivity Perturbations
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2025)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2025)
Auditing Gender Presentation Differences in Text-to-Image Models
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2023)
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2023)
KERMIT: Knowledge Graph Completion of Enhanced Relation Modeling with Inverse Transformation
von: Li, Haotian, et al.
Veröffentlicht: (2023)
von: Li, Haotian, et al.
Veröffentlicht: (2023)
Perturbation-Restrained Sequential Model Editing
von: Ma, Jun-Yu, et al.
Veröffentlicht: (2024)
von: Ma, Jun-Yu, et al.
Veröffentlicht: (2024)
Are Female Carpenters like Blue Bananas? A Corpus Investigation of Occupation Gender Typicality
von: Ju, Da, et al.
Veröffentlicht: (2024)
von: Ju, Da, et al.
Veröffentlicht: (2024)
Getting Your Indices in a Row: Full-Text Search for LLM Training Data for Real World
von: Marinas, Ines Altemir, et al.
Veröffentlicht: (2025)
von: Marinas, Ines Altemir, et al.
Veröffentlicht: (2025)
Taming Sensitive Weights : Noise Perturbation Fine-tuning for Robust LLM Quantization
von: Wang, Dongwei, et al.
Veröffentlicht: (2024)
von: Wang, Dongwei, et al.
Veröffentlicht: (2024)
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
von: Shu, Huizhen, et al.
Veröffentlicht: (2025)
von: Shu, Huizhen, et al.
Veröffentlicht: (2025)
Vulnerability of LLMs to Vertically Aligned Text Manipulations
von: Li, Zhecheng, et al.
Veröffentlicht: (2024)
von: Li, Zhecheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Leveraging Language Models to Detect Greenwashing
von: Vinella, Avalon, et al.
Veröffentlicht: (2023) -
KPEval: Towards Fine-Grained Semantic-Based Keyphrase Evaluation
von: Wu, Di, et al.
Veröffentlicht: (2023) -
IYKYK (But AI Doesn't): Automated Content Moderation Does Not Capture Communities' Heterogeneous Attitudes Towards Reclaimed Language
von: Chance, Christina, et al.
Veröffentlicht: (2026) -
FairytaleQA Translated: Enabling Educational Question and Answer Generation in Less-Resourced Languages
von: Leite, Bernardo, et al.
Veröffentlicht: (2024) -
Reinforcing Stereotypes of Anger: Emotion AI on African American Vernacular English
von: Dorn, Rebecca, et al.
Veröffentlicht: (2025)