Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Dogra, Atharvan, Ghosal, Soumya Suvra, Deshpande, Ameet, Kalyan, Ashwin, Manocha, Dinesh |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Language Models can Subtly Deceive Without Lying: A Case Study on Strategic Phrasing in Legislation
di: Dogra, Atharvan, et al.
Pubblicazione: (2024)
di: Dogra, Atharvan, et al.
Pubblicazione: (2024)
PromptRefine: Enhancing Few-Shot Performance on Low-Resource Indic Languages with Example Selection from Related Example Banks
di: Ghosal, Soumya Suvra, et al.
Pubblicazione: (2024)
di: Ghosal, Soumya Suvra, et al.
Pubblicazione: (2024)
Safety Recovery in Reasoning Models Is Only a Few Early Steering Steps Away
di: Ghosal, Soumya Suvra, et al.
Pubblicazione: (2026)
di: Ghosal, Soumya Suvra, et al.
Pubblicazione: (2026)
Relic: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples
di: Ghosal, Soumya Suvra, et al.
Pubblicazione: (2025)
di: Ghosal, Soumya Suvra, et al.
Pubblicazione: (2025)
IntCoOp: Interpretability-Aware Vision-Language Prompt Tuning
di: Ghosal, Soumya Suvra, et al.
Pubblicazione: (2024)
di: Ghosal, Soumya Suvra, et al.
Pubblicazione: (2024)
Probing AI Safety with Source Code
di: Narayan, Ujwal, et al.
Pubblicazione: (2025)
di: Narayan, Ujwal, et al.
Pubblicazione: (2025)
Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time
di: Chehade, Mohamad, et al.
Pubblicazione: (2025)
di: Chehade, Mohamad, et al.
Pubblicazione: (2025)
Transfer Q Star: Principled Decoding for LLM Alignment
di: Chakraborty, Souradip, et al.
Pubblicazione: (2024)
di: Chakraborty, Souradip, et al.
Pubblicazione: (2024)
Does Thinking More always Help? Mirage of Test-Time Scaling in Reasoning Models
di: Ghosal, Soumya Suvra, et al.
Pubblicazione: (2025)
di: Ghosal, Soumya Suvra, et al.
Pubblicazione: (2025)
Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs
di: Gupta, Shashank, et al.
Pubblicazione: (2023)
di: Gupta, Shashank, et al.
Pubblicazione: (2023)
QualEval: Qualitative Evaluation for Model Improvement
di: Murahari, Vishvak, et al.
Pubblicazione: (2023)
di: Murahari, Vishvak, et al.
Pubblicazione: (2023)
Collab: Controlled Decoding using Mixture of Agents for LLM Alignment
di: Chakraborty, Souradip, et al.
Pubblicazione: (2025)
di: Chakraborty, Souradip, et al.
Pubblicazione: (2025)
KITE: Kernelized and Information Theoretic Exemplars for In-Context Learning
di: Singh, Vaibhav, et al.
Pubblicazione: (2025)
di: Singh, Vaibhav, et al.
Pubblicazione: (2025)
Agent Context Protocols Enhance Collective Inference
di: Bhardwaj, Devansh, et al.
Pubblicazione: (2025)
di: Bhardwaj, Devansh, et al.
Pubblicazione: (2025)
Compromesso! Italian Many-Shot Jailbreaks Undermine the Safety of Large Language Models
di: Pernisi, Fabio, et al.
Pubblicazione: (2024)
di: Pernisi, Fabio, et al.
Pubblicazione: (2024)
Text Is Not All You Need: Multimodal Prompting Helps LLMs Understand Humor
di: Baluja, Ashwin
Pubblicazione: (2024)
di: Baluja, Ashwin
Pubblicazione: (2024)
RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
di: Chaudhari, Shreyas, et al.
Pubblicazione: (2024)
di: Chaudhari, Shreyas, et al.
Pubblicazione: (2024)
PersonaGym: Evaluating Persona Agents and LLMs
di: Samuel, Vinay, et al.
Pubblicazione: (2024)
di: Samuel, Vinay, et al.
Pubblicazione: (2024)
Augmenting Lateral Thinking in Language Models with Humor and Riddle Data for the BRAINTEASER Task
di: Ghashami, Mina, et al.
Pubblicazione: (2024)
di: Ghashami, Mina, et al.
Pubblicazione: (2024)
Analyzing the Safety of Japanese Large Language Models in Stereotype-Triggering Prompts
di: Nakanishi, Akito, et al.
Pubblicazione: (2025)
di: Nakanishi, Akito, et al.
Pubblicazione: (2025)
Reading Between the Prompts: How Stereotypes Shape LLM's Implicit Personalization
di: Neplenbroek, Vera, et al.
Pubblicazione: (2025)
di: Neplenbroek, Vera, et al.
Pubblicazione: (2025)
Do Vision-Language Models Understand Compound Nouns?
di: Kumar, Sonal, et al.
Pubblicazione: (2024)
di: Kumar, Sonal, et al.
Pubblicazione: (2024)
The Reliability Paradox: Exploring How Shortcut Learning Undermines Language Model Calibration
di: Bihani, Geetanjali, et al.
Pubblicazione: (2024)
di: Bihani, Geetanjali, et al.
Pubblicazione: (2024)
Quantifying Stereotypes in Language
di: Liu, Yang
Pubblicazione: (2024)
di: Liu, Yang
Pubblicazione: (2024)
HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models
di: Ajayi, Edward, et al.
Pubblicazione: (2026)
di: Ajayi, Edward, et al.
Pubblicazione: (2026)
HumorGen: Cognitive Synergy for Humor Generation in Large Language Models via Persona-Based Distillation
di: Ajayi, Edward, et al.
Pubblicazione: (2026)
di: Ajayi, Edward, et al.
Pubblicazione: (2026)
Fragile Mastery: Are Domain-Specific Trade-Offs Undermining On-Device Language Models?
di: Jha, Basab, et al.
Pubblicazione: (2025)
di: Jha, Basab, et al.
Pubblicazione: (2025)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
di: Cheng, Stephen, et al.
Pubblicazione: (2026)
di: Cheng, Stephen, et al.
Pubblicazione: (2026)
Measuring Stereotype and Deviation Biases in Large Language Models
di: Wang, Daniel, et al.
Pubblicazione: (2025)
di: Wang, Daniel, et al.
Pubblicazione: (2025)
Systematic Offensive Stereotyping (SOS) Bias in Language Models
di: Elsafoury, Fatma
Pubblicazione: (2023)
di: Elsafoury, Fatma
Pubblicazione: (2023)
Bypassing Safety Guardrails in LLMs Using Humor
di: Cisneros-Velarde, Pedro
Pubblicazione: (2025)
di: Cisneros-Velarde, Pedro
Pubblicazione: (2025)
Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
di: Wu, Yulong, et al.
Pubblicazione: (2025)
di: Wu, Yulong, et al.
Pubblicazione: (2025)
Do Multilingual Large Language Models Mitigate Stereotype Bias?
di: Nie, Shangrui, et al.
Pubblicazione: (2024)
di: Nie, Shangrui, et al.
Pubblicazione: (2024)
ELITE: Enhanced Language-Image Toxicity Evaluation for Safety
di: Lee, Wonjun, et al.
Pubblicazione: (2025)
di: Lee, Wonjun, et al.
Pubblicazione: (2025)
Getting Serious about Humor: Crafting Humor Datasets with Unfunny Large Language Models
di: Horvitz, Zachary, et al.
Pubblicazione: (2024)
di: Horvitz, Zachary, et al.
Pubblicazione: (2024)
Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages
di: Hu, Yujia, et al.
Pubblicazione: (2025)
di: Hu, Yujia, et al.
Pubblicazione: (2025)
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation
di: Zhang, Zhibo, et al.
Pubblicazione: (2025)
di: Zhang, Zhibo, et al.
Pubblicazione: (2025)
How Toxic Can You Get? Search-based Toxicity Testing for Large Language Models
di: Corbo, Simone, et al.
Pubblicazione: (2025)
di: Corbo, Simone, et al.
Pubblicazione: (2025)
CFunModel: A "Funny" Language Model Capable of Chinese Humor Generation and Processing
di: Yu, Zhenghan, et al.
Pubblicazione: (2025)
di: Yu, Zhenghan, et al.
Pubblicazione: (2025)
Women Are Beautiful, Men Are Leaders: Gender Stereotypes in Machine Translation and Language Modeling
di: Pikuliak, Matúš, et al.
Pubblicazione: (2023)
di: Pikuliak, Matúš, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Language Models can Subtly Deceive Without Lying: A Case Study on Strategic Phrasing in Legislation
di: Dogra, Atharvan, et al.
Pubblicazione: (2024) -
PromptRefine: Enhancing Few-Shot Performance on Low-Resource Indic Languages with Example Selection from Related Example Banks
di: Ghosal, Soumya Suvra, et al.
Pubblicazione: (2024) -
Safety Recovery in Reasoning Models Is Only a Few Early Steering Steps Away
di: Ghosal, Soumya Suvra, et al.
Pubblicazione: (2026) -
Relic: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples
di: Ghosal, Soumya Suvra, et al.
Pubblicazione: (2025) -
IntCoOp: Interpretability-Aware Vision-Language Prompt Tuning
di: Ghosal, Soumya Suvra, et al.
Pubblicazione: (2024)