Probability of Differentiation Reveals Brittleness of Homogeneity Bias in GPT-4
Fuente:
arXiv
Guardado en:
| Autores principales: | Lee, Messi H. J., Lai, Calvin K. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Examining the Robustness of Homogeneity Bias to Hyperparameter Adjustments in GPT-4
por: Lee, Messi H. J.
Publicado: (2025)
por: Lee, Messi H. J.
Publicado: (2025)
Large Language Models Portray Socially Subordinate Groups as More Homogeneous, Consistent with a Bias Observed in Humans
por: Lee, Messi H. J., et al.
Publicado: (2024)
por: Lee, Messi H. J., et al.
Publicado: (2024)
Implicit Bias-Like Patterns in Reasoning Models
por: Lee, Messi H. J., et al.
Publicado: (2025)
por: Lee, Messi H. J., et al.
Publicado: (2025)
Token-Level Entropy Reveals Demographic Disparities in Language Models
por: Lee, Messi H. J.
Publicado: (2025)
por: Lee, Messi H. J.
Publicado: (2025)
More Distinctively Black and Feminine Faces Lead to Increased Stereotyping in Vision-Language Models
por: Lee, Messi H. J., et al.
Publicado: (2024)
por: Lee, Messi H. J., et al.
Publicado: (2024)
Vision-Language Models Generate More Homogeneous Stories for Phenotypically Black Individuals
por: Lee, Messi H. J., et al.
Publicado: (2024)
por: Lee, Messi H. J., et al.
Publicado: (2024)
Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases
por: Gao, Bufan, et al.
Publicado: (2025)
por: Gao, Bufan, et al.
Publicado: (2025)
Bias Attribution in Filipino Language Models: Extending a Bias Interpretability Metric for Application on Agglutinative Languages
por: Gamboa, Lance Calvin Lim, et al.
Publicado: (2025)
por: Gamboa, Lance Calvin Lim, et al.
Publicado: (2025)
Robust Bias Evaluation with FilBBQ: A Filipino Bias Benchmark for Question-Answering Language Models
por: Gamboa, Lance Calvin Lim, et al.
Publicado: (2026)
por: Gamboa, Lance Calvin Lim, et al.
Publicado: (2026)
Filipino Benchmarks for Measuring Sexist and Homophobic Bias in Multilingual Language Models from Southeast Asia
por: Gamboa, Lance Calvin Lim, et al.
Publicado: (2024)
por: Gamboa, Lance Calvin Lim, et al.
Publicado: (2024)
Social Bias in Multilingual Language Models: A Survey
por: Gamboa, Lance Calvin Lim, et al.
Publicado: (2025)
por: Gamboa, Lance Calvin Lim, et al.
Publicado: (2025)
Visual Cues of Gender and Race are Associated with Stereotyping in Vision-Language Models
por: Lee, Messi H. J., et al.
Publicado: (2025)
por: Lee, Messi H. J., et al.
Publicado: (2025)
A Novel Interpretability Metric for Explaining Bias in Language Models: Applications on Multilingual Models from Southeast Asia
por: Gamboa, Lance Calvin Lim, et al.
Publicado: (2024)
por: Gamboa, Lance Calvin Lim, et al.
Publicado: (2024)
Weird Generalization is Weirdly Brittle
por: Wanner, Miriam, et al.
Publicado: (2026)
por: Wanner, Miriam, et al.
Publicado: (2026)
Are Humans as Brittle as Large Language Models?
por: Li, Jiahui, et al.
Publicado: (2025)
por: Li, Jiahui, et al.
Publicado: (2025)
Systematic Diagnosis of Brittle Reasoning in Large Language Models
por: Parupudi, V. S. Raghu
Publicado: (2025)
por: Parupudi, V. S. Raghu
Publicado: (2025)
On the Brittleness of LLMs: A Journey around Set Membership
por: Hergert, Lea, et al.
Publicado: (2025)
por: Hergert, Lea, et al.
Publicado: (2025)
Token Homogenization under Positional Bias
por: Yusupov, Viacheslav, et al.
Publicado: (2025)
por: Yusupov, Viacheslav, et al.
Publicado: (2025)
Certain but not Probable? Differentiating Certainty from Probability in LLM Token Outputs for Probabilistic Scenarios
por: Toney-Wails, Autumn, et al.
Publicado: (2025)
por: Toney-Wails, Autumn, et al.
Publicado: (2025)
LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance
por: Haller, Patrick, et al.
Publicado: (2025)
por: Haller, Patrick, et al.
Publicado: (2025)
Detecting Bias in Large Language Models: Fine-tuned KcBERT
por: Lee, J. K., et al.
Publicado: (2024)
por: Lee, J. K., et al.
Publicado: (2024)
LLMs Show Surface-Form Brittleness Under Paraphrase Stress Tests
por: Carranza, Juan Miguel Navarro
Publicado: (2025)
por: Carranza, Juan Miguel Navarro
Publicado: (2025)
Examining Multimodal Gender and Content Bias in ChatGPT-4o
por: Balestri, Roberto
Publicado: (2024)
por: Balestri, Roberto
Publicado: (2024)
Textual Entailment is not a Better Bias Metric than Token Probability
por: Felkner, Virginia K., et al.
Publicado: (2025)
por: Felkner, Virginia K., et al.
Publicado: (2025)
An Evaluation of GPT-4V for Transcribing the Urban Renewal Hand-Written Collection
por: Lee, Myeong, et al.
Publicado: (2024)
por: Lee, Myeong, et al.
Publicado: (2024)
Linguistic Bias in ChatGPT: Language Models Reinforce Dialect Discrimination
por: Fleisig, Eve, et al.
Publicado: (2024)
por: Fleisig, Eve, et al.
Publicado: (2024)
Generalization or Memorization? Brittleness Testing for Chess-Trained Language Models
por: Tang, Ethan
Publicado: (2026)
por: Tang, Ethan
Publicado: (2026)
On the Brittle Foundations of ReAct Prompting for Agentic Large Language Models
por: Verma, Mudit, et al.
Publicado: (2024)
por: Verma, Mudit, et al.
Publicado: (2024)
Brittle Minds, Fixable Activations: Understanding Belief Representations in Language Models
por: Bortoletto, Matteo, et al.
Publicado: (2024)
por: Bortoletto, Matteo, et al.
Publicado: (2024)
Brittleness and Promise: Knowledge Graph Based Reward Modeling for Diagnostic Reasoning
por: Khatwani, Saksham, et al.
Publicado: (2025)
por: Khatwani, Saksham, et al.
Publicado: (2025)
Benchmarking Llama2, Mistral, Gemma and GPT for Factuality, Toxicity, Bias and Propensity for Hallucinations
por: Nadeau, David, et al.
Publicado: (2024)
por: Nadeau, David, et al.
Publicado: (2024)
Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature
por: Bao, Guangsheng, et al.
Publicado: (2023)
por: Bao, Guangsheng, et al.
Publicado: (2023)
How Prevalent is Gender Bias in ChatGPT? -- Exploring German and English ChatGPT Responses
por: Urchs, Stefanie, et al.
Publicado: (2023)
por: Urchs, Stefanie, et al.
Publicado: (2023)
Revealing User Familiarity Bias in Task-Oriented Dialogue via Interactive Evaluation
por: Kim, Takyoung, et al.
Publicado: (2023)
por: Kim, Takyoung, et al.
Publicado: (2023)
Different Time, Different Language: Revisiting the Bias Against Non-Native Speakers in GPT Detectors
por: Ali, Adnan Al, et al.
Publicado: (2026)
por: Ali, Adnan Al, et al.
Publicado: (2026)
Your Agent is More Brittle Than You Think: Uncovering Indirect Injection Vulnerabilities in Agentic LLMs
por: Zhu, Wenhui, et al.
Publicado: (2026)
por: Zhu, Wenhui, et al.
Publicado: (2026)
Does GPT-4 pass the Turing test?
por: Jones, Cameron R., et al.
Publicado: (2023)
por: Jones, Cameron R., et al.
Publicado: (2023)
ChatGPT v.s. Media Bias: A Comparative Study of GPT-3.5 and Fine-tuned Language Models
por: Wen, Zehao, et al.
Publicado: (2024)
por: Wen, Zehao, et al.
Publicado: (2024)
Using GPT-4 to Augment Unbalanced Data for Automatic Scoring
por: Fang, Luyang, et al.
Publicado: (2023)
por: Fang, Luyang, et al.
Publicado: (2023)
100% Elimination of Hallucinations on RAGTruth for GPT-4 and GPT-3.5 Turbo
por: Wood, Michael C., et al.
Publicado: (2024)
por: Wood, Michael C., et al.
Publicado: (2024)
Ejemplares similares
-
Examining the Robustness of Homogeneity Bias to Hyperparameter Adjustments in GPT-4
por: Lee, Messi H. J.
Publicado: (2025) -
Large Language Models Portray Socially Subordinate Groups as More Homogeneous, Consistent with a Bias Observed in Humans
por: Lee, Messi H. J., et al.
Publicado: (2024) -
Implicit Bias-Like Patterns in Reasoning Models
por: Lee, Messi H. J., et al.
Publicado: (2025) -
Token-Level Entropy Reveals Demographic Disparities in Language Models
por: Lee, Messi H. J.
Publicado: (2025) -
More Distinctively Black and Feminine Faces Lead to Increased Stereotyping in Vision-Language Models
por: Lee, Messi H. J., et al.
Publicado: (2024)