BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Islam, Sekh Mainul, Borenstein, Nadav, Pawar, Siddhesh Milind, Yu, Haeun, Arora, Arnav, Augenstein, Isabelle |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Investigating Human Values in Online Communities
von: Borenstein, Nadav, et al.
Veröffentlicht: (2024)
von: Borenstein, Nadav, et al.
Veröffentlicht: (2024)
Presumed Cultural Identity: How Names Shape LLM Responses
von: Pawar, Siddhesh, et al.
Veröffentlicht: (2025)
von: Pawar, Siddhesh, et al.
Veröffentlicht: (2025)
Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement
von: Islam, Sekh Mainul, et al.
Veröffentlicht: (2025)
von: Islam, Sekh Mainul, et al.
Veröffentlicht: (2025)
Revealing Fine-Grained Values and Opinions in Large Language Models
von: Wright, Dustin, et al.
Veröffentlicht: (2024)
von: Wright, Dustin, et al.
Veröffentlicht: (2024)
Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models
von: Yu, Haeun, et al.
Veröffentlicht: (2025)
von: Yu, Haeun, et al.
Veröffentlicht: (2025)
Evaluation Framework for Highlight Explanations of Context Utilisation in Language Models
von: Sun, Jingyi, et al.
Veröffentlicht: (2025)
von: Sun, Jingyi, et al.
Veröffentlicht: (2025)
Not What, But How: A Communicative Audit of LLM Response Framing
von: Pawar, Siddhesh Milind, et al.
Veröffentlicht: (2026)
von: Pawar, Siddhesh Milind, et al.
Veröffentlicht: (2026)
Can Community Notes Replace Professional Fact-Checkers?
von: Borenstein, Nadav, et al.
Veröffentlicht: (2025)
von: Borenstein, Nadav, et al.
Veröffentlicht: (2025)
Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution Methods
von: Yu, Haeun, et al.
Veröffentlicht: (2024)
von: Yu, Haeun, et al.
Veröffentlicht: (2024)
Can Transformers Learn $n$-gram Language Models?
von: Svete, Anej, et al.
Veröffentlicht: (2024)
von: Svete, Anej, et al.
Veröffentlicht: (2024)
Probing Pre-Trained Language Models for Cross-Cultural Differences in Values
von: Arora, Arnav, et al.
Veröffentlicht: (2022)
von: Arora, Arnav, et al.
Veröffentlicht: (2022)
Why Should This Article Be Deleted? Transparent Stance Detection in Multilingual Wikipedia Editor Discussions
von: Kaffee, Lucie-Aimée, et al.
Veröffentlicht: (2023)
von: Kaffee, Lucie-Aimée, et al.
Veröffentlicht: (2023)
Revisiting Noise in Natural Language Processing for Computational Social Science
von: Borenstein, Nadav
Veröffentlicht: (2025)
von: Borenstein, Nadav
Veröffentlicht: (2025)
Survey of Cultural Awareness in Language Models: Text and Beyond
von: Pawar, Siddhesh, et al.
Veröffentlicht: (2024)
von: Pawar, Siddhesh, et al.
Veröffentlicht: (2024)
A Reality Check on Context Utilisation for Retrieval-Augmented Generation
von: Hagström, Lovisa, et al.
Veröffentlicht: (2024)
von: Hagström, Lovisa, et al.
Veröffentlicht: (2024)
Multi-Modal Framing Analysis of News
von: Arora, Arnav, et al.
Veröffentlicht: (2025)
von: Arora, Arnav, et al.
Veröffentlicht: (2025)
What Languages are Easy to Language-Model? A Perspective from Learning Probabilistic Regular Languages
von: Borenstein, Nadav, et al.
Veröffentlicht: (2024)
von: Borenstein, Nadav, et al.
Veröffentlicht: (2024)
Quantifying Gender Biases Towards Politicians on Reddit
von: Marjanovic, Sara, et al.
Veröffentlicht: (2021)
von: Marjanovic, Sara, et al.
Veröffentlicht: (2021)
Specializing Large Language Models to Simulate Survey Response Distributions for Global Populations
von: Cao, Yong, et al.
Veröffentlicht: (2025)
von: Cao, Yong, et al.
Veröffentlicht: (2025)
Factcheck-Bench: Fine-Grained Evaluation Benchmark for Automatic Fact-checkers
von: Wang, Yuxia, et al.
Veröffentlicht: (2023)
von: Wang, Yuxia, et al.
Veröffentlicht: (2023)
Evaluating Input Feature Explanations through a Unified Diagnostic Evaluation Framework
von: Sun, Jingyi, et al.
Veröffentlicht: (2024)
von: Sun, Jingyi, et al.
Veröffentlicht: (2024)
Understanding the Interplay between LLMs' Utilisation of Parametric and Contextual Knowledge: A keynote at ECIR 2025
von: Augenstein, Isabelle
Veröffentlicht: (2026)
von: Augenstein, Isabelle
Veröffentlicht: (2026)
DYNAMICQA: Tracing Internal Knowledge Conflicts in Language Models
von: Marjanović, Sara Vera, et al.
Veröffentlicht: (2024)
von: Marjanović, Sara Vera, et al.
Veröffentlicht: (2024)
ReqElicitGym: An Evaluation Environment for Interview Competence in Conversational Requirements Elicitation
von: Jin, Dongming, et al.
Veröffentlicht: (2026)
von: Jin, Dongming, et al.
Veröffentlicht: (2026)
Social Bias Probing: Fairness Benchmarking for Language Models
von: Manerba, Marta Marchiori, et al.
Veröffentlicht: (2023)
von: Manerba, Marta Marchiori, et al.
Veröffentlicht: (2023)
CUB: Benchmarking Context Utilisation Techniques for Language Models
von: Hagström, Lovisa, et al.
Veröffentlicht: (2025)
von: Hagström, Lovisa, et al.
Veröffentlicht: (2025)
Invisible Women in Digital Diplomacy: A Multidimensional Framework for Online Gender Bias Against Women Ambassadors Worldwide
von: Golovchenko, Yevgeniy, et al.
Veröffentlicht: (2023)
von: Golovchenko, Yevgeniy, et al.
Veröffentlicht: (2023)
Aggregating Soft Labels from Crowd Annotations Improves Uncertainty Estimation Under Distribution Shift
von: Wright, Dustin, et al.
Veröffentlicht: (2022)
von: Wright, Dustin, et al.
Veröffentlicht: (2022)
TOWARDS INTEGRATING RETIMING IN VEHICLE TYPE SCHEDULING PROBLEM
von: Denis Borenstein
Veröffentlicht: (2017)
von: Denis Borenstein
Veröffentlicht: (2017)
Un‐clearance
von: Marc Borenstein
Veröffentlicht: (2025)
von: Marc Borenstein
Veröffentlicht: (2025)
Effects of Soaking, Roasting, and Germination on Saponin Reduction and Nutritional Enhancement in Quinoa ( Chenopodium quinoa )
von: Pooja Rani, et al.
Veröffentlicht: (2026)
von: Pooja Rani, et al.
Veröffentlicht: (2026)
Beyond the Battlefield: Framing Analysis of Media Coverage in Conflict Reporting
von: Kaur, Avneet, et al.
Veröffentlicht: (2025)
von: Kaur, Avneet, et al.
Veröffentlicht: (2025)
Quantitative Currency Evaluation in Low-Resource Settings through Pattern Analysis to Assist Visually Impaired Users
von: Ovi, Md Sultanul Islam, et al.
Veröffentlicht: (2025)
von: Ovi, Md Sultanul Islam, et al.
Veröffentlicht: (2025)
CausalGym: Benchmarking causal interpretability methods on linguistic tasks
von: Arora, Aryaman, et al.
Veröffentlicht: (2024)
von: Arora, Aryaman, et al.
Veröffentlicht: (2024)
Show Me the Work: Fact-Checkers' Requirements for Explainable Automated Fact-Checking
von: Warren, Greta, et al.
Veröffentlicht: (2025)
von: Warren, Greta, et al.
Veröffentlicht: (2025)
Semantic Sensitivities and Inconsistent Predictions: Measuring the Fragility of NLI Models
von: Arakelyan, Erik, et al.
Veröffentlicht: (2024)
von: Arakelyan, Erik, et al.
Veröffentlicht: (2024)
Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation
von: Resck, Lucas, et al.
Veröffentlicht: (2026)
von: Resck, Lucas, et al.
Veröffentlicht: (2026)
Expanding Computation Spaces of LLMs at Inference Time
von: Jang, Yoonna, et al.
Veröffentlicht: (2025)
von: Jang, Yoonna, et al.
Veröffentlicht: (2025)
JBBQ: Japanese Bias Benchmark for Analyzing Social Biases in Large Language Models
von: Yanaka, Hitomi, et al.
Veröffentlicht: (2024)
von: Yanaka, Hitomi, et al.
Veröffentlicht: (2024)
BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models
von: Lee, Isack, et al.
Veröffentlicht: (2024)
von: Lee, Isack, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Investigating Human Values in Online Communities
von: Borenstein, Nadav, et al.
Veröffentlicht: (2024) -
Presumed Cultural Identity: How Names Shape LLM Responses
von: Pawar, Siddhesh, et al.
Veröffentlicht: (2025) -
Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement
von: Islam, Sekh Mainul, et al.
Veröffentlicht: (2025) -
Revealing Fine-Grained Values and Opinions in Large Language Models
von: Wright, Dustin, et al.
Veröffentlicht: (2024) -
Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models
von: Yu, Haeun, et al.
Veröffentlicht: (2025)