Soft-prompt Tuning for Large Language Models to Evaluate Bias
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tian, Jacob-Junqi, Emerson, David, Miyandoab, Sevil Zanjani, Pandya, Deval, Seyyed-Kalantari, Laleh, Khattak, Faiza Khan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Template-Based Probes Are Imperfect Lenses for Counterfactual Bias Evaluation in LLMs
von: Kohankhaki, Farnaz, et al.
Veröffentlicht: (2024)
von: Kohankhaki, Farnaz, et al.
Veröffentlicht: (2024)
On The Role of Reasoning in the Identification of Subtle Stereotypes in Natural Language
von: Tian, Jacob-Junqi, et al.
Veröffentlicht: (2023)
von: Tian, Jacob-Junqi, et al.
Veröffentlicht: (2023)
Red-Teaming for Inducing Societal Bias in Large Language Models
von: Luo, Chu Fei, et al.
Veröffentlicht: (2024)
von: Luo, Chu Fei, et al.
Veröffentlicht: (2024)
Multi-objective Binary Coordinate Search for Feature Selection
von: Miyandoab, Sevil Zanjani, et al.
Veröffentlicht: (2024)
von: Miyandoab, Sevil Zanjani, et al.
Veröffentlicht: (2024)
Compact NSGA-II for Multi-objective Feature Selection
von: Miyandoab, Sevil Zanjani, et al.
Veröffentlicht: (2024)
von: Miyandoab, Sevil Zanjani, et al.
Veröffentlicht: (2024)
We Politely Insist: Your LLM Must Learn the Persian Art of Taarof
von: Sadr, Nikta Gohari, et al.
Veröffentlicht: (2025)
von: Sadr, Nikta Gohari, et al.
Veröffentlicht: (2025)
FAIR Enough: How Can We Develop and Assess a FAIR-Compliant Dataset for Large Language Models' Training?
von: Raza, Shaina, et al.
Veröffentlicht: (2024)
von: Raza, Shaina, et al.
Veröffentlicht: (2024)
Evaluation of Attribution Bias in Generator-Aware Retrieval-Augmented Large Language Models
von: Abolghasemi, Amin, et al.
Veröffentlicht: (2024)
von: Abolghasemi, Amin, et al.
Veröffentlicht: (2024)
Enhancing Diversity in Multi-objective Feature Selection
von: Miyandoab, Sevil Zanjani, et al.
Veröffentlicht: (2024)
von: Miyandoab, Sevil Zanjani, et al.
Veröffentlicht: (2024)
PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
von: Hosseini, Mohammad, et al.
Veröffentlicht: (2025)
von: Hosseini, Mohammad, et al.
Veröffentlicht: (2025)
Mitigating Social Biases in Language Models through Unlearning
von: Dige, Omkar, et al.
Veröffentlicht: (2024)
von: Dige, Omkar, et al.
Veröffentlicht: (2024)
Training Artificial Neural Networks by Coordinate Search Algorithm
von: Rokhsatyazdi, Ehsan, et al.
Veröffentlicht: (2024)
von: Rokhsatyazdi, Ehsan, et al.
Veröffentlicht: (2024)
Evaluating and Mitigating Social Bias for Large Language Models in Open-ended Settings
von: Liu, Zhao, et al.
Veröffentlicht: (2024)
von: Liu, Zhao, et al.
Veröffentlicht: (2024)
BiasCause: Evaluate Socially Biased Causal Reasoning of Large Language Models
von: Xie, Tian, et al.
Veröffentlicht: (2025)
von: Xie, Tian, et al.
Veröffentlicht: (2025)
Evaluating Gender Bias in Large Language Models
von: Döll, Michael, et al.
Veröffentlicht: (2024)
von: Döll, Michael, et al.
Veröffentlicht: (2024)
Mitigating the Bias of Large Language Model Evaluation
von: Zhou, Hongli, et al.
Veröffentlicht: (2024)
von: Zhou, Hongli, et al.
Veröffentlicht: (2024)
Nemesis: Normalizing the Soft-prompt Vectors of Vision-Language Models
von: Fu, Shuai, et al.
Veröffentlicht: (2024)
von: Fu, Shuai, et al.
Veröffentlicht: (2024)
Cultural Alignment in Large Language Models Using Soft Prompt Tuning
von: Masoud, Reem I., et al.
Veröffentlicht: (2025)
von: Masoud, Reem I., et al.
Veröffentlicht: (2025)
Just as Humans Need Vaccines, So Do Models: Model Immunization to Combat Falsehoods
von: Raza, Shaina, et al.
Veröffentlicht: (2025)
von: Raza, Shaina, et al.
Veröffentlicht: (2025)
The Few-shot Dilemma: Over-prompting Large Language Models
von: Tang, Yongjian, et al.
Veröffentlicht: (2025)
von: Tang, Yongjian, et al.
Veröffentlicht: (2025)
DeepSeek's WEIRD Behavior: The cultural alignment of Large Language Models and the effects of prompt language and cultural prompting
von: Luther, James, et al.
Veröffentlicht: (2025)
von: Luther, James, et al.
Veröffentlicht: (2025)
McBE: A Multi-task Chinese Bias Evaluation Benchmark for Large Language Models
von: Lan, Tian, et al.
Veröffentlicht: (2025)
von: Lan, Tian, et al.
Veröffentlicht: (2025)
Promptception: How Sensitive Are Large Multimodal Models to Prompts?
von: Ismithdeen, Mohamed Insaf, et al.
Veröffentlicht: (2025)
von: Ismithdeen, Mohamed Insaf, et al.
Veröffentlicht: (2025)
Talent or Luck? Evaluating Attribution Bias in Large Language Models
von: Raj, Chahat, et al.
Veröffentlicht: (2025)
von: Raj, Chahat, et al.
Veröffentlicht: (2025)
Bias in Large Language Models: Origin, Evaluation, and Mitigation
von: Guo, Yufei, et al.
Veröffentlicht: (2024)
von: Guo, Yufei, et al.
Veröffentlicht: (2024)
FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain
von: Deb, Rohan, et al.
Veröffentlicht: (2025)
von: Deb, Rohan, et al.
Veröffentlicht: (2025)
QA-prompting: Improving Summarization with Large Language Models using Question-Answering
von: Sinha, Neelabh
Veröffentlicht: (2025)
von: Sinha, Neelabh
Veröffentlicht: (2025)
Exploring the Influence of Label Aggregation on Minority Voices: Implications for Dataset Bias and Model Training
von: Pandya, Mugdha, et al.
Veröffentlicht: (2024)
von: Pandya, Mugdha, et al.
Veröffentlicht: (2024)
Mis-prompt: Benchmarking Large Language Models for Proactive Error Handling
von: Zeng, Jiayi, et al.
Veröffentlicht: (2025)
von: Zeng, Jiayi, et al.
Veröffentlicht: (2025)
Evaluating Large Language Models on Urdu Idiom Translation
von: Khan, Muhammad Farmal, et al.
Veröffentlicht: (2025)
von: Khan, Muhammad Farmal, et al.
Veröffentlicht: (2025)
Quantile Regression with Large Language Models for Price Prediction
von: Vedula, Nikhita, et al.
Veröffentlicht: (2025)
von: Vedula, Nikhita, et al.
Veröffentlicht: (2025)
Evaluating Nuanced Bias in Large Language Model Free Response Answers
von: Healey, Jennifer, et al.
Veröffentlicht: (2024)
von: Healey, Jennifer, et al.
Veröffentlicht: (2024)
Evaluating the Presence of Sex Bias in Clinical Reasoning by Large Language Models
von: Tsintsiper, Isabel, et al.
Veröffentlicht: (2026)
von: Tsintsiper, Isabel, et al.
Veröffentlicht: (2026)
Social Bias Evaluation for Large Language Models Requires Prompt Variations
von: Hida, Rem, et al.
Veröffentlicht: (2024)
von: Hida, Rem, et al.
Veröffentlicht: (2024)
Soft Prompt Tuning for Augmenting Dense Retrieval with Large Language Models
von: Peng, Zhiyuan, et al.
Veröffentlicht: (2023)
von: Peng, Zhiyuan, et al.
Veröffentlicht: (2023)
Bias in, Bias out: Annotation Bias in Multilingual Large Language Models
von: Cui, Xia, et al.
Veröffentlicht: (2025)
von: Cui, Xia, et al.
Veröffentlicht: (2025)
Likelihood-based Mitigation of Evaluation Bias in Large Language Models
von: Oi, Masanari, et al.
Veröffentlicht: (2024)
von: Oi, Masanari, et al.
Veröffentlicht: (2024)
LFTF: Locating First and Then Fine-Tuning for Mitigating Gender Bias in Large Language Models
von: Qin, Zhanyue, et al.
Veröffentlicht: (2025)
von: Qin, Zhanyue, et al.
Veröffentlicht: (2025)
What is Your Favorite Gender, MLM? Gender Bias Evaluation in Multilingual Masked Language Models
von: Yu, Jeongrok, et al.
Veröffentlicht: (2024)
von: Yu, Jeongrok, et al.
Veröffentlicht: (2024)
No LLM is Free From Bias: A Comprehensive Study of Bias Evaluation in Large Language Models
von: Kumar, Charaka Vinayak, et al.
Veröffentlicht: (2025)
von: Kumar, Charaka Vinayak, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Template-Based Probes Are Imperfect Lenses for Counterfactual Bias Evaluation in LLMs
von: Kohankhaki, Farnaz, et al.
Veröffentlicht: (2024) -
On The Role of Reasoning in the Identification of Subtle Stereotypes in Natural Language
von: Tian, Jacob-Junqi, et al.
Veröffentlicht: (2023) -
Red-Teaming for Inducing Societal Bias in Large Language Models
von: Luo, Chu Fei, et al.
Veröffentlicht: (2024) -
Multi-objective Binary Coordinate Search for Feature Selection
von: Miyandoab, Sevil Zanjani, et al.
Veröffentlicht: (2024) -
Compact NSGA-II for Multi-objective Feature Selection
von: Miyandoab, Sevil Zanjani, et al.
Veröffentlicht: (2024)