Fine-Tuned LLMs are "Time Capsules" for Tracking Societal Bias Through Books
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Madhusudan, Sangmitra, Morabito, Robert, Reid, Skye, Sadr, Nikta Gohari, Emami, Ali |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Which Words Matter Most in Zero-Shot Prompts?
von: Sadr, Nikta Gohari, et al.
Veröffentlicht: (2025)
von: Sadr, Nikta Gohari, et al.
Veröffentlicht: (2025)
STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions
von: Morabito, Robert, et al.
Veröffentlicht: (2024)
von: Morabito, Robert, et al.
Veröffentlicht: (2024)
The Dog the Cat Chased Stumped the Model: Measuring When Language Models Abandon Structure for Shortcuts
von: Madhusudan, Sangmitra, et al.
Veröffentlicht: (2025)
von: Madhusudan, Sangmitra, et al.
Veröffentlicht: (2025)
We Politely Insist: Your LLM Must Learn the Persian Art of Taarof
von: Sadr, Nikta Gohari, et al.
Veröffentlicht: (2025)
von: Sadr, Nikta Gohari, et al.
Veröffentlicht: (2025)
Common to Whom? Regional Cultural Commonsense and LLM Bias in India
von: Madhusudan, Sangmitra, et al.
Veröffentlicht: (2026)
von: Madhusudan, Sangmitra, et al.
Veröffentlicht: (2026)
Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models
von: Kumar, Abhishek, et al.
Veröffentlicht: (2024)
von: Kumar, Abhishek, et al.
Veröffentlicht: (2024)
Translate With Care: Addressing Gender Bias, Neutrality, and Reasoning in Large Language Model Translations
von: Zahraei, Pardis Sadat, et al.
Veröffentlicht: (2025)
von: Zahraei, Pardis Sadat, et al.
Veröffentlicht: (2025)
Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging
von: Farn, Hua, et al.
Veröffentlicht: (2024)
von: Farn, Hua, et al.
Veröffentlicht: (2024)
Tracking Universal Features Through Fine-Tuning and Model Merging
von: Horn, Niels, et al.
Veröffentlicht: (2024)
von: Horn, Niels, et al.
Veröffentlicht: (2024)
CoBia: Constructed Conversations Can Trigger Otherwise Concealed Societal Biases in LLMs
von: Nikeghbal, Nafiseh, et al.
Veröffentlicht: (2025)
von: Nikeghbal, Nafiseh, et al.
Veröffentlicht: (2025)
Augmented Fine-Tuned LLMs for Enhanced Recruitment Automation
von: Younes, Mohamed T., et al.
Veröffentlicht: (2025)
von: Younes, Mohamed T., et al.
Veröffentlicht: (2025)
Fine-Tuning LLMs for Reliable Medical Question-Answering Services
von: Anaissi, Ali, et al.
Veröffentlicht: (2024)
von: Anaissi, Ali, et al.
Veröffentlicht: (2024)
Trace-of-Thought Prompting: Investigating Prompt-Based Knowledge Distillation Through Question Decomposition
von: McDonald, Tyler, et al.
Veröffentlicht: (2025)
von: McDonald, Tyler, et al.
Veröffentlicht: (2025)
On The Conceptualization and Societal Impact of Cross-Cultural Bias
von: Bhandari, Vitthal
Veröffentlicht: (2025)
von: Bhandari, Vitthal
Veröffentlicht: (2025)
Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language Models
von: Kumar, Abhishek, et al.
Veröffentlicht: (2024)
von: Kumar, Abhishek, et al.
Veröffentlicht: (2024)
Fine-Tuning LLMs for Low-Resource Dialect Translation: The Case of Lebanese
von: Yakhni, Silvana, et al.
Veröffentlicht: (2025)
von: Yakhni, Silvana, et al.
Veröffentlicht: (2025)
DART: Mitigating Harm Drift in Difference-Aware LLMs via Distill-Audit-Repair Training
von: Pan, Ziwen, et al.
Veröffentlicht: (2026)
von: Pan, Ziwen, et al.
Veröffentlicht: (2026)
Memory Dial: A Training Framework for Controllable Memorization in Language Models
von: Zhang, Xiangbo, et al.
Veröffentlicht: (2026)
von: Zhang, Xiangbo, et al.
Veröffentlicht: (2026)
Position-Aware Parameter Efficient Fine-Tuning Approach for Reducing Positional Bias in LLMs
von: Zhang, Zheng, et al.
Veröffentlicht: (2024)
von: Zhang, Zheng, et al.
Veröffentlicht: (2024)
Metamorphic Testing for Fairness Evaluation in Large Language Models: Identifying Intersectional Bias in LLaMA and GPT
von: Reddy, Harishwar, et al.
Veröffentlicht: (2025)
von: Reddy, Harishwar, et al.
Veröffentlicht: (2025)
KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language Models
von: Kim, Seorin, et al.
Veröffentlicht: (2025)
von: Kim, Seorin, et al.
Veröffentlicht: (2025)
EvoGrad: A Dynamic Take on the Winograd Schema Challenge with Human Adversaries
von: Sun, Jing Han, et al.
Veröffentlicht: (2024)
von: Sun, Jing Han, et al.
Veröffentlicht: (2024)
Red-Teaming for Inducing Societal Bias in Large Language Models
von: Luo, Chu Fei, et al.
Veröffentlicht: (2024)
von: Luo, Chu Fei, et al.
Veröffentlicht: (2024)
Order-Independence Without Fine Tuning
von: McIlroy-Young, Reid, et al.
Veröffentlicht: (2024)
von: McIlroy-Young, Reid, et al.
Veröffentlicht: (2024)
QEFT: Quantization for Efficient Fine-Tuning of LLMs
von: Lee, Changhun, et al.
Veröffentlicht: (2024)
von: Lee, Changhun, et al.
Veröffentlicht: (2024)
CALF: Aligning LLMs for Time Series Forecasting via Cross-modal Fine-Tuning
von: Liu, Peiyuan, et al.
Veröffentlicht: (2024)
von: Liu, Peiyuan, et al.
Veröffentlicht: (2024)
Towards Pedagogical LLMs with Supervised Fine Tuning for Computing Education
von: Vassar, Alexandra, et al.
Veröffentlicht: (2024)
von: Vassar, Alexandra, et al.
Veröffentlicht: (2024)
Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
von: Gekhman, Zorik, et al.
Veröffentlicht: (2024)
von: Gekhman, Zorik, et al.
Veröffentlicht: (2024)
Parameter-Efficient Fine-Tuning of LLMs with Mixture of Space Experts
von: Zhang, Buze, et al.
Veröffentlicht: (2026)
von: Zhang, Buze, et al.
Veröffentlicht: (2026)
Dataset Scale and Societal Consistency Mediate Facial Impression Bias in Vision-Language AI
von: Wolfe, Robert, et al.
Veröffentlicht: (2024)
von: Wolfe, Robert, et al.
Veröffentlicht: (2024)
Towards Better Understanding of Cybercrime: The Role of Fine-Tuned LLMs in Translation
von: Valeros, Veronica, et al.
Veröffentlicht: (2024)
von: Valeros, Veronica, et al.
Veröffentlicht: (2024)
Language Bias under Conflicting Information in Multilingual LLMs
von: Östling, Robert, et al.
Veröffentlicht: (2026)
von: Östling, Robert, et al.
Veröffentlicht: (2026)
Supervised Fine-Tuning LLMs to Behave as Pedagogical Agents in Programming Education
von: Ross, Emily, et al.
Veröffentlicht: (2025)
von: Ross, Emily, et al.
Veröffentlicht: (2025)
Controlling Distributional Bias in Multi-Round LLM Generation via KL-Optimized Fine-Tuning
von: Jiang, Yanbei, et al.
Veröffentlicht: (2026)
von: Jiang, Yanbei, et al.
Veröffentlicht: (2026)
Fine-Tuning LLMs with Fine-Grained Human Feedback on Text Spans
von: CH-Wang, Sky, et al.
Veröffentlicht: (2025)
von: CH-Wang, Sky, et al.
Veröffentlicht: (2025)
A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models
von: Sathe, Ashutosh, et al.
Veröffentlicht: (2024)
von: Sathe, Ashutosh, et al.
Veröffentlicht: (2024)
Knowledge Capsules: Structured Nonparametric Memory Units for LLMs
von: Ju, Bin, et al.
Veröffentlicht: (2026)
von: Ju, Bin, et al.
Veröffentlicht: (2026)
From 'Showgirls' to 'Performers': Fine-tuning with Gender-inclusive Language for Bias Reduction in LLMs
von: Bartl, Marion, et al.
Veröffentlicht: (2024)
von: Bartl, Marion, et al.
Veröffentlicht: (2024)
IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages
von: Hanif, Ikhlasul Akmal, et al.
Veröffentlicht: (2026)
von: Hanif, Ikhlasul Akmal, et al.
Veröffentlicht: (2026)
Quantifying and Mitigating Selection Bias in LLMs: A Transferable LoRA Fine-Tuning and Efficient Majority Voting Approach
von: Guda, Blessed, et al.
Veröffentlicht: (2025)
von: Guda, Blessed, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Which Words Matter Most in Zero-Shot Prompts?
von: Sadr, Nikta Gohari, et al.
Veröffentlicht: (2025) -
STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions
von: Morabito, Robert, et al.
Veröffentlicht: (2024) -
The Dog the Cat Chased Stumped the Model: Measuring When Language Models Abandon Structure for Shortcuts
von: Madhusudan, Sangmitra, et al.
Veröffentlicht: (2025) -
We Politely Insist: Your LLM Must Learn the Persian Art of Taarof
von: Sadr, Nikta Gohari, et al.
Veröffentlicht: (2025) -
Common to Whom? Regional Cultural Commonsense and LLM Bias in India
von: Madhusudan, Sangmitra, et al.
Veröffentlicht: (2026)