Mitigating LLM biases toward spurious social contexts using direct preference optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nam, Hyunji, Demszky, Dorottya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Edu-ConvoKit: An Open-Source Library for Education Conversation Data
von: Wang, Rose E., et al.
Veröffentlicht: (2024)
von: Wang, Rose E., et al.
Veröffentlicht: (2024)
TeachLM: Post-Training LLMs for Education Using Authentic Learning Data
von: Perczel, Janos, et al.
Veröffentlicht: (2025)
von: Perczel, Janos, et al.
Veröffentlicht: (2025)
IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations
von: Nam, Hyunji, et al.
Veröffentlicht: (2025)
von: Nam, Hyunji, et al.
Veröffentlicht: (2025)
Mapping the Methodological Space of Classroom Interaction Research: Scale, Duration, and Modality in an Age of AI
von: Demszky, Dorottya, et al.
Veröffentlicht: (2026)
von: Demszky, Dorottya, et al.
Veröffentlicht: (2026)
Problem-Oriented Segmentation and Retrieval: Case Study on Tutoring Conversations
von: Wang, Rose E., et al.
Veröffentlicht: (2024)
von: Wang, Rose E., et al.
Veröffentlicht: (2024)
Bridging the Novice-Expert Gap via Models of Decision-Making: A Case Study on Remediating Math Mistakes
von: Wang, Rose E., et al.
Veröffentlicht: (2023)
von: Wang, Rose E., et al.
Veröffentlicht: (2023)
Efficient RL for optimizing conversation level outcomes with an LLM-based tutor
von: Nam, Hyunji, et al.
Veröffentlicht: (2025)
von: Nam, Hyunji, et al.
Veröffentlicht: (2025)
Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional Data
von: Nam, Hyunji, et al.
Veröffentlicht: (2026)
von: Nam, Hyunji, et al.
Veröffentlicht: (2026)
Continued Pretraining for Domain Adaptation of Wav2vec2.0 in Automatic Speech Recognition for Elementary Math Classroom Settings
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2024)
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2024)
Preference learning in shades of gray: Interpretable and bias-aware reward modeling for human preferences
von: Oprea, Simona-Vasilica, et al.
Veröffentlicht: (2026)
von: Oprea, Simona-Vasilica, et al.
Veröffentlicht: (2026)
Marked Pedagogies: Examining Linguistic Biases in Personalized Automated Writing Feedback
von: Tan, Mei, et al.
Veröffentlicht: (2026)
von: Tan, Mei, et al.
Veröffentlicht: (2026)
REFA: Reference Free Alignment for multi-preference optimization
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
CORG: Generating Answers from Complex, Interrelated Contexts
von: Lee, Hyunji, et al.
Veröffentlicht: (2025)
von: Lee, Hyunji, et al.
Veröffentlicht: (2025)
Fine-Grained Bias Detection in LLM: Enhancing detection mechanisms for nuanced biases
von: Mohanty, Suvendu
Veröffentlicht: (2025)
von: Mohanty, Suvendu
Veröffentlicht: (2025)
EduCoder: An Open-Source Annotation System for Education Transcript Data
von: Ashraf, Saad, et al.
Veröffentlicht: (2025)
von: Ashraf, Saad, et al.
Veröffentlicht: (2025)
Exploring the generalization of LLM truth directions on conversational formats
von: Ichmoukhamedov, Timour, et al.
Veröffentlicht: (2025)
von: Ichmoukhamedov, Timour, et al.
Veröffentlicht: (2025)
Search Wisely: Mitigating Sub-optimal Agentic Searches By Reducing Uncertainty
von: Wu, Peilin, et al.
Veröffentlicht: (2025)
von: Wu, Peilin, et al.
Veröffentlicht: (2025)
AI evaluation may bias perceptions: The importance of context in interpreting academic writing
von: Wu, Shang, et al.
Veröffentlicht: (2026)
von: Wu, Shang, et al.
Veröffentlicht: (2026)
Exploring LLM biases to manipulate AI search overview
von: Smirnov, Roman
Veröffentlicht: (2026)
von: Smirnov, Roman
Veröffentlicht: (2026)
Generalized Correctness Models: Learning Calibrated and Model-Agnostic Correctness Predictors from Historical Patterns
von: Xiao, Hanqi, et al.
Veröffentlicht: (2025)
von: Xiao, Hanqi, et al.
Veröffentlicht: (2025)
Amphista: Bi-directional Multi-head Decoding for Accelerating LLM Inference
von: Li, Zeping, et al.
Veröffentlicht: (2024)
von: Li, Zeping, et al.
Veröffentlicht: (2024)
Training With "Paraphrasing the Original Text" Teaches LLM to Better Retrieve in Long-context Tasks
von: Yu, Yijiong, et al.
Veröffentlicht: (2023)
von: Yu, Yijiong, et al.
Veröffentlicht: (2023)
Communication and Verification in LLM Agents towards Collaboration under Information Asymmetry
von: Peng, Run, et al.
Veröffentlicht: (2025)
von: Peng, Run, et al.
Veröffentlicht: (2025)
MARS: toward more efficient multi-agent collaboration for LLM reasoning
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
Meaningless is better: hashing bias-inducing words in LLM prompts improves performance in logical reasoning and statistical learning
von: Chadimová, Milena, et al.
Veröffentlicht: (2024)
von: Chadimová, Milena, et al.
Veröffentlicht: (2024)
I-CALM: Incentivizing Confidence-Aware Abstention for LLM Hallucination Mitigation
von: Zong, Haotian, et al.
Veröffentlicht: (2026)
von: Zong, Haotian, et al.
Veröffentlicht: (2026)
Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity
von: Zhang, Jiayi, et al.
Veröffentlicht: (2025)
von: Zhang, Jiayi, et al.
Veröffentlicht: (2025)
LLM-Guided Synthetic Augmentation (LGSA) for Mitigating Bias in AI Systems
von: Karri, Sai Suhruth Reddy, et al.
Veröffentlicht: (2025)
von: Karri, Sai Suhruth Reddy, et al.
Veröffentlicht: (2025)
Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge
von: Fujinuma, Yoshinari
Veröffentlicht: (2025)
von: Fujinuma, Yoshinari
Veröffentlicht: (2025)
Differential Information Distribution: A Bayesian Perspective on Direct Preference Optimization
von: Won, Yunjae, et al.
Veröffentlicht: (2025)
von: Won, Yunjae, et al.
Veröffentlicht: (2025)
Long-context LLMs Struggle with Long In-context Learning
von: Li, Tianle, et al.
Veröffentlicht: (2024)
von: Li, Tianle, et al.
Veröffentlicht: (2024)
Backtracking When It Strays: Mitigating Dual Exposure Biases in LLM Reasoning Distillation
von: Wang, Bing, et al.
Veröffentlicht: (2026)
von: Wang, Bing, et al.
Veröffentlicht: (2026)
Mitigating Translationese Bias in Multilingual LLM-as-a-Judge via Disentangled Information Bottleneck
von: Zhang, Hongbin, et al.
Veröffentlicht: (2026)
von: Zhang, Hongbin, et al.
Veröffentlicht: (2026)
HDLCoRe: A Training-Free Framework for Mitigating Hallucinations in LLM-Generated HDL
von: Ping, Heng, et al.
Veröffentlicht: (2025)
von: Ping, Heng, et al.
Veröffentlicht: (2025)
UniBias: Unveiling and Mitigating LLM Bias through Internal Attention and FFN Manipulation
von: Zhou, Hanzhang, et al.
Veröffentlicht: (2024)
von: Zhou, Hanzhang, et al.
Veröffentlicht: (2024)
Read Before You Think: Mitigating LLM Comprehension Failures with Step-by-Step Reading
von: Han, Feijiang, et al.
Veröffentlicht: (2025)
von: Han, Feijiang, et al.
Veröffentlicht: (2025)
How Well Do Large Language Models Truly Ground?
von: Lee, Hyunji, et al.
Veröffentlicht: (2023)
von: Lee, Hyunji, et al.
Veröffentlicht: (2023)
Semiparametric Token-Sequence Co-Supervision
von: Lee, Hyunji, et al.
Veröffentlicht: (2024)
von: Lee, Hyunji, et al.
Veröffentlicht: (2024)
MOSLIM:Align with diverse preferences in prompts through reward classification
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
To Bias or Not to Bias: Detecting bias in News with bias-detector
von: Ghosh, Himel, et al.
Veröffentlicht: (2025)
von: Ghosh, Himel, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Edu-ConvoKit: An Open-Source Library for Education Conversation Data
von: Wang, Rose E., et al.
Veröffentlicht: (2024) -
TeachLM: Post-Training LLMs for Education Using Authentic Learning Data
von: Perczel, Janos, et al.
Veröffentlicht: (2025) -
IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations
von: Nam, Hyunji, et al.
Veröffentlicht: (2025) -
Mapping the Methodological Space of Classroom Interaction Research: Scale, Duration, and Modality in an Age of AI
von: Demszky, Dorottya, et al.
Veröffentlicht: (2026) -
Problem-Oriented Segmentation and Retrieval: Case Study on Tutoring Conversations
von: Wang, Rose E., et al.
Veröffentlicht: (2024)