C2PO: Diagnosing and Disentangling Bias Shortcuts in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Xuan, An, Bo, Gu, Tianlong, Chang, Liang, Hao, Fengrui, Yu, Peipeng, Zhao, Shuai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Debias: Self-correcting for Debiasing Large Language Models
by: Feng, Xuan, et al.
Published: (2026)
by: Feng, Xuan, et al.
Published: (2026)
Learning from Mistakes: Self-correct Adversarial Training for Chinese Unnatural Text Correction
by: Feng, Xuan, et al.
Published: (2024)
by: Feng, Xuan, et al.
Published: (2024)
CHEAT: A Large-scale Dataset for Detecting ChatGPT-writtEn AbsTracts
by: Yu, Peipeng, et al.
Published: (2023)
by: Yu, Peipeng, et al.
Published: (2023)
Do LLMs Overcome Shortcut Learning? An Evaluation of Shortcut Challenges in Large Language Models
by: Yuan, Yu, et al.
Published: (2024)
by: Yuan, Yu, et al.
Published: (2024)
Chimera: Diagnosing Shortcut Learning in Visual-Language Understanding
by: Chi, Ziheng, et al.
Published: (2025)
by: Chi, Ziheng, et al.
Published: (2025)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
by: Marioriyad, Arash, et al.
Published: (2025)
by: Marioriyad, Arash, et al.
Published: (2025)
Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs
by: Ling Team, et al.
Published: (2025)
by: Ling Team, et al.
Published: (2025)
Diagnosing and Mitigating System Bias in Self-Rewarding RL
by: Tan, Chuyi, et al.
Published: (2025)
by: Tan, Chuyi, et al.
Published: (2025)
HiPO: Hybrid Policy Optimization for Dynamic Reasoning in LLMs
by: Deng, Ken, et al.
Published: (2025)
by: Deng, Ken, et al.
Published: (2025)
Disclosure and Mitigation of Gender Bias in LLMs
by: Dong, Xiangjue, et al.
Published: (2024)
by: Dong, Xiangjue, et al.
Published: (2024)
MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs
by: Taghanaki, Saeid Asgari, et al.
Published: (2024)
by: Taghanaki, Saeid Asgari, et al.
Published: (2024)
Automate Knowledge Concept Tagging on Math Questions with LLMs
by: Li, Hang, et al.
Published: (2024)
by: Li, Hang, et al.
Published: (2024)
The Bias is in the Details: An Assessment of Cognitive Bias in LLMs
by: Knipper, R. Alexander, et al.
Published: (2025)
by: Knipper, R. Alexander, et al.
Published: (2025)
LLMs Are Biased Towards Output Formats! Systematically Evaluating and Mitigating Output Format Bias of LLMs
by: Long, Do Xuan, et al.
Published: (2024)
by: Long, Do Xuan, et al.
Published: (2024)
Cross-Cultural Expert-Level Art Critique Evaluation with Vision-Language Models
by: Yu, Haorui, et al.
Published: (2026)
by: Yu, Haorui, et al.
Published: (2026)
Bot Meets Shortcut: How Can LLMs Aid in Handling Unknown Invariance OOD Scenarios?
by: Zheng, Shiyan, et al.
Published: (2025)
by: Zheng, Shiyan, et al.
Published: (2025)
Shortcut Learning in In-Context Learning: A Survey
by: Song, Rui, et al.
Published: (2024)
by: Song, Rui, et al.
Published: (2024)
Short-circuiting Shortcuts: Mechanistic Investigation of Shortcuts in Text Classification
by: Eshuijs, Leon, et al.
Published: (2025)
by: Eshuijs, Leon, et al.
Published: (2025)
Rectifying Demonstration Shortcut in In-Context Learning
by: Jang, Joonwon, et al.
Published: (2024)
by: Jang, Joonwon, et al.
Published: (2024)
DLF: Disentangled-Language-Focused Multimodal Sentiment Analysis
by: Wang, Pan, et al.
Published: (2024)
by: Wang, Pan, et al.
Published: (2024)
Leveraging LLMs for Predicting Unknown Diagnoses from Clinical Notes
by: Albassam, Dina, et al.
Published: (2025)
by: Albassam, Dina, et al.
Published: (2025)
The Biased Oracle: Assessing LLMs' Understandability and Empathy in Medical Diagnoses
by: Yao, Jianzhou, et al.
Published: (2025)
by: Yao, Jianzhou, et al.
Published: (2025)
DLO: Dynamic Layer Operation for Efficient Vertical Scaling of LLMs
by: Tan, Zhen, et al.
Published: (2024)
by: Tan, Zhen, et al.
Published: (2024)
Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis
by: Zhu, Kejian, et al.
Published: (2025)
by: Zhu, Kejian, et al.
Published: (2025)
Understanding and Mitigating Gender Bias in LLMs via Interpretable Neuron Editing
by: Yu, Zeping, et al.
Published: (2025)
by: Yu, Zeping, et al.
Published: (2025)
Speculative Reward Model Boosts Decision Making Ability of LLMs Cost-Effectively
by: Gu, Jiawei, et al.
Published: (2025)
by: Gu, Jiawei, et al.
Published: (2025)
Disentangling Dialect from Social Bias via Multitask Learning to Improve Fairness
by: Spliethöver, Maximilian, et al.
Published: (2024)
by: Spliethöver, Maximilian, et al.
Published: (2024)
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
by: Yan, Lecheng, et al.
Published: (2026)
by: Yan, Lecheng, et al.
Published: (2026)
Shattering the Shortcut: A Topology-Regularized Benchmark for Multi-hop Medical Reasoning in LLMs
by: Zi, Xing, et al.
Published: (2026)
by: Zi, Xing, et al.
Published: (2026)
Evaluating the Bias in LLMs for Surveying Opinion and Decision Making in Healthcare
by: Khaokaew, Yonchanok, et al.
Published: (2025)
by: Khaokaew, Yonchanok, et al.
Published: (2025)
Probing Social Identity Bias in Chinese LLMs with Gendered Pronouns and Social Groups
by: Liu, Geng, et al.
Published: (2025)
by: Liu, Geng, et al.
Published: (2025)
The Gold Medals in an Empty Room: Diagnosing Metalinguistic Reasoning in LLMs with Camlang
by: Liu, Fenghua, et al.
Published: (2025)
by: Liu, Fenghua, et al.
Published: (2025)
From Parameters to Prompts: Understanding and Mitigating the Factuality Gap between Fine-Tuned LLMs
by: Gong, Xuan, et al.
Published: (2025)
by: Gong, Xuan, et al.
Published: (2025)
Affective-ROPTester: Capability and Bias Analysis of LLMs in Predicting Retinopathy of Prematurity
by: Zhao, Shuai, et al.
Published: (2025)
by: Zhao, Shuai, et al.
Published: (2025)
MIST: Towards Multi-dimensional Implicit BiaS Evaluation of LLMs for Theory of Mind
by: Li, Yanlin, et al.
Published: (2025)
by: Li, Yanlin, et al.
Published: (2025)
When Less Language is More: Language-Reasoning Disentanglement Makes LLMs Better Multilingual Reasoners
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
FocalPO: Enhancing Preference Optimizing by Focusing on Correct Preference Rankings
by: Liu, Tong, et al.
Published: (2025)
by: Liu, Tong, et al.
Published: (2025)
Implicit Reasoning in Transformers is Reasoning through Shortcuts
by: Lin, Tianhe, et al.
Published: (2025)
by: Lin, Tianhe, et al.
Published: (2025)
BiasAlert: A Plug-and-play Tool for Social Bias Detection in LLMs
by: Fan, Zhiting, et al.
Published: (2024)
by: Fan, Zhiting, et al.
Published: (2024)
Empowering LLMs with Parameterized Skills for Adversarial Long-Horizon Planning
by: Cui, Sijia, et al.
Published: (2025)
by: Cui, Sijia, et al.
Published: (2025)
Similar Items
-
Self-Debias: Self-correcting for Debiasing Large Language Models
by: Feng, Xuan, et al.
Published: (2026) -
Learning from Mistakes: Self-correct Adversarial Training for Chinese Unnatural Text Correction
by: Feng, Xuan, et al.
Published: (2024) -
CHEAT: A Large-scale Dataset for Detecting ChatGPT-writtEn AbsTracts
by: Yu, Peipeng, et al.
Published: (2023) -
Do LLMs Overcome Shortcut Learning? An Evaluation of Shortcut Challenges in Large Language Models
by: Yuan, Yu, et al.
Published: (2024) -
Chimera: Diagnosing Shortcut Learning in Visual-Language Understanding
by: Chi, Ziheng, et al.
Published: (2025)