Post-hoc Reward Calibration: A Case Study on Length Bias
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Zeyu, Qiu, Zihan, Wang, Zili, Ponti, Edoardo M., Titov, Ivan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
di: Huang, Zeyu, et al.
Pubblicazione: (2025)
di: Huang, Zeyu, et al.
Pubblicazione: (2025)
Understanding Post-hoc Explainers: The Case of Anchors
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2023)
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2023)
Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods
di: Dhaini, Mahdi, et al.
Pubblicazione: (2025)
di: Dhaini, Mahdi, et al.
Pubblicazione: (2025)
Emergent Communication Pretraining for Few-Shot Machine Translation
di: Li, Yaoyiran, et al.
Pubblicazione: (2020)
di: Li, Yaoyiran, et al.
Pubblicazione: (2020)
Probing the Emergence of Cross-lingual Alignment during LLM Training
di: Wang, Hetong, et al.
Pubblicazione: (2024)
di: Wang, Hetong, et al.
Pubblicazione: (2024)
Unlearning Traces the Influential Training Data of Language Models
di: Isonuma, Masaru, et al.
Pubblicazione: (2024)
di: Isonuma, Masaru, et al.
Pubblicazione: (2024)
M-Wanda: Improving One-Shot Pruning for Multilingual LLMs
di: Choenni, Rochelle, et al.
Pubblicazione: (2025)
di: Choenni, Rochelle, et al.
Pubblicazione: (2025)
Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Models
di: Zhao, Dachuan, et al.
Pubblicazione: (2025)
di: Zhao, Dachuan, et al.
Pubblicazione: (2025)
Layerwise Recurrent Router for Mixture-of-Experts
di: Qiu, Zihan, et al.
Pubblicazione: (2024)
di: Qiu, Zihan, et al.
Pubblicazione: (2024)
Scaling Sparse Fine-Tuning to Large Language Models
di: Ansell, Alan, et al.
Pubblicazione: (2024)
di: Ansell, Alan, et al.
Pubblicazione: (2024)
Controlling What You Share: Assessing Language Model Adherence to Privacy Preferences
di: Ramírez, Guillem, et al.
Pubblicazione: (2025)
di: Ramírez, Guillem, et al.
Pubblicazione: (2025)
What's New in My Data? Novelty Exploration via Contrastive Generation
di: Isonuma, Masaru, et al.
Pubblicazione: (2024)
di: Isonuma, Masaru, et al.
Pubblicazione: (2024)
Bootstrapping Action-Grounded Visual Dynamics in Unified Vision-Language Models
di: Qiu, Yifu, et al.
Pubblicazione: (2025)
di: Qiu, Yifu, et al.
Pubblicazione: (2025)
Analysing Chain of Thought Dynamics: Active Guidance or Unfaithful Post-hoc Rationalisation?
di: Lewis-Lim, Samuel, et al.
Pubblicazione: (2025)
di: Lewis-Lim, Samuel, et al.
Pubblicazione: (2025)
A Survey on Multilingual Large Language Models: Corpora, Alignment, and Bias
di: Xu, Yuemei, et al.
Pubblicazione: (2024)
di: Xu, Yuemei, et al.
Pubblicazione: (2024)
AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
di: Liu, Zihan, et al.
Pubblicazione: (2024)
di: Liu, Zihan, et al.
Pubblicazione: (2024)
Spectral Editing of Activations for Large Language Model Alignment
di: Qiu, Yifu, et al.
Pubblicazione: (2024)
di: Qiu, Yifu, et al.
Pubblicazione: (2024)
More Thinking, More Bias: Length-Driven Position Bias in Reasoning Models
di: Wang, Xiao
Pubblicazione: (2026)
di: Wang, Xiao
Pubblicazione: (2026)
Anthropomimetic Uncertainty: What Verbalized Uncertainty in Language Models is Missing
di: Ulmer, Dennis, et al.
Pubblicazione: (2025)
di: Ulmer, Dennis, et al.
Pubblicazione: (2025)
AdaThink-Med: Medical Adaptive Thinking with Uncertainty-Guided Length Calibration
di: Rui, Shaohao, et al.
Pubblicazione: (2025)
di: Rui, Shaohao, et al.
Pubblicazione: (2025)
Post-hoc Utterance Refining Method by Entity Mining for Faithful Knowledge Grounded Conversations
di: Jang, Yoonna, et al.
Pubblicazione: (2024)
di: Jang, Yoonna, et al.
Pubblicazione: (2024)
Act-Adaptive Margin: Dynamically Calibrating Reward Models for Subjective Ambiguity
di: Fang, Feiteng, et al.
Pubblicazione: (2025)
di: Fang, Feiteng, et al.
Pubblicazione: (2025)
Mitigating Length Bias in RLHF through a Causal Lens
di: Kim, Hyeonji, et al.
Pubblicazione: (2025)
di: Kim, Hyeonji, et al.
Pubblicazione: (2025)
Detecting Prefix Bias in LLM-based Reward Models
di: Kumar, Ashwin, et al.
Pubblicazione: (2025)
di: Kumar, Ashwin, et al.
Pubblicazione: (2025)
Explaining Black-box Language Models with Knowledge Probing Systems: A Post-hoc Explanation Perspective
di: Zhao, Yunxiao, et al.
Pubblicazione: (2025)
di: Zhao, Yunxiao, et al.
Pubblicazione: (2025)
Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-Training
di: Du, Wenyu, et al.
Pubblicazione: (2024)
di: Du, Wenyu, et al.
Pubblicazione: (2024)
One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models
di: Fein, Daniel, et al.
Pubblicazione: (2026)
di: Fein, Daniel, et al.
Pubblicazione: (2026)
Shared Doubt: Zero-shot Cross-Lingual Confidence Estimation for Language Models
di: Kyriakou, Athina, et al.
Pubblicazione: (2026)
di: Kyriakou, Athina, et al.
Pubblicazione: (2026)
Joint Localization and Activation Editing for Low-Resource Fine-Tuning
di: Lai, Wen, et al.
Pubblicazione: (2025)
di: Lai, Wen, et al.
Pubblicazione: (2025)
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
di: Liu, Wei, et al.
Pubblicazione: (2025)
di: Liu, Wei, et al.
Pubblicazione: (2025)
Scaling Unverifiable Rewards: A Case Study on Visual Insights
di: Gan, Shuyu, et al.
Pubblicazione: (2025)
di: Gan, Shuyu, et al.
Pubblicazione: (2025)
Self-Improving World Modelling with Latent Actions
di: Qiu, Yifu, et al.
Pubblicazione: (2026)
di: Qiu, Yifu, et al.
Pubblicazione: (2026)
The Effect of Model Size on LLM Post-hoc Explainability via LIME
di: Heyen, Henning, et al.
Pubblicazione: (2024)
di: Heyen, Henning, et al.
Pubblicazione: (2024)
CIRAG: Construction-Integration Retrieval and Adaptive Generation for Multi-hop Question Answering
di: Wei, Zili, et al.
Pubblicazione: (2026)
di: Wei, Zili, et al.
Pubblicazione: (2026)
Understanding the Interplay of Scale, Data, and Bias in Language Models: A Case Study with BERT
di: Ali, Muhammad, et al.
Pubblicazione: (2024)
di: Ali, Muhammad, et al.
Pubblicazione: (2024)
Mixtures of In-Context Learners
di: Hong, Giwon, et al.
Pubblicazione: (2024)
di: Hong, Giwon, et al.
Pubblicazione: (2024)
Guarded Repair for Harm-Aware Post-hoc Replacement of LLM Mathematical Reasoning
di: Xia, Haizhou
Pubblicazione: (2026)
di: Xia, Haizhou
Pubblicazione: (2026)
AGR: Age Group fairness Reward for Bias Mitigation in LLMs
di: Cao, Shuirong, et al.
Pubblicazione: (2024)
di: Cao, Shuirong, et al.
Pubblicazione: (2024)
Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language Models
di: Bani-Harouni, David, et al.
Pubblicazione: (2025)
di: Bani-Harouni, David, et al.
Pubblicazione: (2025)
DR.GAP: Mitigating Bias in Large Language Models using Gender-Aware Prompting with Demonstration and Reasoning
di: Qiu, Hongye, et al.
Pubblicazione: (2025)
di: Qiu, Hongye, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
di: Huang, Zeyu, et al.
Pubblicazione: (2025) -
Understanding Post-hoc Explainers: The Case of Anchors
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2023) -
Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods
di: Dhaini, Mahdi, et al.
Pubblicazione: (2025) -
Emergent Communication Pretraining for Few-Shot Machine Translation
di: Li, Yaoyiran, et al.
Pubblicazione: (2020) -
Probing the Emergence of Cross-lingual Alignment during LLM Training
di: Wang, Hetong, et al.
Pubblicazione: (2024)