Chained Tuning Leads to Biased Forgetting
Fuente:
arXiv
Saved in:
| Main Authors: | Ung, Megan, Sun, Alicia, Bell, Samuel J., Radharapu, Bhaktipriya, Sagun, Levent, Williams, Adina |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Arbiters of Ambivalence: Challenges of Using LLMs in No-Consensus Tasks
by: Radharapu, Bhaktipriya, et al.
Published: (2025)
by: Radharapu, Bhaktipriya, et al.
Published: (2025)
Debiasing Text Safety Classifiers through a Fairness-Aware Ensemble
by: Sturman, Olivia, et al.
Published: (2024)
by: Sturman, Olivia, et al.
Published: (2024)
LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance
by: Haller, Patrick, et al.
Published: (2025)
by: Haller, Patrick, et al.
Published: (2025)
Changing Answer Order Can Decrease MMLU Accuracy
by: Gupta, Vipul, et al.
Published: (2024)
by: Gupta, Vipul, et al.
Published: (2024)
Beg to Differ: Understanding Reasoning-Answer Misalignment Across Languages
by: Ovalle, Anaelia, et al.
Published: (2025)
by: Ovalle, Anaelia, et al.
Published: (2025)
On the Role of Speech Data in Reducing Toxicity Detection Bias
by: Bell, Samuel J., et al.
Published: (2024)
by: Bell, Samuel J., et al.
Published: (2024)
RealSeal: Revolutionizing Media Authentication with Real-Time Realism Scoring
by: Radharapu, Bhaktipriya, et al.
Published: (2024)
by: Radharapu, Bhaktipriya, et al.
Published: (2024)
The Root Shapes the Fruit: On the Persistence of Gender-Exclusive Harms in Aligned Language Models
by: Ovalle, Anaelia, et al.
Published: (2024)
by: Ovalle, Anaelia, et al.
Published: (2024)
Improving Model Evaluation using SMART Filtering of Benchmark Datasets
by: Gupta, Vipul, et al.
Published: (2024)
by: Gupta, Vipul, et al.
Published: (2024)
Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
by: Radharapu, Bhaktipriya, et al.
Published: (2025)
by: Radharapu, Bhaktipriya, et al.
Published: (2025)
An Effective Theory of Bias Amplification
by: Subramonian, Arjun, et al.
Published: (2024)
by: Subramonian, Arjun, et al.
Published: (2024)
Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language Models
by: Shaib, Chantal, et al.
Published: (2025)
by: Shaib, Chantal, et al.
Published: (2025)
Reassessing the Validity of Spurious Correlations Benchmarks
by: Bell, Samuel J., et al.
Published: (2024)
by: Bell, Samuel J., et al.
Published: (2024)
ShieldGemma: Generative AI Content Moderation Based on Gemma
by: Zeng, Wenjun, et al.
Published: (2024)
by: Zeng, Wenjun, et al.
Published: (2024)
Weisfeiler and Leman Go Measurement Modeling: Probing the Validity of the WL Test
by: Subramonian, Arjun, et al.
Published: (2023)
by: Subramonian, Arjun, et al.
Published: (2023)
Stuffed Mamba: Oversized States Lead to the Inability to Forget
by: Chen, Yingfa, et al.
Published: (2024)
by: Chen, Yingfa, et al.
Published: (2024)
Are Female Carpenters like Blue Bananas? A Corpus Investigation of Occupation Gender Typicality
by: Ju, Da, et al.
Published: (2024)
by: Ju, Da, et al.
Published: (2024)
Domain Regeneration: How well do LLMs match syntactic properties of text domains?
by: Ju, Da, et al.
Published: (2025)
by: Ju, Da, et al.
Published: (2025)
Mitigating Forgetting in LLM Fine-Tuning via Low-Perplexity Token Learning
by: Wu, Chao-Chung, et al.
Published: (2025)
by: Wu, Chao-Chung, et al.
Published: (2025)
Networked Inequality: Preferential Attachment Bias in Graph Neural Network Link Prediction
by: Subramonian, Arjun, et al.
Published: (2023)
by: Subramonian, Arjun, et al.
Published: (2023)
Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought
by: Chua, James, et al.
Published: (2024)
by: Chua, James, et al.
Published: (2024)
Revisiting Catastrophic Forgetting in Large Language Model Tuning
by: Li, Hongyu, et al.
Published: (2024)
by: Li, Hongyu, et al.
Published: (2024)
Joint Flashback Adaptation for Forgetting-Resistant Instruction Tuning
by: Zhao, Yukun, et al.
Published: (2025)
by: Zhao, Yukun, et al.
Published: (2025)
Preventing Catastrophic Forgetting: Behavior-Aware Sampling for Safer Language Model Fine-Tuning
by: Pham, Anh, et al.
Published: (2025)
by: Pham, Anh, et al.
Published: (2025)
What makes a good metric? Evaluating automatic metrics for text-to-image consistency
by: Ross, Candace, et al.
Published: (2024)
by: Ross, Candace, et al.
Published: (2024)
Analyzing and Reducing Catastrophic Forgetting in Parameter Efficient Tuning
by: Ren, Weijieying, et al.
Published: (2024)
by: Ren, Weijieying, et al.
Published: (2024)
Exploring Forgetting in Large Language Model Pre-Training
by: Liao, Chonghua, et al.
Published: (2024)
by: Liao, Chonghua, et al.
Published: (2024)
The Causal Influence of Grammatical Gender on Distributional Semantics
by: Stańczak, Karolina, et al.
Published: (2023)
by: Stańczak, Karolina, et al.
Published: (2023)
Fine-Tune, Don't Prompt, Your Language Model to Identify Biased Language in Clinical Notes
by: Landi, Isotta, et al.
Published: (2026)
by: Landi, Isotta, et al.
Published: (2026)
Measuring Catastrophic Forgetting in Cross-Lingual Transfer Paradigms: Exploring Tuning Strategies
by: Koloski, Boshko, et al.
Published: (2023)
by: Koloski, Boshko, et al.
Published: (2023)
Improved Supervised Fine-Tuning for Large Language Models to Mitigate Catastrophic Forgetting
by: Ding, Fei, et al.
Published: (2025)
by: Ding, Fei, et al.
Published: (2025)
A Latent-Variable Model for Intrinsic Probing
by: Stańczak, Karolina, et al.
Published: (2022)
by: Stańczak, Karolina, et al.
Published: (2022)
On the Impact of Fine-Tuning on Chain-of-Thought Reasoning
by: Lobo, Elita, et al.
Published: (2024)
by: Lobo, Elita, et al.
Published: (2024)
"Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas
by: Ding, Junchen, et al.
Published: (2025)
by: Ding, Junchen, et al.
Published: (2025)
OPLoRA: Orthogonal Projection LoRA Prevents Catastrophic Forgetting during Parameter-Efficient Fine-Tuning
by: Xiong, Yifeng, et al.
Published: (2025)
by: Xiong, Yifeng, et al.
Published: (2025)
Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training
by: Lai, Song, et al.
Published: (2025)
by: Lai, Song, et al.
Published: (2025)
Entropy-Adaptive Fine-Tuning: Resolving Confident Conflicts to Mitigate Forgetting
by: Diao, Muxi, et al.
Published: (2026)
by: Diao, Muxi, et al.
Published: (2026)
White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMs
by: Wan, Yixin, et al.
Published: (2024)
by: Wan, Yixin, et al.
Published: (2024)
Do different prompting methods yield a common task representation in language models?
by: Davidson, Guy, et al.
Published: (2025)
by: Davidson, Guy, et al.
Published: (2025)
Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention Models
by: Lee, Chungpa, et al.
Published: (2026)
by: Lee, Chungpa, et al.
Published: (2026)
Similar Items
-
Arbiters of Ambivalence: Challenges of Using LLMs in No-Consensus Tasks
by: Radharapu, Bhaktipriya, et al.
Published: (2025) -
Debiasing Text Safety Classifiers through a Fairness-Aware Ensemble
by: Sturman, Olivia, et al.
Published: (2024) -
LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance
by: Haller, Patrick, et al.
Published: (2025) -
Changing Answer Order Can Decrease MMLU Accuracy
by: Gupta, Vipul, et al.
Published: (2024) -
Beg to Differ: Understanding Reasoning-Answer Misalignment Across Languages
by: Ovalle, Anaelia, et al.
Published: (2025)