Fairness Dynamics During Training
Fuente:
arXiv
Saved in:
| Main Authors: | Patel, Krishna, Sivakumar, Nivedha, Theobald, Barry-John, Zappella, Luca, Apostoloff, Nicholas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bias after Prompting: Persistent Discrimination in Large Language Models
by: Sivakumar, Nivedha, et al.
Published: (2025)
by: Sivakumar, Nivedha, et al.
Published: (2025)
Evaluating Gender Bias Transfer between Pre-trained and Prompt-Adapted Language Models
by: Mackraz, Natalie, et al.
Published: (2024)
by: Mackraz, Natalie, et al.
Published: (2024)
DSO: Direct Steering Optimization for Bias Mitigation
by: Paes, Lucas Monteiro, et al.
Published: (2025)
by: Paes, Lucas Monteiro, et al.
Published: (2025)
Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs
by: Wang, Yinong Oliver, et al.
Published: (2025)
by: Wang, Yinong Oliver, et al.
Published: (2025)
Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution
by: Khan, Falaah Arif, et al.
Published: (2025)
by: Khan, Falaah Arif, et al.
Published: (2025)
Whispering Experts: Neural Interventions for Toxicity Mitigation in Language Models
by: Suau, Xavier, et al.
Published: (2024)
by: Suau, Xavier, et al.
Published: (2024)
Controlling Language and Diffusion Models by Transporting Activations
by: Rodriguez, Pau, et al.
Published: (2024)
by: Rodriguez, Pau, et al.
Published: (2024)
Aligning LLMs by Predicting Preferences from User Writing Samples
by: Aroca-Ouellette, Stéphane, et al.
Published: (2025)
by: Aroca-Ouellette, Stéphane, et al.
Published: (2025)
Theoretical Limits of Language Model Alignment
by: Paes, Lucas Monteiro, et al.
Published: (2026)
by: Paes, Lucas Monteiro, et al.
Published: (2026)
CoRet: Improved Retriever for Code Editing
by: Fehr, Fabio, et al.
Published: (2025)
by: Fehr, Fabio, et al.
Published: (2025)
ExpertLens: Activation steering features are highly interpretable
by: Fedzechkina, Masha, et al.
Published: (2025)
by: Fedzechkina, Masha, et al.
Published: (2025)
How Do Large Language Models Learn Concepts During Continual Pre-Training?
by: Yao, Barry Menglong, et al.
Published: (2026)
by: Yao, Barry Menglong, et al.
Published: (2026)
Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models
by: Sundar, Anirudh, et al.
Published: (2025)
by: Sundar, Anirudh, et al.
Published: (2025)
Uncertainty Quantification for LLM Function-Calling
by: Ye, Zihuiwen, et al.
Published: (2026)
by: Ye, Zihuiwen, et al.
Published: (2026)
Attention to Mamba: A Recipe for Cross-Architecture Distillation
by: Moudgil, Abhinav, et al.
Published: (2026)
by: Moudgil, Abhinav, et al.
Published: (2026)
Insights from the Inverse: Reconstructing LLM Training Goals Through Inverse Reinforcement Learning
by: Joselowitz, Jared, et al.
Published: (2024)
by: Joselowitz, Jared, et al.
Published: (2024)
Are Models Trained on Indian Legal Data Fair?
by: Girhepuje, Sahil, et al.
Published: (2023)
by: Girhepuje, Sahil, et al.
Published: (2023)
RAGVUE: A Diagnostic View for Explainable and Automated Evaluation of Retrieval-Augmented Generation
by: Murugaraj, Keerthana, et al.
Published: (2025)
by: Murugaraj, Keerthana, et al.
Published: (2025)
LinEAS: End-to-end Learning of Activation Steering with a Distributional Loss
by: Rodriguez, Pau, et al.
Published: (2025)
by: Rodriguez, Pau, et al.
Published: (2025)
On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization
by: Lin, Yong, et al.
Published: (2024)
by: Lin, Yong, et al.
Published: (2024)
Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results
by: Santilli, Andrea, et al.
Published: (2025)
by: Santilli, Andrea, et al.
Published: (2025)
How to Leverage Digit Embeddings to Represent Numbers?
by: Sivakumar, Jasivan Alex, et al.
Published: (2024)
by: Sivakumar, Jasivan Alex, et al.
Published: (2024)
Investigating Adversarial Trigger Transfer in Large Language Models
by: Meade, Nicholas, et al.
Published: (2024)
by: Meade, Nicholas, et al.
Published: (2024)
BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training
by: Chizhov, Pavel, et al.
Published: (2024)
by: Chizhov, Pavel, et al.
Published: (2024)
Mitigating Paraphrase Attacks on Machine-Text Detectors via Paraphrase Inversion
by: Soto, Rafael Rivera, et al.
Published: (2024)
by: Soto, Rafael Rivera, et al.
Published: (2024)
Cost-Effective Hallucination Detection for LLMs
by: Valentin, Simon, et al.
Published: (2024)
by: Valentin, Simon, et al.
Published: (2024)
Low-Resource Authorship Style Transfer: Can Non-Famous Authors Be Imitated?
by: Patel, Ajay, et al.
Published: (2022)
by: Patel, Ajay, et al.
Published: (2022)
MCEval: A Dynamic Framework for Fair Multilingual Cultural Evaluation of LLMs
by: Huang, Shulin, et al.
Published: (2025)
by: Huang, Shulin, et al.
Published: (2025)
FairSteer: Inference Time Debiasing for LLMs with Dynamic Activation Steering
by: Li, Yichen, et al.
Published: (2025)
by: Li, Yichen, et al.
Published: (2025)
Linear Recency Bias During Training Improves Transformers' Fit to Reading Times
by: Clark, Christian, et al.
Published: (2024)
by: Clark, Christian, et al.
Published: (2024)
Automating Historical Insight Extraction from Large-Scale Newspaper Archives via Neural Topic Modeling
by: Murugaraj, Keerthana, et al.
Published: (2025)
by: Murugaraj, Keerthana, et al.
Published: (2025)
Decoding News Narratives: A Critical Analysis of Large Language Models in Framing Detection
by: Pastorino, Valeria, et al.
Published: (2024)
by: Pastorino, Valeria, et al.
Published: (2024)
Can Large Language Models Simulate Human Responses? A Case Study of Stated Preference Experiments in the Context of Heating-related Choices
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
Characterizing Learning Curves During Language Model Pre-Training: Learning, Forgetting, and Stability
by: Chang, Tyler A., et al.
Published: (2023)
by: Chang, Tyler A., et al.
Published: (2023)
Empowering Interdisciplinary Research with BERT-Based Models: An Approach Through SciBERT-CNN with Topic Modeling
by: Likhareva, Darya, et al.
Published: (2024)
by: Likhareva, Darya, et al.
Published: (2024)
RAG-X: Systematic Diagnosis of Retrieval-Augmented Generation for Medical Question Answering
by: Sivakumar, Aswini, et al.
Published: (2026)
by: Sivakumar, Aswini, et al.
Published: (2026)
CLUES: Collaborative High-Quality Data Selection for LLMs via Training Dynamics
by: Zhao, Wanru, et al.
Published: (2025)
by: Zhao, Wanru, et al.
Published: (2025)
Value Drifts: Tracing Value Alignment During LLM Post-Training
by: Bhatia, Mehar, et al.
Published: (2025)
by: Bhatia, Mehar, et al.
Published: (2025)
Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model
by: Cheng, Sheng, et al.
Published: (2024)
by: Cheng, Sheng, et al.
Published: (2024)
SMITE: Enhancing Fairness in LLMs through Optimal In-Context Example Selection via Dynamic Validation
by: Chhikara, Garima, et al.
Published: (2025)
by: Chhikara, Garima, et al.
Published: (2025)
Similar Items
-
Bias after Prompting: Persistent Discrimination in Large Language Models
by: Sivakumar, Nivedha, et al.
Published: (2025) -
Evaluating Gender Bias Transfer between Pre-trained and Prompt-Adapted Language Models
by: Mackraz, Natalie, et al.
Published: (2024) -
DSO: Direct Steering Optimization for Bias Mitigation
by: Paes, Lucas Monteiro, et al.
Published: (2025) -
Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs
by: Wang, Yinong Oliver, et al.
Published: (2025) -
Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution
by: Khan, Falaah Arif, et al.
Published: (2025)