Intra-Fairness Dynamics: The Bias Spillover Effect in Targeted LLM Alignment
Fuente:
arXiv
Salvato in:
| Autori principali: | Paraschou, Eva, Clemmensen, Line Harder, Das, Sneha |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Cost-Benefit of Interdisciplinarity in AI for Mental Health
di: Drakos, Katerina, et al.
Pubblicazione: (2025)
di: Drakos, Katerina, et al.
Pubblicazione: (2025)
Evaluation of Stress Detection as Time Series Events -- A Novel Window-Based F1-Metric
di: Skat-Rørdam, Harald Vilhelm, et al.
Pubblicazione: (2025)
di: Skat-Rørdam, Harald Vilhelm, et al.
Pubblicazione: (2025)
Mind the XAI Gap: A Human-Centered LLM Framework for Democratizing Explainable AI
di: Paraschou, Eva, et al.
Pubblicazione: (2025)
di: Paraschou, Eva, et al.
Pubblicazione: (2025)
Beyond Word Error Rate: Auditing the Diversity Tax in Speech Recognition through Dataset Cartography
di: Cheng, Ting-Hui, et al.
Pubblicazione: (2026)
di: Cheng, Ting-Hui, et al.
Pubblicazione: (2026)
Alignment Dynamics in LLM Fine-Tuning
di: Huang, Yuhan, et al.
Pubblicazione: (2026)
di: Huang, Yuhan, et al.
Pubblicazione: (2026)
Fair Clustering via Alignment
di: Kim, Kunwoong, et al.
Pubblicazione: (2025)
di: Kim, Kunwoong, et al.
Pubblicazione: (2025)
FairTargetSim: An Interactive Simulator for Understanding and Explaining the Fairness Effects of Target Variable Definition
di: Gala, Dalia, et al.
Pubblicazione: (2024)
di: Gala, Dalia, et al.
Pubblicazione: (2024)
ECLIPTICA -- A Framework for Switchable LLM Alignment via CITA - Contrastive Instruction-Tuned Alignment
di: Wanaskar, Kapil, et al.
Pubblicazione: (2026)
di: Wanaskar, Kapil, et al.
Pubblicazione: (2026)
A Self-Organizing Clustering System for Unsupervised Distribution Shift Detection
di: Basterrech, Sebastián, et al.
Pubblicazione: (2024)
di: Basterrech, Sebastián, et al.
Pubblicazione: (2024)
RLTHF: Targeted Human Feedback for LLM Alignment
di: Xu, Yifei, et al.
Pubblicazione: (2025)
di: Xu, Yifei, et al.
Pubblicazione: (2025)
Emergent Bias and Fairness in Multi-Agent Decision Systems
di: Madigan, Maeve, et al.
Pubblicazione: (2025)
di: Madigan, Maeve, et al.
Pubblicazione: (2025)
Fairness-Driven LLM-based Causal Discovery with Active Learning and Dynamic Scoring
di: Zanna, Khadija, et al.
Pubblicazione: (2025)
di: Zanna, Khadija, et al.
Pubblicazione: (2025)
Unveiling and Mitigating Bias in Large Language Model Recommendations: A Path to Fairness
di: Das, Anindya Bijoy, et al.
Pubblicazione: (2024)
di: Das, Anindya Bijoy, et al.
Pubblicazione: (2024)
Fair Overlap Number of Balls (Fair-ONB): A Data-Morphology-based Undersampling Method for Bias Reduction
di: Pascual-Triana, José Daniel, et al.
Pubblicazione: (2024)
di: Pascual-Triana, José Daniel, et al.
Pubblicazione: (2024)
Entropy-Guided Dynamic Tokens for Graph-LLM Alignment in Molecular Understanding
di: Jing, Zihao, et al.
Pubblicazione: (2026)
di: Jing, Zihao, et al.
Pubblicazione: (2026)
seeBias: A Comprehensive Tool for Assessing and Visualizing AI Fairness
di: Ning, Yilin, et al.
Pubblicazione: (2025)
di: Ning, Yilin, et al.
Pubblicazione: (2025)
Fair Dataset Distillation via Cross-Group Barycenter Alignment
di: Moslemi, Mohammad Hossein, et al.
Pubblicazione: (2026)
di: Moslemi, Mohammad Hossein, et al.
Pubblicazione: (2026)
FairAgent: Democratizing Fairness-Aware Machine Learning with LLM-Powered Agents
di: Dai, Yucong, et al.
Pubblicazione: (2025)
di: Dai, Yucong, et al.
Pubblicazione: (2025)
From Bias to Balance: Fairness-Aware Paper Recommendation for Equitable Peer Review
di: Oyshi, Uttamasha Anjally, et al.
Pubblicazione: (2026)
di: Oyshi, Uttamasha Anjally, et al.
Pubblicazione: (2026)
ProbLog4Fairness: A Neurosymbolic Approach to Modeling and Mitigating Bias
di: Adriaensen, Rik, et al.
Pubblicazione: (2025)
di: Adriaensen, Rik, et al.
Pubblicazione: (2025)
BMFT: Achieving Fairness via Bias-based Weight Masking Fine-tuning
di: Xue, Yuyang, et al.
Pubblicazione: (2024)
di: Xue, Yuyang, et al.
Pubblicazione: (2024)
Ensuring Equitable Financial Decisions: Leveraging Counterfactual Fairness and Deep Learning for Bias
di: Shinde, Saish
Pubblicazione: (2024)
di: Shinde, Saish
Pubblicazione: (2024)
APPA: Adaptive Preference Pluralistic Alignment for Fair Federated RLHF of LLMs
di: Srewa, Mahmoud, et al.
Pubblicazione: (2026)
di: Srewa, Mahmoud, et al.
Pubblicazione: (2026)
FairCauseSyn: Towards Causally Fair LLM-Augmented Synthetic Data Generation
di: Nagesh, Nitish, et al.
Pubblicazione: (2025)
di: Nagesh, Nitish, et al.
Pubblicazione: (2025)
Automatic Causal Fairness Analysis with LLM-Generated Reporting
di: Berarducci, Alessia, et al.
Pubblicazione: (2026)
di: Berarducci, Alessia, et al.
Pubblicazione: (2026)
Bt-GAN: Generating Fair Synthetic Healthdata via Bias-transforming Generative Adversarial Networks
di: Ramachandranpillai, Resmi, et al.
Pubblicazione: (2024)
di: Ramachandranpillai, Resmi, et al.
Pubblicazione: (2024)
IntraSlice: Towards High-Performance Structural Pruning with Block-Intra PCA for LLMs
di: Li, Meng, et al.
Pubblicazione: (2026)
di: Li, Meng, et al.
Pubblicazione: (2026)
Tokenized Bandit for LLM Decoding and Alignment
di: Shin, Suho, et al.
Pubblicazione: (2025)
di: Shin, Suho, et al.
Pubblicazione: (2025)
Benchmarking Fairness in Spiking Neural Networks: Data Bias, Spurious Features, and Hardware Effects
di: He, Hudi, et al.
Pubblicazione: (2026)
di: He, Hudi, et al.
Pubblicazione: (2026)
Steering LLM Reasoning Through Bias-Only Adaptation
di: Sinii, Viacheslav, et al.
Pubblicazione: (2025)
di: Sinii, Viacheslav, et al.
Pubblicazione: (2025)
Conformal Feedback Alignment: Quantifying Answer-Level Reliability for Robust LLM Alignment
di: Chen, Tiejin, et al.
Pubblicazione: (2026)
di: Chen, Tiejin, et al.
Pubblicazione: (2026)
Explanations as Bias Detectors: A Critical Study of Local Post-hoc XAI Methods for Fairness Exploration
di: Papanikou, Vasiliki, et al.
Pubblicazione: (2025)
di: Papanikou, Vasiliki, et al.
Pubblicazione: (2025)
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
di: Cheng, Ruoxi, et al.
Pubblicazione: (2025)
di: Cheng, Ruoxi, et al.
Pubblicazione: (2025)
BiasGuard: Guardrailing Fairness in Machine Learning Production Systems
di: Cohen-Inger, Nurit, et al.
Pubblicazione: (2025)
di: Cohen-Inger, Nurit, et al.
Pubblicazione: (2025)
Adversarial Preference Learning for Robust LLM Alignment
di: Wang, Yuanfu, et al.
Pubblicazione: (2025)
di: Wang, Yuanfu, et al.
Pubblicazione: (2025)
Evaluation of Large Language Models: STEM education and Gender Stereotypes
di: Due, Smilla, et al.
Pubblicazione: (2024)
di: Due, Smilla, et al.
Pubblicazione: (2024)
Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs
di: Taraghi, Mina, et al.
Pubblicazione: (2025)
di: Taraghi, Mina, et al.
Pubblicazione: (2025)
Classification of Tennis Actions Using Deep Learning
di: Hovad, Emil, et al.
Pubblicazione: (2024)
di: Hovad, Emil, et al.
Pubblicazione: (2024)
Adaptive Repetition for Mitigating Position Bias in LLM-Based Ranking
di: Vardasbi, Ali, et al.
Pubblicazione: (2025)
di: Vardasbi, Ali, et al.
Pubblicazione: (2025)
FairNet: Dynamic Fairness Correction without Performance Loss via Contrastive Conditional LoRA
di: Zhou, Songqi, et al.
Pubblicazione: (2025)
di: Zhou, Songqi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
The Cost-Benefit of Interdisciplinarity in AI for Mental Health
di: Drakos, Katerina, et al.
Pubblicazione: (2025) -
Evaluation of Stress Detection as Time Series Events -- A Novel Window-Based F1-Metric
di: Skat-Rørdam, Harald Vilhelm, et al.
Pubblicazione: (2025) -
Mind the XAI Gap: A Human-Centered LLM Framework for Democratizing Explainable AI
di: Paraschou, Eva, et al.
Pubblicazione: (2025) -
Beyond Word Error Rate: Auditing the Diversity Tax in Speech Recognition through Dataset Cartography
di: Cheng, Ting-Hui, et al.
Pubblicazione: (2026) -
Alignment Dynamics in LLM Fine-Tuning
di: Huang, Yuhan, et al.
Pubblicazione: (2026)