fairBERTs: Erasing Sensitive Information Through Semantic and Fairness-aware Perturbations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Jinfeng, Chen, Yuefeng, Liu, Xiangyu, Huang, Longtao, Zhang, Rong, Xue, Hui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BERTs are Generative In-Context Learners
von: Samuel, David
Veröffentlicht: (2024)
von: Samuel, David
Veröffentlicht: (2024)
Exploring Linguistic Properties of Monolingual BERTs with Typological Classification among Languages
von: Ruzzetti, Elena Sofia, et al.
Veröffentlicht: (2023)
von: Ruzzetti, Elena Sofia, et al.
Veröffentlicht: (2023)
QExplorer: Large Language Model Based Query Extraction for Toxic Content Exploration
von: Ren, Shaola, et al.
Veröffentlicht: (2025)
von: Ren, Shaola, et al.
Veröffentlicht: (2025)
Towards the Resistance of Neural Network Watermarking to Fine-tuning
von: Tang, Ling, et al.
Veröffentlicht: (2025)
von: Tang, Ling, et al.
Veröffentlicht: (2025)
UniErase: Towards Balanced and Precise Unlearning in Language Models
von: Yu, Miao, et al.
Veröffentlicht: (2025)
von: Yu, Miao, et al.
Veröffentlicht: (2025)
General Phrase Debiaser: Debiasing Masked Language Models at a Multi-Token Level
von: Shi, Bingkang, et al.
Veröffentlicht: (2023)
von: Shi, Bingkang, et al.
Veröffentlicht: (2023)
Confidence-aware Self-Semantic Distillation on Knowledge Graph Embedding
von: Liu, Yichen, et al.
Veröffentlicht: (2022)
von: Liu, Yichen, et al.
Veröffentlicht: (2022)
DOPRA: Decoding Over-accumulation Penalization and Re-allocation in Specific Weighting Layer
von: Wei, Jinfeng, et al.
Veröffentlicht: (2024)
von: Wei, Jinfeng, et al.
Veröffentlicht: (2024)
Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs
von: Wang, Ziliang, et al.
Veröffentlicht: (2025)
von: Wang, Ziliang, et al.
Veröffentlicht: (2025)
Improving Fairness in LLMs Through Testing-Time Adversaries
von: Gregio, Isabela Pereira, et al.
Veröffentlicht: (2025)
von: Gregio, Isabela Pereira, et al.
Veröffentlicht: (2025)
How LoRA Remembers? A Parametric Memory Law for LLM Finetuning
von: Xu, Ziwen, et al.
Veröffentlicht: (2026)
von: Xu, Ziwen, et al.
Veröffentlicht: (2026)
RedacBench: Can AI Erase Your Secrets?
von: Jeon, Hyunjun, et al.
Veröffentlicht: (2026)
von: Jeon, Hyunjun, et al.
Veröffentlicht: (2026)
Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation
von: Cheng, Yinjie, et al.
Veröffentlicht: (2025)
von: Cheng, Yinjie, et al.
Veröffentlicht: (2025)
SePer: Measure Retrieval Utility Through The Lens Of Semantic Perplexity Reduction
von: Dai, Lu, et al.
Veröffentlicht: (2025)
von: Dai, Lu, et al.
Veröffentlicht: (2025)
Enhancing the QA Model through a Multi-domain Debiasing Framework
von: Wang, Yuefeng, et al.
Veröffentlicht: (2026)
von: Wang, Yuefeng, et al.
Veröffentlicht: (2026)
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
von: Huang, Kaixuan, et al.
Veröffentlicht: (2025)
von: Huang, Kaixuan, et al.
Veröffentlicht: (2025)
LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
von: Yang, Junxiao, et al.
Veröffentlicht: (2026)
von: Yang, Junxiao, et al.
Veröffentlicht: (2026)
Fairness of Automatic Speech Recognition: Looking Through a Philosophical Lens
von: Choi, Anna Seo Gyeong, et al.
Veröffentlicht: (2025)
von: Choi, Anna Seo Gyeong, et al.
Veröffentlicht: (2025)
VISLA Benchmark: Evaluating Embedding Sensitivity to Semantic and Lexical Alterations
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
UP5: Unbiased Foundation Model for Fairness-aware Recommendation
von: Hua, Wenyue, et al.
Veröffentlicht: (2023)
von: Hua, Wenyue, et al.
Veröffentlicht: (2023)
LLM-CAS: Dynamic Neuron Perturbation for Real-Time Hallucination Correction
von: Zhang, Jensen, et al.
Veröffentlicht: (2025)
von: Zhang, Jensen, et al.
Veröffentlicht: (2025)
Understanding Fairness-Accuracy Trade-offs in Machine Learning Models: Does Promoting Fairness Undermine Performance?
von: Liu, Junhua, et al.
Veröffentlicht: (2024)
von: Liu, Junhua, et al.
Veröffentlicht: (2024)
True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning
von: Tan, Weihao, et al.
Veröffentlicht: (2024)
von: Tan, Weihao, et al.
Veröffentlicht: (2024)
Are AI-Generated Text Detectors Robust to Adversarial Perturbations?
von: Huang, Guanhua, et al.
Veröffentlicht: (2024)
von: Huang, Guanhua, et al.
Veröffentlicht: (2024)
DepFlow: Disentangled Speech Generation to Mitigate Semantic Bias in Depression Detection
von: Li, Yuxin, et al.
Veröffentlicht: (2026)
von: Li, Yuxin, et al.
Veröffentlicht: (2026)
Outraged AI: Large language models prioritise emotion over cost in fairness enforcement
von: Liu, Hao, et al.
Veröffentlicht: (2025)
von: Liu, Hao, et al.
Veröffentlicht: (2025)
DiscoSum: Discourse-aware News Summarization
von: Spangher, Alexander, et al.
Veröffentlicht: (2025)
von: Spangher, Alexander, et al.
Veröffentlicht: (2025)
Advancing and Benchmarking Personalized Tool Invocation for LLMs
von: Huang, Xu, et al.
Veröffentlicht: (2025)
von: Huang, Xu, et al.
Veröffentlicht: (2025)
Taming Sensitive Weights : Noise Perturbation Fine-tuning for Robust LLM Quantization
von: Wang, Dongwei, et al.
Veröffentlicht: (2024)
von: Wang, Dongwei, et al.
Veröffentlicht: (2024)
Causal-aware Large Language Models: Enhancing Decision-Making Through Learning, Adapting and Acting
von: Chen, Wei, et al.
Veröffentlicht: (2025)
von: Chen, Wei, et al.
Veröffentlicht: (2025)
On Fairness of Unified Multimodal Large Language Model for Image Generation
von: Liu, Ming, et al.
Veröffentlicht: (2025)
von: Liu, Ming, et al.
Veröffentlicht: (2025)
$μ$KE: Matryoshka Unstructured Knowledge Editing of Large Language Models
von: Su, Zian, et al.
Veröffentlicht: (2025)
von: Su, Zian, et al.
Veröffentlicht: (2025)
Enhancing Essay Scoring with Adversarial Weights Perturbation and Metric-specific AttentionPooling
von: Huang, Jiaxin, et al.
Veröffentlicht: (2024)
von: Huang, Jiaxin, et al.
Veröffentlicht: (2024)
ProtSAE: Disentangling and Interpreting Protein Language Models via Semantically-Guided Sparse Autoencoders
von: Liu, Xiangyu, et al.
Veröffentlicht: (2025)
von: Liu, Xiangyu, et al.
Veröffentlicht: (2025)
ICDM 2020 Knowledge Graph Contest: Consumer Event-Cause Extraction
von: He, Congqing, et al.
Veröffentlicht: (2021)
von: He, Congqing, et al.
Veröffentlicht: (2021)
The Right Time Matters: Data Arrangement Affects Zero-Shot Generalization in Instruction Tuning
von: He, Bingxiang, et al.
Veröffentlicht: (2024)
von: He, Bingxiang, et al.
Veröffentlicht: (2024)
The Other Side of the Coin: Exploring Fairness in Retrieval-Augmented Generation
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
von: Xu, Haolei, et al.
Veröffentlicht: (2026)
von: Xu, Haolei, et al.
Veröffentlicht: (2026)
OTESGN: Optimal Transport-Enhanced Syntactic-Semantic Graph Networks for Aspect-Based Sentiment Analysis
von: Liao, Xinfeng, et al.
Veröffentlicht: (2025)
von: Liao, Xinfeng, et al.
Veröffentlicht: (2025)
RoSA: Enhancing Parameter-Efficient Fine-Tuning via RoPE-aware Selective Adaptation in Large Language Models
von: Pan, Dayan, et al.
Veröffentlicht: (2025)
von: Pan, Dayan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
BERTs are Generative In-Context Learners
von: Samuel, David
Veröffentlicht: (2024) -
Exploring Linguistic Properties of Monolingual BERTs with Typological Classification among Languages
von: Ruzzetti, Elena Sofia, et al.
Veröffentlicht: (2023) -
QExplorer: Large Language Model Based Query Extraction for Toxic Content Exploration
von: Ren, Shaola, et al.
Veröffentlicht: (2025) -
Towards the Resistance of Neural Network Watermarking to Fine-tuning
von: Tang, Ling, et al.
Veröffentlicht: (2025) -
UniErase: Towards Balanced and Precise Unlearning in Language Models
von: Yu, Miao, et al.
Veröffentlicht: (2025)