Mitigating Self-Preference by Authorship Obfuscation
Fuente:
arXiv
Guardado en:
| Autores principales: | Mahbub, Taslim, Feng, Shi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
JAMDEC: Unsupervised Authorship Obfuscation using Constrained Decoding over Small Language Models
por: Fisher, Jillian, et al.
Publicado: (2024)
por: Fisher, Jillian, et al.
Publicado: (2024)
ALISON: Fast and Effective Stylometric Authorship Obfuscation
por: Xing, Eric, et al.
Publicado: (2024)
por: Xing, Eric, et al.
Publicado: (2024)
Consistency Training while Mitigating Obfuscation via Rate Matching
por: Imran, Sohaib, et al.
Publicado: (2026)
por: Imran, Sohaib, et al.
Publicado: (2026)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
por: Yang, Jinming, et al.
Publicado: (2026)
por: Yang, Jinming, et al.
Publicado: (2026)
CAVE: Controllable Authorship Verification Explanations
por: Ramnath, Sahana, et al.
Publicado: (2024)
por: Ramnath, Sahana, et al.
Publicado: (2024)
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
por: Roytburg, Dani, et al.
Publicado: (2025)
por: Roytburg, Dani, et al.
Publicado: (2025)
Self-supervised Preference Optimization: Enhance Your Language Model with Preference Degree Awareness
por: Li, Jian, et al.
Publicado: (2024)
por: Li, Jian, et al.
Publicado: (2024)
Personalized Author Obfuscation with Large Language Models
por: Shokri, Mohammad, et al.
Publicado: (2025)
por: Shokri, Mohammad, et al.
Publicado: (2025)
PARSI: Persian Authorship Recognition via Stylometric Integration
por: Shahnazari, Kourosh, et al.
Publicado: (2025)
por: Shahnazari, Kourosh, et al.
Publicado: (2025)
LLM one-shot style transfer for Authorship Attribution and Verification
por: Miralles-González, Pablo, et al.
Publicado: (2025)
por: Miralles-González, Pablo, et al.
Publicado: (2025)
Advancing NLP Security by Leveraging LLMs as Adversarial Engines
por: Srinivasan, Sudarshan, et al.
Publicado: (2024)
por: Srinivasan, Sudarshan, et al.
Publicado: (2024)
Inclusion-of-Thoughts: Mitigating Preference Instability via Purifying the Decision Space
por: Madani, Mohammad Reza Ghasemi, et al.
Publicado: (2026)
por: Madani, Mohammad Reza Ghasemi, et al.
Publicado: (2026)
Can LLMs Obfuscate Code? A Systematic Analysis of Large Language Models into Assembly Code Obfuscation
por: Mohseni, Seyedreza, et al.
Publicado: (2024)
por: Mohseni, Seyedreza, et al.
Publicado: (2024)
Every Step Counts: Decoding Trajectories as Authorship Fingerprints of dLLMs
por: Li, Qi, et al.
Publicado: (2025)
por: Li, Qi, et al.
Publicado: (2025)
LINGOLY-TOO: Disentangling Reasoning from Knowledge with Templatised Orthographic Obfuscation
por: Khouja, Jude, et al.
Publicado: (2025)
por: Khouja, Jude, et al.
Publicado: (2025)
CollabStory: Multi-LLM Collaborative Story Generation and Authorship Analysis
por: Venkatraman, Saranya, et al.
Publicado: (2024)
por: Venkatraman, Saranya, et al.
Publicado: (2024)
SGPO: Self-Generated Preference Optimization based on Self-Improver
por: Lee, Hyeonji, et al.
Publicado: (2025)
por: Lee, Hyeonji, et al.
Publicado: (2025)
A Bayesian Approach to Harnessing the Power of LLMs in Authorship Attribution
por: Hu, Zhengmian, et al.
Publicado: (2024)
por: Hu, Zhengmian, et al.
Publicado: (2024)
AuthorMix: Modular Authorship Style Transfer via Layer-wise Adapter Mixing
por: Thillainathan, Sarubi, et al.
Publicado: (2026)
por: Thillainathan, Sarubi, et al.
Publicado: (2026)
Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation
por: Zhang, Xiaoying, et al.
Publicado: (2024)
por: Zhang, Xiaoying, et al.
Publicado: (2024)
Preference Consistency Matters: Enhancing Preference Learning in Language Models with Automated Self-Curation of Training Corpora
por: Lee, JoonHo, et al.
Publicado: (2024)
por: Lee, JoonHo, et al.
Publicado: (2024)
Authorship Without Writing: Large Language Models and the Senior Author Analogy
por: Hurshman, Clint, et al.
Publicado: (2025)
por: Hurshman, Clint, et al.
Publicado: (2025)
BARD10: A New Benchmark Reveals Significance of Bangla Stop-Words in Authorship Attribution
por: Moosa, Abdullah Muhammad, et al.
Publicado: (2025)
por: Moosa, Abdullah Muhammad, et al.
Publicado: (2025)
Self-supervised Attribute-aware Dynamic Preference Ranking Alignment
por: Yang, Hongyu, et al.
Publicado: (2025)
por: Yang, Hongyu, et al.
Publicado: (2025)
Self-Boosting Large Language Models with Synthetic Preference Data
por: Dong, Qingxiu, et al.
Publicado: (2024)
por: Dong, Qingxiu, et al.
Publicado: (2024)
MATCHED: Multimodal Authorship-Attribution To Combat Human Trafficking in Escort-Advertisement Data
por: Saxena, Vageesh, et al.
Publicado: (2024)
por: Saxena, Vageesh, et al.
Publicado: (2024)
Counterfactual Probing for Hallucination Detection and Mitigation in Large Language Models
por: Feng, Yijun
Publicado: (2025)
por: Feng, Yijun
Publicado: (2025)
Self-Consistency Preference Optimization
por: Prasad, Archiki, et al.
Publicado: (2024)
por: Prasad, Archiki, et al.
Publicado: (2024)
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models
por: Pombal, José, et al.
Publicado: (2026)
por: Pombal, José, et al.
Publicado: (2026)
CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations
por: Zhang, Mike, et al.
Publicado: (2026)
por: Zhang, Mike, et al.
Publicado: (2026)
SelfAug: Mitigating Catastrophic Forgetting in Retrieval-Augmented Generation via Distribution Self-Alignment
por: Huang, Yuqing, et al.
Publicado: (2025)
por: Huang, Yuqing, et al.
Publicado: (2025)
Spontaneous Reward Hacking in Iterative Self-Refinement
por: Pan, Jane, et al.
Publicado: (2024)
por: Pan, Jane, et al.
Publicado: (2024)
Safer-Instruct: Aligning Language Models with Automated Preference Data
por: Shi, Taiwei, et al.
Publicado: (2023)
por: Shi, Taiwei, et al.
Publicado: (2023)
AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
por: Gupta, Taneesh, et al.
Publicado: (2025)
por: Gupta, Taneesh, et al.
Publicado: (2025)
Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference Optimization
por: Wu, Jiulong, et al.
Publicado: (2025)
por: Wu, Jiulong, et al.
Publicado: (2025)
Subtle Errors in Reasoning: Preference Learning via Error-injected Self-editing
por: Xu, Kaishuai, et al.
Publicado: (2024)
por: Xu, Kaishuai, et al.
Publicado: (2024)
Multi-Faceted Self-Consistent Preference Alignment for Query Rewriting in Conversational Search
por: Cao, Zhiyu, et al.
Publicado: (2026)
por: Cao, Zhiyu, et al.
Publicado: (2026)
Towards Understanding the Influence of Reward Margin on Preference Model Performance
por: Qin, Bowen, et al.
Publicado: (2024)
por: Qin, Bowen, et al.
Publicado: (2024)
PLaD: Preference-based Large Language Model Distillation with Pseudo-Preference Pairs
por: Zhang, Rongzhi, et al.
Publicado: (2024)
por: Zhang, Rongzhi, et al.
Publicado: (2024)
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering
por: Mohamed, Anas, et al.
Publicado: (2025)
por: Mohamed, Anas, et al.
Publicado: (2025)
Ejemplares similares
-
JAMDEC: Unsupervised Authorship Obfuscation using Constrained Decoding over Small Language Models
por: Fisher, Jillian, et al.
Publicado: (2024) -
ALISON: Fast and Effective Stylometric Authorship Obfuscation
por: Xing, Eric, et al.
Publicado: (2024) -
Consistency Training while Mitigating Obfuscation via Rate Matching
por: Imran, Sohaib, et al.
Publicado: (2026) -
Quantifying and Mitigating Self-Preference Bias of LLM Judges
por: Yang, Jinming, et al.
Publicado: (2026) -
CAVE: Controllable Authorship Verification Explanations
por: Ramnath, Sahana, et al.
Publicado: (2024)