Guardado en:
| Autores principales: | Springer, Max, Lee, Chung Peng, Metevier, Blossom, Castleman, Jane, Turbal, Bohdan, Jung, Hayoung, Shen, Zeyu, Korolova, Aleksandra |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2602.15799 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Measuring Validity in LLM-based Resume Screening
por: Castleman, Jane, et al.
Publicado: (2026)
por: Castleman, Jane, et al.
Publicado: (2026)
Why am I Still Seeing This: Measuring the Effectiveness Of Ad Controls and Explanations in AI-Mediated Ad Targeting Systems
por: Castleman, Jane, et al.
Publicado: (2024)
por: Castleman, Jane, et al.
Publicado: (2024)
Adultification Bias in LLMs and Text-to-Image Models
por: Castleman, Jane, et al.
Publicado: (2025)
por: Castleman, Jane, et al.
Publicado: (2025)
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
por: Chittepu, Yaswanth, et al.
Publicado: (2025)
por: Chittepu, Yaswanth, et al.
Publicado: (2025)
ElliCE: Efficient and Provably Robust Algorithmic Recourse via the Rashomon Sets
por: Turbal, Bohdan, et al.
Publicado: (2026)
por: Turbal, Bohdan, et al.
Publicado: (2026)
When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models
por: Hossain, Ismail, et al.
Publicado: (2026)
por: Hossain, Ismail, et al.
Publicado: (2026)
External Evaluation of Discrimination Mitigation Efforts in Meta's Ad Delivery
por: Imana, Basileal, et al.
Publicado: (2025)
por: Imana, Basileal, et al.
Publicado: (2025)
On Adversarial Robustness of Language Models in Transfer Learning
por: Turbal, Bohdan, et al.
Publicado: (2024)
por: Turbal, Bohdan, et al.
Publicado: (2024)
Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs
por: Roh, Jaechul, et al.
Publicado: (2026)
por: Roh, Jaechul, et al.
Publicado: (2026)
ReliabilityRAG: Effective and Provably Robust Defense for RAG-based Web-Search
por: Shen, Zeyu, et al.
Publicado: (2025)
por: Shen, Zeyu, et al.
Publicado: (2025)
Auditing for Racial Discrimination in the Delivery of Education Ads
por: Imana, Basileal, et al.
Publicado: (2024)
por: Imana, Basileal, et al.
Publicado: (2024)
Auditing for Bias in Ad Delivery Using Inferred Demographic Attributes
por: Imana, Basileal, et al.
Publicado: (2024)
por: Imana, Basileal, et al.
Publicado: (2024)
Multilevel Analysis of Cryptocurrency News using RAG Approach with Fine-Tuned Mistral Large Language Model
por: Pavlyshenko, Bohdan M.
Publicado: (2025)
por: Pavlyshenko, Bohdan M.
Publicado: (2025)
Stability and Multigroup Fairness in Ranking with Uncertain Predictions
por: Devic, Siddartha, et al.
Publicado: (2024)
por: Devic, Siddartha, et al.
Publicado: (2024)
On the Use of Proxies in Political Ad Targeting
por: Sapiezynski, Piotr, et al.
Publicado: (2024)
por: Sapiezynski, Piotr, et al.
Publicado: (2024)
Paremias of the Latvians and the Russians in Latgale: From the Holy Scripture to Modern Existence
por: Jelena Korolova
Publicado: (2020)
por: Jelena Korolova
Publicado: (2020)
Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models
por: Goel, Anmol, et al.
Publicado: (2026)
por: Goel, Anmol, et al.
Publicado: (2026)
Self-Refining Topology Optimization via an LLM-Based Multi-Agent Framework
por: Park, Hyunjee, et al.
Publicado: (2026)
por: Park, Hyunjee, et al.
Publicado: (2026)
"I Have a Dream, Too!": The American Dream in Coretta Scott King Award-Winning Books
por: Parsons, Linda T., et al.
Publicado: (2011)
por: Parsons, Linda T., et al.
Publicado: (2011)
Before 2000: Funding Technology in New Jersey's Schools and Public Libraries by the End of the Century.
por: Peretz, Blossom A.
Publicado: (1997)
por: Peretz, Blossom A.
Publicado: (1997)
RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
por: Asif, Sadia, et al.
Publicado: (2026)
por: Asif, Sadia, et al.
Publicado: (2026)
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
por: Xiao, Yuxin, et al.
Publicado: (2025)
por: Xiao, Yuxin, et al.
Publicado: (2025)
Multi-Selection for Recommendation Systems
por: Sarmasarkar, Sahasrajit, et al.
Publicado: (2025)
por: Sarmasarkar, Sahasrajit, et al.
Publicado: (2025)
An External Fairness Evaluation of LinkedIn Talent Search
por: Behzad, Tina, et al.
Publicado: (2025)
por: Behzad, Tina, et al.
Publicado: (2025)
Differential Privacy with Multiple Selections
por: Goel, Ashish, et al.
Publicado: (2024)
por: Goel, Ashish, et al.
Publicado: (2024)
Exploring the Impact of Childhood Trauma Profiles on Social Competence and Self‐Stigma of Seeking Help in Early Adulthood
por: Hayoung Jung, et al.
Publicado: (2025)
por: Hayoung Jung, et al.
Publicado: (2025)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
por: Sahoo, Subramanyam
Publicado: (2026)
por: Sahoo, Subramanyam
Publicado: (2026)
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning
por: Hossain, Saad, et al.
Publicado: (2025)
por: Hossain, Saad, et al.
Publicado: (2025)
Examining the Influence of Varied Levels of Domain Knowledge Base Inclusion in GPT-based Intelligent Tutors
por: Castleman, Blake, et al.
Publicado: (2023)
por: Castleman, Blake, et al.
Publicado: (2023)
Modulation of Cell Cycle Kinases by Kaposi's Sarcoma‐Associated Herpesvirus
por: Steven Longworth, et al.
Publicado: (2025)
por: Steven Longworth, et al.
Publicado: (2025)
Narrow Fine-Tuning Erodes Safety Alignment in Vision-Language Agents
por: Gulati, Idhant, et al.
Publicado: (2026)
por: Gulati, Idhant, et al.
Publicado: (2026)
Stealth Fine-Tuning: Efficiently Breaking Alignment in RVLMs Using Self-Generated CoT
por: Yu, Le, et al.
Publicado: (2025)
por: Yu, Le, et al.
Publicado: (2025)
When Is Rank-1 Enough? Geometry-Guided Initialization for Parameter-Efficient Fine-Tuning
por: Zhao, Haoran, et al.
Publicado: (2026)
por: Zhao, Haoran, et al.
Publicado: (2026)
Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry
por: Lan, Wenhao, et al.
Publicado: (2026)
por: Lan, Wenhao, et al.
Publicado: (2026)
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
por: Hsiung, Lei, et al.
Publicado: (2025)
por: Hsiung, Lei, et al.
Publicado: (2025)
GR-SAP: Generative Replay for Safety Alignment Preservation during Fine-Tuning
por: Fang, Zhouxiang, et al.
Publicado: (2026)
por: Fang, Zhouxiang, et al.
Publicado: (2026)
PREFINE: Preference-Based Implicit Reward and Cost Fine-Tuning for Safety Alignment
por: Verma, Richa, et al.
Publicado: (2026)
por: Verma, Richa, et al.
Publicado: (2026)
Targeted Vaccine: Safety Alignment for Large Language Models against Harmful Fine-Tuning via Layer-wise Perturbation
por: Liu, Guozhi, et al.
Publicado: (2024)
por: Liu, Guozhi, et al.
Publicado: (2024)
Algorithmic Behaviors Across Regions: A Geolocation Audit of YouTube Search for COVID-19 Misinformation Between the United States and South Africa
por: Jung, Hayoung, et al.
Publicado: (2024)
por: Jung, Hayoung, et al.
Publicado: (2024)
Safety Geometry Collapse in Multimodal LLMs and Adaptive Drift Correction
por: Guo, Jiahe, et al.
Publicado: (2026)
por: Guo, Jiahe, et al.
Publicado: (2026)
Ejemplares similares
-
Measuring Validity in LLM-based Resume Screening
por: Castleman, Jane, et al.
Publicado: (2026) -
Why am I Still Seeing This: Measuring the Effectiveness Of Ad Controls and Explanations in AI-Mediated Ad Targeting Systems
por: Castleman, Jane, et al.
Publicado: (2024) -
Adultification Bias in LLMs and Text-to-Image Models
por: Castleman, Jane, et al.
Publicado: (2025) -
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
por: Chittepu, Yaswanth, et al.
Publicado: (2025) -
ElliCE: Efficient and Provably Robust Algorithmic Recourse via the Rashomon Sets
por: Turbal, Bohdan, et al.
Publicado: (2026)