Safer or Luckier? LLMs as Safety Evaluators Are Not Robust to Artifacts
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Hongyu, Goldfarb-Tarrant, Seraphina |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them
por: Rajani, Neel, et al.
Publicado: (2025)
por: Rajani, Neel, et al.
Publicado: (2025)
MultiContrievers: Analysis of Dense Retrieval Representations
por: Goldfarb-Tarrant, Seraphina, et al.
Publicado: (2024)
por: Goldfarb-Tarrant, Seraphina, et al.
Publicado: (2024)
Small Changes, Large Consequences: Analyzing the Allocational Fairness of LLMs in Hiring Contexts
por: Seshadri, Preethi, et al.
Publicado: (2025)
por: Seshadri, Preethi, et al.
Publicado: (2025)
The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm
por: Aakanksha, et al.
Publicado: (2024)
por: Aakanksha, et al.
Publicado: (2024)
Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations
por: Seshadri, Preethi, et al.
Publicado: (2026)
por: Seshadri, Preethi, et al.
Publicado: (2026)
Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism
por: Orgad, Hadas, et al.
Publicado: (2026)
por: Orgad, Hadas, et al.
Publicado: (2026)
The Multilingual Divide and Its Impact on Global AI Safety
por: Peppin, Aidan, et al.
Publicado: (2025)
por: Peppin, Aidan, et al.
Publicado: (2025)
Are Smarter LLMs Safer? Exploring Safety-Reasoning Trade-offs in Prompting and Fine-Tuning
por: Li, Ang, et al.
Publicado: (2025)
por: Li, Ang, et al.
Publicado: (2025)
RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models
por: An, Bang, et al.
Publicado: (2025)
por: An, Bang, et al.
Publicado: (2025)
Learning is Forgetting: LLM Training As Lossy Compression
por: Conklin, Henry C., et al.
Publicado: (2026)
por: Conklin, Henry C., et al.
Publicado: (2026)
STAR-1: Safer Alignment of Reasoning LLMs with 1K Data
por: Wang, Zijun, et al.
Publicado: (2025)
por: Wang, Zijun, et al.
Publicado: (2025)
Models That Know How Evaluations Are Designed Score Safer
por: Deckenbach, Katharina, et al.
Publicado: (2026)
por: Deckenbach, Katharina, et al.
Publicado: (2026)
Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs
por: Hua, Andong, et al.
Publicado: (2025)
por: Hua, Andong, et al.
Publicado: (2025)
Speech LLMs in Low-Resource Scenarios: Data Volume Requirements and the Impact of Pretraining on High-Resource Languages
por: Fong, Seraphina, et al.
Publicado: (2025)
por: Fong, Seraphina, et al.
Publicado: (2025)
Safer-Instruct: Aligning Language Models with Automated Preference Data
por: Shi, Taiwei, et al.
Publicado: (2023)
por: Shi, Taiwei, et al.
Publicado: (2023)
Supporting Artifact Evaluation with LLMs: A Study with Published Security Research Papers
por: Heye, David, et al.
Publicado: (2026)
por: Heye, David, et al.
Publicado: (2026)
A SMART Mnemonic Sounds like "Glue Tonic": Mixing LLMs with Student Feedback to Make Mnemonic Learning Stick
por: Balepur, Nishant, et al.
Publicado: (2024)
por: Balepur, Nishant, et al.
Publicado: (2024)
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
por: Lunardi, Riccardo, et al.
Publicado: (2025)
por: Lunardi, Riccardo, et al.
Publicado: (2025)
Mix Data or Merge Models? Optimizing for Diverse Multi-Task Learning
por: Aakanksha, et al.
Publicado: (2024)
por: Aakanksha, et al.
Publicado: (2024)
InvThink: Premortem Reasoning for Safer Language Models
por: Kim, Yubin, et al.
Publicado: (2025)
por: Kim, Yubin, et al.
Publicado: (2025)
Deliberative Alignment: Reasoning Enables Safer Language Models
por: Guan, Melody Y., et al.
Publicado: (2024)
por: Guan, Melody Y., et al.
Publicado: (2024)
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
por: Thomas, Rohan Subramanian, et al.
Publicado: (2026)
por: Thomas, Rohan Subramanian, et al.
Publicado: (2026)
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
por: Liu, Shudong, et al.
Publicado: (2025)
por: Liu, Shudong, et al.
Publicado: (2025)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
por: Siu, Vincent, et al.
Publicado: (2025)
por: Siu, Vincent, et al.
Publicado: (2025)
The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts
por: Shen, Lingfeng, et al.
Publicado: (2024)
por: Shen, Lingfeng, et al.
Publicado: (2024)
MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine
por: Kong, Shufeng, et al.
Publicado: (2025)
por: Kong, Shufeng, et al.
Publicado: (2025)
Visuospatial Perspective Taking in Multimodal Language Models
por: Prunty, Jonathan, et al.
Publicado: (2026)
por: Prunty, Jonathan, et al.
Publicado: (2026)
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
por: Liu, Songyang, et al.
Publicado: (2025)
por: Liu, Songyang, et al.
Publicado: (2025)
Preventing Catastrophic Forgetting: Behavior-Aware Sampling for Safer Language Model Fine-Tuning
por: Pham, Anh, et al.
Publicado: (2025)
por: Pham, Anh, et al.
Publicado: (2025)
JT-Safe: Intrinsically Enhancing the Safety and Trustworthiness of LLMs
por: Feng, Junlan, et al.
Publicado: (2025)
por: Feng, Junlan, et al.
Publicado: (2025)
Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity
por: Wu, Xinwei, et al.
Publicado: (2025)
por: Wu, Xinwei, et al.
Publicado: (2025)
Trans-EnV: A Framework for Evaluating the Linguistic Robustness of LLMs Against English Varieties
por: Lee, Jiyoung, et al.
Publicado: (2025)
por: Lee, Jiyoung, et al.
Publicado: (2025)
MoralBench: Moral Evaluation of LLMs
por: Ji, Jianchao, et al.
Publicado: (2024)
por: Ji, Jianchao, et al.
Publicado: (2024)
Cross-Platform Evaluation of Large Language Model Safety in Pediatric Consultations: Evolution of Adversarial Robustness and the Scale Paradox
por: Zolfaghari, Vahideh
Publicado: (2025)
por: Zolfaghari, Vahideh
Publicado: (2025)
A Multi-Domain Red Teaming Framework for Safety, Robustness, and Fairness Evaluation of Medical Large Language Models
por: Feier, Andrei Marian, et al.
Publicado: (2026)
por: Feier, Andrei Marian, et al.
Publicado: (2026)
Identity Lock: Locking API Fine-tuned LLMs With Identity-based Wake Words
por: Su, Hongyu, et al.
Publicado: (2025)
por: Su, Hongyu, et al.
Publicado: (2025)
From Informal to Formal -- Incorporating and Evaluating LLMs on Natural Language Requirements to Verifiable Formal Proofs
por: Cao, Jialun, et al.
Publicado: (2025)
por: Cao, Jialun, et al.
Publicado: (2025)
Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
por: Aldahoul, Nouar, et al.
Publicado: (2025)
por: Aldahoul, Nouar, et al.
Publicado: (2025)
LongSafety: Evaluating Long-Context Safety of Large Language Models
por: Lu, Yida, et al.
Publicado: (2025)
por: Lu, Yida, et al.
Publicado: (2025)
Food Noise & False Safety: A Systematic Evaluation of How LLMs Fail to Adapt to Eating Disorder Queries with Clinician Feedback
por: Pucci, Giulia, et al.
Publicado: (2026)
por: Pucci, Giulia, et al.
Publicado: (2026)
Ejemplares similares
-
Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them
por: Rajani, Neel, et al.
Publicado: (2025) -
MultiContrievers: Analysis of Dense Retrieval Representations
por: Goldfarb-Tarrant, Seraphina, et al.
Publicado: (2024) -
Small Changes, Large Consequences: Analyzing the Allocational Fairness of LLMs in Hiring Contexts
por: Seshadri, Preethi, et al.
Publicado: (2025) -
The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm
por: Aakanksha, et al.
Publicado: (2024) -
Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations
por: Seshadri, Preethi, et al.
Publicado: (2026)