CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety
Fuente:
arXiv
Salvato in:
| Autori principali: | An, Heajun, Zhang, Qi, Achanta, Vedanth, Cho, Jin-Hee |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LLM Can be a Dangerous Persuader: Empirical Study of Persuasion Safety in Large Language Models
di: Liu, Minqian, et al.
Pubblicazione: (2025)
di: Liu, Minqian, et al.
Pubblicazione: (2025)
Does GenAI Rewrite How We Write? An Empirical Study on Two-Million Preprints
di: Qi, Minfeng, et al.
Pubblicazione: (2025)
di: Qi, Minfeng, et al.
Pubblicazione: (2025)
Unmasking the Canvas: A Dynamic Benchmark for Image Generation Jailbreaking and LLM Content Safety
di: Nair, Variath Madhupal Gautham, et al.
Pubblicazione: (2025)
di: Nair, Variath Madhupal Gautham, et al.
Pubblicazione: (2025)
Building Effective Safety Guardrails in AI Education Tools
di: Clark, Hannah-Beth, et al.
Pubblicazione: (2025)
di: Clark, Hannah-Beth, et al.
Pubblicazione: (2025)
CodeGuard: Improving LLM Guardrails in CS Education
di: Raihan, Nishat, et al.
Pubblicazione: (2026)
di: Raihan, Nishat, et al.
Pubblicazione: (2026)
Ensuring Safety and Trust: Analyzing the Risks of Large Language Models in Medicine
di: Yang, Yifan, et al.
Pubblicazione: (2024)
di: Yang, Yifan, et al.
Pubblicazione: (2024)
A Lightweight Explainable Guardrail for Prompt Safety
di: Islam, Md Asiful, et al.
Pubblicazione: (2026)
di: Islam, Md Asiful, et al.
Pubblicazione: (2026)
LLM Nepotism in Organizational Governance
di: Mao, Shunqi, et al.
Pubblicazione: (2026)
di: Mao, Shunqi, et al.
Pubblicazione: (2026)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
di: Zhou, Zhenhong, et al.
Pubblicazione: (2024)
di: Zhou, Zhenhong, et al.
Pubblicazione: (2024)
Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
di: Dong, Zhichen, et al.
Pubblicazione: (2024)
di: Dong, Zhichen, et al.
Pubblicazione: (2024)
Beyond Translation: LLM-Based Data Generation for Multilingual Fact-Checking
di: Chung, Yi-Ling, et al.
Pubblicazione: (2025)
di: Chung, Yi-Ling, et al.
Pubblicazione: (2025)
Place Matters: Comparing LLM Hallucination Rates for Place-Based Legal Queries
di: Curran, Damian, et al.
Pubblicazione: (2025)
di: Curran, Damian, et al.
Pubblicazione: (2025)
Safety Evaluation and Enhancement of DeepSeek Models in Chinese Contexts
di: Zhang, Wenjing, et al.
Pubblicazione: (2025)
di: Zhang, Wenjing, et al.
Pubblicazione: (2025)
When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
di: Backmann, Steffen, et al.
Pubblicazione: (2025)
di: Backmann, Steffen, et al.
Pubblicazione: (2025)
Evaluating Psychological Safety of Large Language Models
di: Li, Xingxuan, et al.
Pubblicazione: (2022)
di: Li, Xingxuan, et al.
Pubblicazione: (2022)
Empirical Evidence for Alignment Faking in a Small LLM and Prompt-Based Mitigation Techniques
di: Koorndijk, Jeanice
Pubblicazione: (2025)
di: Koorndijk, Jeanice
Pubblicazione: (2025)
SGuard-v1: Safety Guardrail for Large Language Models
di: Lee, JoonHo, et al.
Pubblicazione: (2025)
di: Lee, JoonHo, et al.
Pubblicazione: (2025)
Psychometric Comparability of LLM-Based Digital Twins
di: Zhang, Yufei, et al.
Pubblicazione: (2025)
di: Zhang, Yufei, et al.
Pubblicazione: (2025)
Building Trust: Foundations of Security, Safety and Transparency in AI
di: Sidhpurwala, Huzaifa, et al.
Pubblicazione: (2024)
di: Sidhpurwala, Huzaifa, et al.
Pubblicazione: (2024)
Measuring What Matters -- or What's Convenient?: Robustness of LLM-Based Scoring Systems to Construct-Irrelevant Factors
di: Walsh, Cole, et al.
Pubblicazione: (2026)
di: Walsh, Cole, et al.
Pubblicazione: (2026)
Gender and Positional Biases in LLM-Based Hiring Decisions: Evidence from Comparative CV/Résumé Evaluations
di: Rozado, David
Pubblicazione: (2025)
di: Rozado, David
Pubblicazione: (2025)
The Homogenization Problem in LLMs: Towards Meaningful Diversity in AI Safety
di: Rios-Sialer, Ian
Pubblicazione: (2026)
di: Rios-Sialer, Ian
Pubblicazione: (2026)
Building Guardrails for Large Language Models
di: Dong, Yi, et al.
Pubblicazione: (2024)
di: Dong, Yi, et al.
Pubblicazione: (2024)
TherapyGym: Evaluating and Aligning Clinical Fidelity and Safety in Therapy Chatbots
di: Huang, Fangrui, et al.
Pubblicazione: (2026)
di: Huang, Fangrui, et al.
Pubblicazione: (2026)
Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Models
di: Sreedhar, Makesh Narsimhan, et al.
Pubblicazione: (2025)
di: Sreedhar, Makesh Narsimhan, et al.
Pubblicazione: (2025)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
di: Tang, Zeyu, et al.
Pubblicazione: (2026)
di: Tang, Zeyu, et al.
Pubblicazione: (2026)
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
di: Nghiem, Huy, et al.
Pubblicazione: (2025)
di: Nghiem, Huy, et al.
Pubblicazione: (2025)
From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
di: Yuan, Yuan, et al.
Pubblicazione: (2025)
di: Yuan, Yuan, et al.
Pubblicazione: (2025)
Deep Research with Open-Domain Evaluation and Multi-Stage Guardrails for Safety
di: Huang, Wei-Chieh, et al.
Pubblicazione: (2025)
di: Huang, Wei-Chieh, et al.
Pubblicazione: (2025)
PICKT: Practical Interlinked Concept Knowledge Tracing for Personalized Learning using Knowledge Map Concept Relations
di: Lee, Wonbeen, et al.
Pubblicazione: (2025)
di: Lee, Wonbeen, et al.
Pubblicazione: (2025)
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
di: Lee, Jaehyeok, et al.
Pubblicazione: (2026)
di: Lee, Jaehyeok, et al.
Pubblicazione: (2026)
Law in Silico: Simulating Legal Society with LLM-Based Agents
di: Wang, Yiding, et al.
Pubblicazione: (2025)
di: Wang, Yiding, et al.
Pubblicazione: (2025)
Safety Analysis in the Era of Large Language Models: A Case Study of STPA using ChatGPT
di: Qi, Yi, et al.
Pubblicazione: (2023)
di: Qi, Yi, et al.
Pubblicazione: (2023)
Why Do Safety Guardrails Degrade Across Languages?
di: Zhang, Max, et al.
Pubblicazione: (2026)
di: Zhang, Max, et al.
Pubblicazione: (2026)
United in Diversity? Contextual Biases in LLM-Based Predictions of the 2024 European Parliament Elections
di: von der Heyde, Leah, et al.
Pubblicazione: (2024)
di: von der Heyde, Leah, et al.
Pubblicazione: (2024)
A Scoping Review of LLM-as-a-Judge in Healthcare and the MedJUDGE Framework
di: Li, Chenyu, et al.
Pubblicazione: (2026)
di: Li, Chenyu, et al.
Pubblicazione: (2026)
Clinical Note Bloat Reduction for Efficient LLM Use
di: Cahoon, Jordan L., et al.
Pubblicazione: (2026)
di: Cahoon, Jordan L., et al.
Pubblicazione: (2026)
Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
di: Hawkins, John, et al.
Pubblicazione: (2025)
di: Hawkins, John, et al.
Pubblicazione: (2025)
Navigating LLM Ethics: Advancements, Challenges, and Future Directions
di: Jiao, Junfeng, et al.
Pubblicazione: (2024)
di: Jiao, Junfeng, et al.
Pubblicazione: (2024)
Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
di: Kwon, Jea, et al.
Pubblicazione: (2025)
di: Kwon, Jea, et al.
Pubblicazione: (2025)
Documenti analoghi
-
LLM Can be a Dangerous Persuader: Empirical Study of Persuasion Safety in Large Language Models
di: Liu, Minqian, et al.
Pubblicazione: (2025) -
Does GenAI Rewrite How We Write? An Empirical Study on Two-Million Preprints
di: Qi, Minfeng, et al.
Pubblicazione: (2025) -
Unmasking the Canvas: A Dynamic Benchmark for Image Generation Jailbreaking and LLM Content Safety
di: Nair, Variath Madhupal Gautham, et al.
Pubblicazione: (2025) -
Building Effective Safety Guardrails in AI Education Tools
di: Clark, Hannah-Beth, et al.
Pubblicazione: (2025) -
CodeGuard: Improving LLM Guardrails in CS Education
di: Raihan, Nishat, et al.
Pubblicazione: (2026)