MindGuard: Guardrail Classifiers for Multi-Turn Mental Health Support
Fuente:
arXiv
Salvato in:
| Autori principali: | Farinhas, António, Guerreiro, Nuno M., Pombal, José, Martins, Pedro Henrique, Melton, Laura, Conway, Alex, Dochat, Cara, D'Eon, Maya, Rei, Ricardo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MindEval: Benchmarking Language Models on Multi-turn Mental Health Support
di: Pombal, José, et al.
Pubblicazione: (2025)
di: Pombal, José, et al.
Pubblicazione: (2025)
MindGuard: Towards Accessible and Sitgma-free Mental Health First Aid via Edge LLM
di: Ji, Sijie, et al.
Pubblicazione: (2024)
di: Ji, Sijie, et al.
Pubblicazione: (2024)
Adding Chocolate to Mint: Mitigating Metric Interference in Machine Translation
di: Pombal, José, et al.
Pubblicazione: (2025)
di: Pombal, José, et al.
Pubblicazione: (2025)
Zero-shot Benchmarking: A Framework for Flexible and Scalable Automatic Evaluation of Language Models
di: Pombal, José, et al.
Pubblicazione: (2025)
di: Pombal, José, et al.
Pubblicazione: (2025)
Translate Smart, not Hard: Cascaded Translation Systems with Quality-Aware Deferral
di: Farinhas, António, et al.
Pubblicazione: (2025)
di: Farinhas, António, et al.
Pubblicazione: (2025)
MindGuard: Intrinsic Decision Inspection for Securing LLM Agents Against Metadata Poisoning
di: Wang, Zhiqiang, et al.
Pubblicazione: (2025)
di: Wang, Zhiqiang, et al.
Pubblicazione: (2025)
Tower+: Bridging Generality and Translation Specialization in Multilingual LLMs
di: Rei, Ricardo, et al.
Pubblicazione: (2025)
di: Rei, Ricardo, et al.
Pubblicazione: (2025)
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models
di: Pombal, José, et al.
Pubblicazione: (2026)
di: Pombal, José, et al.
Pubblicazione: (2026)
Can Automatic Metrics Assess High-Quality Translations?
di: Agrawal, Sweta, et al.
Pubblicazione: (2024)
di: Agrawal, Sweta, et al.
Pubblicazione: (2024)
Modeling User Preferences with Automatic Metrics: Creating a High-Quality Preference Dataset for Machine Translation
di: Agrawal, Sweta, et al.
Pubblicazione: (2024)
di: Agrawal, Sweta, et al.
Pubblicazione: (2024)
Tradução, adaptação cultural e validação da versão portuguesa do Children with Special Health Care Needs Screener
di: Fernanda Pombal
Pubblicazione: (2024)
di: Fernanda Pombal
Pubblicazione: (2024)
xTower: A Multilingual LLM for Explaining and Correcting Translation Errors
di: Treviso, Marcos, et al.
Pubblicazione: (2024)
di: Treviso, Marcos, et al.
Pubblicazione: (2024)
SEQUOR: A Multi-Turn Benchmark for Realistic Constraint Following
di: Canaverde, Beatriz, et al.
Pubblicazione: (2026)
di: Canaverde, Beatriz, et al.
Pubblicazione: (2026)
CodeGuard: Improving LLM Guardrails in CS Education
di: Raihan, Nishat, et al.
Pubblicazione: (2026)
di: Raihan, Nishat, et al.
Pubblicazione: (2026)
PoseGuard: Pose-Guided Generation with Safety Guardrails
di: Wang, Kongxin, et al.
Pubblicazione: (2025)
di: Wang, Kongxin, et al.
Pubblicazione: (2025)
Communication is Translation, or, How to Mind the Gap
di: Kyle Conway
Pubblicazione: (2017)
di: Kyle Conway
Pubblicazione: (2017)
Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
di: Giarrusso, Francesco, et al.
Pubblicazione: (2025)
di: Giarrusso, Francesco, et al.
Pubblicazione: (2025)
WebGuard: Building a Generalizable Guardrail for Web Agents
di: Zheng, Boyuan, et al.
Pubblicazione: (2025)
di: Zheng, Boyuan, et al.
Pubblicazione: (2025)
OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
di: Zhu, Boyu, et al.
Pubblicazione: (2025)
di: Zhu, Boyu, et al.
Pubblicazione: (2025)
BiasGuard: Guardrailing Fairness in Machine Learning Production Systems
di: Cohen-Inger, Nurit, et al.
Pubblicazione: (2025)
di: Cohen-Inger, Nurit, et al.
Pubblicazione: (2025)
ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails
di: Wen, Xiaofei, et al.
Pubblicazione: (2025)
di: Wen, Xiaofei, et al.
Pubblicazione: (2025)
LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models
di: Elesedy, Hayder, et al.
Pubblicazione: (2024)
di: Elesedy, Hayder, et al.
Pubblicazione: (2024)
EuroLLM: Multilingual Language Models for Europe
di: Martins, Pedro Henrique, et al.
Pubblicazione: (2024)
di: Martins, Pedro Henrique, et al.
Pubblicazione: (2024)
Tower: An Open Multilingual Large Language Model for Translation-Related Tasks
di: Alves, Duarte M., et al.
Pubblicazione: (2024)
di: Alves, Duarte M., et al.
Pubblicazione: (2024)
The Slow Drift of Support: Boundary Failures in Multi-Turn Mental Health LLM Dialogues
di: Cheng, Youyou, et al.
Pubblicazione: (2026)
di: Cheng, Youyou, et al.
Pubblicazione: (2026)
Walkability and Mental Health Resiliency During the COVID‐19 Pandemic
di: Karen Smith Conway, et al.
Pubblicazione: (2025)
di: Karen Smith Conway, et al.
Pubblicazione: (2025)
ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails
di: Wang, Yan, et al.
Pubblicazione: (2026)
di: Wang, Yan, et al.
Pubblicazione: (2026)
MrGuard: A Multilingual Reasoning Guardrail for Universal LLM Safety
di: Yang, Yahan, et al.
Pubblicazione: (2025)
di: Yang, Yahan, et al.
Pubblicazione: (2025)
SentGuard: Sentence-Level Streaming Guardrails for Large Language Models
di: Yu, Jiaqi, et al.
Pubblicazione: (2026)
di: Yu, Jiaqi, et al.
Pubblicazione: (2026)
Australia's U‐Turn on Chinese Investment: A Neoclassical Realist Perspective
di: Rei Koga
Pubblicazione: (2025)
di: Rei Koga
Pubblicazione: (2025)
Be Mindful of Students' Mental Health Needs
di: Dawn Z. Hodges
Pubblicazione: (2025)
di: Dawn Z. Hodges
Pubblicazione: (2025)
Be Mindful of Students’ Mental Health Needs
di: Dawn Z. Hodges
Pubblicazione: (2025)
di: Dawn Z. Hodges
Pubblicazione: (2025)
Minimum Wages, the Earned Income Tax Credit, and Mental Health Around Pregnancy
di: Bryce J. Stanley, et al.
Pubblicazione: (2025)
di: Bryce J. Stanley, et al.
Pubblicazione: (2025)
Poly-Guard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset
di: Kang, Mintong, et al.
Pubblicazione: (2025)
di: Kang, Mintong, et al.
Pubblicazione: (2025)
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
di: Li, Hao, et al.
Pubblicazione: (2024)
di: Li, Hao, et al.
Pubblicazione: (2024)
EuroLLM-9B: Technical Report
di: Martins, Pedro Henrique, et al.
Pubblicazione: (2025)
di: Martins, Pedro Henrique, et al.
Pubblicazione: (2025)
Analyzing Context Contributions in LLM-based Machine Translation
di: Zaranis, Emmanouil, et al.
Pubblicazione: (2024)
di: Zaranis, Emmanouil, et al.
Pubblicazione: (2024)
Reranking Laws for Language Generation: A Communication-Theoretic Perspective
di: Farinhas, António, et al.
Pubblicazione: (2024)
di: Farinhas, António, et al.
Pubblicazione: (2024)
Beyond Red-Teaming: Formal Guarantees of LLM Guardrail Classifiers
di: Kezins, Nikita, et al.
Pubblicazione: (2026)
di: Kezins, Nikita, et al.
Pubblicazione: (2026)
ConceptGuard: Neuro-Symbolic Safety Guardrails via Sparse Interpretable Jailbreak Concepts
di: Aswal, Darpan, et al.
Pubblicazione: (2025)
di: Aswal, Darpan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
MindEval: Benchmarking Language Models on Multi-turn Mental Health Support
di: Pombal, José, et al.
Pubblicazione: (2025) -
MindGuard: Towards Accessible and Sitgma-free Mental Health First Aid via Edge LLM
di: Ji, Sijie, et al.
Pubblicazione: (2024) -
Adding Chocolate to Mint: Mitigating Metric Interference in Machine Translation
di: Pombal, José, et al.
Pubblicazione: (2025) -
Zero-shot Benchmarking: A Framework for Flexible and Scalable Automatic Evaluation of Language Models
di: Pombal, José, et al.
Pubblicazione: (2025) -
Translate Smart, not Hard: Cascaded Translation Systems with Quality-Aware Deferral
di: Farinhas, António, et al.
Pubblicazione: (2025)