Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Yue, Zhuang, Haomin, Ye, Jiayi, Bao, Han, Wang, Yanbo, Hua, Hang, Wu, Siyuan, Chen, Pin-Yu, Zhang, Xiangliang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Granite Guardian
by: Padhi, Inkit, et al.
Published: (2024)
by: Padhi, Inkit, et al.
Published: (2024)
Adaptive Distraction: Probing LLM Contextual Robustness with Automated Tree Search
by: Wang, Yanbo, et al.
Published: (2025)
by: Wang, Yanbo, et al.
Published: (2025)
Guardians of Discourse: Evaluating LLMs on Multilingual Offensive Language Detection
by: He, Jianfei, et al.
Published: (2024)
by: He, Jianfei, et al.
Published: (2024)
Dual Optimal: Make Your LLM Peer-like with Dignity
by: Wang, Xiangqi, et al.
Published: (2026)
by: Wang, Xiangqi, et al.
Published: (2026)
Dissecting Logical Reasoning in LLMs: A Fine-Grained Evaluation and Supervision Study
by: Zhou, Yujun, et al.
Published: (2025)
by: Zhou, Yujun, et al.
Published: (2025)
SLM as Guardian: Pioneering AI Safety with Small Language Models
by: Kwon, Ohjoon, et al.
Published: (2024)
by: Kwon, Ohjoon, et al.
Published: (2024)
Social Science Meets LLMs: How Reliable Are Large Language Models in Social Simulations?
by: Huang, Yue, et al.
Published: (2024)
by: Huang, Yue, et al.
Published: (2024)
DynaGuard: A Dynamic Guardian Model With User-Defined Policies
by: Hoover, Monte, et al.
Published: (2025)
by: Hoover, Monte, et al.
Published: (2025)
Reliable Control-Point Selection for Steering Reasoning in Large Language Models
by: Zhuang, Haomin, et al.
Published: (2026)
by: Zhuang, Haomin, et al.
Published: (2026)
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
Guardians of the Machine Translation Meta-Evaluation: Sentinel Metrics Fall In!
by: Perrella, Stefano, et al.
Published: (2024)
by: Perrella, Stefano, et al.
Published: (2024)
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
by: Xu, Zixiang, et al.
Published: (2025)
by: Xu, Zixiang, et al.
Published: (2025)
Exploring Multi-Temperature Strategies for Token- and Rollout-Level Control in RLVR
by: Zhuang, Haomin, et al.
Published: (2025)
by: Zhuang, Haomin, et al.
Published: (2025)
The Privacy Guardian Agent: Towards Trustworthy AI Privacy Agents
by: Freiberger, Vincent
Published: (2026)
by: Freiberger, Vincent
Published: (2026)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
by: Zhuang, Haomin, et al.
Published: (2024)
by: Zhuang, Haomin, et al.
Published: (2024)
ChemOrch: Empowering LLMs with Chemical Intelligence via Synthetic Instructions
by: Huang, Yue, et al.
Published: (2025)
by: Huang, Yue, et al.
Published: (2025)
Large language models for newspaper sentiment analysis during COVID-19: The Guardian
by: Chandra, Rohitash, et al.
Published: (2024)
by: Chandra, Rohitash, et al.
Published: (2024)
Equality's Guardians
Published: (2025)
Published: (2025)
Guardians of the Sea
Published: (2018)
Published: (2018)
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
by: Ye, Jiayi, et al.
Published: (2024)
by: Ye, Jiayi, et al.
Published: (2024)
SWAP: Study Work Advisor Program: Handbook for the Project Director, Sponsor, Employer, Parent or Guardian.
by: Boeyink, Joann, et al.
Published: (1973)
by: Boeyink, Joann, et al.
Published: (1973)
UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models
by: Lin, Huawei, et al.
Published: (2025)
by: Lin, Huawei, et al.
Published: (2025)
How to Train Your Advisor: Steering Black-Box LLMs with Advisor Models
by: Asawa, Parth, et al.
Published: (2025)
by: Asawa, Parth, et al.
Published: (2025)
Cross-Lingual Pitfalls: Automatic Probing Cross-Lingual Weakness of Multilingual Large Language Models
by: Xu, Zixiang, et al.
Published: (2025)
by: Xu, Zixiang, et al.
Published: (2025)
The Guardian and The NY Times
by: Anonymous Author
Published: (2025)
by: Anonymous Author
Published: (2025)
Economic Research Guardian
Published: (2013)
Published: (2013)
Guardians of Land and Water
Published: (2025)
Published: (2025)
Guardians of Public Value
by: Boin, Arjen, et al.
Published: (2020)
by: Boin, Arjen, et al.
Published: (2020)
Guardians of the Quantum GAN
by: Ghosh, Archisman, et al.
Published: (2024)
by: Ghosh, Archisman, et al.
Published: (2024)
1+1>2: Can Large Language Models Serve as Cross-Lingual Knowledge Aggregators?
by: Huang, Yue, et al.
Published: (2024)
by: Huang, Yue, et al.
Published: (2024)
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
by: Wang, Yanbo, et al.
Published: (2025)
by: Wang, Yanbo, et al.
Published: (2025)
Emergent Social Intelligence Risks in Generative Multi-Agent Systems
by: Huang, Yue, et al.
Published: (2026)
by: Huang, Yue, et al.
Published: (2026)
Guardian Lion of Route 66
by: Scan-the-World
Published: (2026)
by: Scan-the-World
Published: (2026)
Guardián de la conciencia
by: Egan, Linda
Published: (2008)
by: Egan, Linda
Published: (2008)
Guardianes de la red
by: Villalobos, Jorge
Published: (2007)
by: Villalobos, Jorge
Published: (2007)
AI Alignment Breaks at the Edge
by: Bao, Han, et al.
Published: (2026)
by: Bao, Han, et al.
Published: (2026)
TrustLLM: Trustworthiness in Large Language Models
by: Huang, Yue, et al.
Published: (2024)
by: Huang, Yue, et al.
Published: (2024)
DataGen: Unified Synthetic Dataset Generation via Large Language Models
by: Huang, Yue, et al.
Published: (2024)
by: Huang, Yue, et al.
Published: (2024)
LegalGuardian: A Privacy-Preserving Framework for Secure Integration of Large Language Models in Legal Practice
by: Demir, M. Mikail, et al.
Published: (2025)
by: Demir, M. Mikail, et al.
Published: (2025)
SenseMath: Do LLMs Have Number Sense? Evaluating Shortcut Use, Judgment, and Generation
by: Zhuang, Haomin, et al.
Published: (2026)
by: Zhuang, Haomin, et al.
Published: (2026)
Similar Items
-
Granite Guardian
by: Padhi, Inkit, et al.
Published: (2024) -
Adaptive Distraction: Probing LLM Contextual Robustness with Automated Tree Search
by: Wang, Yanbo, et al.
Published: (2025) -
Guardians of Discourse: Evaluating LLMs on Multilingual Offensive Language Detection
by: He, Jianfei, et al.
Published: (2024) -
Dual Optimal: Make Your LLM Peer-like with Dignity
by: Wang, Xiangqi, et al.
Published: (2026) -
Dissecting Logical Reasoning in LLMs: A Fine-Grained Evaluation and Supervision Study
by: Zhou, Yujun, et al.
Published: (2025)