Customize Multi-modal RAI Guardrails with Precedent-based predictions
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Cheng-Fu, Tran, Thanh, Christodoulopoulos, Christos, Ruan, Weitong, Gupta, Rahul, Chang, Kai-Wei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SWAN: Semantic Watermarking with Abstract Meaning Representation
by: Ye, Ziping, et al.
Published: (2026)
by: Ye, Ziping, et al.
Published: (2026)
AI Ethics by Design: Implementing Customizable Guardrails for Responsible AI Development
by: Šekrst, Kristina, et al.
Published: (2024)
by: Šekrst, Kristina, et al.
Published: (2024)
CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety
by: An, Heajun, et al.
Published: (2026)
by: An, Heajun, et al.
Published: (2026)
News Source Citing Patterns in AI Search Systems
by: Yang, Kai-Cheng
Published: (2025)
by: Yang, Kai-Cheng
Published: (2025)
White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMs
by: Wan, Yixin, et al.
Published: (2024)
by: Wan, Yixin, et al.
Published: (2024)
Unfair TOS: An Automated Approach using Customized BERT
by: Akash, Bathini Sai, et al.
Published: (2024)
by: Akash, Bathini Sai, et al.
Published: (2024)
InsideOut: Measuring and Mitigating Insider-Outsider Bias in Interview Script Generation
by: Wan, Yixin, et al.
Published: (2025)
by: Wan, Yixin, et al.
Published: (2025)
Accuracy and Political Bias of News Source Credibility Ratings by Large Language Models
by: Yang, Kai-Cheng, et al.
Published: (2023)
by: Yang, Kai-Cheng, et al.
Published: (2023)
Agree to Disagree? A Meta-Evaluation of LLM Misgendering
by: Subramonian, Arjun, et al.
Published: (2025)
by: Subramonian, Arjun, et al.
Published: (2025)
PROMPTEVALS: A Dataset of Assertions and Guardrails for Custom Production Large Language Model Pipelines
by: Vir, Reya, et al.
Published: (2025)
by: Vir, Reya, et al.
Published: (2025)
How Far Are LLMs from Believable AI? A Benchmark for Evaluating the Believability of Human Behavior Simulation
by: Xiao, Yang, et al.
Published: (2023)
by: Xiao, Yang, et al.
Published: (2023)
Improving Graduate Outcomes by Identifying Skills Gaps and Recommending Courses Based on Career Interests
by: Soni, Rahul, et al.
Published: (2025)
by: Soni, Rahul, et al.
Published: (2025)
A Big Data-empowered System for Real-time Detection of Regional Discriminatory Comments on Vietnamese Social Media
by: Huynh, An Nghiep, et al.
Published: (2024)
by: Huynh, An Nghiep, et al.
Published: (2024)
BeliefShift: Benchmarking Temporal Belief Consistency and Opinion Drift in LLM Agents
by: Myakala, Praveen Kumar, et al.
Published: (2026)
by: Myakala, Praveen Kumar, et al.
Published: (2026)
Leveraging Large Language Models for Predictive Analysis of Human Misery
by: Seal, Bishanka, et al.
Published: (2025)
by: Seal, Bishanka, et al.
Published: (2025)
Training-Free Cultural Alignment of Large Language Models via Persona Disagreement
by: Kiet, Huynh Trung, et al.
Published: (2026)
by: Kiet, Huynh Trung, et al.
Published: (2026)
PediatricsMQA: a Multi-modal Pediatrics Question Answering Benchmark
by: Bahaj, Adil, et al.
Published: (2025)
by: Bahaj, Adil, et al.
Published: (2025)
War and Peace (WarAgent): Large Language Model-based Multi-Agent Simulation of World Wars
by: Hua, Wenyue, et al.
Published: (2023)
by: Hua, Wenyue, et al.
Published: (2023)
None of the Above, Less of the Right: Parallel Patterns between Humans and LLMs on Multi-Choice Questions Answering
by: Tam, Zhi Rui, et al.
Published: (2025)
by: Tam, Zhi Rui, et al.
Published: (2025)
Unveiling the Truth and Facilitating Change: Towards Agent-based Large-scale Social Movement Simulation
by: Mou, Xinyi, et al.
Published: (2024)
by: Mou, Xinyi, et al.
Published: (2024)
RAGAT-Mind: A Multi-Granular Modeling Approach for Rumor Detection Based on MindSpore
by: Qin, Zhenkai, et al.
Published: (2025)
by: Qin, Zhenkai, et al.
Published: (2025)
LangLingual: A Personalised, Exercise-oriented English Language Learning Tool Leveraging Large Language Models
by: Gupta, Sammriddh, et al.
Published: (2025)
by: Gupta, Sammriddh, et al.
Published: (2025)
When Actions Go Off-Task: Detecting and Correcting Misaligned Actions in Computer-Use Agents
by: Ning, Yuting, et al.
Published: (2026)
by: Ning, Yuting, et al.
Published: (2026)
Large Language Models Require Curated Context for Reliable Political Fact-Checking -- Even with Reasoning and Web Search
by: DeVerna, Matthew R., et al.
Published: (2025)
by: DeVerna, Matthew R., et al.
Published: (2025)
Bias and Volatility: A Statistical Framework for Evaluating Large Language Model's Stereotypes and the Associated Generation Inconsistency
by: Liu, Yiran, et al.
Published: (2024)
by: Liu, Yiran, et al.
Published: (2024)
Compounding Disadvantage: Auditing Intersectional Bias in LLM-Generated Explanations Across Indian and American STEM Education
by: Gupta, Amogh, et al.
Published: (2026)
by: Gupta, Amogh, et al.
Published: (2026)
"Mirror" Language AI Models of Depression are Criterion-Contaminated
by: Li, Tong, et al.
Published: (2025)
by: Li, Tong, et al.
Published: (2025)
Mitigating Bias for Question Answering Models by Tracking Bias Influence
by: Ma, Mingyu Derek, et al.
Published: (2023)
by: Ma, Mingyu Derek, et al.
Published: (2023)
Responsible Artificial Intelligence (RAI) in U.S. Federal Government : Principles, Policies, and Practices
by: Rawal, Atul, et al.
Published: (2025)
by: Rawal, Atul, et al.
Published: (2025)
Reinforcing Stereotypes of Anger: Emotion AI on African American Vernacular English
by: Dorn, Rebecca, et al.
Published: (2025)
by: Dorn, Rebecca, et al.
Published: (2025)
Comprehensive Study on Sentiment Analysis: From Rule-based to modern LLM based system
by: Gupta, Shailja, et al.
Published: (2024)
by: Gupta, Shailja, et al.
Published: (2024)
Large Language Models in the Abuse Detection Pipeline
by: Kath, Suraj, et al.
Published: (2026)
by: Kath, Suraj, et al.
Published: (2026)
CPsyCoun: A Report-based Multi-turn Dialogue Reconstruction and Evaluation Framework for Chinese Psychological Counseling
by: Zhang, Chenhao, et al.
Published: (2024)
by: Zhang, Chenhao, et al.
Published: (2024)
The Factuality Tax of Diversity-Intervened Text-to-Image Generation: Benchmark and Fact-Augmented Intervention
by: Wan, Yixin, et al.
Published: (2024)
by: Wan, Yixin, et al.
Published: (2024)
Locating Risk: Task Designers and the Challenge of Risk Disclosure in RAI Content Work
by: Qian, Alice, et al.
Published: (2025)
by: Qian, Alice, et al.
Published: (2025)
SocialNLP Fake-EmoReact 2021 Challenge Overview: Predicting Fake Tweets from Their Replies and GIFs
by: Huang, Chien-Kun, et al.
Published: (2024)
by: Huang, Chien-Kun, et al.
Published: (2024)
The Staircase of Ethics: Probing LLM Value Priorities through Multi-Step Induction to Complex Moral Dilemmas
by: Wu, Ya, et al.
Published: (2025)
by: Wu, Ya, et al.
Published: (2025)
Worker Discretion Advised: Co-designing Risk Disclosure in Crowdsourced Responsible AI (RAI) Content Work
by: Qian, Alice, et al.
Published: (2025)
by: Qian, Alice, et al.
Published: (2025)
Not Just Novelty: A Longitudinal Study on Utility and Customization of an AI Workflow
by: Long, Tao, et al.
Published: (2024)
by: Long, Tao, et al.
Published: (2024)
SALAD: Smart AI Language Assistant Daily
by: Nihal, Ragib Amin, et al.
Published: (2024)
by: Nihal, Ragib Amin, et al.
Published: (2024)
Similar Items
-
SWAN: Semantic Watermarking with Abstract Meaning Representation
by: Ye, Ziping, et al.
Published: (2026) -
AI Ethics by Design: Implementing Customizable Guardrails for Responsible AI Development
by: Šekrst, Kristina, et al.
Published: (2024) -
CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety
by: An, Heajun, et al.
Published: (2026) -
News Source Citing Patterns in AI Search Systems
by: Yang, Kai-Cheng
Published: (2025) -
White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMs
by: Wan, Yixin, et al.
Published: (2024)