Pluralistic Behavior Suite: Stress-Testing Multi-Turn Adherence to Custom Behavioral Policies
Fuente:
arXiv
Saved in:
| Main Authors: | Varshney, Prasoon, Sreedhar, Makesh Narsimhan, Jiang, Liwei, Rebedea, Traian, Parisien, Christopher |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Models
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2025)
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2025)
Unsupervised Extraction of Dialogue Policies from Conversations
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2024)
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2024)
Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails
by: Ghosh, Shaona, et al.
Published: (2025)
by: Ghosh, Shaona, et al.
Published: (2025)
CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2024)
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2024)
Defenses & Enablers For Skill Injection Attacks on Terminal Based Agents
by: Fujinuma, Yoshinari, et al.
Published: (2026)
by: Fujinuma, Yoshinari, et al.
Published: (2026)
Towards Inference-time Category-wise Safety Steering for Large Language Models
by: Bhattacharjee, Amrita, et al.
Published: (2024)
by: Bhattacharjee, Amrita, et al.
Published: (2024)
Semi-Supervised Learning for Large Language Models Safety and Content Moderation
by: Dinuta, Eduard Stefan, et al.
Published: (2025)
by: Dinuta, Eduard Stefan, et al.
Published: (2025)
HelpSteer2: Open-source dataset for training top-performing reward models
by: Wang, Zhilin, et al.
Published: (2024)
by: Wang, Zhilin, et al.
Published: (2024)
Improving Legal Judgement Prediction in Romanian with Long Text Encoders
by: Masala, Mihai, et al.
Published: (2024)
by: Masala, Mihai, et al.
Published: (2024)
AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
by: Ghosh, Shaona, et al.
Published: (2024)
by: Ghosh, Shaona, et al.
Published: (2024)
MultiMatch: Multihead Consistency Regularization Matching for Semi-Supervised Text Classification
by: Sirbu, Iustin, et al.
Published: (2025)
by: Sirbu, Iustin, et al.
Published: (2025)
Training a General Purpose Automated Red Teaming Model
by: Padmakumar, Aishwarya, et al.
Published: (2026)
by: Padmakumar, Aishwarya, et al.
Published: (2026)
Meta-learning how to Share Credit among Macro-Actions
by: Hosu, Ionel-Alexandru, et al.
Published: (2025)
by: Hosu, Ionel-Alexandru, et al.
Published: (2025)
Complexity-based code embeddings
by: Folea, Rares, et al.
Published: (2026)
by: Folea, Rares, et al.
Published: (2026)
Selective Deficits in LLM Mental Self-Modeling in a Behavior-Based Test of Theory of Mind
by: Ackerman, Christopher
Published: (2026)
by: Ackerman, Christopher
Published: (2026)
Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
by: Ding, Yifeng, et al.
Published: (2025)
by: Ding, Yifeng, et al.
Published: (2025)
PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier
by: Jiang, Yuhua, et al.
Published: (2025)
by: Jiang, Yuhua, et al.
Published: (2025)
A Roadmap to Pluralistic Alignment
by: Sorensen, Taylor, et al.
Published: (2024)
by: Sorensen, Taylor, et al.
Published: (2024)
Addressing LLM Diversity by Infusing Random Concepts
by: Agrawal, Pulin, et al.
Published: (2026)
by: Agrawal, Pulin, et al.
Published: (2026)
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
by: Liu, Mingjie, et al.
Published: (2025)
by: Liu, Mingjie, et al.
Published: (2025)
Interpreting and Controlling Model Behavior via Constitutions for Atomic Concept Edits
by: Kalibhat, Neha, et al.
Published: (2026)
by: Kalibhat, Neha, et al.
Published: (2026)
Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time
by: Zhang, Zhenyu, et al.
Published: (2025)
by: Zhang, Zhenyu, et al.
Published: (2025)
MixDPO: Modeling Preference Strength for Pluralistic Alignment
by: Imai, Saki, et al.
Published: (2026)
by: Imai, Saki, et al.
Published: (2026)
Pluralistic Alignment for Healthcare: A Role-Driven Framework
by: Zhong, Jiayou, et al.
Published: (2025)
by: Zhong, Jiayou, et al.
Published: (2025)
Learning to Reason from Feedback at Test-Time
by: Li, Yanyang, et al.
Published: (2025)
by: Li, Yanyang, et al.
Published: (2025)
VITAL: A New Dataset for Benchmarking Pluralistic Alignment in Healthcare
by: Shetty, Anudeex, et al.
Published: (2025)
by: Shetty, Anudeex, et al.
Published: (2025)
VISPA: Pluralistic Alignment via Automatic Value Selection and Activation
by: Zheng, Shenyan, et al.
Published: (2026)
by: Zheng, Shenyan, et al.
Published: (2026)
Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
by: Zheng, Mingqian, et al.
Published: (2026)
by: Zheng, Mingqian, et al.
Published: (2026)
Models Recall What They Violate: Constraint Adherence in Multi-Turn LLM Ideation
by: Kruthof, Garvin
Published: (2026)
by: Kruthof, Garvin
Published: (2026)
X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
by: Rahman, Salman, et al.
Published: (2025)
by: Rahman, Salman, et al.
Published: (2025)
Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agents
by: Wang, Guoqing, et al.
Published: (2025)
by: Wang, Guoqing, et al.
Published: (2025)
A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI
by: Karagoz, Atahan
Published: (2026)
by: Karagoz, Atahan
Published: (2026)
Efficient Model-Agnostic Multi-Group Equivariant Networks
by: Baltaji, Razan, et al.
Published: (2023)
by: Baltaji, Razan, et al.
Published: (2023)
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
by: Sorensen, Taylor, et al.
Published: (2023)
by: Sorensen, Taylor, et al.
Published: (2023)
Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
by: Huang, Jing, et al.
Published: (2025)
by: Huang, Jing, et al.
Published: (2025)
BARRED: Synthetic Training of Custom Policy Guardrails via Asymmetric Debate
by: Mazza, Arnon, et al.
Published: (2026)
by: Mazza, Arnon, et al.
Published: (2026)
ToolACE-MT: Non-Autoregressive Generation for Agentic Multi-Turn Interaction
by: Zeng, Xingshan, et al.
Published: (2025)
by: Zeng, Xingshan, et al.
Published: (2025)
Inner Speech as Behavior Guides: Steerable Imitation of Diverse Behaviors for Human-AI coordination
by: Trivedi, Rakshit, et al.
Published: (2026)
by: Trivedi, Rakshit, et al.
Published: (2026)
Aligning Machiavellian Agents: Behavior Steering via Test-Time Policy Shaping
by: Mujtaba, Dena, et al.
Published: (2025)
by: Mujtaba, Dena, et al.
Published: (2025)
Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings
by: Wu, Yuning, et al.
Published: (2026)
by: Wu, Yuning, et al.
Published: (2026)
Similar Items
-
Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Models
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2025) -
Unsupervised Extraction of Dialogue Policies from Conversations
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2024) -
Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails
by: Ghosh, Shaona, et al.
Published: (2025) -
CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2024) -
Defenses & Enablers For Skill Injection Attacks on Terminal Based Agents
by: Fujinuma, Yoshinari, et al.
Published: (2026)