Auditing Agent Harness Safety
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Liu, Chengzhi, Guo, Yichen, Liu, Yepeng, Yang, Yuzhe, Yan, Qianqi, Zhao, Xuandong, Hua, Wenyue, Liu, Sheng, Li, Sharon, Bu, Yuheng, Wang, Xin Eric |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Dataset Protection via Watermarked Canaries in Retrieval-Augmented LLMs
par: Liu, Yepeng, et autres
Publié: (2025)
par: Liu, Yepeng, et autres
Publié: (2025)
In-Context Watermarks for Large Language Models
par: Liu, Yepeng, et autres
Publié: (2025)
par: Liu, Yepeng, et autres
Publié: (2025)
Adaptive Text Watermark for Large Language Models
par: Liu, Yepeng, et autres
Publié: (2024)
par: Liu, Yepeng, et autres
Publié: (2024)
Position: LLM Watermarking Should Align Stakeholders' Incentives for Practical Adoption
par: Liu, Yepeng, et autres
Publié: (2025)
par: Liu, Yepeng, et autres
Publié: (2025)
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
par: Liu, Chengzhi, et autres
Publié: (2026)
par: Liu, Chengzhi, et autres
Publié: (2026)
Multimodal Situational Safety
par: Zhou, Kaiwen, et autres
Publié: (2024)
par: Zhou, Kaiwen, et autres
Publié: (2024)
The Hidden Risks of Large Reasoning Models: A Safety Assessment of R1
par: Zhou, Kaiwen, et autres
Publié: (2025)
par: Zhou, Kaiwen, et autres
Publié: (2025)
Beyond the Safety Bundle: Auditing the Helpful and Harmless Dataset
par: Chehbouni, Khaoula, et autres
Publié: (2024)
par: Chehbouni, Khaoula, et autres
Publié: (2024)
Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space
par: Liu, Chengzhi, et autres
Publié: (2025)
par: Liu, Chengzhi, et autres
Publié: (2025)
War and Peace (WarAgent): Large Language Model-based Multi-Agent Simulation of World Wars
par: Hua, Wenyue, et autres
Publié: (2023)
par: Hua, Wenyue, et autres
Publié: (2023)
The Silicon Ceiling: Auditing GPT's Race and Gender Biases in Hiring
par: Armstrong, Lena, et autres
Publié: (2024)
par: Armstrong, Lena, et autres
Publié: (2024)
Defending LLM Watermarking Against Spoofing Attacks with Contrastive Representation Learning
par: An, Li, et autres
Publié: (2025)
par: An, Li, et autres
Publié: (2025)
Exploring Novelty Differences between Industry and Academia: A Knowledge Entity-centric Perspective
par: Zhao, Hongye, et autres
Publié: (2026)
par: Zhao, Hongye, et autres
Publié: (2026)
A Multimodal, Multilingual, and Multidimensional Pipeline for Fine-grained Crowdsourcing Earthquake Damage Evaluation
par: Ma, Zihui, et autres
Publié: (2025)
par: Ma, Zihui, et autres
Publié: (2025)
The Better Angels of Machine Personality: How Personality Relates to LLM Safety
par: Zhang, Jie, et autres
Publié: (2024)
par: Zhang, Jie, et autres
Publié: (2024)
Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling
par: Zhang, Zhen, et autres
Publié: (2026)
par: Zhang, Zhen, et autres
Publié: (2026)
Assessing Judging Bias in Large Reasoning Models: An Empirical Study
par: Wang, Qian, et autres
Publié: (2025)
par: Wang, Qian, et autres
Publié: (2025)
Analyzing the Safety of Japanese Large Language Models in Stereotype-Triggering Prompts
par: Nakanishi, Akito, et autres
Publié: (2025)
par: Nakanishi, Akito, et autres
Publié: (2025)
From Attribution to Abstention: Training-Free Attention-Based Auditing for Clinical Summarization
par: Yan, Qianqi, et autres
Publié: (2026)
par: Yan, Qianqi, et autres
Publié: (2026)
Auditing Stance Asymmetry in Generative Explanations
par: Han, Jiarui
Publié: (2026)
par: Han, Jiarui
Publié: (2026)
Presenting a Paper is an Art: Self-Improvement Aesthetic Agents for Academic Presentations
par: Liu, Chengzhi, et autres
Publié: (2025)
par: Liu, Chengzhi, et autres
Publié: (2025)
AuditWen:An Open-Source Large Language Model for Audit
par: Huang, Jiajia, et autres
Publié: (2024)
par: Huang, Jiajia, et autres
Publié: (2024)
Exploring Gender Biases in Language Patterns of Human-Conversational Agent Conversations
par: Liu, Weizi
Publié: (2024)
par: Liu, Weizi
Publié: (2024)
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
par: Zhang, Chi, et autres
Publié: (2025)
par: Zhang, Chi, et autres
Publié: (2025)
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
par: Gomaa, Amr, et autres
Publié: (2025)
par: Gomaa, Amr, et autres
Publié: (2025)
EmpathyAgent: Can Embodied Agents Conduct Empathetic Actions?
par: Chen, Xinyan, et autres
Publié: (2025)
par: Chen, Xinyan, et autres
Publié: (2025)
REALM: A Dataset of Real-World LLM Use Cases
par: Cheng, Jingwen, et autres
Publié: (2025)
par: Cheng, Jingwen, et autres
Publié: (2025)
Persuasion Dynamics in LLMs: Investigating Robustness and Adaptability in Knowledge and Safety with DuET-PD
par: Tan, Bryan Chen Zhengyu, et autres
Publié: (2025)
par: Tan, Bryan Chen Zhengyu, et autres
Publié: (2025)
Safety Evaluation and Enhancement of DeepSeek Models in Chinese Contexts
par: Zhang, Wenjing, et autres
Publié: (2025)
par: Zhang, Wenjing, et autres
Publié: (2025)
Characterizing Selective Refusal Bias in Large Language Models
par: Khorramrouz, Adel, et autres
Publié: (2025)
par: Khorramrouz, Adel, et autres
Publié: (2025)
EcoLANG: Efficient and Effective Agent Communication Language Induction for Social Simulation
par: Mou, Xinyi, et autres
Publié: (2025)
par: Mou, Xinyi, et autres
Publié: (2025)
AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios
par: Mou, Xinyi, et autres
Publié: (2024)
par: Mou, Xinyi, et autres
Publié: (2024)
A Comparative Evaluation of Structural Topic Models and BERTopic for Short, Open-Ended Survey Responses
par: Jiang, Yan, et autres
Publié: (2026)
par: Jiang, Yan, et autres
Publié: (2026)
Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMs
par: Hartmann, David, et autres
Publié: (2026)
par: Hartmann, David, et autres
Publié: (2026)
The Staircase of Ethics: Probing LLM Value Priorities through Multi-Step Induction to Complex Moral Dilemmas
par: Wu, Ya, et autres
Publié: (2025)
par: Wu, Ya, et autres
Publié: (2025)
OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs
par: Yan, Qianqi, et autres
Publié: (2026)
par: Yan, Qianqi, et autres
Publié: (2026)
SOTOPIA-$Ω$: Dynamic Strategy Injection Learning and Social Instruction Following Evaluation for Social Agents
par: Zhang, Wenyuan, et autres
Publié: (2025)
par: Zhang, Wenyuan, et autres
Publié: (2025)
On the Suitability of LLM-Driven Agents for Dark Pattern Audits
par: Sun, Chen, et autres
Publié: (2026)
par: Sun, Chen, et autres
Publié: (2026)
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race
par: Sun, Lihao, et autres
Publié: (2025)
par: Sun, Lihao, et autres
Publié: (2025)
How Did We Get Here? Summarizing Conversation Dynamics
par: Hua, Yilun, et autres
Publié: (2024)
par: Hua, Yilun, et autres
Publié: (2024)
Documents similaires
-
Dataset Protection via Watermarked Canaries in Retrieval-Augmented LLMs
par: Liu, Yepeng, et autres
Publié: (2025) -
In-Context Watermarks for Large Language Models
par: Liu, Yepeng, et autres
Publié: (2025) -
Adaptive Text Watermark for Large Language Models
par: Liu, Yepeng, et autres
Publié: (2024) -
Position: LLM Watermarking Should Align Stakeholders' Incentives for Practical Adoption
par: Liu, Yepeng, et autres
Publié: (2025) -
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
par: Liu, Chengzhi, et autres
Publié: (2026)