Can We Trust a Black-box LLM? LLM Untrustworthy Boundary Detection via Bias-Diffusion and Multi-Agent Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Xiaotian, Tang, Di, Wang, Xiaofeng, Liu, Xiaozhong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can We Trust LLM Detectors?
by: Sandhan, Jivnesh, et al.
Published: (2026)
by: Sandhan, Jivnesh, et al.
Published: (2026)
LLM-REVal: Can We Trust LLM Reviewers Yet?
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
Knowledge-Infused Legal Wisdom: Navigating LLM Consultation through the Lens of Diagnostics and Positive-Unlabeled Reinforcement Learning
by: Wu, Yang, et al.
Published: (2024)
by: Wu, Yang, et al.
Published: (2024)
CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation
by: Sinha, Aarush, et al.
Published: (2026)
by: Sinha, Aarush, et al.
Published: (2026)
Nautilus Compass: Black-box Persona Drift Detection for Production LLM Agents
by: Wang, Chunxiao
Published: (2026)
by: Wang, Chunxiao
Published: (2026)
Crabs: Consuming Resource via Auto-generation for LLM-DoS Attack under Black-box Settings
by: Zhang, Yuanhe, et al.
Published: (2024)
by: Zhang, Yuanhe, et al.
Published: (2024)
ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering
by: Liu, Zexi, et al.
Published: (2025)
by: Liu, Zexi, et al.
Published: (2025)
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?
by: Wang, Leyao, et al.
Published: (2026)
by: Wang, Leyao, et al.
Published: (2026)
Epistemic Context Learning: Building Trust the Right Way in LLM-Based Multi-Agent Systems
by: Zhou, Ruiwen, et al.
Published: (2026)
by: Zhou, Ruiwen, et al.
Published: (2026)
RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Towards LLM-guided Causal Explainability for Black-box Text Classifiers
by: Bhattacharjee, Amrita, et al.
Published: (2023)
by: Bhattacharjee, Amrita, et al.
Published: (2023)
WorkForceAgent-R1: Incentivizing Reasoning Capability in LLM-based Web Agents via Reinforcement Learning
by: Zhuang, Yuchen, et al.
Published: (2025)
by: Zhuang, Yuchen, et al.
Published: (2025)
TrustAgent: Towards Safe and Trustworthy LLM-based Agents
by: Hua, Wenyue, et al.
Published: (2024)
by: Hua, Wenyue, et al.
Published: (2024)
Reinforce LLM Reasoning through Multi-Agent Reflection
by: Yuan, Yurun, et al.
Published: (2025)
by: Yuan, Yurun, et al.
Published: (2025)
LLM-Generated Black-box Explanations Can Be Adversarially Helpful
by: Ajwani, Rohan, et al.
Published: (2024)
by: Ajwani, Rohan, et al.
Published: (2024)
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
by: Ning, Yansong, et al.
Published: (2025)
by: Ning, Yansong, et al.
Published: (2025)
ERABAL: Enhancing Role-Playing Agents through Boundary-Aware Learning
by: Tang, Yihong, et al.
Published: (2024)
by: Tang, Yihong, et al.
Published: (2024)
Black-box Prompt Tuning with Subspace Learning
by: Zheng, Yuanhang, et al.
Published: (2023)
by: Zheng, Yuanhang, et al.
Published: (2023)
Detecting Prefix Bias in LLM-based Reward Models
by: Kumar, Ashwin, et al.
Published: (2025)
by: Kumar, Ashwin, et al.
Published: (2025)
AHP-Powered LLM Reasoning for Multi-Criteria Evaluation of Open-Ended Responses
by: Lu, Xiaotian, et al.
Published: (2024)
by: Lu, Xiaotian, et al.
Published: (2024)
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
by: Dong, Guanting, et al.
Published: (2025)
by: Dong, Guanting, et al.
Published: (2025)
LLM-based Multi-Agent Reinforcement Learning: Current and Future Directions
by: Sun, Chuanneng, et al.
Published: (2024)
by: Sun, Chuanneng, et al.
Published: (2024)
Reading Between the Lines: Towards Reliable Black-box LLM Fingerprinting via Zeroth-order Gradient Estimation
by: Shao, Shuo, et al.
Published: (2025)
by: Shao, Shuo, et al.
Published: (2025)
Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
by: Zhang, Zhiwei, et al.
Published: (2025)
by: Zhang, Zhiwei, et al.
Published: (2025)
RealVul: Can We Detect Vulnerabilities in Web Applications with LLM?
by: Cao, Di, et al.
Published: (2024)
by: Cao, Di, et al.
Published: (2024)
Dynamic Generation of Multi-LLM Agents Communication Topologies with Graph Diffusion Models
by: Jiang, Eric Hanchen, et al.
Published: (2025)
by: Jiang, Eric Hanchen, et al.
Published: (2025)
LLM-MRD: LLM-Guided Multi-View Reasoning Distillation for Fake News Detection
by: Zhou, Weilin, et al.
Published: (2026)
by: Zhou, Weilin, et al.
Published: (2026)
Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning
by: Wu, Yuhang, et al.
Published: (2026)
by: Wu, Yuhang, et al.
Published: (2026)
Persuasiveness and Bias in LLM: Investigating the Impact of Persuasiveness and Reinforcement of Bias in Language Models
by: Roy, Saumya
Published: (2025)
by: Roy, Saumya
Published: (2025)
Can We Still Hear the Accent? Investigating the Resilience of Native Language Signals in the LLM Era
by: Utami, Nabelanita, et al.
Published: (2026)
by: Utami, Nabelanita, et al.
Published: (2026)
Auto-Drafting Police Reports from Noisy ASR Outputs: A Trust-Centered LLM Approach
by: Kulkarni, Param, et al.
Published: (2025)
by: Kulkarni, Param, et al.
Published: (2025)
In Agents We Trust, but Who Do Agents Trust? Latent Source Preferences Steer LLM Generations
by: Khan, Mohammad Aflah, et al.
Published: (2026)
by: Khan, Mohammad Aflah, et al.
Published: (2026)
AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
Instruction Learning Paradigms: A Dual Perspective on White-box and Black-box LLMs
by: Ren, Yanwei, et al.
Published: (2025)
by: Ren, Yanwei, et al.
Published: (2025)
Citations and Trust in LLM Generated Responses
by: Ding, Yifan, et al.
Published: (2025)
by: Ding, Yifan, et al.
Published: (2025)
Offline Reinforcement Learning for LLM Multi-Step Reasoning
by: Wang, Huaijie, et al.
Published: (2024)
by: Wang, Huaijie, et al.
Published: (2024)
Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?
by: Prandi, Matteo, et al.
Published: (2025)
by: Prandi, Matteo, et al.
Published: (2025)
DiscourseFlip: An Oblique Discourse-Level Opinion Manipulation Attack against Black-box Retrieval-Augmented Generation
by: Gong, Yuyang, et al.
Published: (2026)
by: Gong, Yuyang, et al.
Published: (2026)
Zodiac: A Cardiologist-Level LLM Framework for Multi-Agent Diagnostics
by: Zhou, Yuan, et al.
Published: (2024)
by: Zhou, Yuan, et al.
Published: (2024)
Trust, Lies, and Long Memories: Emergent Social Dynamics and Reputation in Multi-Round Avalon with LLM Agents
by: Ellawela, Suveen
Published: (2026)
by: Ellawela, Suveen
Published: (2026)
Similar Items
-
Can We Trust LLM Detectors?
by: Sandhan, Jivnesh, et al.
Published: (2026) -
LLM-REVal: Can We Trust LLM Reviewers Yet?
by: Li, Rui, et al.
Published: (2025) -
Knowledge-Infused Legal Wisdom: Navigating LLM Consultation through the Lens of Diagnostics and Positive-Unlabeled Reinforcement Learning
by: Wu, Yang, et al.
Published: (2024) -
CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation
by: Sinha, Aarush, et al.
Published: (2026) -
Nautilus Compass: Black-box Persona Drift Detection for Production LLM Agents
by: Wang, Chunxiao
Published: (2026)