Are Your Agents Upward Deceivers?
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Dadi, Liu, Qingyu, Liu, Dongrui, Ren, Qihan, Shao, Shuai, Qiu, Tianyi, Li, Haoran, Fung, Yi R., Ba, Zhongjie, Dai, Juntao, Ji, Jiaming, Chen, Zhikai, Tao, Jialing, Yang, Yaodong, Shao, Jing, Hu, Xia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?
by: Guo, Dadi, et al.
Published: (2026)
by: Guo, Dadi, et al.
Published: (2026)
Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm
by: Guo, Dadi, et al.
Published: (2025)
by: Guo, Dadi, et al.
Published: (2025)
Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
by: Shao, Shuai, et al.
Published: (2025)
by: Shao, Shuai, et al.
Published: (2025)
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
by: Ren, Qihan, et al.
Published: (2026)
by: Ren, Qihan, et al.
Published: (2026)
The Why Behind the Action: Unveiling Internal Drivers via Agentic Attribution
by: Qian, Chen, et al.
Published: (2026)
by: Qian, Chen, et al.
Published: (2026)
LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
by: Hu, Xuhao, et al.
Published: (2025)
by: Hu, Xuhao, et al.
Published: (2025)
RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
by: Yang, Jingyi, et al.
Published: (2025)
by: Yang, Jingyi, et al.
Published: (2025)
Tree-frog-inspired osmocapillary adhesive bonding to diverse substrates
by: Shao, Zefan, et al.
Published: (2025)
by: Shao, Zefan, et al.
Published: (2025)
Sequence to Sequence Reward Modeling: Improving RLHF by Language Feedback
by: Zhou, Jiayi, et al.
Published: (2024)
by: Zhou, Jiayi, et al.
Published: (2024)
Independent characterization of the elastic and the mixing parts of hydrogel osmotic pressure
by: Shao, Zefan, et al.
Published: (2023)
by: Shao, Zefan, et al.
Published: (2023)
R$^2$BD: A Reconstruction-Based Method for Generalizable and Efficient Detection of Fake Images
by: Liu, Qingyu, et al.
Published: (2026)
by: Liu, Qingyu, et al.
Published: (2026)
AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
by: Liu, Dongrui, et al.
Published: (2026)
by: Liu, Dongrui, et al.
Published: (2026)
COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation
by: Zhou, Tianyi, et al.
Published: (2026)
by: Zhou, Tianyi, et al.
Published: (2026)
SafeMCP: Proactive Power Regulation for LLM Agent Defense via Environment-Grounded Look-Ahead Reasoning
by: Wang, Lichao, et al.
Published: (2026)
by: Wang, Lichao, et al.
Published: (2026)
"Training robust watermarking model may hurt authentication!'' Exploring and Mitigating the Identity Leakage in Robust Watermarking
by: Zhang, Xinyu, et al.
Published: (2026)
by: Zhang, Xinyu, et al.
Published: (2026)
Language Models Resist Alignment: Evidence From Data Compression
by: Ji, Jiaming, et al.
Published: (2024)
by: Ji, Jiaming, et al.
Published: (2024)
Aligner: Efficient Alignment by Learning to Correct
by: Ji, Jiaming, et al.
Published: (2024)
by: Ji, Jiaming, et al.
Published: (2024)
Exposing the Deception: Uncovering More Forgery Clues for Deepfake Detection
by: Ba, Zhongjie, et al.
Published: (2024)
by: Ba, Zhongjie, et al.
Published: (2024)
VISA: Value Injection via Shielded Adaptation for Personalized LLM Alignment
by: Chen, Jiawei, et al.
Published: (2026)
by: Chen, Jiawei, et al.
Published: (2026)
A Game-Theoretic Negotiation Framework for Cross-Cultural Consensus in LLMs
by: Zhang, Guoxi, et al.
Published: (2025)
by: Zhang, Guoxi, et al.
Published: (2025)
Attack-Resistant Watermarking for AIGC Image Forensics via Diffusion-based Semantic Deflection
by: Liu, Qingyu, et al.
Published: (2026)
by: Liu, Qingyu, et al.
Published: (2026)
Attributing Emergence in Million-Agent Systems
by: Tang, Ling, et al.
Published: (2026)
by: Tang, Ling, et al.
Published: (2026)
ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety Alignment
by: Bu, Yuyan, et al.
Published: (2026)
by: Bu, Yuyan, et al.
Published: (2026)
Loop as a Bridge: Can Looped Transformers Truly Link Representation Space and Natural Language Outputs?
by: Chen, Guanxu, et al.
Published: (2026)
by: Chen, Guanxu, et al.
Published: (2026)
ProgressGym: Alignment with a Millennium of Moral Progress
by: Qiu, Tianyi, et al.
Published: (2024)
by: Qiu, Tianyi, et al.
Published: (2024)
Conditional Advantage Estimation for Reinforcement Learning in Large Reasoning Models
by: Chen, Guanxu, et al.
Published: (2025)
by: Chen, Guanxu, et al.
Published: (2025)
Towards the Dynamics of a DNN Learning Symbolic Interactions
by: Ren, Qihan, et al.
Published: (2024)
by: Ren, Qihan, et al.
Published: (2024)
Harnessing Frequency Spectrum Insights for Image Copyright Protection Against Diffusion Models
by: Liu, Zhenguang, et al.
Published: (2025)
by: Liu, Zhenguang, et al.
Published: (2025)
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation
by: Dai, Juntao, et al.
Published: (2024)
by: Dai, Juntao, et al.
Published: (2024)
ThinkPatterns-21k: A Systematic Study on the Impact of Thinking Patterns in LLMs
by: Wen, Pengcheng, et al.
Published: (2025)
by: Wen, Pengcheng, et al.
Published: (2025)
PRISM: Preference-Aware Influence Function Based Data Selection Method for Efficient Fine-Tuning
by: Lin, Qihao, et al.
Published: (2026)
by: Lin, Qihao, et al.
Published: (2026)
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
by: Ji, Jiaming, et al.
Published: (2024)
by: Ji, Jiaming, et al.
Published: (2024)
SafeDreamer: Safe Reinforcement Learning with World Models
by: Huang, Weidong, et al.
Published: (2023)
by: Huang, Weidong, et al.
Published: (2023)
Evaluating the Correctness of Inference Patterns Used by LLMs for Judgment
by: Chen, Lu, et al.
Published: (2024)
by: Chen, Lu, et al.
Published: (2024)
When Detectors Forget Forensics: Blocking Semantic Shortcuts for Generalizable AI-Generated Image Detection
by: Shuai, Chao, et al.
Published: (2026)
by: Shuai, Chao, et al.
Published: (2026)
Views Can Be Deceiving: Improved SSL Through Feature Space Augmentation
by: Hamidieh, Kimia, et al.
Published: (2024)
by: Hamidieh, Kimia, et al.
Published: (2024)
SafeMT: Multi-turn Safety for Multimodal Language Models
by: Zhu, Han, et al.
Published: (2025)
by: Zhu, Han, et al.
Published: (2025)
TradeTrap: Are LLM-based Trading Agents Truly Reliable and Faithful?
by: Yan, Lewen, et al.
Published: (2025)
by: Yan, Lewen, et al.
Published: (2025)
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
by: Zhang, Borong, et al.
Published: (2025)
by: Zhang, Borong, et al.
Published: (2025)
Similar Items
-
Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?
by: Guo, Dadi, et al.
Published: (2026) -
Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm
by: Guo, Dadi, et al.
Published: (2025) -
Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
by: Shao, Shuai, et al.
Published: (2025) -
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
by: Ren, Qihan, et al.
Published: (2026) -
The Why Behind the Action: Unveiling Internal Drivers via Agentic Attribution
by: Qian, Chen, et al.
Published: (2026)