The Why Behind the Action: Unveiling Internal Drivers via Agentic Attribution
Fuente:
arXiv
Saved in:
| Main Authors: | Qian, Chen, Wang, Peng, Liu, Dongrui, Yang, Junyao, Guo, Dadi, Tang, Ling, Mei, Jilin, Ren, Qihan, Shao, Shuai, Liu, Yong, Fu, Jie, Shao, Jing, Hu, Xia |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Attributing Emergence in Million-Agent Systems
by: Tang, Ling, et al.
Published: (2026)
by: Tang, Ling, et al.
Published: (2026)
Interpreting Emergent Extreme Events in Multi-Agent Systems
by: Tang, Ling, et al.
Published: (2026)
by: Tang, Ling, et al.
Published: (2026)
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
by: Ren, Qihan, et al.
Published: (2026)
by: Ren, Qihan, et al.
Published: (2026)
Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
by: Shao, Shuai, et al.
Published: (2025)
by: Shao, Shuai, et al.
Published: (2025)
Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm
by: Guo, Dadi, et al.
Published: (2025)
by: Guo, Dadi, et al.
Published: (2025)
ReasonAny: Incorporating Reasoning Capability to Any Model via Simple and Effective Model Merging
by: Yang, Junyao, et al.
Published: (2026)
by: Yang, Junyao, et al.
Published: (2026)
What Do EEG Foundation Models Capture from Human Brain Signals?
by: Tang, Ling, et al.
Published: (2026)
by: Tang, Ling, et al.
Published: (2026)
The Tug of War Within: Mitigating the Fairness-Privacy Conflicts in Large Language Models
by: Qian, Chen, et al.
Published: (2024)
by: Qian, Chen, et al.
Published: (2024)
RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
by: Yang, Jingyi, et al.
Published: (2025)
by: Yang, Jingyi, et al.
Published: (2025)
TradeTrap: Are LLM-based Trading Agents Truly Reliable and Faithful?
by: Yan, Lewen, et al.
Published: (2025)
by: Yan, Lewen, et al.
Published: (2025)
Entropy-Gradient Inversion: Moving Toward Internal Mechanism of Large Reasoning Models
by: Yang, Junyao, et al.
Published: (2026)
by: Yang, Junyao, et al.
Published: (2026)
ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
VLSBench: Unveiling Visual Leakage in Multimodal Safety
by: Hu, Xuhao, et al.
Published: (2024)
by: Hu, Xuhao, et al.
Published: (2024)
Are Your Agents Upward Deceivers?
by: Guo, Dadi, et al.
Published: (2025)
by: Guo, Dadi, et al.
Published: (2025)
Independent characterization of the elastic and the mixing parts of hydrogel osmotic pressure
by: Shao, Zefan, et al.
Published: (2023)
by: Shao, Zefan, et al.
Published: (2023)
Conditional Advantage Estimation for Reinforcement Learning in Large Reasoning Models
by: Chen, Guanxu, et al.
Published: (2025)
by: Chen, Guanxu, et al.
Published: (2025)
REEF: Representation Encoding Fingerprints for Large Language Models
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
The Better Angels of Machine Personality: How Personality Relates to LLM Safety
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
Loop as a Bridge: Can Looped Transformers Truly Link Representation Space and Natural Language Outputs?
by: Chen, Guanxu, et al.
Published: (2026)
by: Chen, Guanxu, et al.
Published: (2026)
Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning
by: Qian, Chen, et al.
Published: (2025)
by: Qian, Chen, et al.
Published: (2025)
Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?
by: Guo, Dadi, et al.
Published: (2026)
by: Guo, Dadi, et al.
Published: (2026)
Towards Tracing Trustworthiness Dynamics: Revisiting Pre-training Period of Large Language Models
by: Qian, Chen, et al.
Published: (2024)
by: Qian, Chen, et al.
Published: (2024)
COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation
by: Zhou, Tianyi, et al.
Published: (2026)
by: Zhou, Tianyi, et al.
Published: (2026)
LED-Merging: Mitigating Safety-Utility Conflicts in Model Merging with Location-Election-Disjoint
by: Ma, Qianli, et al.
Published: (2025)
by: Ma, Qianli, et al.
Published: (2025)
Tree-frog-inspired osmocapillary adhesive bonding to diverse substrates
by: Shao, Zefan, et al.
Published: (2025)
by: Shao, Zefan, et al.
Published: (2025)
PRISM: Preference-Aware Influence Function Based Data Selection Method for Efficient Fine-Tuning
by: Lin, Qihao, et al.
Published: (2026)
by: Lin, Qihao, et al.
Published: (2026)
RvB: Automating AI System Hardening via Iterative Red-Blue Games
by: Huang, Lige, et al.
Published: (2026)
by: Huang, Lige, et al.
Published: (2026)
PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
by: Liu, Zicheng, et al.
Published: (2025)
by: Liu, Zicheng, et al.
Published: (2025)
AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
by: Liu, Dongrui, et al.
Published: (2026)
by: Liu, Dongrui, et al.
Published: (2026)
Towards the Dynamics of a DNN Learning Symbolic Interactions
by: Ren, Qihan, et al.
Published: (2024)
by: Ren, Qihan, et al.
Published: (2024)
X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Multi-Turn Jailbreaks without Compromising Usability
by: Lu, Xiaoya, et al.
Published: (2025)
by: Lu, Xiaoya, et al.
Published: (2025)
INFA-Guard: Mitigating Malicious Propagation via Infection-Aware Safeguarding in LLM-Based Multi-Agent Systems
by: Zhou, Yijin, et al.
Published: (2026)
by: Zhou, Yijin, et al.
Published: (2026)
Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks
by: Wang, Yu, et al.
Published: (2026)
by: Wang, Yu, et al.
Published: (2026)
Evaluating the Correctness of Inference Patterns Used by LLMs for Judgment
by: Chen, Lu, et al.
Published: (2024)
by: Chen, Lu, et al.
Published: (2024)
Beyond AME: A Novel Connection between Quantum Secret Sharing Schemes and $k$-Uniform States
by: Liu, Xuhong, et al.
Published: (2025)
by: Liu, Xuhong, et al.
Published: (2025)
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring
by: Chen, Guanxu, et al.
Published: (2025)
by: Chen, Guanxu, et al.
Published: (2025)
The LLM Already Knows: Estimating LLM-Perceived Question Difficulty via Hidden Representations
by: Zhu, Yubo, et al.
Published: (2025)
by: Zhu, Yubo, et al.
Published: (2025)
LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
by: Hu, Xuhao, et al.
Published: (2025)
by: Hu, Xuhao, et al.
Published: (2025)
Rethinking Entropy Regularization in Large Reasoning Models
by: Jiang, Yuxian, et al.
Published: (2025)
by: Jiang, Yuxian, et al.
Published: (2025)
PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
by: Zhang, Mingxuan, et al.
Published: (2026)
by: Zhang, Mingxuan, et al.
Published: (2026)
Similar Items
-
Attributing Emergence in Million-Agent Systems
by: Tang, Ling, et al.
Published: (2026) -
Interpreting Emergent Extreme Events in Multi-Agent Systems
by: Tang, Ling, et al.
Published: (2026) -
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
by: Ren, Qihan, et al.
Published: (2026) -
Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
by: Shao, Shuai, et al.
Published: (2025) -
Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm
by: Guo, Dadi, et al.
Published: (2025)