The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Shasha, Carroll, Fiona, Bentley, Barry L. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Causal Impact of Tool Affordance on Safety Alignment in LLM Agents
by: Yu, Shasha, et al.
Published: (2026)
by: Yu, Shasha, et al.
Published: (2026)
An Executable Benchmarking Suite for Tool-Using Agents
by: Zhong, Zhiqing, et al.
Published: (2026)
by: Zhong, Zhiqing, et al.
Published: (2026)
Deploy-Master: Automating the Deployment of 50,000+ Agent-Ready Scientific Tools in One Day
by: Wang, Yi, et al.
Published: (2026)
by: Wang, Yi, et al.
Published: (2026)
Rethinking the Role of Entropy in Optimizing Tool-Use Behaviors for Large Language Model Agents
by: Li, Zeping, et al.
Published: (2026)
by: Li, Zeping, et al.
Published: (2026)
Learning Correct Behavior from Examples: Validating Sequential Execution in Autonomous Agents
by: Sharma, Reshabh K, et al.
Published: (2026)
by: Sharma, Reshabh K, et al.
Published: (2026)
A Tool for Generating Exceptional Behavior Tests With Large Language Models
by: Zhong, Linghan, et al.
Published: (2025)
by: Zhong, Linghan, et al.
Published: (2025)
Repairing Tool Calls Using Post-tool Execution Reflection and RAG
by: Tsay, Jason, et al.
Published: (2025)
by: Tsay, Jason, et al.
Published: (2025)
Teaching LLMs to Learn Tool Trialing and Execution through Environment Interaction
by: Gao, Xingjie, et al.
Published: (2026)
by: Gao, Xingjie, et al.
Published: (2026)
Automating Android Build Repair: Bridging the Reasoning-Execution Gap in LLM Agents with Domain-Specific Tools
by: Son, Ha Min, et al.
Published: (2025)
by: Son, Ha Min, et al.
Published: (2025)
Class-Level Code Generation from Natural Language Using Iterative, Tool-Enhanced Reasoning over Repository
by: Deshpande, Ajinkya, et al.
Published: (2024)
by: Deshpande, Ajinkya, et al.
Published: (2024)
Constraint-Guided Multi-Agent Decompilation for Executable Binary Recovery
by: Zhang, Yifan, et al.
Published: (2026)
by: Zhang, Yifan, et al.
Published: (2026)
ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents
by: Li, Dawei, et al.
Published: (2026)
by: Li, Dawei, et al.
Published: (2026)
SHERPA: A Model-Driven Framework for Large Language Model Execution
by: Chen, Boqi, et al.
Published: (2025)
by: Chen, Boqi, et al.
Published: (2025)
ToolFuzz -- Automated Agent Tool Testing
by: Milev, Ivan, et al.
Published: (2025)
by: Milev, Ivan, et al.
Published: (2025)
Z-Space: A Multi-Agent Tool Orchestration Framework for Enterprise-Grade LLM Automation
by: He, Qingsong, et al.
Published: (2025)
by: He, Qingsong, et al.
Published: (2025)
Toward Executable Repository-Level Code Generation via Environment Alignment
by: Pan, Ruwei, et al.
Published: (2026)
by: Pan, Ruwei, et al.
Published: (2026)
RA-Gen: A Controllable Code Generation Framework Using ReAct for Multi-Agent Task Execution
by: Liu, Aofan, et al.
Published: (2025)
by: Liu, Aofan, et al.
Published: (2025)
Rethinking AI Literacy Education in Higher Education: Bridging Risk Perception and Responsible Adoption
by: Yu, Shasha, et al.
Published: (2026)
by: Yu, Shasha, et al.
Published: (2026)
JTPRO: A Joint Tool-Prompt Reflective Optimization Framework for Language Agents
by: Ghoshal, Sandip, et al.
Published: (2026)
by: Ghoshal, Sandip, et al.
Published: (2026)
AgentForge: Execution-Grounded Multi-Agent LLM Framework for Autonomous Software Engineering
by: Kumar, Rajesh, et al.
Published: (2026)
by: Kumar, Rajesh, et al.
Published: (2026)
DynamicsLLM: a Dynamic Analysis-based Tool for Generating Intelligent Execution Traces Using LLMs to Detect Android Behavioural Code Smells
by: Cherief, Houcine Abdelkader, et al.
Published: (2026)
by: Cherief, Houcine Abdelkader, et al.
Published: (2026)
REDO: Execution-Free Runtime Error Detection for COding Agents
by: Li, Shou, et al.
Published: (2024)
by: Li, Shou, et al.
Published: (2024)
Open, Reliable, and Collective: A Community-Driven Framework for Tool-Using AI Agents
by: Dang, Hy, et al.
Published: (2026)
by: Dang, Hy, et al.
Published: (2026)
MathViz-E: A Case-study in Domain-Specialized Tool-Using Agents
by: Bulusu, Arya, et al.
Published: (2024)
by: Bulusu, Arya, et al.
Published: (2024)
RedCode: Risky Code Execution and Generation Benchmark for Code Agents
by: Guo, Chengquan, et al.
Published: (2024)
by: Guo, Chengquan, et al.
Published: (2024)
exLong: Generating Exceptional Behavior Tests with Large Language Models
by: Zhang, Jiyang, et al.
Published: (2024)
by: Zhang, Jiyang, et al.
Published: (2024)
DynaFix: Iterative Automated Program Repair Driven by Execution-Level Dynamic Information
by: Huang, Zhili, et al.
Published: (2025)
by: Huang, Zhili, et al.
Published: (2025)
ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution
by: Wang, Yubang, et al.
Published: (2026)
by: Wang, Yubang, et al.
Published: (2026)
Towards Structured, State-Aware, and Execution-Grounded Reasoning for Software Engineering Agents
by: Tse-Hsun, et al.
Published: (2026)
by: Tse-Hsun, et al.
Published: (2026)
ContextCov: Deriving and Enforcing Executable Constraints from Agent Instruction Files
by: Sharma, Reshabh K
Published: (2026)
by: Sharma, Reshabh K
Published: (2026)
You Name It, I Run It: An LLM Agent to Execute Tests of Arbitrary Projects
by: Bouzenia, Islem, et al.
Published: (2024)
by: Bouzenia, Islem, et al.
Published: (2024)
SOEN-101: Code Generation by Emulating Software Process Models Using Large Language Model Agents
by: Lin, Feng, et al.
Published: (2024)
by: Lin, Feng, et al.
Published: (2024)
Semantic Tool Discovery for Large Language Models: A Vector-Based Approach to MCP Tool Selection
by: Mudunuri, Sarat, et al.
Published: (2026)
by: Mudunuri, Sarat, et al.
Published: (2026)
An Execution-Verified Multi-Language Benchmark for Code Semantic Reasoning
by: Li, Yikun, et al.
Published: (2026)
by: Li, Yikun, et al.
Published: (2026)
TransAgent: Enhancing LLM-Based Code Translation via Fine-Grained Execution Alignment
by: Yuan, Zhiqiang, et al.
Published: (2024)
by: Yuan, Zhiqiang, et al.
Published: (2024)
Ambiguity Detection and Elimination in Automated Executable Process Modeling
by: Matei, Ion, et al.
Published: (2026)
by: Matei, Ion, et al.
Published: (2026)
AgentSLA : Towards a Service Level Agreement for AI Agents
by: Jouneaux, Gwendal, et al.
Published: (2025)
by: Jouneaux, Gwendal, et al.
Published: (2025)
ParaTool: Shifting Tool Representations from Context to Parameters
by: Yu, Zekai, et al.
Published: (2026)
by: Yu, Zekai, et al.
Published: (2026)
ASA: Training-Free Representation Engineering for Tool-Calling Agents
by: Wang, Youjin, et al.
Published: (2026)
by: Wang, Youjin, et al.
Published: (2026)
Squeez: Task-Conditioned Tool-Output Pruning for Coding Agents
by: Kovács, Ádám
Published: (2026)
by: Kovács, Ádám
Published: (2026)
Similar Items
-
The Causal Impact of Tool Affordance on Safety Alignment in LLM Agents
by: Yu, Shasha, et al.
Published: (2026) -
An Executable Benchmarking Suite for Tool-Using Agents
by: Zhong, Zhiqing, et al.
Published: (2026) -
Deploy-Master: Automating the Deployment of 50,000+ Agent-Ready Scientific Tools in One Day
by: Wang, Yi, et al.
Published: (2026) -
Rethinking the Role of Entropy in Optimizing Tool-Use Behaviors for Large Language Model Agents
by: Li, Zeping, et al.
Published: (2026) -
Learning Correct Behavior from Examples: Validating Sequential Execution in Autonomous Agents
by: Sharma, Reshabh K, et al.
Published: (2026)