Teaching LLMs to Learn Tool Trialing and Execution through Environment Interaction
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Gao, Xingjie, Huang, Pengcheng, Liu, Zhenghao, Yan, Yukun, Wang, Shuo, Chen, Zulong, Qian, Chen, Yu, Ge, Gu, Yu |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
COAST: Enhancing the Code Debugging Ability of LLMs through Communicative Agent Based Data Synthesis
par: Yang, Weiqing, et autres
Publié: (2024)
par: Yang, Weiqing, et autres
Publié: (2024)
An Empirical Study of Interaction Bugs in ROS-based Software
par: Chen, Zhixiang, et autres
Publié: (2025)
par: Chen, Zhixiang, et autres
Publié: (2025)
INTERVENOR: Prompting the Coding Ability of Large Language Models with the Interactive Chain of Repair
par: Wang, Hanbin, et autres
Publié: (2023)
par: Wang, Hanbin, et autres
Publié: (2023)
Sifting through the Chaff: On Utilizing Execution Feedback for Ranking the Generated Code Candidates
par: Sun, Zhihong, et autres
Publié: (2024)
par: Sun, Zhihong, et autres
Publié: (2024)
CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification
par: Tian, Yuchen, et autres
Publié: (2024)
par: Tian, Yuchen, et autres
Publié: (2024)
Simulating Complex Multi-Turn Tool Calling Interactions in Stateless Execution Environments
par: Crouse, Maxwell, et autres
Publié: (2026)
par: Crouse, Maxwell, et autres
Publié: (2026)
Teaching Code LLMs to Use Autocompletion Tools in Repository-Level Code Generation
par: Wang, Chong, et autres
Publié: (2024)
par: Wang, Chong, et autres
Publié: (2024)
Integrating Symbolic Execution with LLMs for Automated Generation of Program Specifications
par: Yang, Fanpeng, et autres
Publié: (2025)
par: Yang, Fanpeng, et autres
Publié: (2025)
Can Large Language Models Solve Path Constraints in Symbolic Execution?
par: Wang, Wenhan, et autres
Publié: (2025)
par: Wang, Wenhan, et autres
Publié: (2025)
Hyperion: Unveiling DApp Inconsistencies using LLM and Dataflow-Guided Symbolic Execution
par: Yang, Shuo, et autres
Publié: (2024)
par: Yang, Shuo, et autres
Publié: (2024)
Logging Like Humans for LLMs: Rethinking Logging via Execution and Runtime Feedback
par: Wang, Xin, et autres
Publié: (2026)
par: Wang, Xin, et autres
Publié: (2026)
SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs
par: Xia, Hongfei, et autres
Publié: (2025)
par: Xia, Hongfei, et autres
Publié: (2025)
SelfPiCo: Self-Guided Partial Code Execution with LLMs
par: Xue, Zhipeng, et autres
Publié: (2024)
par: Xue, Zhipeng, et autres
Publié: (2024)
Compiling Code LLMs into Lightweight Executables
par: Shi, Jieke, et autres
Publié: (2026)
par: Shi, Jieke, et autres
Publié: (2026)
CodeScore: Evaluating Code Generation by Learning Code Execution
par: Dong, Yihong, et autres
Publié: (2023)
par: Dong, Yihong, et autres
Publié: (2023)
NumScout: Unveiling Numerical Defects in Smart Contracts using LLM-Pruning Symbolic Execution
par: Chen, Jiachi, et autres
Publié: (2025)
par: Chen, Jiachi, et autres
Publié: (2025)
Mokav: Execution-driven Differential Testing with LLMs
par: Etemadi, Khashayar, et autres
Publié: (2024)
par: Etemadi, Khashayar, et autres
Publié: (2024)
An Executable Benchmarking Suite for Tool-Using Agents
par: Zhong, Zhiqing, et autres
Publié: (2026)
par: Zhong, Zhiqing, et autres
Publié: (2026)
Teaching LLMs Program Semantics via Symbolic Execution Traces
par: Bayer, Jonas, et autres
Publié: (2026)
par: Bayer, Jonas, et autres
Publié: (2026)
VulInstruct: Teaching LLMs Root-Cause Reasoning for Vulnerability Detection via Security Specifications
par: Zhu, Hao, et autres
Publié: (2025)
par: Zhu, Hao, et autres
Publié: (2025)
A Hybrid Execution Environment for Computer-Interpretable Guidelines in PROforma
par: Kogan, Alexandra, et autres
Publié: (2024)
par: Kogan, Alexandra, et autres
Publié: (2024)
Python Symbolic Execution with LLM-powered Code Generation
par: Wang, Wenhan, et autres
Publié: (2024)
par: Wang, Wenhan, et autres
Publié: (2024)
Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages
par: Wu, Fan, et autres
Publié: (2026)
par: Wu, Fan, et autres
Publié: (2026)
A Tool for In-depth Analysis of Code Execution Reasoning of Large Language Models
par: Liu, Changshu, et autres
Publié: (2025)
par: Liu, Changshu, et autres
Publié: (2025)
Learning Adaptive Parallel Execution for Efficient Code Localization
par: Xu, Ke, et autres
Publié: (2026)
par: Xu, Ke, et autres
Publié: (2026)
The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment
par: Yu, Shasha, et autres
Publié: (2026)
par: Yu, Shasha, et autres
Publié: (2026)
Seismology modeling agent: A smart assistant for geophysical researchers
par: Ren, Yukun, et autres
Publié: (2025)
par: Ren, Yukun, et autres
Publié: (2025)
ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution
par: Wang, Yubang, et autres
Publié: (2026)
par: Wang, Yubang, et autres
Publié: (2026)
Disproving Program Equivalence with LLMs
par: Allamanis, Miltiadis, et autres
Publié: (2025)
par: Allamanis, Miltiadis, et autres
Publié: (2025)
Leveraging LLMs for Dynamic IoT Systems Generation through Mixed-Initiative Interaction
par: Adnan, Bassam, et autres
Publié: (2025)
par: Adnan, Bassam, et autres
Publié: (2025)
SynthCoder: A Synthetical Strategy to Tune LLMs for Code Completion
par: Yu, Dongjun, et autres
Publié: (2025)
par: Yu, Dongjun, et autres
Publié: (2025)
ExeCoder: Empowering Large Language Models with Executability Representation for Code Translation
par: He, Minghua, et autres
Publié: (2025)
par: He, Minghua, et autres
Publié: (2025)
Can LLMs Recover Program Semantics? A Systematic Evaluation with Symbolic Execution
par: Feng, Rong, et autres
Publié: (2025)
par: Feng, Rong, et autres
Publié: (2025)
Development of a Evaluation Tool for Age-Appropriate Software in Aging Environments: A Delphi Study
par: Bai, Zhenggang, et autres
Publié: (2024)
par: Bai, Zhenggang, et autres
Publié: (2024)
Identifying Process Improvement Opportunities through Process Execution Benchmarking
par: Abb, Luka, et autres
Publié: (2025)
par: Abb, Luka, et autres
Publié: (2025)
Execution-free Program Repair
par: Huang, Li, et autres
Publié: (2024)
par: Huang, Li, et autres
Publié: (2024)
NExT: Teaching Large Language Models to Reason about Code Execution
par: Ni, Ansong, et autres
Publié: (2024)
par: Ni, Ansong, et autres
Publié: (2024)
Analyzing C/C++ Library Migrations at the Package-level: Prevalence, Domains, Targets and Rationals across Seven Package Management Tools
par: Gu, Haiqiao, et autres
Publié: (2025)
par: Gu, Haiqiao, et autres
Publié: (2025)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
par: Chen, Zaoyu, et autres
Publié: (2026)
par: Chen, Zaoyu, et autres
Publié: (2026)
DOne: Decoupling Structure and Rendering for High-Fidelity Design-to-Code Generation
par: Huang, Xinhao, et autres
Publié: (2026)
par: Huang, Xinhao, et autres
Publié: (2026)
Documents similaires
-
COAST: Enhancing the Code Debugging Ability of LLMs through Communicative Agent Based Data Synthesis
par: Yang, Weiqing, et autres
Publié: (2024) -
An Empirical Study of Interaction Bugs in ROS-based Software
par: Chen, Zhixiang, et autres
Publié: (2025) -
INTERVENOR: Prompting the Coding Ability of Large Language Models with the Interactive Chain of Repair
par: Wang, Hanbin, et autres
Publié: (2023) -
Sifting through the Chaff: On Utilizing Execution Feedback for Ranking the Generated Code Candidates
par: Sun, Zhihong, et autres
Publié: (2024) -
CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification
par: Tian, Yuchen, et autres
Publié: (2024)