Simulating Complex Multi-Turn Tool Calling Interactions in Stateless Execution Environments
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Crouse, Maxwell, Abdelaziz, Ibrahim, Fadnis, Kshitij, Patel, Siva Sankalp, Basu, Kinjal, Gunasekara, Chulaka, Kumaravel, Sadhana, Munawar, Asim, Kapanipathi, Pavan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMs
von: Basu, Kinjal, et al.
Veröffentlicht: (2024)
von: Basu, Kinjal, et al.
Veröffentlicht: (2024)
NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls
von: Basu, Kinjal, et al.
Veröffentlicht: (2024)
von: Basu, Kinjal, et al.
Veröffentlicht: (2024)
Formally Specifying the High-Level Behavior of LLM-Based Agents
von: Crouse, Maxwell, et al.
Veröffentlicht: (2023)
von: Crouse, Maxwell, et al.
Veröffentlicht: (2023)
ToolRM: Outcome Reward Models for Tool-Calling Large Language Models
von: Agarwal, Mayank, et al.
Veröffentlicht: (2025)
von: Agarwal, Mayank, et al.
Veröffentlicht: (2025)
Granite-Function Calling Model: Introducing Function Calling Abilities via Multi-task Learning of Granular Tasks
von: Abdelaziz, Ibrahim, et al.
Veröffentlicht: (2024)
von: Abdelaziz, Ibrahim, et al.
Veröffentlicht: (2024)
R2D2: Remembering, Replaying and Dynamic Decision Making with a Reflective Agentic Memory
von: Huang, Tenghao, et al.
Veröffentlicht: (2025)
von: Huang, Tenghao, et al.
Veröffentlicht: (2025)
MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
von: Katsis, Yannis, et al.
Veröffentlicht: (2025)
von: Katsis, Yannis, et al.
Veröffentlicht: (2025)
InspectorRAGet: An Introspection Platform for RAG Evaluation
von: Fadnis, Kshitij, et al.
Veröffentlicht: (2024)
von: Fadnis, Kshitij, et al.
Veröffentlicht: (2024)
Multi-Document Grounded Multi-Turn Synthetic Dialog Generation
von: Lee, Young-Suk, et al.
Veröffentlicht: (2024)
von: Lee, Young-Suk, et al.
Veröffentlicht: (2024)
Reducing the Scope of Language Models
von: Yunis, David, et al.
Veröffentlicht: (2024)
von: Yunis, David, et al.
Veröffentlicht: (2024)
Live API-Bench: 2500+ Live APIs for Testing Multi-Step Tool Calling
von: Elder, Benjamin, et al.
Veröffentlicht: (2025)
von: Elder, Benjamin, et al.
Veröffentlicht: (2025)
RAGAPHENE: A RAG Annotation Platform with Human Enhancements and Edits
von: Fadnis, Kshitij, et al.
Veröffentlicht: (2025)
von: Fadnis, Kshitij, et al.
Veröffentlicht: (2025)
LongFuncEval: Measuring the effectiveness of long context models for function calling
von: Kate, Kiran, et al.
Veröffentlicht: (2025)
von: Kate, Kiran, et al.
Veröffentlicht: (2025)
Transient Turn Injection: Exposing Stateless Multi-Turn Vulnerabilities in Large Language Models
von: Rayhan, Naheed, et al.
Veröffentlicht: (2026)
von: Rayhan, Naheed, et al.
Veröffentlicht: (2026)
Semileptonic decay and form factors of $Ω_b^- \rightarrow Ω_c^0\,e\,\bar{ν_e}$
von: Patel, Kinjal, et al.
Veröffentlicht: (2025)
von: Patel, Kinjal, et al.
Veröffentlicht: (2025)
Electromagnetic and weak decay of singly Heavy Baryons (Qqq)
von: Patel, Kinjal, et al.
Veröffentlicht: (2025)
von: Patel, Kinjal, et al.
Veröffentlicht: (2025)
Transition properties of Doubly Heavy Baryons
von: Patel, Kinjal, et al.
Veröffentlicht: (2024)
von: Patel, Kinjal, et al.
Veröffentlicht: (2024)
Semileptonic decay form factors of $Ξ_b^0 \rightarrow Ξ_c^+\ell\barν_{\ell}$ in HQET
von: Patel, Kinjal, et al.
Veröffentlicht: (2026)
von: Patel, Kinjal, et al.
Veröffentlicht: (2026)
Putting It All into Context: Simplifying Agents with LCLMs
von: Jiang, Mingjian, et al.
Veröffentlicht: (2025)
von: Jiang, Mingjian, et al.
Veröffentlicht: (2025)
Tool Calling for Arabic LLMs: Data Strategies and Instruction Tuning
von: Ersoy, Asim, et al.
Veröffentlicht: (2025)
von: Ersoy, Asim, et al.
Veröffentlicht: (2025)
ToolWeave: Structured Synthesis of Complex Multi-Turn Tool-Calling Dialogues
von: Khandelwal, Dinesh, et al.
Veröffentlicht: (2026)
von: Khandelwal, Dinesh, et al.
Veröffentlicht: (2026)
Automatic Detection and Classification of Corona Infection (COVID-19) from X-ray Images Using Convolution Neural Network
von: Patel, Kinjal A, et al.
Veröffentlicht: (2024)
von: Patel, Kinjal A, et al.
Veröffentlicht: (2024)
Tragico varamento masivo en Luisiana confirma la presencia de tortugas Kempis. / Deborah Crouse
von: Crouse, Deborah
Veröffentlicht: (1993)
von: Crouse, Deborah
Veröffentlicht: (1993)
On the Nature of Discrete Space-Time Part 2: Special Relativity in Discrete Space-Time
von: Crouse, David
Veröffentlicht: (2024)
von: Crouse, David
Veröffentlicht: (2024)
Textbooks 101: Textbook Collection at the University of Minnesota
von: Crouse, Caroline
Veröffentlicht: (2007)
von: Crouse, Caroline
Veröffentlicht: (2007)
What Makes Local Updates Effective: The Role of Data Heterogeneity and Smoothness
von: Patel, Kumar Kshitij
Veröffentlicht: (2025)
von: Patel, Kumar Kshitij
Veröffentlicht: (2025)
Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models
von: Luo, Zheng, et al.
Veröffentlicht: (2026)
von: Luo, Zheng, et al.
Veröffentlicht: (2026)
Repairing Tool Calls Using Post-tool Execution Reflection and RAG
von: Tsay, Jason, et al.
Veröffentlicht: (2025)
von: Tsay, Jason, et al.
Veröffentlicht: (2025)
Stateless Citizenship
von: Molavi, Shourideh C.
Veröffentlicht: (2018)
von: Molavi, Shourideh C.
Veröffentlicht: (2018)
Activated LoRA: Fine-tuned LLMs for Intrinsics
von: Greenewald, Kristjan, et al.
Veröffentlicht: (2025)
von: Greenewald, Kristjan, et al.
Veröffentlicht: (2025)
Multi-Turn Reinforcement Learning for Tool-Calling Agents with Iterative Reward Calibration
von: Modecrua, Wachiravit, et al.
Veröffentlicht: (2026)
von: Modecrua, Wachiravit, et al.
Veröffentlicht: (2026)
ECG-Agent: On-Device Tool-Calling Agent for ECG Multi-Turn Dialogue
von: Chung, Hyunseung, et al.
Veröffentlicht: (2026)
von: Chung, Hyunseung, et al.
Veröffentlicht: (2026)
Analyzing child health and water, sanitation, hygiene facilities in Punjab, Pakistan: A multilevel and spatial approach
von: Muhammad Munawar Hussain, et al.
Veröffentlicht: (2024)
von: Muhammad Munawar Hussain, et al.
Veröffentlicht: (2024)
Water, Sanitation, Hygiene Facilities and Economic Well‐Being: A Multilevel and Spatial Analysis in Punjab, Pakistan
von: Muhammad Munawar Hussain, et al.
Veröffentlicht: (2024)
von: Muhammad Munawar Hussain, et al.
Veröffentlicht: (2024)
The Six Sigma Agent: Achieving Enterprise-Grade Reliability in LLM Systems Through Consensus-Driven Decomposed Execution
von: Patel, Khush, et al.
Veröffentlicht: (2026)
von: Patel, Khush, et al.
Veröffentlicht: (2026)
Beyond State Machines: Executing Network Procedures with Agentic Tool-Calling Sequences
von: Garigipati, Purna Sai, et al.
Veröffentlicht: (2026)
von: Garigipati, Purna Sai, et al.
Veröffentlicht: (2026)
Numerical investigation of fracture behaviour of polyurethane adhesives under the influence of moisture
von: Josyula, Siva Pavan, et al.
Veröffentlicht: (2024)
von: Josyula, Siva Pavan, et al.
Veröffentlicht: (2024)
Statelessness in Public Law
von: Pudzianowska, Dorota
Veröffentlicht: (2024)
von: Pudzianowska, Dorota
Veröffentlicht: (2024)
MiGrATe: Mixed-Policy GRPO for Adaptation at Test-Time
von: Phan, Peter, et al.
Veröffentlicht: (2025)
von: Phan, Peter, et al.
Veröffentlicht: (2025)
An analysis of project provenance through a novel method of framework analysis
von: Larson, T., et al.
Veröffentlicht: (2025)
von: Larson, T., et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMs
von: Basu, Kinjal, et al.
Veröffentlicht: (2024) -
NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls
von: Basu, Kinjal, et al.
Veröffentlicht: (2024) -
Formally Specifying the High-Level Behavior of LLM-Based Agents
von: Crouse, Maxwell, et al.
Veröffentlicht: (2023) -
ToolRM: Outcome Reward Models for Tool-Calling Large Language Models
von: Agarwal, Mayank, et al.
Veröffentlicht: (2025) -
Granite-Function Calling Model: Introducing Function Calling Abilities via Multi-task Learning of Granular Tasks
von: Abdelaziz, Ibrahim, et al.
Veröffentlicht: (2024)