How Good Are LLMs at Processing Tool Outputs?
Fuente:
arXiv
Saved in:
| Main Authors: | Kate, Kiran, Rizk, Yara, Ghosh, Poulami, Gulati, Ashu, Chakraborti, Tathagata, Wright, Zidane, Agarwal, Mayank |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Repairing Tool Calls Using Post-tool Execution Reflection and RAG
by: Tsay, Jason, et al.
Published: (2025)
by: Tsay, Jason, et al.
Published: (2025)
ALIGN-FL: Architecture-independent Learning through Invariant Generative component sharing in Federated Learning
by: Gulati, Mayank, et al.
Published: (2025)
by: Gulati, Mayank, et al.
Published: (2025)
HiSpec: Hierarchical Speculative Decoding for LLMs
by: Kumar, Avinash, et al.
Published: (2025)
by: Kumar, Avinash, et al.
Published: (2025)
Multi-Output Distributional Fairness via Post-Processing
by: Li, Gang, et al.
Published: (2024)
by: Li, Gang, et al.
Published: (2024)
Aligners: Decoupling LLMs and Alignment
by: Ngweta, Lilian, et al.
Published: (2024)
by: Ngweta, Lilian, et al.
Published: (2024)
Derivation of Output Correlation Inferences for Multi-Output (aka Multi-Task) Gaussian Process
by: Watanabe, Shuhei
Published: (2025)
by: Watanabe, Shuhei
Published: (2025)
Agent Lifecycle Toolkit (ALTK): Reusable Middleware Components for Robust AI Agents
by: Wright, Zidane, et al.
Published: (2026)
by: Wright, Zidane, et al.
Published: (2026)
When Agents go Astray: Course-Correcting SWE Agents with PRMs
by: Gandhi, Shubham, et al.
Published: (2025)
by: Gandhi, Shubham, et al.
Published: (2025)
Dialogue Without Limits: Constant-Sized KV Caches for Extended Responses in LLMs
by: Ghadia, Ravi, et al.
Published: (2025)
by: Ghadia, Ravi, et al.
Published: (2025)
Enhancing Training Efficiency Using Packing with Flash Attention
by: Kundu, Achintya, et al.
Published: (2024)
by: Kundu, Achintya, et al.
Published: (2024)
It Takes a Good Model to Train a Good Model: Generalized Gaussian Priors for Optimized LLMs
by: Wu, Jun, et al.
Published: (2025)
by: Wu, Jun, et al.
Published: (2025)
Toward Carbon-Neutral Human AI: Rethinking Data, Computation, and Learning Paradigms for Sustainable Intelligence
by: Santosh, KC, et al.
Published: (2025)
by: Santosh, KC, et al.
Published: (2025)
AI-CARE: Carbon-Aware Reporting Evaluation Metric for AI Models
by: Santosh, KC, et al.
Published: (2026)
by: Santosh, KC, et al.
Published: (2026)
Learning Displacement-Robust Representations for Landslide Early Warning under Rainfall Forecast Uncertainty
by: Ozeki, Ren, et al.
Published: (2026)
by: Ozeki, Ren, et al.
Published: (2026)
Tube Loss: A Novel Approach for Prediction Interval Estimation
by: Anand, Pritam, et al.
Published: (2024)
by: Anand, Pritam, et al.
Published: (2024)
Geometric Scaling of Bayesian Inference in LLMs
by: Agarwal, Naman, et al.
Published: (2025)
by: Agarwal, Naman, et al.
Published: (2025)
Are LLMs Good Cryptic Crossword Solvers?
by: Sadallah, Abdelrahman, et al.
Published: (2024)
by: Sadallah, Abdelrahman, et al.
Published: (2024)
Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations
by: Aravindan, Ashwath Vaithinathan, et al.
Published: (2026)
by: Aravindan, Ashwath Vaithinathan, et al.
Published: (2026)
NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls
by: Basu, Kinjal, et al.
Published: (2024)
by: Basu, Kinjal, et al.
Published: (2024)
Are LLMs Good For Quantum Software, Architecture, and System Design?
by: Wawdhane, Sourish, et al.
Published: (2026)
by: Wawdhane, Sourish, et al.
Published: (2026)
RL-Struct: A Lightweight Reinforcement Learning Framework for Reliable Structured Output in LLMs
by: Hu, Ruike, et al.
Published: (2025)
by: Hu, Ruike, et al.
Published: (2025)
MEADOW: Memory-efficient Dataflow and Data Packing for Low Power Edge LLMs
by: Moitra, Abhishek, et al.
Published: (2025)
by: Moitra, Abhishek, et al.
Published: (2025)
Tool Unlearning for Tool-Augmented LLMs
by: Cheng, Jiali, et al.
Published: (2025)
by: Cheng, Jiali, et al.
Published: (2025)
Agentic Misalignment: How LLMs Could Be Insider Threats
by: Lynch, Aengus, et al.
Published: (2025)
by: Lynch, Aengus, et al.
Published: (2025)
Tool Zero: Training Tool-Augmented LLMs via Pure RL from Scratch
by: Zeng, Yirong, et al.
Published: (2025)
by: Zeng, Yirong, et al.
Published: (2025)
CodeTool: Enhancing Programmatic Tool Invocation of LLMs via Process Supervision
by: Lu, Yifei, et al.
Published: (2025)
by: Lu, Yifei, et al.
Published: (2025)
Throughput Optimization as a Strategic Lever in Large-Scale AI Systems: Evidence from Dataloader and Memory Profiling Innovations
by: Jha, Mayank
Published: (2026)
by: Jha, Mayank
Published: (2026)
Gradient Dynamics of Attention: How Cross-Entropy Sculpts Bayesian Manifolds
by: Agarwal, Naman, et al.
Published: (2025)
by: Agarwal, Naman, et al.
Published: (2025)
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
by: Nikolić, Kristina, et al.
Published: (2025)
by: Nikolić, Kristina, et al.
Published: (2025)
Mamba State-Space Models Are Lyapunov-Stable Learners
by: Halloran, John T., et al.
Published: (2024)
by: Halloran, John T., et al.
Published: (2024)
Decentralized Adversarial Training over Graphs
by: Cao, Ying, et al.
Published: (2023)
by: Cao, Ying, et al.
Published: (2023)
One Supervisor, Many Modalities: Adaptive Tool Orchestration for Autonomous Queries
by: Saini, Mayank, et al.
Published: (2026)
by: Saini, Mayank, et al.
Published: (2026)
Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments
by: Vishwakarma, Harsh, et al.
Published: (2025)
by: Vishwakarma, Harsh, et al.
Published: (2025)
Fuzzy Rule based Intelligent Cardiovascular Disease Prediction using Complex Event Processing
by: Kumar, Shashi Shekhar, et al.
Published: (2024)
by: Kumar, Shashi Shekhar, et al.
Published: (2024)
A Diagnosis and Treatment of Liver Diseases: Integrating Batch Processing, Rule-Based Event Detection and Explainable Artificial Intelligence
by: Chandra, Ritesh, et al.
Published: (2023)
by: Chandra, Ritesh, et al.
Published: (2023)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
by: Chaudhary, Maheep, et al.
Published: (2025)
by: Chaudhary, Maheep, et al.
Published: (2025)
Precision Tracked Transformer via Kalman Filtering, Kriging and Process Noise
by: Long, Bo, et al.
Published: (2026)
by: Long, Bo, et al.
Published: (2026)
GraphTool-Instruction: Revolutionizing Graph Reasoning in LLMs through Decomposed Subtask Instruction
by: Wang, Rongzheng, et al.
Published: (2024)
by: Wang, Rongzheng, et al.
Published: (2024)
Efficiency Boost in Decentralized Optimization: Reimagining Neighborhood Aggregation with Minimal Overhead
by: Kalwar, Durgesh, et al.
Published: (2025)
by: Kalwar, Durgesh, et al.
Published: (2025)
ConDiSim: Conditional Diffusion Models for Simulation Based Inference
by: Nautiyal, Mayank, et al.
Published: (2025)
by: Nautiyal, Mayank, et al.
Published: (2025)
Similar Items
-
Repairing Tool Calls Using Post-tool Execution Reflection and RAG
by: Tsay, Jason, et al.
Published: (2025) -
ALIGN-FL: Architecture-independent Learning through Invariant Generative component sharing in Federated Learning
by: Gulati, Mayank, et al.
Published: (2025) -
HiSpec: Hierarchical Speculative Decoding for LLMs
by: Kumar, Avinash, et al.
Published: (2025) -
Multi-Output Distributional Fairness via Post-Processing
by: Li, Gang, et al.
Published: (2024) -
Aligners: Decoupling LLMs and Alignment
by: Ngweta, Lilian, et al.
Published: (2024)