Saved in:
| Main Author: | Terdalkar, Hrishikesh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2406.18276 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration
by: Xu, Haoyuan, et al.
Published: (2026)
by: Xu, Haoyuan, et al.
Published: (2026)
ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering
by: Liu, Marianne Menglin, et al.
Published: (2025)
by: Liu, Marianne Menglin, et al.
Published: (2025)
ChartMark: A Structured Grammar for Chart Annotation
by: Chen, Yiyu, et al.
Published: (2025)
by: Chen, Yiyu, et al.
Published: (2025)
MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
by: Huang, Yue, et al.
Published: (2023)
by: Huang, Yue, et al.
Published: (2023)
From Prediction to Application: Language Model-based Code Knowledge Tracing with Domain Adaptive Pre-Training and Automatic Feedback System with Pedagogical Prompting for Comprehensive Programming Education
by: Lee, Unggi, et al.
Published: (2024)
by: Lee, Unggi, et al.
Published: (2024)
Benchmarking Failures in Tool-Augmented Language Models
by: Treviño, Eduardo, et al.
Published: (2025)
by: Treviño, Eduardo, et al.
Published: (2025)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
by: Fu, Lingyue, et al.
Published: (2025)
by: Fu, Lingyue, et al.
Published: (2025)
Self-reflection in Automated Qualitative Coding: Improving Text Annotation through Secondary LLM Critique
by: Dunivin, Zackary Okun, et al.
Published: (2026)
by: Dunivin, Zackary Okun, et al.
Published: (2026)
DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
by: Chen, Shiqi, et al.
Published: (2026)
by: Chen, Shiqi, et al.
Published: (2026)
MORTAR: Multi-turn Metamorphic Testing for LLM-based Dialogue Systems
by: Guo, Guoxiang, et al.
Published: (2024)
by: Guo, Guoxiang, et al.
Published: (2024)
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios
by: Huang, Shiting, et al.
Published: (2025)
by: Huang, Shiting, et al.
Published: (2025)
Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs
by: Lu, Yuxuan, et al.
Published: (2026)
by: Lu, Yuxuan, et al.
Published: (2026)
Beyond Accuracy: A Cognitive Load Framework for Mapping the Capability Boundaries of Tool-use Agents
by: Wang, Qihao, et al.
Published: (2026)
by: Wang, Qihao, et al.
Published: (2026)
RCAgent: Cloud Root Cause Analysis by Autonomous Agents with Tool-Augmented Large Language Models
by: Wang, Zefan, et al.
Published: (2023)
by: Wang, Zefan, et al.
Published: (2023)
Tool-Aware Planning in Contact Center AI: Evaluating LLMs through Lineage-Guided Query Decomposition
by: Nathan, Varun, et al.
Published: (2026)
by: Nathan, Varun, et al.
Published: (2026)
ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints
by: Choi, Hyeonje, et al.
Published: (2026)
by: Choi, Hyeonje, et al.
Published: (2026)
CodeUpdateArena: Benchmarking Knowledge Editing on API Updates
by: Liu, Zeyu Leo, et al.
Published: (2024)
by: Liu, Zeyu Leo, et al.
Published: (2024)
Adapting Large Language Models to Log Analysis with Interpretable Domain Knowledge
by: Ji, Yuhe, et al.
Published: (2024)
by: Ji, Yuhe, et al.
Published: (2024)
KCoEvo: A Knowledge Graph Augmented Framework for Evolutionary Code Generation
by: Kang, Jiazhen, et al.
Published: (2026)
by: Kang, Jiazhen, et al.
Published: (2026)
Knowledge-aware Alert Aggregation in Large-scale Cloud Systems: a Hybrid Approach
by: Kuang, Jinxi, et al.
Published: (2024)
by: Kuang, Jinxi, et al.
Published: (2024)
Automating Computational Reproducibility in Social Science: Comparing Prompt-Based and Agent-Based Approaches
by: Shah, Syed Mehtab Hussain, et al.
Published: (2026)
by: Shah, Syed Mehtab Hussain, et al.
Published: (2026)
Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability
by: Chen, Zejian, et al.
Published: (2026)
by: Chen, Zejian, et al.
Published: (2026)
Knowledge Boundary Probing and Demand-Guided Intervention for LLM-Based Power System Code Generation
by: Wu, Hui, et al.
Published: (2026)
by: Wu, Hui, et al.
Published: (2026)
Using Large Language Models for Student-Code Guided Test Case Generation in Computer Science Education
by: Kumar, Nischal Ashok, et al.
Published: (2024)
by: Kumar, Nischal Ashok, et al.
Published: (2024)
LogLM: From Task-based to Instruction-based Automated Log Analysis
by: Liu, Yilun, et al.
Published: (2024)
by: Liu, Yilun, et al.
Published: (2024)
Software-Based Dialogue Systems: Survey, Taxonomy and Challenges
by: Motger, Quim, et al.
Published: (2021)
by: Motger, Quim, et al.
Published: (2021)
Quality Matters: Evaluating Synthetic Data for Tool-Using LLMs
by: Iskander, Shadi, et al.
Published: (2024)
by: Iskander, Shadi, et al.
Published: (2024)
Contextualized Data-Wrangling Code Generation in Computational Notebooks
by: Huang, Junjie, et al.
Published: (2024)
by: Huang, Junjie, et al.
Published: (2024)
RePair: Automated Program Repair with Process-based Feedback
by: Zhao, Yuze, et al.
Published: (2024)
by: Zhao, Yuze, et al.
Published: (2024)
A Reference Architecture for Designing Foundation Model based Systems
by: Lu, Qinghua, et al.
Published: (2023)
by: Lu, Qinghua, et al.
Published: (2023)
RITFIS: Robust input testing framework for LLMs-based intelligent software
by: Xiao, Mingxuan, et al.
Published: (2024)
by: Xiao, Mingxuan, et al.
Published: (2024)
A Comprehensive Survey on Benchmarks and Solutions in Software Engineering of LLM-Empowered Agentic System
by: Guo, Jiale, et al.
Published: (2025)
by: Guo, Jiale, et al.
Published: (2025)
CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification
by: Tian, Yuchen, et al.
Published: (2024)
by: Tian, Yuchen, et al.
Published: (2024)
CodeBenchGen: Creating Scalable Execution-based Code Generation Benchmarks
by: Xie, Yiqing, et al.
Published: (2024)
by: Xie, Yiqing, et al.
Published: (2024)
Simulating Complex Multi-Turn Tool Calling Interactions in Stateless Execution Environments
by: Crouse, Maxwell, et al.
Published: (2026)
by: Crouse, Maxwell, et al.
Published: (2026)
Breaking the Cycle of Recurring Failures: Applying Generative AI to Root Cause Analysis in Legacy Banking Systems
by: Jin, Siyuan, et al.
Published: (2024)
by: Jin, Siyuan, et al.
Published: (2024)
To Diff or Not to Diff? Structure-Aware and Adaptive Output Formats for Efficient LLM-based Code Editing
by: Cheng, Wei, et al.
Published: (2026)
by: Cheng, Wei, et al.
Published: (2026)
SCOPE: Tree-based Self-Correcting Online Log Parsing via Syntactic-Semantic Collaboration
by: Fan, Dongyi, et al.
Published: (2026)
by: Fan, Dongyi, et al.
Published: (2026)
ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents
by: Jeon, YoungHoon, et al.
Published: (2026)
by: Jeon, YoungHoon, et al.
Published: (2026)
Similar Items
-
The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration
by: Xu, Haoyuan, et al.
Published: (2026) -
ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering
by: Liu, Marianne Menglin, et al.
Published: (2025) -
ChartMark: A Structured Grammar for Chart Annotation
by: Chen, Yiyu, et al.
Published: (2025) -
MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
by: Huang, Yue, et al.
Published: (2023) -
From Prediction to Application: Language Model-based Code Knowledge Tracing with Domain Adaptive Pre-Training and Automatic Feedback System with Pedagogical Prompting for Comprehensive Programming Education
by: Lee, Unggi, et al.
Published: (2024)