MathViz-E: A Case-study in Domain-Specialized Tool-Using Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Bulusu, Arya, Man, Brandon, Jagmohan, Ashish, Vempaty, Aditya, Mari-Wyka, Jennifer, Akkil, Deepak |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multimodal Auto Validation For Self-Refinement in Web Agents
by: Azam, Ruhana, et al.
Published: (2024)
by: Azam, Ruhana, et al.
Published: (2024)
Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic Systems
by: Abuelsaad, Tamer, et al.
Published: (2024)
by: Abuelsaad, Tamer, et al.
Published: (2024)
Learning API Functionality from In-Context Demonstrations for Tool-based Agents
by: Patel, Bhrij, et al.
Published: (2025)
by: Patel, Bhrij, et al.
Published: (2025)
Reflection-Based Memory For Web navigation Agents
by: Azam, Ruhana, et al.
Published: (2025)
by: Azam, Ruhana, et al.
Published: (2025)
MathDuels: Evaluating LLMs as Problem Posers and Solvers
by: Xu, Zhiqiu, et al.
Published: (2026)
by: Xu, Zhiqiu, et al.
Published: (2026)
Towards Online Code Specialization of Systems
by: Anand, Vaastav, et al.
Published: (2025)
by: Anand, Vaastav, et al.
Published: (2025)
A Tool for Test Case Scenarios Generation Using Large Language Models
by: Sami, Abdul Malik, et al.
Published: (2024)
by: Sami, Abdul Malik, et al.
Published: (2024)
SEAL: Suite for Evaluating API-use of LLMs
by: Kim, Woojeong, et al.
Published: (2024)
by: Kim, Woojeong, et al.
Published: (2024)
SolAgent: A Specialized Multi-Agent Framework for Solidity Code Generation
by: Chen, Wei, et al.
Published: (2026)
by: Chen, Wei, et al.
Published: (2026)
Design and Implementation of a Domain-specific Language for Modelling Evacuation Scenarios Using Eclipse EMG/GMF Tool
by: Banerjee, Heerok
Published: (2025)
by: Banerjee, Heerok
Published: (2025)
When Agents Fail: A Comprehensive Study of Bugs in LLM Agents with Automated Labeling
by: Islam, Niful, et al.
Published: (2026)
by: Islam, Niful, et al.
Published: (2026)
Managing Uncertainty in LLM-based Multi-Agent System Operation
by: Zhang, Man, et al.
Published: (2026)
by: Zhang, Man, et al.
Published: (2026)
LeakageDetector: An Open Source Data Leakage Analysis Tool in Machine Learning Pipelines
by: AlOmar, Eman Abdullah, et al.
Published: (2025)
by: AlOmar, Eman Abdullah, et al.
Published: (2025)
Towards Verifiably Safe Tool Use for LLM Agents
by: Doshi, Aarya, et al.
Published: (2026)
by: Doshi, Aarya, et al.
Published: (2026)
ToolFuzz -- Automated Agent Tool Testing
by: Milev, Ivan, et al.
Published: (2025)
by: Milev, Ivan, et al.
Published: (2025)
Automating Android Build Repair: Bridging the Reasoning-Execution Gap in LLM Agents with Domain-Specific Tools
by: Son, Ha Min, et al.
Published: (2025)
by: Son, Ha Min, et al.
Published: (2025)
SBFT Tool Competition 2024 -- Python Test Case Generation Track
by: Erni, Nicolas, et al.
Published: (2024)
by: Erni, Nicolas, et al.
Published: (2024)
SBFT Tool Competition 2025 -- Java Test Case Generation Track
by: Kifetew, Fitsum, et al.
Published: (2025)
by: Kifetew, Fitsum, et al.
Published: (2025)
Towards Supporting Quality Architecture Evaluation with LLM Tools
by: Capilla, Rafael, et al.
Published: (2026)
by: Capilla, Rafael, et al.
Published: (2026)
Property Testing for Ocean Models. Can We Specify It? (Invited Talk)
by: Cherian, Deepak A.
Published: (2025)
by: Cherian, Deepak A.
Published: (2025)
CompileAgent: Automated Real-World Repo-Level Compilation with Tool-Integrated LLM-based Agent System
by: Hu, Li, et al.
Published: (2025)
by: Hu, Li, et al.
Published: (2025)
SWITCH: An Exemplar for Evaluating Self-Adaptive ML-Enabled Systems
by: Marda, Arya, et al.
Published: (2024)
by: Marda, Arya, et al.
Published: (2024)
An Executable Benchmarking Suite for Tool-Using Agents
by: Zhong, Zhiqing, et al.
Published: (2026)
by: Zhong, Zhiqing, et al.
Published: (2026)
CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challenges
by: Zhang, Kechi, et al.
Published: (2024)
by: Zhang, Kechi, et al.
Published: (2024)
Tool for Supporting Debugging and Understanding of Normative Requirements Using LLMs
by: Kleijwegt, Alex, et al.
Published: (2025)
by: Kleijwegt, Alex, et al.
Published: (2025)
Open, Reliable, and Collective: A Community-Driven Framework for Tool-Using AI Agents
by: Dang, Hy, et al.
Published: (2026)
by: Dang, Hy, et al.
Published: (2026)
Class-Level Code Generation from Natural Language Using Iterative, Tool-Enhanced Reasoning over Repository
by: Deshpande, Ajinkya, et al.
Published: (2024)
by: Deshpande, Ajinkya, et al.
Published: (2024)
DADL: A Declarative Description Language for Enterprise Tool Libraries in LLM Agent Systems
by: Dunkel, Axel
Published: (2026)
by: Dunkel, Axel
Published: (2026)
Test Case Specification Techniques and System Testing Tools in the Automotive Industry: A Review
by: Zyberaj, Denesa, et al.
Published: (2025)
by: Zyberaj, Denesa, et al.
Published: (2025)
Toward an Understanding of Developer Behaviour while Using Bug Localization Tools
by: Pedreira, Pablo Diaz, et al.
Published: (2026)
by: Pedreira, Pablo Diaz, et al.
Published: (2026)
The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration
by: Xu, Haoyuan, et al.
Published: (2026)
by: Xu, Haoyuan, et al.
Published: (2026)
DomAgent: Leveraging Knowledge Graphs and Case-Based Reasoning for Domain-Specific Code Generation
by: Wang, Shuai, et al.
Published: (2026)
by: Wang, Shuai, et al.
Published: (2026)
Test Case Generation for Simulink Models: An Experience from the E-Bike Domain
by: Marzella, Michael, et al.
Published: (2025)
by: Marzella, Michael, et al.
Published: (2025)
ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering
by: Liu, Marianne Menglin, et al.
Published: (2025)
by: Liu, Marianne Menglin, et al.
Published: (2025)
ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents
by: Li, Dawei, et al.
Published: (2026)
by: Li, Dawei, et al.
Published: (2026)
Analyzing C/C++ Library Migrations at the Package-level: Prevalence, Domains, Targets and Rationals across Seven Package Management Tools
by: Gu, Haiqiao, et al.
Published: (2025)
by: Gu, Haiqiao, et al.
Published: (2025)
Traceability and Accountability in Role-Specialized Multi-Agent LLM Pipelines
by: Barrak, Amine
Published: (2025)
by: Barrak, Amine
Published: (2025)
An Event-Driven Tool for Context-Aware Code Smell Detection Using SmellDSL
by: Viegas, Matheus dos Santos, et al.
Published: (2026)
by: Viegas, Matheus dos Santos, et al.
Published: (2026)
Exploring the Output of Software Testing Tools through a Visual Comparative Analysis
by: Lit, Brandon, et al.
Published: (2026)
by: Lit, Brandon, et al.
Published: (2026)
ToolRosella: Translating Code Repositories into Standardized Tools for Scientific Agents
by: Di, Shimin, et al.
Published: (2026)
by: Di, Shimin, et al.
Published: (2026)
Similar Items
-
Multimodal Auto Validation For Self-Refinement in Web Agents
by: Azam, Ruhana, et al.
Published: (2024) -
Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic Systems
by: Abuelsaad, Tamer, et al.
Published: (2024) -
Learning API Functionality from In-Context Demonstrations for Tool-based Agents
by: Patel, Bhrij, et al.
Published: (2025) -
Reflection-Based Memory For Web navigation Agents
by: Azam, Ruhana, et al.
Published: (2025) -
MathDuels: Evaluating LLMs as Problem Posers and Solvers
by: Xu, Zhiqiu, et al.
Published: (2026)