Salvato in:
| Autori principali: | Nathan, Varun, Guha, Shreyas, Kumar, Ayush |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2602.14955 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Planning-Aware Code Infilling via Horizon-Length Prediction
di: Ding, Yifeng, et al.
Pubblicazione: (2024)
di: Ding, Yifeng, et al.
Pubblicazione: (2024)
ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering
di: Liu, Marianne Menglin, et al.
Pubblicazione: (2025)
di: Liu, Marianne Menglin, et al.
Pubblicazione: (2025)
Quality Matters: Evaluating Synthetic Data for Tool-Using LLMs
di: Iskander, Shadi, et al.
Pubblicazione: (2024)
di: Iskander, Shadi, et al.
Pubblicazione: (2024)
PERC: Plan-As-Query Example Retrieval for Underrepresented Code Generation
di: Yoo, Jaeseok, et al.
Pubblicazione: (2024)
di: Yoo, Jaeseok, et al.
Pubblicazione: (2024)
Evaluation of Code LLMs on Geospatial Code Generation
di: Gramacki, Piotr, et al.
Pubblicazione: (2024)
di: Gramacki, Piotr, et al.
Pubblicazione: (2024)
LeDex: Training LLMs to Better Self-Debug and Explain Code
di: Jiang, Nan, et al.
Pubblicazione: (2024)
di: Jiang, Nan, et al.
Pubblicazione: (2024)
MathDuels: Evaluating LLMs as Problem Posers and Solvers
di: Xu, Zhiqiu, et al.
Pubblicazione: (2026)
di: Xu, Zhiqiu, et al.
Pubblicazione: (2026)
AgentPack: A Dataset of Code Changes, Co-Authored by Agents and Humans
di: Zi, Yangtian, et al.
Pubblicazione: (2025)
di: Zi, Yangtian, et al.
Pubblicazione: (2025)
FairCoder: Evaluating Social Bias of LLMs in Code Generation
di: Du, Yongkang, et al.
Pubblicazione: (2025)
di: Du, Yongkang, et al.
Pubblicazione: (2025)
The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration
di: Xu, Haoyuan, et al.
Pubblicazione: (2026)
di: Xu, Haoyuan, et al.
Pubblicazione: (2026)
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios
di: Huang, Shiting, et al.
Pubblicazione: (2025)
di: Huang, Shiting, et al.
Pubblicazione: (2025)
AetherCode: Evaluating LLMs' Ability to Win In Premier Programming Competitions
di: Wang, Zihan, et al.
Pubblicazione: (2025)
di: Wang, Zihan, et al.
Pubblicazione: (2025)
Using Large Language Models for Student-Code Guided Test Case Generation in Computer Science Education
di: Kumar, Nischal Ashok, et al.
Pubblicazione: (2024)
di: Kumar, Nischal Ashok, et al.
Pubblicazione: (2024)
CodeScout: Contextual Problem Statement Enhancement for Software Agents
di: Suri, Manan, et al.
Pubblicazione: (2026)
di: Suri, Manan, et al.
Pubblicazione: (2026)
Privacy Policy Analysis through Prompt Engineering for LLMs
di: Goknil, Arda, et al.
Pubblicazione: (2024)
di: Goknil, Arda, et al.
Pubblicazione: (2024)
MATCH: Task-Driven Code Evaluation through Contrastive Learning
di: Ghoummaid, Marah, et al.
Pubblicazione: (2025)
di: Ghoummaid, Marah, et al.
Pubblicazione: (2025)
Evaluation of LLMs on Syntax-Aware Code Fill-in-the-Middle Tasks
di: Gong, Linyuan, et al.
Pubblicazione: (2024)
di: Gong, Linyuan, et al.
Pubblicazione: (2024)
MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
di: Huang, Yue, et al.
Pubblicazione: (2023)
di: Huang, Yue, et al.
Pubblicazione: (2023)
EffiCoder: Enhancing Code Generation in Large Language Models through Efficiency-Aware Fine-tuning
di: Huang, Dong, et al.
Pubblicazione: (2024)
di: Huang, Dong, et al.
Pubblicazione: (2024)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
di: Fu, Lingyue, et al.
Pubblicazione: (2025)
di: Fu, Lingyue, et al.
Pubblicazione: (2025)
Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination
di: Zheng, Jiasheng, et al.
Pubblicazione: (2026)
di: Zheng, Jiasheng, et al.
Pubblicazione: (2026)
Benchmarking Failures in Tool-Augmented Language Models
di: Treviño, Eduardo, et al.
Pubblicazione: (2025)
di: Treviño, Eduardo, et al.
Pubblicazione: (2025)
From Output to Evaluation: Does Raw Instruction-Tuned Code LLMs Output Suffice for Fill-in-the-Middle Code Generation?
di: Ahmad, Wasi Uddin, et al.
Pubblicazione: (2025)
di: Ahmad, Wasi Uddin, et al.
Pubblicazione: (2025)
Evaluating Plan Compliance in Autonomous Programming Agents
di: Liu, Shuyang, et al.
Pubblicazione: (2026)
di: Liu, Shuyang, et al.
Pubblicazione: (2026)
CodeTool: Enhancing Programmatic Tool Invocation of LLMs via Process Supervision
di: Lu, Yifei, et al.
Pubblicazione: (2025)
di: Lu, Yifei, et al.
Pubblicazione: (2025)
Sanskrit Knowledge-based Systems: Annotation and Computational Tools
di: Terdalkar, Hrishikesh
Pubblicazione: (2024)
di: Terdalkar, Hrishikesh
Pubblicazione: (2024)
ToolRegistry: A Protocol-Agnostic Tool Management Library for Function-Calling LLMs
di: Ding, Peng, et al.
Pubblicazione: (2025)
di: Ding, Peng, et al.
Pubblicazione: (2025)
A Review of Prominent Paradigms for LLM-Based Agents: Tool Use (Including RAG), Planning, and Feedback Learning
di: Li, Xinzhe
Pubblicazione: (2024)
di: Li, Xinzhe
Pubblicazione: (2024)
Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure
di: Yang, Zheyuan, et al.
Pubblicazione: (2025)
di: Yang, Zheyuan, et al.
Pubblicazione: (2025)
Dont Stop Early: Scalable Enterprise Deep Research with Controlled Information Flow and Evidence-Aware Termination
di: Choubey, Prafulla Kumar, et al.
Pubblicazione: (2026)
di: Choubey, Prafulla Kumar, et al.
Pubblicazione: (2026)
JEDI: Java Evaluation of Declarative and Imperative Queries
di: Schiavio, Filippo, et al.
Pubblicazione: (2026)
di: Schiavio, Filippo, et al.
Pubblicazione: (2026)
Characterizing and Evaluating the Reliability of LLMs against Jailbreak Attacks
di: Chen, Kexin, et al.
Pubblicazione: (2024)
di: Chen, Kexin, et al.
Pubblicazione: (2024)
SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
di: Chen, Shiqi, et al.
Pubblicazione: (2026)
di: Chen, Shiqi, et al.
Pubblicazione: (2026)
Multi-Programming Language Sandbox for LLMs
di: Dou, Shihan, et al.
Pubblicazione: (2024)
di: Dou, Shihan, et al.
Pubblicazione: (2024)
Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval
di: Wang, Jiexin, et al.
Pubblicazione: (2024)
di: Wang, Jiexin, et al.
Pubblicazione: (2024)
LLMs in Mobile Apps: Practices, Challenges, and Opportunities
di: Hau, Kimberly, et al.
Pubblicazione: (2025)
di: Hau, Kimberly, et al.
Pubblicazione: (2025)
Learning Code Preference via Synthetic Evolution
di: Liu, Jiawei, et al.
Pubblicazione: (2024)
di: Liu, Jiawei, et al.
Pubblicazione: (2024)
SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations
di: Wang, Shuaiqi, et al.
Pubblicazione: (2026)
di: Wang, Shuaiqi, et al.
Pubblicazione: (2026)
Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs
di: Lu, Yuxuan, et al.
Pubblicazione: (2026)
di: Lu, Yuxuan, et al.
Pubblicazione: (2026)
Beyond Accuracy: A Cognitive Load Framework for Mapping the Capability Boundaries of Tool-use Agents
di: Wang, Qihao, et al.
Pubblicazione: (2026)
di: Wang, Qihao, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Planning-Aware Code Infilling via Horizon-Length Prediction
di: Ding, Yifeng, et al.
Pubblicazione: (2024) -
ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering
di: Liu, Marianne Menglin, et al.
Pubblicazione: (2025) -
Quality Matters: Evaluating Synthetic Data for Tool-Using LLMs
di: Iskander, Shadi, et al.
Pubblicazione: (2024) -
PERC: Plan-As-Query Example Retrieval for Underrepresented Code Generation
di: Yoo, Jaeseok, et al.
Pubblicazione: (2024) -
Evaluation of Code LLMs on Geospatial Code Generation
di: Gramacki, Piotr, et al.
Pubblicazione: (2024)