Towards Reliable Benchmarking: A Contamination Free, Controllable Evaluation Framework for Multi-step LLM Function Calling
Fuente:
arXiv
Saved in:
| Main Authors: | Maekawa, Seiji, Hassell, Jackson, Pezeshkpour, Pouya, Mitchell, Tom, Hruschka, Estevam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs
by: Pezeshkpour, Pouya, et al.
Published: (2026)
by: Pezeshkpour, Pouya, et al.
Published: (2026)
Align then Train: Efficient Retrieval Adapter Learning
by: Maekawa, Seiji, et al.
Published: (2026)
by: Maekawa, Seiji, et al.
Published: (2026)
From Task Solving to Robust Real-World Adaptation in LLM Agents
by: Pezeshkpour, Pouya, et al.
Published: (2026)
by: Pezeshkpour, Pouya, et al.
Published: (2026)
Multi-Conditional Ranking with Large Language Models
by: Pezeshkpour, Pouya, et al.
Published: (2024)
by: Pezeshkpour, Pouya, et al.
Published: (2024)
Learning Beyond the Surface: How Far Can Continual Pre-Training with LoRA Enhance LLMs' Domain-Specific Insight Learning?
by: Pezeshkpour, Pouya, et al.
Published: (2025)
by: Pezeshkpour, Pouya, et al.
Published: (2025)
Insight-RAG: Enhancing LLMs with Insight-Driven Augmentation
by: Pezeshkpour, Pouya, et al.
Published: (2025)
by: Pezeshkpour, Pouya, et al.
Published: (2025)
From Single to Multi: How LLMs Hallucinate in Multi-Document Summarization
by: Belem, Catarina G., et al.
Published: (2024)
by: Belem, Catarina G., et al.
Published: (2024)
Evaluating Bias in LLMs for Job-Resume Matching: Gender, Race, and Education
by: Iso, Hayate, et al.
Published: (2025)
by: Iso, Hayate, et al.
Published: (2025)
Reasoning Capacity in Multi-Agent Systems: Limitations, Challenges and Human-Centered Solutions
by: Pezeshkpour, Pouya, et al.
Published: (2024)
by: Pezeshkpour, Pouya, et al.
Published: (2024)
From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models
by: Bayat, Farima Fatahi, et al.
Published: (2025)
by: Bayat, Farima Fatahi, et al.
Published: (2025)
Less is More for Long Document Summary Evaluation by LLMs
by: Wu, Yunshu, et al.
Published: (2023)
by: Wu, Yunshu, et al.
Published: (2023)
Learning from Supervision with Semantic and Episodic Memory: A Reflective Approach to Agent Adaptation
by: Hassell, Jackson, et al.
Published: (2025)
by: Hassell, Jackson, et al.
Published: (2025)
Mixed Signals: Decoding VLMs' Reasoning and Underlying Bias in Vision-Language Conflict
by: Pezeshkpour, Pouya, et al.
Published: (2025)
by: Pezeshkpour, Pouya, et al.
Published: (2025)
Do Agents Need to Plan Step-by-Step? Rethinking Planning Horizon in Data-Centric Tool Calling
by: Otani, Naoki, et al.
Published: (2026)
by: Otani, Naoki, et al.
Published: (2026)
Natural Language Processing for Human Resources: A Survey
by: Otani, Naoki, et al.
Published: (2024)
by: Otani, Naoki, et al.
Published: (2024)
The Rarity Blind Spot: A Framework for Evaluating Statistical Reasoning in LLMs
by: Maekawa, Seiji, et al.
Published: (2025)
by: Maekawa, Seiji, et al.
Published: (2025)
Towards Probabilistic Question Answering Over Tabular Data
by: Shen, Chen, et al.
Published: (2025)
by: Shen, Chen, et al.
Published: (2025)
LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs
by: Davoodi, Arash Gholami, et al.
Published: (2024)
by: Davoodi, Arash Gholami, et al.
Published: (2024)
FactLens: Benchmarking Fine-Grained Fact Verification
by: Mitra, Kushan, et al.
Published: (2024)
by: Mitra, Kushan, et al.
Published: (2024)
Functional Consistency of LLM Code Embeddings: A Self-Evolving Data Synthesis Framework for Benchmarking
by: Li, Zhuohao, et al.
Published: (2025)
by: Li, Zhuohao, et al.
Published: (2025)
Holistic Reasoning with Long-Context LMs: A Benchmark for Database Operations on Massive Textual Data
by: Maekawa, Seiji, et al.
Published: (2024)
by: Maekawa, Seiji, et al.
Published: (2024)
A Dynamic Self-Evolving Extraction System
by: Amin-Naseri, Moin, et al.
Published: (2026)
by: Amin-Naseri, Moin, et al.
Published: (2026)
An LLM-Tool Compiler for Fused Parallel Function Calling
by: Singh, Simranjit, et al.
Published: (2024)
by: Singh, Simranjit, et al.
Published: (2024)
Same Content, Different Representations: A Controlled Study for Table QA
by: Zhang, Yue, et al.
Published: (2025)
by: Zhang, Yue, et al.
Published: (2025)
Verification-Aware Planning for Multi-Agent Systems
by: Xu, Tianyang, et al.
Published: (2025)
by: Xu, Tianyang, et al.
Published: (2025)
RECAP: REwriting Conversations for Intent Understanding in Agentic Planning
by: Mitra, Kushan, et al.
Published: (2025)
by: Mitra, Kushan, et al.
Published: (2025)
Visualizing the Evaluation of Functional Programs for Debugging
by: Whitington, John, et al.
Published: (2024)
by: Whitington, John, et al.
Published: (2024)
Reasoning about External Calls
by: Drossopoulou, Sophia, et al.
Published: (2025)
by: Drossopoulou, Sophia, et al.
Published: (2025)
SwiftEval: Developing a Language-Specific Benchmark for LLM-generated Code Evaluation
by: Petrukha, Ivan, et al.
Published: (2025)
by: Petrukha, Ivan, et al.
Published: (2025)
FreeCHR: An Algebraic Framework for CHR-Embeddings
by: Rechenberger, Sascha, et al.
Published: (2023)
by: Rechenberger, Sascha, et al.
Published: (2023)
Controllable and Reliable Knowledge-Intensive Task-Oriented Conversational Agents with Declarative Genie Worksheets
by: Joshi, Harshit, et al.
Published: (2024)
by: Joshi, Harshit, et al.
Published: (2024)
Towards Contamination Resistant Benchmarks
by: Musawi, Rahmatullah, et al.
Published: (2025)
by: Musawi, Rahmatullah, et al.
Published: (2025)
Dual-Numbers Reverse AD for Functional Array Languages
by: Smeding, Tom, et al.
Published: (2025)
by: Smeding, Tom, et al.
Published: (2025)
Effects and Coeffects in Call-By-Push-Value (Extended Version)
by: Torczon, Cassia, et al.
Published: (2023)
by: Torczon, Cassia, et al.
Published: (2023)
Characterizing Large Language Models as Rationalizers of Knowledge-intensive Tasks
by: Mishra, Aditi, et al.
Published: (2023)
by: Mishra, Aditi, et al.
Published: (2023)
Towards Automated Verification of LLM-Synthesized C Programs
by: Mukherjee, Prasita, et al.
Published: (2024)
by: Mukherjee, Prasita, et al.
Published: (2024)
Does Task Complexity Moderate the Benefits of Liveness? A Controlled Experiment
by: Rein, Patrick, et al.
Published: (2024)
by: Rein, Patrick, et al.
Published: (2024)
Don't Call Us, We'll Call You: Towards Mixed-Initiative Interactive Proof Assistants for Programming Language Theory
by: Verter, Jan Liam, et al.
Published: (2024)
by: Verter, Jan Liam, et al.
Published: (2024)
Towards General Loop Invariant Generation: A Benchmark of Programs with Memory Manipulation
by: Liu, Chang, et al.
Published: (2023)
by: Liu, Chang, et al.
Published: (2023)
A2H-MAS: An Algorithm-to-HLS Multi-Agent System for Automated and Reliable FPGA Implementation
by: Lei, Jie, et al.
Published: (2025)
by: Lei, Jie, et al.
Published: (2025)
Similar Items
-
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs
by: Pezeshkpour, Pouya, et al.
Published: (2026) -
Align then Train: Efficient Retrieval Adapter Learning
by: Maekawa, Seiji, et al.
Published: (2026) -
From Task Solving to Robust Real-World Adaptation in LLM Agents
by: Pezeshkpour, Pouya, et al.
Published: (2026) -
Multi-Conditional Ranking with Large Language Models
by: Pezeshkpour, Pouya, et al.
Published: (2024) -
Learning Beyond the Surface: How Far Can Continual Pre-Training with LoRA Enhance LLMs' Domain-Specific Insight Learning?
by: Pezeshkpour, Pouya, et al.
Published: (2025)