LongFuncEval: Measuring the effectiveness of long context models for function calling
Fuente:
arXiv
Guardado en:
| Autores principales: | Kate, Kiran, Pedapati, Tejaswini, Basu, Kinjal, Rizk, Yara, Chenthamarakshan, Vijil, Chaudhury, Subhajit, Agarwal, Mayank, Abdelaziz, Ibrahim |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ToolRM: Outcome Reward Models for Tool-Calling Large Language Models
por: Agarwal, Mayank, et al.
Publicado: (2025)
por: Agarwal, Mayank, et al.
Publicado: (2025)
NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls
por: Basu, Kinjal, et al.
Publicado: (2024)
por: Basu, Kinjal, et al.
Publicado: (2024)
TabSketchFM: Sketch-based Tabular Representation Learning for Data Discovery over Data Lakes
por: Khatiwada, Aamod, et al.
Publicado: (2024)
por: Khatiwada, Aamod, et al.
Publicado: (2024)
Sparse Gradient Compression for Fine-Tuning Large Language Models
por: Yang, David H., et al.
Publicado: (2025)
por: Yang, David H., et al.
Publicado: (2025)
How Good Are LLMs at Processing Tool Outputs?
por: Kate, Kiran, et al.
Publicado: (2025)
por: Kate, Kiran, et al.
Publicado: (2025)
EXPLORER: Exploration-guided Reasoning for Textual Reinforcement Learning
por: Basu, Kinjal, et al.
Publicado: (2024)
por: Basu, Kinjal, et al.
Publicado: (2024)
Towards LLMs Robustness to Changes in Prompt Format Styles
por: Ngweta, Lilian, et al.
Publicado: (2025)
por: Ngweta, Lilian, et al.
Publicado: (2025)
API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMs
por: Basu, Kinjal, et al.
Publicado: (2024)
por: Basu, Kinjal, et al.
Publicado: (2024)
ZoomR: Memory Efficient Reasoning through Multi-Granularity Key Value Retrieval
por: Yang, David H., et al.
Publicado: (2026)
por: Yang, David H., et al.
Publicado: (2026)
EpMAN: Episodic Memory AttentioN for Generalizing to Longer Contexts
por: Chaudhury, Subhajit, et al.
Publicado: (2025)
por: Chaudhury, Subhajit, et al.
Publicado: (2025)
When Agents go Astray: Course-Correcting SWE Agents with PRMs
por: Gandhi, Shubham, et al.
Publicado: (2025)
por: Gandhi, Shubham, et al.
Publicado: (2025)
Large Language Models can be Strong Self-Detoxifiers
por: Ko, Ching-Yun, et al.
Publicado: (2024)
por: Ko, Ching-Yun, et al.
Publicado: (2024)
R2D2: Remembering, Replaying and Dynamic Decision Making with a Reflective Agentic Memory
por: Huang, Tenghao, et al.
Publicado: (2025)
por: Huang, Tenghao, et al.
Publicado: (2025)
From PEFT to DEFT: Parameter Efficient Finetuning for Reducing Activation Density in Transformers
por: Runwal, Bharat, et al.
Publicado: (2024)
por: Runwal, Bharat, et al.
Publicado: (2024)
Larimar: Large Language Models with Episodic Memory Control
por: Das, Payel, et al.
Publicado: (2024)
por: Das, Payel, et al.
Publicado: (2024)
Repairing Tool Calls Using Post-tool Execution Reflection and RAG
por: Tsay, Jason, et al.
Publicado: (2025)
por: Tsay, Jason, et al.
Publicado: (2025)
Structure-Informed Protein Language Model
por: Zhang, Zuobai, et al.
Publicado: (2024)
por: Zhang, Zuobai, et al.
Publicado: (2024)
ProtIR: Iterative Refinement between Retrievers and Predictors for Protein Function Annotation
por: Zhang, Zuobai, et al.
Publicado: (2024)
por: Zhang, Zuobai, et al.
Publicado: (2024)
Modular Prompt Learning Improves Vision-Language Models
por: Huang, Zhenhan, et al.
Publicado: (2025)
por: Huang, Zhenhan, et al.
Publicado: (2025)
Differentiable Prompt Learning for Vision Language Models
por: Huang, Zhenhan, et al.
Publicado: (2024)
por: Huang, Zhenhan, et al.
Publicado: (2024)
Intermediate Representations are Strong AI-Generated Image Detectors
por: Huang, Zhenhan, et al.
Publicado: (2026)
por: Huang, Zhenhan, et al.
Publicado: (2026)
FuncEvalGMN: Evaluating Functional Correctness of SQL via Graph Matching Network
por: Zhan, Yi, et al.
Publicado: (2024)
por: Zhan, Yi, et al.
Publicado: (2024)
EvalAssist: A Human-Centered Tool for LLM-as-a-Judge
por: Ashktorab, Zahra, et al.
Publicado: (2025)
por: Ashktorab, Zahra, et al.
Publicado: (2025)
MrBlankness/HyFunc: HyFunc -- KDD'26
por: Weibin Liao
Publicado: (2026)
por: Weibin Liao
Publicado: (2026)
Live API-Bench: 2500+ Live APIs for Testing Multi-Step Tool Calling
por: Elder, Benjamin, et al.
Publicado: (2025)
por: Elder, Benjamin, et al.
Publicado: (2025)
Large Language Model Confidence Estimation via Black-Box Access
por: Pedapati, Tejaswini, et al.
Publicado: (2024)
por: Pedapati, Tejaswini, et al.
Publicado: (2024)
Generation Constraint Scaling Can Mitigate Hallucination
por: Kollias, Georgios, et al.
Publicado: (2024)
por: Kollias, Georgios, et al.
Publicado: (2024)
GP-MoLFormer: A Foundation Model For Molecular Generation
por: Ross, Jerret, et al.
Publicado: (2024)
por: Ross, Jerret, et al.
Publicado: (2024)
GP-MoLFormer-Sim: Test Time Molecular Optimization through Contextual Similarity Guidance
por: Navratil, Jiri, et al.
Publicado: (2025)
por: Navratil, Jiri, et al.
Publicado: (2025)
Aligning Protein Conformation Ensemble Generation with Physical Feedback
por: Lu, Jiarui, et al.
Publicado: (2025)
por: Lu, Jiarui, et al.
Publicado: (2025)
Graph is all you need? Lightweight data-agnostic neural architecture search without training
por: Huang, Zhenhan, et al.
Publicado: (2024)
por: Huang, Zhenhan, et al.
Publicado: (2024)
Multi-Scale Representation Learning for Protein Fitness Prediction
por: Zhang, Zuobai, et al.
Publicado: (2024)
por: Zhang, Zuobai, et al.
Publicado: (2024)
Simulating Complex Multi-Turn Tool Calling Interactions in Stateless Execution Environments
por: Crouse, Maxwell, et al.
Publicado: (2026)
por: Crouse, Maxwell, et al.
Publicado: (2026)
Niveles de empatía según la escala de Jefferson en estudiantes de Medicina, Enfermería y Odontología de Honduras
por: Herman Rozengway Vijil
Publicado: (2016)
por: Herman Rozengway Vijil
Publicado: (2016)
Algunos factores claves del contexto poselectoral y del gobierno de Xiomara Castro en Honduras
por: Rolando Canizales Vijil
Publicado: (2022)
por: Rolando Canizales Vijil
Publicado: (2022)
Pautas de exámenes inteligentes
por: Herman Rozengway Vijil
Publicado: (2018)
por: Herman Rozengway Vijil
Publicado: (2018)
quadram-institute-bioscience/FuncDarkMatt: FuncDarkMatt v1.0.1
por: Sumeet Tiwari
Publicado: (2026)
por: Sumeet Tiwari
Publicado: (2026)
Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences
por: Ashktorab, Zahra, et al.
Publicado: (2024)
por: Ashktorab, Zahra, et al.
Publicado: (2024)
Formally Specifying the High-Level Behavior of LLM-Based Agents
por: Crouse, Maxwell, et al.
Publicado: (2023)
por: Crouse, Maxwell, et al.
Publicado: (2023)
On the Effects of Fine-tuning Language Models for Text-Based Reinforcement Learning
por: Gruppi, Mauricio, et al.
Publicado: (2024)
por: Gruppi, Mauricio, et al.
Publicado: (2024)
Ejemplares similares
-
ToolRM: Outcome Reward Models for Tool-Calling Large Language Models
por: Agarwal, Mayank, et al.
Publicado: (2025) -
NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls
por: Basu, Kinjal, et al.
Publicado: (2024) -
TabSketchFM: Sketch-based Tabular Representation Learning for Data Discovery over Data Lakes
por: Khatiwada, Aamod, et al.
Publicado: (2024) -
Sparse Gradient Compression for Fine-Tuning Large Language Models
por: Yang, David H., et al.
Publicado: (2025) -
How Good Are LLMs at Processing Tool Outputs?
por: Kate, Kiran, et al.
Publicado: (2025)