FireBench: Evaluating Instruction Following in Enterprise and API-Driven LLM Applications
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Yunfan, Bei, Yijie, Ravi, Jetashree, Garbacki, Pawel |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation
par: Pysklo, Hubert M., et autres
Publié: (2026)
par: Pysklo, Hubert M., et autres
Publié: (2026)
ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents
par: Jeon, YoungHoon, et autres
Publié: (2026)
par: Jeon, YoungHoon, et autres
Publié: (2026)
A Conceptual Framework for API Refactoring in Enterprise Application Architectures
par: Montesi, Fabrizio, et autres
Publié: (2024)
par: Montesi, Fabrizio, et autres
Publié: (2024)
ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
par: Zhang, Chenchen, et autres
Publié: (2025)
par: Zhang, Chenchen, et autres
Publié: (2025)
AutoMonitor-Bench: Evaluating the Reliability of LLM-Based Misbehavior Monitor
par: Yang, Shu, et autres
Publié: (2026)
par: Yang, Shu, et autres
Publié: (2026)
The Stability Trap: Evaluating the Reliability of LLM-Based Instruction Adherence Auditing
par: Shergadwala, Murtuza N.
Publié: (2026)
par: Shergadwala, Murtuza N.
Publié: (2026)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
par: Fu, Lingyue, et autres
Publié: (2025)
par: Fu, Lingyue, et autres
Publié: (2025)
Learning Selective LLM Autonomy from Copilot Feedback in Enterprise Customer Support Workflows
par: Borovkov, Nikita, et autres
Publié: (2026)
par: Borovkov, Nikita, et autres
Publié: (2026)
Revisiting the Reliability of Language Models in Instruction-Following
par: Dong, Jianshuo, et autres
Publié: (2025)
par: Dong, Jianshuo, et autres
Publié: (2025)
A Comparative Study on the Impact of Test-Driven Development (TDD) and Behavior-Driven Development (BDD) on Enterprise Software Delivery Effectiveness
par: Cui, Jun
Publié: (2024)
par: Cui, Jun
Publié: (2024)
CodeIF-Bench: Evaluating Instruction-Following Capabilities of Large Language Models in Interactive Code Generation
par: Wang, Peiding, et autres
Publié: (2025)
par: Wang, Peiding, et autres
Publié: (2025)
CodeUpdateArena: Benchmarking Knowledge Editing on API Updates
par: Liu, Zeyu Leo, et autres
Publié: (2024)
par: Liu, Zeyu Leo, et autres
Publié: (2024)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
par: Li, Jia, et autres
Publié: (2024)
par: Li, Jia, et autres
Publié: (2024)
SoAy: A Solution-based LLM API-using Methodology for Academic Information Seeking
par: Wang, Yuanchun, et autres
Publié: (2024)
par: Wang, Yuanchun, et autres
Publié: (2024)
Sphinx: Benchmarking and Modeling for LLM-Driven Pull Request Review
par: Zhang, Daoan, et autres
Publié: (2026)
par: Zhang, Daoan, et autres
Publié: (2026)
Enhancing Project-Specific Code Completion by Inferring Internal API Information
par: Deng, Le, et autres
Publié: (2025)
par: Deng, Le, et autres
Publié: (2025)
FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
par: Li, Wei, et autres
Publié: (2025)
par: Li, Wei, et autres
Publié: (2025)
Evaluating LLMs on Sequential API Call Through Automated Test Generation
par: Huang, Yuheng, et autres
Publié: (2025)
par: Huang, Yuheng, et autres
Publié: (2025)
Evaluating and Achieving Controllable Code Completion in Code LLM
par: Zhang, Jiajun, et autres
Publié: (2026)
par: Zhang, Jiajun, et autres
Publié: (2026)
Adaptable and Precise: Enterprise-Scenario LLM Function-Calling Capability Training Pipeline
par: Zeng, Guancheng, et autres
Publié: (2024)
par: Zeng, Guancheng, et autres
Publié: (2024)
UA-Code-Bench: A Competitive Programming Benchmark for Evaluating LLM Code Generation in Ukrainian
par: Syromiatnikov, Mykyta, et autres
Publié: (2025)
par: Syromiatnikov, Mykyta, et autres
Publié: (2025)
SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation
par: Ma, Zeyao, et autres
Publié: (2024)
par: Ma, Zeyao, et autres
Publié: (2024)
When "Better" Prompts Hurt: Evaluation-Driven Iteration for LLM Applications
par: Commey, Daniel
Publié: (2026)
par: Commey, Daniel
Publié: (2026)
Revolutionizing API Documentation through Summarization
par: Naghshzan, AmirHossein, et autres
Publié: (2024)
par: Naghshzan, AmirHossein, et autres
Publié: (2024)
Evaluating Retrieval-Augmented Generation Variants for Natural Language-Based SQL and API Call Generation
par: Marketsmüller, Michael, et autres
Publié: (2026)
par: Marketsmüller, Michael, et autres
Publié: (2026)
EffiBench: Benchmarking the Efficiency of Automatically Generated Code
par: Huang, Dong, et autres
Publié: (2024)
par: Huang, Dong, et autres
Publié: (2024)
BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
par: Zhuo, Terry Yue, et autres
Publié: (2024)
par: Zhuo, Terry Yue, et autres
Publié: (2024)
Comparing Developer and LLM Biases in Code Evaluation
par: Mittal, Aditya, et autres
Publié: (2026)
par: Mittal, Aditya, et autres
Publié: (2026)
Dont Stop Early: Scalable Enterprise Deep Research with Controlled Information Flow and Evidence-Aware Termination
par: Choubey, Prafulla Kumar, et autres
Publié: (2026)
par: Choubey, Prafulla Kumar, et autres
Publié: (2026)
AutoIOT: LLM-Driven Automated Natural Language Programming for AIoT Applications
par: Shen, Leming, et autres
Publié: (2025)
par: Shen, Leming, et autres
Publié: (2025)
ToolFactory: Automating Tool Generation by Leveraging LLM to Understand REST API Documentations
par: Ni, Xinyi, et autres
Publié: (2025)
par: Ni, Xinyi, et autres
Publié: (2025)
MATCH: Task-Driven Code Evaluation through Contrastive Learning
par: Ghoummaid, Marah, et autres
Publié: (2025)
par: Ghoummaid, Marah, et autres
Publié: (2025)
SWE-Dev: Evaluating and Training Autonomous Feature-Driven Software Development
par: Du, Yaxin, et autres
Publié: (2025)
par: Du, Yaxin, et autres
Publié: (2025)
Compositional API Recommendation for Library-Oriented Code Generation
par: Ma, Zexiong, et autres
Publié: (2024)
par: Ma, Zexiong, et autres
Publié: (2024)
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
par: Chou, Jason, et autres
Publié: (2025)
par: Chou, Jason, et autres
Publié: (2025)
From Output to Evaluation: Does Raw Instruction-Tuned Code LLMs Output Suffice for Fill-in-the-Middle Code Generation?
par: Ahmad, Wasi Uddin, et autres
Publié: (2025)
par: Ahmad, Wasi Uddin, et autres
Publié: (2025)
BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity
par: Diddee, Harshita, et autres
Publié: (2026)
par: Diddee, Harshita, et autres
Publié: (2026)
CommitBench: A Benchmark for Commit Message Generation
par: Schall, Maximilian, et autres
Publié: (2024)
par: Schall, Maximilian, et autres
Publié: (2024)
CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation
par: Wang, Sizhe, et autres
Publié: (2025)
par: Wang, Sizhe, et autres
Publié: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
par: Jiang, Hongchao, et autres
Publié: (2025)
par: Jiang, Hongchao, et autres
Publié: (2025)
Documents similaires
-
Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation
par: Pysklo, Hubert M., et autres
Publié: (2026) -
ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents
par: Jeon, YoungHoon, et autres
Publié: (2026) -
A Conceptual Framework for API Refactoring in Enterprise Application Architectures
par: Montesi, Fabrizio, et autres
Publié: (2024) -
ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
par: Zhang, Chenchen, et autres
Publié: (2025) -
AutoMonitor-Bench: Evaluating the Reliability of LLM-Based Misbehavior Monitor
par: Yang, Shu, et autres
Publié: (2026)