Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Yuxuan, Wang, Ziyi, Lu, Yingzhou, Sang, Yisi, Gesi, Jiri, Tang, Xianfeng, Zhang, Yimeng, Dai, Zhenwei, Liu, Hui, Lu, Hanqing, Luo, Chen, He, Qi, Dumoulin, Benoit, Huang, Jing, Wang, Dakuo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Trajectory2Task: Training Robust Tool-Calling Agents with Synthesized Yet Verifiable Data for Complex User Intents
by: Wang, Ziyi, et al.
Published: (2026)
by: Wang, Ziyi, et al.
Published: (2026)
WEBSERV: A Full-Stack and RL-Ready Web Environment for Training Web Agents at Scale
by: Lu, Yuxuan, et al.
Published: (2025)
by: Lu, Yuxuan, et al.
Published: (2025)
Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior Data
by: Lu, Yuxuan, et al.
Published: (2025)
by: Lu, Yuxuan, et al.
Published: (2025)
Customer-R1: Personalized Simulation of Human Behaviors via RL-based LLM Agent in Online Shopping
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
Adaptive Auto-Harness: Sustained Self-Improvement for Agentic System Deployment on Open-Ended Task Streams
by: Liu, Zewen, et al.
Published: (2026)
by: Liu, Zewen, et al.
Published: (2026)
See, Think, Act: Online Shopper Behavior Simulation with VLM Agents
by: Zhang, Yimeng, et al.
Published: (2025)
by: Zhang, Yimeng, et al.
Published: (2025)
TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
by: He, Pengfei, et al.
Published: (2025)
by: He, Pengfei, et al.
Published: (2025)
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
by: Chen, Jiaju, et al.
Published: (2025)
by: Chen, Jiaju, et al.
Published: (2025)
LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
by: Sun, Lu, et al.
Published: (2025)
by: Sun, Lu, et al.
Published: (2025)
UXAgent: A System for Simulating Usability Testing of Web Design with LLM Agents
by: Lu, Yuxuan, et al.
Published: (2025)
by: Lu, Yuxuan, et al.
Published: (2025)
UXAgent: An LLM Agent-Based Usability Testing Framework for Web Design
by: Lu, Yuxuan, et al.
Published: (2025)
by: Lu, Yuxuan, et al.
Published: (2025)
Shop-R1: Rewarding LLMs to Simulate Human Behavior in Online Shopping via Reinforcement Learning
by: Zhang, Yimeng, et al.
Published: (2025)
by: Zhang, Yimeng, et al.
Published: (2025)
Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
by: Lin, Minhua, et al.
Published: (2026)
by: Lin, Minhua, et al.
Published: (2026)
Beyond Self-learned Attention: Mitigating Attention Bias in Transformer-based Models Using Attention Guidance
by: Gesi, Jiri, et al.
Published: (2024)
by: Gesi, Jiri, et al.
Published: (2024)
More Samples or More Prompts? Exploring Effective In-Context Sampling for LLM Few-Shot Prompt Engineering
by: Yao, Bingsheng, et al.
Published: (2023)
by: Yao, Bingsheng, et al.
Published: (2023)
Position: Agentic Evolution is the Path to Evolving LLMs
by: Lin, Minhua, et al.
Published: (2026)
by: Lin, Minhua, et al.
Published: (2026)
CDFCI: High-Performance Parallel Software for Many-Body Large-Scale Eigenvalue Problems
by: Zhang, Yuejia, et al.
Published: (2026)
by: Zhang, Yuejia, et al.
Published: (2026)
DPRF: A Generalizable Dynamic Persona Refinement Framework for Optimizing Behavior Alignment Between Personalized LLM Role-Playing Agents and Humans
by: Yao, Bingsheng, et al.
Published: (2025)
by: Yao, Bingsheng, et al.
Published: (2025)
Does More Advice Help? The Effects of Second Opinions in AI-Assisted Decision Making
by: Lu, Zhuoran, et al.
Published: (2024)
by: Lu, Zhuoran, et al.
Published: (2024)
Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
by: Liu, Zhining, et al.
Published: (2025)
by: Liu, Zhining, et al.
Published: (2025)
OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
PinChecker: Identifying Unsound Safe Abstractions of Rust Pinning APIs
by: Dai, Yuxuan, et al.
Published: (2025)
by: Dai, Yuxuan, et al.
Published: (2025)
Live API-Bench: 2500+ Live APIs for Testing Multi-Step Tool Calling
by: Elder, Benjamin, et al.
Published: (2025)
by: Elder, Benjamin, et al.
Published: (2025)
Landscape Analysis of Excited States Calculation over Quantum Computers
by: Chen, Hengzhun, et al.
Published: (2025)
by: Chen, Hengzhun, et al.
Published: (2025)
Adaptive Test-Time Reasoning via Reward-Guided Dual-Phase Search
by: Cui, Yingqian, et al.
Published: (2025)
by: Cui, Yingqian, et al.
Published: (2025)
Qubit Count Reduction by Orthogonally-Constrained Orbital Optimization for Variational Quantum Excited States Solvers
by: Bierman, Joel, et al.
Published: (2023)
by: Bierman, Joel, et al.
Published: (2023)
DrugCLIP: Contrastive Drug-Disease Interaction For Drug Repurposing
by: Lu, Yingzhou, et al.
Published: (2024)
by: Lu, Yingzhou, et al.
Published: (2024)
StableToolBench-MirrorAPI: Modeling Tool Environments as Mirrors of 7,000+ Real-World APIs
by: Guo, Zhicheng, et al.
Published: (2025)
by: Guo, Zhicheng, et al.
Published: (2025)
Alignment for Efficient Tool Calling of Large Language Models
by: Xu, Hongshen, et al.
Published: (2025)
by: Xu, Hongshen, et al.
Published: (2025)
StorySparkQA: Expert-Annotated QA Pairs with Real-World Knowledge for Children's Story-Based Learning
by: Chen, Jiaju, et al.
Published: (2023)
by: Chen, Jiaju, et al.
Published: (2023)
ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
Verifiable Generation with Subsentence-Level Fine-Grained Citations
by: Cao, Shuyang, et al.
Published: (2024)
by: Cao, Shuyang, et al.
Published: (2024)
Bernstein-von Mises Theorem for Sparse Generalized Linear Model
by: Li, Hanqing, et al.
Published: (2026)
by: Li, Hanqing, et al.
Published: (2026)
Towards Robustness Analysis of E-Commerce Ranking System
by: Wang, Ningfei, et al.
Published: (2024)
by: Wang, Ningfei, et al.
Published: (2024)
TransGI: Real-Time Dynamic Global Illumination With Object-Centric Neural Transfer Model
by: Deng, Yijie, et al.
Published: (2025)
by: Deng, Yijie, et al.
Published: (2025)
Firefly Parasites and Predators
by: Lloyd, James E
Published: (1973)
by: Lloyd, James E
Published: (1973)
LLM Assisted Alpha Fairness for 6 GHz WiFi and NR_U Coexistence: An Agentic Orchestrator for Throughput, Energy, and SLA
by: Wang, Qun, et al.
Published: (2025)
by: Wang, Qun, et al.
Published: (2025)
ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling
by: Lin, Jianghao, et al.
Published: (2025)
by: Lin, Jianghao, et al.
Published: (2025)
Quantum Circuit for Non-Unitary Linear Transformation of Basis Sets
by: Zhu, Guorui, et al.
Published: (2025)
by: Zhu, Guorui, et al.
Published: (2025)
State-Specific Orbital Optimization for Enhanced Excited-States Calculation on Quantum Computers
by: Zhu, Guorui, et al.
Published: (2025)
by: Zhu, Guorui, et al.
Published: (2025)
Similar Items
-
Trajectory2Task: Training Robust Tool-Calling Agents with Synthesized Yet Verifiable Data for Complex User Intents
by: Wang, Ziyi, et al.
Published: (2026) -
WEBSERV: A Full-Stack and RL-Ready Web Environment for Training Web Agents at Scale
by: Lu, Yuxuan, et al.
Published: (2025) -
Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior Data
by: Lu, Yuxuan, et al.
Published: (2025) -
Customer-R1: Personalized Simulation of Human Behaviors via RL-based LLM Agent in Online Shopping
by: Wang, Ziyi, et al.
Published: (2025) -
Adaptive Auto-Harness: Sustained Self-Improvement for Agentic System Deployment on Open-Ended Task Streams
by: Liu, Zewen, et al.
Published: (2026)