AppBench: Planning of Multiple APIs from Various APPs for Complex User Instruction
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Hongru, Wang, Rui, Xue, Boyang, Xia, Heming, Cao, Jingtao, Liu, Zeming, Pan, Jeff Z., Wong, Kam-Fai |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UniRetriever: Multi-task Candidates Selection for Various Context-Adaptive Conversational Retrieval
by: Wang, Hongru, et al.
Published: (2024)
by: Wang, Hongru, et al.
Published: (2024)
Enhancing Large Language Models Against Inductive Instructions with Dual-critique Prompting
by: Wang, Rui, et al.
Published: (2023)
by: Wang, Rui, et al.
Published: (2023)
VLEU: a Method for Automatic Evaluation for Generalizability of Text-to-Image Models
by: Cao, Jingtao, et al.
Published: (2024)
by: Cao, Jingtao, et al.
Published: (2024)
Self-Reasoning Language Models: Unfold Hidden Reasoning Chains with Few Reasoning Catalyst
by: Wang, Hongru, et al.
Published: (2025)
by: Wang, Hongru, et al.
Published: (2025)
OSPC: Detecting Harmful Memes with Large Language Model as a Catalyst
by: Cao, Jingtao, et al.
Published: (2024)
by: Cao, Jingtao, et al.
Published: (2024)
MlingConf: A Comprehensive Study of Multilingual Confidence Estimation on Large Language Models
by: Xue, Boyang, et al.
Published: (2024)
by: Xue, Boyang, et al.
Published: (2024)
Role Prompting Guided Domain Adaptation with General Capability Preserve for Large Language Models
by: Wang, Rui, et al.
Published: (2024)
by: Wang, Rui, et al.
Published: (2024)
Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models
by: Wang, Rui, et al.
Published: (2025)
by: Wang, Rui, et al.
Published: (2025)
MlingConf: A Comprehensive Study of Multilingual Confidence Estimation on Large Language Models
by: Xue, Boyang, et al.
Published: (2024)
by: Xue, Boyang, et al.
Published: (2024)
UAlign: Leveraging Uncertainty Estimations for Factuality Alignment on Large Language Models
by: Xue, Boyang, et al.
Published: (2024)
by: Xue, Boyang, et al.
Published: (2024)
Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges
by: Wang, Hongru, et al.
Published: (2025)
by: Wang, Hongru, et al.
Published: (2025)
UniMS-RAG: A Unified Multi-source Retrieval-Augmented Generation for Personalized Dialogue Systems
by: Wang, Hongru, et al.
Published: (2024)
by: Wang, Hongru, et al.
Published: (2024)
Position: Agent Should Invoke External Tools ONLY When Epistemically Necessary
by: Wang, Hongru, et al.
Published: (2025)
by: Wang, Hongru, et al.
Published: (2025)
DAST: Difficulty-Aware Self-Training on Large Language Models
by: Xue, Boyang, et al.
Published: (2025)
by: Xue, Boyang, et al.
Published: (2025)
ReliableMath: Benchmark of Reliable Mathematical Reasoning on Large Language Models
by: Xue, Boyang, et al.
Published: (2025)
by: Xue, Boyang, et al.
Published: (2025)
SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs
by: Xia, Hongfei, et al.
Published: (2025)
by: Xia, Hongfei, et al.
Published: (2025)
Self-DC: When to Reason and When to Act? Self Divide-and-Conquer for Compositional Unknown Questions
by: Wang, Hongru, et al.
Published: (2024)
by: Wang, Hongru, et al.
Published: (2024)
A Survey of the Evolution of Language Model-Based Dialogue Systems: Data, Task and Models
by: Wang, Hongru, et al.
Published: (2023)
by: Wang, Hongru, et al.
Published: (2023)
HomeBench: Evaluating LLMs in Smart Homes with Valid and Invalid Instructions Across Single and Multiple Devices
by: Li, Silin, et al.
Published: (2025)
by: Li, Silin, et al.
Published: (2025)
TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments
by: Lu, Yuheng, et al.
Published: (2025)
by: Lu, Yuheng, et al.
Published: (2025)
WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation Models
by: Wang, Rui, et al.
Published: (2025)
by: Wang, Rui, et al.
Published: (2025)
Self-Guard: Empower the LLM to Safeguard Itself
by: Wang, Zezhong, et al.
Published: (2023)
by: Wang, Zezhong, et al.
Published: (2023)
PerLTQA: A Personal Long-Term Memory Dataset for Memory Classification, Retrieval, and Synthesis in Question Answering
by: Du, Yiming, et al.
Published: (2024)
by: Du, Yiming, et al.
Published: (2024)
Nonlinear optical analogues of quantum phase transitions in a squeezing-enhanced LMG model
by: Kam, Chon-Fai
Published: (2025)
by: Kam, Chon-Fai
Published: (2025)
Majorana Constellations: A Geometric Lens on Multipartite Entanglement and Geometric Phases
by: Kam, Chon-Fai
Published: (2026)
by: Kam, Chon-Fai
Published: (2026)
Three-Axis Spin Squeezed States Associated with Excited-State Quantum Phase Transitions
by: Kam, Chon-Fai
Published: (2025)
by: Kam, Chon-Fai
Published: (2025)
Nonlinear optical realization of non-integrable phases accompanying quantum phase transitions
by: Kam, Chon-Fai
Published: (2025)
by: Kam, Chon-Fai
Published: (2025)
Counterspeech for Mitigating the Influence of Media Bias: Comparing Human and LLM-Generated Responses
by: Lin, Luyang, et al.
Published: (2025)
by: Lin, Luyang, et al.
Published: (2025)
Investigating Bias in LLM-Based Bias Detection: Disparities between LLMs and Human Perception
by: Lin, Luyang, et al.
Published: (2024)
by: Lin, Luyang, et al.
Published: (2024)
Confused ChatGPT: Cross-App Context Poisoning via First-Party APIs
by: Wang, Chao, et al.
Published: (2026)
by: Wang, Chao, et al.
Published: (2026)
MobileSteward: Integrating Multiple App-Oriented Agents with Self-Evolution to Automate Cross-App Instructions
by: Liu, Yuxuan, et al.
Published: (2025)
by: Liu, Yuxuan, et al.
Published: (2025)
LLM+Reasoning+Planning for Supporting Incomplete User Queries in Presence of APIs
by: Agarwal, Sudhir, et al.
Published: (2024)
by: Agarwal, Sudhir, et al.
Published: (2024)
Rethinking Memory in LLM based Agents: Representations, Operations, and Emerging Topics
by: Du, Yiming, et al.
Published: (2025)
by: Du, Yiming, et al.
Published: (2025)
Memory-T1: Reinforcement Learning for Temporal Reasoning in Multi-session Agents
by: Du, Yiming, et al.
Published: (2025)
by: Du, Yiming, et al.
Published: (2025)
Terminal-World: Scaling Terminal-Agent Environments via Agent Skills
by: Cheng, Zihao, et al.
Published: (2026)
by: Cheng, Zihao, et al.
Published: (2026)
MVISU-Bench: Benchmarking Mobile Agents for Real-World Tasks by Multi-App, Vague, Interactive, Single-App and Unethical Instructions
by: Huang, Zeyu, et al.
Published: (2025)
by: Huang, Zeyu, et al.
Published: (2025)
SeRTS: Self-Rewarding Tree Search for Biomedical Retrieval-Augmented Generation
by: Hu, Minda, et al.
Published: (2024)
by: Hu, Minda, et al.
Published: (2024)
Acting Less is Reasoning More! Teaching Model to Act Efficiently
by: Wang, Hongru, et al.
Published: (2025)
by: Wang, Hongru, et al.
Published: (2025)
Selective Forgetting: Advancing Machine Unlearning Techniques and Evaluation in Language Models
by: Wang, Lingzhi, et al.
Published: (2024)
by: Wang, Lingzhi, et al.
Published: (2024)
IndiTag: An Online Media Bias Analysis System Using Fine-Grained Bias Indicators
by: Lin, Luyang, et al.
Published: (2024)
by: Lin, Luyang, et al.
Published: (2024)
Similar Items
-
UniRetriever: Multi-task Candidates Selection for Various Context-Adaptive Conversational Retrieval
by: Wang, Hongru, et al.
Published: (2024) -
Enhancing Large Language Models Against Inductive Instructions with Dual-critique Prompting
by: Wang, Rui, et al.
Published: (2023) -
VLEU: a Method for Automatic Evaluation for Generalizability of Text-to-Image Models
by: Cao, Jingtao, et al.
Published: (2024) -
Self-Reasoning Language Models: Unfold Hidden Reasoning Chains with Few Reasoning Catalyst
by: Wang, Hongru, et al.
Published: (2025) -
OSPC: Detecting Harmful Memes with Large Language Model as a Catalyst
by: Cao, Jingtao, et al.
Published: (2024)