ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yuxiang, Chen, Jing, Wang, Junjie, Liu, Yaxin, Yang, Cheng, Shi, Chufan, Zhu, Xinyu, Lin, Zihao, Wan, Hanwen, Yang, Yujiu, Sakai, Tetsuya, Feng, Tian, Yamana, Hayato |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Data-Efficient Massive Tool Retrieval: A Reinforcement Learning Approach for Query-Tool Alignment with Language Models
by: Zhang, Yuxiang, et al.
Published: (2024)
by: Zhang, Yuxiang, et al.
Published: (2024)
HoLLMwood: Unleashing the Creativity of Large Language Models in Screenwriting via Role Playing
by: Chen, Jing, et al.
Published: (2024)
by: Chen, Jing, et al.
Published: (2024)
LiFi: Lightweight Controlled Text Generation with Fine-Grained Control Codes
by: Shi, Chufan, et al.
Published: (2024)
by: Shi, Chufan, et al.
Published: (2024)
ContextVis: Envision Contextual Learning and Interaction with Generative Models
by: Shui, Bo, et al.
Published: (2024)
by: Shui, Bo, et al.
Published: (2024)
ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation
by: Yang, Cheng, et al.
Published: (2024)
by: Yang, Cheng, et al.
Published: (2024)
AgroTools: A Benchmark for Tool-Augmented Multimodal Agents in Agriculture
by: Ye, Zi, et al.
Published: (2026)
by: Ye, Zi, et al.
Published: (2026)
News Recommendation with Category Description by a Large Language Model
by: Yada, Yuki, et al.
Published: (2024)
by: Yada, Yuki, et al.
Published: (2024)
Fully-Optimized Quantum Metrology: Framework, Tools, and Applications
by: Liu, Qiushi, et al.
Published: (2024)
by: Liu, Qiushi, et al.
Published: (2024)
Fully‐Optimized Quantum Metrology: Framework, Tools, and Applications
by: Qiushi Liu, et al.
Published: (2024)
by: Qiushi Liu, et al.
Published: (2024)
LLM2: Let Large Language Models Harness System 2 Reasoning
by: Yang, Cheng, et al.
Published: (2024)
by: Yang, Cheng, et al.
Published: (2024)
MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decoding
by: Deng, Jingyuan, et al.
Published: (2025)
by: Deng, Jingyuan, et al.
Published: (2025)
Solving Math Word Problems via Cooperative Reasoning induced Language Models
by: Zhu, Xinyu, et al.
Published: (2022)
by: Zhu, Xinyu, et al.
Published: (2022)
Unchosen Experts Can Contribute Too: Unleashing MoE Models' Power by Self-Contrast
by: Shi, Chufan, et al.
Published: (2024)
by: Shi, Chufan, et al.
Published: (2024)
Benchmarking LLM Tool-Use in the Wild
by: Yu, Peijie, et al.
Published: (2026)
by: Yu, Peijie, et al.
Published: (2026)
A Thorough Examination of Decoding Methods in the Era of LLMs
by: Shi, Chufan, et al.
Published: (2024)
by: Shi, Chufan, et al.
Published: (2024)
OpenLVLM-MIA: A Controlled Benchmark Revealing the Limits of Membership Inference Attacks on Large Vision-Language Models
by: Miyamoto, Ryoto, et al.
Published: (2025)
by: Miyamoto, Ryoto, et al.
Published: (2025)
ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints
by: Choi, Hyeonje, et al.
Published: (2026)
by: Choi, Hyeonje, et al.
Published: (2026)
SciToolAgent: A Knowledge Graph-Driven Scientific Agent for Multi-Tool Integration
by: Ding, Keyan, et al.
Published: (2025)
by: Ding, Keyan, et al.
Published: (2025)
Tool-MCoT: Tool Augmented Multimodal Chain-of-Thought for Content Safety Moderation
by: Zhang, Shutong, et al.
Published: (2026)
by: Zhang, Shutong, et al.
Published: (2026)
Internal Representations as Indicators of Hallucinations in Agent Tool Selection
by: Healy, Kait, et al.
Published: (2026)
by: Healy, Kait, et al.
Published: (2026)
TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object Understanding
by: Liu, Yun, et al.
Published: (2024)
by: Liu, Yun, et al.
Published: (2024)
Benchmarking Failures in Tool-Augmented Language Models
by: Treviño, Eduardo, et al.
Published: (2025)
by: Treviño, Eduardo, et al.
Published: (2025)
GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces
by: Geng, Xinyu, et al.
Published: (2026)
by: Geng, Xinyu, et al.
Published: (2026)
Mitigating the Reasoning Tax in Vision-Language Fine-Tuning with Input-Adaptive Depth Aggregation
by: Ren, Yiming, et al.
Published: (2026)
by: Ren, Yiming, et al.
Published: (2026)
InsCL: A Data-efficient Continual Learning Paradigm for Fine-tuning Large Language Models with Instructions
by: Wang, Yifan, et al.
Published: (2024)
by: Wang, Yifan, et al.
Published: (2024)
Hint-enhanced In-Context Learning wakes Large Language Models up for knowledge-intensive tasks
by: Wang, Yifan, et al.
Published: (2023)
by: Wang, Yifan, et al.
Published: (2023)
BeHonest: Benchmarking Honesty in Large Language Models
by: Chern, Steffi, et al.
Published: (2024)
by: Chern, Steffi, et al.
Published: (2024)
Personalized Fashion Recommendation with Image Attributes and Aesthetics Assessment
by: Chen, Chongxian, et al.
Published: (2025)
by: Chen, Chongxian, et al.
Published: (2025)
Why is the User Interface a Dark Pattern? : Explainable Auto-Detection and its Analysis
by: Yada, Yuki, et al.
Published: (2023)
by: Yada, Yuki, et al.
Published: (2023)
Tool-Augmented Reward Modeling
by: Li, Lei, et al.
Published: (2023)
by: Li, Lei, et al.
Published: (2023)
Pass@k Metric for RLVR: A Diagnostic Tool of Exploration, But Not an Objective
by: Yu, Yang
Published: (2025)
by: Yu, Yang
Published: (2025)
Tool Unlearning for Tool-Augmented LLMs
by: Cheng, Jiali, et al.
Published: (2025)
by: Cheng, Jiali, et al.
Published: (2025)
HonestLLM: Toward an Honest and Helpful Large Language Model
by: Gao, Chujie, et al.
Published: (2024)
by: Gao, Chujie, et al.
Published: (2024)
Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
by: Gao, Xin, et al.
Published: (2026)
by: Gao, Xin, et al.
Published: (2026)
GeoAgentBench: A Dynamic Execution Benchmark for Tool-Augmented Agents in Spatial Analysis
by: Yu, Bo, et al.
Published: (2026)
by: Yu, Bo, et al.
Published: (2026)
ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs
by: Kokane, Shirley, et al.
Published: (2024)
by: Kokane, Shirley, et al.
Published: (2024)
MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
by: Huang, Yue, et al.
Published: (2023)
by: Huang, Yue, et al.
Published: (2023)
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
by: Gou, Zhibin, et al.
Published: (2023)
by: Gou, Zhibin, et al.
Published: (2023)
BrowseMaster: Towards Scalable Web Browsing via Tool-Augmented Programmatic Agent Pair
by: Pang, Xianghe, et al.
Published: (2025)
by: Pang, Xianghe, et al.
Published: (2025)
FamilyTool: A Multi-hop Personalized Tool Use Benchmark
by: Wang, Yuxin, et al.
Published: (2025)
by: Wang, Yuxin, et al.
Published: (2025)
Similar Items
-
Data-Efficient Massive Tool Retrieval: A Reinforcement Learning Approach for Query-Tool Alignment with Language Models
by: Zhang, Yuxiang, et al.
Published: (2024) -
HoLLMwood: Unleashing the Creativity of Large Language Models in Screenwriting via Role Playing
by: Chen, Jing, et al.
Published: (2024) -
LiFi: Lightweight Controlled Text Generation with Fine-Grained Control Codes
by: Shi, Chufan, et al.
Published: (2024) -
ContextVis: Envision Contextual Learning and Interaction with Generative Models
by: Shui, Bo, et al.
Published: (2024) -
ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation
by: Yang, Cheng, et al.
Published: (2024)