HSCodeComp: A Realistic and Expert-level Benchmark for Deep Search Agents in Hierarchical Rule Application

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Yiqian, Lan, Tian, Jia, Qianghuai, Zhu, Li, Jiang, Hui, Zhu, Hang, Wang, Longyue, Luo, Weihua, Zhang, Kaifu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918165744713728
author Yang, Yiqian
Lan, Tian
Jia, Qianghuai
Zhu, Li
Jiang, Hui
Zhu, Hang
Wang, Longyue
Luo, Weihua
Zhang, Kaifu
author_facet Yang, Yiqian
Lan, Tian
Jia, Qianghuai
Zhu, Li
Jiang, Hui
Zhu, Hang
Wang, Longyue
Luo, Weihua
Zhang, Kaifu
contents Effective deep search agents must not only access open-domain and domain-specific knowledge but also apply complex rules-such as legal clauses, medical manuals and tariff rules. These rules often feature vague boundaries and implicit logic relationships, making precise application challenging for agents. However, this critical capability is largely overlooked by current agent benchmarks. To fill this gap, we introduce HSCodeComp, the first realistic, expert-level e-commerce benchmark designed to evaluate deep search agents in hierarchical rule application. In this task, the deep reasoning process of agents is guided by these rules to predict 10-digit Harmonized System Code (HSCode) of products with noisy but realistic descriptions. These codes, established by the World Customs Organization, are vital for global supply chain efficiency. Built from real-world data collected from large-scale e-commerce platforms, our proposed HSCodeComp comprises 632 product entries spanning diverse product categories, with these HSCodes annotated by several human experts. Extensive experimental results on several state-of-the-art LLMs, open-source, and closed-source agents reveal a huge performance gap: best agent achieves only 46.8% 10-digit accuracy, far below human experts at 95.0%. Besides, detailed analysis demonstrates the challenges of hierarchical rule application, and test-time scaling fails to improve performance further.
format Preprint
id arxiv_https___arxiv_org_abs_2510_19631
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HSCodeComp: A Realistic and Expert-level Benchmark for Deep Search Agents in Hierarchical Rule Application
Yang, Yiqian
Lan, Tian
Jia, Qianghuai
Zhu, Li
Jiang, Hui
Zhu, Hang
Wang, Longyue
Luo, Weihua
Zhang, Kaifu
Artificial Intelligence
Computation and Language
Multiagent Systems
Effective deep search agents must not only access open-domain and domain-specific knowledge but also apply complex rules-such as legal clauses, medical manuals and tariff rules. These rules often feature vague boundaries and implicit logic relationships, making precise application challenging for agents. However, this critical capability is largely overlooked by current agent benchmarks. To fill this gap, we introduce HSCodeComp, the first realistic, expert-level e-commerce benchmark designed to evaluate deep search agents in hierarchical rule application. In this task, the deep reasoning process of agents is guided by these rules to predict 10-digit Harmonized System Code (HSCode) of products with noisy but realistic descriptions. These codes, established by the World Customs Organization, are vital for global supply chain efficiency. Built from real-world data collected from large-scale e-commerce platforms, our proposed HSCodeComp comprises 632 product entries spanning diverse product categories, with these HSCodes annotated by several human experts. Extensive experimental results on several state-of-the-art LLMs, open-source, and closed-source agents reveal a huge performance gap: best agent achieves only 46.8% 10-digit accuracy, far below human experts at 95.0%. Besides, detailed analysis demonstrates the challenges of hierarchical rule application, and test-time scaling fails to improve performance further.
title HSCodeComp: A Realistic and Expert-level Benchmark for Deep Search Agents in Hierarchical Rule Application
topic Artificial Intelligence
Computation and Language
Multiagent Systems
url https://arxiv.org/abs/2510.19631