T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Qinsi, Ye, Hancheng, Kim, Jinhee, Ke, Jinghan, Wang, Yifei, Kuo, Martin, Shao, Zishan, Li, Dongting, Lin, Yueqian, Jiang, Ting, Wei, Chiyue, Qian, Qi, Wen, Wei, Li, Helen, Chen, Yiran
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918368927285248
author Wang, Qinsi
Ye, Hancheng
Kim, Jinhee
Ke, Jinghan
Wang, Yifei
Kuo, Martin
Shao, Zishan
Li, Dongting
Lin, Yueqian
Jiang, Ting
Wei, Chiyue
Qian, Qi
Wen, Wei
Li, Helen
Chen, Yiran
author_facet Wang, Qinsi
Ye, Hancheng
Kim, Jinhee
Ke, Jinghan
Wang, Yifei
Kuo, Martin
Shao, Zishan
Li, Dongting
Lin, Yueqian
Jiang, Ting
Wei, Chiyue
Qian, Qi
Wen, Wei
Li, Helen
Chen, Yiran
contents Think about how human handles complex reading tasks: marking key points, inferring their relationships, and structuring information to guide understanding and responses. Likewise, can a large language model benefit from text structure to enhance text-processing performance? To explore it, in this work, we first introduce Structure of Thought (SoT), a prompting technique that explicitly guides models to construct intermediate text structures, consistently boosting performance across eight tasks and three model families. Building upon this insight, we present T2S-Bench, the first benchmark designed to evaluate and improve text-to-structure capabilities of models. T2S-Bench includes 1.8K samples across 6 scientific domains and 32 structural types, rigorously constructed to ensure accuracy, fairness, and quality. Evaluation on 45 mainstream models reveals substantial improvement potential: the average accuracy on the multi-hop reasoning task is only 52.1%, and even the most advanced model achieves 58.1% node accuracy in end-to-end extraction. Furthermore, on Qwen2.5-7B-Instruct, SoT alone yields an average +5.7% improvement across eight diverse text-processing tasks, and fine-tuning on T2S-Bench further increases this gain to +8.6%. These results highlight the value of explicit text structuring and the complementary contributions of SoT and T2S-Bench. Dataset and eval code have been released at https://t2s-bench.github.io/T2S-Bench-Page/.
format Preprint
id arxiv_https___arxiv_org_abs_2603_03790
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasoning
Wang, Qinsi
Ye, Hancheng
Kim, Jinhee
Ke, Jinghan
Wang, Yifei
Kuo, Martin
Shao, Zishan
Li, Dongting
Lin, Yueqian
Jiang, Ting
Wei, Chiyue
Qian, Qi
Wen, Wei
Li, Helen
Chen, Yiran
Computation and Language
Artificial Intelligence
Think about how human handles complex reading tasks: marking key points, inferring their relationships, and structuring information to guide understanding and responses. Likewise, can a large language model benefit from text structure to enhance text-processing performance? To explore it, in this work, we first introduce Structure of Thought (SoT), a prompting technique that explicitly guides models to construct intermediate text structures, consistently boosting performance across eight tasks and three model families. Building upon this insight, we present T2S-Bench, the first benchmark designed to evaluate and improve text-to-structure capabilities of models. T2S-Bench includes 1.8K samples across 6 scientific domains and 32 structural types, rigorously constructed to ensure accuracy, fairness, and quality. Evaluation on 45 mainstream models reveals substantial improvement potential: the average accuracy on the multi-hop reasoning task is only 52.1%, and even the most advanced model achieves 58.1% node accuracy in end-to-end extraction. Furthermore, on Qwen2.5-7B-Instruct, SoT alone yields an average +5.7% improvement across eight diverse text-processing tasks, and fine-tuning on T2S-Bench further increases this gain to +8.6%. These results highlight the value of explicit text structuring and the complementary contributions of SoT and T2S-Bench. Dataset and eval code have been released at https://t2s-bench.github.io/T2S-Bench-Page/.
title T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasoning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2603.03790