Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ahmed, Ammar, Khan, Azal Ahmad, Ahmad, Ayaan, Di, Sheng, Liu, Zirui, Anwar, Ali
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913108773044224
author Ahmed, Ammar
Khan, Azal Ahmad
Ahmad, Ayaan
Di, Sheng
Liu, Zirui
Anwar, Ali
author_facet Ahmed, Ammar
Khan, Azal Ahmad
Ahmad, Ayaan
Di, Sheng
Liu, Zirui
Anwar, Ali
contents Large reasoning models improve accuracy by producing long reasoning traces, but this inflates latency and cost, motivating inference-time efficiency. We propose Retrieval-of-Thought (RoT), which reuses prior reasoning as composable ``thought" steps to guide new problems. RoT organizes steps into a thought graph with sequential and semantic edges to enable fast retrieval and flexible recombination. At inference, RoT retrieves query-relevant nodes and applies reward-guided traversal to assemble a problem-specific template that guides generation. This dynamic template reuse reduces redundant exploration and, therefore, reduces output tokens while preserving accuracy. We evaluate RoT on reasoning benchmarks with multiple models, measuring accuracy, token usage, latency, and memory overhead. Findings show small prompt growth but substantial efficiency gains, with RoT reducing output tokens by up to 40%, inference latency by 82%, and cost by 59% while maintaining accuracy. RoT establishes a scalable paradigm for efficient LRM reasoning via dynamic template construction through retrieval.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21743
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts
Ahmed, Ammar
Khan, Azal Ahmad
Ahmad, Ayaan
Di, Sheng
Liu, Zirui
Anwar, Ali
Artificial Intelligence
Machine Learning
Large reasoning models improve accuracy by producing long reasoning traces, but this inflates latency and cost, motivating inference-time efficiency. We propose Retrieval-of-Thought (RoT), which reuses prior reasoning as composable ``thought" steps to guide new problems. RoT organizes steps into a thought graph with sequential and semantic edges to enable fast retrieval and flexible recombination. At inference, RoT retrieves query-relevant nodes and applies reward-guided traversal to assemble a problem-specific template that guides generation. This dynamic template reuse reduces redundant exploration and, therefore, reduces output tokens while preserving accuracy. We evaluate RoT on reasoning benchmarks with multiple models, measuring accuracy, token usage, latency, and memory overhead. Findings show small prompt growth but substantial efficiency gains, with RoT reducing output tokens by up to 40%, inference latency by 82%, and cost by 59% while maintaining accuracy. RoT establishes a scalable paradigm for efficient LRM reasoning via dynamic template construction through retrieval.
title Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.21743