Reflection Pretraining Enables Token-Level Self-Correction in Biological Sequence Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Xiang, Wei, Jiaqi, Yang, Yuejin, Qiu, Zijie, Chen, Yuhan, Gao, Zhiqiang, Abdul-Mageed, Muhammad, Lakshmanan, Laks V. S., Ouyang, Wanli, You, Chenyu, Sun, Siqi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912787791347712
author Zhang, Xiang
Wei, Jiaqi
Yang, Yuejin
Qiu, Zijie
Chen, Yuhan
Gao, Zhiqiang
Abdul-Mageed, Muhammad
Lakshmanan, Laks V. S.
Ouyang, Wanli
You, Chenyu
Sun, Siqi
author_facet Zhang, Xiang
Wei, Jiaqi
Yang, Yuejin
Qiu, Zijie
Chen, Yuhan
Gao, Zhiqiang
Abdul-Mageed, Muhammad
Lakshmanan, Laks V. S.
Ouyang, Wanli
You, Chenyu
Sun, Siqi
contents Chain-of-Thought (CoT) prompting has significantly advanced task-solving capabilities in natural language processing with large language models. Unlike standard prompting, CoT encourages the model to generate intermediate reasoning steps, non-answer tokens, that help guide the model toward more accurate final outputs. These intermediate steps enable more complex reasoning processes such as error correction, memory management, future planning, and self-reflection. However, applying CoT to non-natural language domains, such as protein and RNA language models, is not yet possible, primarily due to the limited expressiveness of their token spaces (e.g., amino acid tokens). In this work, we propose and define the concept of language expressiveness: the ability of a given language, using its tokens and grammar, to encode information. We show that the limited expressiveness of protein language severely restricts the applicability of CoT-style reasoning. To overcome this, we introduce reflection pretraining, for the first time in a biological sequence model, which enables the model to engage in intermediate reasoning through the generation of auxiliary "thinking tokens" beyond simple answer tokens. Theoretically, we demonstrate that our augmented token set significantly enhances biological language expressiveness, thereby improving the overall reasoning capacity of the model. Experimentally, our pretraining approach teaches protein models to self-correct and leads to substantial performance gains compared to standard pretraining.
format Preprint
id arxiv_https___arxiv_org_abs_2512_20954
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reflection Pretraining Enables Token-Level Self-Correction in Biological Sequence Models
Zhang, Xiang
Wei, Jiaqi
Yang, Yuejin
Qiu, Zijie
Chen, Yuhan
Gao, Zhiqiang
Abdul-Mageed, Muhammad
Lakshmanan, Laks V. S.
Ouyang, Wanli
You, Chenyu
Sun, Siqi
Computation and Language
Artificial Intelligence
Chain-of-Thought (CoT) prompting has significantly advanced task-solving capabilities in natural language processing with large language models. Unlike standard prompting, CoT encourages the model to generate intermediate reasoning steps, non-answer tokens, that help guide the model toward more accurate final outputs. These intermediate steps enable more complex reasoning processes such as error correction, memory management, future planning, and self-reflection. However, applying CoT to non-natural language domains, such as protein and RNA language models, is not yet possible, primarily due to the limited expressiveness of their token spaces (e.g., amino acid tokens). In this work, we propose and define the concept of language expressiveness: the ability of a given language, using its tokens and grammar, to encode information. We show that the limited expressiveness of protein language severely restricts the applicability of CoT-style reasoning. To overcome this, we introduce reflection pretraining, for the first time in a biological sequence model, which enables the model to engage in intermediate reasoning through the generation of auxiliary "thinking tokens" beyond simple answer tokens. Theoretically, we demonstrate that our augmented token set significantly enhances biological language expressiveness, thereby improving the overall reasoning capacity of the model. Experimentally, our pretraining approach teaches protein models to self-correct and leads to substantial performance gains compared to standard pretraining.
title Reflection Pretraining Enables Token-Level Self-Correction in Biological Sequence Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2512.20954