Two Heads are Better than One: Distilling Large Language Model Features Into Small Models with Feature Decomposition and Mixture

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fu, Tianhao, Xu, Xinxin, Xu, Weichen, Chen, Jue, Ren, Ruilong, Deng, Bowen, Zhao, Xinyu, Cao, Jian, Cao, Xixin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914290735251456
author Fu, Tianhao
Xu, Xinxin
Xu, Weichen
Chen, Jue
Ren, Ruilong
Deng, Bowen
Zhao, Xinyu
Cao, Jian
Cao, Xixin
author_facet Fu, Tianhao
Xu, Xinxin
Xu, Weichen
Chen, Jue
Ren, Ruilong
Deng, Bowen
Zhao, Xinyu
Cao, Jian
Cao, Xixin
contents Market making (MM) through Reinforcement Learning (RL) has attracted significant attention in financial trading. With the development of Large Language Models (LLMs), more and more attempts are being made to apply LLMs to financial areas. A simple, direct application of LLM as an agent shows significant performance. Such methods are hindered by their slow inference speed, while most of the current research has not studied LLM distillation for this specific task. To address this, we first propose the normalized fluorescent probe to study the mechanism of the LLM's feature. Based on the observation found by our investigation, we propose Cooperative Market Making (CMM), a novel framework that decouples LLM features across three orthogonal dimensions: layer, task, and data. Various student models collaboratively learn simple LLM features along with different dimensions, with each model responsible for a distinct feature to achieve knowledge distillation. Furthermore, CMM introduces an Hájek-MoE to integrate the output of the student models by investigating the contribution of different models in a kernel function-generated common feature space. Extensive experimental results on four real-world market datasets demonstrate the superiority of CMM over the current distillation method and RL-based market-making strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2511_07110
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Two Heads are Better than One: Distilling Large Language Model Features Into Small Models with Feature Decomposition and Mixture
Fu, Tianhao
Xu, Xinxin
Xu, Weichen
Chen, Jue
Ren, Ruilong
Deng, Bowen
Zhao, Xinyu
Cao, Jian
Cao, Xixin
Artificial Intelligence
Market making (MM) through Reinforcement Learning (RL) has attracted significant attention in financial trading. With the development of Large Language Models (LLMs), more and more attempts are being made to apply LLMs to financial areas. A simple, direct application of LLM as an agent shows significant performance. Such methods are hindered by their slow inference speed, while most of the current research has not studied LLM distillation for this specific task. To address this, we first propose the normalized fluorescent probe to study the mechanism of the LLM's feature. Based on the observation found by our investigation, we propose Cooperative Market Making (CMM), a novel framework that decouples LLM features across three orthogonal dimensions: layer, task, and data. Various student models collaboratively learn simple LLM features along with different dimensions, with each model responsible for a distinct feature to achieve knowledge distillation. Furthermore, CMM introduces an Hájek-MoE to integrate the output of the student models by investigating the contribution of different models in a kernel function-generated common feature space. Extensive experimental results on four real-world market datasets demonstrate the superiority of CMM over the current distillation method and RL-based market-making strategies.
title Two Heads are Better than One: Distilling Large Language Model Features Into Small Models with Feature Decomposition and Mixture
topic Artificial Intelligence
url https://arxiv.org/abs/2511.07110