Filtered not Mixed: Stochastic Filtering-Based Online Gating for Mixture of Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Saqur, Raeid, Kratsios, Anastasis, Krach, Florian, Limmer, Yannick, Tian, Jacob-Junqi, Willes, John, Horvath, Blanka, Rudzicz, Frank
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929723115831296
author Saqur, Raeid
Kratsios, Anastasis
Krach, Florian
Limmer, Yannick
Tian, Jacob-Junqi
Willes, John
Horvath, Blanka
Rudzicz, Frank
author_facet Saqur, Raeid
Kratsios, Anastasis
Krach, Florian
Limmer, Yannick
Tian, Jacob-Junqi
Willes, John
Horvath, Blanka
Rudzicz, Frank
contents We propose MoE-F - a formalized mechanism for combining $N$ pre-trained Large Language Models (LLMs) for online time-series prediction by adaptively forecasting the best weighting of LLM predictions at every time step. Our mechanism leverages the conditional information in each expert's running performance to forecast the best combination of LLMs for predicting the time series in its next step. Diverging from static (learned) Mixture of Experts (MoE) methods, our approach employs time-adaptive stochastic filtering techniques to combine experts. By framing the expert selection problem as a finite state-space, continuous-time Hidden Markov model (HMM), we can leverage the Wohman-Shiryaev filter. Our approach first constructs N parallel filters corresponding to each of the $N$ individual LLMs. Each filter proposes its best combination of LLMs, given the information that they have access to. Subsequently, the N filter outputs are optimally aggregated to maximize their robust predictive power, and this update is computed efficiently via a closed-form expression, generating our ensemble predictor. Our contributions are: **(I)** the MoE-F plug-and-play filtering harness algorithm, **(II)** theoretical optimality guarantees of the proposed filtering-based gating algorithm (via optimality guarantees for its parallel Bayesian filtering and its robust aggregation steps), and **(III)** empirical evaluation and ablative results using state-of-the-art foundational and MoE LLMs on a real-world __Financial Market Movement__ task where MoE-F attains a remarkable 17\% absolute and 48.5\% relative F1 measure improvement over the next best performing individual LLM expert predicting short-horizon market movement based on streaming news. Further, we provide empirical evidence of substantial performance gains in applying MoE-F over specialized models in the long-horizon time-series forecasting domain.
format Preprint
id arxiv_https___arxiv_org_abs_2406_02969
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Filtered not Mixed: Stochastic Filtering-Based Online Gating for Mixture of Large Language Models
Saqur, Raeid
Kratsios, Anastasis
Krach, Florian
Limmer, Yannick
Tian, Jacob-Junqi
Willes, John
Horvath, Blanka
Rudzicz, Frank
Machine Learning
Artificial Intelligence
Computation and Language
Computational Finance
Mathematical Finance
60J05, 60G35, 68T20, 68T42, 68T50
I.2.6; I.2.7; G.3
We propose MoE-F - a formalized mechanism for combining $N$ pre-trained Large Language Models (LLMs) for online time-series prediction by adaptively forecasting the best weighting of LLM predictions at every time step. Our mechanism leverages the conditional information in each expert's running performance to forecast the best combination of LLMs for predicting the time series in its next step. Diverging from static (learned) Mixture of Experts (MoE) methods, our approach employs time-adaptive stochastic filtering techniques to combine experts. By framing the expert selection problem as a finite state-space, continuous-time Hidden Markov model (HMM), we can leverage the Wohman-Shiryaev filter. Our approach first constructs N parallel filters corresponding to each of the $N$ individual LLMs. Each filter proposes its best combination of LLMs, given the information that they have access to. Subsequently, the N filter outputs are optimally aggregated to maximize their robust predictive power, and this update is computed efficiently via a closed-form expression, generating our ensemble predictor. Our contributions are: **(I)** the MoE-F plug-and-play filtering harness algorithm, **(II)** theoretical optimality guarantees of the proposed filtering-based gating algorithm (via optimality guarantees for its parallel Bayesian filtering and its robust aggregation steps), and **(III)** empirical evaluation and ablative results using state-of-the-art foundational and MoE LLMs on a real-world __Financial Market Movement__ task where MoE-F attains a remarkable 17\% absolute and 48.5\% relative F1 measure improvement over the next best performing individual LLM expert predicting short-horizon market movement based on streaming news. Further, we provide empirical evidence of substantial performance gains in applying MoE-F over specialized models in the long-horizon time-series forecasting domain.
title Filtered not Mixed: Stochastic Filtering-Based Online Gating for Mixture of Large Language Models
topic Machine Learning
Artificial Intelligence
Computation and Language
Computational Finance
Mathematical Finance
60J05, 60G35, 68T20, 68T42, 68T50
I.2.6; I.2.7; G.3
url https://arxiv.org/abs/2406.02969