DFPE: A Diverse Fingerprint Ensemble for Enhancing LLM Performance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cohen, Seffi, Goldshlager, Niv, Cohen-Inger, Nurit, Shapira, Bracha, Rokach, Lior
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909481974104064
author Cohen, Seffi
Goldshlager, Niv
Cohen-Inger, Nurit
Shapira, Bracha
Rokach, Lior
author_facet Cohen, Seffi
Goldshlager, Niv
Cohen-Inger, Nurit
Shapira, Bracha
Rokach, Lior
contents Large Language Models (LLMs) have shown remarkable capabilities across various natural language processing tasks but often struggle to excel uniformly in diverse or complex domains. We propose a novel ensemble method - Diverse Fingerprint Ensemble (DFPE), which leverages the complementary strengths of multiple LLMs to achieve more robust performance. Our approach involves: (1) clustering models based on response "fingerprints" patterns, (2) applying a quantile-based filtering mechanism to remove underperforming models at a per-subject level, and (3) assigning adaptive weights to remaining models based on their subject-wise validation accuracy. In experiments on the Massive Multitask Language Understanding (MMLU) benchmark, DFPE outperforms the best single model by 3% overall accuracy and 5% in discipline-level accuracy. This method increases the robustness and generalization of LLMs and underscores how model selection, diversity preservation, and performance-driven weighting can effectively address challenging, multi-faceted language understanding tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2501_17479
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DFPE: A Diverse Fingerprint Ensemble for Enhancing LLM Performance
Cohen, Seffi
Goldshlager, Niv
Cohen-Inger, Nurit
Shapira, Bracha
Rokach, Lior
Machine Learning
Artificial Intelligence
Computation and Language
Large Language Models (LLMs) have shown remarkable capabilities across various natural language processing tasks but often struggle to excel uniformly in diverse or complex domains. We propose a novel ensemble method - Diverse Fingerprint Ensemble (DFPE), which leverages the complementary strengths of multiple LLMs to achieve more robust performance. Our approach involves: (1) clustering models based on response "fingerprints" patterns, (2) applying a quantile-based filtering mechanism to remove underperforming models at a per-subject level, and (3) assigning adaptive weights to remaining models based on their subject-wise validation accuracy. In experiments on the Massive Multitask Language Understanding (MMLU) benchmark, DFPE outperforms the best single model by 3% overall accuracy and 5% in discipline-level accuracy. This method increases the robustness and generalization of LLMs and underscores how model selection, diversity preservation, and performance-driven weighting can effectively address challenging, multi-faceted language understanding tasks.
title DFPE: A Diverse Fingerprint Ensemble for Enhancing LLM Performance
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2501.17479