When Are Two Scores Better Than One? Investigating Ensembles of Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Razafindralambo, Raphaël, Sun, Rémy, Precioso, Frédéric, Garreau, Damien, Mattei, Pierre-Alexandre
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917213913481216
author Razafindralambo, Raphaël
Sun, Rémy
Precioso, Frédéric
Garreau, Damien
Mattei, Pierre-Alexandre
author_facet Razafindralambo, Raphaël
Sun, Rémy
Precioso, Frédéric
Garreau, Damien
Mattei, Pierre-Alexandre
contents Diffusion models now generate high-quality, diverse samples, with an increasing focus on more powerful models. Although ensembling is a well-known way to improve supervised models, its application to unconditional score-based diffusion models remains largely unexplored. In this work we investigate whether it provides tangible benefits for generative modelling. We find that while ensembling the scores generally improves the score-matching loss and model likelihood, it fails to consistently enhance perceptual quality metrics such as FID on image datasets. We confirm this observation across a breadth of aggregation rules using Deep Ensembles, Monte Carlo Dropout, on CIFAR-10 and FFHQ. We attempt to explain this discrepancy by investigating possible explanations, such as the link between score estimation and image quality. We also look into tabular data through random forests, and find that one aggregation strategy outperforms the others. Finally, we provide theoretical insights into the summing of score models, which shed light not only on ensembling but also on several model composition techniques (e.g. guidance).
format Preprint
id arxiv_https___arxiv_org_abs_2601_11444
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When Are Two Scores Better Than One? Investigating Ensembles of Diffusion Models
Razafindralambo, Raphaël
Sun, Rémy
Precioso, Frédéric
Garreau, Damien
Mattei, Pierre-Alexandre
Machine Learning
Computer Vision and Pattern Recognition
Statistics Theory
Methodology
62-08 (Primary) 60F10 (Secondary)
G.3
Diffusion models now generate high-quality, diverse samples, with an increasing focus on more powerful models. Although ensembling is a well-known way to improve supervised models, its application to unconditional score-based diffusion models remains largely unexplored. In this work we investigate whether it provides tangible benefits for generative modelling. We find that while ensembling the scores generally improves the score-matching loss and model likelihood, it fails to consistently enhance perceptual quality metrics such as FID on image datasets. We confirm this observation across a breadth of aggregation rules using Deep Ensembles, Monte Carlo Dropout, on CIFAR-10 and FFHQ. We attempt to explain this discrepancy by investigating possible explanations, such as the link between score estimation and image quality. We also look into tabular data through random forests, and find that one aggregation strategy outperforms the others. Finally, we provide theoretical insights into the summing of score models, which shed light not only on ensembling but also on several model composition techniques (e.g. guidance).
title When Are Two Scores Better Than One? Investigating Ensembles of Diffusion Models
topic Machine Learning
Computer Vision and Pattern Recognition
Statistics Theory
Methodology
62-08 (Primary) 60F10 (Secondary)
G.3
url https://arxiv.org/abs/2601.11444