SVGauge: Towards Human-Aligned Evaluation for SVG Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zini, Leonardo, Frigieri, Elia, Aloscari, Sebastiano, Generali, Marcello, Dodi, Lorenzo, Dosen, Robert, Baraldi, Lorenzo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914027896045568
author Zini, Leonardo
Frigieri, Elia
Aloscari, Sebastiano
Generali, Marcello
Dodi, Lorenzo
Dosen, Robert
Baraldi, Lorenzo
author_facet Zini, Leonardo
Frigieri, Elia
Aloscari, Sebastiano
Generali, Marcello
Dodi, Lorenzo
Dosen, Robert
Baraldi, Lorenzo
contents Generated Scalable Vector Graphics (SVG) images demand evaluation criteria tuned to their symbolic and vectorial nature: criteria that existing metrics such as FID, LPIPS, or CLIPScore fail to satisfy. In this paper, we introduce SVGauge, the first human-aligned, reference based metric for text-to-SVG generation. SVGauge jointly measures (i) visual fidelity, obtained by extracting SigLIP image embeddings and refining them with PCA and whitening for domain alignment, and (ii) semantic consistency, captured by comparing BLIP-2-generated captions of the SVGs against the original prompts in the combined space of SBERT and TF-IDF. Evaluation on the proposed SHE benchmark shows that SVGauge attains the highest correlation with human judgments and reproduces system-level rankings of eight zero-shot LLM-based generators more faithfully than existing metrics. Our results highlight the necessity of vector-specific evaluation and provide a practical tool for benchmarking future text-to-SVG generation models.
format Preprint
id arxiv_https___arxiv_org_abs_2509_07127
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SVGauge: Towards Human-Aligned Evaluation for SVG Generation
Zini, Leonardo
Frigieri, Elia
Aloscari, Sebastiano
Generali, Marcello
Dodi, Lorenzo
Dosen, Robert
Baraldi, Lorenzo
Graphics
Artificial Intelligence
Computer Vision and Pattern Recognition
Generated Scalable Vector Graphics (SVG) images demand evaluation criteria tuned to their symbolic and vectorial nature: criteria that existing metrics such as FID, LPIPS, or CLIPScore fail to satisfy. In this paper, we introduce SVGauge, the first human-aligned, reference based metric for text-to-SVG generation. SVGauge jointly measures (i) visual fidelity, obtained by extracting SigLIP image embeddings and refining them with PCA and whitening for domain alignment, and (ii) semantic consistency, captured by comparing BLIP-2-generated captions of the SVGs against the original prompts in the combined space of SBERT and TF-IDF. Evaluation on the proposed SHE benchmark shows that SVGauge attains the highest correlation with human judgments and reproduces system-level rankings of eight zero-shot LLM-based generators more faithfully than existing metrics. Our results highlight the necessity of vector-specific evaluation and provide a practical tool for benchmarking future text-to-SVG generation models.
title SVGauge: Towards Human-Aligned Evaluation for SVG Generation
topic Graphics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.07127