Hacking Neural Evaluation Metrics with Single Hub Text

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Deguchi, Hiroyuki, Chousa, Katsuki, Sakai, Yusuke
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911371164123136
author Deguchi, Hiroyuki
Chousa, Katsuki
Sakai, Yusuke
author_facet Deguchi, Hiroyuki
Chousa, Katsuki
Sakai, Yusuke
contents Strongly human-correlated evaluation metrics serve as an essential compass for the development and improvement of generation models and must be highly reliable and robust. Recent embedding-based neural text evaluation metrics, such as COMET for translation tasks, are widely used in both research and development fields. However, there is no guarantee that they yield reliable evaluation results due to the black-box nature of neural networks. To raise concerns about the reliability and safety of such metrics, we propose a method for finding a single adversarial text in the discrete space that is consistently evaluated as high-quality, regardless of the test cases, to identify the vulnerabilities in evaluation metrics. The single hub text found with our method achieved 79.1 COMET% and 67.8 COMET% in the WMT'24 English-to-Japanese (En--Ja) and English-to-German (En--De) translation tasks, respectively, outperforming translations generated individually for each source sentence by using M2M100, a general translation model. Furthermore, we also confirmed that the hub text found with our method generalizes across multiple language pairs such as Ja--En and De--En.
format Preprint
id arxiv_https___arxiv_org_abs_2512_16323
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hacking Neural Evaluation Metrics with Single Hub Text
Deguchi, Hiroyuki
Chousa, Katsuki
Sakai, Yusuke
Computation and Language
Strongly human-correlated evaluation metrics serve as an essential compass for the development and improvement of generation models and must be highly reliable and robust. Recent embedding-based neural text evaluation metrics, such as COMET for translation tasks, are widely used in both research and development fields. However, there is no guarantee that they yield reliable evaluation results due to the black-box nature of neural networks. To raise concerns about the reliability and safety of such metrics, we propose a method for finding a single adversarial text in the discrete space that is consistently evaluated as high-quality, regardless of the test cases, to identify the vulnerabilities in evaluation metrics. The single hub text found with our method achieved 79.1 COMET% and 67.8 COMET% in the WMT'24 English-to-Japanese (En--Ja) and English-to-German (En--De) translation tasks, respectively, outperforming translations generated individually for each source sentence by using M2M100, a general translation model. Furthermore, we also confirmed that the hub text found with our method generalizes across multiple language pairs such as Ja--En and De--En.
title Hacking Neural Evaluation Metrics with Single Hub Text
topic Computation and Language
url https://arxiv.org/abs/2512.16323