Hacking Neural Evaluation Metrics with Single Hub Text
Fuente:
arXiv
Saved in:
| Main Authors: | Deguchi, Hiroyuki, Chousa, Katsuki, Sakai, Yusuke |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
One Single Hub Text Breaks CLIP: Identifying Vulnerabilities in Cross-Modal Encoders via Hubness
by: Deguchi, Hiroyuki, et al.
Published: (2026)
by: Deguchi, Hiroyuki, et al.
Published: (2026)
JaParaPat: A Large-Scale Japanese-English Parallel Patent Application Corpus
by: Nagata, Masaaki, et al.
Published: (2025)
by: Nagata, Masaaki, et al.
Published: (2025)
mbrs: A Library for Minimum Bayes Risk Decoding
by: Deguchi, Hiroyuki, et al.
Published: (2024)
by: Deguchi, Hiroyuki, et al.
Published: (2024)
A Japanese-Chinese Parallel Corpus Using Crowdsourcing for Web Mining
by: Nagata, Masaaki, et al.
Published: (2024)
by: Nagata, Masaaki, et al.
Published: (2024)
Diversity Explains Inference Scaling Laws: Through a Case Study of Minimum Bayes Risk Decoding
by: Kamigaito, Hidetaka, et al.
Published: (2024)
by: Kamigaito, Hidetaka, et al.
Published: (2024)
Centroid-Based Efficient Minimum Bayes Risk Decoding
by: Deguchi, Hiroyuki, et al.
Published: (2024)
by: Deguchi, Hiroyuki, et al.
Published: (2024)
Agreement-Constrained Probabilistic Minimum Bayes Risk Decoding
by: Natsumi, Koki, et al.
Published: (2025)
by: Natsumi, Koki, et al.
Published: (2025)
Rethinking Evaluation Metrics for Grammatical Error Correction: Why Use a Different Evaluation Process than Human?
by: Goto, Takumi, et al.
Published: (2025)
by: Goto, Takumi, et al.
Published: (2025)
WikiSplit++: Easy Data Refinement for Split and Rephrase
by: Tsukagoshi, Hayato, et al.
Published: (2024)
by: Tsukagoshi, Hayato, et al.
Published: (2024)
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models
by: Taguchi, Shun, et al.
Published: (2025)
by: Taguchi, Shun, et al.
Published: (2025)
Case-Based Decision-Theoretic Decoding with Quality Memories
by: Deguchi, Hiroyuki, et al.
Published: (2025)
by: Deguchi, Hiroyuki, et al.
Published: (2025)
Reliability Crisis of Reference-free Metrics for Grammatical Error Correction
by: Goto, Takumi, et al.
Published: (2025)
by: Goto, Takumi, et al.
Published: (2025)
AdTEC: A Unified Benchmark for Evaluating Text Quality in Search Engine Advertising
by: Zhang, Peinan, et al.
Published: (2024)
by: Zhang, Peinan, et al.
Published: (2024)
Grammatical Error Correction Evaluation by Optimally Transporting Edit Representation
by: Goto, Takumi, et al.
Published: (2026)
by: Goto, Takumi, et al.
Published: (2026)
gec-metrics: A Unified Library for Grammatical Error Correction Evaluation
by: Goto, Takumi, et al.
Published: (2025)
by: Goto, Takumi, et al.
Published: (2025)
SoftMatcha: A Soft and Fast Pattern Matcher for Billion-Scale Corpus Searches
by: Deguchi, Hiroyuki, et al.
Published: (2025)
by: Deguchi, Hiroyuki, et al.
Published: (2025)
Long-Tail Crisis in Nearest Neighbor Language Models
by: Nishida, Yuto, et al.
Published: (2025)
by: Nishida, Yuto, et al.
Published: (2025)
Evaluating Text Style Transfer Evaluation: Are There Any Reliable Metrics?
by: Mukherjee, Sourabrata, et al.
Published: (2025)
by: Mukherjee, Sourabrata, et al.
Published: (2025)
Identifying Reliable Evaluation Metrics for Scientific Text Revision
by: Jourdan, Léane, et al.
Published: (2025)
by: Jourdan, Léane, et al.
Published: (2025)
Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation
by: Kasaei, Seyed Amir, et al.
Published: (2025)
by: Kasaei, Seyed Amir, et al.
Published: (2025)
Reference-free Evaluation Metrics for Text Generation: A Survey
by: Ito, Takumi, et al.
Published: (2025)
by: Ito, Takumi, et al.
Published: (2025)
Evaluation Metrics for Text Data Augmentation in NLP
by: Amadeus, Marcellus, et al.
Published: (2024)
by: Amadeus, Marcellus, et al.
Published: (2024)
Edit-level Majority Voting Mitigates Over-Correction in LLM-based Grammatical Error Correction
by: Goto, Takumi, et al.
Published: (2026)
by: Goto, Takumi, et al.
Published: (2026)
More Than Bits: Multi-Envelope Double Binary Factorization for Extreme Quantization
by: Ichikawa, Yuma, et al.
Published: (2025)
by: Ichikawa, Yuma, et al.
Published: (2025)
KDH-MLTC: Knowledge Distillation for Healthcare Multi-Label Text Classification
by: Sakai, Hajar, et al.
Published: (2025)
by: Sakai, Hajar, et al.
Published: (2025)
Large Language Models for Healthcare Text Classification: A Systematic Review
by: Sakai, Hajar, et al.
Published: (2025)
by: Sakai, Hajar, et al.
Published: (2025)
TextTIGER: Text-based Intelligent Generation with Entity Prompt Refinement for Text-to-Image Generation
by: Ozaki, Shintaro, et al.
Published: (2025)
by: Ozaki, Shintaro, et al.
Published: (2025)
Reproducing the Metric-Based Evaluation of a Set of Controllable Text Generation Techniques
by: Lorandi, Michela, et al.
Published: (2024)
by: Lorandi, Michela, et al.
Published: (2024)
QUAD-LLM-MLTC: Large Language Models Ensemble Learning for Healthcare Text Multi-Label Classification
by: Sakai, Hajar, et al.
Published: (2025)
by: Sakai, Hajar, et al.
Published: (2025)
Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following Ability
by: Sakai, Yusuke, et al.
Published: (2025)
by: Sakai, Yusuke, et al.
Published: (2025)
IMPARA-GED: Grammatical Error Detection is Boosting Reference-free Grammatical Error Quality Estimator
by: Sakai, Yusuke, et al.
Published: (2025)
by: Sakai, Yusuke, et al.
Published: (2025)
Assessing Evaluation Metrics for Neural Test Oracle Generation
by: Shin, Jiho, et al.
Published: (2023)
by: Shin, Jiho, et al.
Published: (2023)
HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application Vulnerabilities
by: Ren, Xiaoxue, et al.
Published: (2025)
by: Ren, Xiaoxue, et al.
Published: (2025)
Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation
by: Dejl, Adam, et al.
Published: (2025)
by: Dejl, Adam, et al.
Published: (2025)
NormEval: A Unified Multi-Metric Framework for Evaluating Semantic Fidelity in Text Normalization
by: Kafi, Md Abdullah Al, et al.
Published: (2025)
by: Kafi, Md Abdullah Al, et al.
Published: (2025)
Analyzing Correlations Between Intrinsic and Extrinsic Bias Metrics of Static Word Embeddings With Their Measuring Biases Aligned
by: Katô, Taisei, et al.
Published: (2024)
by: Katô, Taisei, et al.
Published: (2024)
Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation
by: Baumann, Joachim, et al.
Published: (2025)
by: Baumann, Joachim, et al.
Published: (2025)
Multilingual Dialogue Generation and Localization with Dialogue Act Scripting
by: Vasselli, Justin, et al.
Published: (2025)
by: Vasselli, Justin, et al.
Published: (2025)
Discovery of Rare Causal Knowledge from Financial Statement Summaries
by: Sakaji, Hiroki, et al.
Published: (2024)
by: Sakaji, Hiroki, et al.
Published: (2024)
Simple Hack for Transformers against Heavy Long-Text Classification on a Time- and Memory-Limited GPU Service
by: Mutasodirin, Mirza Alim, et al.
Published: (2024)
by: Mutasodirin, Mirza Alim, et al.
Published: (2024)
Similar Items
-
One Single Hub Text Breaks CLIP: Identifying Vulnerabilities in Cross-Modal Encoders via Hubness
by: Deguchi, Hiroyuki, et al.
Published: (2026) -
JaParaPat: A Large-Scale Japanese-English Parallel Patent Application Corpus
by: Nagata, Masaaki, et al.
Published: (2025) -
mbrs: A Library for Minimum Bayes Risk Decoding
by: Deguchi, Hiroyuki, et al.
Published: (2024) -
A Japanese-Chinese Parallel Corpus Using Crowdsourcing for Web Mining
by: Nagata, Masaaki, et al.
Published: (2024) -
Diversity Explains Inference Scaling Laws: Through a Case Study of Minimum Bayes Risk Decoding
by: Kamigaito, Hidetaka, et al.
Published: (2024)