Controlled Automatic Task-Specific Synthetic Data Generation for Hallucination Detection
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866908755137921024 |
|---|---|
| author | Xie, Yong Aggarwal, Karan Ahmad, Aitzaz Lau, Stephen |
| author_facet | Xie, Yong Aggarwal, Karan Ahmad, Aitzaz Lau, Stephen |
| contents | We present a novel approach to automatically generate non-trivial task-specific synthetic datasets for hallucination detection. Our approach features a two-step generation-selection pipeline, using hallucination pattern guidance and a language style alignment during generation. Hallucination pattern guidance leverages the most important task-specific hallucination patterns while language style alignment aligns the style of the synthetic dataset with benchmark text. To obtain robust supervised detectors from synthetic datasets, we also adopt a data mixture strategy to improve performance robustness and generalization. Our results on three datasets show that our generated hallucination text is more closely aligned with non-hallucinated text versus baselines, to train hallucination detectors with better generalization. Our hallucination detectors trained on synthetic datasets outperform in-context-learning (ICL)-based detectors by a large margin of 32%. Our extensive experiments confirm the benefits of our approach with cross-task and cross-generator generalization. Our data-mixture-based training further improves the generalization and robustness of hallucination detection. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_12278 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Controlled Automatic Task-Specific Synthetic Data Generation for Hallucination Detection Xie, Yong Aggarwal, Karan Ahmad, Aitzaz Lau, Stephen Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language 68T50 I.2.7 We present a novel approach to automatically generate non-trivial task-specific synthetic datasets for hallucination detection. Our approach features a two-step generation-selection pipeline, using hallucination pattern guidance and a language style alignment during generation. Hallucination pattern guidance leverages the most important task-specific hallucination patterns while language style alignment aligns the style of the synthetic dataset with benchmark text. To obtain robust supervised detectors from synthetic datasets, we also adopt a data mixture strategy to improve performance robustness and generalization. Our results on three datasets show that our generated hallucination text is more closely aligned with non-hallucinated text versus baselines, to train hallucination detectors with better generalization. Our hallucination detectors trained on synthetic datasets outperform in-context-learning (ICL)-based detectors by a large margin of 32%. Our extensive experiments confirm the benefits of our approach with cross-task and cross-generator generalization. Our data-mixture-based training further improves the generalization and robustness of hallucination detection. |
| title | Controlled Automatic Task-Specific Synthetic Data Generation for Hallucination Detection |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language 68T50 I.2.7 |
| url | https://arxiv.org/abs/2410.12278 |