inversedMixup: Data Augmentation via Inverting Mixed Embeddings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kong, Fanshuang, Zhang, Richong, Sun, Qiyu, Nie, Zhijie, Deng, Ting, Hu, Chunming
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912883752828928
author Kong, Fanshuang
Zhang, Richong
Sun, Qiyu
Nie, Zhijie
Deng, Ting
Hu, Chunming
author_facet Kong, Fanshuang
Zhang, Richong
Sun, Qiyu
Nie, Zhijie
Deng, Ting
Hu, Chunming
contents Mixup generates augmented samples by linearly interpolating inputs and labels with a controllable ratio. However, since it operates in the latent embedding level, the resulting samples are not human-interpretable. In contrast, LLM-based augmentation methods produce sentences via prompts at the token level, yielding readable outputs but offering limited control over the generation process. Inspired by recent advances in LLM inversion, which reconstructs natural language from embeddings and helps bridge the gap between latent embedding space and discrete token space, we propose inversedMixup, a unified framework that combines the controllability of Mixup with the interpretability of LLM-based generation. Specifically, inversedMixup adopts a three-stage training procedure to align the output embedding space of a task-specific model with the input embedding space of an LLM. Upon successful alignment, inversedMixup can reconstruct mixed embeddings with a controllable mixing ratio into human-interpretable augmented sentences, thereby improving the augmentation performance. Additionally, inversedMixup provides the first empirical evidence of the manifold intrusion phenomenon in text Mixup and introduces a simple yet effective strategy to mitigate it. Extensive experiments demonstrate the effectiveness and generalizability of our approach in both few-shot and fully supervised scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21543
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle inversedMixup: Data Augmentation via Inverting Mixed Embeddings
Kong, Fanshuang
Zhang, Richong
Sun, Qiyu
Nie, Zhijie
Deng, Ting
Hu, Chunming
Computation and Language
Mixup generates augmented samples by linearly interpolating inputs and labels with a controllable ratio. However, since it operates in the latent embedding level, the resulting samples are not human-interpretable. In contrast, LLM-based augmentation methods produce sentences via prompts at the token level, yielding readable outputs but offering limited control over the generation process. Inspired by recent advances in LLM inversion, which reconstructs natural language from embeddings and helps bridge the gap between latent embedding space and discrete token space, we propose inversedMixup, a unified framework that combines the controllability of Mixup with the interpretability of LLM-based generation. Specifically, inversedMixup adopts a three-stage training procedure to align the output embedding space of a task-specific model with the input embedding space of an LLM. Upon successful alignment, inversedMixup can reconstruct mixed embeddings with a controllable mixing ratio into human-interpretable augmented sentences, thereby improving the augmentation performance. Additionally, inversedMixup provides the first empirical evidence of the manifold intrusion phenomenon in text Mixup and introduces a simple yet effective strategy to mitigate it. Extensive experiments demonstrate the effectiveness and generalizability of our approach in both few-shot and fully supervised scenarios.
title inversedMixup: Data Augmentation via Inverting Mixed Embeddings
topic Computation and Language
url https://arxiv.org/abs/2601.21543