LM-mixup: Text Data Augmentation via Language Model based Mixup
Fuente:
arXiv
Saved in:
| Main Authors: | Deng, Zhijie, Shen, Zhouan, Li, Ling, Zhou, Yao, Zhu, Zhaowei, He, Yanji, Wang, Wei, Wei, Jiaheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification
by: Pang, Zirui, et al.
Published: (2025)
by: Pang, Zirui, et al.
Published: (2025)
inversedMixup: Data Augmentation via Inverting Mixed Embeddings
by: Kong, Fanshuang, et al.
Published: (2026)
by: Kong, Fanshuang, et al.
Published: (2026)
OFFSIDE: Benchmarking Unlearning Misinformation in Multimodal Large Language Models
by: Zheng, Hao, et al.
Published: (2025)
by: Zheng, Hao, et al.
Published: (2025)
GUARD: Generation-time LLM Unlearning via Adaptive Restriction and Detection
by: Deng, Zhijie, et al.
Published: (2025)
by: Deng, Zhijie, et al.
Published: (2025)
ENTP: Enhancing Low-Quality SFT Data via Neural-Symbolic Text Purge-Mix
by: Yang, Zile, et al.
Published: (2025)
by: Yang, Zile, et al.
Published: (2025)
Label Smoothing Improves Gradient Ascent in LLM Unlearning
by: Pang, Zirui, et al.
Published: (2025)
by: Pang, Zirui, et al.
Published: (2025)
Improving Data Efficiency via Curating LLM-Driven Rating Systems
by: Pang, Jinlong, et al.
Published: (2024)
by: Pang, Jinlong, et al.
Published: (2024)
DiffLM: Controllable Synthetic Data Generation via Diffusion Language Models
by: Zhou, Ying, et al.
Published: (2024)
by: Zhou, Ying, et al.
Published: (2024)
Token Cleaning: Fine-Grained Data Selection for LLM Supervised Fine-Tuning
by: Pang, Jinlong, et al.
Published: (2025)
by: Pang, Jinlong, et al.
Published: (2025)
ZeroLM: Data-Free Transformer Architecture Search for Language Models
by: Chen, Zhen-Song, et al.
Published: (2025)
by: Chen, Zhen-Song, et al.
Published: (2025)
CAARMA: Class Augmentation with Adversarial Mixup Regularization
by: Baali, Massa, et al.
Published: (2025)
by: Baali, Massa, et al.
Published: (2025)
LegiLM: A Fine-Tuned Legal Language Model for Data Compliance
by: Zhu, Linkai, et al.
Published: (2024)
by: Zhu, Linkai, et al.
Published: (2024)
PrivLM-Bench: A Multi-level Privacy Evaluation Benchmark for Language Models
by: Li, Haoran, et al.
Published: (2023)
by: Li, Haoran, et al.
Published: (2023)
KG-BiLM: Knowledge Graph Embedding via Bidirectional Language Models
by: Chen, Zirui, et al.
Published: (2025)
by: Chen, Zirui, et al.
Published: (2025)
InternLM-Law: An Open Source Chinese Legal Large Language Model
by: Fei, Zhiwei, et al.
Published: (2024)
by: Fei, Zhiwei, et al.
Published: (2024)
SFTMix: Elevating Language Model Instruction Tuning with Mixup Recipe
by: Xiao, Yuxin, et al.
Published: (2024)
by: Xiao, Yuxin, et al.
Published: (2024)
Towards Robustness and Diversity: Continual Learning in Dialog Generation with Text-Mixup and Batch Nuclear-Norm Maximization
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
Progressive Document-level Text Simplification via Large Language Models
by: Fang, Dengzhao, et al.
Published: (2025)
by: Fang, Dengzhao, et al.
Published: (2025)
Simple-Sampling and Hard-Mixup with Prototypes to Rebalance Contrastive Learning for Text Classification
by: Li, Mengyu, et al.
Published: (2024)
by: Li, Mengyu, et al.
Published: (2024)
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
by: Dong, Xiaoyi, et al.
Published: (2024)
by: Dong, Xiaoyi, et al.
Published: (2024)
PonderLM: Pretraining Language Models to Ponder in Continuous Space
by: Zeng, Boyi, et al.
Published: (2025)
by: Zeng, Boyi, et al.
Published: (2025)
ControlLM: Crafting Diverse Personalities for Language Models
by: Weng, Yixuan, et al.
Published: (2024)
by: Weng, Yixuan, et al.
Published: (2024)
Aligning Large Language Models via Fully Self-Synthetic Data
by: Yin, Shangjian, et al.
Published: (2025)
by: Yin, Shangjian, et al.
Published: (2025)
CogLM: Tracking Cognitive Development of Large Language Models
by: Wang, Xinglin, et al.
Published: (2024)
by: Wang, Xinglin, et al.
Published: (2024)
Lost in Translation: Latent Concept Misalignment in Text-to-Image Diffusion Models
by: Zhao, Juntu, et al.
Published: (2024)
by: Zhao, Juntu, et al.
Published: (2024)
CLLMs: Consistency Large Language Models
by: Kou, Siqi, et al.
Published: (2024)
by: Kou, Siqi, et al.
Published: (2024)
HeLM: Highlighted Evidence augmented Language Model for Enhanced Table-to-Text Generation
by: Bian, Junyi, et al.
Published: (2023)
by: Bian, Junyi, et al.
Published: (2023)
Enhancing Multi-Hop Fact Verification with Structured Knowledge-Augmented Large Language Models
by: Cao, Han, et al.
Published: (2025)
by: Cao, Han, et al.
Published: (2025)
Construction of Knowledge Graph based on Language Model
by: Zhu, Qiubai, et al.
Published: (2026)
by: Zhu, Qiubai, et al.
Published: (2026)
DecorateLM: Data Engineering through Corpus Rating, Tagging, and Editing with Language Models
by: Zhao, Ranchi, et al.
Published: (2024)
by: Zhao, Ranchi, et al.
Published: (2024)
Simulating Financial Market via Large Language Model based Agents
by: Gao, Shen, et al.
Published: (2024)
by: Gao, Shen, et al.
Published: (2024)
IDEA: An Interpretable and Editable Decision-Making Framework for LLMs via Verbal-to-Numeric Calibration
by: He, Yanji, et al.
Published: (2026)
by: He, Yanji, et al.
Published: (2026)
DiLM: Distilling Dataset into Language Model for Text-level Dataset Distillation
by: Maekawa, Aru, et al.
Published: (2024)
by: Maekawa, Aru, et al.
Published: (2024)
FltLM: An Intergrated Long-Context Large Language Model for Effective Context Filtering and Understanding
by: Deng, Jingyang, et al.
Published: (2024)
by: Deng, Jingyang, et al.
Published: (2024)
MzansiText and MzansiLM: An Open Corpus and Decoder-Only Language Model for South African Languages
by: Lombard, Anri, et al.
Published: (2026)
by: Lombard, Anri, et al.
Published: (2026)
DATE-LM: Benchmarking Data Attribution Evaluation for Large Language Models
by: Jiao, Cathy, et al.
Published: (2025)
by: Jiao, Cathy, et al.
Published: (2025)
CodecLM: Aligning Language Models with Tailored Synthetic Data
by: Wang, Zifeng, et al.
Published: (2024)
by: Wang, Zifeng, et al.
Published: (2024)
ProxyLM: Predicting Language Model Performance on Multilingual Tasks via Proxy Models
by: Anugraha, David, et al.
Published: (2024)
by: Anugraha, David, et al.
Published: (2024)
Unmasking and Improving Data Credibility: A Study with Datasets for Training Harmless Language Models
by: Zhu, Zhaowei, et al.
Published: (2023)
by: Zhu, Zhaowei, et al.
Published: (2023)
Replicating ReLM Results: Validating Large Language Models with ReLM
by: Adamson, Reece, et al.
Published: (2025)
by: Adamson, Reece, et al.
Published: (2025)
Similar Items
-
When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification
by: Pang, Zirui, et al.
Published: (2025) -
inversedMixup: Data Augmentation via Inverting Mixed Embeddings
by: Kong, Fanshuang, et al.
Published: (2026) -
OFFSIDE: Benchmarking Unlearning Misinformation in Multimodal Large Language Models
by: Zheng, Hao, et al.
Published: (2025) -
GUARD: Generation-time LLM Unlearning via Adaptive Restriction and Detection
by: Deng, Zhijie, et al.
Published: (2025) -
ENTP: Enhancing Low-Quality SFT Data via Neural-Symbolic Text Purge-Mix
by: Yang, Zile, et al.
Published: (2025)