Reformulation for Pretraining Data Augmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Hao, Xintong, Zhu, Ruijie, Zhang, Ge, Shen, Ke, Li, Chenggang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection
by: Hua, Kai, et al.
Published: (2025)
by: Hua, Kai, et al.
Published: (2025)
I Could've Asked That: Reformulating Unanswerable Questions
by: Zhao, Wenting, et al.
Published: (2024)
by: Zhao, Wenting, et al.
Published: (2024)
Unveiling Downstream Performance Scaling of LLMs: A Clustering-Based Perspective
by: Xu, Chengyin, et al.
Published: (2025)
by: Xu, Chengyin, et al.
Published: (2025)
Information-Preserving Reformulation of Reasoning Traces for Antidistillation
by: Ding, Jiayu, et al.
Published: (2025)
by: Ding, Jiayu, et al.
Published: (2025)
Semantic Reformulation Entropy for Robust Hallucination Detection in QA Tasks
by: Tong, Chaodong, et al.
Published: (2025)
by: Tong, Chaodong, et al.
Published: (2025)
Aug2Search: Enhancing Facebook Marketplace Search with LLM-Generated Synthetic Data Augmentation
by: Xi, Ruijie, et al.
Published: (2025)
by: Xi, Ruijie, et al.
Published: (2025)
Towards Pretraining Robust ASR Foundation Model with Acoustic-Aware Data Augmentation
by: Liu, Dancheng, et al.
Published: (2025)
by: Liu, Dancheng, et al.
Published: (2025)
DRS: Deep Question Reformulation With Structured Output
by: Li, Zhecheng, et al.
Published: (2024)
by: Li, Zhecheng, et al.
Published: (2024)
A Survey on Data Synthesis and Augmentation for Large Language Models
by: Wang, Ke, et al.
Published: (2024)
by: Wang, Ke, et al.
Published: (2024)
PretrainZero: Reinforcement Active Pretraining
by: Xing, Xingrun, et al.
Published: (2025)
by: Xing, Xingrun, et al.
Published: (2025)
Reformulating Sequential Recommendation: Learning Dynamic User Interest with Content-enriched Language Modeling
by: Jiang, Junzhe, et al.
Published: (2023)
by: Jiang, Junzhe, et al.
Published: (2023)
MachineLearningLM: Scaling Many-shot In-context Learning via Continued Pretraining
by: Dong, Haoyu, et al.
Published: (2025)
by: Dong, Haoyu, et al.
Published: (2025)
LLMs as Bridges: Reformulating Grounded Multimodal Named Entity Recognition
by: Li, Jinyuan, et al.
Published: (2024)
by: Li, Jinyuan, et al.
Published: (2024)
ConvGQR: Generative Query Reformulation for Conversational Search
by: Mo, Fengran, et al.
Published: (2023)
by: Mo, Fengran, et al.
Published: (2023)
Conversational Query Reformulation with the Guidance of Retrieved Documents
by: Park, Jeonghyun, et al.
Published: (2024)
by: Park, Jeonghyun, et al.
Published: (2024)
Efficient Pretraining Length Scaling
by: Wu, Bohong, et al.
Published: (2025)
by: Wu, Bohong, et al.
Published: (2025)
BiMix: A Bivariate Data Mixing Law for Language Model Pretraining
by: Ge, Ce, et al.
Published: (2024)
by: Ge, Ce, et al.
Published: (2024)
Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration
by: Bai, Tianyi, et al.
Published: (2024)
by: Bai, Tianyi, et al.
Published: (2024)
PAS: Data-Efficient Plug-and-Play Prompt Augmentation System
by: Zheng, Miao, et al.
Published: (2024)
by: Zheng, Miao, et al.
Published: (2024)
Solving Situation Puzzles with Large Language Model and External Reformulation
by: Li, Kun, et al.
Published: (2025)
by: Li, Kun, et al.
Published: (2025)
LM-mixup: Text Data Augmentation via Language Model based Mixup
by: Deng, Zhijie, et al.
Published: (2025)
by: Deng, Zhijie, et al.
Published: (2025)
Simple and Effective Input Reformulations for Translation
by: Yu, Brian, et al.
Published: (2023)
by: Yu, Brian, et al.
Published: (2023)
CoUDA: Coherence Evaluation via Unified Data Augmentation
by: Zhu, Dawei, et al.
Published: (2024)
by: Zhu, Dawei, et al.
Published: (2024)
RAG-DDR: Optimizing Retrieval-Augmented Generation Using Differentiable Data Rewards
by: Li, Xinze, et al.
Published: (2024)
by: Li, Xinze, et al.
Published: (2024)
Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining
by: Zhu, Jinchang, et al.
Published: (2026)
by: Zhu, Jinchang, et al.
Published: (2026)
ZeQR: Zero-shot Query Reformulation for Conversational Search
by: Yang, Dayu, et al.
Published: (2023)
by: Yang, Dayu, et al.
Published: (2023)
From Documents to Segments: A Contextual Reformulation for Topic Assignment
by: Yoon, Hoonsang, et al.
Published: (2026)
by: Yoon, Hoonsang, et al.
Published: (2026)
A Systematic Survey of Semantic Role Labeling in the Era of Pretrained Language Models
by: Chen, Huiyao, et al.
Published: (2025)
by: Chen, Huiyao, et al.
Published: (2025)
BootAug: Boosting Text Augmentation via Hybrid Instance Filtering Framework
by: Yang, Heng, et al.
Published: (2022)
by: Yang, Heng, et al.
Published: (2022)
Data, Data Everywhere: A Guide for Pretraining Dataset Construction
by: Parmar, Jupinder, et al.
Published: (2024)
by: Parmar, Jupinder, et al.
Published: (2024)
Temporal Entailment Pretraining for Clinical Language Models over EHR Data
by: Tanaka, Tatsunori, et al.
Published: (2025)
by: Tanaka, Tatsunori, et al.
Published: (2025)
Cleaner Pretraining Corpus Curation with Neural Web Scraping
by: Xu, Zhipeng, et al.
Published: (2024)
by: Xu, Zhipeng, et al.
Published: (2024)
Canvas: End-to-End Kernel Architecture Search in Neural Networks
by: Zhao, Chenggang, et al.
Published: (2023)
by: Zhao, Chenggang, et al.
Published: (2023)
Transplant Then Regenerate: A New Paradigm for Text Data Augmentation
by: Wang, Guangzhan, et al.
Published: (2025)
by: Wang, Guangzhan, et al.
Published: (2025)
Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks
by: Li, Miaomiao, et al.
Published: (2025)
by: Li, Miaomiao, et al.
Published: (2025)
PMC-InterCPT: Rethinking Biomedical Interleaved Data for Multimodal Continued Pretraining
by: Zhu, Guanghao, et al.
Published: (2026)
by: Zhu, Guanghao, et al.
Published: (2026)
On the importance of Data Scale in Pretraining Arabic Language Models
by: Ghaddar, Abbas, et al.
Published: (2024)
by: Ghaddar, Abbas, et al.
Published: (2024)
Alignment Drift in Long-Term Human-LLM Interaction: A Mechanism-Oriented Framework
by: Yao, Xintong
Published: (2026)
by: Yao, Xintong
Published: (2026)
QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining
by: Liu, Fengze, et al.
Published: (2025)
by: Liu, Fengze, et al.
Published: (2025)
Target-Oriented Pretraining Data Selection via Neuron-Activated Graph
by: Wang, Zijun, et al.
Published: (2026)
by: Wang, Zijun, et al.
Published: (2026)
Similar Items
-
AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection
by: Hua, Kai, et al.
Published: (2025) -
I Could've Asked That: Reformulating Unanswerable Questions
by: Zhao, Wenting, et al.
Published: (2024) -
Unveiling Downstream Performance Scaling of LLMs: A Clustering-Based Perspective
by: Xu, Chengyin, et al.
Published: (2025) -
Information-Preserving Reformulation of Reasoning Traces for Antidistillation
by: Ding, Jiayu, et al.
Published: (2025) -
Semantic Reformulation Entropy for Robust Hallucination Detection in QA Tasks
by: Tong, Chaodong, et al.
Published: (2025)