Reformulation for Pretraining Data Augmentation
Fuente:
arXiv
Salvato in:
| Autori principali: | Hao, Xintong, Zhu, Ruijie, Zhang, Ge, Shen, Ke, Li, Chenggang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection
di: Hua, Kai, et al.
Pubblicazione: (2025)
di: Hua, Kai, et al.
Pubblicazione: (2025)
I Could've Asked That: Reformulating Unanswerable Questions
di: Zhao, Wenting, et al.
Pubblicazione: (2024)
di: Zhao, Wenting, et al.
Pubblicazione: (2024)
Unveiling Downstream Performance Scaling of LLMs: A Clustering-Based Perspective
di: Xu, Chengyin, et al.
Pubblicazione: (2025)
di: Xu, Chengyin, et al.
Pubblicazione: (2025)
Information-Preserving Reformulation of Reasoning Traces for Antidistillation
di: Ding, Jiayu, et al.
Pubblicazione: (2025)
di: Ding, Jiayu, et al.
Pubblicazione: (2025)
Semantic Reformulation Entropy for Robust Hallucination Detection in QA Tasks
di: Tong, Chaodong, et al.
Pubblicazione: (2025)
di: Tong, Chaodong, et al.
Pubblicazione: (2025)
Aug2Search: Enhancing Facebook Marketplace Search with LLM-Generated Synthetic Data Augmentation
di: Xi, Ruijie, et al.
Pubblicazione: (2025)
di: Xi, Ruijie, et al.
Pubblicazione: (2025)
Towards Pretraining Robust ASR Foundation Model with Acoustic-Aware Data Augmentation
di: Liu, Dancheng, et al.
Pubblicazione: (2025)
di: Liu, Dancheng, et al.
Pubblicazione: (2025)
DRS: Deep Question Reformulation With Structured Output
di: Li, Zhecheng, et al.
Pubblicazione: (2024)
di: Li, Zhecheng, et al.
Pubblicazione: (2024)
A Survey on Data Synthesis and Augmentation for Large Language Models
di: Wang, Ke, et al.
Pubblicazione: (2024)
di: Wang, Ke, et al.
Pubblicazione: (2024)
PretrainZero: Reinforcement Active Pretraining
di: Xing, Xingrun, et al.
Pubblicazione: (2025)
di: Xing, Xingrun, et al.
Pubblicazione: (2025)
Reformulating Sequential Recommendation: Learning Dynamic User Interest with Content-enriched Language Modeling
di: Jiang, Junzhe, et al.
Pubblicazione: (2023)
di: Jiang, Junzhe, et al.
Pubblicazione: (2023)
MachineLearningLM: Scaling Many-shot In-context Learning via Continued Pretraining
di: Dong, Haoyu, et al.
Pubblicazione: (2025)
di: Dong, Haoyu, et al.
Pubblicazione: (2025)
LLMs as Bridges: Reformulating Grounded Multimodal Named Entity Recognition
di: Li, Jinyuan, et al.
Pubblicazione: (2024)
di: Li, Jinyuan, et al.
Pubblicazione: (2024)
ConvGQR: Generative Query Reformulation for Conversational Search
di: Mo, Fengran, et al.
Pubblicazione: (2023)
di: Mo, Fengran, et al.
Pubblicazione: (2023)
Conversational Query Reformulation with the Guidance of Retrieved Documents
di: Park, Jeonghyun, et al.
Pubblicazione: (2024)
di: Park, Jeonghyun, et al.
Pubblicazione: (2024)
Efficient Pretraining Length Scaling
di: Wu, Bohong, et al.
Pubblicazione: (2025)
di: Wu, Bohong, et al.
Pubblicazione: (2025)
BiMix: A Bivariate Data Mixing Law for Language Model Pretraining
di: Ge, Ce, et al.
Pubblicazione: (2024)
di: Ge, Ce, et al.
Pubblicazione: (2024)
Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration
di: Bai, Tianyi, et al.
Pubblicazione: (2024)
di: Bai, Tianyi, et al.
Pubblicazione: (2024)
PAS: Data-Efficient Plug-and-Play Prompt Augmentation System
di: Zheng, Miao, et al.
Pubblicazione: (2024)
di: Zheng, Miao, et al.
Pubblicazione: (2024)
Solving Situation Puzzles with Large Language Model and External Reformulation
di: Li, Kun, et al.
Pubblicazione: (2025)
di: Li, Kun, et al.
Pubblicazione: (2025)
LM-mixup: Text Data Augmentation via Language Model based Mixup
di: Deng, Zhijie, et al.
Pubblicazione: (2025)
di: Deng, Zhijie, et al.
Pubblicazione: (2025)
Simple and Effective Input Reformulations for Translation
di: Yu, Brian, et al.
Pubblicazione: (2023)
di: Yu, Brian, et al.
Pubblicazione: (2023)
CoUDA: Coherence Evaluation via Unified Data Augmentation
di: Zhu, Dawei, et al.
Pubblicazione: (2024)
di: Zhu, Dawei, et al.
Pubblicazione: (2024)
RAG-DDR: Optimizing Retrieval-Augmented Generation Using Differentiable Data Rewards
di: Li, Xinze, et al.
Pubblicazione: (2024)
di: Li, Xinze, et al.
Pubblicazione: (2024)
Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining
di: Zhu, Jinchang, et al.
Pubblicazione: (2026)
di: Zhu, Jinchang, et al.
Pubblicazione: (2026)
ZeQR: Zero-shot Query Reformulation for Conversational Search
di: Yang, Dayu, et al.
Pubblicazione: (2023)
di: Yang, Dayu, et al.
Pubblicazione: (2023)
From Documents to Segments: A Contextual Reformulation for Topic Assignment
di: Yoon, Hoonsang, et al.
Pubblicazione: (2026)
di: Yoon, Hoonsang, et al.
Pubblicazione: (2026)
A Systematic Survey of Semantic Role Labeling in the Era of Pretrained Language Models
di: Chen, Huiyao, et al.
Pubblicazione: (2025)
di: Chen, Huiyao, et al.
Pubblicazione: (2025)
BootAug: Boosting Text Augmentation via Hybrid Instance Filtering Framework
di: Yang, Heng, et al.
Pubblicazione: (2022)
di: Yang, Heng, et al.
Pubblicazione: (2022)
Data, Data Everywhere: A Guide for Pretraining Dataset Construction
di: Parmar, Jupinder, et al.
Pubblicazione: (2024)
di: Parmar, Jupinder, et al.
Pubblicazione: (2024)
Temporal Entailment Pretraining for Clinical Language Models over EHR Data
di: Tanaka, Tatsunori, et al.
Pubblicazione: (2025)
di: Tanaka, Tatsunori, et al.
Pubblicazione: (2025)
Cleaner Pretraining Corpus Curation with Neural Web Scraping
di: Xu, Zhipeng, et al.
Pubblicazione: (2024)
di: Xu, Zhipeng, et al.
Pubblicazione: (2024)
Canvas: End-to-End Kernel Architecture Search in Neural Networks
di: Zhao, Chenggang, et al.
Pubblicazione: (2023)
di: Zhao, Chenggang, et al.
Pubblicazione: (2023)
Transplant Then Regenerate: A New Paradigm for Text Data Augmentation
di: Wang, Guangzhan, et al.
Pubblicazione: (2025)
di: Wang, Guangzhan, et al.
Pubblicazione: (2025)
Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks
di: Li, Miaomiao, et al.
Pubblicazione: (2025)
di: Li, Miaomiao, et al.
Pubblicazione: (2025)
PMC-InterCPT: Rethinking Biomedical Interleaved Data for Multimodal Continued Pretraining
di: Zhu, Guanghao, et al.
Pubblicazione: (2026)
di: Zhu, Guanghao, et al.
Pubblicazione: (2026)
On the importance of Data Scale in Pretraining Arabic Language Models
di: Ghaddar, Abbas, et al.
Pubblicazione: (2024)
di: Ghaddar, Abbas, et al.
Pubblicazione: (2024)
Alignment Drift in Long-Term Human-LLM Interaction: A Mechanism-Oriented Framework
di: Yao, Xintong
Pubblicazione: (2026)
di: Yao, Xintong
Pubblicazione: (2026)
QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining
di: Liu, Fengze, et al.
Pubblicazione: (2025)
di: Liu, Fengze, et al.
Pubblicazione: (2025)
Target-Oriented Pretraining Data Selection via Neuron-Activated Graph
di: Wang, Zijun, et al.
Pubblicazione: (2026)
di: Wang, Zijun, et al.
Pubblicazione: (2026)
Documenti analoghi
-
AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection
di: Hua, Kai, et al.
Pubblicazione: (2025) -
I Could've Asked That: Reformulating Unanswerable Questions
di: Zhao, Wenting, et al.
Pubblicazione: (2024) -
Unveiling Downstream Performance Scaling of LLMs: A Clustering-Based Perspective
di: Xu, Chengyin, et al.
Pubblicazione: (2025) -
Information-Preserving Reformulation of Reasoning Traces for Antidistillation
di: Ding, Jiayu, et al.
Pubblicazione: (2025) -
Semantic Reformulation Entropy for Robust Hallucination Detection in QA Tasks
di: Tong, Chaodong, et al.
Pubblicazione: (2025)