Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Maini, Pratyush, Seto, Skyler, Bai, He, Grangier, David, Zhang, Yizhe, Jaitly, Navdeep |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SAGE: Steering Dialog Generation with Future-Aware State-Action Augmentation
by: Zhang, Yizhe, et al.
Published: (2025)
by: Zhang, Yizhe, et al.
Published: (2025)
Assessing the Role of Data Quality in Training Bilingual Language Models
by: Seto, Skyler, et al.
Published: (2025)
by: Seto, Skyler, et al.
Published: (2025)
Training Bilingual LMs with Data Constraints in the Targeted Language
by: Seto, Skyler, et al.
Published: (2024)
by: Seto, Skyler, et al.
Published: (2024)
Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
by: Grangier, David, et al.
Published: (2024)
by: Grangier, David, et al.
Published: (2024)
Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIP
by: Huang, Chen, et al.
Published: (2024)
by: Huang, Chen, et al.
Published: (2024)
KGLens: Towards Efficient and Effective Knowledge Probing of Large Language Models with Knowledge Graphs
by: Zheng, Shangshang, et al.
Published: (2023)
by: Zheng, Shangshang, et al.
Published: (2023)
STAMP Your Content: Proving Dataset Membership via Watermarked Rephrasings
by: Rastogi, Saksham, et al.
Published: (2025)
by: Rastogi, Saksham, et al.
Published: (2025)
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
by: Ablin, Pierre, et al.
Published: (2025)
by: Ablin, Pierre, et al.
Published: (2025)
Probing the Multi-turn Planning Capabilities of LLMs via 20 Question Games
by: Zhang, Yizhe, et al.
Published: (2023)
by: Zhang, Yizhe, et al.
Published: (2023)
Embarrassingly Simple Self-Distillation Improves Code Generation
by: Zhang, Ruixiang, et al.
Published: (2026)
by: Zhang, Ruixiang, et al.
Published: (2026)
PLANNER: Generating Diversified Paragraph via Latent Language Diffusion Model
by: Zhang, Yizhe, et al.
Published: (2023)
by: Zhang, Yizhe, et al.
Published: (2023)
CLaRa: Bridging Retrieval and Generation with Continuous Latent Reasoning
by: He, Jie, et al.
Published: (2025)
by: He, Jie, et al.
Published: (2025)
Optimal Splitting of Language Models from Mixtures to Specialized Domains
by: Seto, Skyler, et al.
Published: (2026)
by: Seto, Skyler, et al.
Published: (2026)
Closing the Gap Between Text and Speech Understanding in LLMs
by: Cuervo, Santiago, et al.
Published: (2025)
by: Cuervo, Santiago, et al.
Published: (2025)
How Far Are We from Intelligent Visual Deductive Reasoning?
by: Zhang, Yizhe, et al.
Published: (2024)
by: Zhang, Yizhe, et al.
Published: (2024)
Divide-or-Conquer? Which Part Should You Distill Your LLM?
by: Wu, Zhuofeng, et al.
Published: (2024)
by: Wu, Zhuofeng, et al.
Published: (2024)
SpeakStream: Streaming Text-to-Speech with Interleaved Data
by: Bai, Richard He, et al.
Published: (2025)
by: Bai, Richard He, et al.
Published: (2025)
DiffuCoder: Understanding and Improving Masked Diffusion Models for Code Generation
by: Gong, Shansan, et al.
Published: (2025)
by: Gong, Shansan, et al.
Published: (2025)
Eliciting In-context Retrieval and Reasoning for Long-context Large Language Models
by: Qiu, Yifu, et al.
Published: (2025)
by: Qiu, Yifu, et al.
Published: (2025)
Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators
by: Bansal, Hritik, et al.
Published: (2025)
by: Bansal, Hritik, et al.
Published: (2025)
Split and Rephrase with Large Language Models
by: Ponce, David, et al.
Published: (2023)
by: Ponce, David, et al.
Published: (2023)
Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive Flows
by: Zhang, Ruixiang, et al.
Published: (2025)
by: Zhang, Ruixiang, et al.
Published: (2025)
What Makes the Preferred Thinking Direction for LLMs in Multiple-choice Questions?
by: Zhang, Yizhe, et al.
Published: (2025)
by: Zhang, Yizhe, et al.
Published: (2025)
Training Software Engineering Agents and Verifiers with SWE-Gym
by: Pan, Jiayi, et al.
Published: (2024)
by: Pan, Jiayi, et al.
Published: (2024)
Chinese Spelling Correction as Rephrasing Language Model
by: Liu, Linfeng, et al.
Published: (2023)
by: Liu, Linfeng, et al.
Published: (2023)
Rephrase and Contrast: Fine-Tuning Language Models for Enhanced Understanding of Communication and Computer Networks
by: Wang, Liujianfu, et al.
Published: (2024)
by: Wang, Liujianfu, et al.
Published: (2024)
Revisiting ASR Error Correction with Specialized Models
by: Gu, Zijin, et al.
Published: (2024)
by: Gu, Zijin, et al.
Published: (2024)
LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning
by: Kang, Haoqiang, et al.
Published: (2025)
by: Kang, Haoqiang, et al.
Published: (2025)
Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions
by: Sedova, Anastasiia, et al.
Published: (2026)
by: Sedova, Anastasiia, et al.
Published: (2026)
Omni-Router: Sharing Routing Decisions in Sparse Mixture-of-Experts for Speech Recognition
by: Gu, Zijin, et al.
Published: (2025)
by: Gu, Zijin, et al.
Published: (2025)
Scaling Laws for Mixture Pretraining Under Data Constraints
by: Sedova, Anastasiia, et al.
Published: (2026)
by: Sedova, Anastasiia, et al.
Published: (2026)
Rephrasing Electronic Health Records for Pretraining Clinical Language Models
by: Liu, Jinghui, et al.
Published: (2024)
by: Liu, Jinghui, et al.
Published: (2024)
dMel: Speech Tokenization made Simple
by: Bai, Richard He, et al.
Published: (2024)
by: Bai, Richard He, et al.
Published: (2024)
LongRecipe: Recipe for Efficient Long Context Generalization in Large Language Models
by: Hu, Zhiyuan, et al.
Published: (2024)
by: Hu, Zhiyuan, et al.
Published: (2024)
ChipChat: Low-Latency Cascaded Conversational Agent in MLX
by: Likhomanenko, Tatiana, et al.
Published: (2025)
by: Likhomanenko, Tatiana, et al.
Published: (2025)
WikiSplit++: Easy Data Refinement for Split and Rephrase
by: Tsukagoshi, Hayato, et al.
Published: (2024)
by: Tsukagoshi, Hayato, et al.
Published: (2024)
Matryoshka Diffusion Models
by: Gu, Jiatao, et al.
Published: (2023)
by: Gu, Jiatao, et al.
Published: (2023)
How Good is Post-Hoc Watermarking With Language Model Rephrasing?
by: Fernandez, Pierre, et al.
Published: (2025)
by: Fernandez, Pierre, et al.
Published: (2025)
PERL: Pinyin Enhanced Rephrasing Language Model for Chinese ASR N-best Error Correction
by: Liang, Junhong, et al.
Published: (2024)
by: Liang, Junhong, et al.
Published: (2024)
Dynamic Gradient Alignment for Online Data Mixing
by: Fan, Simin, et al.
Published: (2024)
by: Fan, Simin, et al.
Published: (2024)
Similar Items
-
SAGE: Steering Dialog Generation with Future-Aware State-Action Augmentation
by: Zhang, Yizhe, et al.
Published: (2025) -
Assessing the Role of Data Quality in Training Bilingual Language Models
by: Seto, Skyler, et al.
Published: (2025) -
Training Bilingual LMs with Data Constraints in the Targeted Language
by: Seto, Skyler, et al.
Published: (2024) -
Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
by: Grangier, David, et al.
Published: (2024) -
Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIP
by: Huang, Chen, et al.
Published: (2024)