Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Mozhi, Tissue, Howe, Wang, Lu, Qiu, Xipeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Law with Learning Rate Annealing
by: Tissue, Howe, et al.
Published: (2024)
by: Tissue, Howe, et al.
Published: (2024)
Learning Dynamics in Continual Pre-Training for Large Language Models
by: Wang, Xingjin, et al.
Published: (2025)
by: Wang, Xingjin, et al.
Published: (2025)
Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance
by: Ye, Jiasheng, et al.
Published: (2024)
by: Ye, Jiasheng, et al.
Published: (2024)
MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining
by: Wen, Bingbing, et al.
Published: (2026)
by: Wen, Bingbing, et al.
Published: (2026)
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training
by: Ge, Albert, et al.
Published: (2025)
by: Ge, Albert, et al.
Published: (2025)
LongSafety: Enhance Safety for Long-Context LLMs
by: Huang, Mianqiu, et al.
Published: (2024)
by: Huang, Mianqiu, et al.
Published: (2024)
OptiMer: Optimal Distribution Vector Merging Is Better than Data Mixing for Continual Pre-Training
by: Song, Haiyue, et al.
Published: (2026)
by: Song, Haiyue, et al.
Published: (2026)
New Encoders for German Trained from Scratch: Comparing ModernGBERT with Converted LLM2Vec Models
by: Wunderle, Julia, et al.
Published: (2025)
by: Wunderle, Julia, et al.
Published: (2025)
Contrastive Learning and Mixture of Experts Enables Precise Vector Embeddings
by: Hallee, Logan, et al.
Published: (2024)
by: Hallee, Logan, et al.
Published: (2024)
Diversity as a Reward: Fine-Tuning LLMs on a Mixture of Domain-Undetermined Data
by: Ling, Zhenqing, et al.
Published: (2025)
by: Ling, Zhenqing, et al.
Published: (2025)
Mini-batch Coresets for Memory-efficient Language Model Training on Data Mixtures
by: Nguyen, Dang, et al.
Published: (2024)
by: Nguyen, Dang, et al.
Published: (2024)
DIDS: Domain Impact-aware Data Sampling for Large Language Model Training
by: Shi, Weijie, et al.
Published: (2025)
by: Shi, Weijie, et al.
Published: (2025)
A Practice of Post-Training on Llama-3 70B with Optimal Selection of Additional Language Mixture Ratio
by: Xi, Ningyuan, et al.
Published: (2024)
by: Xi, Ningyuan, et al.
Published: (2024)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
by: Pan, Bowen, et al.
Published: (2024)
by: Pan, Bowen, et al.
Published: (2024)
RLPR: Extrapolating RLVR to General Domains without Verifiers
by: Yu, Tianyu, et al.
Published: (2025)
by: Yu, Tianyu, et al.
Published: (2025)
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
by: Zeng, Zhiyuan, et al.
Published: (2025)
by: Zeng, Zhiyuan, et al.
Published: (2025)
Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts
by: Gritsch, Nikolas, et al.
Published: (2024)
by: Gritsch, Nikolas, et al.
Published: (2024)
Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs
by: Ling Team, et al.
Published: (2025)
by: Ling Team, et al.
Published: (2025)
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
by: Nakamura, Taishi, et al.
Published: (2025)
by: Nakamura, Taishi, et al.
Published: (2025)
Explicit Multi-head Attention for Inter-head Interaction in Large Language Models
by: Peng, Runyu, et al.
Published: (2026)
by: Peng, Runyu, et al.
Published: (2026)
EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors
by: Banayeeanzade, Amin, et al.
Published: (2026)
by: Banayeeanzade, Amin, et al.
Published: (2026)
MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources
by: Nguyen, Huu, et al.
Published: (2025)
by: Nguyen, Huu, et al.
Published: (2025)
Training Optimal Large Diffusion Language Models
by: Ni, Jinjie, et al.
Published: (2025)
by: Ni, Jinjie, et al.
Published: (2025)
Test-Time Detoxification without Training or Learning Anything
by: Saglam, Baturay, et al.
Published: (2026)
by: Saglam, Baturay, et al.
Published: (2026)
DaMoC: Efficiently Selecting the Optimal Large Language Model for Fine-tuning Domain Tasks Based on Data and Model Compression
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
E2Vec: Feature Embedding with Temporal Information for Analyzing Student Actions in E-Book Systems
by: Miyazaki, Yuma, et al.
Published: (2024)
by: Miyazaki, Yuma, et al.
Published: (2024)
Unmasking and Improving Data Credibility: A Study with Datasets for Training Harmless Language Models
by: Zhu, Zhaowei, et al.
Published: (2023)
by: Zhu, Zhaowei, et al.
Published: (2023)
Diagnosing Medical Datasets with Training Dynamics
by: Wenderoth, Laura
Published: (2024)
by: Wenderoth, Laura
Published: (2024)
Mousse: Rectifying the Geometry of Muon with Curvature-Aware Preconditioning
by: Zhang, Yechen, et al.
Published: (2026)
by: Zhang, Yechen, et al.
Published: (2026)
Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization
by: Nakamura, Taishi, et al.
Published: (2025)
by: Nakamura, Taishi, et al.
Published: (2025)
LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls
by: Zhang, Kangning, et al.
Published: (2025)
by: Zhang, Kangning, et al.
Published: (2025)
ECO: Quantized Training without Full-Precision Master Weights
by: Nikdan, Mahdi, et al.
Published: (2026)
by: Nikdan, Mahdi, et al.
Published: (2026)
MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
by: Fan, Run-Ze, et al.
Published: (2025)
by: Fan, Run-Ze, et al.
Published: (2025)
Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
by: Liu, Yilun, et al.
Published: (2025)
by: Liu, Yilun, et al.
Published: (2025)
How to Synthesize Text Data without Model Collapse?
by: Zhu, Xuekai, et al.
Published: (2024)
by: Zhu, Xuekai, et al.
Published: (2024)
QuantMoE-Bench: Examining Post-Training Quantization for Mixture-of-Experts
by: Li, Pingzhi, et al.
Published: (2024)
by: Li, Pingzhi, et al.
Published: (2024)
LLMSurgeon: Diagnosing Data Mixture of Large Language Models
by: Luo, Yaxin, et al.
Published: (2026)
by: Luo, Yaxin, et al.
Published: (2026)
SMART: Submodular Data Mixture Strategy for Instruction Tuning
by: Renduchintala, H S V N S Kowndinya, et al.
Published: (2024)
by: Renduchintala, H S V N S Kowndinya, et al.
Published: (2024)
Revisiting Multilingual Data Mixtures in Language Model Pretraining
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
Similar Items
-
Scaling Law with Learning Rate Annealing
by: Tissue, Howe, et al.
Published: (2024) -
Learning Dynamics in Continual Pre-Training for Large Language Models
by: Wang, Xingjin, et al.
Published: (2025) -
Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance
by: Ye, Jiasheng, et al.
Published: (2024) -
MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining
by: Wen, Bingbing, et al.
Published: (2026) -
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training
by: Ge, Albert, et al.
Published: (2025)