DataChef: Cooking Up Optimal Data Recipes for LLM Adaptation via Reinforcement Learning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Chen, Yicheng, Ma, Zerun, Xie, Xinchen, Li, Yining, Chen, Kai |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space
par: Chen, Yicheng, et autres
Publié: (2025)
par: Chen, Yicheng, et autres
Publié: (2025)
TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration
par: Ma, Zerun, et autres
Publié: (2026)
par: Ma, Zerun, et autres
Publié: (2026)
Losses that Cook: Topological Optimal Transport for Structured Recipe Generation
par: Ottoborgo, Mattia, et autres
Publié: (2026)
par: Ottoborgo, Mattia, et autres
Publié: (2026)
GTA: A Benchmark for General Tool Agents
par: Wang, Jize, et autres
Publié: (2024)
par: Wang, Jize, et autres
Publié: (2024)
Cooking Up Creativity: Enhancing LLM Creativity through Structured Recombination
par: Mizrahi, Moran, et autres
Publié: (2025)
par: Mizrahi, Moran, et autres
Publié: (2025)
Rethinking Data Synthesis: A Teacher Model Training Recipe with Interpretation
par: Chen, Yifang, et autres
Publié: (2024)
par: Chen, Yifang, et autres
Publié: (2024)
Domain-Specific Data Generation Framework for RAG Adaptation
par: Tian, Chris Xing, et autres
Publié: (2025)
par: Tian, Chris Xing, et autres
Publié: (2025)
Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
par: Sun, Shengyin, et autres
Publié: (2025)
par: Sun, Shengyin, et autres
Publié: (2025)
CGU-ILALab at FoodBench-QA 2026: Comparing Traditional and LLM-based Approaches for Recipe Nutrient Estimation
par: Chen, Wei-Chun, et autres
Publié: (2026)
par: Chen, Wei-Chun, et autres
Publié: (2026)
Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents
par: Zhang, Xing, et autres
Publié: (2026)
par: Zhang, Xing, et autres
Publié: (2026)
CodeScout: An Effective Recipe for Reinforcement Learning of Code Search Agents
par: Sutawika, Lintang, et autres
Publié: (2026)
par: Sutawika, Lintang, et autres
Publié: (2026)
Learning from Response not Preference: A Stackelberg Approach for LLM Detoxification using Non-parallel Data
par: Xie, Xinhong, et autres
Publié: (2024)
par: Xie, Xinhong, et autres
Publié: (2024)
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement
par: Cao, Maosong, et autres
Publié: (2025)
par: Cao, Maosong, et autres
Publié: (2025)
R3-RAG: Learning Step-by-Step Reasoning and Retrieval for LLMs via Reinforcement Learning
par: Li, Yuan, et autres
Publié: (2025)
par: Li, Yuan, et autres
Publié: (2025)
Boosting LLM via Learning from Data Iteratively and Selectively
par: Jia, Qi, et autres
Publié: (2024)
par: Jia, Qi, et autres
Publié: (2024)
FoodSky: A Food-oriented Large Language Model that Passes the Chef and Dietetic Examination
par: Zhou, Pengfei, et autres
Publié: (2024)
par: Zhou, Pengfei, et autres
Publié: (2024)
Ranking Unraveled: Recipes for LLM Rankings in Head-to-Head AI Combat
par: Daynauth, Roland, et autres
Publié: (2024)
par: Daynauth, Roland, et autres
Publié: (2024)
Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
par: Xie, Tian, et autres
Publié: (2025)
par: Xie, Tian, et autres
Publié: (2025)
Enhancing Action and Ingredient Modeling for Semantically Grounded Recipe Generation
par: Liu, Guoshan, et autres
Publié: (2026)
par: Liu, Guoshan, et autres
Publié: (2026)
Reinforcement Learning on Pre-Training Data
par: Li, Siheng, et autres
Publié: (2025)
par: Li, Siheng, et autres
Publié: (2025)
Improving Data Efficiency via Curating LLM-Driven Rating Systems
par: Pang, Jinlong, et autres
Publié: (2024)
par: Pang, Jinlong, et autres
Publié: (2024)
Data Compressibility Quantifies LLM Memorization
par: Huang, Yizhan, et autres
Publié: (2025)
par: Huang, Yizhan, et autres
Publié: (2025)
GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning
par: Yang, Ningyuan, et autres
Publié: (2026)
par: Yang, Ningyuan, et autres
Publié: (2026)
MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes
par: Zhao, Changsheng, et autres
Publié: (2025)
par: Zhao, Changsheng, et autres
Publié: (2025)
LESA: Learnable LLM Layer Scaling-Up
par: Yang, Yifei, et autres
Publié: (2025)
par: Yang, Yifei, et autres
Publié: (2025)
Metacognition as Reward: Reinforcing LLM Reasoning via Knowledge and Regulation Signals
par: Chen, Sirui, et autres
Publié: (2026)
par: Chen, Sirui, et autres
Publié: (2026)
Teaching Language Models to Critique via Reinforcement Learning
par: Xie, Zhihui, et autres
Publié: (2025)
par: Xie, Zhihui, et autres
Publié: (2025)
SeRL: Self-Play Reinforcement Learning for Large Language Models with Limited Data
par: Fang, Wenkai, et autres
Publié: (2025)
par: Fang, Wenkai, et autres
Publié: (2025)
CAST: Achieving Stable LLM-based Text Analysis for Data Analytics
par: Xie, Jinxiang, et autres
Publié: (2026)
par: Xie, Jinxiang, et autres
Publié: (2026)
Safer-Instruct: Aligning Language Models with Automated Preference Data
par: Shi, Taiwei, et autres
Publié: (2023)
par: Shi, Taiwei, et autres
Publié: (2023)
ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
par: Chen, Mingyang, et autres
Publié: (2025)
par: Chen, Mingyang, et autres
Publié: (2025)
ProMind-LLM: Proactive Mental Health Care via Causal Reasoning with Sensor Data
par: Zheng, Xinzhe, et autres
Publié: (2025)
par: Zheng, Xinzhe, et autres
Publié: (2025)
SoftDedup: an Efficient Data Reweighting Method for Speeding Up Language Model Pre-training
par: He, Nan, et autres
Publié: (2024)
par: He, Nan, et autres
Publié: (2024)
Domain Adaptation of LLMs for Process Data
par: Oyamada, Rafael Seidi, et autres
Publié: (2025)
par: Oyamada, Rafael Seidi, et autres
Publié: (2025)
Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout Replay
par: Sun, Yifan, et autres
Publié: (2025)
par: Sun, Yifan, et autres
Publié: (2025)
AddrLLM: Address Rewriting via Large Language Model on Nationwide Logistics Data
par: Yang, Qinchen, et autres
Publié: (2024)
par: Yang, Qinchen, et autres
Publié: (2024)
Optimizing Decomposition for Optimal Claim Verification
par: Lu, Yining, et autres
Publié: (2025)
par: Lu, Yining, et autres
Publié: (2025)
On the Step Length Confounding in LLM Reasoning Data Selection
par: Wang, Bing, et autres
Publié: (2026)
par: Wang, Bing, et autres
Publié: (2026)
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
par: Shen, Maohao, et autres
Publié: (2025)
par: Shen, Maohao, et autres
Publié: (2025)
Learning Multi-Indicator Weights for Data Selection: A Joint Task-Model Adaptation Framework with Efficient Proxies
par: Song, Jingze, et autres
Publié: (2026)
par: Song, Jingze, et autres
Publié: (2026)
Documents similaires
-
MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space
par: Chen, Yicheng, et autres
Publié: (2025) -
TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration
par: Ma, Zerun, et autres
Publié: (2026) -
Losses that Cook: Topological Optimal Transport for Structured Recipe Generation
par: Ottoborgo, Mattia, et autres
Publié: (2026) -
GTA: A Benchmark for General Tool Agents
par: Wang, Jize, et autres
Publié: (2024) -
Cooking Up Creativity: Enhancing LLM Creativity through Structured Recombination
par: Mizrahi, Moran, et autres
Publié: (2025)