Easy Samples Are All You Need: Self-Evolving LLMs via Data-Efficient Reinforcement Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Yu, Zhiyin, Zhang, Bo, Hou, Qibin, Wu, Zhonghai, Luo, Xiao, Bai, Lei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Is Depth All You Need? An Exploration of Iterative Reasoning in LLMs
por: Wu, Zongqian, et al.
Publicado: (2025)
por: Wu, Zongqian, et al.
Publicado: (2025)
Self-Evolving LLMs via Continual Instruction Tuning
por: Kang, Jiazheng, et al.
Publicado: (2025)
por: Kang, Jiazheng, et al.
Publicado: (2025)
A Survey of Reinforcement Learning for Large Language Models under Data Scarcity: Challenges and Solutions
por: Yu, Zhiyin, et al.
Publicado: (2026)
por: Yu, Zhiyin, et al.
Publicado: (2026)
Efficient Deep Learning Board: Training Feedback Is Not All You Need
por: Gong, Lina, et al.
Publicado: (2024)
por: Gong, Lina, et al.
Publicado: (2024)
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective
por: He, Shenghua, et al.
Publicado: (2025)
por: He, Shenghua, et al.
Publicado: (2025)
Is Exploration All You Need? Effective Exploration Characteristics for Transfer in Reinforcement Learning
por: Balloch, Jonathan C., et al.
Publicado: (2024)
por: Balloch, Jonathan C., et al.
Publicado: (2024)
Scaling Is All You Need: Autonomous Driving with JAX-Accelerated Reinforcement Learning
por: Harmel, Moritz, et al.
Publicado: (2023)
por: Harmel, Moritz, et al.
Publicado: (2023)
EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
por: Tang, Hanlin, et al.
Publicado: (2024)
por: Tang, Hanlin, et al.
Publicado: (2024)
LoRA is All You Need for Safety Alignment of Reasoning LLMs
por: Xue, Yihao, et al.
Publicado: (2025)
por: Xue, Yihao, et al.
Publicado: (2025)
Rethinking Data Selection at Scale: Random Selection is Almost All You Need
por: Xia, Tingyu, et al.
Publicado: (2024)
por: Xia, Tingyu, et al.
Publicado: (2024)
Importance Sampling is All You Need: Predict LLM's performance on new benchmark by reusing existing benchmark
por: Shi, Junjie, et al.
Publicado: (2025)
por: Shi, Junjie, et al.
Publicado: (2025)
Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning
por: Yuan, Xinbin, et al.
Publicado: (2025)
por: Yuan, Xinbin, et al.
Publicado: (2025)
Attention Is All You Need for KV Cache in Diffusion LLMs
por: Nguyen-Tri, Quan, et al.
Publicado: (2025)
por: Nguyen-Tri, Quan, et al.
Publicado: (2025)
Attention is All You Need Until You Need Retention
por: Yaslioglu, M. Murat
Publicado: (2025)
por: Yaslioglu, M. Murat
Publicado: (2025)
OffSeeker: Online Reinforcement Learning Is Not All You Need for Deep Research Agents
por: Zhou, Yuhang, et al.
Publicado: (2026)
por: Zhou, Yuhang, et al.
Publicado: (2026)
Transduction is All You Need for Structured Data Workflows
por: Gliozzo, Alfio, et al.
Publicado: (2025)
por: Gliozzo, Alfio, et al.
Publicado: (2025)
Synthetic Data RL: Task Definition Is All You Need
por: Guo, Yiduo, et al.
Publicado: (2025)
por: Guo, Yiduo, et al.
Publicado: (2025)
Beyond Chat: a Framework for LLMs as Human-Centered Support Systems
por: Zhou, Zhiyin
Publicado: (2025)
por: Zhou, Zhiyin
Publicado: (2025)
More Agents Is All You Need
por: Li, Junyou, et al.
Publicado: (2024)
por: Li, Junyou, et al.
Publicado: (2024)
Not All Documents Are What You Need for Extracting Instruction Tuning Data
por: Zhang, Chi, et al.
Publicado: (2025)
por: Zhang, Chi, et al.
Publicado: (2025)
Contrast Is All You Need
por: Kilic, Burak, et al.
Publicado: (2023)
por: Kilic, Burak, et al.
Publicado: (2023)
Context is All You Need
por: Delanois, Jean Erik, et al.
Publicado: (2026)
por: Delanois, Jean Erik, et al.
Publicado: (2026)
SynthDST: Synthetic Data is All You Need for Few-Shot Dialog State Tracking
por: Kulkarni, Atharva, et al.
Publicado: (2024)
por: Kulkarni, Atharva, et al.
Publicado: (2024)
FineVision: Open Data Is All You Need
por: Wiedmann, Luis, et al.
Publicado: (2025)
por: Wiedmann, Luis, et al.
Publicado: (2025)
Common Sense Is All You Need
por: Latapie, Hugo
Publicado: (2025)
por: Latapie, Hugo
Publicado: (2025)
Data-Efficient Training by Evolved Sampling
por: Cheng, Ziheng, et al.
Publicado: (2025)
por: Cheng, Ziheng, et al.
Publicado: (2025)
Self-Verification is All You Need To Pass The Japanese Bar Examination
por: Shin, Andrew
Publicado: (2026)
por: Shin, Andrew
Publicado: (2026)
[MASK] is All You Need
por: Hu, Vincent Tao, et al.
Publicado: (2024)
por: Hu, Vincent Tao, et al.
Publicado: (2024)
Information Gain Is Not All You Need
por: Ericson, Ludvig, et al.
Publicado: (2025)
por: Ericson, Ludvig, et al.
Publicado: (2025)
Memory augment is All You Need for image restoration
por: Zhang, Xiao Feng, et al.
Publicado: (2023)
por: Zhang, Xiao Feng, et al.
Publicado: (2023)
Cooperation Is All You Need
por: Adeel, Ahsan, et al.
Publicado: (2023)
por: Adeel, Ahsan, et al.
Publicado: (2023)
Self-Evolved Reward Learning for LLMs
por: Huang, Chenghua, et al.
Publicado: (2024)
por: Huang, Chenghua, et al.
Publicado: (2024)
Ideal Registration? Segmentation is All You Need
por: Chen, Xiang, et al.
Publicado: (2025)
por: Chen, Xiang, et al.
Publicado: (2025)
Learning to Evolve: A Self-Improving Framework for Multi-Agent Systems via Textual Parameter Graph Optimization
por: He, Shan, et al.
Publicado: (2026)
por: He, Shan, et al.
Publicado: (2026)
Exploitation Is All You Need... for Exploration
por: Rentschler, Micah, et al.
Publicado: (2025)
por: Rentschler, Micah, et al.
Publicado: (2025)
Training on the Benchmark Is Not All You Need
por: Ni, Shiwen, et al.
Publicado: (2024)
por: Ni, Shiwen, et al.
Publicado: (2024)
Rethinking Deep Clustering Paradigms: Self-Supervision Is All You Need
por: Shaheena, Amal, et al.
Publicado: (2025)
por: Shaheena, Amal, et al.
Publicado: (2025)
Rho-1: Not All Tokens Are What You Need
por: Lin, Zhenghao, et al.
Publicado: (2024)
por: Lin, Zhenghao, et al.
Publicado: (2024)
EasyGen: Easing Multimodal Generation with BiDiffuser and LLMs
por: Zhao, Xiangyu, et al.
Publicado: (2023)
por: Zhao, Xiangyu, et al.
Publicado: (2023)
Bayesian Flow Is All You Need to Sample Out-of-Distribution Chemical Spaces
por: Tao, Nianze, et al.
Publicado: (2024)
por: Tao, Nianze, et al.
Publicado: (2024)
Ejemplares similares
-
Is Depth All You Need? An Exploration of Iterative Reasoning in LLMs
por: Wu, Zongqian, et al.
Publicado: (2025) -
Self-Evolving LLMs via Continual Instruction Tuning
por: Kang, Jiazheng, et al.
Publicado: (2025) -
A Survey of Reinforcement Learning for Large Language Models under Data Scarcity: Challenges and Solutions
por: Yu, Zhiyin, et al.
Publicado: (2026) -
Efficient Deep Learning Board: Training Feedback Is Not All You Need
por: Gong, Lina, et al.
Publicado: (2024) -
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective
por: He, Shenghua, et al.
Publicado: (2025)