Rethinking Data Selection at Scale: Random Selection is Almost All You Need
Fuente:
arXiv
Salvato in:
| Autori principali: | Xia, Tingyu, Yu, Bowen, Dang, Kai, Yang, An, Wu, Yuan, Tian, Yuan, Chang, Yi, Lin, Junyang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Language Models can Evaluate Themselves via Probability Discrepancy
di: Xia, Tingyu, et al.
Pubblicazione: (2024)
di: Xia, Tingyu, et al.
Pubblicazione: (2024)
A Survey of RWKV
di: Li, Zhiyuan, et al.
Pubblicazione: (2024)
di: Li, Zhiyuan, et al.
Pubblicazione: (2024)
Tensor Product Attention Is All You Need
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
Not All Layers Need Tuning: Selective Layer Restoration Recovers Diversity
di: Zhang, Bowen, et al.
Pubblicazione: (2026)
di: Zhang, Bowen, et al.
Pubblicazione: (2026)
Not All Documents Are What You Need for Extracting Instruction Tuning Data
di: Zhang, Chi, et al.
Pubblicazione: (2025)
di: Zhang, Chi, et al.
Pubblicazione: (2025)
Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization
di: Dong, Zhijin
Pubblicazione: (2025)
di: Dong, Zhijin
Pubblicazione: (2025)
COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning
di: Bai, Yuelin, et al.
Pubblicazione: (2024)
di: Bai, Yuelin, et al.
Pubblicazione: (2024)
Sparsity May Be All You Need: Sparse Random Parameter Adaptation
di: Rios, Jesus, et al.
Pubblicazione: (2025)
di: Rios, Jesus, et al.
Pubblicazione: (2025)
Rho-1: Not All Tokens Are What You Need
di: Lin, Zhenghao, et al.
Pubblicazione: (2024)
di: Lin, Zhenghao, et al.
Pubblicazione: (2024)
Prediction Is All MoE Needs: Expert Load Distribution Goes from Fluctuating to Stabilizing
di: Cong, Peizhuang, et al.
Pubblicazione: (2024)
di: Cong, Peizhuang, et al.
Pubblicazione: (2024)
Training on the Benchmark Is Not All You Need
di: Ni, Shiwen, et al.
Pubblicazione: (2024)
di: Ni, Shiwen, et al.
Pubblicazione: (2024)
Contrast Is All You Need
di: Kilic, Burak, et al.
Pubblicazione: (2023)
di: Kilic, Burak, et al.
Pubblicazione: (2023)
Synthetic Data RL: Task Definition Is All You Need
di: Guo, Yiduo, et al.
Pubblicazione: (2025)
di: Guo, Yiduo, et al.
Pubblicazione: (2025)
More Agents Is All You Need
di: Li, Junyou, et al.
Pubblicazione: (2024)
di: Li, Junyou, et al.
Pubblicazione: (2024)
Is Depth All You Need? An Exploration of Iterative Reasoning in LLMs
di: Wu, Zongqian, et al.
Pubblicazione: (2025)
di: Wu, Zongqian, et al.
Pubblicazione: (2025)
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective
di: He, Shenghua, et al.
Pubblicazione: (2025)
di: He, Shenghua, et al.
Pubblicazione: (2025)
Not All Tokens Are What You Need In Thinking
di: Yuan, Hang, et al.
Pubblicazione: (2025)
di: Yuan, Hang, et al.
Pubblicazione: (2025)
Learning from the Best, Differently: A Diversity-Driven Rethinking on Data Selection
di: He, Hongyi, et al.
Pubblicazione: (2025)
di: He, Hongyi, et al.
Pubblicazione: (2025)
Agents Are All You Need for LLM Unlearning
di: Sanyal, Debdeep, et al.
Pubblicazione: (2025)
di: Sanyal, Debdeep, et al.
Pubblicazione: (2025)
Selected Languages are All You Need for Cross-lingual Truthfulness Transfer
di: Liu, Weihao, et al.
Pubblicazione: (2024)
di: Liu, Weihao, et al.
Pubblicazione: (2024)
Towards Supporting Legal Argumentation with NLP: Is More Data Really All You Need?
di: Santosh, T. Y. S. S, et al.
Pubblicazione: (2024)
di: Santosh, T. Y. S. S, et al.
Pubblicazione: (2024)
Soft Adaptive Policy Optimization
di: Gao, Chang, et al.
Pubblicazione: (2025)
di: Gao, Chang, et al.
Pubblicazione: (2025)
Length-Controlled Margin-Based Preference Optimization without Reference Model
di: Li, Gengxu, et al.
Pubblicazione: (2025)
di: Li, Gengxu, et al.
Pubblicazione: (2025)
Large Language Model Evaluation via Matrix Nuclear-Norm
di: Li, Yahan, et al.
Pubblicazione: (2024)
di: Li, Yahan, et al.
Pubblicazione: (2024)
Selective Prompt Anchoring for Code Generation
di: Tian, Yuan, et al.
Pubblicazione: (2024)
di: Tian, Yuan, et al.
Pubblicazione: (2024)
Attention Smoothing Is All You Need For Unlearning
di: Zade, Saleh Zare, et al.
Pubblicazione: (2026)
di: Zade, Saleh Zare, et al.
Pubblicazione: (2026)
SynthDST: Synthetic Data is All You Need for Few-Shot Dialog State Tracking
di: Kulkarni, Atharva, et al.
Pubblicazione: (2024)
di: Kulkarni, Atharva, et al.
Pubblicazione: (2024)
Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design
di: Sun, Lin, et al.
Pubblicazione: (2025)
di: Sun, Lin, et al.
Pubblicazione: (2025)
Boosting Explainability through Selective Rationalization in Pre-trained Language Models
di: Yuan, Libing, et al.
Pubblicazione: (2025)
di: Yuan, Libing, et al.
Pubblicazione: (2025)
Think When You Need: Self-Adaptive Chain-of-Thought Learning
di: Yang, Junjie, et al.
Pubblicazione: (2025)
di: Yang, Junjie, et al.
Pubblicazione: (2025)
Self-Verification is All You Need To Pass The Japanese Bar Examination
di: Shin, Andrew
Pubblicazione: (2026)
di: Shin, Andrew
Pubblicazione: (2026)
Unlocking the Power of LLM Uncertainty for Active In-Context Example Selection
di: Huang, Hsiu-Yuan, et al.
Pubblicazione: (2024)
di: Huang, Hsiu-Yuan, et al.
Pubblicazione: (2024)
Select2Reason: Efficient Instruction-Tuning Data Selection for Long-CoT Reasoning
di: Yang, Cehao, et al.
Pubblicazione: (2025)
di: Yang, Cehao, et al.
Pubblicazione: (2025)
SCALE: Selective Resource Allocation for Overcoming Performance Bottlenecks in Mathematical Test-time Scaling
di: Xiao, Yang, et al.
Pubblicazione: (2025)
di: Xiao, Yang, et al.
Pubblicazione: (2025)
Communication is All You Need: Persuasion Dataset Construction via Multi-LLM Communication
di: Ma, Weicheng, et al.
Pubblicazione: (2025)
di: Ma, Weicheng, et al.
Pubblicazione: (2025)
Scaling Law in LLM Simulated Personality: More Detailed and Realistic Persona Profile Is All You Need
di: Bai, Yuqi, et al.
Pubblicazione: (2025)
di: Bai, Yuqi, et al.
Pubblicazione: (2025)
Similarity is Not All You Need: Endowing Retrieval Augmented Generation with Multi Layered Thoughts
di: Gan, Chunjing, et al.
Pubblicazione: (2024)
di: Gan, Chunjing, et al.
Pubblicazione: (2024)
CrowdSelect: Synthetic Instruction Data Selection with Multi-LLM Wisdom
di: Li, Yisen, et al.
Pubblicazione: (2025)
di: Li, Yisen, et al.
Pubblicazione: (2025)
Attention Is All You Need for KV Cache in Diffusion LLMs
di: Nguyen-Tri, Quan, et al.
Pubblicazione: (2025)
di: Nguyen-Tri, Quan, et al.
Pubblicazione: (2025)
Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM
di: Lu, Xiaoding, et al.
Pubblicazione: (2024)
di: Lu, Xiaoding, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Language Models can Evaluate Themselves via Probability Discrepancy
di: Xia, Tingyu, et al.
Pubblicazione: (2024) -
A Survey of RWKV
di: Li, Zhiyuan, et al.
Pubblicazione: (2024) -
Tensor Product Attention Is All You Need
di: Zhang, Yifan, et al.
Pubblicazione: (2025) -
Not All Layers Need Tuning: Selective Layer Restoration Recovers Diversity
di: Zhang, Bowen, et al.
Pubblicazione: (2026) -
Not All Documents Are What You Need for Extracting Instruction Tuning Data
di: Zhang, Chi, et al.
Pubblicazione: (2025)