Unified Data Selection for LLM Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Xiaoyuan, Ma, Yubo, Li, Chengpeng, Zhu, Fengbin, Yu, Yiyao, Bao, Keqin, Wang, Wenjie, Feng, Fuli, Liu, Dayiheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation
di: Li, Xiaoyuan, et al.
Pubblicazione: (2025)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2025)
On Predicting the Post-training Potential of Pre-trained LLMs
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code
di: Bao, Keqin, et al.
Pubblicazione: (2025)
di: Bao, Keqin, et al.
Pubblicazione: (2025)
SAGE: Scalable Automated Robustness Augmentation for LLM Knowledge Evaluation
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
HellaSwag-Pro: A Large-Scale Bilingual Benchmark for Evaluating the Robustness of LLMs in Commonsense Reasoning
di: Li, Xiaoyuan, et al.
Pubblicazione: (2025)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2025)
MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning
di: Li, Xiaoyuan, et al.
Pubblicazione: (2025)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2025)
One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment
di: Cai, Hongru, et al.
Pubblicazione: (2026)
di: Cai, Hongru, et al.
Pubblicazione: (2026)
TAT-LLM: A Specialized Language Model for Discrete Reasoning over Tabular and Textual Data
di: Zhu, Fengbin, et al.
Pubblicazione: (2024)
di: Zhu, Fengbin, et al.
Pubblicazione: (2024)
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
di: Fang, Yi, et al.
Pubblicazione: (2026)
di: Fang, Yi, et al.
Pubblicazione: (2026)
CrAM: Credibility-Aware Attention Modification in LLMs for Combating Misinformation in RAG
di: Deng, Boyi, et al.
Pubblicazione: (2024)
di: Deng, Boyi, et al.
Pubblicazione: (2024)
Teaching Language Models to Reason with Tools
di: Li, Chengpeng, et al.
Pubblicazione: (2025)
di: Li, Chengpeng, et al.
Pubblicazione: (2025)
Think Twice Before Trusting: Self-Detection for Large Language Models through Comprehensive Answer Reflection
di: Li, Moxin, et al.
Pubblicazione: (2024)
di: Li, Moxin, et al.
Pubblicazione: (2024)
CoRT: Code-integrated Reasoning within Thinking
di: Li, Chengpeng, et al.
Pubblicazione: (2025)
di: Li, Chengpeng, et al.
Pubblicazione: (2025)
Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction
di: Li, Xiaoyuan, et al.
Pubblicazione: (2024)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2024)
Latent Inter-User Difference Modeling for LLM Personalization
di: Qiu, Yilun, et al.
Pubblicazione: (2025)
di: Qiu, Yilun, et al.
Pubblicazione: (2025)
Doc2SoarGraph: Discrete Reasoning over Visually-Rich Table-Text Documents via Semantic-Oriented Hierarchical Graphs
di: Zhu, Fengbin, et al.
Pubblicazione: (2023)
di: Zhu, Fengbin, et al.
Pubblicazione: (2023)
DotaMath: Decomposition of Thought with Code Assistance and Self-correction for Mathematical Reasoning
di: Li, Chengpeng, et al.
Pubblicazione: (2024)
di: Li, Chengpeng, et al.
Pubblicazione: (2024)
Optimizing Knowledge Integration in Retrieval-Augmented Generation with Self-Selection
di: Weng, Yan, et al.
Pubblicazione: (2025)
di: Weng, Yan, et al.
Pubblicazione: (2025)
PERM: Psychology-grounded Empathetic Reward Modeling for Large Language Models
di: Wang, Chengbing, et al.
Pubblicazione: (2026)
di: Wang, Chengbing, et al.
Pubblicazione: (2026)
Chain of Execution Supervision Promotes General Reasoning in Large Language Models
di: Chen, Nuo, et al.
Pubblicazione: (2025)
di: Chen, Nuo, et al.
Pubblicazione: (2025)
START: Self-taught Reasoner with Tools
di: Li, Chengpeng, et al.
Pubblicazione: (2025)
di: Li, Chengpeng, et al.
Pubblicazione: (2025)
Importance-Aware Data Selection for Efficient LLM Instruction Tuning
di: Jiang, Tingyu, et al.
Pubblicazione: (2025)
di: Jiang, Tingyu, et al.
Pubblicazione: (2025)
Prospect Personalized Recommendation on Large Language Model-based Agent Platform
di: Zhang, Jizhi, et al.
Pubblicazione: (2024)
di: Zhang, Jizhi, et al.
Pubblicazione: (2024)
Navigating Through Paper Flood: Advancing LLM-based Paper Evaluation through Domain-Aware Retrieval and Latent Reasoning
di: Zheng, Wuqiang, et al.
Pubblicazione: (2025)
di: Zheng, Wuqiang, et al.
Pubblicazione: (2025)
IGD: Token Decisiveness Modeling via Information Gain in LLMs for Personalized Recommendation
di: Lin, Zijie, et al.
Pubblicazione: (2025)
di: Lin, Zijie, et al.
Pubblicazione: (2025)
LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection
di: Wu, Jian, et al.
Pubblicazione: (2025)
di: Wu, Jian, et al.
Pubblicazione: (2025)
Counterfactual Debating with Preset Stances for Hallucination Elimination of LLMs
di: Fang, Yi, et al.
Pubblicazione: (2024)
di: Fang, Yi, et al.
Pubblicazione: (2024)
Assistant-Guided Mitigation of Teacher Preference Bias in LLM-as-a-Judge
di: Liu, Zhuo, et al.
Pubblicazione: (2025)
di: Liu, Zhuo, et al.
Pubblicazione: (2025)
TacoMAS: Test-Time Co-Evolution of Topology and Capability in LLM-based Multi-Agent Systems
di: Xu, Chen, et al.
Pubblicazione: (2026)
di: Xu, Chen, et al.
Pubblicazione: (2026)
Less is More: Improving LLM Alignment via Preference Data Selection
di: Deng, Xun, et al.
Pubblicazione: (2025)
di: Deng, Xun, et al.
Pubblicazione: (2025)
Think-While-Generating: On-the-Fly Reasoning for Personalized Long-Form Generation
di: Wang, Chengbing, et al.
Pubblicazione: (2025)
di: Wang, Chengbing, et al.
Pubblicazione: (2025)
UniICL: An Efficient Unified Framework Unifying Compression, Selection, and Generation
di: Gao, Jun, et al.
Pubblicazione: (2024)
di: Gao, Jun, et al.
Pubblicazione: (2024)
Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking
di: Li, Mingxin, et al.
Pubblicazione: (2026)
di: Li, Mingxin, et al.
Pubblicazione: (2026)
Large Language Models Empowered Personalized Web Agents
di: Cai, Hongru, et al.
Pubblicazione: (2024)
di: Cai, Hongru, et al.
Pubblicazione: (2024)
CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning
di: Zhu, Xinyu, et al.
Pubblicazione: (2026)
di: Zhu, Xinyu, et al.
Pubblicazione: (2026)
On the Step Length Confounding in LLM Reasoning Data Selection
di: Wang, Bing, et al.
Pubblicazione: (2026)
di: Wang, Bing, et al.
Pubblicazione: (2026)
CALM Before the STORM: Unlocking Native Reasoning for Optimization Modeling
di: Tang, Zhengyang, et al.
Pubblicazione: (2025)
di: Tang, Zhengyang, et al.
Pubblicazione: (2025)
SASFT: Sparse Autoencoder-guided Supervised Finetuning to Mitigate Unexpected Code-Switching in LLMs
di: Deng, Boyi, et al.
Pubblicazione: (2025)
di: Deng, Boyi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026) -
SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026) -
MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation
di: Li, Xiaoyuan, et al.
Pubblicazione: (2025) -
On Predicting the Post-training Potential of Pre-trained LLMs
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026) -
Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code
di: Bao, Keqin, et al.
Pubblicazione: (2025)