Learn More with Less: Uncertainty Consistency Guided Query Selection for RLVR
Fuente:
arXiv
Saved in:
| Main Authors: | Yi, Hao, Hu, Yulan, Li, Xin, Ouyang, Sheng, Ding, Lizhong, Liu, Yong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VIGraph: Generative Self-supervised Learning for Class-Imbalanced Node Classification
by: Hu, Yulan, et al.
Published: (2023)
by: Hu, Yulan, et al.
Published: (2023)
Refining Latent Representations: A Generative SSL Approach for Heterogeneous Graph Learning
by: Hu, Yulan, et al.
Published: (2023)
by: Hu, Yulan, et al.
Published: (2023)
Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning
by: Hu, Yulan, et al.
Published: (2025)
by: Hu, Yulan, et al.
Published: (2025)
GUNDAM: Aligning Large Language Models with Graph Understanding
by: Ouyang, Sheng, et al.
Published: (2024)
by: Ouyang, Sheng, et al.
Published: (2024)
Preserving Node Distinctness in Graph Autoencoders via Similarity Distillation
by: Chen, Ge, et al.
Published: (2024)
by: Chen, Ge, et al.
Published: (2024)
No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning
by: Li, Zhicong, et al.
Published: (2026)
by: Li, Zhicong, et al.
Published: (2026)
SPPD: Self-training with Process Preference Learning Using Dynamic Value Margin
by: Yi, Hao, et al.
Published: (2025)
by: Yi, Hao, et al.
Published: (2025)
Towards Reward Fairness in RLHF: From a Resource Allocation Perspective
by: Ouyang, Sheng, et al.
Published: (2025)
by: Ouyang, Sheng, et al.
Published: (2025)
Beyond Variance: Prompt-Efficient RLVR via Rare-Event Amplification and Bidirectional Pairing
by: Pang, Yujuan, et al.
Published: (2026)
by: Pang, Yujuan, et al.
Published: (2026)
Exploring Task Unification in Graph Representation Learning via Generative Approach
by: Hu, Yulan, et al.
Published: (2024)
by: Hu, Yulan, et al.
Published: (2024)
Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes
by: Bauer, Justin, et al.
Published: (2026)
by: Bauer, Justin, et al.
Published: (2026)
Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning
by: Li, Xuchen, et al.
Published: (2025)
by: Li, Xuchen, et al.
Published: (2025)
Learning More from Less: Unlocking Internal Representations for Benchmark Compression
by: Zhang, Yueqi, et al.
Published: (2026)
by: Zhang, Yueqi, et al.
Published: (2026)
LIMI: Less is More for Agency
by: Xiao, Yang, et al.
Published: (2025)
by: Xiao, Yang, et al.
Published: (2025)
Less is More for Improving Automatic Evaluation of Factual Consistency
by: Wang, Tong, et al.
Published: (2024)
by: Wang, Tong, et al.
Published: (2024)
Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
by: Li, Xuchen, et al.
Published: (2025)
by: Li, Xuchen, et al.
Published: (2025)
Learn More, Forget Less: A Gradient-Aware Data Selection Approach for LLM
by: Liu, Yibai, et al.
Published: (2025)
by: Liu, Yibai, et al.
Published: (2025)
Graph Ranking Contrastive Learning: A Extremely Simple yet Efficient Method
by: Hu, Yulan, et al.
Published: (2023)
by: Hu, Yulan, et al.
Published: (2023)
Less is More: Efficient Brain-Inspired Learning for Autonomous Driving Trajectory Prediction
by: Liao, Haicheng, et al.
Published: (2024)
by: Liao, Haicheng, et al.
Published: (2024)
LIMO: Less is More for Reasoning
by: Ye, Yixin, et al.
Published: (2025)
by: Ye, Yixin, et al.
Published: (2025)
Towards Comprehensive Preference Data Collection for Reward Modeling
by: Hu, Yulan, et al.
Published: (2024)
by: Hu, Yulan, et al.
Published: (2024)
Less is More: Improving LLM Reasoning with Minimal Test-Time Intervention
by: Yang, Zhen, et al.
Published: (2025)
by: Yang, Zhen, et al.
Published: (2025)
LIMR: Less is More for RL Scaling
by: Li, Xuefeng, et al.
Published: (2025)
by: Li, Xuefeng, et al.
Published: (2025)
Less is More: Selective Reflection for Compatible and Efficient Knowledge Distillation in Large Language Models
by: Liu, Lingyuan, et al.
Published: (2025)
by: Liu, Lingyuan, et al.
Published: (2025)
Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning
by: Xie, Can, et al.
Published: (2025)
by: Xie, Can, et al.
Published: (2025)
Gradient Coupling: The Hidden Barrier to Generalization in Agentic Reinforcement Learning
by: Liu, Jingyu, et al.
Published: (2025)
by: Liu, Jingyu, et al.
Published: (2025)
Less is More: Resource-Efficient Low-Rank Adaptation
by: Tian, Chunlin, et al.
Published: (2025)
by: Tian, Chunlin, et al.
Published: (2025)
Less is More: Multimodal Region Representation via Pairwise Inter-view Learning
by: Namgung, Min, et al.
Published: (2025)
by: Namgung, Min, et al.
Published: (2025)
Transformer Multivariate Forecasting: Less is More?
by: Xu, Jingjing, et al.
Published: (2023)
by: Xu, Jingjing, et al.
Published: (2023)
Less is More: Label-Guided Summarization of Procedural and Instructional Videos
by: Rajpal, Shreya, et al.
Published: (2026)
by: Rajpal, Shreya, et al.
Published: (2026)
The Path Not Taken: RLVR Provably Learns Off the Principals
by: Zhu, Hanqing, et al.
Published: (2025)
by: Zhu, Hanqing, et al.
Published: (2025)
Forget Less, Generalize More: Unifying Temporal and Structural Adaptation for Dynamic Graphs
by: Chang, Qian, et al.
Published: (2026)
by: Chang, Qian, et al.
Published: (2026)
Compositional Generalization from Learned Skills via CoT Training: A Theoretical and Structural Analysis for Reasoning
by: Yao, Xinhao, et al.
Published: (2025)
by: Yao, Xinhao, et al.
Published: (2025)
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
by: Liu, Chengzhi, et al.
Published: (2025)
by: Liu, Chengzhi, et al.
Published: (2025)
Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
Learning More with Less: A Dynamic Dual-Level Down-Sampling Framework for Efficient Policy Optimization
by: Wang, Chao, et al.
Published: (2025)
by: Wang, Chao, et al.
Published: (2025)
Train Less, Learn More: Adaptive Efficient Rollout Optimization for Group-Based Reinforcement Learning
by: Zhang, Zhi, et al.
Published: (2026)
by: Zhang, Zhi, et al.
Published: (2026)
NeocorRAG: Less Irrelevant Information, More Explicit Evidence, and More Effective Recall via Evidence Chains
by: Peng, Shiyao, et al.
Published: (2026)
by: Peng, Shiyao, et al.
Published: (2026)
LIME: Less Is More for MLLM Evaluation
by: Zhu, King, et al.
Published: (2024)
by: Zhu, King, et al.
Published: (2024)
Less Noise, More Voice: Reinforcement Learning for Reasoning via Instruction Purification
by: Guo, Yiju, et al.
Published: (2026)
by: Guo, Yiju, et al.
Published: (2026)
Similar Items
-
VIGraph: Generative Self-supervised Learning for Class-Imbalanced Node Classification
by: Hu, Yulan, et al.
Published: (2023) -
Refining Latent Representations: A Generative SSL Approach for Heterogeneous Graph Learning
by: Hu, Yulan, et al.
Published: (2023) -
Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning
by: Hu, Yulan, et al.
Published: (2025) -
GUNDAM: Aligning Large Language Models with Graph Understanding
by: Ouyang, Sheng, et al.
Published: (2024) -
Preserving Node Distinctness in Graph Autoencoders via Similarity Distillation
by: Chen, Ge, et al.
Published: (2024)