Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Bi, Baolong, Liu, Shenghua, Wang, Yiwei, Tong, Siqian, Mei, Lingrui, Ge, Yuyao, Xu, Yilong, Guo, Jiafeng, Cheng, Xueqi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Innate Reasoning is Not Enough: In-Context Learning Enhances Reasoning Large Language Models with Less Overthinking
by: Ge, Yuyao, et al.
Published: (2025)
by: Ge, Yuyao, et al.
Published: (2025)
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
by: Ge, Yuyao, et al.
Published: (2025)
by: Ge, Yuyao, et al.
Published: (2025)
Prism-$Δ$: Differential Subspace Steering for Prompt Highlighting in Large Language Models
by: Ge, Yuyao, et al.
Published: (2026)
by: Ge, Yuyao, et al.
Published: (2026)
SLANG: New Concept Comprehension of Large Language Models
by: Mei, Lingrui, et al.
Published: (2024)
by: Mei, Lingrui, et al.
Published: (2024)
LPNL: Scalable Link Prediction with Large Language Models
by: Bi, Baolong, et al.
Published: (2024)
by: Bi, Baolong, et al.
Published: (2024)
Gated Differentiable Working Memory for Long-Context Language Modeling
by: Mei, Lingrui, et al.
Published: (2026)
by: Mei, Lingrui, et al.
Published: (2026)
Parameters vs. Context: Fine-Grained Control of Knowledge Reliance in Language Models
by: Bi, Baolong, et al.
Published: (2025)
by: Bi, Baolong, et al.
Published: (2025)
Adaptive Token Biaser: Knowledge Editing via Biasing Key Entities
by: Bi, Baolong, et al.
Published: (2024)
by: Bi, Baolong, et al.
Published: (2024)
Can Graph Descriptive Order Affect Solving Graph Problems with LLMs?
by: Ge, Yuyao, et al.
Published: (2024)
by: Ge, Yuyao, et al.
Published: (2024)
Who is in the Spotlight: The Hidden Bias Undermining Multimodal Retrieval-Augmented Generation
by: Yao, Jiayu, et al.
Published: (2025)
by: Yao, Jiayu, et al.
Published: (2025)
a1: Steep Test-time Scaling Law via Environment Augmented Generation
by: Mei, Lingrui, et al.
Published: (2025)
by: Mei, Lingrui, et al.
Published: (2025)
"Not Aligned" is Not "Malicious": Being Careful about Hallucinations of Large Language Models' Jailbreak
by: Mei, Lingrui, et al.
Published: (2024)
by: Mei, Lingrui, et al.
Published: (2024)
Decoding by Contrasting Knowledge: Enhancing LLMs' Confidence on Edited Facts
by: Bi, Baolong, et al.
Published: (2024)
by: Bi, Baolong, et al.
Published: (2024)
HiddenGuard: Fine-Grained Safe Generation with Specialized Representation Router
by: Mei, Lingrui, et al.
Published: (2024)
by: Mei, Lingrui, et al.
Published: (2024)
PromptCD: Test-Time Behavior Enhancement via Polarity-Prompt Contrastive Decoding
by: Bi, Baolong, et al.
Published: (2026)
by: Bi, Baolong, et al.
Published: (2026)
StruEdit: Structured Outputs Enable the Fast and Accurate Knowledge Editing for Large Language Models
by: Bi, Baolong, et al.
Published: (2024)
by: Bi, Baolong, et al.
Published: (2024)
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs
by: Bi, Baolong, et al.
Published: (2025)
by: Bi, Baolong, et al.
Published: (2025)
Not in Sync: Unveiling Temporal Bias in Audio Chat Models
by: Yao, Jiayu, et al.
Published: (2025)
by: Yao, Jiayu, et al.
Published: (2025)
Is Factuality Enhancement a Free Lunch For LLMs? Better Factuality Can Lead to Worse Context-Faithfulness
by: Bi, Baolong, et al.
Published: (2024)
by: Bi, Baolong, et al.
Published: (2024)
HighlightBench: Benchmarking Markup-Driven Table Reasoning in Scientific Documents
by: Wang, Lexin, et al.
Published: (2026)
by: Wang, Lexin, et al.
Published: (2026)
A Survey of Vibe Coding with Large Language Models
by: Ge, Yuyao, et al.
Published: (2025)
by: Ge, Yuyao, et al.
Published: (2025)
AuTAgent: A Reinforcement Learning Framework for Tool-Augmented Audio Reasoning
by: Tong, Siqian, et al.
Published: (2026)
by: Tong, Siqian, et al.
Published: (2026)
MORE: Multi-mOdal REtrieval Augmented Generative Commonsense Reasoning
by: Cui, Wanqing, et al.
Published: (2024)
by: Cui, Wanqing, et al.
Published: (2024)
Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary Perception
by: Ni, Shiyu, et al.
Published: (2025)
by: Ni, Shiyu, et al.
Published: (2025)
A Survey of Context Engineering for Large Language Models
by: Mei, Lingrui, et al.
Published: (2025)
by: Mei, Lingrui, et al.
Published: (2025)
Bridging Queries and Tables through Entities in Table Retrieval
by: Li, Da, et al.
Published: (2025)
by: Li, Da, et al.
Published: (2025)
Reproducibility Analysis and Enhancements for Multi-Aspect Dense Retriever with Aspect Learning
by: Bi, Keping, et al.
Published: (2024)
by: Bi, Keping, et al.
Published: (2024)
ALiiCE: Evaluating Positional Fine-grained Citation Generation
by: Xu, Yilong, et al.
Published: (2024)
by: Xu, Yilong, et al.
Published: (2024)
Rethinking All Evidence: Enhancing Trustworthy Retrieval-Augmented Generation via Conflict-Driven Summarization
by: Chen, Juan, et al.
Published: (2025)
by: Chen, Juan, et al.
Published: (2025)
Estimating Commonsense Plausibility through Semantic Shifts
by: Cui, Wanqing, et al.
Published: (2025)
by: Cui, Wanqing, et al.
Published: (2025)
An Iterative Utility Judgment Framework Inspired by Philosophical Relevance via LLMs
by: Zhang, Hengran, et al.
Published: (2024)
by: Zhang, Hengran, et al.
Published: (2024)
Tailoring Table Retrieval from a Field-aware Hybrid Matching Perspective
by: Li, Da, et al.
Published: (2025)
by: Li, Da, et al.
Published: (2025)
How Knowledge Popularity Influences and Enhances LLM Knowledge Boundary Perception
by: Ni, Shiyu, et al.
Published: (2025)
by: Ni, Shiyu, et al.
Published: (2025)
CIR at the NTCIR-17 ULTRE-2 Task
by: Yu, Lulu, et al.
Published: (2023)
by: Yu, Lulu, et al.
Published: (2023)
When Do LLMs Need Retrieval Augmentation? Mitigating LLMs' Overconfidence Helps Retrieval Augmentation
by: Ni, Shiyu, et al.
Published: (2024)
by: Ni, Shiyu, et al.
Published: (2024)
Iterative Structured Pruning for Large Language Models with Multi-Domain Calibration
by: Wu, Guangxin, et al.
Published: (2026)
by: Wu, Guangxin, et al.
Published: (2026)
Training a Utility-based Retriever Through Shared Context Attribution for Retrieval-Augmented Language Models
by: Xu, Yilong, et al.
Published: (2025)
by: Xu, Yilong, et al.
Published: (2025)
Classifier Guidance Enhances Diffusion-based Adversarial Purification by Preserving Predictive Information
by: Zhang, Mingkun, et al.
Published: (2024)
by: Zhang, Mingkun, et al.
Published: (2024)
MVAM: Multi-View Attention Method for Fine-grained Image-Text Matching
by: Cui, Wanqing, et al.
Published: (2024)
by: Cui, Wanqing, et al.
Published: (2024)
Generating Data-Driven Reasoning Rubrics for Domain-Adaptive Reward Modeling
by: Sanders, Kate, et al.
Published: (2026)
by: Sanders, Kate, et al.
Published: (2026)
Similar Items
-
Innate Reasoning is Not Enough: In-Context Learning Enhances Reasoning Large Language Models with Less Overthinking
by: Ge, Yuyao, et al.
Published: (2025) -
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
by: Ge, Yuyao, et al.
Published: (2025) -
Prism-$Δ$: Differential Subspace Steering for Prompt Highlighting in Large Language Models
by: Ge, Yuyao, et al.
Published: (2026) -
SLANG: New Concept Comprehension of Large Language Models
by: Mei, Lingrui, et al.
Published: (2024) -
LPNL: Scalable Link Prediction with Large Language Models
by: Bi, Baolong, et al.
Published: (2024)