RHyVE: Competence-Aware Verification and Phase-Aware Deployment for LLM-Generated Reward Hypotheses
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Feiyu, Zheng, Xu, Wang, Zhuocheng, Dai, Yi ming, Li, Hui |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
NH-CROP: Robust Pricing for Governed Language Data Assets under Cost Uncertainty
di: Zheng, Xu, et al.
Pubblicazione: (2026)
di: Zheng, Xu, et al.
Pubblicazione: (2026)
SCARV: Structure-Constrained Aggregation for Stable Sample Ranking in Redundant NLP Datasets
di: Zheng, Xu, et al.
Pubblicazione: (2026)
di: Zheng, Xu, et al.
Pubblicazione: (2026)
Grounding Generative Planners in Verifiable Logic: A Hybrid Architecture for Trustworthy Embodied AI
di: Wu, Feiyu, et al.
Pubblicazione: (2026)
di: Wu, Feiyu, et al.
Pubblicazione: (2026)
LLM With Tools: A Survey
di: Shen, Zhuocheng
Pubblicazione: (2024)
di: Shen, Zhuocheng
Pubblicazione: (2024)
PEAR: Phase Entropy Aware Reward for Efficient Reasoning
di: Huang, Chen, et al.
Pubblicazione: (2025)
di: Huang, Chen, et al.
Pubblicazione: (2025)
Towards Socially and Morally Aware RL agent: Reward Design With LLM
di: Wang, Zhaoyue
Pubblicazione: (2024)
di: Wang, Zhaoyue
Pubblicazione: (2024)
HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment
di: Mekky, Ali, et al.
Pubblicazione: (2025)
di: Mekky, Ali, et al.
Pubblicazione: (2025)
Teaching According to Talents! Instruction Tuning LLMs with Competence-Aware Curriculum Learning
di: Li, Yangning, et al.
Pubblicazione: (2025)
di: Li, Yangning, et al.
Pubblicazione: (2025)
Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment
di: Zhang, Hongbin, et al.
Pubblicazione: (2025)
di: Zhang, Hongbin, et al.
Pubblicazione: (2025)
VASR: Variance-Aware Systematic Resampling for Reward-Guided Diffusion
di: Shekhar, Shivanshu, et al.
Pubblicazione: (2026)
di: Shekhar, Shivanshu, et al.
Pubblicazione: (2026)
Persona-Aware Alignment Framework for Personalized Dialogue Generation
di: Li, Guanrong, et al.
Pubblicazione: (2025)
di: Li, Guanrong, et al.
Pubblicazione: (2025)
What Are Research Hypotheses?
di: Wu, Jian, et al.
Pubblicazione: (2025)
di: Wu, Jian, et al.
Pubblicazione: (2025)
Fairness Aware Reward Optimization
di: Choi, Ching Lam, et al.
Pubblicazione: (2026)
di: Choi, Ching Lam, et al.
Pubblicazione: (2026)
Planner-Centric Reinforcement Learning for Deep Research with Structure-Aware Reward
di: Hussain, Mustafa Anis, et al.
Pubblicazione: (2026)
di: Hussain, Mustafa Anis, et al.
Pubblicazione: (2026)
Expected Value Alignment for Generative Reward Modeling in Formal Mathematics Verification
di: Ji, Shihao, et al.
Pubblicazione: (2026)
di: Ji, Shihao, et al.
Pubblicazione: (2026)
Distribution-Aware Reward: Reinforcement Learning over Predictive Distributions for LLM Regression
di: Park, Jungsoo, et al.
Pubblicazione: (2026)
di: Park, Jungsoo, et al.
Pubblicazione: (2026)
V-CAGE: Context-Aware Generation and Verification for Scalable Long-Horizon Embodied Tasks
di: Liu, Yaru, et al.
Pubblicazione: (2026)
di: Liu, Yaru, et al.
Pubblicazione: (2026)
Uncertainty-Aware Reward-Free Exploration with General Function Approximation
di: Zhang, Junkai, et al.
Pubblicazione: (2024)
di: Zhang, Junkai, et al.
Pubblicazione: (2024)
GroundedPRM: Tree-Guided and Fidelity-Aware Process Reward Modeling for Step-Level Reasoning
di: Zhang, Yao, et al.
Pubblicazione: (2025)
di: Zhang, Yao, et al.
Pubblicazione: (2025)
Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation
di: Thebaud, Thomas, et al.
Pubblicazione: (2026)
di: Thebaud, Thomas, et al.
Pubblicazione: (2026)
Co-Sight: Enhancing LLM-Based Agents via Conflict-Aware Meta-Verification and Trustworthy Reasoning with Structured Facts
di: Zhang, Hongwei, et al.
Pubblicazione: (2025)
di: Zhang, Hongwei, et al.
Pubblicazione: (2025)
SeqRoute: Global Budget-Aware Sequential LLM Routing via Offline Reinforcement Learning
di: Xu, Zhongling, et al.
Pubblicazione: (2026)
di: Xu, Zhongling, et al.
Pubblicazione: (2026)
Generating Diverse Hypotheses for Inductive Reasoning
di: Lee, Kang-il, et al.
Pubblicazione: (2024)
di: Lee, Kang-il, et al.
Pubblicazione: (2024)
Call-Chain-Aware LLM-Based Test Generation for Java Projects
di: Wang, Guancheng, et al.
Pubblicazione: (2026)
di: Wang, Guancheng, et al.
Pubblicazione: (2026)
Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks
di: Luo, Zhimeng, et al.
Pubblicazione: (2025)
di: Luo, Zhimeng, et al.
Pubblicazione: (2025)
GRAIL: AI translation for scientists application workflow on satellite data
di: Shang, Zhuocheng, et al.
Pubblicazione: (2026)
di: Shang, Zhuocheng, et al.
Pubblicazione: (2026)
DenoiseFlow: Uncertainty-Aware Denoising for Reliable LLM Agentic Workflows
di: Yan, Yandong, et al.
Pubblicazione: (2026)
di: Yan, Yandong, et al.
Pubblicazione: (2026)
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
di: Zhang, Tuo, et al.
Pubblicazione: (2025)
di: Zhang, Tuo, et al.
Pubblicazione: (2025)
Scaling-Aware Adapter for Structure-Grounded LLM Reasoning
di: Jing, Zihao, et al.
Pubblicazione: (2026)
di: Jing, Zihao, et al.
Pubblicazione: (2026)
FairAgent: Democratizing Fairness-Aware Machine Learning with LLM-Powered Agents
di: Dai, Yucong, et al.
Pubblicazione: (2025)
di: Dai, Yucong, et al.
Pubblicazione: (2025)
RHiOTS: A Framework for Evaluating Hierarchical Time Series Forecasting Algorithms
di: Roque, Luis, et al.
Pubblicazione: (2024)
di: Roque, Luis, et al.
Pubblicazione: (2024)
Graph of Verification: Structured Verification of LLM Reasoning with Directed Acyclic Graphs
di: Fang, Jiwei, et al.
Pubblicazione: (2025)
di: Fang, Jiwei, et al.
Pubblicazione: (2025)
Boosting Universal LLM Reward Design through Heuristic Reward Observation Space Evolution
di: Heng, Zen Kit, et al.
Pubblicazione: (2025)
di: Heng, Zen Kit, et al.
Pubblicazione: (2025)
EmoMAS: Emotion-Aware Multi-Agent System for High-Stakes Edge-Deployable Negotiation with Bayesian Orchestration
di: Long, Yunbo, et al.
Pubblicazione: (2026)
di: Long, Yunbo, et al.
Pubblicazione: (2026)
PositionID: LLMs can Control Lengths, Copy and Paste with Explicit Positional Awareness
di: Wang, Zekun, et al.
Pubblicazione: (2024)
di: Wang, Zekun, et al.
Pubblicazione: (2024)
Planning with Sketch-Guided Verification for Physics-Aware Video Generation
di: Huang, Yidong, et al.
Pubblicazione: (2025)
di: Huang, Yidong, et al.
Pubblicazione: (2025)
Dafny as Verification-Aware Intermediate Language for Code Generation
di: Li, Yue Chen, et al.
Pubblicazione: (2025)
di: Li, Yue Chen, et al.
Pubblicazione: (2025)
Context-Aware Probabilistic Modeling with LLM for Multimodal Time Series Forecasting
di: Yao, Yueyang, et al.
Pubblicazione: (2025)
di: Yao, Yueyang, et al.
Pubblicazione: (2025)
Verification-Aware Planning for Multi-Agent Systems
di: Xu, Tianyang, et al.
Pubblicazione: (2025)
di: Xu, Tianyang, et al.
Pubblicazione: (2025)
HESTIA: A Hessian-Guided Differentiable Quantization-Aware Training Framework for Extremely Low-Bit LLMs
di: Wang, Guoan, et al.
Pubblicazione: (2026)
di: Wang, Guoan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
NH-CROP: Robust Pricing for Governed Language Data Assets under Cost Uncertainty
di: Zheng, Xu, et al.
Pubblicazione: (2026) -
SCARV: Structure-Constrained Aggregation for Stable Sample Ranking in Redundant NLP Datasets
di: Zheng, Xu, et al.
Pubblicazione: (2026) -
Grounding Generative Planners in Verifiable Logic: A Hybrid Architecture for Trustworthy Embodied AI
di: Wu, Feiyu, et al.
Pubblicazione: (2026) -
LLM With Tools: A Survey
di: Shen, Zhuocheng
Pubblicazione: (2024) -
PEAR: Phase Entropy Aware Reward for Efficient Reasoning
di: Huang, Chen, et al.
Pubblicazione: (2025)