Evaluate-as-Action: Self-Evaluated Process Rewards for Retrieval-Augmented Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shu, Jiangming, Zhang, Yuxiang, Ma, Ye, Lin, Xueyuan, Sang, Jitao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2025)
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2025)
Agent models: Internalizing Chain-of-Action Generation into Reasoning models
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2025)
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2025)
OpenRFT: Adapting Reasoning Foundation Model for Domain-specific Tasks with Reinforcement Fine-Tuning
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2024)
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2024)
o1-Coder: an o1 Replication for Coding
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2024)
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2024)
Retrieval-Augmented Process Reward Model for Generalizable Mathematical Reasoning
von: Zhu, Jiachen, et al.
Veröffentlicht: (2025)
von: Zhu, Jiachen, et al.
Veröffentlicht: (2025)
Named Entity Recognition in COVID-19 tweets with Entity Knowledge Augmentation
von: Zhang, Xuankang, et al.
Veröffentlicht: (2025)
von: Zhang, Xuankang, et al.
Veröffentlicht: (2025)
CSPO: Alleviating Reward Ambiguity for Structured Table-to-LaTeX Generation
von: Yang, Yunfan, et al.
Veröffentlicht: (2026)
von: Yang, Yunfan, et al.
Veröffentlicht: (2026)
KG-FPQ: Evaluating Factuality Hallucination in LLMs with Knowledge Graph-based False Premise Questions
von: Zhu, Yanxu, et al.
Veröffentlicht: (2024)
von: Zhu, Yanxu, et al.
Veröffentlicht: (2024)
A Disguised Wolf Is More Harmful Than a Toothless Tiger: Adaptive Malicious Code Injection Backdoor Attack Leveraging User Behavior as Triggers
von: Wu, Shangxi, et al.
Veröffentlicht: (2024)
von: Wu, Shangxi, et al.
Veröffentlicht: (2024)
WebSynthesis: World-Model-Guided MCTS for Efficient WebUI-Trajectory Synthesis
von: Gao, Yifei, et al.
Veröffentlicht: (2025)
von: Gao, Yifei, et al.
Veröffentlicht: (2025)
GUITestScape: Towards Open-set Evaluation on Exploratory GUI Testing
von: Chen, Xiaoyi, et al.
Veröffentlicht: (2026)
von: Chen, Xiaoyi, et al.
Veröffentlicht: (2026)
ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents
von: Li, Dawei, et al.
Veröffentlicht: (2026)
von: Li, Dawei, et al.
Veröffentlicht: (2026)
Deepchecks: Evaluating Retrieval-Augmented Generation (RAG)
von: Gerner, Assaf, et al.
Veröffentlicht: (2026)
von: Gerner, Assaf, et al.
Veröffentlicht: (2026)
Evaluation of Retrieval-Augmented Generation: A Survey
von: Yu, Hao, et al.
Veröffentlicht: (2024)
von: Yu, Hao, et al.
Veröffentlicht: (2024)
HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation
von: Wu, Peilin, et al.
Veröffentlicht: (2025)
von: Wu, Peilin, et al.
Veröffentlicht: (2025)
GUITester: Enabling GUI Agents for Exploratory Defect Discovery
von: Gao, Yifei, et al.
Veröffentlicht: (2026)
von: Gao, Yifei, et al.
Veröffentlicht: (2026)
Hybrid Differential Reward: Combining Temporal Difference and Action Gradients for Efficient Multi-Agent Reinforcement Learning in Cooperative Driving
von: Han, Ye, et al.
Veröffentlicht: (2025)
von: Han, Ye, et al.
Veröffentlicht: (2025)
Self-Guided Defense: Adaptive Safety Alignment for Reasoning Models via Synthesized Guidelines
von: Wang, Yuhang, et al.
Veröffentlicht: (2025)
von: Wang, Yuhang, et al.
Veröffentlicht: (2025)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
von: Lù, Xing Han, et al.
Veröffentlicht: (2025)
von: Lù, Xing Han, et al.
Veröffentlicht: (2025)
Exploring the Privacy Protection Capabilities of Chinese Large Language Models
von: Yang, Yuqi, et al.
Veröffentlicht: (2024)
von: Yang, Yuqi, et al.
Veröffentlicht: (2024)
Reasoning Shapes Alignment: Investigating Cultural Alignment in Large Reasoning Models with Cultural Norms
von: Wang, Yuhang, et al.
Veröffentlicht: (2025)
von: Wang, Yuhang, et al.
Veröffentlicht: (2025)
How Reliable is Your Simulator? Analysis on the Limitations of Current LLM-based User Simulators for Conversational Recommendation
von: Zhu, Lixi, et al.
Veröffentlicht: (2024)
von: Zhu, Lixi, et al.
Veröffentlicht: (2024)
ReInAgent: A Context-Aware GUI Agent Enabling Human-in-the-Loop Mobile Task Navigation
von: Jia, Haitao, et al.
Veröffentlicht: (2025)
von: Jia, Haitao, et al.
Veröffentlicht: (2025)
NAP-Tuning: Neural Augmented Prompt Tuning for Adversarially Robust Vision-Language Models
von: Zhang, Jiaming, et al.
Veröffentlicht: (2025)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2025)
AnyAttack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-language Models
von: Zhang, Jiaming, et al.
Veröffentlicht: (2024)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2024)
RewardHackingAgents: Benchmarking Evaluation Integrity for LLM ML-Engineering Agents
von: Atinafu, Yonas, et al.
Veröffentlicht: (2026)
von: Atinafu, Yonas, et al.
Veröffentlicht: (2026)
DICE: Discrete Interpretable Comparative Evaluation with Probabilistic Scoring for Retrieval-Augmented Generation
von: Liu, Shiyan, et al.
Veröffentlicht: (2025)
von: Liu, Shiyan, et al.
Veröffentlicht: (2025)
RPM-MCTS: Knowledge-Retrieval as Process Reward Model with Monte Carlo Tree Search for Code Generation
von: Lin, Yuanyuan, et al.
Veröffentlicht: (2025)
von: Lin, Yuanyuan, et al.
Veröffentlicht: (2025)
Retrieval Augmented Generation (RAG) for Fintech: Agentic Design and Evaluation
von: Cook, Thomas, et al.
Veröffentlicht: (2025)
von: Cook, Thomas, et al.
Veröffentlicht: (2025)
StepMathAgent: A Step-Wise Agent for Evaluating Mathematical Processes through Tree-of-Error
von: Yang, Shu-Xun, et al.
Veröffentlicht: (2025)
von: Yang, Shu-Xun, et al.
Veröffentlicht: (2025)
Privacy in Action: Towards Realistic Privacy Mitigation and Evaluation for LLM-Powered Agents
von: Wang, Shouju, et al.
Veröffentlicht: (2025)
von: Wang, Shouju, et al.
Veröffentlicht: (2025)
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
Inference-Time Rule Eraser: Fair Recognition via Distilling and Removing Biased Rules
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
GUI-PRA: Process Reward Agent for GUI Tasks
von: Xiong, Tao, et al.
Veröffentlicht: (2025)
von: Xiong, Tao, et al.
Veröffentlicht: (2025)
Evaluating Retrieval-Augmented Generation Agents for Autonomous Scientific Discovery in Astrophysics
von: Xu, Xueqing, et al.
Veröffentlicht: (2025)
von: Xu, Xueqing, et al.
Veröffentlicht: (2025)
Process-based Self-Rewarding Language Models
von: Zhang, Shimao, et al.
Veröffentlicht: (2025)
von: Zhang, Shimao, et al.
Veröffentlicht: (2025)
ReasonSTL: Bridging Natural Language and Signal Temporal Logic via Tool-Augmented Process-Rewarded Learning
von: Ye, Bowen, et al.
Veröffentlicht: (2026)
von: Ye, Bowen, et al.
Veröffentlicht: (2026)
RAGe: A Retrieval-Augmented Generation Evaluation Framework
von: Guder, Larissa, et al.
Veröffentlicht: (2026)
von: Guder, Larissa, et al.
Veröffentlicht: (2026)
ITDR: An Instruction Tuning Dataset for Enhancing Large Language Models in Recommendations
von: Liu, Zekun, et al.
Veröffentlicht: (2025)
von: Liu, Zekun, et al.
Veröffentlicht: (2025)
Unifying Perplexing Behaviors in Modified BP Attributions through Alignment Perspective
von: Zheng, Guanhua, et al.
Veröffentlicht: (2025)
von: Zheng, Guanhua, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2025) -
Agent models: Internalizing Chain-of-Action Generation into Reasoning models
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2025) -
OpenRFT: Adapting Reasoning Foundation Model for Domain-specific Tasks with Reinforcement Fine-Tuning
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2024) -
o1-Coder: an o1 Replication for Coding
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2024) -
Retrieval-Augmented Process Reward Model for Generalizable Mathematical Reasoning
von: Zhu, Jiachen, et al.
Veröffentlicht: (2025)