Web-Shepherd: Advancing PRMs for Reinforcing Web Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Chae, Hyungjoo, Kim, Sunghwan, Cho, Junhee, Kim, Seungone, Moon, Seungjun, Hwangbo, Gyeom, Lim, Dongha, Kim, Minjin, Hwang, Yeonjun, Gwak, Minju, Choi, Dongwook, Kang, Minseok, Im, Gwanhoon, Cho, ByeongUng, Kim, Hyojun, Han, Jun Hee, Kwon, Taeyoon, Kim, Minju, Kwak, Beong-woo, Kang, Dongjin, Yeo, Jinyoung |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ToolHaystack: Stress-Testing Tool-Augmented Language Models in Realistic Long-Term Interactions
by: Kwak, Beong-woo, et al.
Published: (2025)
by: Kwak, Beong-woo, et al.
Published: (2025)
Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation
by: Chae, Hyungjoo, et al.
Published: (2024)
by: Chae, Hyungjoo, et al.
Published: (2024)
Embodied Agents Meet Personalization: Investigating Challenges and Solutions Through the Lens of Memory Utilization
by: Kwon, Taeyoon, et al.
Published: (2025)
by: Kwon, Taeyoon, et al.
Published: (2025)
Can You Share Your Story? Modeling Clients' Metacognition and Openness for LLM Therapist Evaluation
by: Kim, Minju, et al.
Published: (2025)
by: Kim, Minju, et al.
Published: (2025)
Pearl: A Review-driven Persona-Knowledge Grounded Conversational Recommendation Dataset
by: Kim, Minjin, et al.
Published: (2024)
by: Kim, Minjin, et al.
Published: (2024)
EMBGuard: Constructing Hazard-Aware Guardrails for Safe Planning in Embodied Agents
by: Choi, Dongwook, et al.
Published: (2026)
by: Choi, Dongwook, et al.
Published: (2026)
Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
by: Kim, Sunghwan, et al.
Published: (2025)
by: Kim, Sunghwan, et al.
Published: (2025)
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length
by: Kim, Sunghwan, et al.
Published: (2026)
by: Kim, Sunghwan, et al.
Published: (2026)
CONDESION-BENCH: Conditional Decision-Making of Large Language Models in Compositional Action Space
by: Hwang, Yeonjun, et al.
Published: (2026)
by: Hwang, Yeonjun, et al.
Published: (2026)
Language Models as Compilers: Simulating Pseudocode Execution Improves Algorithmic Reasoning in Language Models
by: Chae, Hyungjoo, et al.
Published: (2024)
by: Chae, Hyungjoo, et al.
Published: (2024)
Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support Conversation
by: Kang, Dongjin, et al.
Published: (2024)
by: Kang, Dongjin, et al.
Published: (2024)
Evaluating Robustness of Reward Models for Mathematical Reasoning
by: Kim, Sunghwan, et al.
Published: (2024)
by: Kim, Sunghwan, et al.
Published: (2024)
Revisiting the Uniform Information Density Hypothesis in LLM Reasoning
by: Gwak, Minju, et al.
Published: (2025)
by: Gwak, Minju, et al.
Published: (2025)
Revisiting the UID Hypothesis in LLM Reasoning Traces
by: Gwak, Minju, et al.
Published: (2025)
by: Gwak, Minju, et al.
Published: (2025)
Towards Lifelong Dialogue Agents via Timeline-based Memory Management
by: Ong, Kai Tzu-iunn, et al.
Published: (2024)
by: Ong, Kai Tzu-iunn, et al.
Published: (2024)
LLM Meets Scene Graph: Can Large Language Models Understand and Generate Scene Graphs? A Benchmark and Empirical Study
by: Yang, Dongil, et al.
Published: (2025)
by: Yang, Dongil, et al.
Published: (2025)
Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics
by: Lee, Seungbeen, et al.
Published: (2024)
by: Lee, Seungbeen, et al.
Published: (2024)
PAC-BENCH: Evaluating Multi-Agent Collaboration under Privacy Constraints
by: Park, Minjun, et al.
Published: (2026)
by: Park, Minjun, et al.
Published: (2026)
Designing Memory-Augmented AR Agents for Spatiotemporal Reasoning in Personalized Task Assistance
by: Choi, Dongwook, et al.
Published: (2025)
by: Choi, Dongwook, et al.
Published: (2025)
Semi-Supervised 3D Object Detection with Channel Augmentation using Transformation Equivariance
by: Kang, Minju, et al.
Published: (2024)
by: Kang, Minju, et al.
Published: (2024)
Convergence of orbital integrals on unitary groups in positive characteristic
by: Kim, Wansu, et al.
Published: (2026)
by: Kim, Wansu, et al.
Published: (2026)
PaP-NF: Probabilistic Long-Term Time Series Forecasting via Prefix-as-Prompt Reprogramming and Normalizing Flows
by: Kim, Minju, et al.
Published: (2026)
by: Kim, Minju, et al.
Published: (2026)
Sounds of Hidden Agents: The Development of Causal Reasoning About Musical Sounds
by: Minju Kim, et al.
Published: (2025)
by: Minju Kim, et al.
Published: (2025)
PRINCIPLES: Synthetic Strategy Memory for Proactive Dialogue Agents
by: Kim, Namyoung, et al.
Published: (2025)
by: Kim, Namyoung, et al.
Published: (2025)
Towards Direct Evaluation of Harness Optimizers via Priority Ranking
by: Ong, Kai Tzu-iunn, et al.
Published: (2026)
by: Ong, Kai Tzu-iunn, et al.
Published: (2026)
AgenticShop: Benchmarking Agentic Product Curation for Personalized Web Shopping
by: Kim, Sunghwan, et al.
Published: (2026)
by: Kim, Sunghwan, et al.
Published: (2026)
Comparison of Corneal Endothelial Imaging Techniques by Specular Microscopy in Unsedated Healthy Dogs
by: Hyunwoo Suk, et al.
Published: (2026)
by: Hyunwoo Suk, et al.
Published: (2026)
Fine-Grained and Thematic Evaluation of LLMs in Social Deduction Game
by: Kim, Byungjun, et al.
Published: (2024)
by: Kim, Byungjun, et al.
Published: (2024)
Leveraging Large Language Models for Active Merchant Non-player Characters
by: Kim, Byungjun, et al.
Published: (2024)
by: Kim, Byungjun, et al.
Published: (2024)
PHISH in MESH: Korean Adversarial Phonetic Substitution and Phonetic-Semantic Feature Integration Defense
by: Kim, Byungjun, et al.
Published: (2025)
by: Kim, Byungjun, et al.
Published: (2025)
Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History
by: Kim, Serin, et al.
Published: (2026)
by: Kim, Serin, et al.
Published: (2026)
Horocycles in hyperbolic 3-manifolds with round Sierpiński limit sets
by: Kim, Dongryul M., et al.
Published: (2025)
by: Kim, Dongryul M., et al.
Published: (2025)
Multiple Photon Subtraction on Light with Tunable Intensity Correlations
by: Minju Kim, et al.
Published: (2026)
by: Minju Kim, et al.
Published: (2026)
Lower Trapezius Transfer Using the Retrograde Keyhole Technique With an Achilles Tendon–Bone Allograft
by: Hyung‐gyu Cho, et al.
Published: (2025)
by: Hyung‐gyu Cho, et al.
Published: (2025)
Cactus: Towards Psychological Counseling Conversations using Cognitive Behavioral Theory
by: Lee, Suyeon, et al.
Published: (2024)
by: Lee, Suyeon, et al.
Published: (2024)
LEGO-Eval: Towards Fine-Grained Evaluation on Synthesizing 3D Embodied Environments with Tool Augmentation
by: Hwangbo, Gyeom, et al.
Published: (2025)
by: Hwangbo, Gyeom, et al.
Published: (2025)
One Missing Piece for Open-Source Reasoning Models: A Dataset to Mitigate Cold-Starting Short CoT LLMs in RL
by: Chae, Hyungjoo, et al.
Published: (2025)
by: Chae, Hyungjoo, et al.
Published: (2025)
Therapeutic Efficacy of Endoscope‐Guided Coblation for Tongue Base Reduction as Part of Multilevel Surgery in Moderate to Severe OSA Patients
by: Minju Kim, et al.
Published: (2026)
by: Minju Kim, et al.
Published: (2026)
Coffee-Gym: An Environment for Evaluating and Improving Natural Language Feedback on Erroneous Code
by: Chae, Hyungjoo, et al.
Published: (2024)
by: Chae, Hyungjoo, et al.
Published: (2024)
Can LLMs and humans be friends? Uncovering factors affecting human-AI intimacy formation
by: Hong, Yeseon, et al.
Published: (2025)
by: Hong, Yeseon, et al.
Published: (2025)
Similar Items
-
ToolHaystack: Stress-Testing Tool-Augmented Language Models in Realistic Long-Term Interactions
by: Kwak, Beong-woo, et al.
Published: (2025) -
Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation
by: Chae, Hyungjoo, et al.
Published: (2024) -
Embodied Agents Meet Personalization: Investigating Challenges and Solutions Through the Lens of Memory Utilization
by: Kwon, Taeyoon, et al.
Published: (2025) -
Can You Share Your Story? Modeling Clients' Metacognition and Openness for LLM Therapist Evaluation
by: Kim, Minju, et al.
Published: (2025) -
Pearl: A Review-driven Persona-Knowledge Grounded Conversational Recommendation Dataset
by: Kim, Minjin, et al.
Published: (2024)