Guiding VLM Agents with Process Rewards at Inference Time for GUI Navigation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Zhiyuan, Xiong, Shiyun, Zhang, Yifan, Ng, See-Kiong, Luu, Anh Tuan, An, Bo, Yan, Shuicheng, Hooi, Bryan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026)
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
Collaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026)
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
von: Cao, Tri, et al.
Veröffentlicht: (2026)
von: Cao, Tri, et al.
Veröffentlicht: (2026)
Uncertainty of Thoughts: Uncertainty-Aware Planning Enhances Information Seeking in Large Language Models
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2024)
Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding
von: Nguyen, Thong, et al.
Veröffentlicht: (2025)
von: Nguyen, Thong, et al.
Veröffentlicht: (2025)
LongRecipe: Recipe for Efficient Long Context Generalization in Large Language Models
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2024)
Spatio-Temporal Foundation Models: Vision, Challenges, and Opportunities
von: Goodge, Adam, et al.
Veröffentlicht: (2025)
von: Goodge, Adam, et al.
Veröffentlicht: (2025)
Encoding and Controlling Global Semantics for Long-form Video Question Answering
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
Mercury: A Code Efficiency Benchmark for Code Large Language Models
von: Du, Mingzhe, et al.
Veröffentlicht: (2024)
von: Du, Mingzhe, et al.
Veröffentlicht: (2024)
SemRoDe: Macro Adversarial Training to Learn Representations That are Robust to Word-Level Attacks
von: Formento, Brian, et al.
Veröffentlicht: (2024)
von: Formento, Brian, et al.
Veröffentlicht: (2024)
Vision-and-Language Pretraining
von: Nguyen, Thong, et al.
Veröffentlicht: (2022)
von: Nguyen, Thong, et al.
Veröffentlicht: (2022)
Multi-Scale Contrastive Learning for Video Temporal Grounding
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift
von: Le, Khoi, et al.
Veröffentlicht: (2026)
von: Le, Khoi, et al.
Veröffentlicht: (2026)
READ: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language Modeling
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
GUI-PRA: Process Reward Agent for GUI Tasks
von: Xiong, Tao, et al.
Veröffentlicht: (2025)
von: Xiong, Tao, et al.
Veröffentlicht: (2025)
Inference-Time Attribute Distribution Alignment for Unconditional Diffusion
von: Luan, Hao, et al.
Veröffentlicht: (2026)
von: Luan, Hao, et al.
Veröffentlicht: (2026)
How Does Response Length Affect Long-Form Factuality
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
Topic Modeling as Multi-Objective Contrastive Optimization
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
Geneshift: Impact of different scenario shift on Jailbreaking LLM
von: Wu, Tianyi, et al.
Veröffentlicht: (2025)
von: Wu, Tianyi, et al.
Veröffentlicht: (2025)
Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
MAMA: Meta-optimized Angular Margin Contrastive Framework for Video-Language Representation Learning
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
Paper Espresso: From Paper Overload to Research Insight
von: Du, Mingzhe, et al.
Veröffentlicht: (2026)
von: Du, Mingzhe, et al.
Veröffentlicht: (2026)
CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs
von: Fu, Jinlan, et al.
Veröffentlicht: (2025)
von: Fu, Jinlan, et al.
Veröffentlicht: (2025)
CodeArena: A Collective Evaluation Platform for LLM Code Generation
von: Du, Mingzhe, et al.
Veröffentlicht: (2025)
von: Du, Mingzhe, et al.
Veröffentlicht: (2025)
Hint-before-Solving Prompting: Guiding LLMs to Effectively Utilize Encoded Knowledge
von: Fu, Jinlan, et al.
Veröffentlicht: (2024)
von: Fu, Jinlan, et al.
Veröffentlicht: (2024)
Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning
von: Wu, Tianyi, et al.
Veröffentlicht: (2025)
von: Wu, Tianyi, et al.
Veröffentlicht: (2025)
Mem-W: Latent Memory-Native GUI Agents
von: Zhang, Guibin, et al.
Veröffentlicht: (2026)
von: Zhang, Guibin, et al.
Veröffentlicht: (2026)
Image Can Bring Your Memory Back: A Novel Multi-Modal Guided Attack against Image Generation Model Unlearning
von: Liu, Renyang, et al.
Veröffentlicht: (2025)
von: Liu, Renyang, et al.
Veröffentlicht: (2025)
Unsupervised Hallucination Detection by Inspecting Reasoning Processes
von: Srey, Ponhvoan, et al.
Veröffentlicht: (2025)
von: Srey, Ponhvoan, et al.
Veröffentlicht: (2025)
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
TRACE: TRansformer-based Attribution using Contrastive Embeddings in LLMs
von: Wang, Cheng, et al.
Veröffentlicht: (2024)
von: Wang, Cheng, et al.
Veröffentlicht: (2024)
Multi-Agent Sampling: Scaling Inference Compute for Data Synthesis with Tree Search-Based Agentic Collaboration
von: Ye, Hai, et al.
Veröffentlicht: (2024)
von: Ye, Hai, et al.
Veröffentlicht: (2024)
MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration
von: Xu, Lin, et al.
Veröffentlicht: (2023)
von: Xu, Lin, et al.
Veröffentlicht: (2023)
MASim: Multilingual Agent-Based Simulation for Social Science
von: Zhang, Xuan, et al.
Veröffentlicht: (2025)
von: Zhang, Xuan, et al.
Veröffentlicht: (2025)
AgentStudio: A Toolkit for Building General Virtual Agents
von: Zheng, Longtao, et al.
Veröffentlicht: (2024)
von: Zheng, Longtao, et al.
Veröffentlicht: (2024)
Ferret: Federated Full-Parameter Tuning at Scale for Large Language Models
von: Shu, Yao, et al.
Veröffentlicht: (2024)
von: Shu, Yao, et al.
Veröffentlicht: (2024)
Projected Coupled Diffusion for Test-Time Constrained Joint Generation
von: Luan, Hao, et al.
Veröffentlicht: (2025)
von: Luan, Hao, et al.
Veröffentlicht: (2025)
Confidence Elicitation: A New Attack Vector for Large Language Models
von: Formento, Brian, et al.
Veröffentlicht: (2025)
von: Formento, Brian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026) -
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
von: Zhao, James Xu, et al.
Veröffentlicht: (2025) -
Collaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026) -
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
von: Cao, Tri, et al.
Veröffentlicht: (2026) -
Uncertainty of Thoughts: Uncertainty-Aware Planning Enhances Information Seeking in Large Language Models
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2024)