Praxis-VLM: Vision-Grounded Decision Making via Text-Driven Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Zhe, Li, Jing, Pu, Zhongzhu, Chan, Hou Pong, Yin, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VIVA: A Benchmark for Vision-Grounded Decision-Making with Human Values
by: Hu, Zhe, et al.
Published: (2024)
by: Hu, Zhe, et al.
Published: (2024)
A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation
by: Zhao, Yi, et al.
Published: (2026)
by: Zhao, Yi, et al.
Published: (2026)
Debate-to-Write: A Persona-Driven Multi-Agent Framework for Diverse Argument Generation
by: Hu, Zhe, et al.
Published: (2024)
by: Hu, Zhe, et al.
Published: (2024)
Scaling Language-Centric Omnimodal Representation Learning
by: Xiao, Chenghao, et al.
Published: (2025)
by: Xiao, Chenghao, et al.
Published: (2025)
Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
by: Zhai, Yuexiang, et al.
Published: (2024)
by: Zhai, Yuexiang, et al.
Published: (2024)
AMERICANO: Argument Generation with Discourse-driven Decomposition and Agent Interaction
by: Hu, Zhe, et al.
Published: (2023)
by: Hu, Zhe, et al.
Published: (2023)
SleepVLM: Explainable and Rule-Grounded Sleep Staging via a Vision-Language Model
by: Deng, Guifeng, et al.
Published: (2026)
by: Deng, Guifeng, et al.
Published: (2026)
HPE-CogVLM: Advancing Vision Language Models with a Head Pose Grounding Task
by: Tian, Yu, et al.
Published: (2024)
by: Tian, Yu, et al.
Published: (2024)
OViP: Online Vision-Language Preference Learning for VLM Hallucination
by: Liu, Shujun, et al.
Published: (2025)
by: Liu, Shujun, et al.
Published: (2025)
VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
by: Yuan, Ruifeng, et al.
Published: (2025)
by: Yuan, Ruifeng, et al.
Published: (2025)
G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
by: Hu, Wenbo, et al.
Published: (2025)
by: Hu, Wenbo, et al.
Published: (2025)
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
by: Fallah, Forouzan, et al.
Published: (2025)
by: Fallah, Forouzan, et al.
Published: (2025)
Multimedia Generative Script Learning for Task Planning
by: Wang, Qingyun, et al.
Published: (2022)
by: Wang, Qingyun, et al.
Published: (2022)
Systematic Reward Gap Optimization for Mitigating VLM Hallucinations
by: He, Lehan, et al.
Published: (2024)
by: He, Lehan, et al.
Published: (2024)
GeoDANO: Geometric VLM with Domain Agnostic Vision Encoder
by: Cho, Seunghyuk, et al.
Published: (2025)
by: Cho, Seunghyuk, et al.
Published: (2025)
MIRL: Mutual Information-Guided Reinforcement Learning for Vision-Language Models
by: Zhang, Yin, et al.
Published: (2026)
by: Zhang, Yin, et al.
Published: (2026)
Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models
by: Shao, Zhenwei, et al.
Published: (2025)
by: Shao, Zhenwei, et al.
Published: (2025)
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
by: Fan, Zhiwen, et al.
Published: (2025)
by: Fan, Zhiwen, et al.
Published: (2025)
GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning
by: Ebouky, Brown, et al.
Published: (2026)
by: Ebouky, Brown, et al.
Published: (2026)
Allegory of the Cave: Measurement-Grounded Vision-Language Learning
by: Xu, Kepeng, et al.
Published: (2026)
by: Xu, Kepeng, et al.
Published: (2026)
Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions
by: Hu, Zhe, et al.
Published: (2024)
by: Hu, Zhe, et al.
Published: (2024)
Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars
by: Zhang, Youliang, et al.
Published: (2026)
by: Zhang, Youliang, et al.
Published: (2026)
From Behavioral Performance to Internal Competence: Interpreting Vision-Language Models with VLM-Lens
by: Sheta, Hala, et al.
Published: (2025)
by: Sheta, Hala, et al.
Published: (2025)
IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding
by: Li, Junxian, et al.
Published: (2025)
by: Li, Junxian, et al.
Published: (2025)
Test-Time Reinforcement Learning for GUI Grounding via Region Consistency
by: Du, Yong, et al.
Published: (2025)
by: Du, Yong, et al.
Published: (2025)
ColorConceptBench: A Benchmark for Probabilistic Color-Concept Understanding in Text-to-Image Models
by: Ruan, Chenxi, et al.
Published: (2026)
by: Ruan, Chenxi, et al.
Published: (2026)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
by: Wu, Yin, et al.
Published: (2025)
by: Wu, Yin, et al.
Published: (2025)
SCoPE VLM: Selective Context Processing for Efficient Document Navigation in Vision-Language Models
by: Lim, Gyubeum, et al.
Published: (2025)
by: Lim, Gyubeum, et al.
Published: (2025)
PatientVLM Meets DocVLM: Pre-Consultation Dialogue Between Vision-Language Models for Efficient Diagnosis
by: Lokesh, K, et al.
Published: (2026)
by: Lokesh, K, et al.
Published: (2026)
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
by: Zhang, Di, et al.
Published: (2024)
by: Zhang, Di, et al.
Published: (2024)
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
by: Shen, Haozhan, et al.
Published: (2025)
by: Shen, Haozhan, et al.
Published: (2025)
SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
by: Liu, Zheng, et al.
Published: (2024)
by: Liu, Zheng, et al.
Published: (2024)
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
by: Luo, Fuwen, et al.
Published: (2025)
by: Luo, Fuwen, et al.
Published: (2025)
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
by: Yang, Senqiao, et al.
Published: (2025)
by: Yang, Senqiao, et al.
Published: (2025)
Teaching Vision-Language Models to Ask: Resolving Ambiguity in Visual Questions
by: Jian, Pu, et al.
Published: (2025)
by: Jian, Pu, et al.
Published: (2025)
Vision-Language Modeling in PET/CT for Visual Grounding of Positive Findings
by: Huemann, Zachary, et al.
Published: (2025)
by: Huemann, Zachary, et al.
Published: (2025)
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
by: Chaffin, Antoine, et al.
Published: (2024)
by: Chaffin, Antoine, et al.
Published: (2024)
CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception
by: Carvalho, Miguel, et al.
Published: (2025)
by: Carvalho, Miguel, et al.
Published: (2025)
Q-GroundCAM: Quantifying Grounding in Vision Language Models via GradCAM
by: Rajabi, Navid, et al.
Published: (2024)
by: Rajabi, Navid, et al.
Published: (2024)
IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
by: Lu, Xiaoya, et al.
Published: (2025)
by: Lu, Xiaoya, et al.
Published: (2025)
Similar Items
-
VIVA: A Benchmark for Vision-Grounded Decision-Making with Human Values
by: Hu, Zhe, et al.
Published: (2024) -
A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation
by: Zhao, Yi, et al.
Published: (2026) -
Debate-to-Write: A Persona-Driven Multi-Agent Framework for Diverse Argument Generation
by: Hu, Zhe, et al.
Published: (2024) -
Scaling Language-Centric Omnimodal Representation Learning
by: Xiao, Chenghao, et al.
Published: (2025) -
Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
by: Zhai, Yuexiang, et al.
Published: (2024)