COVR:Collaborative Optimization of VLMs and RL Agent for Visual-Based Control
Fuente:
arXiv
Saved in:
| Main Authors: | Xia, Canming, Peng, Peixi, Tan, Guang, Su, Zhan, Xu, Haoran, Liu, Zhenxian, Li, Luntong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deflickering Vision-Based Occupancy Networks through Lightweight Spatio-Temporal Correlation
by: Yu, Fengcheng, et al.
Published: (2025)
by: Yu, Fengcheng, et al.
Published: (2025)
Learning to Search: A Decision-Based Agent for Knowledge-Based Visual Question Answering
by: Chen, Zhuohong, et al.
Published: (2026)
by: Chen, Zhuohong, et al.
Published: (2026)
ProphetDWM: A Driving World Model for Rolling Out Future Actions and Videos
by: Wang, Xiaodong, et al.
Published: (2025)
by: Wang, Xiaodong, et al.
Published: (2025)
FreeGen: Feed-Forward Reconstruction-Generation Co-Training for Free-Viewpoint Driving Scene Synthesis
by: Chen, Shijie, et al.
Published: (2025)
by: Chen, Shijie, et al.
Published: (2025)
MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning
by: Xia, Peng, et al.
Published: (2025)
by: Xia, Peng, et al.
Published: (2025)
LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs
by: Wang, Xiaodong, et al.
Published: (2025)
by: Wang, Xiaodong, et al.
Published: (2025)
When Does RL Help Medical VLMs? Disentangling Vision, SFT, and RL Gains
by: Jeddi, Ahmadreza, et al.
Published: (2026)
by: Jeddi, Ahmadreza, et al.
Published: (2026)
LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model
by: Wang, Xiaodong, et al.
Published: (2025)
by: Wang, Xiaodong, et al.
Published: (2025)
Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs
by: Elmansoury, Sary, et al.
Published: (2025)
by: Elmansoury, Sary, et al.
Published: (2025)
LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding
by: Wang, Xiaodong, et al.
Published: (2026)
by: Wang, Xiaodong, et al.
Published: (2026)
Navigating Heat Exposure: Simulation of Route Planning Based on Visual Language Model Agents
by: Ma, Haoran, et al.
Published: (2025)
by: Ma, Haoran, et al.
Published: (2025)
DASK: Distribution Rehearsing via Adaptive Style Kernel Learning for Exemplar-Free Lifelong Person Re-Identification
by: Xu, Kunlun, et al.
Published: (2024)
by: Xu, Kunlun, et al.
Published: (2024)
Can VLMs Truly Forget? Benchmarking Training-Free Visual Concept Unlearning
by: Tan, Zhangyun, et al.
Published: (2026)
by: Tan, Zhangyun, et al.
Published: (2026)
VACoT: Rethinking Visual Data Augmentation with VLMs
by: Xu, Zhengzhuo, et al.
Published: (2025)
by: Xu, Zhengzhuo, et al.
Published: (2025)
Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning
by: Yang, Jingru, et al.
Published: (2024)
by: Yang, Jingru, et al.
Published: (2024)
SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration
by: Yang, Zhongyu, et al.
Published: (2026)
by: Yang, Zhongyu, et al.
Published: (2026)
IE-NeRF: Inpainting Enhanced Neural Radiance Fields in the Wild
by: Wang, Shuaixian, et al.
Published: (2024)
by: Wang, Shuaixian, et al.
Published: (2024)
LVC: A Lightweight Compression Framework for Enhancing VLMs in Long Video Understanding
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
by: Berman, Shmuel, et al.
Published: (2025)
by: Berman, Shmuel, et al.
Published: (2025)
Multimodal Backdoor Attack on VLMs for Autonomous Driving via Graffiti and Cross-Lingual Triggers
by: Wang, Jiancheng, et al.
Published: (2026)
by: Wang, Jiancheng, et al.
Published: (2026)
Probing Visual Language Priors in VLMs
by: Luo, Tiange, et al.
Published: (2024)
by: Luo, Tiange, et al.
Published: (2024)
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
by: Törtei, Brigitta Malagurski, et al.
Published: (2025)
by: Törtei, Brigitta Malagurski, et al.
Published: (2025)
AgentVista: Evaluating Multimodal Agents in Ultra-Challenging Realistic Visual Scenarios
by: Su, Zhaochen, et al.
Published: (2026)
by: Su, Zhaochen, et al.
Published: (2026)
Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
by: Liu, Zhining, et al.
Published: (2025)
by: Liu, Zhining, et al.
Published: (2025)
Towards Implicit Aggregation: Robust Image Representation for Place Recognition in the Transformer Era
by: Lu, Feng, et al.
Published: (2025)
by: Lu, Feng, et al.
Published: (2025)
Vision Verification Enhanced Fusion of VLMs for Efficient Visual Reasoning
by: Tekin, Selim Furkan, et al.
Published: (2026)
by: Tekin, Selim Furkan, et al.
Published: (2026)
Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth
by: Wu, Yuhuan, et al.
Published: (2026)
by: Wu, Yuhuan, et al.
Published: (2026)
Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual Illusions
by: Sun, Xiaoxiao, et al.
Published: (2026)
by: Sun, Xiaoxiao, et al.
Published: (2026)
Efficient Event-Based Semantic Segmentation via Exploiting Frame-Event Fusion: A Hybrid Neural Network Approach
by: Li, Hebei, et al.
Published: (2025)
by: Li, Hebei, et al.
Published: (2025)
LightZeroNav: Zero-Shot Vision Language Navigation in Continuous Environments Based on Lightweight VLMs
by: Luo, Kun, et al.
Published: (2026)
by: Luo, Kun, et al.
Published: (2026)
FlowErase-RL: Rethinking Concept Erasure as Reward Optimization in Flow Matching Models
by: Sun, Yi, et al.
Published: (2026)
by: Sun, Yi, et al.
Published: (2026)
One-shot Optimized Steering Vector for Hallucination Mitigation for VLMs
by: Shi, Youxu, et al.
Published: (2026)
by: Shi, Youxu, et al.
Published: (2026)
T2T-VICL: Unlocking the Boundaries of Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs
by: Xia, Shao-Jun, et al.
Published: (2025)
by: Xia, Shao-Jun, et al.
Published: (2025)
HIVTP: A Training-Free Method to Improve VLMs Efficiency via Hierarchical Visual Token Pruning Using Middle-Layer-Based Importance Score
by: Xu, Jingqi, et al.
Published: (2025)
by: Xu, Jingqi, et al.
Published: (2025)
vMFCoOp: Towards Equilibrium on a Unified Hyperspherical Manifold for Prompting Biomedical VLMs
by: Shao, Minye, et al.
Published: (2025)
by: Shao, Minye, et al.
Published: (2025)
Adaptive Time-step Training for Enhancing Spike-Based Neural Radiance Fields
by: Lin, Ranxi, et al.
Published: (2025)
by: Lin, Ranxi, et al.
Published: (2025)
VisualActBench: Can VLMs See and Act like a Human?
by: Zhang, Daoan, et al.
Published: (2025)
by: Zhang, Daoan, et al.
Published: (2025)
Beyond Shortcuts: Mitigating Visual Illusions in Frozen VLMs via Qualitative Reasoning
by: Guo, Hao, et al.
Published: (2026)
by: Guo, Hao, et al.
Published: (2026)
Can VLMs be used on videos for action recognition? LLMs are Visual Reasoning Coordinators
by: Lunia, Harsh
Published: (2024)
by: Lunia, Harsh
Published: (2024)
Generalizing Consistency Policy to Visual RL with Prioritized Proximal Experience Regularization
by: Li, Haoran, et al.
Published: (2024)
by: Li, Haoran, et al.
Published: (2024)
Similar Items
-
Deflickering Vision-Based Occupancy Networks through Lightweight Spatio-Temporal Correlation
by: Yu, Fengcheng, et al.
Published: (2025) -
Learning to Search: A Decision-Based Agent for Knowledge-Based Visual Question Answering
by: Chen, Zhuohong, et al.
Published: (2026) -
ProphetDWM: A Driving World Model for Rolling Out Future Actions and Videos
by: Wang, Xiaodong, et al.
Published: (2025) -
FreeGen: Feed-Forward Reconstruction-Generation Co-Training for Free-Viewpoint Driving Scene Synthesis
by: Chen, Shijie, et al.
Published: (2025) -
MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning
by: Xia, Peng, et al.
Published: (2025)