CPPO: Contrastive Perception Policy Optimization for VLM Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rezaei, Ahmad, Gholami, Mohsen, Alvar, Saeed Ranjbar, Cannons, Kevin, Hossain, Mohammad Asiful, Weimin, Zhou, Zhang, Yong, Akbari, Mohammad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CASP: Compression of Large Multimodal Models Based on Attention Sparsity
von: Gholami, Mohsen, et al.
Veröffentlicht: (2025)
von: Gholami, Mohsen, et al.
Veröffentlicht: (2025)
LaWa: Using Latent Space for In-Generation Image Watermarking
von: Rezaei, Ahmad, et al.
Veröffentlicht: (2024)
von: Rezaei, Ahmad, et al.
Veröffentlicht: (2024)
Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
von: Gholami, Mohsen, et al.
Veröffentlicht: (2025)
von: Gholami, Mohsen, et al.
Veröffentlicht: (2025)
DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models
von: Alvar, Saeed Ranjbar, et al.
Veröffentlicht: (2025)
von: Alvar, Saeed Ranjbar, et al.
Veröffentlicht: (2025)
Compressive Feature Selection for Remote Visual Multi-Task Inference
von: Alvar, Saeed Ranjbar, et al.
Veröffentlicht: (2024)
von: Alvar, Saeed Ranjbar, et al.
Veröffentlicht: (2024)
Continual Learning: Less Forgetting, More OOD Generalization via Adaptive Contrastive Replay
von: Rezaei, Hossein, et al.
Veröffentlicht: (2024)
von: Rezaei, Hossein, et al.
Veröffentlicht: (2024)
No-reference Quality Assessment of Contrast-distorted Images using Contrast-enhanced Pseudo Reference
von: Mahmoudpour, Mohammad-Ali, et al.
Veröffentlicht: (2025)
von: Mahmoudpour, Mohammad-Ali, et al.
Veröffentlicht: (2025)
AMUSE: Adaptive Multi-Segment Encoding for Dataset Watermarking
von: Alvar, Saeed Ranjbar, et al.
Veröffentlicht: (2024)
von: Alvar, Saeed Ranjbar, et al.
Veröffentlicht: (2024)
Training-Free Diffusion Framework for Stylized Image Generation with Identity Preservation
von: Rezaei, Mohammad Ali, et al.
Veröffentlicht: (2025)
von: Rezaei, Mohammad Ali, et al.
Veröffentlicht: (2025)
SpaceVLM: Sub-Space Modeling of Negation in Vision-Language Models
von: Ranjbar, Sepehr Kazemi, et al.
Veröffentlicht: (2025)
von: Ranjbar, Sepehr Kazemi, et al.
Veröffentlicht: (2025)
Beyond Overall Accuracy: Pose- and Occlusion-driven Fairness Analysis in Pedestrian Detection for Autonomous Driving
von: Khoshkdahan, Mohammad, et al.
Veröffentlicht: (2025)
von: Khoshkdahan, Mohammad, et al.
Veröffentlicht: (2025)
ConDiSR: Contrastive Disentanglement and Style Regularization for Single Domain Generalization
von: Matsun, Aleksandr, et al.
Veröffentlicht: (2024)
von: Matsun, Aleksandr, et al.
Veröffentlicht: (2024)
Towards Secure and Usable 3D Assets: A Novel Framework for Automatic Visible Watermarking
von: Singh, Gursimran, et al.
Veröffentlicht: (2024)
von: Singh, Gursimran, et al.
Veröffentlicht: (2024)
Automated Defect Detection and Grading of Piarom Dates Using Deep Learning
von: Azimi, Nasrin, et al.
Veröffentlicht: (2024)
von: Azimi, Nasrin, et al.
Veröffentlicht: (2024)
EgoVLM: Policy Optimization for Egocentric Video Understanding
von: Vinod, Ashwin, et al.
Veröffentlicht: (2025)
von: Vinod, Ashwin, et al.
Veröffentlicht: (2025)
TrojVLM: Backdoor Attack Against Vision Language Models
von: Lyu, Weimin, et al.
Veröffentlicht: (2024)
von: Lyu, Weimin, et al.
Veröffentlicht: (2024)
Multi-Rationale Explainable Object Recognition via Contrastive Conditional Inference
von: Rasekh, Ali, et al.
Veröffentlicht: (2025)
von: Rasekh, Ali, et al.
Veröffentlicht: (2025)
Towards Pixel-Level VLM Perception via Simple Points Prediction
von: Song, Tianhui, et al.
Veröffentlicht: (2026)
von: Song, Tianhui, et al.
Veröffentlicht: (2026)
APPO: Attention-guided Perception Policy Optimization for Video Reasoning
von: Du, Henghui, et al.
Veröffentlicht: (2026)
von: Du, Henghui, et al.
Veröffentlicht: (2026)
DriveAgent-R1: Advancing VLM-based Autonomous Driving with Active Perception and Hybrid Thinking
von: Zheng, Weicheng, et al.
Veröffentlicht: (2025)
von: Zheng, Weicheng, et al.
Veröffentlicht: (2025)
HiconAgent: History Context-aware Policy Optimization for GUI Agents
von: Zhou, Xurui, et al.
Veröffentlicht: (2025)
von: Zhou, Xurui, et al.
Veröffentlicht: (2025)
DynRsl-VLM: Enhancing Autonomous Driving Perception with Dynamic Resolution Vision-Language Models
von: Zhou, Xirui, et al.
Veröffentlicht: (2025)
von: Zhou, Xirui, et al.
Veröffentlicht: (2025)
GAP3D: Generative Alignment of VLM Latents to Patch-Level Embeddings for 3D Generation
von: Gkotsi, Polytimi Anna, et al.
Veröffentlicht: (2026)
von: Gkotsi, Polytimi Anna, et al.
Veröffentlicht: (2026)
FETAL-GAUGE: A Benchmark for Assessing Vision-Language Models in Fetal Ultrasound
von: Alasmawi, Hussain, et al.
Veröffentlicht: (2025)
von: Alasmawi, Hussain, et al.
Veröffentlicht: (2025)
Text-driven 3D Human Generation via Contrastive Preference Optimization
von: Zhou, Pengfei, et al.
Veröffentlicht: (2025)
von: Zhou, Pengfei, et al.
Veröffentlicht: (2025)
HyCoVAD: A Hybrid SSL-LLM Model for Complex Video Anomaly Detection
von: Hemmatyar, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
von: Hemmatyar, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
MAFM^3: Modular Adaptation of Foundation Models for Multi-Modal Medical AI
von: Qazi, Mohammad Areeb, et al.
Veröffentlicht: (2025)
von: Qazi, Mohammad Areeb, et al.
Veröffentlicht: (2025)
Secure Diagnostics: Adversarial Robustness Meets Clinical Interpretability
von: Najafi, Mohammad Hossein, et al.
Veröffentlicht: (2025)
von: Najafi, Mohammad Hossein, et al.
Veröffentlicht: (2025)
ROOT: VLM based System for Indoor Scene Understanding and Beyond
von: Wang, Yonghui, et al.
Veröffentlicht: (2024)
von: Wang, Yonghui, et al.
Veröffentlicht: (2024)
AllWeatherNet:Unified Image Enhancement for Autonomous Driving under Adverse Weather and Lowlight-conditions
von: Qian, Chenghao, et al.
Veröffentlicht: (2024)
von: Qian, Chenghao, et al.
Veröffentlicht: (2024)
Taming the Randomness: Towards Label-Preserving Cropping in Contrastive Learning
von: Hassan, Mohamed, et al.
Veröffentlicht: (2025)
von: Hassan, Mohamed, et al.
Veröffentlicht: (2025)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
von: Xu, Runsen, et al.
Veröffentlicht: (2024)
von: Xu, Runsen, et al.
Veröffentlicht: (2024)
Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models
von: Rezaei, Parham, et al.
Veröffentlicht: (2025)
von: Rezaei, Parham, et al.
Veröffentlicht: (2025)
Diffusion Beats Autoregressive: An Evaluation of Compositional Generation in Text-to-Image Models
von: Marioriyad, Arash, et al.
Veröffentlicht: (2024)
von: Marioriyad, Arash, et al.
Veröffentlicht: (2024)
Guidance Contrastive Token Credit Assignment for Discrete Policy Optimization
von: Li, Shufan, et al.
Veröffentlicht: (2026)
von: Li, Shufan, et al.
Veröffentlicht: (2026)
Weakly-supervised VLM-guided Partial Contrastive Learning for Visual Language Navigation
von: Wang, Ruoyu, et al.
Veröffentlicht: (2025)
von: Wang, Ruoyu, et al.
Veröffentlicht: (2025)
LeOCLR: Leveraging Original Images for Contrastive Learning of Visual Representations
von: Alkhalefi, Mohammad, et al.
Veröffentlicht: (2024)
von: Alkhalefi, Mohammad, et al.
Veröffentlicht: (2024)
BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning
von: Zhan, Qiusi, et al.
Veröffentlicht: (2025)
von: Zhan, Qiusi, et al.
Veröffentlicht: (2025)
Contrastive learning-based video quality assessment-jointed video vision transformer for video recognition
von: Sun, Jian, et al.
Veröffentlicht: (2026)
von: Sun, Jian, et al.
Veröffentlicht: (2026)
A Contrastive Learning Framework for Breast Cancer Detection
von: Saeed, Samia, et al.
Veröffentlicht: (2025)
von: Saeed, Samia, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CASP: Compression of Large Multimodal Models Based on Attention Sparsity
von: Gholami, Mohsen, et al.
Veröffentlicht: (2025) -
LaWa: Using Latent Space for In-Generation Image Watermarking
von: Rezaei, Ahmad, et al.
Veröffentlicht: (2024) -
Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
von: Gholami, Mohsen, et al.
Veröffentlicht: (2025) -
DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models
von: Alvar, Saeed Ranjbar, et al.
Veröffentlicht: (2025) -
Compressive Feature Selection for Remote Visual Multi-Task Inference
von: Alvar, Saeed Ranjbar, et al.
Veröffentlicht: (2024)