Vision-Language Models Provide Promptable Representations for Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, William, Mees, Oier, Kumar, Aviral, Levine, Sergey |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Ingredients for Robotic Diffusion Transformers
von: Dasari, Sudeep, et al.
Veröffentlicht: (2024)
von: Dasari, Sudeep, et al.
Veröffentlicht: (2024)
Training Diffusion Models with Reinforcement Learning
von: Black, Kevin, et al.
Veröffentlicht: (2023)
von: Black, Kevin, et al.
Veröffentlicht: (2023)
mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
von: Pai, Jonas, et al.
Veröffentlicht: (2025)
von: Pai, Jonas, et al.
Veröffentlicht: (2025)
Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models
von: Blank, Nils, et al.
Veröffentlicht: (2024)
von: Blank, Nils, et al.
Veröffentlicht: (2024)
Multimodal Spatial Language Maps for Robot Navigation and Manipulation
von: Huang, Chenguang, et al.
Veröffentlicht: (2025)
von: Huang, Chenguang, et al.
Veröffentlicht: (2025)
GHIL-Glue: Hierarchical Control with Filtered Subgoal Images
von: Hatch, Kyle B., et al.
Veröffentlicht: (2024)
von: Hatch, Kyle B., et al.
Veröffentlicht: (2024)
Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
von: Zhai, Yuexiang, et al.
Veröffentlicht: (2024)
von: Zhai, Yuexiang, et al.
Veröffentlicht: (2024)
Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations
von: Srivastava, Archita, et al.
Veröffentlicht: (2025)
von: Srivastava, Archita, et al.
Veröffentlicht: (2025)
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
von: Yang, Senqiao, et al.
Veröffentlicht: (2025)
von: Yang, Senqiao, et al.
Veröffentlicht: (2025)
Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2025)
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2025)
CURLing the Dream: Contrastive Representations for World Modeling in Reinforcement Learning
von: Kich, Victor Augusto, et al.
Veröffentlicht: (2024)
von: Kich, Victor Augusto, et al.
Veröffentlicht: (2024)
Driver Activity Classification Using Generalizable Representations from Vision-Language Models
von: Greer, Ross, et al.
Veröffentlicht: (2024)
von: Greer, Ross, et al.
Veröffentlicht: (2024)
Decipher-MR: A Vision-Language Foundation Model for 3D MRI Representations
von: Yang, Zhijian, et al.
Veröffentlicht: (2025)
von: Yang, Zhijian, et al.
Veröffentlicht: (2025)
The Surprising Ineffectiveness of Pre-Trained Visual Representations for Model-Based Reinforcement Learning
von: Schneider, Moritz, et al.
Veröffentlicht: (2024)
von: Schneider, Moritz, et al.
Veröffentlicht: (2024)
FastVLM: Efficient Vision Encoding for Vision Language Models
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2024)
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2024)
BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak Attacks
von: Zhao, Yunhan, et al.
Veröffentlicht: (2024)
von: Zhao, Yunhan, et al.
Veröffentlicht: (2024)
Opportunistic Promptable Segmentation: Leveraging Routine Radiological Annotations to Guide 3D CT Lesion Segmentation
von: Church, Samuel, et al.
Veröffentlicht: (2026)
von: Church, Samuel, et al.
Veröffentlicht: (2026)
Towards Principled Representation Learning from Videos for Reinforcement Learning
von: Misra, Dipendra, et al.
Veröffentlicht: (2024)
von: Misra, Dipendra, et al.
Veröffentlicht: (2024)
Compositional Entailment Learning for Hyperbolic Vision-Language Models
von: Pal, Avik, et al.
Veröffentlicht: (2024)
von: Pal, Avik, et al.
Veröffentlicht: (2024)
Tree of Attributes Prompt Learning for Vision-Language Models
von: Ding, Tong, et al.
Veröffentlicht: (2024)
von: Ding, Tong, et al.
Veröffentlicht: (2024)
Parallel In-context Learning for Large Vision Language Models
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2026)
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2026)
Differentiable Prompt Learning for Vision Language Models
von: Huang, Zhenhan, et al.
Veröffentlicht: (2024)
von: Huang, Zhenhan, et al.
Veröffentlicht: (2024)
Efficient Few-Shot Learning in Remote Sensing: Fusing Vision and Vision-Language Models
von: Chua, Jia Yun, et al.
Veröffentlicht: (2025)
von: Chua, Jia Yun, et al.
Veröffentlicht: (2025)
Continual Learning in Vision-Language Models via Aligned Model Merging
von: Sokar, Ghada, et al.
Veröffentlicht: (2025)
von: Sokar, Ghada, et al.
Veröffentlicht: (2025)
Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
von: Lin, Zinan, et al.
Veröffentlicht: (2025)
von: Lin, Zinan, et al.
Veröffentlicht: (2025)
AAPL: Adding Attributes to Prompt Learning for Vision-Language Models
von: Kim, Gahyeon, et al.
Veröffentlicht: (2024)
von: Kim, Gahyeon, et al.
Veröffentlicht: (2024)
Lightweight Unsupervised Federated Learning with Pretrained Vision Language Model
von: Yan, Hao, et al.
Veröffentlicht: (2024)
von: Yan, Hao, et al.
Veröffentlicht: (2024)
Decoupling Augmentation Bias in Prompt Learning for Vision-Language Models
von: Kim, Gahyeon, et al.
Veröffentlicht: (2025)
von: Kim, Gahyeon, et al.
Veröffentlicht: (2025)
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
von: Pach, Mateusz, et al.
Veröffentlicht: (2025)
von: Pach, Mateusz, et al.
Veröffentlicht: (2025)
Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
von: Grover, Shresth, et al.
Veröffentlicht: (2025)
von: Grover, Shresth, et al.
Veröffentlicht: (2025)
AutoCLIP: Auto-tuning Zero-Shot Classifiers for Vision-Language Models
von: Metzen, Jan Hendrik, et al.
Veröffentlicht: (2023)
von: Metzen, Jan Hendrik, et al.
Veröffentlicht: (2023)
Malicious Path Manipulations via Exploitation of Representation Vulnerabilities of Vision-Language Navigation Systems
von: Islam, Chashi Mahiul, et al.
Veröffentlicht: (2024)
von: Islam, Chashi Mahiul, et al.
Veröffentlicht: (2024)
An Investigation into Pre-Training Object-Centric Representations for Reinforcement Learning
von: Yoon, Jaesik, et al.
Veröffentlicht: (2023)
von: Yoon, Jaesik, et al.
Veröffentlicht: (2023)
VLLFL: A Vision-Language Model Based Lightweight Federated Learning Framework for Smart Agriculture
von: Li, Long, et al.
Veröffentlicht: (2025)
von: Li, Long, et al.
Veröffentlicht: (2025)
Self-Predictive Dynamics for Generalization of Vision-based Reinforcement Learning
von: Kim, Kyungsoo, et al.
Veröffentlicht: (2025)
von: Kim, Kyungsoo, et al.
Veröffentlicht: (2025)
FuRL: Visual-Language Models as Fuzzy Rewards for Reinforcement Learning
von: Fu, Yuwei, et al.
Veröffentlicht: (2024)
von: Fu, Yuwei, et al.
Veröffentlicht: (2024)
A Retrospect to Multi-prompt Learning across Vision and Language
von: Chen, Ziliang, et al.
Veröffentlicht: (2025)
von: Chen, Ziliang, et al.
Veröffentlicht: (2025)
VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
von: Yu, Xinlei, et al.
Veröffentlicht: (2025)
von: Yu, Xinlei, et al.
Veröffentlicht: (2025)
From Pixels to Predicates: Learning Symbolic World Models via Pretrained Vision-Language Models
von: Athalye, Ashay, et al.
Veröffentlicht: (2024)
von: Athalye, Ashay, et al.
Veröffentlicht: (2024)
Vision-Language Meets the Skeleton: Progressively Distillation with Cross-Modal Knowledge for 3D Action Representation Learning
von: Chen, Yang, et al.
Veröffentlicht: (2024)
von: Chen, Yang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Ingredients for Robotic Diffusion Transformers
von: Dasari, Sudeep, et al.
Veröffentlicht: (2024) -
Training Diffusion Models with Reinforcement Learning
von: Black, Kevin, et al.
Veröffentlicht: (2023) -
mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
von: Pai, Jonas, et al.
Veröffentlicht: (2025) -
Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models
von: Blank, Nils, et al.
Veröffentlicht: (2024) -
Multimodal Spatial Language Maps for Robot Navigation and Manipulation
von: Huang, Chenguang, et al.
Veröffentlicht: (2025)