QueryAdapter: Rapid Adaptation of Vision-Language Models in Response to Natural Language Queries
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chapman, Nicolas Harvey, Dayoub, Feras, Browne, Will, Lehnert, Christopher |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhancing Embodied Object Detection through Language-Image Pre-training and Implicit Object Memory
von: Chapman, Nicolas Harvey, et al.
Veröffentlicht: (2024)
von: Chapman, Nicolas Harvey, et al.
Veröffentlicht: (2024)
Vision Foundation Models for Domain Generalisable Cross-View Localisation in Planetary Ground-Aerial Robotic Teams
von: Holden, Lachlan, et al.
Veröffentlicht: (2026)
von: Holden, Lachlan, et al.
Veröffentlicht: (2026)
Embodied Domain Adaptation for Object Detection
von: Shi, Xiangyu, et al.
Veröffentlicht: (2025)
von: Shi, Xiangyu, et al.
Veröffentlicht: (2025)
A Physical Agentic Loop for Language-Guided Grasping with Execution-State Monitoring
von: Wang, Wenze, et al.
Veröffentlicht: (2026)
von: Wang, Wenze, et al.
Veröffentlicht: (2026)
SmartWay: Enhanced Waypoint Prediction and Backtracking for Zero-Shot Vision-and-Language Navigation
von: Shi, Xiangyu, et al.
Veröffentlicht: (2025)
von: Shi, Xiangyu, et al.
Veröffentlicht: (2025)
KITE: Keyframe-Indexed Tokenized Evidence for VLM-Based Robot Failure Analysis
von: Hosseinzadeh, Mehdi, et al.
Veröffentlicht: (2026)
von: Hosseinzadeh, Mehdi, et al.
Veröffentlicht: (2026)
To Ask or Not to Ask? Detecting Absence of Information in Vision and Language Navigation
von: Abraham, Savitha Sam, et al.
Veröffentlicht: (2024)
von: Abraham, Savitha Sam, et al.
Veröffentlicht: (2024)
QueryOcc: Query-based Self-Supervision for 3D Semantic Occupancy
von: Lilja, Adam, et al.
Veröffentlicht: (2025)
von: Lilja, Adam, et al.
Veröffentlicht: (2025)
FlexVLN: Flexible Adaptation for Diverse Vision-and-Language Navigation Tasks
von: Zhang, Siqi, et al.
Veröffentlicht: (2025)
von: Zhang, Siqi, et al.
Veröffentlicht: (2025)
Turning Adaptation into Assets: Cross-Domain Bridging for Online Vision-Language Navigation
von: Hu, Zixuan, et al.
Veröffentlicht: (2026)
von: Hu, Zixuan, et al.
Veröffentlicht: (2026)
Temporal Attention for Cross-View Sequential Image Localization
von: Yuan, Dong, et al.
Veröffentlicht: (2024)
von: Yuan, Dong, et al.
Veröffentlicht: (2024)
Wasserstein Distance-based Expansion of Low-Density Latent Regions for Unknown Class Detection
von: Mallick, Prakash, et al.
Veröffentlicht: (2024)
von: Mallick, Prakash, et al.
Veröffentlicht: (2024)
HarvestFlex: Strawberry Harvesting via Vision-Language-Action Policy Adaptation in the Wild
von: Zhao, Ziyang, et al.
Veröffentlicht: (2026)
von: Zhao, Ziyang, et al.
Veröffentlicht: (2026)
Unified Vision-Language-Action Model
von: Wang, Yuqi, et al.
Veröffentlicht: (2025)
von: Wang, Yuqi, et al.
Veröffentlicht: (2025)
Improving Online Source-free Domain Adaptation for Object Detection by Unsupervised Data Acquisition
von: Shi, Xiangyu, et al.
Veröffentlicht: (2023)
von: Shi, Xiangyu, et al.
Veröffentlicht: (2023)
LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation
von: Miao, Bo, et al.
Veröffentlicht: (2026)
von: Miao, Bo, et al.
Veröffentlicht: (2026)
DC-VLAQ: Query-Residual Aggregation for Robust Visual Place Recognition
von: Zhu, Hanyu, et al.
Veröffentlicht: (2026)
von: Zhu, Hanyu, et al.
Veröffentlicht: (2026)
Generalized Robot 3D Vision-Language Model with Fast Rendering and Pre-Training Vision-Language Alignment
von: Liu, Kangcheng, et al.
Veröffentlicht: (2023)
von: Liu, Kangcheng, et al.
Veröffentlicht: (2023)
Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
von: Han, Xiaofeng, et al.
Veröffentlicht: (2025)
von: Han, Xiaofeng, et al.
Veröffentlicht: (2025)
Detecting Precise Hand Touch Moments in Egocentric Video
von: Nguyen, Huy Anh, et al.
Veröffentlicht: (2026)
von: Nguyen, Huy Anh, et al.
Veröffentlicht: (2026)
Query3D: LLM-Powered Open-Vocabulary Scene Segmentation with Language Embedded 3D Gaussian
von: Chahe, Amirhosein, et al.
Veröffentlicht: (2024)
von: Chahe, Amirhosein, et al.
Veröffentlicht: (2024)
Attention Hijacking: Response Manipulation Across Queries in Vision-Language Models
von: Wang, Zhiqiang, et al.
Veröffentlicht: (2026)
von: Wang, Zhiqiang, et al.
Veröffentlicht: (2026)
TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals
von: Podgorski, Stefan, et al.
Veröffentlicht: (2025)
von: Podgorski, Stefan, et al.
Veröffentlicht: (2025)
InVDriver: Intra-Instance Aware Vectorized Query-Based Autonomous Driving Transformer
von: Zhang, Bo, et al.
Veröffentlicht: (2025)
von: Zhang, Bo, et al.
Veröffentlicht: (2025)
QueST: Persistent Queries as Semantic Monitors for Drift Suppression in Long-Horizon Tracking
von: Anand, Mayank, et al.
Veröffentlicht: (2026)
von: Anand, Mayank, et al.
Veröffentlicht: (2026)
Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models
von: Yang, Yurou, et al.
Veröffentlicht: (2026)
von: Yang, Yurou, et al.
Veröffentlicht: (2026)
NEARL-CLIP: Interacted Query Adaptation with Orthogonal Regularization for Medical Vision-Language Understanding
von: Peng, Zelin, et al.
Veröffentlicht: (2025)
von: Peng, Zelin, et al.
Veröffentlicht: (2025)
DexVLG: Dexterous Vision-Language-Grasp Model at Scale
von: He, Jiawei, et al.
Veröffentlicht: (2025)
von: He, Jiawei, et al.
Veröffentlicht: (2025)
LLaDA-VLA: Vision Language Diffusion Action Models
von: Wen, Yuqing, et al.
Veröffentlicht: (2025)
von: Wen, Yuqing, et al.
Veröffentlicht: (2025)
BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
von: Bhat, Vineet, et al.
Veröffentlicht: (2025)
von: Bhat, Vineet, et al.
Veröffentlicht: (2025)
AffordanceLLM: Grounding Affordance from Vision Language Models
von: Qian, Shengyi, et al.
Veröffentlicht: (2024)
von: Qian, Shengyi, et al.
Veröffentlicht: (2024)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023)
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023)
Towards Open-World Grasping with Large Vision-Language Models
von: Tziafas, Georgios, et al.
Veröffentlicht: (2024)
von: Tziafas, Georgios, et al.
Veröffentlicht: (2024)
Structured Observation Language for Efficient and Generalizable Vision-Language Navigation
von: Peng, Daojie, et al.
Veröffentlicht: (2026)
von: Peng, Daojie, et al.
Veröffentlicht: (2026)
Query-Calibrated Segmental Admission for Descriptor-Agnostic LiDAR Loop Closure in Repetitive Environments
von: Kim, Jaehyun, et al.
Veröffentlicht: (2025)
von: Kim, Jaehyun, et al.
Veröffentlicht: (2025)
StreamLTS: Query-based Temporal-Spatial LiDAR Fusion for Cooperative Object Detection
von: Yuan, Yunshuang, et al.
Veröffentlicht: (2024)
von: Yuan, Yunshuang, et al.
Veröffentlicht: (2024)
MAG-VLAQ: Multi-modal Aerial-Ground Query Aggregation for Cross-View Place Recognition
von: Xu, Zhengyi, et al.
Veröffentlicht: (2026)
von: Xu, Zhengyi, et al.
Veröffentlicht: (2026)
Natural Language Instructions for Scene-Responsive Human-in-the-Loop Motion Planning in Autonomous Driving using Vision-Language-Action Models
von: Martinez-Sanchez, Angel, et al.
Veröffentlicht: (2026)
von: Martinez-Sanchez, Angel, et al.
Veröffentlicht: (2026)
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
von: Won, John, et al.
Veröffentlicht: (2025)
von: Won, John, et al.
Veröffentlicht: (2025)
WMNav: Integrating Vision-Language Models into World Models for Object Goal Navigation
von: Nie, Dujun, et al.
Veröffentlicht: (2025)
von: Nie, Dujun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Enhancing Embodied Object Detection through Language-Image Pre-training and Implicit Object Memory
von: Chapman, Nicolas Harvey, et al.
Veröffentlicht: (2024) -
Vision Foundation Models for Domain Generalisable Cross-View Localisation in Planetary Ground-Aerial Robotic Teams
von: Holden, Lachlan, et al.
Veröffentlicht: (2026) -
Embodied Domain Adaptation for Object Detection
von: Shi, Xiangyu, et al.
Veröffentlicht: (2025) -
A Physical Agentic Loop for Language-Guided Grasping with Execution-State Monitoring
von: Wang, Wenze, et al.
Veröffentlicht: (2026) -
SmartWay: Enhanced Waypoint Prediction and Backtracking for Zero-Shot Vision-and-Language Navigation
von: Shi, Xiangyu, et al.
Veröffentlicht: (2025)