A ROS 2 Wrapper for Florence-2: Multi-Mode Local Vision-Language Inference for Robotic Systems
Fuente:
arXiv
Saved in:
| Main Author: | Domínguez-Vidal, J. E. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation
by: Li, Zaijing, et al.
Published: (2026)
by: Li, Zaijing, et al.
Published: (2026)
Physically Grounded Vision-Language Models for Robotic Manipulation
by: Gao, Jensen, et al.
Published: (2023)
by: Gao, Jensen, et al.
Published: (2023)
AgriVLN: Vision-and-Language Navigation for Agricultural Robots
by: Zhao, Xiaobei, et al.
Published: (2025)
by: Zhao, Xiaobei, et al.
Published: (2025)
ROSA: Harnessing Robot States for Vision-Language and Action Alignment
by: Wen, Yuqing, et al.
Published: (2025)
by: Wen, Yuqing, et al.
Published: (2025)
Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation
by: Chen, Shizhe, et al.
Published: (2025)
by: Chen, Shizhe, et al.
Published: (2025)
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
by: Yuan, Wentao, et al.
Published: (2024)
by: Yuan, Wentao, et al.
Published: (2024)
ManiSoft: Towards Vision-Language Manipulation for Soft Continuum Robotics
by: Wei, Ziyu, et al.
Published: (2026)
by: Wei, Ziyu, et al.
Published: (2026)
Explainable Adversarial-Robust Vision-Language-Action Model for Robotic Manipulation
by: Kim, Ju-Young, et al.
Published: (2025)
by: Kim, Ju-Young, et al.
Published: (2025)
Visual Perception Engine: Fast and Flexible Multi-Head Inference for Robotic Vision Tasks
by: Łucki, Jakub, et al.
Published: (2025)
by: Łucki, Jakub, et al.
Published: (2025)
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
by: Zhou, Enshen, et al.
Published: (2025)
by: Zhou, Enshen, et al.
Published: (2025)
TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
by: Han, Yi, et al.
Published: (2025)
by: Han, Yi, et al.
Published: (2025)
Unifying 2D and 3D Vision-Language Understanding
by: Jain, Ayush, et al.
Published: (2025)
by: Jain, Ayush, et al.
Published: (2025)
Advances and Innovations in the Multi-Agent Robotic System (MARS) Challenge
by: Kang, Li, et al.
Published: (2026)
by: Kang, Li, et al.
Published: (2026)
Stable Multi-Drone GNSS Tracking System for Marine Robots
by: Wen, Shuo, et al.
Published: (2025)
by: Wen, Shuo, et al.
Published: (2025)
PRISM: A Multi-View Multi-Capability Retail Video Dataset for Embodied Vision-Language Models
by: Rouhi, Amirreza, et al.
Published: (2026)
by: Rouhi, Amirreza, et al.
Published: (2026)
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
by: Song, Chan Hee, et al.
Published: (2024)
by: Song, Chan Hee, et al.
Published: (2024)
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
by: Chen, Xinyi, et al.
Published: (2025)
by: Chen, Xinyi, et al.
Published: (2025)
VitaTouch: Property-Aware Vision-Tactile-Language Model for Robotic Quality Inspection in Manufacturing
by: Zong, Junyi, et al.
Published: (2026)
by: Zong, Junyi, et al.
Published: (2026)
SKT: Integrating State-Aware Keypoint Trajectories with Vision-Language Models for Robotic Garment Manipulation
by: Li, Xin, et al.
Published: (2024)
by: Li, Xin, et al.
Published: (2024)
Robotic Environmental State Recognition with Pre-Trained Vision-Language Models and Black-Box Optimization
by: Kawaharazuka, Kento, et al.
Published: (2024)
by: Kawaharazuka, Kento, et al.
Published: (2024)
On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations
by: Guo, Jianing, et al.
Published: (2025)
by: Guo, Jianing, et al.
Published: (2025)
RoboFlamingo-Plus: Fusion of Depth and RGB Perception with Vision-Language Models for Enhanced Robotic Manipulation
by: Wang, Sheng
Published: (2025)
by: Wang, Sheng
Published: (2025)
StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems
by: Ye, Jinhui, et al.
Published: (2026)
by: Ye, Jinhui, et al.
Published: (2026)
Privacy-Preserving Multi-Stage Fall Detection Framework with Semi-supervised Federated Learning and Robotic Vision Confirmation
by: Azghadi, Seyed Alireza Rahimi, et al.
Published: (2025)
by: Azghadi, Seyed Alireza Rahimi, et al.
Published: (2025)
PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation
by: Fan, Qingyu, et al.
Published: (2026)
by: Fan, Qingyu, et al.
Published: (2026)
Autonomous Robot for Disaster Mapping and Victim Localization
by: Potter, Michael, et al.
Published: (2024)
by: Potter, Michael, et al.
Published: (2024)
A Large Vision-Language Model based Environment Perception System for Visually Impaired People
by: Chen, Zezhou, et al.
Published: (2025)
by: Chen, Zezhou, et al.
Published: (2025)
Integrating Deep RL and Bayesian Inference for ObjectNav in Mobile Robotics
by: Castelo-Branco, João, et al.
Published: (2026)
by: Castelo-Branco, João, et al.
Published: (2026)
Robotic State Recognition with Image-to-Text Retrieval Task of Pre-Trained Vision-Language Model and Black-Box Optimization
by: Kawaharazuka, Kento, et al.
Published: (2024)
by: Kawaharazuka, Kento, et al.
Published: (2024)
ActiveVLN: Towards Active Exploration via Multi-Turn RL in Vision-and-Language Navigation
by: Zhang, Zekai, et al.
Published: (2025)
by: Zhang, Zekai, et al.
Published: (2025)
A Systematic Literature Review of Computer Vision Applications in Robotized Wire Harness Assembly
by: Wang, Hao, et al.
Published: (2023)
by: Wang, Hao, et al.
Published: (2023)
A Brief Survey on Leveraging Large Scale Vision Models for Enhanced Robot Grasping
by: Kamboj, Abhi, et al.
Published: (2024)
by: Kamboj, Abhi, et al.
Published: (2024)
Robotic Grasping of Harvested Tomato Trusses Using Vision and Online Learning
by: Bent, Luuk van den, et al.
Published: (2023)
by: Bent, Luuk van den, et al.
Published: (2023)
Probing Collision Grounding in Vision-Language Models for Safe Human-Robot Collaboration
by: Wang, Jun, et al.
Published: (2026)
by: Wang, Jun, et al.
Published: (2026)
Zero-Shot Peg Insertion: Identifying Mating Holes and Estimating SE(2) Poses with Vision-Language Models
by: Yajima, Masaru, et al.
Published: (2025)
by: Yajima, Masaru, et al.
Published: (2025)
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
by: Zhou, Zhongyi, et al.
Published: (2025)
by: Zhou, Zhongyi, et al.
Published: (2025)
Vision-Language Navigation with Embodied Intelligence: A Survey
by: Gao, Peng, et al.
Published: (2024)
by: Gao, Peng, et al.
Published: (2024)
A Navigation Framework Utilizing Vision-Language Models
by: Duan, Yicheng, et al.
Published: (2025)
by: Duan, Yicheng, et al.
Published: (2025)
Multi-Modal World Model for Physical Robot Interactions: Simultaneous Visual and Tactile Predictions for Enhanced Accuracy
by: Mandil, Willow, et al.
Published: (2023)
by: Mandil, Willow, et al.
Published: (2023)
A Survey on Vision-Language-Action Models for Autonomous Driving
by: Jiang, Sicong, et al.
Published: (2025)
by: Jiang, Sicong, et al.
Published: (2025)
Similar Items
-
Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation
by: Li, Zaijing, et al.
Published: (2026) -
Physically Grounded Vision-Language Models for Robotic Manipulation
by: Gao, Jensen, et al.
Published: (2023) -
AgriVLN: Vision-and-Language Navigation for Agricultural Robots
by: Zhao, Xiaobei, et al.
Published: (2025) -
ROSA: Harnessing Robot States for Vision-Language and Action Alignment
by: Wen, Yuqing, et al.
Published: (2025) -
Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation
by: Chen, Shizhe, et al.
Published: (2025)