Guardado en:
| Autores principales: | Miao, Bo, Liu, Weijia, Luo, Jun, Shinnick, Lachlan, Liu, Jian, Hamilton-Smith, Thomas, Yang, Yuhe, Wu, Zijie, Videnovic, Vanja, Dayoub, Feras, Hengel, Anton van den |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2602.02220 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals
por: Podgorski, Stefan, et al.
Publicado: (2025)
por: Podgorski, Stefan, et al.
Publicado: (2025)
RoboHop: Segment-based Topological Map Representation for Open-World Visual Navigation
por: Garg, Sourav, et al.
Publicado: (2024)
por: Garg, Sourav, et al.
Publicado: (2024)
Procedural Pretraining: Warming Up Language Models with Abstract Data
por: Jiang, Liangze, et al.
Publicado: (2026)
por: Jiang, Liangze, et al.
Publicado: (2026)
Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers
por: Shinnick, Zachary, et al.
Publicado: (2025)
por: Shinnick, Zachary, et al.
Publicado: (2025)
Transformers Pretrained on Procedural Data Contain Modular Structures for Algorithmic Reasoning
por: Shinnick, Zachary, et al.
Publicado: (2025)
por: Shinnick, Zachary, et al.
Publicado: (2025)
To Ask or Not to Ask? Detecting Absence of Information in Vision and Language Navigation
por: Abraham, Savitha Sam, et al.
Publicado: (2024)
por: Abraham, Savitha Sam, et al.
Publicado: (2024)
AARK: An Open Toolkit for Autonomous Racing Research
por: Bockman, James, et al.
Publicado: (2024)
por: Bockman, James, et al.
Publicado: (2024)
ObjectReact: Learning Object-Relative Control for Visual Navigation
por: Garg, Sourav, et al.
Publicado: (2025)
por: Garg, Sourav, et al.
Publicado: (2025)
LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part Segmentation
por: Miao, Yang, et al.
Publicado: (2025)
por: Miao, Yang, et al.
Publicado: (2025)
SceneEdited: A City-Scale Benchmark for 3D HD Map Updating via Image-Guided Change Detection
por: Lin, Chun-Jung, et al.
Publicado: (2025)
por: Lin, Chun-Jung, et al.
Publicado: (2025)
Vision Foundation Models for Domain Generalisable Cross-View Localisation in Planetary Ground-Aerial Robotic Teams
por: Holden, Lachlan, et al.
Publicado: (2026)
por: Holden, Lachlan, et al.
Publicado: (2026)
Embodied Domain Adaptation for Object Detection
por: Shi, Xiangyu, et al.
Publicado: (2025)
por: Shi, Xiangyu, et al.
Publicado: (2025)
Temporal Attention for Cross-View Sequential Image Localization
por: Yuan, Dong, et al.
Publicado: (2024)
por: Yuan, Dong, et al.
Publicado: (2024)
A Physical Agentic Loop for Language-Guided Grasping with Execution-State Monitoring
por: Wang, Wenze, et al.
Publicado: (2026)
por: Wang, Wenze, et al.
Publicado: (2026)
Wasserstein Distance-based Expansion of Low-Density Latent Regions for Unknown Class Detection
por: Mallick, Prakash, et al.
Publicado: (2024)
por: Mallick, Prakash, et al.
Publicado: (2024)
AIMC-Spec: A Benchmark Dataset for Automatic Intrapulse Modulation Classification under Variable Noise Conditions
por: Cocks, Sebastian L., et al.
Publicado: (2026)
por: Cocks, Sebastian L., et al.
Publicado: (2026)
Hybrid Navigation Acceptability and Safety
por: Clement, Benoit, et al.
Publicado: (2024)
por: Clement, Benoit, et al.
Publicado: (2024)
KITE: Keyframe-Indexed Tokenized Evidence for VLM-Based Robot Failure Analysis
por: Hosseinzadeh, Mehdi, et al.
Publicado: (2026)
por: Hosseinzadeh, Mehdi, et al.
Publicado: (2026)
Detecting Precise Hand Touch Moments in Egocentric Video
por: Nguyen, Huy Anh, et al.
Publicado: (2026)
por: Nguyen, Huy Anh, et al.
Publicado: (2026)
OVAL: Open-Vocabulary Augmented Memory Model for Lifelong Object Goal Navigation
por: Pei, Jiahua, et al.
Publicado: (2026)
por: Pei, Jiahua, et al.
Publicado: (2026)
Improving Online Source-free Domain Adaptation for Object Detection by Unsupervised Data Acquisition
por: Shi, Xiangyu, et al.
Publicado: (2023)
por: Shi, Xiangyu, et al.
Publicado: (2023)
QueryAdapter: Rapid Adaptation of Vision-Language Models in Response to Natural Language Queries
por: Chapman, Nicolas Harvey, et al.
Publicado: (2025)
por: Chapman, Nicolas Harvey, et al.
Publicado: (2025)
Enhancing Embodied Object Detection through Language-Image Pre-training and Implicit Object Memory
por: Chapman, Nicolas Harvey, et al.
Publicado: (2024)
por: Chapman, Nicolas Harvey, et al.
Publicado: (2024)
HM3D-OVON: A Dataset and Benchmark for Open-Vocabulary Object Goal Navigation
por: Yokoyama, Naoki, et al.
Publicado: (2024)
por: Yokoyama, Naoki, et al.
Publicado: (2024)
SmartWay: Enhanced Waypoint Prediction and Backtracking for Zero-Shot Vision-and-Language Navigation
por: Shi, Xiangyu, et al.
Publicado: (2025)
por: Shi, Xiangyu, et al.
Publicado: (2025)
Physically Embodied Gaussian Splatting: A Realtime Correctable World Model for Robotics
por: Abou-Chakra, Jad, et al.
Publicado: (2024)
por: Abou-Chakra, Jad, et al.
Publicado: (2024)
Segment Beyond View: Handling Partially Missing Modality for Audio-Visual Semantic Segmentation
por: Wu, Renjie, et al.
Publicado: (2023)
por: Wu, Renjie, et al.
Publicado: (2023)
OVSegDT: Segmenting Transformer for Open-Vocabulary Object Goal Navigation
por: Zemskova, Tatiana, et al.
Publicado: (2025)
por: Zemskova, Tatiana, et al.
Publicado: (2025)
Grounded Vision-Language Navigation for UAVs with Open-Vocabulary Goal Understanding
por: Zhang, Yuhang, et al.
Publicado: (2025)
por: Zhang, Yuhang, et al.
Publicado: (2025)
Uncertainty-Informed Active Perception for Open Vocabulary Object Goal Navigation
por: Bajpai, Utkarsh, et al.
Publicado: (2025)
por: Bajpai, Utkarsh, et al.
Publicado: (2025)
Points-to-3D: Structure-Aware 3D Generation with Point Cloud Priors
por: Xia, Jiatong, et al.
Publicado: (2026)
por: Xia, Jiatong, et al.
Publicado: (2026)
Robust Scene Change Detection Using Visual Foundation Models and Cross-Attention Mechanisms
por: Lin, Chun-Jung, et al.
Publicado: (2024)
por: Lin, Chun-Jung, et al.
Publicado: (2024)
GoalSwarm: Multi-UAV Semantic Coordination for Open-Vocabulary Object Navigation
por: James, MoniJesu Wonders, et al.
Publicado: (2026)
por: James, MoniJesu Wonders, et al.
Publicado: (2026)
Hierarchical Process Reward Models are Symbolic Vision Learners
por: Zhang, Shan, et al.
Publicado: (2025)
por: Zhang, Shan, et al.
Publicado: (2025)
Query-Based Knowledge Sharing for Open-Vocabulary Multi-Label Classification
por: Zhu, Xuelin, et al.
Publicado: (2024)
por: Zhu, Xuelin, et al.
Publicado: (2024)
UAV-ON: A Benchmark for Open-World Object Goal Navigation with Aerial Agents
por: Xiao, Jianqiang, et al.
Publicado: (2025)
por: Xiao, Jianqiang, et al.
Publicado: (2025)
Let Your Video Listen to Your Music!
por: Zhang, Xinyu, et al.
Publicado: (2025)
por: Zhang, Xinyu, et al.
Publicado: (2025)
MuseBarControl: Enhancing Fine-Grained Control in Symbolic Music Generation through Pre-Training and Counterfactual Loss
por: Shu, Yangyang, et al.
Publicado: (2024)
por: Shu, Yangyang, et al.
Publicado: (2024)
Scaling up Multi-domain Semantic Segmentation with Sentence Embeddings
por: Yin, Wei, et al.
Publicado: (2022)
por: Yin, Wei, et al.
Publicado: (2022)
VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation
por: Lin, Sihao, et al.
Publicado: (2025)
por: Lin, Sihao, et al.
Publicado: (2025)
Ejemplares similares
-
TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals
por: Podgorski, Stefan, et al.
Publicado: (2025) -
RoboHop: Segment-based Topological Map Representation for Open-World Visual Navigation
por: Garg, Sourav, et al.
Publicado: (2024) -
Procedural Pretraining: Warming Up Language Models with Abstract Data
por: Jiang, Liangze, et al.
Publicado: (2026) -
Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers
por: Shinnick, Zachary, et al.
Publicado: (2025) -
Transformers Pretrained on Procedural Data Contain Modular Structures for Algorithmic Reasoning
por: Shinnick, Zachary, et al.
Publicado: (2025)