Hierarchical Instruction-aware Embodied Visual Tracking
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Kui, Chen, Hao, Wang, Churan, Karray, Fakhri, Li, Zhoujun, Wang, Yizhou, Zhong, Fangwei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models
by: Wu, Kui, et al.
Published: (2025)
by: Wu, Kui, et al.
Published: (2025)
Empowering Embodied Visual Tracking with Visual Foundation Models and Offline RL
by: Zhong, Fangwei, et al.
Published: (2024)
by: Zhong, Fangwei, et al.
Published: (2024)
UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI
by: Zhong, Fangwei, et al.
Published: (2024)
by: Zhong, Fangwei, et al.
Published: (2024)
RescueBench: Can Embodied Agents Save Lives in the Wild ?
by: Wu, Kui, et al.
Published: (2026)
by: Wu, Kui, et al.
Published: (2026)
Clinical Inspired MRI Lesion Segmentation
by: Yan, Lijun, et al.
Published: (2025)
by: Yan, Lijun, et al.
Published: (2025)
TrackVLA: Embodied Visual Tracking in the Wild
by: Wang, Shaoan, et al.
Published: (2025)
by: Wang, Shaoan, et al.
Published: (2025)
TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking
by: Liu, Jiahang, et al.
Published: (2025)
by: Liu, Jiahang, et al.
Published: (2025)
Projected Gradient Unlearning for Text-to-Image Diffusion Models: Defending Against Concept Revival Attacks
by: Aladawi, Aljalila, et al.
Published: (2026)
by: Aladawi, Aljalila, et al.
Published: (2026)
ADAM-Dehaze: Adaptive Density-Aware Multi-Stage Dehazing for Improved Object Detection in Foggy Conditions
by: AlHindaassi, Fatmah, et al.
Published: (2025)
by: AlHindaassi, Fatmah, et al.
Published: (2025)
Autoregressive Sequence Modeling for 3D Medical Image Representation
by: Wang, Siwen, et al.
Published: (2024)
by: Wang, Siwen, et al.
Published: (2024)
EmbRACE-3K: Embodied Reasoning and Action in Complex Environments
by: Lin, Mingxian, et al.
Published: (2025)
by: Lin, Mingxian, et al.
Published: (2025)
FaceAnonyMixer: Cancelable Faces via Identity Consistent Latent Space Mixing
by: Alam, Mohammed Talha, et al.
Published: (2025)
by: Alam, Mohammed Talha, et al.
Published: (2025)
AstroSpy: On detecting Fake Images in Astronomy via Joint Image-Spectral Representations
by: Alam, Mohammed Talha, et al.
Published: (2024)
by: Alam, Mohammed Talha, et al.
Published: (2024)
FLARE up your data: Diffusion-based Augmentation Method in Astronomical Imaging
by: Alam, Mohammed Talha, et al.
Published: (2024)
by: Alam, Mohammed Talha, et al.
Published: (2024)
Cross-Dimensional Medical Self-Supervised Representation Learning Based on a Pseudo-3D Transformation
by: Gao, Fei, et al.
Published: (2024)
by: Gao, Fei, et al.
Published: (2024)
Clinically-Grounded Counterfactual Reasoning for Medical Video Diagnosis
by: Gao, Jianzhe, et al.
Published: (2026)
by: Gao, Jianzhe, et al.
Published: (2026)
CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging
by: Imam, Raza, et al.
Published: (2024)
by: Imam, Raza, et al.
Published: (2024)
LLM Enhanced Action Recognition via Hierarchical Global-Local Skeleton-Language Model
by: Wang, Ruosi, et al.
Published: (2026)
by: Wang, Ruosi, et al.
Published: (2026)
NL-MambaXCT: Self-Supervised Nested-Learning Mamba for Nomex Honeycomb X-ray CT Defect Classification
by: Aldoboni, Ghaleb, et al.
Published: (2026)
by: Aldoboni, Ghaleb, et al.
Published: (2026)
Vision Language Models for Dynamic Human Activity Recognition in Healthcare Settings
by: Abid, Abderrazek, et al.
Published: (2025)
by: Abid, Abderrazek, et al.
Published: (2025)
CRAVES: Controlling Robotic Arm with a Vision-based Economic System
by: Zuo, Yiming, et al.
Published: (2018)
by: Zuo, Yiming, et al.
Published: (2018)
Enhanced MRI Representation via Cross-series Masking
by: Wang, Churan, et al.
Published: (2024)
by: Wang, Churan, et al.
Published: (2024)
Confusing Pair Correction Based on Category Prototype for Domain Adaptation under Noisy Environments
by: Zhi, Churan, et al.
Published: (2024)
by: Zhi, Churan, et al.
Published: (2024)
Towards Context-aware Convolutional Network for Image Restoration
by: Hao, Fangwei, et al.
Published: (2024)
by: Hao, Fangwei, et al.
Published: (2024)
Semi- and Weakly-Supervised Learning for Mammogram Mass Segmentation with Limited Annotations
by: Xiong, Xinyu, et al.
Published: (2024)
by: Xiong, Xinyu, et al.
Published: (2024)
RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented Generation
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
EmbodiedPlace: Learning Mixture-of-Features with Embodied Constraints for Visual Place Recognition
by: Liu, Bingxi, et al.
Published: (2025)
by: Liu, Bingxi, et al.
Published: (2025)
UrbanGS: Semantic-Guided Gaussian Splatting for Urban Scene Reconstruction
by: Li, Ziwen, et al.
Published: (2024)
by: Li, Ziwen, et al.
Published: (2024)
SAUCE: Selective Concept Unlearning in Vision-Language Models with Sparse Autoencoders
by: Li, Qing, et al.
Published: (2025)
by: Li, Qing, et al.
Published: (2025)
AdaptPrompt: Parameter-Efficient Adaptation of VLMs for Generalizable Deepfake Detection
by: Jiang, Yichen, et al.
Published: (2025)
by: Jiang, Yichen, et al.
Published: (2025)
REMONI: An Autonomous System Integrating Wearables and Multimodal Large Language Models for Enhanced Remote Health Monitoring
by: Ho, Thanh Cong, et al.
Published: (2025)
by: Ho, Thanh Cong, et al.
Published: (2025)
AdaTracker: Learning Adaptive In-Context Policy for Cross-Embodiment Active Visual Tracking
by: Wu, Kui, et al.
Published: (2026)
by: Wu, Kui, et al.
Published: (2026)
LUMOS: Universal Semi-Supervised OCT Retinal Layer Segmentation with Hierarchical Reliable Mutual Learning
by: Fang, Yizhou, et al.
Published: (2026)
by: Fang, Yizhou, et al.
Published: (2026)
Envision: Embodied Visual Planning via Goal-Imagery Video Diffusion
by: Gu, Yuming, et al.
Published: (2025)
by: Gu, Yuming, et al.
Published: (2025)
Internal Activation Revision: Safeguarding Vision Language Models Without Parameter Update
by: Li, Qing, et al.
Published: (2025)
by: Li, Qing, et al.
Published: (2025)
Multi-Object Tracking by Hierarchical Visual Representations
by: Cao, Jinkun, et al.
Published: (2024)
by: Cao, Jinkun, et al.
Published: (2024)
Explicit Visual Prompts for Visual Object Tracking
by: Shi, Liangtao, et al.
Published: (2024)
by: Shi, Liangtao, et al.
Published: (2024)
Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning
by: Dong, Qihua, et al.
Published: (2026)
by: Dong, Qihua, et al.
Published: (2026)
MovieDreamer: Hierarchical Generation for Coherent Long Visual Sequence
by: Zhao, Canyu, et al.
Published: (2024)
by: Zhao, Canyu, et al.
Published: (2024)
Less is More: Token Context-aware Learning for Object Tracking
by: Xu, Chenlong, et al.
Published: (2025)
by: Xu, Chenlong, et al.
Published: (2025)
Similar Items
-
VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models
by: Wu, Kui, et al.
Published: (2025) -
Empowering Embodied Visual Tracking with Visual Foundation Models and Offline RL
by: Zhong, Fangwei, et al.
Published: (2024) -
UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI
by: Zhong, Fangwei, et al.
Published: (2024) -
RescueBench: Can Embodied Agents Save Lives in the Wild ?
by: Wu, Kui, et al.
Published: (2026) -
Clinical Inspired MRI Lesion Segmentation
by: Yan, Lijun, et al.
Published: (2025)