Towards Visual Query Segmentation in the Wild
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fan, Bing, Li, Minghao, Zhang, Hanzhi, Dong, Shaohua, Mareedu, Naga Prudhvi, Shi, Weishi, Feng, Yunhe, Huang, Yan, Fan, Heng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PRVQL: Progressive Knowledge-guided Refinement for Robust Egocentric Visual Query Localization
von: Fan, Bing, et al.
Veröffentlicht: (2025)
von: Fan, Bing, et al.
Veröffentlicht: (2025)
LoReTrack: Efficient and Accurate Low-Resolution Transformer Tracking
von: Dong, Shaohua, et al.
Veröffentlicht: (2024)
von: Dong, Shaohua, et al.
Veröffentlicht: (2024)
Towards Long-Form Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2026)
von: Gu, Xin, et al.
Veröffentlicht: (2026)
VastTrack: Vast Category Visual Object Tracking
von: Peng, Liang, et al.
Veröffentlicht: (2024)
von: Peng, Liang, et al.
Veröffentlicht: (2024)
Benchmarking the Robustness of UAV Tracking Against Common Corruptions
von: Liu, Xiaoqiong, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoqiong, et al.
Veröffentlicht: (2024)
OmniSTVG: Toward Spatio-Temporal Omni-Object Video Grounding
von: Yao, Jiali, et al.
Veröffentlicht: (2025)
von: Yao, Jiali, et al.
Veröffentlicht: (2025)
DECO: Unleashing the Potential of ConvNets for Query-based Detection and Segmentation
von: Chen, Xinghao, et al.
Veröffentlicht: (2023)
von: Chen, Xinghao, et al.
Veröffentlicht: (2023)
GSOT3D: Towards Generic 3D Single Object Tracking in the Wild
von: Jiao, Yifan, et al.
Veröffentlicht: (2024)
von: Jiao, Yifan, et al.
Veröffentlicht: (2024)
Towards Visual Query Localization in the 3D World
von: Peng, Liang, et al.
Veröffentlicht: (2026)
von: Peng, Liang, et al.
Veröffentlicht: (2026)
DMTrack: Spatio-Temporal Multimodal Tracking via Dual-Adapter
von: Li, Weihong, et al.
Veröffentlicht: (2025)
von: Li, Weihong, et al.
Veröffentlicht: (2025)
Show Me When and Where: Towards Referring Video Object Segmentation in the Wild
von: Gao, Mingqi, et al.
Veröffentlicht: (2026)
von: Gao, Mingqi, et al.
Veröffentlicht: (2026)
Diagnosing Shortcut-Induced Rigidity in Continual Learning: The Einstellung Rigidity Index (ERI)
von: Gu, Kai, et al.
Veröffentlicht: (2025)
von: Gu, Kai, et al.
Veröffentlicht: (2025)
Beyond MOT: Semantic Multi-Object Tracking
von: Li, Yunhao, et al.
Veröffentlicht: (2024)
von: Li, Yunhao, et al.
Veröffentlicht: (2024)
Med-Query: Steerable Parsing of 9-DoF Medical Anatomies with Query Embedding
von: Guo, Heng, et al.
Veröffentlicht: (2022)
von: Guo, Heng, et al.
Veröffentlicht: (2022)
Learning to Segment Liquids in Real-world Images
von: Li, Jonas, et al.
Veröffentlicht: (2026)
von: Li, Jonas, et al.
Veröffentlicht: (2026)
Brought a Gun to a Knife Fight: Modern VFM Baselines Outgun Specialized Detectors on In-the-Wild AI Image Detection
von: Zhou, Yue, et al.
Veröffentlicht: (2025)
von: Zhou, Yue, et al.
Veröffentlicht: (2025)
VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation
von: Chen, Hanzhi, et al.
Veröffentlicht: (2025)
von: Chen, Hanzhi, et al.
Veröffentlicht: (2025)
Visual Text Generation in the Wild
von: Zhu, Yuanzhi, et al.
Veröffentlicht: (2024)
von: Zhu, Yuanzhi, et al.
Veröffentlicht: (2024)
Thinking in 360°: Humanoid Visual Search in the Wild
von: Yu, Heyang, et al.
Veröffentlicht: (2025)
von: Yu, Heyang, et al.
Veröffentlicht: (2025)
Adaptive Context Matters: Towards Provable Multi-Modality Guidance for Super-Resolution
von: Luo, Jinyi, et al.
Veröffentlicht: (2026)
von: Luo, Jinyi, et al.
Veröffentlicht: (2026)
SegviGen: Repurposing 3D Generative Model for Part Segmentation
von: Li, Lin, et al.
Veröffentlicht: (2026)
von: Li, Lin, et al.
Veröffentlicht: (2026)
Training Like a Medical Resident: Context-Prior Learning Toward Universal Medical Image Segmentation
von: Gao, Yunhe, et al.
Veröffentlicht: (2023)
von: Gao, Yunhe, et al.
Veröffentlicht: (2023)
Context-Guided Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2024)
von: Gu, Xin, et al.
Veröffentlicht: (2024)
Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling
von: Movva, Prahitha, et al.
Veröffentlicht: (2025)
von: Movva, Prahitha, et al.
Veröffentlicht: (2025)
Correspondence as Video: Test-Time Adaption on SAM2 for Reference Segmentation in the Wild
von: Wang, Haoran, et al.
Veröffentlicht: (2025)
von: Wang, Haoran, et al.
Veröffentlicht: (2025)
Robust Ego-Exo Correspondence with Long-Term Memory
von: Hu, Yijun, et al.
Veröffentlicht: (2025)
von: Hu, Yijun, et al.
Veröffentlicht: (2025)
Simplicity Prevails: The Emergence of Generalizable AIGI Detection in Visual Foundation Models
von: Zhou, Yue, et al.
Veröffentlicht: (2026)
von: Zhou, Yue, et al.
Veröffentlicht: (2026)
Continuous-Multiple Image Outpainting in One-Step via Positional Query and A Diffusion-based Approach
von: Zhang, Shaofeng, et al.
Veröffentlicht: (2024)
von: Zhang, Shaofeng, et al.
Veröffentlicht: (2024)
Mitigating Query Selection Bias in Referring Video Object Segmentation
von: Zhang, Dingwei, et al.
Veröffentlicht: (2025)
von: Zhang, Dingwei, et al.
Veröffentlicht: (2025)
Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens
von: Ma, Fan, et al.
Veröffentlicht: (2023)
von: Ma, Fan, et al.
Veröffentlicht: (2023)
DQFormer: Towards Unified LiDAR Panoptic Segmentation with Decoupled Queries
von: Yang, Yu, et al.
Veröffentlicht: (2024)
von: Yang, Yu, et al.
Veröffentlicht: (2024)
Aligning Human Knowledge with Visual Concepts Towards Explainable Medical Image Classification
von: Gao, Yunhe, et al.
Veröffentlicht: (2024)
von: Gao, Yunhe, et al.
Veröffentlicht: (2024)
Anatomy-Aware Text-Visual Fusion with Dual-Perspective Prompts for Fine-Grained Lumbar Spine Segmentation
von: Lian, Sheng, et al.
Veröffentlicht: (2025)
von: Lian, Sheng, et al.
Veröffentlicht: (2025)
DVIS-DAQ: Improving Video Segmentation via Dynamic Anchor Queries
von: Zhou, Yikang, et al.
Veröffentlicht: (2024)
von: Zhou, Yikang, et al.
Veröffentlicht: (2024)
Extending Adaptive Cruise Control with Machine Learning Intrusion Detection Systems
von: Othmane, Lotfi Ben, et al.
Veröffentlicht: (2026)
von: Othmane, Lotfi Ben, et al.
Veröffentlicht: (2026)
GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding
von: Fan, Rong, et al.
Veröffentlicht: (2026)
von: Fan, Rong, et al.
Veröffentlicht: (2026)
Towards Unified Video Quality Assessment
von: Feng, Chen, et al.
Veröffentlicht: (2025)
von: Feng, Chen, et al.
Veröffentlicht: (2025)
Towards Flexible Visual Relationship Segmentation
von: Zhu, Fangrui, et al.
Veröffentlicht: (2024)
von: Zhu, Fangrui, et al.
Veröffentlicht: (2024)
Query-guided Prototype Evolution Network for Few-Shot Segmentation
von: Cong, Runmin, et al.
Veröffentlicht: (2024)
von: Cong, Runmin, et al.
Veröffentlicht: (2024)
SynRES: Towards Referring Expression Segmentation in the Wild via Synthetic Data
von: Kim, Dong-Hee, et al.
Veröffentlicht: (2025)
von: Kim, Dong-Hee, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PRVQL: Progressive Knowledge-guided Refinement for Robust Egocentric Visual Query Localization
von: Fan, Bing, et al.
Veröffentlicht: (2025) -
LoReTrack: Efficient and Accurate Low-Resolution Transformer Tracking
von: Dong, Shaohua, et al.
Veröffentlicht: (2024) -
Towards Long-Form Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2026) -
VastTrack: Vast Category Visual Object Tracking
von: Peng, Liang, et al.
Veröffentlicht: (2024) -
Benchmarking the Robustness of UAV Tracking Against Common Corruptions
von: Liu, Xiaoqiong, et al.
Veröffentlicht: (2024)