Beyond Literal Descriptions: Understanding and Locating Open-World Objects Aligned with Human Intentions
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Wenxuan, Zhang, Yisi, He, Xingjian, Yan, Yichen, Zhao, Zijia, Wang, Xinlong, Liu, Jing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unveiling Parts Beyond Objects:Towards Finer-Granularity Referring Expression Segmentation
by: Wang, Wenxuan, et al.
Published: (2023)
by: Wang, Wenxuan, et al.
Published: (2023)
EAVL: Explicitly Align Vision and Language for Referring Image Segmentation
by: Yan, Yichen, et al.
Published: (2023)
by: Yan, Yichen, et al.
Published: (2023)
Image Difference Grounding with Natural Language
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
Towards Unified Referring Expression Segmentation Across Omni-Level Visual Target Granularities
by: Liu, Jing, et al.
Published: (2025)
by: Liu, Jing, et al.
Published: (2025)
CM-MaskSD: Cross-Modality Masked Self-Distillation for Referring Image Segmentation
by: Wang, Wenxuan, et al.
Published: (2023)
by: Wang, Wenxuan, et al.
Published: (2023)
The Instance-centric Transformer for the RVOS Track of LSVOS Challenge: 3rd Place Solution
by: Cao, Bin, et al.
Published: (2024)
by: Cao, Bin, et al.
Published: (2024)
OpenDlign: Open-World Point Cloud Understanding with Depth-Aligned Images
by: Mao, Ye, et al.
Published: (2024)
by: Mao, Ye, et al.
Published: (2024)
2nd Place Solution for MeViS Track in CVPR 2024 PVUW Workshop: Motion Expression guided Video Segmentation
by: Cao, Bin, et al.
Published: (2024)
by: Cao, Bin, et al.
Published: (2024)
Calibration & Reconstruction: Deep Integrated Language for Referring Image Segmentation
by: Yan, Yichen, et al.
Published: (2024)
by: Yan, Yichen, et al.
Published: (2024)
Fuse & Calibrate: A bi-directional Vision-Language Guided Framework for Referring Image Segmentation
by: Yan, Yichen, et al.
Published: (2024)
by: Yan, Yichen, et al.
Published: (2024)
Open 3D World in Autonomous Driving
by: Cheng, Xinlong, et al.
Published: (2024)
by: Cheng, Xinlong, et al.
Published: (2024)
Beyond Flat Unknown Labels in Open-World Object Detection
by: Zhang, Yuchen, et al.
Published: (2025)
by: Zhang, Yuchen, et al.
Published: (2025)
AI Sees Your Location, But With A Bias Toward The Wealthy World
by: Huang, Jingyuan, et al.
Published: (2025)
by: Huang, Jingyuan, et al.
Published: (2025)
Beyond the Individual: Introducing Group Intention Forecasting with SHOT Dataset
by: Zhang, Ruixu, et al.
Published: (2025)
by: Zhang, Ruixu, et al.
Published: (2025)
LLM-Guided Agentic Object Detection for Open-World Understanding
by: Mumcu, Furkan, et al.
Published: (2025)
by: Mumcu, Furkan, et al.
Published: (2025)
Diffusion Feedback Helps CLIP See Better
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
Open World Object Detection: A Survey
by: Li, Yiming, et al.
Published: (2024)
by: Li, Yiming, et al.
Published: (2024)
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
by: Zhu, Minjie, et al.
Published: (2025)
by: Zhu, Minjie, et al.
Published: (2025)
Eyes on Target: Gaze-Aware Object Detection in Egocentric Video
by: Lall, Vishakha, et al.
Published: (2025)
by: Lall, Vishakha, et al.
Published: (2025)
Towards 3D Objectness Learning in an Open World
by: Liu, Taichi, et al.
Published: (2025)
by: Liu, Taichi, et al.
Published: (2025)
Towards Efficient 3D Object Detection for Vehicle-Infrastructure Collaboration via Risk-Intent Selection
by: Wang, Li, et al.
Published: (2026)
by: Wang, Li, et al.
Published: (2026)
SC-Tune: Unleashing Self-Consistent Referential Comprehension in Large Vision Language Models
by: Yue, Tongtian, et al.
Published: (2024)
by: Yue, Tongtian, et al.
Published: (2024)
Intent3D: 3D Object Detection in RGB-D Scans Based on Human Intention
by: Kang, Weitai, et al.
Published: (2024)
by: Kang, Weitai, et al.
Published: (2024)
SPAN: Continuous Modeling of Suspicion Progression for Temporal Intention Localization
by: Hu, Xinyi, et al.
Published: (2025)
by: Hu, Xinyi, et al.
Published: (2025)
OpenAD: Open-World Autonomous Driving Benchmark for 3D Object Detection
by: Xia, Zhongyu, et al.
Published: (2024)
by: Xia, Zhongyu, et al.
Published: (2024)
WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors
by: Wu, Keming, et al.
Published: (2026)
by: Wu, Keming, et al.
Published: (2026)
Open-World Human-Object Interaction Detection via Multi-modal Prompts
by: Yang, Jie, et al.
Published: (2024)
by: Yang, Jie, et al.
Published: (2024)
DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation
by: Zhang, Wenchao, et al.
Published: (2025)
by: Zhang, Wenchao, et al.
Published: (2025)
Understanding the Implicit User Intention via Reasoning with Large Language Model for Image Editing
by: Wang, Yijia, et al.
Published: (2025)
by: Wang, Yijia, et al.
Published: (2025)
DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval
by: He, Xinwei, et al.
Published: (2026)
by: He, Xinwei, et al.
Published: (2026)
Transferable Physical-World Adversarial Patches Against Object Detection in Autonomous Driving
by: Zhu, Zihui, et al.
Published: (2026)
by: Zhu, Zihui, et al.
Published: (2026)
Open-Set Object Detection By Aligning Known Class Representations
by: Sarkar, Hiran, et al.
Published: (2024)
by: Sarkar, Hiran, et al.
Published: (2024)
Detailed Object Description with Controllable Dimensions
by: Wang, Xinran, et al.
Published: (2024)
by: Wang, Xinran, et al.
Published: (2024)
YOLO-World: Real-Time Open-Vocabulary Object Detection
by: Cheng, Tianheng, et al.
Published: (2024)
by: Cheng, Tianheng, et al.
Published: (2024)
Aligning Object Detector Bounding Boxes with Human Preference
by: Strafforello, Ombretta, et al.
Published: (2024)
by: Strafforello, Ombretta, et al.
Published: (2024)
HOID-R1: Reinforcement Learning for Open-World Human-Object Interaction Detection Reasoning with Multimodal Large Language Model
by: Zhang, Zhenhao, et al.
Published: (2025)
by: Zhang, Zhenhao, et al.
Published: (2025)
End-to-End Vision Tokenizer Tuning
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
IOR: Inversed Objects Replay for Incremental Object Detection
by: An, Zijia, et al.
Published: (2024)
by: An, Zijia, et al.
Published: (2024)
FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning
by: Hu, Shiyu, et al.
Published: (2024)
by: Hu, Shiyu, et al.
Published: (2024)
Similar Items
-
Unveiling Parts Beyond Objects:Towards Finer-Granularity Referring Expression Segmentation
by: Wang, Wenxuan, et al.
Published: (2023) -
EAVL: Explicitly Align Vision and Language for Referring Image Segmentation
by: Yan, Yichen, et al.
Published: (2023) -
Image Difference Grounding with Natural Language
by: Wang, Wenxuan, et al.
Published: (2025) -
Towards Unified Referring Expression Segmentation Across Omni-Level Visual Target Granularities
by: Liu, Jing, et al.
Published: (2025) -
CM-MaskSD: Cross-Modality Masked Self-Distillation for Referring Image Segmentation
by: Wang, Wenxuan, et al.
Published: (2023)