Beyond Object Categories: Multi-Attribute Reference Understanding for Visual Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Hao, Zhu, Jianfei, Fan, Wei, Yi, Chunzhi, Jiang, Feng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AD-DINO: Attention-Dynamic DINO for Distance-Aware Embodied Reference Understanding
by: Guo, Hao, et al.
Published: (2024)
by: Guo, Hao, et al.
Published: (2024)
OSCAR: Object Status and Contextual Awareness for Recipes to Support Non-Visual Cooking
by: Li, Franklin Mingzhe, et al.
Published: (2025)
by: Li, Franklin Mingzhe, et al.
Published: (2025)
L-WISE: Boosting Human Visual Category Learning Through Model-Based Image Selection and Enhancement
by: Talbot, Morgan B., et al.
Published: (2024)
by: Talbot, Morgan B., et al.
Published: (2024)
PoseDriver: A Unified Approach to Multi-Category Skeleton Detection for Autonomous Driving
by: Borhani, Yasamin, et al.
Published: (2026)
by: Borhani, Yasamin, et al.
Published: (2026)
AttributionScanner: A Visual Analytics System for Model Validation with Metadata-Free Slice Finding
by: Xuan, Xiwei, et al.
Published: (2024)
by: Xuan, Xiwei, et al.
Published: (2024)
Visually Grounded Narratives: Reducing Cognitive Burden in Researcher-Participant Interaction
by: Wu, Runtong, et al.
Published: (2025)
by: Wu, Runtong, et al.
Published: (2025)
ChatStitch: Visualizing Through Structures via Surround-View Unsupervised Deep Image Stitching with Collaborative LLM-Agents
by: Liang, Hao, et al.
Published: (2025)
by: Liang, Hao, et al.
Published: (2025)
OW-CLIP: Data-Efficient Visual Supervision for Open-World Object Detection via Human-AI Collaboration
by: Duan, Junwen, et al.
Published: (2025)
by: Duan, Junwen, et al.
Published: (2025)
Panda or not Panda? Understanding Adversarial Attacks with Interactive Visualization
by: You, Yuzhe, et al.
Published: (2023)
by: You, Yuzhe, et al.
Published: (2023)
Referring Human Pose and Mask Estimation in the Wild
by: Miao, Bo, et al.
Published: (2024)
by: Miao, Bo, et al.
Published: (2024)
SimVecVis: A Dataset for Enhancing MLLMs in Visualization Understanding
by: Liu, Can, et al.
Published: (2025)
by: Liu, Can, et al.
Published: (2025)
A Multimodal Depth-Aware Method For Embodied Reference Understanding
by: Eyiokur, Fevziye Irem, et al.
Published: (2025)
by: Eyiokur, Fevziye Irem, et al.
Published: (2025)
Designing Multi-Robot Ground Video Sensemaking with Public Safety Professionals
by: Zhou, Puqi, et al.
Published: (2026)
by: Zhou, Puqi, et al.
Published: (2026)
ObjectGS: Object-aware Scene Reconstruction and Scene Understanding via Gaussian Splatting
by: Zhu, Ruijie, et al.
Published: (2025)
by: Zhu, Ruijie, et al.
Published: (2025)
ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People
by: Liu, Ruiping, et al.
Published: (2024)
by: Liu, Ruiping, et al.
Published: (2024)
CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
by: Verma, Arnav, et al.
Published: (2025)
by: Verma, Arnav, et al.
Published: (2025)
Talk to Parallel LiDARs: A Human-LiDAR Interaction Method Based on 3D Visual Grounding
by: Liu, Yuhang, et al.
Published: (2024)
by: Liu, Yuhang, et al.
Published: (2024)
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
by: You, Keen, et al.
Published: (2024)
by: You, Keen, et al.
Published: (2024)
CoEditor++: Instruction-based Visual Editing via Cognitive Reasoning
by: Ni, Minheng, et al.
Published: (2026)
by: Ni, Minheng, et al.
Published: (2026)
Category-aware EEG image generation based on wavelet transform and contrast semantic loss
by: Zhang, Enshang, et al.
Published: (2025)
by: Zhang, Enshang, et al.
Published: (2025)
Deep Learning-based Lightweight RGB Object Tracking for Augmented Reality Devices
by: Smith, Alice, et al.
Published: (2025)
by: Smith, Alice, et al.
Published: (2025)
SpriteHand: Real-Time Versatile Hand-Object Interaction with Autoregressive Video Generation
by: Li, Zisu, et al.
Published: (2025)
by: Li, Zisu, et al.
Published: (2025)
Exploring Object Status Recognition for Recipe Progress Tracking in Non-Visual Cooking
by: Li, Franklin Mingzhe, et al.
Published: (2025)
by: Li, Franklin Mingzhe, et al.
Published: (2025)
Face-LLaVA: Facial Expression and Attribute Understanding through Instruction Tuning
by: Chaubey, Ashutosh, et al.
Published: (2025)
by: Chaubey, Ashutosh, et al.
Published: (2025)
Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
by: Fei, Hao, et al.
Published: (2024)
by: Fei, Hao, et al.
Published: (2024)
VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation
by: Pan, Bo, et al.
Published: (2025)
by: Pan, Bo, et al.
Published: (2025)
Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision
by: Li, Chentao, et al.
Published: (2026)
by: Li, Chentao, et al.
Published: (2026)
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
by: Wang, Siting, et al.
Published: (2025)
by: Wang, Siting, et al.
Published: (2025)
GenColor: Generative Color-Concept Association in Visual Design
by: Hou, Yihan, et al.
Published: (2025)
by: Hou, Yihan, et al.
Published: (2025)
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
by: Sun, Boyuan, et al.
Published: (2026)
by: Sun, Boyuan, et al.
Published: (2026)
VisionGPT: LLM-Assisted Real-Time Anomaly Detection for Safe Visual Navigation
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
ConceptFactory: Facilitate 3D Object Knowledge Annotation with Object Conceptualization
by: Sun, Jianhua, et al.
Published: (2024)
by: Sun, Jianhua, et al.
Published: (2024)
GroundUp: Rapid Sketch-Based 3D City Massing
by: Unlu, Gizem Esra, et al.
Published: (2024)
by: Unlu, Gizem Esra, et al.
Published: (2024)
Visual Neural Decoding via Improved Visual-EEG Semantic Consistency
by: Chen, Hongzhou, et al.
Published: (2024)
by: Chen, Hongzhou, et al.
Published: (2024)
Efficient 3D Reconstruction, Streaming and Visualization of Static and Dynamic Scene Parts for Multi-client Live-telepresence in Large-scale Environments
by: Van Holland, Leif, et al.
Published: (2022)
by: Van Holland, Leif, et al.
Published: (2022)
AccessLens: Auto-detecting Inaccessibility of Everyday Objects
by: Kwon, Nahyun, et al.
Published: (2024)
by: Kwon, Nahyun, et al.
Published: (2024)
MM-SCALE: Grounded Multimodal Moral Reasoning via Scalar Judgment and Listwise Alignment
by: Park, Eunkyu, et al.
Published: (2026)
by: Park, Eunkyu, et al.
Published: (2026)
Computer Vision for Objects used in Group Work: Challenges and Opportunities
by: Jung, Changsoo, et al.
Published: (2025)
by: Jung, Changsoo, et al.
Published: (2025)
Tell Me Without Telling Me: Two-Way Prediction of Visualization Literacy and Visual Attention
by: Chang, Minsuk, et al.
Published: (2025)
by: Chang, Minsuk, et al.
Published: (2025)
MAGE: A Multi-task Architecture for Gaze Estimation with an Efficient Calibration Module
by: Huang, Haoming, et al.
Published: (2025)
by: Huang, Haoming, et al.
Published: (2025)
Similar Items
-
AD-DINO: Attention-Dynamic DINO for Distance-Aware Embodied Reference Understanding
by: Guo, Hao, et al.
Published: (2024) -
OSCAR: Object Status and Contextual Awareness for Recipes to Support Non-Visual Cooking
by: Li, Franklin Mingzhe, et al.
Published: (2025) -
L-WISE: Boosting Human Visual Category Learning Through Model-Based Image Selection and Enhancement
by: Talbot, Morgan B., et al.
Published: (2024) -
PoseDriver: A Unified Approach to Multi-Category Skeleton Detection for Autonomous Driving
by: Borhani, Yasamin, et al.
Published: (2026) -
AttributionScanner: A Visual Analytics System for Model Validation with Metadata-Free Slice Finding
by: Xuan, Xiwei, et al.
Published: (2024)