Binding Visual Features Point by Point
Fuente:
arXiv
Saved in:
| Main Authors: | Haputhanthri, Udith, Campbell, Declan, Assouel, Rim, Cohen, Jonathan D., Webb, Taylor W. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visual symbolic mechanisms: Emergent symbol processing in vision language models
by: Assouel, Rim, et al.
Published: (2025)
by: Assouel, Rim, et al.
Published: (2025)
Object-centric Binding in Contrastive Language-Image Pretraining
by: Assouel, Rim, et al.
Published: (2025)
by: Assouel, Rim, et al.
Published: (2025)
PGT: Procedurally Generated Tasks for improving visual grounding in MLLMs
by: Assouel, Rim, et al.
Published: (2026)
by: Assouel, Rim, et al.
Published: (2026)
Understanding the Limits of Vision Language Models Through the Lens of the Binding Problem
by: Campbell, Declan, et al.
Published: (2024)
by: Campbell, Declan, et al.
Published: (2024)
The Geometry of Representational Failures in Vision Language Models
by: Savietto, Daniele, et al.
Published: (2026)
by: Savietto, Daniele, et al.
Published: (2026)
Mining and Transferring Feature-Geometry Coherence for Unsupervised Point Cloud Registration
by: Xiong, Kezheng, et al.
Published: (2024)
by: Xiong, Kezheng, et al.
Published: (2024)
Poivre: Self-Refining Visual Pointing with Reinforcement Learning
by: Yang, Wenjie, et al.
Published: (2025)
by: Yang, Wenjie, et al.
Published: (2025)
Complementary Pseudo Multimodal Feature for Point Cloud Anomaly Detection
by: Cao, Yunkang, et al.
Published: (2023)
by: Cao, Yunkang, et al.
Published: (2023)
Point2RBox: Combine Knowledge from Synthetic Visual Patterns for End-to-end Oriented Object Detection with Single Point Supervision
by: Yu, Yi, et al.
Published: (2023)
by: Yu, Yi, et al.
Published: (2023)
All in One: Visual-Description-Guided Unified Point Cloud Segmentation
by: Han, Zongyan, et al.
Published: (2025)
by: Han, Zongyan, et al.
Published: (2025)
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
by: Clark, Christopher, et al.
Published: (2026)
by: Clark, Christopher, et al.
Published: (2026)
Point-DAE: Denoising Autoencoders for Self-supervised Point Cloud Learning
by: Zhang, Yabin, et al.
Published: (2022)
by: Zhang, Yabin, et al.
Published: (2022)
Point-SAM: Promptable 3D Segmentation Model for Point Clouds
by: Zhou, Yuchen, et al.
Published: (2024)
by: Zhou, Yuchen, et al.
Published: (2024)
FPQVAR: Floating Point Quantization for Visual Autoregressive Model with FPGA Hardware Co-design
by: Wei, Renjie, et al.
Published: (2025)
by: Wei, Renjie, et al.
Published: (2025)
Triple Point Masking
by: Liu, Jiaming, et al.
Published: (2024)
by: Liu, Jiaming, et al.
Published: (2024)
PointSplit: Towards On-device 3D Object Detection with Heterogeneous Low-power Accelerators
by: Park, Keondo, et al.
Published: (2025)
by: Park, Keondo, et al.
Published: (2025)
Pic@Point: Cross-Modal Learning by Local and Global Point-Picture Correspondence
by: Herzog, Vencia, et al.
Published: (2024)
by: Herzog, Vencia, et al.
Published: (2024)
Points-to-3D: Structure-Aware 3D Generation with Point Cloud Priors
by: Xia, Jiatong, et al.
Published: (2026)
by: Xia, Jiatong, et al.
Published: (2026)
PointCG: Self-supervised Point Cloud Learning via Joint Completion and Generation
by: Liu, Yun, et al.
Published: (2024)
by: Liu, Yun, et al.
Published: (2024)
ResAgent: Entropy-based Prior Point Discovery and Visual Reasoning for Referring Expression Segmentation
by: Wang, Yihao, et al.
Published: (2026)
by: Wang, Yihao, et al.
Published: (2026)
Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Multi-Stage Visual Grounding
by: Xue, Haotian, et al.
Published: (2025)
by: Xue, Haotian, et al.
Published: (2025)
MPVO: Motion-Prior based Visual Odometry for PointGoal Navigation
by: Paul, Sayan, et al.
Published: (2024)
by: Paul, Sayan, et al.
Published: (2024)
DiffPoint: Single and Multi-view Point Cloud Reconstruction with ViT Based Diffusion Model
by: Feng, Yu, et al.
Published: (2024)
by: Feng, Yu, et al.
Published: (2024)
Point2Quad: Generating Quad Meshes from Point Clouds via Face Prediction
by: Li, Zezeng, et al.
Published: (2025)
by: Li, Zezeng, et al.
Published: (2025)
PointOBB-v3: Expanding Performance Boundaries of Single Point-Supervised Oriented Object Detection
by: Zhang, Peiyuan, et al.
Published: (2025)
by: Zhang, Peiyuan, et al.
Published: (2025)
SpatialPoint: Spatial-aware Point Prediction for Embodied Localization
by: Zhu, Qiming, et al.
Published: (2026)
by: Zhu, Qiming, et al.
Published: (2026)
Rectified Point Flow: Generic Point Cloud Pose Estimation
by: Sun, Tao, et al.
Published: (2025)
by: Sun, Tao, et al.
Published: (2025)
Single-Frame Point-Pixel Registration via Supervised Cross-Modal Feature Matching
by: Han, Yu, et al.
Published: (2025)
by: Han, Yu, et al.
Published: (2025)
Slot Abstractors: Toward Scalable Abstract Visual Reasoning
by: Mondal, Shanka Subhra, et al.
Published: (2024)
by: Mondal, Shanka Subhra, et al.
Published: (2024)
PointGS: Semantic-Consistent Unsupervised 3D Point Cloud Segmentation with 3D Gaussian Splatting
by: Song, Yixiao, et al.
Published: (2026)
by: Song, Yixiao, et al.
Published: (2026)
PointOBB-v2: Towards Simpler, Faster, and Stronger Single Point Supervised Oriented Object Detection
by: Ren, Botao, et al.
Published: (2024)
by: Ren, Botao, et al.
Published: (2024)
PointLLM: Empowering Large Language Models to Understand Point Clouds
by: Xu, Runsen, et al.
Published: (2023)
by: Xu, Runsen, et al.
Published: (2023)
Caption This, Reason That: VLMs Caught in the Middle
by: Weng, Zihan, et al.
Published: (2025)
by: Weng, Zihan, et al.
Published: (2025)
Point2RBox-v2: Rethinking Point-supervised Oriented Object Detection with Spatial Layout Among Instances
by: Yu, Yi, et al.
Published: (2025)
by: Yu, Yi, et al.
Published: (2025)
Stake the Points: Structure-Faithful Instance Unlearning
by: Hong, Kiseong, et al.
Published: (2026)
by: Hong, Kiseong, et al.
Published: (2026)
Equivariant Flow Matching for Point Cloud Assembly
by: Wang, Ziming, et al.
Published: (2025)
by: Wang, Ziming, et al.
Published: (2025)
FBPT: A Fully Binary Point Transformer
by: Hou, Zhixing, et al.
Published: (2024)
by: Hou, Zhixing, et al.
Published: (2024)
Point2RBox-v3: Self-Bootstrapping from Point Annotations via Integrated Pseudo-Label Refinement and Utilization
by: Zhang, Teng, et al.
Published: (2025)
by: Zhang, Teng, et al.
Published: (2025)
The Point, the Vision and the Text: Does Point Cloud Boost Spatial Reasoning of Large Language Models? A Bias-Controlled Study
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
Point-DETR3D: Leveraging Imagery Data with Spatial Point Prior for Weakly Semi-supervised 3D Object Detection
by: Gao, Hongzhi, et al.
Published: (2024)
by: Gao, Hongzhi, et al.
Published: (2024)
Similar Items
-
Visual symbolic mechanisms: Emergent symbol processing in vision language models
by: Assouel, Rim, et al.
Published: (2025) -
Object-centric Binding in Contrastive Language-Image Pretraining
by: Assouel, Rim, et al.
Published: (2025) -
PGT: Procedurally Generated Tasks for improving visual grounding in MLLMs
by: Assouel, Rim, et al.
Published: (2026) -
Understanding the Limits of Vision Language Models Through the Lens of the Binding Problem
by: Campbell, Declan, et al.
Published: (2024) -
The Geometry of Representational Failures in Vision Language Models
by: Savietto, Daniele, et al.
Published: (2026)