Towards Cross-View Point Correspondence in Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Yipu, Ji, Yuheng, Liu, Yuyang, Zhou, Enshen, Yang, Ziqiang, Tian, Yuxuan, Qin, Ziheng, Liu, Yue, Tan, Huajie, Chi, Cheng, Ma, Zhiyuan, Zeng, Daniel Dajun, Zheng, Xiaolong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes
di: Qin, Ziheng, et al.
Pubblicazione: (2025)
di: Qin, Ziheng, et al.
Pubblicazione: (2025)
VisualTrans: A Benchmark for Real-World Visual Transformation Reasoning
di: Ji, Yuheng, et al.
Pubblicazione: (2025)
di: Ji, Yuheng, et al.
Pubblicazione: (2025)
Alleviating Performance Disparity in Adversarial Spatiotemporal Graph Learning Under Zero-Inflated Distribution
di: Bai, Songran, et al.
Pubblicazione: (2025)
di: Bai, Songran, et al.
Pubblicazione: (2025)
MathSticks: A Benchmark for Visual Symbolic Compositional Reasoning with Matchstick Puzzles
di: Ji, Yuheng, et al.
Pubblicazione: (2025)
di: Ji, Yuheng, et al.
Pubblicazione: (2025)
PRM-as-a-Judge: A Dense Evaluation Paradigm for Fine-Grained Robotic Auditing
di: Ji, Yuheng, et al.
Pubblicazione: (2026)
di: Ji, Yuheng, et al.
Pubblicazione: (2026)
Deep Causal Learning: Representation, Discovery and Inference
di: Deng, Zizhen, et al.
Pubblicazione: (2022)
di: Deng, Zizhen, et al.
Pubblicazione: (2022)
Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language Models
di: Tan, Huajie, et al.
Pubblicazione: (2025)
di: Tan, Huajie, et al.
Pubblicazione: (2025)
CRC-SGAD: Conformal Risk Control for Supervised Graph Anomaly Detection
di: Bai, Songran, et al.
Pubblicazione: (2025)
di: Bai, Songran, et al.
Pubblicazione: (2025)
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
di: Liu, Mengzhen, et al.
Pubblicazione: (2026)
di: Liu, Mengzhen, et al.
Pubblicazione: (2026)
RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics
di: Zhou, Enshen, et al.
Pubblicazione: (2025)
di: Zhou, Enshen, et al.
Pubblicazione: (2025)
Projectively Wakamatsu Tilting Modules over One-Point Extensions
di: Liu, Dajun, et al.
Pubblicazione: (2026)
di: Liu, Dajun, et al.
Pubblicazione: (2026)
Learning Instance-Aware Correspondences for Robust Multi-Instance Point Cloud Registration in Cluttered Scenes
di: Yu, Zhiyuan, et al.
Pubblicazione: (2024)
di: Yu, Zhiyuan, et al.
Pubblicazione: (2024)
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
di: Zhou, Enshen, et al.
Pubblicazione: (2025)
di: Zhou, Enshen, et al.
Pubblicazione: (2025)
Enhancing Adversarial Robustness of Vision-Language Models through Low-Rank Adaptation
di: Ji, Yuheng, et al.
Pubblicazione: (2024)
di: Ji, Yuheng, et al.
Pubblicazione: (2024)
CASA: Class-Agnostic Shared Attributes in Vision-Language Models for Efficient Incremental Object Detection
di: Guo, Mingyi, et al.
Pubblicazione: (2024)
di: Guo, Mingyi, et al.
Pubblicazione: (2024)
Front‐Footed Defense: Leveraging Early Counsel Intervention for Expedited Justice
di: Chengchen He, et al.
Pubblicazione: (2026)
di: Chengchen He, et al.
Pubblicazione: (2026)
MambaSOD: Dual Mamba-Driven Cross-Modal Fusion Network for RGB-D Salient Object Detection
di: Zhan, Yue, et al.
Pubblicazione: (2024)
di: Zhan, Yue, et al.
Pubblicazione: (2024)
Cross-View Completion Models are Zero-shot Correspondence Estimators
di: An, Honggyu, et al.
Pubblicazione: (2024)
di: An, Honggyu, et al.
Pubblicazione: (2024)
C3Po: Cross-View Cross-Modality Correspondence by Pointmap Prediction
di: Huang, Kuan Wei, et al.
Pubblicazione: (2025)
di: Huang, Kuan Wei, et al.
Pubblicazione: (2025)
DSNet: A Computer Vision‐Based Detection and Corrosion Segmentation Network for Corroded Bolt Detection in Tunnel
di: Lei Tan, et al.
Pubblicazione: (2024)
di: Lei Tan, et al.
Pubblicazione: (2024)
CorrespondentDream: Enhancing 3D Fidelity of Text-to-3D using Cross-View Correspondences
di: Kim, Seungwook, et al.
Pubblicazione: (2024)
di: Kim, Seungwook, et al.
Pubblicazione: (2024)
RKHS-BA: A Robust Correspondence-Free Multi-View Registration Framework with Semantic Point Clouds
di: Zhang, Ray, et al.
Pubblicazione: (2024)
di: Zhang, Ray, et al.
Pubblicazione: (2024)
Pic@Point: Cross-Modal Learning by Local and Global Point-Picture Correspondence
di: Herzog, Vencia, et al.
Pubblicazione: (2024)
di: Herzog, Vencia, et al.
Pubblicazione: (2024)
CLNet: Cross-View Correspondence Makes a Stronger Geo-Localizationer
di: Cao, Xianwei, et al.
Pubblicazione: (2025)
di: Cao, Xianwei, et al.
Pubblicazione: (2025)
Unveiling Factual Recall Behaviors of Large Language Models through Knowledge Neurons
di: Wang, Yifei, et al.
Pubblicazione: (2024)
di: Wang, Yifei, et al.
Pubblicazione: (2024)
CBV: Clean-label Backdoor Attacks on Vision Language Models via Diffusion Models
di: Guo, Ji, et al.
Pubblicazione: (2026)
di: Guo, Ji, et al.
Pubblicazione: (2026)
UniABG: Unified Adversarial View Bridging and Graph Correspondence for Unsupervised Cross-View Geo-Localization
di: Chen, Cuiqun, et al.
Pubblicazione: (2025)
di: Chen, Cuiqun, et al.
Pubblicazione: (2025)
CrossVL: Complexity-Aware Feature Routing and Paired Curriculum for Cross-View Vision-Language Detection
di: Liu, Zhipeng, et al.
Pubblicazione: (2026)
di: Liu, Zhipeng, et al.
Pubblicazione: (2026)
Cascaded Dual Vision Transformer for Accurate Facial Landmark Detection
di: Dang, Ziqiang, et al.
Pubblicazione: (2024)
di: Dang, Ziqiang, et al.
Pubblicazione: (2024)
Learning Cross-View Object Correspondence via Cycle-Consistent Mask Prediction
di: Yan, Shannan, et al.
Pubblicazione: (2026)
di: Yan, Shannan, et al.
Pubblicazione: (2026)
TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
di: Han, Yi, et al.
Pubblicazione: (2025)
di: Han, Yi, et al.
Pubblicazione: (2025)
COM3D: Leveraging Cross-View Correspondence and Cross-Modal Mining for 3D Retrieval
di: Wu, Hao, et al.
Pubblicazione: (2024)
di: Wu, Hao, et al.
Pubblicazione: (2024)
"What I Sign Is Not What I See": Towards Explainable and Trustworthy Cryptocurrency Wallet Signatures
di: Qin, Yuyang, et al.
Pubblicazione: (2026)
di: Qin, Yuyang, et al.
Pubblicazione: (2026)
Towards Class-agnostic Tracking Using Feature Decorrelation in Point Clouds
di: Tian, Shengjing, et al.
Pubblicazione: (2022)
di: Tian, Shengjing, et al.
Pubblicazione: (2022)
RoCo Challenge at AAAI 2026: Benchmarking Robotic Collaborative Manipulation for Assembly Towards Industrial Automation
di: Liu, Haichao, et al.
Pubblicazione: (2026)
di: Liu, Haichao, et al.
Pubblicazione: (2026)
CoE: Deep Coupled Embedding for Non-Rigid Point Cloud Correspondences
di: Zeng, Huajian, et al.
Pubblicazione: (2024)
di: Zeng, Huajian, et al.
Pubblicazione: (2024)
Consistency of tug-of-war type operators on random data clouds
di: Han, Jeongmin, et al.
Pubblicazione: (2025)
di: Han, Jeongmin, et al.
Pubblicazione: (2025)
IDMR: Towards Instance-Driven Precise Visual Correspondence in Multimodal Retrieval
di: Liu, Bangwei, et al.
Pubblicazione: (2025)
di: Liu, Bangwei, et al.
Pubblicazione: (2025)
Scalable Vision-Based 3D Object Detection and Monocular Depth Estimation for Autonomous Driving
di: Liu, Yuxuan
Pubblicazione: (2024)
di: Liu, Yuxuan
Pubblicazione: (2024)
Efficient Circuit-Based Quantum State Tomography via Sparse Entry Optimization
di: Li, Chi-Kwong, et al.
Pubblicazione: (2024)
di: Li, Chi-Kwong, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes
di: Qin, Ziheng, et al.
Pubblicazione: (2025) -
VisualTrans: A Benchmark for Real-World Visual Transformation Reasoning
di: Ji, Yuheng, et al.
Pubblicazione: (2025) -
Alleviating Performance Disparity in Adversarial Spatiotemporal Graph Learning Under Zero-Inflated Distribution
di: Bai, Songran, et al.
Pubblicazione: (2025) -
MathSticks: A Benchmark for Visual Symbolic Compositional Reasoning with Matchstick Puzzles
di: Ji, Yuheng, et al.
Pubblicazione: (2025) -
PRM-as-a-Judge: A Dense Evaluation Paradigm for Fine-Grained Robotic Auditing
di: Ji, Yuheng, et al.
Pubblicazione: (2026)