Cross-Attentive Multiview Fusion of Vision-Language Embeddings
Fuente:
arXiv
Saved in:
| Main Authors: | Martins, Tomas Berriel, Oswald, Martin R., Civera, Javier |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Open-Vocabulary Online Semantic Mapping for SLAM
by: Martins, Tomas Berriel, et al.
Published: (2024)
by: Martins, Tomas Berriel, et al.
Published: (2024)
Feature Splatting for Better Novel View Synthesis with Low Overlap
by: Martins, T. Berriel, et al.
Published: (2024)
by: Martins, T. Berriel, et al.
Published: (2024)
SLAM&Render: A Benchmark for the Intersection Between Neural Rendering, Gaussian Splatting and SLAM
by: Cerezo, Samuel, et al.
Published: (2025)
by: Cerezo, Samuel, et al.
Published: (2025)
Alignment Scores: Robust Metrics for Multiview Pose Accuracy Evaluation
by: Lee, Seong Hun, et al.
Published: (2024)
by: Lee, Seong Hun, et al.
Published: (2024)
GNSS-Inertial State Initialization Using Inter-Epoch Baseline Residuals
by: Cerezo, Samuel, et al.
Published: (2025)
by: Cerezo, Samuel, et al.
Published: (2025)
Close, But Not There: Boosting Geographic Distance Sensitivity in Visual Place Recognition
by: Izquierdo, Sergio, et al.
Published: (2024)
by: Izquierdo, Sergio, et al.
Published: (2024)
Optimal Transport Aggregation for Visual Place Recognition
by: Izquierdo, Sergio, et al.
Published: (2023)
by: Izquierdo, Sergio, et al.
Published: (2023)
Camera Motion Estimation from RGB-D-Inertial Scene Flow
by: Cerezo, Samuel, et al.
Published: (2024)
by: Cerezo, Samuel, et al.
Published: (2024)
AnyCalib: On-Manifold Learning for Model-Agnostic Single-View Camera Calibration
by: Tirado-Garín, Javier, et al.
Published: (2025)
by: Tirado-Garín, Javier, et al.
Published: (2025)
From Correspondences to Pose: Non-minimal Certifiably Optimal Relative Pose without Disambiguation
by: Tirado-Garín, Javier, et al.
Published: (2023)
by: Tirado-Garín, Javier, et al.
Published: (2023)
Multiview Manifold Evidential Fusion for PolSAR Image Classification
by: Shi, Junfei, et al.
Published: (2025)
by: Shi, Junfei, et al.
Published: (2025)
Facial Emotion Learning with Text-Guided Multiview Fusion via Vision-Language Model for 3D/4D Facial Expression Recognition
by: Behzad, Muzammil
Published: (2025)
by: Behzad, Muzammil
Published: (2025)
Deformable Attentive Visual Enhancement for Referring Segmentation Using Vision-Language Model
by: Dalaq, Alaa, et al.
Published: (2025)
by: Dalaq, Alaa, et al.
Published: (2025)
ADEM-VL: Adaptive and Embedded Fusion for Efficient Vision-Language Tuning
by: Hao, Zhiwei, et al.
Published: (2024)
by: Hao, Zhiwei, et al.
Published: (2024)
Multi-Modal Building Inspection via Perceiver IO Fusion of Satellite and Street-Level Imagery
by: Sombekke, Niels, et al.
Published: (2026)
by: Sombekke, Niels, et al.
Published: (2026)
DefVINS: Visual-Inertial Odometry for Deformable Scenes
by: Cerezo, Samuel, et al.
Published: (2026)
by: Cerezo, Samuel, et al.
Published: (2026)
An Interpretable Cross-Attentive Multi-modal MRI Fusion Framework for Schizophrenia Diagnosis
by: Zhou, Ziyu, et al.
Published: (2024)
by: Zhou, Ziyu, et al.
Published: (2024)
Multimodal and Multiview Deep Fusion for Autonomous Marine Navigation
by: Dagdilelis, Dimitrios, et al.
Published: (2025)
by: Dagdilelis, Dimitrios, et al.
Published: (2025)
Learning Neural Implicit through Volume Rendering with Attentive Depth Fusion Priors
by: Hu, Pengchong, et al.
Published: (2023)
by: Hu, Pengchong, et al.
Published: (2023)
Self-supervised 3D Patient Modeling with Multi-modal Attentive Fusion
by: Zheng, Meng, et al.
Published: (2024)
by: Zheng, Meng, et al.
Published: (2024)
Audio Deepfake Detection with Half-Truth Localisation Using Cross-Attentive Feature Fusion
by: Sutharya, S., et al.
Published: (2026)
by: Sutharya, S., et al.
Published: (2026)
Robust Single Rotation Averaging Revisited
by: Lee, Seong Hun, et al.
Published: (2023)
by: Lee, Seong Hun, et al.
Published: (2023)
What's Wrong with the Absolute Trajectory Error?
by: Lee, Seong Hun, et al.
Published: (2022)
by: Lee, Seong Hun, et al.
Published: (2022)
Cross-Layer Attentive Feature Upsampling for Low-latency Semantic Segmentation
by: Cheng, Tianheng, et al.
Published: (2026)
by: Cheng, Tianheng, et al.
Published: (2026)
CAST: Cross-Attentive Spatio-Temporal feature fusion for deepfake detection
by: Thakre, Aryan, et al.
Published: (2025)
by: Thakre, Aryan, et al.
Published: (2025)
FusionCell: Cross-Attentive Fusion of Layout Geometry and Netlist Topology for Standard-Cell Performance Prediction
by: Zhang, Haoyi, et al.
Published: (2026)
by: Zhang, Haoyi, et al.
Published: (2026)
Multiview Scene Graph
by: Zhang, Juexiao, et al.
Published: (2024)
by: Zhang, Juexiao, et al.
Published: (2024)
ACE-LoRA: Graph-Attentive Context Enhancement for Parameter-Efficient Adaptation of Medical Vision-Language Models
by: Aydın, M. Arda, et al.
Published: (2026)
by: Aydın, M. Arda, et al.
Published: (2026)
Fourier-Attentive Representation Learning: A Fourier-Guided Framework for Few-Shot Generalization in Vision-Language Models
by: Pham, Hieu Dinh Trung, et al.
Published: (2025)
by: Pham, Hieu Dinh Trung, et al.
Published: (2025)
Accurate and Scalable Multimodal Pathology Retrieval via Attentive Vision-Language Alignment
by: Wang, Hongyi, et al.
Published: (2025)
by: Wang, Hongyi, et al.
Published: (2025)
Cross-Modal Redundancy and the Geometry of Vision-Language Embeddings
by: Dhimoïla, Grégoire, et al.
Published: (2026)
by: Dhimoïla, Grégoire, et al.
Published: (2026)
Non-Minimal Sampling and Consensus for Prohibitively Large Datasets
by: Lee, Seong Hun, et al.
Published: (2026)
by: Lee, Seong Hun, et al.
Published: (2026)
P3P Made Easy
by: Lee, Seong Hun, et al.
Published: (2025)
by: Lee, Seong Hun, et al.
Published: (2025)
Ego-1K -- A Large-Scale Multiview Video Dataset for Egocentric Vision
by: Lee, Jae Yong, et al.
Published: (2026)
by: Lee, Jae Yong, et al.
Published: (2026)
Addressing the challenges of loop detection in agricultural environments
by: Soncini, Nicolás, et al.
Published: (2024)
by: Soncini, Nicolás, et al.
Published: (2024)
Multiview Image-Based Localization
by: Fiore, Cameron, et al.
Published: (2025)
by: Fiore, Cameron, et al.
Published: (2025)
BcQLM: Efficient Vision-Language Understanding with Distilled Q-Gated Cross-Modal Fusion
by: Xiang, Sike, et al.
Published: (2025)
by: Xiang, Sike, et al.
Published: (2025)
MVD$^2$: Efficient Multiview 3D Reconstruction for Multiview Diffusion
by: Zheng, Xin-Yang, et al.
Published: (2024)
by: Zheng, Xin-Yang, et al.
Published: (2024)
VSLAM-LAB: A Comprehensive Framework for Visual SLAM Methods and Datasets
by: Fontan, Alejandro, et al.
Published: (2025)
by: Fontan, Alejandro, et al.
Published: (2025)
MAG-VLAQ: Multi-modal Aerial-Ground Query Aggregation for Cross-View Place Recognition
by: Xu, Zhengyi, et al.
Published: (2026)
by: Xu, Zhengyi, et al.
Published: (2026)
Similar Items
-
Open-Vocabulary Online Semantic Mapping for SLAM
by: Martins, Tomas Berriel, et al.
Published: (2024) -
Feature Splatting for Better Novel View Synthesis with Low Overlap
by: Martins, T. Berriel, et al.
Published: (2024) -
SLAM&Render: A Benchmark for the Intersection Between Neural Rendering, Gaussian Splatting and SLAM
by: Cerezo, Samuel, et al.
Published: (2025) -
Alignment Scores: Robust Metrics for Multiview Pose Accuracy Evaluation
by: Lee, Seong Hun, et al.
Published: (2024) -
GNSS-Inertial State Initialization Using Inter-Epoch Baseline Residuals
by: Cerezo, Samuel, et al.
Published: (2025)