FusionBERT: Multi-View Image-3D Retrieval via Cross-Attention Visual Fusion and Normal-Aware 3D Encoder
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Wei, Ren, Yufan, Jiang, Hanqing, Ding, Jianhui, Peng, Zhen, Feng, Leman, Shentu, Yichun, Xu, Guoqiang, Sun, Baigui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CVFusion: Cross-View Fusion of 4D Radar and Camera for 3D Object Detection
von: Zhong, Hanzhi, et al.
Veröffentlicht: (2025)
von: Zhong, Hanzhi, et al.
Veröffentlicht: (2025)
CorrespondentDream: Enhancing 3D Fidelity of Text-to-3D using Cross-View Correspondences
von: Kim, Seungwook, et al.
Veröffentlicht: (2024)
von: Kim, Seungwook, et al.
Veröffentlicht: (2024)
Wavelet-based Multi-View Fusion of 4D Radar Tensor and Camera for Robust 3D Object Detection
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
MV-SAM3D: Adaptive Multi-View Fusion for Layout-Aware 3D Generation
von: Li, Baicheng, et al.
Veröffentlicht: (2026)
von: Li, Baicheng, et al.
Veröffentlicht: (2026)
3D Crowd Counting via Geometric Attention-guided Multi-View Fusion
von: Zhang, Qi, et al.
Veröffentlicht: (2020)
von: Zhang, Qi, et al.
Veröffentlicht: (2020)
EditSplat: Multi-View Fusion and Attention-Guided Optimization for View-Consistent 3D Scene Editing with 3D Gaussian Splatting
von: Lee, Dong In, et al.
Veröffentlicht: (2024)
von: Lee, Dong In, et al.
Veröffentlicht: (2024)
WildFusion: Learning 3D-Aware Latent Diffusion Models in View Space
von: Schwarz, Katja, et al.
Veröffentlicht: (2023)
von: Schwarz, Katja, et al.
Veröffentlicht: (2023)
Inconsistency-Aware Cross-Attention for Audio-Visual Fusion in Dimensional Emotion Recognition
von: Rajasekhar, G, et al.
Veröffentlicht: (2024)
von: Rajasekhar, G, et al.
Veröffentlicht: (2024)
Simplicial Approximation of Deforming 3D Spaces for Visualizing Fusion Plasma Simulation Data
von: Ren, Congrong, et al.
Veröffentlicht: (2023)
von: Ren, Congrong, et al.
Veröffentlicht: (2023)
Adaptive 3D Convolution for Remote Sensing Image Fusion
von: Peng, Siran, et al.
Veröffentlicht: (2026)
von: Peng, Siran, et al.
Veröffentlicht: (2026)
PCaM: A Progressive Focus Attention-Based Information Fusion Method for Improving Vision Transformer Domain Adaptation
von: Zang, Zelin, et al.
Veröffentlicht: (2025)
von: Zang, Zelin, et al.
Veröffentlicht: (2025)
Improving Audio-Visual Speech Recognition by Lip-Subword Correlation Based Visual Pre-training and Cross-Modal Fusion Encoder
von: Dai, Yusheng, et al.
Veröffentlicht: (2023)
von: Dai, Yusheng, et al.
Veröffentlicht: (2023)
Fusion to Enhance: Fusion Visual Encoder to Enhance Multimodal Language Model
von: She, Yifei, et al.
Veröffentlicht: (2025)
von: She, Yifei, et al.
Veröffentlicht: (2025)
GraphFusion3D: Dynamic Graph Attention Convolution with Adaptive Cross-Modal Transformer for 3D Object Detection
von: Mia, Md Sohag, et al.
Veröffentlicht: (2025)
von: Mia, Md Sohag, et al.
Veröffentlicht: (2025)
AFN: Adaptive Fusion Normalization via an Encoder-Decoder Framework
von: Zhou, Zikai, et al.
Veröffentlicht: (2023)
von: Zhou, Zikai, et al.
Veröffentlicht: (2023)
AdaptiveFusion: Adaptive Multi-Modal Multi-View Fusion for 3D Human Body Reconstruction
von: Chen, Anjun, et al.
Veröffentlicht: (2024)
von: Chen, Anjun, et al.
Veröffentlicht: (2024)
ProFuse: Efficient Cross-View Context Fusion for Open-Vocabulary 3D Gaussian Splatting
von: Chiou, Yen-Jen, et al.
Veröffentlicht: (2026)
von: Chiou, Yen-Jen, et al.
Veröffentlicht: (2026)
Intrinsic Image Fusion for Multi-View 3D Material Reconstruction
von: Kocsis, Peter, et al.
Veröffentlicht: (2025)
von: Kocsis, Peter, et al.
Veröffentlicht: (2025)
OccFusion: Depth Estimation Free Multi-sensor Fusion for 3D Occupancy Prediction
von: Zhang, Ji, et al.
Veröffentlicht: (2024)
von: Zhang, Ji, et al.
Veröffentlicht: (2024)
Uplifting Range-View-based 3D Semantic Segmentation in Real-Time with Multi-Sensor Fusion
von: Tan, Shiqi, et al.
Veröffentlicht: (2024)
von: Tan, Shiqi, et al.
Veröffentlicht: (2024)
COM3D: Leveraging Cross-View Correspondence and Cross-Modal Mining for 3D Retrieval
von: Wu, Hao, et al.
Veröffentlicht: (2024)
von: Wu, Hao, et al.
Veröffentlicht: (2024)
HyperPointFormer: Multimodal Fusion in 3D Space with Dual-Branch Cross-Attention Transformers
von: Rizaldy, Aldino, et al.
Veröffentlicht: (2025)
von: Rizaldy, Aldino, et al.
Veröffentlicht: (2025)
Consistent-1-to-3: Consistent Image to 3D View Synthesis via Geometry-aware Diffusion Models
von: Ye, Jianglong, et al.
Veröffentlicht: (2023)
von: Ye, Jianglong, et al.
Veröffentlicht: (2023)
Cross Attentional Audio-Visual Fusion for Dimensional Emotion Recognition
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2021)
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2021)
RefineFormer3D: Efficient 3D Medical Image Segmentation via Adaptive Multi-Scale Transformer with Cross Attention Fusion
von: Tyagi, Kavyansh, et al.
Veröffentlicht: (2026)
von: Tyagi, Kavyansh, et al.
Veröffentlicht: (2026)
MonoFusion: Sparse-View 4D Reconstruction via Monocular Fusion
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
A Novel Audio-Visual Information Fusion System for Mental Disorders Detection
von: Li, Yichun, et al.
Veröffentlicht: (2024)
von: Li, Yichun, et al.
Veröffentlicht: (2024)
3D Lymphoma Segmentation on PET/CT Images via Multi-Scale Information Fusion with Cross-Attention
von: Huang, Huan, et al.
Veröffentlicht: (2024)
von: Huang, Huan, et al.
Veröffentlicht: (2024)
Multimodal 3D Fusion and In-Situ Learning for Spatially Aware AI
von: Xu, Chengyuan, et al.
Veröffentlicht: (2024)
von: Xu, Chengyuan, et al.
Veröffentlicht: (2024)
MVLight: Relightable Text-to-3D Generation via Light-conditioned Multi-View Diffusion
von: Shim, Dongseok, et al.
Veröffentlicht: (2024)
von: Shim, Dongseok, et al.
Veröffentlicht: (2024)
PVAFN: Point-Voxel Attention Fusion Network with Multi-Pooling Enhancing for 3D Object Detection
von: Li, Yidi, et al.
Veröffentlicht: (2024)
von: Li, Yidi, et al.
Veröffentlicht: (2024)
BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing
von: Chen, Jiacheng, et al.
Veröffentlicht: (2025)
von: Chen, Jiacheng, et al.
Veröffentlicht: (2025)
Cross-Modal Registration Between 3D and 2D Fingerprints via Pose-Aware Unwrapping and Point-Cloud Fusion
von: Guan, Xiongjun, et al.
Veröffentlicht: (2026)
von: Guan, Xiongjun, et al.
Veröffentlicht: (2026)
Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Prediction
von: Chen, Dubing, et al.
Veröffentlicht: (2025)
von: Chen, Dubing, et al.
Veröffentlicht: (2025)
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition
von: Cong, Kaixuan, et al.
Veröffentlicht: (2025)
von: Cong, Kaixuan, et al.
Veröffentlicht: (2025)
Adaptive Transformer Attention and Multi-Scale Fusion for Spine 3D Segmentation
von: Xiang, Yanlin, et al.
Veröffentlicht: (2025)
von: Xiang, Yanlin, et al.
Veröffentlicht: (2025)
2D_3D Feature Fusion via Cross-Modal Latent Synthesis and Attention Guided Restoration for Industrial Anomaly Detection
von: Ali, Usman, et al.
Veröffentlicht: (2025)
von: Ali, Usman, et al.
Veröffentlicht: (2025)
Hierarchical Modeling for Medical Visual Question Answering with Cross-Attention Fusion
von: Zhang, Junkai, et al.
Veröffentlicht: (2025)
von: Zhang, Junkai, et al.
Veröffentlicht: (2025)
DepthFusion: Depth-Aware Hybrid Feature Fusion for LiDAR-Camera 3D Object Detection
von: Ji, Mingqian, et al.
Veröffentlicht: (2025)
von: Ji, Mingqian, et al.
Veröffentlicht: (2025)
3D Multi-frame Fusion for Video Stabilization
von: Peng, Zhan, et al.
Veröffentlicht: (2024)
von: Peng, Zhan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CVFusion: Cross-View Fusion of 4D Radar and Camera for 3D Object Detection
von: Zhong, Hanzhi, et al.
Veröffentlicht: (2025) -
CorrespondentDream: Enhancing 3D Fidelity of Text-to-3D using Cross-View Correspondences
von: Kim, Seungwook, et al.
Veröffentlicht: (2024) -
Wavelet-based Multi-View Fusion of 4D Radar Tensor and Camera for Robust 3D Object Detection
von: Guan, Runwei, et al.
Veröffentlicht: (2025) -
MV-SAM3D: Adaptive Multi-View Fusion for Layout-Aware 3D Generation
von: Li, Baicheng, et al.
Veröffentlicht: (2026) -
3D Crowd Counting via Geometric Attention-guided Multi-View Fusion
von: Zhang, Qi, et al.
Veröffentlicht: (2020)