Transformer-based stereo-aware 3D object detection from binocular images
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Hanqing, Pang, Yanwei, Cao, Jiale, Xie, Jin, Li, Xuelong |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VFMM3D: Releasing the Potential of Image by Vision Foundation Model for Monocular 3D Object Detection
by: Ding, Bonan, et al.
Published: (2024)
by: Ding, Bonan, et al.
Published: (2024)
Deep Intra-Image Contrastive Learning for Weakly Supervised One-Step Person Search
by: Wang, Jiabei, et al.
Published: (2023)
by: Wang, Jiabei, et al.
Published: (2023)
iSeg: An Iterative Refinement-based Framework for Training-free Segmentation
by: Sun, Lin, et al.
Published: (2024)
by: Sun, Lin, et al.
Published: (2024)
Implicit and Explicit Language Guidance for Diffusion-based Visual Perception
by: Wang, Hefeng, et al.
Published: (2024)
by: Wang, Hefeng, et al.
Published: (2024)
CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation
by: Sun, Lin, et al.
Published: (2024)
by: Sun, Lin, et al.
Published: (2024)
SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic Segmentation
by: Xie, Bin, et al.
Published: (2023)
by: Xie, Bin, et al.
Published: (2023)
SNNSIR: A Simple Spiking Neural Network for Stereo Image Restoration
by: Xu, Ronghua, et al.
Published: (2025)
by: Xu, Ronghua, et al.
Published: (2025)
CLIP-VIS: Adapting CLIP for Open-Vocabulary Video Instance Segmentation
by: Zhu, Wenqi, et al.
Published: (2024)
by: Zhu, Wenqi, et al.
Published: (2024)
Joint stereo 3D object detection and implicit surface reconstruction
by: Li, Shichao, et al.
Published: (2021)
by: Li, Shichao, et al.
Published: (2021)
Underlying Semantic Diffusion for Effective and Efficient In-Context Learning
by: Ji, Zhong, et al.
Published: (2025)
by: Ji, Zhong, et al.
Published: (2025)
Detail-aware multi-view stereo network for depth estimation
by: Tian, Haitao, et al.
Published: (2025)
by: Tian, Haitao, et al.
Published: (2025)
YASMOT: Yet another stereo image multi-object tracker
by: Malde, Ketil
Published: (2025)
by: Malde, Ketil
Published: (2025)
An evaluation of Deep Learning based stereo dense matching dataset shift from aerial images and a large scale stereo dataset
by: Wu, Teng, et al.
Published: (2024)
by: Wu, Teng, et al.
Published: (2024)
Hierarchical Matching and Reasoning for Multi-Query Image Retrieval
by: Ji, Zhong, et al.
Published: (2023)
by: Ji, Zhong, et al.
Published: (2023)
The Sampling-Gaussian for stereo matching
by: Pan, Baiyu, et al.
Published: (2024)
by: Pan, Baiyu, et al.
Published: (2024)
Optimal Transport Adapter Tuning for Bridging Modality Gaps in Few-Shot Remote Sensing Scene Classification
by: Ji, Zhong, et al.
Published: (2025)
by: Ji, Zhong, et al.
Published: (2025)
Multi-scale direction-aware SAR object detection network via global information fusion
by: Cao, Mingxiang, et al.
Published: (2023)
by: Cao, Mingxiang, et al.
Published: (2023)
Multi-scale interaction network for stereo image super-resolution
by: Xu, Liyi, et al.
Published: (2026)
by: Xu, Liyi, et al.
Published: (2026)
SSLFusion: Scale & Space Aligned Latent Fusion Model for Multimodal 3D Object Detection
by: Ding, Bonan, et al.
Published: (2025)
by: Ding, Bonan, et al.
Published: (2025)
Individuation of 3D perceptual units from neurogeometry of binocular cells
by: Bolelli, Maria Virginia, et al.
Published: (2024)
by: Bolelli, Maria Virginia, et al.
Published: (2024)
U(PM)$^2$:Unsupervised polygon matching with pre-trained models for challenging stereo images
by: Li, Chang, et al.
Published: (2025)
by: Li, Chang, et al.
Published: (2025)
DiffTF++: 3D-aware Diffusion Transformer for Large-Vocabulary 3D Generation
by: Cao, Ziang, et al.
Published: (2024)
by: Cao, Ziang, et al.
Published: (2024)
Language-to-Space Programming for Training-Free 3D Visual Grounding
by: Mi, Boyu, et al.
Published: (2025)
by: Mi, Boyu, et al.
Published: (2025)
3D StreetUnveiler with Semantic-aware 2DGS -- a simple baseline
by: Xu, Jingwei, et al.
Published: (2024)
by: Xu, Jingwei, et al.
Published: (2024)
MVSFormer++: Revealing the Devil in Transformer's Details for Multi-View Stereo
by: Cao, Chenjie, et al.
Published: (2024)
by: Cao, Chenjie, et al.
Published: (2024)
Sub-Image Recapture for Multi-View 3D Reconstruction
by: Wang, Yanwei
Published: (2025)
by: Wang, Yanwei
Published: (2025)
HAIR: Hypernetworks-based All-in-One Image Restoration
by: Cao, Jin, et al.
Published: (2024)
by: Cao, Jin, et al.
Published: (2024)
Transferable 3D Adversarial Shape Completion using Diffusion Models
by: Dai, Xuelong, et al.
Published: (2024)
by: Dai, Xuelong, et al.
Published: (2024)
Restereo: Diffusion stereo video generation and restoration
by: Huang, Xingchang, et al.
Published: (2025)
by: Huang, Xingchang, et al.
Published: (2025)
Illicit object detection in X-ray images using Vision Transformers
by: Cani, Jorgen, et al.
Published: (2024)
by: Cani, Jorgen, et al.
Published: (2024)
Analysis of different disparity estimation techniques on aerial stereo image datasets
by: Narayan, Ishan, et al.
Published: (2024)
by: Narayan, Ishan, et al.
Published: (2024)
MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression Segmentation
by: Wu, Changli, et al.
Published: (2026)
by: Wu, Changli, et al.
Published: (2026)
AppleGrowthVision: A large-scale stereo dataset for phenological analysis, fruit detection, and 3D reconstruction in apple orchards
by: von Hirschhausen, Laura-Sophia, et al.
Published: (2025)
by: von Hirschhausen, Laura-Sophia, et al.
Published: (2025)
MVInpainter: Learning Multi-View Consistent Inpainting to Bridge 2D and 3D Editing
by: Cao, Chenjie, et al.
Published: (2024)
by: Cao, Chenjie, et al.
Published: (2024)
MinD-3D: Reconstruct High-quality 3D objects in Human Brain
by: Gao, Jianxiong, et al.
Published: (2023)
by: Gao, Jianxiong, et al.
Published: (2023)
LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
by: Zhu, Chenming, et al.
Published: (2024)
by: Zhu, Chenming, et al.
Published: (2024)
Dehaze-then-Splat: Generative Dehazing with Physics-Informed 3D Gaussian Splatting for Smoke-Free Novel View Synthesis
by: Chen, Boss, et al.
Published: (2026)
by: Chen, Boss, et al.
Published: (2026)
Motion-guided small MAV detection in complex and non-planar scenes
by: Guo, Hanqing, et al.
Published: (2024)
by: Guo, Hanqing, et al.
Published: (2024)
3DPillars: Pillar-based two-stage 3D object detection
by: Noh, Jongyoun, et al.
Published: (2025)
by: Noh, Jongyoun, et al.
Published: (2025)
ARM3D: Attention-based relation module for indoor 3D object detection
by: Lan, Yuqing, et al.
Published: (2022)
by: Lan, Yuqing, et al.
Published: (2022)
Similar Items
-
VFMM3D: Releasing the Potential of Image by Vision Foundation Model for Monocular 3D Object Detection
by: Ding, Bonan, et al.
Published: (2024) -
Deep Intra-Image Contrastive Learning for Weakly Supervised One-Step Person Search
by: Wang, Jiabei, et al.
Published: (2023) -
iSeg: An Iterative Refinement-based Framework for Training-free Segmentation
by: Sun, Lin, et al.
Published: (2024) -
Implicit and Explicit Language Guidance for Diffusion-based Visual Perception
by: Wang, Hefeng, et al.
Published: (2024) -
CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation
by: Sun, Lin, et al.
Published: (2024)